Financial Reasoning on S&P 500 Scenario-based MCQs Stage I
87.14AccuracyDeepSeek-v3.1
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-v3.1Provider=DeepSeek-AI2026.04 | 87.14 | |
| Claude-4.5-SonnetProvider=Anthropic2026.04 | 86.19 | |
| Ours (Full Model)Backbone=Llama-3.1-8B-Instruct, Fine-tuning=Q-LoRA, CORA=true, DARA=true, Verification=true2026.04 | 82.38 | |
| GPT-5.1Provider=OpenAI2026.04 | 80.95 | |
| Ours w/o DARABackbone=Llama-3.1-8B-Instruct, Fine-tuning=Q-LoRA, CORA=true2026.04 | 76.67 | |
| Qwen3-8BParameters=8B, Provider=Alibaba2026.04 | 71.9 | |
| Ours w/o CORA, DARA & VerificationBackbone=Llama-3.1-8B-Instruct, Fine-tuning=Q-LoRA2026.04 | 70.95 | |
| Mistral-7B-Instruct-v0.3Parameters=7B2026.04 | 64.29 | |
| Llama-3.1-8B-InstructParameters=8B, Provider=Meta2026.04 | 57.62 | |
| Ours w/o CORABackbone=Llama-3.1-8B-Instruct, Fine-tuning=Q-LoRA, DARA=true2026.04 | 46.19 |