Question Answering on SciDQA
64.3Accuracy (Table)DeepSeek-R1
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| DeepSeek-R1RAG Framework=NeuSym-RAG2025.05 | 64.3 | 64.6 | 63.9 | 64.5 | — | |
| DeepSeek-R1RAG Framework=Classic-RAG2025.05 | 63.9 | 61.3 | 61.7 | 62.4 | — | |
| DeepSeek-R1RAG Strategy=Classic RAG, Training Strategy=Prompting2026.01 | 63.9 | 61.3 | 61.7 | 62.4 | — | |
| GPT-4o-miniRAG Framework=NeuSym-RAG2025.05 | 63 | 63.6 | 62.5 | 63 | — | |
| GPT-4VRAG Framework=NeuSym-RAG2025.05 | 62.6 | 63.5 | 63.2 | 63.1 | — | |
| Qwen2.5-VL-72B-InstructRAG Framework=NeuSym-RAG2025.05 | 60.2 | 60.6 | 61.8 | 60.5 | — | |
| GPT-4o-miniRAG Framework=Classic-RAG2025.05 | 59.4 | 60.4 | 59.3 | 59.8 | — | |
| GPT-4o-miniRAG Strategy=Classic RAG, Training Strategy=Prompting2026.01 | 59.4 | 60.4 | 59.3 | 59.8 | — | |
| Llama-3.3-70B-InstructRAG Framework=Classic-RAG2025.05 | 56.8 | 58.8 | 58.9 | 58 | — | |
| GPT-4VRAG Framework=Classic-RAG2025.05 | 56.6 | 56.8 | 58.1 | 57.4 | — | |
| Llama-3.3-70B-InstructRAG Framework=NeuSym-RAG2025.05 | 55.5 | 57.3 | 56.6 | 56.4 | — | |
| Qwen2.5-VL-72B-InstructRAG Framework=Classic-RAG2025.05 | 54.8 | 56.9 | 56.3 | 56.2 | — | |
| Qwen2.5-VL-72B-InstructRAG Strategy=Classic RAG, Training Strategy=Prompting2026.01 | 54.8 | 56.9 | 56.3 | 56.2 | — | |
| Qwen2.5-7B-Instruct (DTFT) + DFPO (PaperGuide)RAG Strategy=NeuSym RAG, Backbone=Qwen2.5-7B-Instruct, Fine-tuning=DTFT, RL Method=DFPO2026.01 | 49.5 | 48.8 | 45.5 | 48.3 | 37.4 | |
| Qwen2.5-7B-Instruct (SFT) + DAPORAG Strategy=NeuSym RAG, Backbone=Qwen2.5-7B-Instruct, Fine-tuning=SFT, RL Method=DAPO2026.01 | 48.1 | 47.8 | 45.1 | 47.1 | 34.5 | |
| Qwen2.5-7B-Instruct (DTFT)RAG Strategy=NeuSym RAG, Backbone=Qwen2.5-7B-Instruct, Fine-tuning=DTFT2026.01 | 46.9 | 48.4 | 43.3 | 47 | 34.9 | |
| Qwen2.5-7B-Instruct (SFT) + M-GRPORAG Strategy=NeuSym RAG, Backbone=Qwen2.5-7B-Instruct, Fine-tuning=SFT, RL Method=M-GRPO2026.01 | 46.8 | 47.7 | 41.5 | 46.1 | 32 | |
| Qwen2.5-7B-Instruct (SFT)RAG Strategy=NeuSym RAG, Backbone=Qwen2.5-7B-Instruct, Fine-tuning=SFT2026.01 | 46.5 | 45.9 | 42.7 | 45.6 | 33.2 | |
| Qwen2.5-7B-InstructRAG Strategy=Classic RAG, Training Strategy=Prompting2026.01 | 45.1 | 44.9 | 45.4 | 45.1 | — | |
| Qwen2.5-3B-Instruct (SFT)RAG Strategy=NeuSym RAG, Backbone=Qwen2.5-3B-Instruct, Fine-tuning=SFT2026.01 | 44.8 | 44.5 | 38.8 | 43.3 | 28.4 | |
| Qwen2.5-3B-Instruct (SFT) + DAPORAG Strategy=NeuSym RAG, Backbone=Qwen2.5-3B-Instruct, Fine-tuning=SFT, RL Method=DAPO2026.01 | 44.6 | 44.7 | 41.7 | 43.9 | 29.3 | |
| Qwen2.5-3B-Instruct (DTFT) + DFPO (PaperGuide)RAG Strategy=NeuSym RAG, Backbone=Qwen2.5-3B-Instruct, Fine-tuning=DTFT, RL Method=DFPO2026.01 | 43.6 | 44.2 | 41.9 | 43.5 | 36.1 | |
| Qwen2.5-3B-Instruct (DTFT)RAG Strategy=NeuSym RAG, Backbone=Qwen2.5-3B-Instruct, Fine-tuning=DTFT2026.01 | 41.7 | 41.5 | 38.3 | 40.9 | 30.3 | |
| Qwen2.5-3B-Instruct (SFT) + M-GRPORAG Strategy=NeuSym RAG, Backbone=Qwen2.5-3B-Instruct, Fine-tuning=SFT, RL Method=M-GRPO2026.01 | 39.5 | 38.6 | 39.2 | 39.1 | 26 | |
| Qwen2.5-3B-InstructRAG Strategy=Classic RAG, Training Strategy=Prompting2026.01 | 35.6 | 37 | 33.9 | 35.2 | — |