Science Question Answering on ARC-C (Accuracy and Gap)
92.5AccuracyEvo 8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Evo 8BPost-training=SFT, Number of shots=02026.02 | 92.5 | — | |
| LLaDA 8BPost-training=SFT, Number of shots=02026.02 | 88.5 | — | |
| Qwen2.5-Omni-7BSystem Type=End-to-end system2025.10 | 87.1 | 1.3 | |
| BD3-LM 7BPost-training=SFT+RL, Number of shots=02026.02 | 86.7 | — | |
| ASR + Qwen2.5-7BSystem Type=Cascaded topline, ASR=Whisper-v3-Large2025.10 | 86.5 | 1.9 | |
| SALAD-7BTraining Stage=Stage II2025.10 | 84 | 4.4 | |
| AceMADBase Model=Qwen3-235B-A22B-Instruct, Iteration Rounds (T)=2, Number of Agents (N)=52026.03 | 82.43 | — | |
| LLaMA3 8BPost-training=SFT+RL, Number of shots=02026.02 | 82.4 | — | |
| SALAD-7BTraining Stage=Stage I2025.10 | 82.3 | 6.1 | |
| AceMADBase Model=Qwen3-235B-A22B-Instruct, Iteration Rounds (T)=3, Number of Agents (N)=52026.03 | 80.56 | — | |
| SALAD-3BTraining Stage=Stage II2025.10 | 79.9 | 1.9 | |
| AceMADBase Model=Qwen3-235B-A22B-Instruct, Iteration Rounds (T)=5, Number of Agents (N)=52026.03 | 77.78 | — | |
| ASR + Qwen2.5-3BSystem Type=Cascaded topline, ASR=Whisper-v3-Large2025.10 | 77.7 | 4.2 | |
| SALAD-3BTraining Stage=Stage I2025.10 | 75.6 | 6.2 | |
| Majority VotingBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 69.44 | — | |
| Single AgentBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 65.74 | — | |
| GLM-4-Voice-9BSystem Type=End-to-end system2025.10 | 64.6 | 28.7 | |
| AceMADBase Model=GPT-4o-mini, Iteration Rounds (T)=5, Number of Agents (N)=52026.03 | 59.26 | — | |
| ARD 7BPost-training=SFT+RL, Number of shots=02026.02 | 56.9 | — | |
| AceMADBase Model=GPT-4o-mini, Iteration Rounds (T)=2, Number of Agents (N)=52026.03 | 56.48 | — | |
| AceMADBase Model=GPT-4o-mini, Iteration Rounds (T)=3, Number of Agents (N)=52026.03 | 56.48 | — | |
| Decentralized MADBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 49.07 | — | |
| DiVA-Llama3.1-8BSystem Type=End-to-end system2025.10 | 45.9 | 36 | |
| Sparse MADBase Model=GPT-4o-mini, Number of Agents (N)=52026.03 | 43.52 | — | |
| Qwen2-Audio-7BSystem Type=End-to-end system2025.10 | 43.5 | 28.5 | |
| Decentralized MADBase Model=GPT-4o-mini, Number of Agents (N)=52026.03 | 41.67 | — | |
| Centralized MADBase Model=GPT-4o-mini, Number of Agents (N)=52026.03 | 41.67 | — | |
| Random2025.10 | 25 | — | |
| Majority VotingBase Model=GPT-4o-mini, Number of Agents (N)=52026.03 | 23.15 | — | |
| Single AgentBase Model=GPT-4o-mini, Number of Agents (N)=52026.03 | 20.37 | — | |
| Centralized MADBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 13.89 | — | |
| Sparse MADBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 13.89 | — |