Accuracy on GPQA
71.43AccuracyGRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPOInference Setting=Scaffold-free2026.06 | 71.43 | |
| ParaBridgeInference Setting=Scaffold-free2026.06 | 71.43 | |
| BaselineInference Setting=Scaffold-free2026.06 | 71.34 | |
| RFTInference Setting=Scaffold-free2026.06 | 68.45 | |
| Claude 3.5 SonnetEvaluation Protocol=0-shot CoT2026.03 | 59.4 | |
| DyLANBase Model=gpt-4o, Number of Agents=82026.04 | 58.89 | |
| HuggingGPTImplementation Framework=original2026.03 | 56.67 | |
| GoAMaxBase Model=gpt-4o, Number of Agents=62026.04 | 56.57 | |
| Layer-wise CapreseModel Size=14B, Sparse Algorithm=GRIFFIN, 0-shot=true2025.05 | 55.05 | |
| GoAMaxBase Model=gpt-4o, Number of Agents=32026.04 | 55.05 | |
| RefineBase Model=gpt-4o, Number of Agents=62026.04 | 54.98 | |
| SCBase Model=gpt-4o, Number of Agents=62026.04 | 54.27 | |
| GPT-4oEvaluation Protocol=0-shot CoT2026.03 | 53.6 | |
| DebateBase Model=gpt-4o, Number of Agents=62026.04 | 53.03 | |
| ReconcileBase Model=gpt-4o, Number of Agents=62026.04 | 53.03 | |
| DS-Qwen 14BModel Size=14B, 0-shot=true2025.05 | 52.53 | |
| Llama 3 405BEvaluation Protocol=0-shot CoT2026.03 | 51.1 | |
| E2E CapreseModel Size=14B, Sparse Algorithm=CATS, 0-shot=true2025.05 | 51.01 | |
| Layer-wise CapreseModel Size=14B, Sparse Algorithm=CATS, 0-shot=true2025.05 | 50.51 | |
| MoABase Model=gpt-4o, Number of Agents=62026.04 | 50.51 | |
| AlphaRL-enhanced DAPOAlgorithm=DAPO, Stage=40%2025.10 | 49.25 | |
| DAPOAlgorithm=DAPO, Stage=Fully Trained2025.10 | 48.23 | |
| Qwen3-4B + TTSVModel=Qwen3-4B, Adaptation=TTSV2025.12 | 47.98 | |
| gpt-4oBase Model=gpt-4o, Number of Agents=12026.04 | 47.47 | |
| HuggingGPTImplementation Framework=MASFactory2026.03 | 47.32 | |
| GRPOAlgorithm=GRPO, Stage=Fully Trained2025.10 | 47.1 | |
| Llama 3 70BEvaluation Protocol=0-shot CoT2026.03 | 46.7 | |
| Qwen3-4B-Instruct-2507Model=Qwen3-4B-Instruct-2507, Training Method=Instruct2026.03 | 46.6 | |
| RLOOAlgorithm=RLOO, Stage=Fully Trained2025.10 | 45.82 | |
| HeRLModel=Qwen3-4B-Instruct-2507, Training Method=HeRL2026.03 | 45.3 | |
| CATSModel Size=14B, 0-shot=true2025.05 | 44.95 | |
| AlphaRL-enhanced RLOOAlgorithm=RLOO, Stage=10%2025.10 | 44.95 | |
| AlphaRL-enhanced RLOOAlgorithm=RLOO, Stage=40%2025.10 | 44.4 | |
| Relational AggregationCalls=3, Tokens(k)=0.663, Time(s)=0.202026.05 | 43.64 | |
| Relational AggregationCategory=Ours2026.05 | 43.64 | |
| E2E CapreseModel Size=14B, Sparse Algorithm=GRIFFIN, 0-shot=true2025.05 | 43.43 | |
| AlphaRL-enhanced GRPOAlgorithm=GRPO, Stage=40%2025.10 | 43.13 | |
| Anchor (Refined Belief)Category=Ours2026.05 | 42.14 | |
| RLOOAlgorithm=RLOO, Stage=40%2025.10 | 42.05 | |
| Qwen3-4BModel=Qwen3-4B2025.12 | 41.92 | |
| GRIFFINModel Size=14B, 0-shot=true2025.05 | 41.92 | |
| DAPOAlgorithm=DAPO, Stage=40%2025.10 | 41.67 | |
| AlphaRL-enhanced DAPOAlgorithm=DAPO, Stage=10%2025.10 | 41.54 | |
| GRPOAlgorithm=GRPO, Stage=40%2025.10 | 41.16 | |
| Anchor (Initial Belief)Category=Ours2026.05 | 40.61 | |
| GoAMeanCategory=Multi-Agent2026.05 | 40.54 | |
| GoA-MaxCalls=11, Tokens(k)=17.32, Time(s)=88.152026.05 | 39.98 | |
| GoAMaxCategory=Multi-Agent2026.05 | 39.98 | |
| Vibe Graphing-Task SpecificImplementation Framework=Vibe Graphing2026.03 | 39.51 | |
| Majority (Refined Belief)Category=Ours2026.05 | 39.41 | |
| CharacterFlywheel V7Evaluation Protocol=0-shot CoT2026.03 | 39.3 | |
| RefineCategory=Multi-Agent2026.05 | 38.92 | |
| RLOOAlgorithm=RLOO, Stage=10%2025.10 | 38.65 | |
| AgentVerseImplementation Framework=original2026.03 | 38.39 | |
| DS-Qwen 7BModel Size=7B, 0-shot=true2025.05 | 38.38 | |
| Layer-wise CapreseModel Size=7B, Sparse Algorithm=CATS, 0-shot=true2025.05 | 38.38 | |
| AgentVerseImplementation Framework=MASFactory2026.03 | 37.5 | |
| Majority (Initial Belief)Category=Ours2026.05 | 37.26 | |
| GRPOAlgorithm=GRPO, Stage=10%2025.10 | 36.74 | |
| AlphaRL-enhanced GRPOAlgorithm=GRPO, Stage=10%2025.10 | 36.74 | |
| DAPOAlgorithm=DAPO, Stage=10%2025.10 | 36.66 | |
| SCCategory=Multi-Agent2026.05 | 36.36 | |
| E2E CapreseModel Size=7B, Sparse Algorithm=CATS, 0-shot=true2025.05 | 35.86 | |
| ReConcileCategory=Multi-Agent2026.05 | 34.34 | |
| Qwen2.5-7B-InstructModel=Qwen2.5-7B-Instruct, Training Method=Instruct2026.03 | 34.1 | |
| HeRLModel=Qwen2.5-7B-Instruct, Training Method=HeRL2026.03 | 33.9 | |
| CodeCategory=Single-Agent2026.05 | 33.84 | |
| Self-MoACategory=Multi-Agent2026.05 | 33.84 | |
| CATSModel Size=7B, 0-shot=true2025.05 | 32.83 | |
| MoACalls=19, Tokens(k)=56.87, Time(s)=245.652026.05 | 32.83 | |
| GeneralCategory=Single-Agent2026.05 | 32.83 | |
| MoACategory=Multi-Agent2026.05 | 32.83 | |
| Llama 3 8BEvaluation Protocol=0-shot CoT2026.03 | 32.8 | |
| CAMELImplementation Framework=original2026.03 | 32.59 | |
| MathCategory=Single-Agent2026.05 | 30.81 | |
| LegalCategory=Single-Agent2026.05 | 30.3 | |
| DebateCategory=Multi-Agent2026.05 | 29.29 | |
| HeRLModel=Llama-3.2-3B-Instruct, Training Method=HeRL2026.03 | 29.2 | |
| FinanceCategory=Single-Agent2026.05 | 28.28 | |
| Layer-wise CapreseModel Size=7B, Sparse Algorithm=GRIFFIN, 0-shot=true2025.05 | 27.78 | |
| Llama-3.2-3B-InstructModel=Llama-3.2-3B-Instruct, Training Method=Instruct2026.03 | 26.5 | |
| BiomedicalCategory=Single-Agent2026.05 | 25.25 | |
| CAMELImplementation Framework=MASFactory2026.03 | 24.78 | |
| E2E CapreseModel Size=7B, Sparse Algorithm=GRIFFIN, 0-shot=true2025.05 | 23.74 | |
| E2E CapreseModel Size=1.5B, Sparse Algorithm=CATS, 0-shot=true2025.05 | 22.22 | |
| GRIFFINModel Size=7B, 0-shot=true2025.05 | 21.72 | |
| DS-Qwen 1.5BModel Size=1.5B, 0-shot=true2025.05 | 18.69 | |
| E2E CapreseModel Size=1.5B, Sparse Algorithm=GRIFFIN, 0-shot=true2025.05 | 16.67 | |
| Layer-wise CapreseModel Size=1.5B, Sparse Algorithm=CATS, 0-shot=true2025.05 | 16.67 | |
| Layer-wise CapreseModel Size=1.5B, Sparse Algorithm=GRIFFIN, 0-shot=true2025.05 | 13.13 | |
| GRIFFINModel Size=1.5B, 0-shot=true2025.05 | 11.62 | |
| CATSModel Size=1.5B, 0-shot=true2025.05 | 11.62 |