Multi-hop Reasoning on HotpotQA (Accuracy, Majority Vote)
42.3AccuracyGRPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPOContext=Continued-training, Sampling Level=Trajectory2026.06 | 42.3 | 43.5 | |
| Orch-RMContext=Continued-training, Sampling Level=Orchestration2026.06 | 41.63 | 42.5 | |
| Orch-RMContext=Training-from-scratch, Sampling Level=Orchestration2026.06 | 41.06 | 42.5 | |
| DPOContext=Continued-training2026.06 | 40.88 | 41 | |
| MAS-OrchestraContext=Training-from-scratch, Sampling Level=Trajectory2026.06 | 40.63 | 42.5 | |
| LLM-as-a-judgeContext=Continued-training, Sampling Level=Orchestration2026.06 | 40.5 | 41.5 | |
| RFTContext=Continued-training, Sampling Level=Trajectory2026.06 | 40.19 | 40 | |
| Qwen2.5-7B-InstructContext=Base orchestrator2026.06 | 30.43 | 34 |