Question Answering on LongHealth
90.8AccuracyQwen3.5-397B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3.5-397BBase Model=Qwen3.5-397B2026.05 | 90.8 | — | |
| Kimi-K2.5Base Model=Kimi-K2.52026.05 | 90.4 | — | |
| Qwen3.5-122BBase Model=Qwen3.5-122B2026.05 | 90.2 | — | |
| Qwen3.5-35BBase Model=Qwen3.5-35B2026.05 | 89.8 | — | |
| GLM-5.1Base Model=GLM-5.12026.05 | 89.7 | — | |
| Qwen3.5-9B + MaRBase Model=Qwen3.5-9B, Optimization Method=MaR2026.05 | 89.3 | — | |
| Deepseek-V3.2Base Model=Deepseek-V3.22026.05 | 89.2 | — | |
| Qwen3.5-4B + MaRBase Model=Qwen3.5-4B, Optimization Method=MaR2026.05 | 88.5 | — | |
| Qwen3.5-9B + DAPOBase Model=Qwen3.5-9B, Optimization Method=DAPO2026.05 | 88 | — | |
| GPT-OSS-120BBase Model=GPT-OSS-120B2026.05 | 87.5 | — | |
| Qwen3.5-4B + DAPOBase Model=Qwen3.5-4B, Optimization Method=DAPO2026.05 | 87.4 | — | |
| Oracle ContextComp.=1×2026.06 | 87.4 | — | |
| Qwen3.5-9BBase Model=Qwen3.5-9B2026.05 | 87.1 | — | |
| Qwen3.5-4BBase Model=Qwen3.5-4B2026.05 | 86.8 | — | |
| Synthetic Mixed Training + Focal RewritingModel=Qwen, RAG=Enabled2026.03 | 82.3 | 8.6 | |
| CAS Training (Single-Cartridge / Doc.)Comp.=2×2026.06 | 81.1 | — | |
| Synthetic Mixed Training + Focal RewritingModel=Llama, RAG=Enabled2026.03 | 80.3 | 10.3 | |
| CAS Training (Single-Cartridge / Doc.)Comp.=10×2026.06 | 80.1 | — | |
| CAS Training (Single-Cartridge / Doc.)Comp.=5×2026.06 | 79.9 | — | |
| CAS Training (Single-Cartridge / Doc.)Comp.=20×2026.06 | 78.9 | — | |
| CAS Training (Single-Cartridge / Doc.)Comp.=100×2026.06 | 77.3 | — | |
| CAS Training (Single-Cartridge / Doc.)Comp.=50×2026.06 | 76.9 | — | |
| Synthetic Mixed Training + Focal RewritingModel=Qwen, RAG=Disabled2026.03 | 76.7 | — | |
| Vanilla RAGModel=Qwen, RAG=Enabled2026.03 | 73.7 | — | |
| Synthetic Mixed Training + Focal RewritingModel=Llama, RAG=Disabled2026.03 | 71.2 | — | |
| Vanilla RAGModel=Llama, RAG=Enabled2026.03 | 70 | — | |
| No ContextComp.=–2026.06 | 37.5 | — |