Coding on LiveCodeBench (Accuracy, Output Token Count)
90.7AccuracyGemini-3.0
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini-3.0variant=Pro, protocol=Pass@1-COT2025.12 | 90.7 | 13,000 | — | |
| DeepSeek-V3.2variant=Speciale, protocol=Pass@1-COT2025.12 | 88.7 | 27,000 | — | |
| GPT-5variant=High, protocol=Pass@1-COT2025.12 | 84.5 | 13,000 | — | |
| DeepSeek-V3.2variant=Thinking, protocol=Pass@1-COT2025.12 | 83.3 | 16,000 | — | |
| Kimi-K2variant=Thinking, protocol=Pass@1-COT2025.12 | 82.6 | 29,000 | — | |
| RAMModel Category=Merged Models2026.01 | 31.96 | — | 47.72 | |
| TAModel Category=Merged Models2026.01 | 31.95 | — | 46.69 | |
| DARE+TAModel Category=Merged Models2026.01 | 31.95 | — | 46.69 | |
| RAM+Model Category=Merged Models2026.01 | 31.6 | — | 46.84 | |
| FisherModel Category=Merged Models2026.01 | 30.87 | — | 45.89 | |
| TIESModel Category=Merged Models2026.01 | 30.63 | — | 46.32 | |
| CURE (Coding)Model Category=Base and Task Models2026.01 | 30.23 | — | 45.76 | |
| DARE+TIESModel Category=Merged Models2026.01 | 29.26 | — | 39.53 | |
| MemAgent (Memory)Model Category=Base and Task Models2026.01 | 28.92 | — | 44.8 | |
| ToolRL (Tool)Model Category=Base and Task Models2026.01 | 26.76 | — | 42.05 | |
| BaseModel Category=Base and Task Models2026.01 | 23.43 | — | 36.42 |