Mathematical Reasoning on BeyondAIME (accuracy)
82.5AccuracyTRICE-30B
Evaluation Results
| Method | Links | |
|---|---|---|
| TRICE-30BTool Usage=true, Parameter Scale=~30B, Evaluation Protocol=unified2026.05 | 82.5 | |
| TRICE-30BTool Use=Yes2026.05 | 82.5 | |
| GLM-4.7-Flash w/ recipeTool Usage=true, Parameter Scale=~30B, Evaluation Protocol=unified2026.05 | 81 | |
| Nemotron-3-Nano-30B-A3BTool Usage=true, Parameter Scale=~30B, Evaluation Protocol=unified2026.05 | 80 | |
| DeepSeek-V3.2-ThinkingTool Use=No2026.05 | 76.8 | |
| GLM-4.7-FlashTool Usage=true, Parameter Scale=~30B, Evaluation Protocol=unified2026.05 | 76 | |
| Qwen3.5-35B-A3BTool Usage=false, Parameter Scale=~30B, Evaluation Protocol=unified2026.05 | 72.5 | |
| Qwen3-235B-A22B-ThinkingTool Use=No2026.05 | 71.8 | |
| TRICE-4BTool Usage=true, Parameter Scale=<10B, Evaluation Protocol=unified2026.05 | 71.3 | |
| TRICE-30BTool Usage=false, Parameter Scale=~30B, Evaluation Protocol=unified2026.05 | 71 | |
| Qwen3.5-9BTool Usage=false, Parameter Scale=<10B, Evaluation Protocol=unified2026.05 | 67.3 | |
| Qwen3-30B-A3B-Thinking-2507Tool Usage=false, Parameter Scale=~30B, Evaluation Protocol=unified2026.05 | 65.9 | |
| GPT-OSS-20BTool Usage=true, Parameter Scale=~30B, Evaluation Protocol=unified2026.05 | 63 | |
| ASTER-4B†Tool Usage=true, Parameter Scale=<10B, Evaluation Protocol=original2026.05 | 61.7 | |
| Qwen3.5-4BTool Usage=false, Parameter Scale=<10B, Evaluation Protocol=unified2026.05 | 58.8 | |
| TRICE-4BTool Usage=false, Parameter Scale=<10B, Evaluation Protocol=unified2026.05 | 58.5 | |
| Qwen3-4B-Thinking-2507Tool Usage=false, Parameter Scale=<10B, Evaluation Protocol=unified2026.05 | 54.3 | |
| Qwen3-30B-A3B-Instruct-2507Tool Usage=false, Parameter Scale=~30B, Evaluation Protocol=unified2026.05 | 51.3 |