Complex Multi-step Reasoning on Big-Bench Hard
85.7Hard AccuracyDeepCompress-Zero-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepCompress-Zero-7BParameters=7B2025.10 | 85.7 | |
| DeepMath-Zero-7BParameters=7B2025.10 | 85 | |
| Open-Reasoner-Zero-7BParameters=7B2025.10 | 83.2 | |
| Qwen-2.5-7B-SimpleRL-ZooParameters=7B, Variant=SimpleRL-Zoo2025.10 | 74.9 | |
| DeepCompress-Zero-3BParameters=3B2025.10 | 73.7 | |
| DeepMath-Zero-3BParameters=3B2025.10 | 71.6 | |
| Qwen-2.5-3B-InstructParameters=3B, Variant=Instruct2025.10 | 71.3 | |
| Qwen-2.5-3BParameters=3B2025.10 | 48.1 | |
| Qwen-2.5-7BParameters=7B2025.10 | 41.5 | |
| NITPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 28.67 | |
| NTPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 28.07 | |
| NITPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 26.14 | |
| NTPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 21.92 | |
| NITPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 19.11 | |
| NTPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 17.7 |