Reasoning on BIG-Bench Hard (train)
91.9AccuracyR1-CI
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| R1-CIModel=Qwen2.5-14B-Instruct-1M, Evaluation Strategy=R1-CI2025.05 | 91.9 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-7B-Instruct-1M, Evaluation Strategy=R1-CI2025.05 | 89.2 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-14B-Instruct-1M, Evaluation Strategy=R1-CI wo IP2025.05 | 88.6 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-14B-Instruct-1M, Evaluation Strategy=R1-CI wo CL2025.05 | 87.9 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-14B-Instruct-1M, Evaluation Strategy=R1-CI wo GRPO2025.05 | 87.2 | — | — | — | — | — | — | — | — | — | |
| GPT-4oModel=GPT-4o, Evaluation Strategy=All Text2025.05 | 86.7 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-7B-Instruct-1M, Evaluation Strategy=R1-CI wo IP2025.05 | 86.5 | — | — | — | — | — | — | — | — | — | |
| Qwen3-14BModel=Qwen3-14B, Evaluation Strategy=CodeSteer2025.05 | 86 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-3B-Instruct, Evaluation Strategy=R1-CI2025.05 | 86 | — | — | — | — | — | — | — | — | — | |
| Qwen3-14BModel=Qwen3-14B, Evaluation Strategy=All Text2025.05 | 85.5 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-7B-Instruct-1M, Evaluation Strategy=R1-CI wo CL2025.05 | 85.2 | — | — | — | — | — | — | — | — | — | |
| GPT-4oModel=GPT-4o, Evaluation Strategy=Code Interpreter2025.05 | 85 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-7B-Instruct-1M, Evaluation Strategy=R1-CI wo GRPO2025.05 | 83.7 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-3B-Instruct, Evaluation Strategy=R1-CI wo IP2025.05 | 83.2 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-3B-Instruct, Evaluation Strategy=R1-CI wo CL2025.05 | 81.1 | — | — | — | — | — | — | — | — | — | |
| R1-CIModel=Qwen2.5-3B-Instruct, Evaluation Strategy=R1-CI wo GRPO2025.05 | 78.6 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-14B-Instruct-1MModel=Qwen2.5-14B-Instruct-1M, Evaluation Strategy=All Code2025.05 | 78.1 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-14B-Instruct-1MModel=Qwen2.5-14B-Instruct-1M, Evaluation Strategy=All Text2025.05 | 77.7 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-14B-Instruct-1MModel=Qwen2.5-14B-Instruct-1M, Evaluation Strategy=CodeSteer2025.05 | 75.6 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-7B-Instruct-1MModel=Qwen2.5-7B-Instruct-1M, Evaluation Strategy=All Code2025.05 | 70.2 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-7B-Instruct-1MModel=Qwen2.5-7B-Instruct-1M, Evaluation Strategy=CodeSteer2025.05 | 69.2 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-7B-Instruct-1MModel=Qwen2.5-7B-Instruct-1M, Evaluation Strategy=Code Agent (CI wo Fine-tune)2025.05 | 69 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-7B-Instruct-1MModel=Qwen2.5-7B-Instruct-1M, Evaluation Strategy=All Text2025.05 | 66 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-14B-Instruct-1MModel=Qwen2.5-14B-Instruct-1M, Evaluation Strategy=Code Agent (CI wo Fine-tune)2025.05 | 64.6 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-3B-InstructModel=Qwen2.5-3B-Instruct, Evaluation Strategy=CodeSteer2025.05 | 62.4 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-3B-InstructModel=Qwen2.5-3B-Instruct, Evaluation Strategy=All Code2025.05 | 60.4 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-3B-InstructModel=Qwen2.5-3B-Instruct, Evaluation Strategy=Code Agent (CI wo Fine-tune)2025.05 | 56.9 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-3B-InstructModel=Qwen2.5-3B-Instruct, Evaluation Strategy=All Text2025.05 | 55.1 | — | — | — | — | — | — | — | — | — | |
| Evolve 'Step-by-step'search budget=50 evaluations2023.11 | — | 63.1 | 57.1 | 64.6 | 60.6 | 84.9 | 18.25 | 86.2 | 54.5 | 61.2 | |
| Genetic Algorithmsearch budget=50 evaluations, pool size=42023.11 | — | 63.1 | 58.7 | 64.8 | 63.5 | 82.8 | 20.2 | 85.1 | 49.3 | 60.94 | |
| Greedysearch budget=50 evaluations, pool size=12023.11 | — | 63 | 59.5 | 66.7 | 68.5 | 83.5 | 19 | 85.1 | 53.9 | 62.4 | |
| Original Prompt2023.11 | — | 58.9 | 54.4 | 58 | 60 | 74.4 | 16 | 82 | 38.8 | 55.3 | |
| Our Methodsearch budget=50 evaluations, pool size=4, backbone=T5-encoder2023.11 | — | 67.7 | 60 | 70.1 | 70.4 | 88.4 | 22.8 | 85.4 | 59.8 | 65.6 |