General Reasoning on BBH (test)
86.5AccuracySelf-Consistency
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Self-ConsistencyCategory=Advanced Single-Agent2026.05 | 86.5 | — | — | — | — | |
| AgentPSOCategory=Multi-Agent2026.05 | 86.5 | — | — | — | — | |
| Chain-of-ThoughtCategory=Vanilla Single-Agent2026.05 | 85.5 | — | — | — | — | |
| DMADCategory=Multi-Agent2026.05 | 85 | — | — | — | — | |
| MADCategory=Multi-Agent2026.05 | 84 | — | — | — | — | |
| ReflectionCategory=Advanced Single-Agent2026.05 | 83 | — | — | — | — | |
| Step-Back PromptingCategory=Vanilla Single-Agent2026.05 | 82.5 | — | — | — | — | |
| Self-RefineCategory=Advanced Single-Agent2026.05 | 82.5 | — | — | — | — | |
| ICL+FTModel Size=27B, Number of training examples=1502025.12 | 81.8 | — | — | — | — | |
| ICL+FTModel Size=9B, Number of training examples=1502025.12 | 78.7 | — | — | — | — | |
| ToTCategory=Advanced Single-Agent2026.05 | 78.5 | — | — | — | — | |
| MoACategory=Multi-Agent2026.05 | 77.5 | — | — | — | — | |
| Gemini 1.5 Pro (ICL-Only)Number of training examples=1502025.12 | 76.6 | — | — | — | — | |
| Gemini 1.5 Pro (ICL-Only)Number of training examples=302025.12 | 72.8 | — | — | — | — | |
| FT-OnlyModel Size=27B, Number of training examples=1502025.12 | 72.6 | — | — | — | — | |
| ICL+FTModel Size=27B, Number of training examples=302025.12 | 72.3 | — | — | — | — | |
| FT-OnlyModel Size=9B, Number of training examples=1502025.12 | 69.7 | — | — | — | — | |
| ICL+FTModel Size=9B, Number of training examples=302025.12 | 68.7 | — | — | — | — | |
| ICL+FTModel Size=2B, Number of training examples=1502025.12 | 67.5 | — | — | — | — | |
| ICL-OnlyModel Size=27B, Number of training examples=30, Selection=Best performing 1, 3, 5, 10, 15, 30, or 100-shot2025.12 | 64.6 | — | — | — | — | |
| ICL-OnlyModel Size=27B, Number of training examples=150, Selection=Best performing 1, 3, 5, 10, 15, 30, or 100-shot2025.12 | 64.2 | — | — | — | — | |
| FT-OnlyModel Size=27B, Number of training examples=302025.12 | 62.6 | — | — | — | — | |
| ICL-OnlyModel Size=9B, Number of training examples=30, Selection=Best performing 1, 3, 5, 10, 15, 30, or 100-shot2025.12 | 57.1 | — | — | — | — | |
| ICL-OnlyModel Size=9B, Number of training examples=150, Selection=Best performing 1, 3, 5, 10, 15, 30, or 100-shot2025.12 | 56.6 | — | — | — | — | |
| ICL+FTModel Size=2B, Number of training examples=302025.12 | 55.3 | — | — | — | — | |
| FT-OnlyModel Size=9B, Number of training examples=302025.12 | 52.8 | — | — | — | — | |
| FT-OnlyModel Size=2B, Number of training examples=1502025.12 | 50.7 | — | — | — | — | |
| ICL-OnlyModel Size=2B, Number of training examples=150, Selection=Best performing 1, 3, 5, 10, 15, 30, or 100-shot2025.12 | 37.6 | — | — | — | — | |
| ICL-OnlyModel Size=2B, Number of training examples=30, Selection=Best performing 1, 3, 5, 10, 15, 30, or 100-shot2025.12 | 37.2 | — | — | — | — | |
| FT-OnlyModel Size=2B, Number of training examples=302025.12 | 27.3 | — | — | — | — | |
| 0-shotBackbone=GPT-32023.06 | — | 60.8 | 64.1 | 56.4 | 45.9 | |
| 0-shotBackbone=GPT-42023.06 | — | 64 | 74 | 79.2 | 68.5 | |
| 16-shotBackbone=GPT-42023.06 | — | 93.3 | 75.5 | 80.9 | 66.4 | |
| 5-shotBackbone=GPT-32023.06 | — | 55.6 | 56.5 | 62.1 | 36.7 | |
| 5-shotBackbone=GPT-42023.06 | — | 88.4 | 75.7 | 79.3 | 62.8 | |
| APE-15Backbone=GPT-32023.06 | — | 68.5 | 67.3 | 32.1 | 45.5 | |
| APE-400Backbone=GPT-32023.06 | — | 65.5 | 56.9 | 23.5 | 45.6 | |
| DLN-1Backbone=GPT-32023.06 | — | 91.9 | 68.5 | 55.7 | 47.5 | |
| DLN-1Backbone=GPT-42023.06 | — | 95.2 | 77.1 | 76.7 | 69.1 | |
| KATEBackbone=GPT-32023.06 | — | 71.1 | 56.9 | 61.1 | 44.4 |