Reading Comprehension Accuracy and Calibration on RACE
74.93AccuracyCORAL
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| CORALModel=Mistral-7B-Instruct-v0.32026.02 | 74.93 | 3.08 | 3.18 | 55.06 | 15.63 | |
| CORALModel=Deepseek-7B-Chat2026.02 | 73.06 | 2.62 | 3.08 | 57.69 | 16.06 | |
| CORALModel=Qwen2.5-7B-Instruct2026.02 | 71.04 | 3.46 | 3.35 | 54.4 | 16.88 | |
| ITIModel=Mistral-7B-Instruct-v0.32026.02 | 61.45 | 11.97 | 7.54 | 68.43 | 17.77 | |
| FewshotModel=Mistral-7B-Instruct-v0.32026.02 | 61.43 | 3.18 | 3.72 | 57.59 | 19.72 | |
| CCPSModel=Mistral-7B-Instruct-v0.32026.02 | 60.56 | 2.45 | 3.29 | 57.8 | 19.77 | |
| CCPSModel=Qwen2.5-7B-Instruct2026.02 | 58.38 | 5.11 | 6.44 | 55.22 | 15.55 | |
| FewshotModel=Qwen2.5-7B-Instruct2026.02 | 58.37 | 10.6 | 6.44 | 64.84 | 21.48 | |
| ITIModel=Qwen2.5-7B-Instruct2026.02 | 58.34 | 15.6 | 13.06 | 70.45 | 19.36 | |
| FewshotModel=Deepseek-7B-Chat2026.02 | 55.67 | 4.4 | 4.16 | 61.59 | 21.18 | |
| ITIModel=Deepseek-7B-Chat2026.02 | 55.63 | 16.56 | 8.44 | 73.18 | 19.09 | |
| CCPSModel=Deepseek-7B-Chat2026.02 | 54.9 | 3 | 4.25 | 60.36 | 20.84 | |
| SteerConfModel=Deepseek-7B-Chat2026.02 | 50.83 | 10.66 | 8.52 | 70.15 | 25.1 | |
| B3D-RWKV-7.2BModel Type=Strictly Causal Diffusion LM, Few-shot examples=02026.05 | 49.7 | — | — | — | — | |
| SteerConfModel=Qwen2.5-7B-Instruct2026.02 | 45.69 | 21.45 | 10.87 | 104.68 | 27.11 | |
| Dream-7BModel Type=Diffusion LM, Few-shot examples=02026.05 | 44.7 | — | — | — | — | |
| RWKV-7-7.2BModel Type=Causal LM, Few-shot examples=02026.05 | 43.5 | — | — | — | — | |
| LLaMA3-8BModel Type=Causal LM, Few-shot examples=02026.05 | 41.9 | — | — | — | — | |
| SteerConfModel=Mistral-7B-Instruct-v0.32026.02 | 40.83 | 4.67 | 4.78 | 122.27 | 26.52 | |
| PonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 39 | — | — | — | — | |
| LLaDA-8BModel Type=Diffusion LM, Few-shot examples=02026.05 | 38.7 | — | — | — | — | |
| AdaPonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 38.5 | — | — | — | — | |
| OPT-2.7BShots=5-shot, Training tokens=300B2026.03 | 37.5 | — | — | — | — | |
| PonderLM-1.4BShots=5-shot, Training tokens=300B2026.03 | 37.1 | — | — | — | — | |
| Pythia-6.9BShots=0-shot, Training tokens=300B2026.03 | 36.9 | — | — | — | — | |
| Pythia-6.9BShots=5-shot, Training tokens=300B2026.03 | 36.7 | — | — | — | — | |
| Ponder-2.8BShots=0-shot, Training tokens=300B2026.03 | 36.5 | — | — | — | — | |
| Tinyllama-1.1BShots=0-shot, Training tokens=3T2026.03 | 36.4 | — | — | — | — | |
| Tinyllama-1.1BShots=5-shot, Training tokens=3T2026.03 | 36.4 | — | — | — | — | |
| AdaPonderLM-2.8BShots=0-shot, Training tokens=305B2026.03 | 36.3 | — | — | — | — | |
| AdaPonderLM-1.4BShots=5-shot, Training tokens=310B2026.03 | 36.3 | — | — | — | — | |
| AR-1.7BModel Architecture=Autoregressive, Parameters=1.7B, Training Data=Nemotron-Pre-Training-Dataset, Training Tokens=2.1T2026.02 | 36.2 | — | — | — | — | |
| OPT-2.7BShots=0-shot, Training tokens=300B2026.03 | 36.2 | — | — | — | — | |
| Pythia-2.8BShots=5-shot, Training tokens=300B2026.03 | 35.9 | — | — | — | — | |
| GPTneo-2.7BShots=5-shot, Training tokens=300B2026.03 | 35.5 | — | — | — | — | |
| OPT-1.3BShots=5-shot, Training tokens=300B2026.03 | 35.4 | — | — | — | — | |
| Ponder-1.4BShots=0-shot, Training tokens=300B2026.03 | 35.2 | — | — | — | — | |
| Bloom-3BShots=0-shot, Training tokens=366B2026.03 | 35.2 | — | — | — | — | |
| GPTneo-2.7BShots=0-shot, Training tokens=300B2026.03 | 35.1 | — | — | — | — | |
| Duo-1.7BModel Architecture=Uniform-state Diffusion, Parameters=1.7B, Training Data=Nemotron-Pre-Training-Dataset, Training Tokens=2.1T2026.02 | 35 | — | — | — | — | |
| Pythia-2.8BShots=0-shot, Training tokens=300B2026.03 | 34.9 | — | — | — | — | |
| MDLM-1.7BModel Architecture=Masked Diffusion, Parameters=1.7B, Training Data=Nemotron-Pre-Training-Dataset, Training Tokens=2.1T2026.02 | 34.7 | — | — | — | — | |
| AdaPonderLM-1.4BShots=0-shot, Training tokens=312B2026.03 | 34.6 | — | — | — | — | |
| Pythia-1.4BShots=5-shot, Training tokens=300B2026.03 | 34.6 | — | — | — | — | |
| Bloom-3BShots=5-shot, Training tokens=366B2026.03 | 34.6 | — | — | — | — | |
| OPT-1.3BShots=0-shot, Training tokens=300B2026.03 | 34.3 | — | — | — | — | |
| Pythia-1.4BShots=0-shot, Training tokens=300B2026.03 | 34.1 | — | — | — | — | |
| Bloom-1.7BShots=5-shot, Training tokens=366B2026.03 | 33.5 | — | — | — | — | |
| Bloom-1.7BShots=0-shot, Training tokens=366B2026.03 | 33.2 | — | — | — | — | |
| Loop Transformer-410MShots=0-shot, Training tokens=25B2026.03 | 31.2 | — | — | — | — | |
| Pythia-410MShots=0-shot, Training tokens=300B2026.03 | 30.9 | — | — | — | — | |
| Loop Transformer-410MShots=5-shot, Training tokens=25B2026.03 | 30.5 | — | — | — | — | |
| Pythia-410MShots=5-shot, Training tokens=300B2026.03 | 30.4 | — | — | — | — | |
| Pause Token-410MShots=0-shot, Training tokens=25B2026.03 | 30.3 | — | — | — | — | |
| AdaPonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 30.3 | — | — | — | — | |
| AdaPonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 30.2 | — | — | — | — | |
| PonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 30.1 | — | — | — | — | |
| PonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 29.9 | — | — | — | — | |
| SMDM-1BParameters=1B, Training Data=slim-pajama2026.02 | 29.3 | — | — | — | — | |
| Pause Token-410MShots=5-shot, Training tokens=25B2026.03 | 29 | — | — | — | — | |
| BaseSparsity=Dense2026.02 | 28.5 | — | — | — | — | |
| SparseGPTSparsity=0.252026.02 | 28.35 | — | — | — | — | |
| Sink-AwareSparsity=0.25, Base Pruner=SparseGPT2026.02 | 28.2 | — | — | — | — | |
| WandaSparsity=0.252026.02 | 28.12 | — | — | — | — | |
| Sink-AwareSparsity=0.25, Base Pruner=Wanda2026.02 | 27.9 | — | — | — | — | |
| SparseGPTSparsity=0.502026.02 | 27.2 | — | — | — | — | |
| Sink-AwareSparsity=0.50, Base Pruner=SparseGPT2026.02 | 26.95 | — | — | — | — | |
| WandaSparsity=0.502026.02 | 26.8 | — | — | — | — | |
| Sink-AwareSparsity=0.50, Base Pruner=Wanda2026.02 | 26.55 | — | — | — | — | |
| Eso-LM-1.7BModel Architecture=Interpolating Diffusion, Parameters=1.7B, Training Data=Nemotron-Pre-Training-Dataset, Training Tokens=2.1T2026.02 | 26.1 | — | — | — | — | |
| SparseGPTSparsity=0.752026.02 | 24.95 | — | — | — | — | |
| Sink-AwareSparsity=0.75, Base Pruner=SparseGPT2026.02 | 24.7 | — | — | — | — | |
| Sink-AwareSparsity=0.75, Base Pruner=Wanda2026.02 | 24.4 | — | — | — | — | |
| Chance2026.02 | 24.2 | — | — | — | — | |
| WandaSparsity=0.752026.02 | 24.15 | — | — | — | — |