Question Answering on ARC-C (accuracy %)
72.44Accuracy (ARC-C)Gemma3
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemma3Parameters=12B2026.03 | 72.44 | |
| LoopRPTParameters=2.6B2026.03 | 66.89 | |
| OuroParameters=2.6B2026.03 | 66.12 | |
| Qwen3Parameters=8B2026.03 | 66.1 | |
| Qwen2.5Parameters=4B2026.03 | 63.65 | |
| Qwen2.5Parameters=7B2026.03 | 63.65 | |
| Qwen3Parameters=4B2026.03 | 60.75 | |
| Llama3.1Parameters=8B2026.03 | 60.75 | |
| Gemma3Parameters=3B2026.03 | 55.46 | |
| Llama3.2Parameters=3B2026.03 | 52.47 | |
| PonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 35.8 | |
| Pythia-6.9BShots=5-shot, Training tokens=300B2026.03 | 35.6 | |
| AdaPonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 35.2 | |
| Ponder-2.8BShots=0-shot, Training tokens=300B2026.03 | 32.5 | |
| PonderLM-1.4BShots=5-shot, Training tokens=300B2026.03 | 32.4 | |
| Bloom-3BShots=5-shot, Training tokens=366B2026.03 | 31.7 | |
| Pythia-6.9BShots=0-shot, Training tokens=300B2026.03 | 31.4 | |
| AdaPonderLM-2.8BShots=0-shot, Training tokens=305B2026.03 | 31.4 | |
| Tinyllama-1.1BShots=5-shot, Training tokens=3T2026.03 | 31.1 | |
| Pythia-2.8BShots=5-shot, Training tokens=300B2026.03 | 31 | |
| AdaPonderLM-1.4BShots=5-shot, Training tokens=310B2026.03 | 30.9 | |
| GPTneo-2.7BShots=5-shot, Training tokens=300B2026.03 | 30.1 | |
| OPT-2.7BShots=5-shot, Training tokens=300B2026.03 | 29.8 | |
| Pythia-2.8BShots=0-shot, Training tokens=300B2026.03 | 29.5 | |
| Pythia-1.4BShots=5-shot, Training tokens=300B2026.03 | 28.8 | |
| Tinyllama-1.1BShots=0-shot, Training tokens=3T2026.03 | 28 | |
| Bloom-3BShots=0-shot, Training tokens=366B2026.03 | 28 | |
| GPTneo-2.7BShots=0-shot, Training tokens=300B2026.03 | 27.5 | |
| AdaPonderLM-1.4BShots=0-shot, Training tokens=312B2026.03 | 27.1 | |
| Ponder-1.4BShots=0-shot, Training tokens=300B2026.03 | 27 | |
| OPT-1.3BShots=5-shot, Training tokens=300B2026.03 | 26.9 | |
| OPT-2.7BShots=0-shot, Training tokens=300B2026.03 | 26.8 | |
| Bloom-1.7BShots=5-shot, Training tokens=366B2026.03 | 26.2 | |
| Pythia-1.4BShots=0-shot, Training tokens=300B2026.03 | 25.9 | |
| Bloom-1.7BShots=0-shot, Training tokens=366B2026.03 | 23.7 | |
| OPT-1.3BShots=0-shot, Training tokens=300B2026.03 | 23.4 | |
| Pythia-410MShots=5-shot, Training tokens=300B2026.03 | 22.3 | |
| PonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 22.2 | |
| Pause Token-410MShots=5-shot, Training tokens=25B2026.03 | 21.8 | |
| Pause Token-410MShots=0-shot, Training tokens=25B2026.03 | 21.6 | |
| Pythia-410MShots=0-shot, Training tokens=300B2026.03 | 21.4 | |
| AdaPonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 21.4 | |
| Loop Transformer-410MShots=5-shot, Training tokens=25B2026.03 | 21.2 | |
| PonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 20.3 | |
| AdaPonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 20.1 | |
| Loop Transformer-410MShots=0-shot, Training tokens=25B2026.03 | 20 |