Science Question Answering on ARC Challenge (Score)
73.89ScoreGIFT
Evaluation Results
| Method | Links | |
|---|---|---|
| GIFTModel Size=32B2025.10 | 73.89 | |
| GRPOModel Size=32B2025.10 | 72.27 | |
| InstructModel Size=32B2025.10 | 71.93 | |
| GIFTModel Size=7B2025.10 | 65.7 | |
| InstructModel Size=7B2025.10 | 64.59 | |
| GRPOModel Size=7B2025.10 | 64.25 | |
| Qwen3.5-4BParameters=4B2026.05 | 54.9 | |
| OLMo-3-7BParameters=7B2026.05 | 53.6 | |
| Mellum 2Parameters=2.5B/12B2026.05 | 53.5 | |
| Qwen2.5-7BParameters=7B2026.05 | 51.3 | |
| Qwen3-4BParameters=4B2026.05 | 51.2 | |
| Nemotron-CC_ASI+Pre-training Curation Strategy=AI Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 49.32 | |
| Nemotron-CC_ASIPre-training Curation Strategy=AI Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 48.41 | |
| DCLMPre-training Curation Strategy=Human Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 45.02 | |
| Ultra-FinewebPre-training Curation Strategy=Human Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 43.77 | |
| Nemotron-CCPre-training Curation Strategy=AI Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 43.52 | |
| Fineweb-EduPre-training Curation Strategy=Human Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 43.45 |