Reasoning on OBQA (leave-one-out setup)
87.7Average AccuracyMashup Learning
Evaluation Results
| Method | Links | |
|---|---|---|
| Mashup LearningModel=Gemma-3 4B, Setup=LoRA2026.03 | 87.7 | |
| From scratchModel=Gemma-3 4B, Setup=LoRA2026.03 | 85.6 | |
| Mashup LearningModel=Gemma-2 2B, Setup=LoRA2026.03 | 85.5 | |
| From scratchModel=Gemma-2 2B, Setup=LoRA2026.03 | 85 | |
| Mashup LearningModel=Gemma-3 4B, Setup=Full FT2026.03 | 84.6 | |
| From scratchModel=Gemma-3 4B, Setup=Full FT2026.03 | 83.9 | |
| Mashup LearningModel=Gemma-2 2B, Setup=Full FT2026.03 | 81.2 | |
| From scratchModel=Gemma-2 2B, Setup=Full FT2026.03 | 80.8 | |
| Mashup LearningModel=Gemma-3 1B, Setup=LoRA2026.03 | 75.6 | |
| Mashup LearningModel=Gemma-3 1B, Setup=Full FT2026.03 | 73.1 | |
| From scratchModel=Gemma-3 1B, Setup=LoRA2026.03 | 70.3 | |
| From scratchModel=Gemma-3 1B, Setup=Full FT2026.03 | 68.3 |