General Reasoning on General Reasoning Tasks Group 2
35TheoremQA ScoreLLaMA3.1-8B Instruct + CFT*
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLaMA3.1-8B Instruct + CFT*Base Model=LLaMA3.1-8B Instruct, Fine-tuning Strategy=CFT*, Teacher Model=GPT-4o2025.05 | 35 | 30.3 | 40.8 | 36.4 | |
| S1.1-3B + Distilled SFTBase Model=S1.1-3B, Fine-tuning Strategy=Distilled SFT2025.05 | 34.9 | 29.3 | 36.4 | 33.5 | |
| LLaMA3.1-8B Instruct + CGDBase Model=LLaMA3.1-8B Instruct, Fine-tuning Strategy=Critique-Guided Distillation (CGD), Teacher Model=LLaMA3.3-70B Instruct2025.05 | 34 | 35.9 | 40.3 | 36.7 | |
| S1.1-3B + CGDBase Model=S1.1-3B, Fine-tuning Strategy=Critique-Guided Distillation (CGD), Teacher Model=LLaMA3.3-70B Instruct2025.05 | 32.8 | 31.8 | 35.7 | 33.4 | |
| LLaMA3.1-8B Instruct + Distilled SFTBase Model=LLaMA3.1-8B Instruct, Fine-tuning Strategy=Distilled SFT2025.05 | 28.9 | 31.8 | 35.1 | 31.9 | |
| LLaMA3.1-8B Instruct + CFTBase Model=LLaMA3.1-8B Instruct, Fine-tuning Strategy=CFT, Teacher Model=LLaMA3.3-70B Instruct2025.05 | 28.2 | 34.3 | 34.2 | 32.4 | |
| LLaMA3.1-8B InstructBase Model=LLaMA3.1-8B Instruct2025.05 | 27.6 | 30.8 | 31.2 | 29.9 | |
| S1.1-3B + CFTBase Model=S1.1-3B, Fine-tuning Strategy=CFT, Teacher Model=LLaMA3.3-70B Instruct2025.05 | 25.9 | 26.7 | 35.9 | 29.5 | |
| S1.1-3B + SFTBase Model=S1.1-3B, Fine-tuning Strategy=SFT2025.05 | 22.8 | 29.8 | 36.9 | 29.8 | |
| LLaMA3.1-8B Instruct + SFTBase Model=LLaMA3.1-8B Instruct, Fine-tuning Strategy=SFT2025.05 | 22.1 | 33.3 | 39.3 | 31.6 | |
| S1.1-3BBase Model=S1.1-3B2025.05 | 21.6 | 16.7 | 13.7 | 17.9 |