General Reasoning on Aggregated Reasoning Tasks 10-Task Combination
52.2Average ScoreQwen2.5-7B + E3-TIR (Ours)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-7B + E3-TIR (Ours)Backbone=Qwen2.5-7B, Training Protocol=E3-TIR2026.04 | 52.2 | 7.5 | |
| Qwen2.5-7B + SFT-then-RLBackbone=Qwen2.5-7B, Training Protocol=SFT-then-RL2026.04 | 49.9 | 7.7 | |
| Qwen2.5-7B + Zero-RLBackbone=Qwen2.5-7B, Training Protocol=Zero-RL2026.04 | 49 | 7.9 | |
| Llama3.1-8B + E3-TIR (Ours)Backbone=Llama3.1-8B, Training Protocol=E3-TIR2026.04 | 47.8 | 7.4 | |
| Qwen2.5-3B + E3-TIR (Ours)Backbone=Qwen2.5-3B, Training Protocol=E3-TIR2026.04 | 46.7 | 6.2 | |
| Llama3.1-8B + SFT-then-RLBackbone=Llama3.1-8B, Training Protocol=SFT-then-RL2026.04 | 45.7 | 7.1 | |
| Llama3.1-8B + Zero-RLBackbone=Llama3.1-8B, Training Protocol=Zero-RL2026.04 | 44.9 | 7.7 | |
| Qwen2.5-3B + SFT-then-RLBackbone=Qwen2.5-3B, Training Protocol=SFT-then-RL2026.04 | 44.2 | 6.1 | |
| Qwen2.5-3B + Zero-RLBackbone=Qwen2.5-3B, Training Protocol=Zero-RL2026.04 | 43.2 | 7.9 | |
| Qwen2.5-7B + Only SFTBackbone=Qwen2.5-7B, Training Protocol=Only SFT2026.04 | 39.2 | 8.2 | |
| Llama3.1-8B + Only SFTBackbone=Llama3.1-8B, Training Protocol=Only SFT2026.04 | 34.8 | 8.4 | |
| Qwen2.5-3B + Only SFTBackbone=Qwen2.5-3B, Training Protocol=Only SFT2026.04 | 32.9 | 7.4 | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B, Training Protocol=Instruct2026.04 | 30.7 | — | |
| Qwen2.5-3B-InstructBackbone=Qwen2.5-3B, Training Protocol=Instruct2026.04 | 26.8 | — | |
| Llama3.1-8B-InstructBackbone=Llama3.1-8B, Training Protocol=Instruct2026.04 | 23.9 | — |