General Reasoning and Code Generation on Combined Benchmarks Avg
33.75Average ScoreDLLG
Evaluation Results
| Method | Links | |
|---|---|---|
| DLLGModel Scale=0.5B, Evaluation=0-shot2026.06 | 33.75 | |
| UniTeModel Scale=0.5B, Evaluation=0-shot2026.06 | 32.9 | |
| GaCModel Scale=0.5B, Evaluation=0-shot2026.06 | 32.77 | |
| Pack of LLMsModel Scale=0.5B, Evaluation=0-shot2026.06 | 31.7 | |
| Qwen2.5-0.5B-InstructModel Scale=0.5B, Evaluation=0-shot2026.06 | 31.41 | |
| Entropy WeightingModel Scale=0.5B, Evaluation=0-shot2026.06 | 30.9 | |
| Token Maj-VotingModel Scale=0.5B, Evaluation=0-shot2026.06 | 30.67 | |
| SLERPModel Scale=0.5B, Evaluation=0-shot2026.06 | 30.58 | |
| EmbedLLMModel Scale=0.5B, Evaluation=0-shot2026.06 | 30.58 | |
| LinearModel Scale=0.5B, Evaluation=0-shot2026.06 | 30.55 | |
| RouterDCModel Scale=0.5B, Evaluation=0-shot2026.06 | 29.94 | |
| Dolphin3.0-Qwen2.5-0.5BModel Scale=0.5B, Evaluation=0-shot2026.06 | 28.13 | |
| Qwen2.5-Coder-0.5B-InstructModel Scale=0.5B, Evaluation=0-shot2026.06 | 27.04 | |
| Task ArithmeticModel Scale=0.5B, Evaluation=0-shot2026.06 | 22.64 |