Rubric satisfaction evaluation on ML
36.7Claude-4 Sonnet ScoreGPT-5-Thinking
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5-ThinkingMode=Thinking2025.12 | 36.7 | 57.2 | 45.5 | |
| GPT-OSS-120B2025.12 | 30.9 | 50.6 | 30.6 | |
| ML-Ft-Qwen-3-30BFine-tuned domain=ML2025.12 | 30.2 | 40.9 | 23.2 | |
| Claude-4.1-Opus2025.12 | 29.4 | 44 | 27.3 | |
| Arxiv-Ft-Qwen-3-30BFine-tuned domain=ArXiv2025.12 | 27.9 | 40 | 23.8 | |
| GPT-OSS-20B2025.12 | 27.5 | 42.7 | 23.9 | |
| Qwen-3-30B-A3B-ThinkingMode=Thinking2025.12 | 27.3 | 34.8 | 17.6 | |
| ML-Ft-Qwen-3-4BFine-tuned domain=ML2025.12 | 26.8 | 38.4 | 19.6 | |
| Grok-42025.12 | 26.8 | 39.9 | 24.8 | |
| Med-Ft-Qwen-3-30BFine-tuned domain=Medical2025.12 | 25.6 | 40.5 | 23.4 | |
| Claude-4-Sonnet2025.12 | 24.2 | 39 | 21.6 | |
| Gemini-2.5-Pro2025.12 | 24.1 | 44.7 | 24.6 | |
| Grok-4-FastSpeed=Fast2025.12 | 23.3 | 42.6 | 24.4 | |
| Qwen-3-30B-A3B2025.12 | 23.2 | 35.6 | 18.1 | |
| ML-Ft-Llama3.1-8BFine-tuned domain=ML2025.12 | 17.9 | 29.3 | 8.3 | |
| Qwen-3-4B-ThinkingMode=Thinking2025.12 | 16.9 | 24.7 | 10.3 | |
| Gemma-3-27B2025.12 | 16.6 | 35.2 | 14.9 | |
| Qwen-3-4B2025.12 | 13.9 | 26.5 | 10.9 | |
| ML-Ft-Gemma3-4BFine-tuned domain=ML2025.12 | 13.7 | 28.4 | 7.9 | |
| Gemma-3-4B2025.12 | 7.6 | 22.2 | 6.2 | |
| Llama-4-Maverick2025.12 | 5.7 | 18.6 | 6.6 | |
| Llama-3.3-70B2025.12 | 4.7 | 16.3 | 5.4 | |
| Llama-4-Scout2025.12 | 3.4 | 13.7 | 4.8 | |
| Llama-3.1-8B2025.12 | 2.3 | 9.8 | 2.5 |