Rubric Satisfaction Evaluation on Medical
50.9Claude-4 Sonnet ScoreGPT-5-Thinking
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5-ThinkingMode=Thinking2025.12 | 50.9 | 69.7 | 56.5 | |
| GPT-OSS-120B2025.12 | 43.9 | 63.2 | 39.1 | |
| Claude-4.1-Opus2025.12 | 43.9 | 62.2 | 38.4 | |
| Med-Ft-Qwen-3-30BFine-tuned domain=Medical2025.12 | 43.5 | 54.4 | 37 | |
| Claude-4-Sonnet2025.12 | 41.5 | 58.8 | 35.2 | |
| ML-Ft-Qwen-3-30BFine-tuned domain=ML2025.12 | 41.3 | 50.6 | 31.6 | |
| GPT-OSS-20B2025.12 | 41 | 56.9 | 33 | |
| Arxiv-Ft-Qwen-3-30BFine-tuned domain=ArXiv2025.12 | 40 | 51 | 31.7 | |
| Qwen-3-30B-A3B-ThinkingMode=Thinking2025.12 | 39.4 | 49.2 | 27.6 | |
| Qwen-3-30B-A3B2025.12 | 38.8 | 50.6 | 29.3 | |
| Grok-4-FastSpeed=Fast2025.12 | 38.8 | 55.9 | 34.8 | |
| Grok-42025.12 | 38.5 | 53.3 | 34 | |
| Gemma-3-27B2025.12 | 32.5 | 48.2 | 24.4 | |
| Gemini-2.5-Pro2025.12 | 32.3 | 53.3 | 31.6 | |
| Qwen-3-4B2025.12 | 31 | 42.4 | 19.9 | |
| Qwen-3-4B-ThinkingMode=Thinking2025.12 | 30.7 | 39.9 | 17.4 | |
| Llama-4-Maverick2025.12 | 19.8 | 35.7 | 16.4 | |
| Gemma-3-4B2025.12 | 18.9 | 33.9 | 12.8 | |
| Llama-3.3-70B2025.12 | 18.3 | 32.7 | 14.4 | |
| Llama-4-Scout2025.12 | 12.1 | 20.5 | 9.2 | |
| Llama-3.1-8B2025.12 | 11 | 23.4 | 8.5 |