Held-out Elo Estimation on LMArena (held-out models)
0.6Beta CoefficientSoft-Elo
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Soft-EloJudge=Qwen3.5-27B2026.06 | 0.6 | 46 | 13.6 | 70 | 0.003 | |
| Soft-EloJudge=GPT-OSS-120B2026.06 | 0.58 | 47.4 | 14.4 | 70 | 0.002 | |
| Soft-EloJudge=Gemma4-26B-A4B2026.06 | 0.54 | 55.9 | 15.7 | 72 | 0.011 | |
| Soft-EloJudge=Qwen3-32B2026.06 | 0.5 | 27.5 | 16.7 | 39 | 0.005 | |
| Soft-EloJudge=GPT-OSS-20B2026.06 | 0.46 | 34.5 | 19.8 | 42 | 0.009 | |
| Soft-EloJudge=Llama-3.3-70B2026.06 | 0.41 | 43.9 | 24.5 | 44 | 0.004 | |
| Soft-EloJudge=Gemma4-E4B2026.06 | 0.38 | 48.2 | 21 | 57 | 0.012 | |
| Soft-EloJudge=DeepSeek-V3.22026.06 | 0.36 | 63.4 | 17.1 | 73 | 0.003 |