Human Preference Ranking on Human Evaluation Elo (test)
1,634Elo Score(ai, ai+1)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| (ai, ai+1)training_pairs=1303, strategy=Stepwise preference pairs2025.12 | 1,634 | — | |
| (ai, ai+1)dstraining_pairs=277, strategy=Stepwise preference pairs (downsampled)2025.12 | 1,620 | — | |
| (a0, a*)training_pairs=277, strategy=Cumulative rewriting preference2025.12 | 1,612 | — | |
| (a0, a1)training_pairs=277, strategy=Single-step edit preference2025.12 | 1,525 | — | |
| (a, b)training_pairs=277, strategy=Standard A/B preference ranking2025.12 | 1,465 | — | |
| basetraining_pairs=02025.12 | 1,383 | — | |
| a* SFTtraining_pairs=277, strategy=Supervised Fine-Tuning2025.12 | 1,377 | — | |
| GPT-4o-05132024.09 | 1,079 | 1 | |
| Molmo-72BNumber of crops=122024.09 | 1,077 | 2 | |
| Gemini 1.5 Pro2024.09 | 1,074 | 3 | |
| Claude-3.5 Sonnet2024.09 | 1,069 | 4 | |
| Llama-3.2V-90B-Instruct2024.09 | 1,063 | 5 | |
| Molmo-7B-DNumber of crops=122024.09 | 1,056 | 6 | |
| Gemini 1.5 Flash2024.09 | 1,054 | 7 | |
| LLaVA OneVision-72B2024.09 | 1,051 | 8 | |
| Molmo-7B-ONumber of crops=122024.09 | 1,051 | 9 | |
| GPT-4V2024.09 | 1,041 | 10 | |
| Llama-3.2V-11B-Instruct2024.09 | 1,040 | 11 | |
| Qwen2-VL-72B2024.09 | 1,037 | 12 | |
| MolmoE-1BNumber of crops=122024.09 | 1,032 | 13 | |
| Qwen2-VL-7B2024.09 | 1,025 | 14 | |
| LLaVA OneVision-7B2024.09 | 1,024 | 15 | |
| InternVL2-Llama-3-76B2024.09 | 1,018 | 16 | |
| Pixtral-12B2024.09 | 1,016 | 17 | |
| Claude-3 Haiku2024.09 | 999 | 18 | |
| Phi3.5-Vision-4B2024.09 | 982 | 19 | |
| xGen-MM-interleave-4B2024.09 | 979 | 20 | |
| Claude-3 Opus2024.09 | 971 | 21 | |
| LLaVA-1.5-13B2024.09 | 960 | 22 | |
| InternVL2-8B2024.09 | 953 | 23 | |
| Cambrian-1-34B2024.09 | 953 | 24 | |
| Cambrian-1-8B2024.09 | 952 | 25 | |
| LLaVA-1.5-7B2024.09 | 951 | 26 | |
| PaliGemma-mix-3B2024.09 | 937 | 27 |