Instruction Following Evaluation on PPE-IFEval
76ScoreRubric-ARROW-voting@5
Evaluation Results
| Method | Links | |
|---|---|---|
| Rubric-ARROW-voting@5Voting Strategy=voting@52026.05 | 76 | |
| RIFL-voting@5Voting Strategy=voting@52026.05 | 75.8 | |
| Gemini-2.5-FlashModel Category=Black-box LLMs2026.05 | 75 | |
| RIFLModel Category=Rubric-based Methods2026.05 | 73.3 | |
| Rubric-ARROWModel Category=Rubric-based Methods2026.05 | 72.6 | |
| Rubric-ARROW w/o RL-voting@5Voting Strategy=voting@5, Training Protocol=w/o RL2026.05 | 72.2 | |
| Rubric-RM-voting@5Voting Strategy=voting@52026.05 | 70.8 | |
| Rubric-ARROW w/o RLTraining Protocol=w/o RL2026.05 | 69.5 | |
| Rubric-RMModel Category=Rubric-based Methods2026.05 | 67 | |
| RM-R1-32B (DeepSeek-Dist)Backbone=DeepSeek-Dist2026.05 | 63.2 | |
| RM-R1-14B (DeepSeek-Dist)Backbone=DeepSeek-Dist2026.05 | 61.2 | |
| API Prompting (Rubric+Judge pairwise)Evaluation Protocol=Rubric+Judge pairwise2026.05 | 61 | |
| RM-R1-32B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst2026.05 | 60.4 | |
| RRM-32BModel Category=Larger White-box LLMs2026.05 | 60.2 | |
| API Prompting (direct Judge)Evaluation Protocol=direct Judge2026.05 | 59.2 | |
| RM-R1-14B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst2026.05 | 59 | |
| Claude-3.5-SonnetModel Category=Black-box LLMs2026.05 | 58 | |
| RM-R1-7B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst2026.05 | 55.2 | |
| Qwen-3-8B (Rubric+Judge pairwise)Backbone=Qwen-3-8B, Evaluation Protocol=Rubric+Judge pairwise2026.05 | 53.8 | |
| Qwen-3-8B (Rubric+Judge pointwise)Backbone=Qwen-3-8B, Evaluation Protocol=Rubric+Judge pointwise2026.05 | 52.6 | |
| RM-R1-7B (DeepSeek-Dist)Backbone=DeepSeek-Dist2026.05 | 51 | |
| RRM-7BModel Category=White-box Judge/Reward LLMs2026.05 | 51 | |
| JudgeLRM-7BModel Category=White-box Judge/Reward LLMs2026.05 | 46 | |
| API Prompting (Rubric+Judge pointwise)Evaluation Protocol=Rubric+Judge pointwise2026.05 | 26.4 |