Reviewer suitability classification on LR-Bench (test)
71.64AccuracyMERIT-Assessor-4B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MERIT-Assessor-4BPrompting strategy=Expertise-aware prompting2026.05 | 71.64 | 73.11 | 62.09 | 81.89 | 70.63 | |
| DeepSeek-V3.2Prompting strategy=Expertise-aware prompting2026.05 | 69.51 | 71 | 60.06 | 79.92 | 68.58 | |
| Qwen3.5-PlusPrompting strategy=Expertise-aware prompting2026.05 | 67.87 | 70.39 | 57.71 | 85.43 | 68.89 | |
| Qwen3-4BPrompting strategy=Expertise-aware prompting2026.05 | 67.08 | 68.81 | 57.61 | 78.45 | 66.15 | |
| DeepSeek-V3.2Prompting strategy=Direct prompting2026.05 | 63.11 | 66.93 | 53.4 | 89.76 | 66.96 | |
| Qwen3.5-PlusPrompting strategy=Direct prompting2026.05 | 59.18 | 64.41 | 50.52 | 95.67 | 66.12 |