Resume quality judgment on Resume screening dataset ground truth GPT-5.1
86.82AccuracyQwen3-8B (AutoScreen-FW)
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-8B (AutoScreen-FW)Few-shot=true, Sampling Strategy=Diversity-based, Shots=5, Sample Type=High-quality only, Attribute Type=Overall Judgement and Dimension Scores2026.03 | 86.82 | |
| Qwen3-8BFew-shot=false, Sampling Strategy=N/A, Shots=0, Sample Type=N/A, Attribute Type=N/A2026.03 | 85.52 | |
| Llama-3.1-8B (AutoScreen-FW)Few-shot=true, Sampling Strategy=Clustering-based, Shots=15, Sample Type=High- and Low- quality, Attribute Type=Overall Judgement and Dimension Scores2026.03 | 84.85 | |
| GPT-5-miniFew-shot=false, Sampling Strategy=N/A, Shots=0, Sample Type=N/A, Attribute Type=N/A2026.03 | 84.45 | |
| GPT-5-nanoFew-shot=false, Sampling Strategy=N/A, Shots=0, Sample Type=N/A, Attribute Type=N/A2026.03 | 83.74 | |
| Llama-3.1-8BFew-shot=false, Sampling Strategy=N/A, Shots=0, Sample Type=N/A, Attribute Type=N/A2026.03 | 79.27 |