Abstract screening on SESR-Eval Study 41 (test)
79.7AccuracyGEPA
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| GEPAStudent model=Qwen-2.5-7B-instruct, Reflection model=GPT-5-mini, max_eval=122026.05 | 79.7 | 79.9 | 90.5 | 84.8 | 0.562 | 0.548 | 70.6 | 85.5 | |
| Structured baselineStudent model=Qwen-2.5-7B-instruct2026.05 | 78.8 | 79.1 | 89.6 | 84 | 0.539 | 0.53 | 70.4 | 84.7 | |
| GEPAStudent model=Qwen-2.5-7B-instruct, Reflection model=GPT-5-mini, max_eval=62026.05 | 78.6 | 79.2 | 90.7 | 84.2 | 0.54 | 0.516 | 71.9 | 84.9 | |
| GEPAStudent model=Qwen-2.5-7B-instruct, Reflection model=GPT-5-mini, max_eval=242026.05 | 78 | 81.1 | 85.5 | 82.9 | 0.529 | 0.519 | 66.1 | 83.2 | |
| GEPAStudent model=Qwen-2.5-7B-instruct, Reflection model=GPT-5-mini, max_eval=22026.05 | 77.6 | 76.2 | 93.8 | 83.9 | 0.518 | 0.482 | 76.8 | 85 |