LLM-as-a-Judge Evaluation using 10-fold Cross-Validation
88CG AccuracyQwen3-14B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-14BTarget=Concepts C2026.06 | 88 | 80 | 82.9 | |
| Gemini-2-FlashTarget=Concepts C2026.06 | 83 | 72 | 80.8 | |
| GPT-OSS 20BTarget=Prediction y^2026.06 | 83 | 78 | 53 | |
| Gemini-2-FlashTarget=Prediction y^2026.06 | 79 | 74 | 57.5 | |
| GPT-OSS 20BTarget=Concepts C2026.06 | 77 | 67 | 86.2 | |
| Qwen3-14BTarget=Prediction y^2026.06 | 64 | 62 | 64.4 | |
| Random BaselineTarget=Prediction y^2026.06 | 50 | 50 | 12 | |
| Random BaselineTarget=Concepts C2026.06 | 25 | 25 | 27 |