Open Question Answering on LegalMC4 (test)
77.2LLM Factual CorrectnessGPT-5 (min. reasoning)
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5 (min. reasoning)Train Data=Base Model2026.01 | 77.2 | |
| GPT-5-mini (min. reasoning)Train Data=Base Model2026.01 | 70.1 | |
| LLaMA 3.1 (8B)Train Data=Difficulty-Graded QA Data2026.01 | 55.4 | |
| Gemma 3 (12B)Train Data=Difficulty-Graded QA Data2026.01 | 54.5 | |
| LLaMA 3.1 (8B)Train Data=Base Model2026.01 | 43 | |
| Gemma 3 (12B)Train Data=Standard Instruction QA Data2026.01 | 41.8 | |
| Gemma 3 (12B)Train Data=Base Model2026.01 | 38.7 | |
| LLaMA 3.1 (8B)Train Data=Standard Instruction QA Data2026.01 | 35.4 |