Open Question Answering on BGB (test)
76.4Factual Correctness (%)Gemma 3 (12B)
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemma 3 (12B)Train Data=Difficulty-Graded QA Data2026.01 | 76.4 | |
| GPT-5 (min. reasoning)Train Data=Base Model2026.01 | 74.7 | |
| GPT-5-mini (min. reasoning)Train Data=Base Model2026.01 | 62.7 | |
| LLaMA 3.1 (8B)Train Data=Difficulty-Graded QA Data2026.01 | 59.2 | |
| Gemma 3 (12B)Train Data=Standard Instruction QA Data2026.01 | 46.6 | |
| LLaMA 3.1 (8B)Train Data=Base Model2026.01 | 39.2 | |
| LLaMA 3.1 (8B)Train Data=Standard Instruction QA Data2026.01 | 37.6 | |
| Gemma 3 (12B)Train Data=Base Model2026.01 | 36.8 |