Question Answering on IRB1K
84.9English CorrectnessGPT-4.1
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4.1Evaluation Protocol=RAG performance2026.02 | 84.9 | 73.4 | 80.1 | 82.4 | 79.1 | 80.6 | 86.4 | 84.5 | 82.5 | 66.9 | 2.8 | |
| GPT-5 mediumEvaluation Protocol=RAG performance2026.02 | 84.9 | 77.6 | 81.4 | 83.5 | 80.9 | 82.6 | 85 | 85.5 | 83.5 | 71 | 50.5 | |
| gpt-oss-120bEvaluation Protocol=RAG performance2026.02 | 84.4 | 69.8 | 79.9 | 79.7 | 78.3 | 78 | 86.2 | 86.4 | 80.4 | 73.4 | 14.5 | |
| GPT-5 mini mediumEvaluation Protocol=RAG performance2026.02 | 83.1 | 72.8 | 77.6 | 81.8 | 77.4 | 80.1 | 84.8 | 82.3 | 81 | 66.1 | 35 | |
| GPT-4.1 miniEvaluation Protocol=RAG performance2026.02 | 82.8 | 72.2 | 76.8 | 81.7 | 77.9 | 79.1 | 85 | 81.8 | 80.6 | 66.1 | 1.5 | |
| DeepSeek-R1Evaluation Protocol=RAG performance2026.02 | 79.9 | 71.4 | 76.2 | 78.1 | 75.1 | 77.6 | 80.6 | 80.5 | 78.5 | 62.1 | 22.8 | |
| Llama-3.3-70BEvaluation Protocol=RAG performance2026.02 | 78.6 | 62.9 | 73.6 | 73.6 | 72.7 | 71.3 | 80.4 | 80.9 | 75.2 | 54.8 | 5 | |
| Llama-4-ScoutEvaluation Protocol=RAG performance2026.02 | 76 | 62.9 | 71.9 | 71.9 | 70.2 | 70.3 | 75.9 | 80 | 73.5 | 52.4 | 20.5 | |
| GPT-5 mediumEvaluation Protocol=Closed-book performance2026.02 | 41.9 | 31.5 | 48.4 | 30.4 | 29.8 | 36.8 | 54.9 | 58.2 | 39.4 | 29 | 24 | |
| GPT-4.1Evaluation Protocol=Closed-book performance2026.02 | 30.6 | 21.4 | 32 | 24.1 | 18.4 | 26.1 | 40.4 | 43.2 | 28.3 | 20.2 | 1 | |
| gpt-oss-120bEvaluation Protocol=Closed-book performance2026.02 | 16.3 | 11.9 | 16.1 | 13.9 | 5.7 | 14.3 | 28 | 28.2 | 15.4 | 8.9 | 11.5 | |
| GPT-5 mini mediumEvaluation Protocol=Closed-book performance2026.02 | 14.8 | 9.1 | 16.1 | 10.4 | 5.8 | 10.8 | 26.2 | 28.2 | 12.9 | 14.5 | 8 | |
| Llama-3.3-70BEvaluation Protocol=Closed-book performance2026.02 | 14.8 | 9.5 | 16.7 | 10.1 | 4.9 | 9.5 | 29 | 34.5 | 13.1 | 12.9 | 12.2 | |
| GPT-4.1 miniEvaluation Protocol=Closed-book performance2026.02 | 11.4 | 9.7 | 11.5 | 10.4 | 5.9 | 10.4 | 20.3 | 22.7 | 11.1 | 8.1 | 10 | |
| DeepSeek-R1Evaluation Protocol=Closed-book performance2026.02 | 11.3 | 8.5 | 13.4 | 7.9 | 7.5 | 9.1 | 14.7 | 17.7 | 10.9 | 4.8 | 39.5 | |
| Llama-4-ScoutEvaluation Protocol=Closed-book performance2026.02 | 8.5 | 4.8 | 10 | 5.1 | 3.2 | 5.1 | 14 | 17.3 | 7.7 | 3.2 | 31 |