Extractive Question Answering on Five Extractive QA datasets aggregated
0.91Calibration Score (C)Mistral-8x22B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Mistral-8x22BParams (B)=222025.12 | 0.91 | 0.78 | 0.73 | 0.81 | — | |
| DeepSeek R1 0528Params (B)=272025.12 | 0.87 | 0.76 | 0.63 | 0.75 | — | |
| Qwen3-235BParams (B)=222025.12 | 0.84 | 0.74 | 0.7 | 0.76 | — | |
| Llama 4 ScoutParams (B)=172025.12 | 0.81 | 0.7 | 0.64 | 0.72 | — | |
| MiniMax-Text-01Params (B)=252025.12 | 0.81 | 0.69 | 0.63 | 0.71 | — | |
| Gemma 2Params (B)=272025.12 | 0.71 | 0.68 | 0.71 | 0.7 | — | |
| Kimi K2Params (B)=152025.12 | 0.68 | 0.66 | 0.67 | 0.67 | — | |
| Mistral-7BParams (B)=72025.12 | 0.52 | 0.65 | 0.58 | 0.63 | — | |
| LLaMA-3-7BParams (B)=72025.12 | 0.16 | 0.54 | 0.44 | 0.57 | — | |
| Falcon-7BParams (B)=72025.12 | 0 | 0.51 | 0.41 | 0.52 | — |