Question Answering on BLENDQA
48.9F1 ScoreSchema-Guided
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Schema-GuidedBackbone=GPT-4.1, Evaluation Setting=zero-shot, Sources=dataset-provided2026.06 | 48.9 | 29.3 | 54.1 | 72.3 | |
| AtomRBackbone=GPT-4.1, Evaluation Setting=zero-shot, Sources=dataset-provided2026.06 | 43.3 | 26.7 | 49.2 | 68.5 | |
| Best publishedEvaluation Setting=zero-shot / no-finetuning2026.06 | 43.3 | — | — | — | |
| single-passBackbone=GPT-4.1, Evaluation Setting=zero-shot, Ablation=removes multi-pass reasoning2026.06 | 41.8 | 25.4 | 47.9 | 67.9 | |
| w/o struc. intel.Backbone=GPT-4.1, Evaluation Setting=zero-shot, Ablation=removes structural intelligence2026.06 | 38.9 | 26.3 | 44.8 | 56 | |
| CoKBackbone=GPT-4.1, Evaluation Setting=zero-shot, Sources=dataset-provided2026.06 | 38.6 | 20.4 | 40.5 | 61.1 | |
| w/o KG ingestionBackbone=GPT-4.1, Evaluation Setting=zero-shot, Ablation=removes KG ingestion2026.06 | 36.5 | 21.3 | 43.1 | 58.2 | |
| ProbTreeBackbone=GPT-4.1, Evaluation Setting=zero-shot, Sources=dataset-provided2026.06 | 33.9 | 20.1 | 36.6 | 56.7 | |
| RAGBackbone=GPT-4.1, Evaluation Setting=zero-shot, Sources=dataset-provided2026.06 | 33.6 | 19.1 | 26.1 | 52.3 | |
| w/o schema guidanceBackbone=GPT-4.1, Evaluation Setting=zero-shot, Ablation=removes schema guidance2026.06 | 32.3 | 20 | 39.3 | 53.3 |