Knowledge Editing on MQuAKE Story
100Fact Accuracy (Easy)SFT
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| SFTBackbone=Llama 3.1, Training strategy=SFT, Evaluation protocol=without CoT, Questions=single-hop2026.02 | 100 | 100 | 20.1 | 84.8 | 31.2 | |
| Learn from factsBackbone=Llama 3.1, Training strategy=Learn from facts, Evaluation protocol=without CoT, Questions=Baseline2026.02 | 100 | 96.2 | 52.1 | 67.2 | 18.8 | |
| Learn from factsBackbone=Llama 3.1, Training strategy=Learn from facts, Evaluation protocol=without CoT, Questions=multi-hop2026.02 | 100 | 100 | 94.7 | 96.6 | 66.7 | |
| Reasoning-centric Training FrameworkBackbone=Llama 3.1, Training strategy=Learn from stories, Evaluation protocol=without CoT, Questions=Baseline2026.02 | 100 | 99.1 | 63 | 75.2 | 19.8 | |
| Reasoning-centric Training FrameworkBackbone=Llama 3.1, Training strategy=Learn from stories, Evaluation protocol=without CoT, Questions=single-hop2026.02 | 100 | 98.2 | 94.6 | 92.9 | 43.8 | |
| Reasoning-centric Training FrameworkBackbone=Llama 3.1, Training strategy=Learn from stories, Evaluation protocol=without CoT, Questions=multi-hop2026.02 | 100 | 99.1 | 94.8 | 96.4 | 66.7 | |
| SFTBackbone=Llama 3.1, Training strategy=SFT, Evaluation protocol=with CoT, Questions=single-hop2026.02 | 100 | 99.4 | 17.1 | 86.7 | 25 | |
| Learn from factsBackbone=Llama 3.1, Training strategy=Learn from facts, Evaluation protocol=with CoT, Questions=single-hop2026.02 | 99.1 | 95.9 | 95.4 | 91.7 | 71.9 | |
| Reasoning-centric Training FrameworkBackbone=Llama 3.1, Training strategy=Learn from stories, Evaluation protocol=with CoT, Questions=Baseline2026.02 | 98.2 | 97.7 | 49.6 | 68.7 | 27.1 | |
| Reasoning-centric Training FrameworkBackbone=Llama 3.1, Training strategy=Learn from stories, Evaluation protocol=with CoT, Questions=single-hop2026.02 | 98.2 | 98.2 | 94.2 | 92.8 | 77.1 | |
| Reasoning-centric Training FrameworkBackbone=Llama 3.1, Training strategy=Learn from stories, Evaluation protocol=with CoT, Questions=multi-hop2026.02 | 98.2 | 97.4 | 96.1 | 96.2 | 95.8 | |
| Learn from factsBackbone=Llama 3.1, Training strategy=Learn from facts, Evaluation protocol=with CoT, Questions=multi-hop2026.02 | 97.4 | 93.6 | 93.8 | 93.8 | 88.5 | |
| Learn from factsBackbone=Llama 3.1, Training strategy=Learn from facts, Evaluation protocol=without CoT, Questions=single-hop2026.02 | 96.5 | 96.8 | 96.1 | 81.6 | 37.5 | |
| Learn from factsBackbone=Llama 3.1, Training strategy=Learn from facts, Evaluation protocol=with CoT, Questions=Baseline2026.02 | 96.5 | 94.2 | 44.5 | 66.1 | 21.9 |