Knowledge Editing on MQuAKE-Story 1.0 (test)
100Fact Accuracy (Easy)Qwen3 (SFT)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Qwen3 (SFT)Training Supervision=SFT, Questions=single-hop, Evaluation Mode=without CoT2026.02 | 100 | 100 | 88.4 | 86.5 | 36.5 | |
| Qwen3 (Learn from facts)Training Supervision=Learn from facts, Questions=single-hop, Evaluation Mode=without CoT2026.02 | 100 | 99.1 | 90.7 | 85.1 | 51 | |
| Qwen3 (Learn from facts)Training Supervision=Learn from facts, Questions=multi-hop, Evaluation Mode=without CoT2026.02 | 100 | 100 | 91.9 | 95.2 | 59.4 | |
| Reasoning-centric Training Framework (Learn from stories)Training Supervision=Learn from stories, Questions=single-hop, Evaluation Mode=without CoT2026.02 | 100 | 98.2 | 91.7 | 88.7 | 62.5 | |
| Reasoning-centric Training Framework (Learn from stories)Training Supervision=Learn from stories, Questions=multi-hop, Evaluation Mode=without CoT2026.02 | 100 | 99.4 | 90.4 | 93.8 | 71.9 | |
| Qwen3 (Learn from facts)Training Supervision=Learn from facts, Questions=Baseline, Evaluation Mode=without CoT2026.02 | 98.2 | 95 | 87.5 | 68.7 | 29.2 | |
| Reasoning-centric Training Framework (Learn from stories)Training Supervision=Learn from stories, Questions=single-hop, Evaluation Mode=with CoT2026.02 | 98.2 | 96.5 | 95.2 | 92 | 84.4 | |
| Reasoning-centric Training Framework (Learn from stories)Training Supervision=Learn from stories, Questions=multi-hop, Evaluation Mode=with CoT2026.02 | 98.2 | 97.1 | 95.4 | 94.4 | 99 | |
| Qwen3 (SFT)Training Supervision=SFT, Questions=single-hop, Evaluation Mode=with CoT2026.02 | 96.5 | 97.7 | 90.1 | 83.5 | 37.5 | |
| Reasoning-centric Training Framework (Learn from stories)Training Supervision=Learn from stories, Questions=Baseline, Evaluation Mode=without CoT2026.02 | 91.2 | 85.7 | 84.4 | 44.3 | 31.2 | |
| Qwen3 (Learn from facts)Training Supervision=Learn from facts, Questions=single-hop, Evaluation Mode=with CoT2026.02 | 86.8 | 81.9 | 95.9 | 80 | 68.8 | |
| Qwen3 (Learn from facts)Training Supervision=Learn from facts, Questions=multi-hop, Evaluation Mode=with CoT2026.02 | 85.1 | 78.4 | 96.2 | 81.7 | 72.9 | |
| Qwen3 (Learn from facts)Training Supervision=Learn from facts, Questions=Baseline, Evaluation Mode=with CoT2026.02 | 83.3 | 76.9 | 91.7 | 50.5 | 45.8 | |
| Reasoning-centric Training Framework (Learn from stories)Training Supervision=Learn from stories, Questions=Baseline, Evaluation Mode=with CoT2026.02 | 70.2 | 61.7 | 91.8 | 37.6 | 41.7 |