Error Attribution on Who&When
8.1Pair µF1Qwen3-8B-Thinking + GRPO
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Qwen3-8B-Thinking + GRPOModel Scale=Medium-Scale, Training Protocol=GRPO2025.09 | 8.1 | 3.19 | 53.12 | 40.52 | 11.25 | 6.91 | 17.58 | — | |
| gpt-oss-120bModel Scale=Large-Scale, Evaluation Protocol=Zero-shot2025.09 | 8.09 | 3.3 | 51.56 | 35.41 | 14.75 | 6.98 | 17.07 | — | |
| o3Model Scale=Large-Scale, Evaluation Protocol=Zero-shot2025.09 | 7.41 | 3.98 | 53.1 | 42.55 | 14.88 | 8.63 | 20.24 | — | |
| Gemini-2.5-FlashModel Scale=Large-Scale, Evaluation Protocol=Zero-shot2025.09 | 7.32 | 3.33 | 55.56 | 36.98 | 11.94 | 7.96 | 19.55 | — | |
| Gemini-2.5-ProModel Scale=Large-Scale, Evaluation Protocol=Zero-shot2025.09 | 6.81 | 2.69 | 53.11 | 34.92 | 11.07 | 8.11 | 18.35 | — | |
| Claude-Sonnet-4Model Scale=Large-Scale, Evaluation Protocol=Zero-shot2025.09 | 6.77 | 2.66 | 44.76 | 37.23 | 13.33 | 9.23 | 18.16 | — | |
| VerifyMAS-7B-SFTBackbone=Qwen2.5-7B, Setting=SFT2026.05 | 5.25 | 2.44 | 42 | 33.69 | 14.85 | 13.25 | — | — | |
| VerifyMAS-8B-SFTBackbone=Qwen3-8B, Setting=SFT2026.05 | 5.21 | 2.85 | 40.09 | 32.9 | 13.35 | 12.32 | — | — | |
| Qwen3-8B-Non-Thinking + SFTModel Scale=Medium-Scale, Training Protocol=SFT2025.09 | 5.17 | 2.33 | 45.48 | 30.77 | 8 | 5.29 | 21.41 | — | |
| Qwen3-8B-SFTBackbone=Qwen3-8B, Setting=SFT2026.05 | 5.17 | 2.33 | 45.48 | 30.77 | 8 | 5.29 | — | — | |
| Qwen2.5-14B-Instruct + Aegis-SFTModel Scale=Medium-Scale, Training Protocol=SFT2025.09 | 4.03 | 2.08 | 51.14 | 36.94 | 9.87 | 7.77 | 26.51 | — | |
| Qwen3-8B-Non-ThinkingModel Scale=Medium-Scale2025.09 | 3.88 | 1.81 | 27.78 | 17.64 | 3.88 | 1.91 | 10.12 | — | |
| Qwen2.5-72B-InstructModel Scale=Large-Scale, Evaluation Protocol=Zero-shot2025.09 | 3.56 | 2.11 | 44.44 | 26.05 | 5.59 | 4.34 | 15.01 | — | |
| GPT-4.1Model Scale=Large-Scale, Evaluation Protocol=Zero-shot2025.09 | 3.36 | 1.16 | 42.29 | 28.93 | 7 | 5.84 | 15.27 | — | |
| Qwen2.5-14B-Instruct + Aegis-GRPOModel Scale=Medium-Scale, Training Protocol=GRPO2025.09 | 2.45 | 1.49 | 54.43 | 40.88 | 4.15 | 2.67 | 18.41 | — | |
| Qwen2.5-7B-InstructModel Scale=Medium-Scale, Evaluation Protocol=Zero-shot2025.09 | 2.31 | 1.14 | 40.92 | 23.5 | 3.64 | 1.77 | 12.43 | — | |
| Qwen2.5-7B-Instruct + GRPOModel Scale=Medium-Scale, Training Protocol=GRPO2025.09 | 2.31 | 1.19 | 50.77 | 30.14 | 3.86 | 2.3 | 14.87 | — | |
| Qwen3-8B-Non-Thinking + GRPOModel Scale=Medium-Scale, Training Protocol=GRPO2025.09 | 2.21 | 1.45 | 50.94 | 38.26 | 2.21 | 1.68 | 17.15 | — | |
| GPT-4o-miniModel Scale=Large-Scale, Evaluation Protocol=Zero-shot2025.09 | 2.11 | 0.98 | 47.42 | 34.21 | 5.26 | 3.33 | 15.83 | — | |
| Qwen3-8B-ThinkingModel Scale=Medium-Scale2025.09 | 1.95 | 1.1 | 37.91 | 27.58 | 4.65 | 2.21 | 13.06 | — | |
| DCLModel Scale=Small-Scale2025.09 | 1.6 | 0.77 | 8.4 | 6.07 | 14.67 | 10.57 | 12.61 | — | |
| DCLSetting=SFT2026.05 | 1.6 | 0.77 | 8.4 | 6.07 | 14.67 | 10.57 | — | — | |
| Qwen2.5-7B-Instruct + SFTModel Scale=Medium-Scale, Training Protocol=SFT2025.09 | 1.26 | 0.52 | 43.51 | 32.51 | 6.77 | 4.2 | 17.99 | — | |
| Qwen2.5-7B-SFTBackbone=Qwen2.5-7B, Setting=SFT2026.05 | 1.26 | 0.52 | 43.51 | 32.51 | 6.77 | 4.2 | — | — | |
| only-mix headModel Scale=Small-Scale, Ablation Setting=only-mix head2025.09 | 1.2 | 0.6 | 9.4 | 6.07 | — | 10.67 | 12.42 | — | |
| w/o intentModel Scale=Small-Scale, Ablation Setting=w/o intent2025.09 | 1.1 | 0.55 | 8 | 5.9 | 11 | 9 | 10.16 | — | |
| only-bilinearModel Scale=Small-Scale, Ablation Setting=only-bilinear2025.09 | 0.6 | 0.53 | 7.03 | 5.43 | 12.97 | 11.4 | 10.01 | — | |
| w/o consistencyModel Scale=Small-Scale, Ablation Setting=w/o consistency2025.09 | 0.5 | 0.43 | 6.27 | 5.2 | 11.8 | 9.5 | 9.52 | — | |
| Random2025.09 | 0.11 | 0.05 | 1.06 | 0.83 | 8.74 | 7.14 | 4.08 | — | |
| Qwen2.5-14B-InstructModel Scale=Medium-Scale, Evaluation Protocol=Zero-shot2025.09 | 0 | 0 | 49.88 | 33.19 | 1.56 | 1.35 | 13.99 | — | |
| CRSVPScoring Function Type=Naive LLM, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 21 | |
| CRSVPScoring Function Type=Role-prompted, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 23 | |
| Left FilteringScoring Function Type=Naive LLM, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 12 | |
| Left FilteringScoring Function Type=Role-prompted, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 12 | |
| Right FilteringScoring Function Type=Naive LLM, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 31 | |
| Right FilteringScoring Function Type=Role-prompted, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 30 | |
| Two-WayScoring Function Type=Naive LLM, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 19 | |
| Two-WayScoring Function Type=Role-prompted, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 22 | |
| VanillaScoring Function Type=Naive LLM, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 20 | |
| VanillaScoring Function Type=Role-prompted, Backbone=gpt-4o-mini2026.05 | — | — | — | — | — | — | — | 19 |