Argument Reconstruction on Arguinas Argument Corpus
100ValidityAAR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AARMethod category=LLM-based Engine, Base LLM=Claude Sonnet 4.5 (no-think)2026.03 | 100 | 21.4 | |
| GAAR (general)Method category=LLM-based Engine, Base LLM=Claude Sonnet 4.5 (no-think)2026.03 | 100 | — | |
| GAAR (specific)Method category=LLM-based Engine, Base LLM=Claude Sonnet 4.5 (no-think)2026.03 | 100 | 46.5 | |
| GPT-5.2Method category=LLM Prompting, Thinking mode=xhigh2026.03 | 80.8 | 48.8 | |
| Claude Sonnet 4.5Method category=LLM Prompting, Thinking mode=think2026.03 | 60.8 | 48 | |
| Claude Sonnet 4.5Method category=LLM Prompting, Thinking mode=no-think2026.03 | 59.2 | 44.6 | |
| GPT-5.2Method category=LLM Prompting, Thinking mode=none2026.03 | 51.7 | 46.2 | |
| gpt-oss-120bMethod category=LLM Prompting2026.03 | 40.8 | 26.4 | |
| Qwen3-235B-A22B-Thinking-2507Method category=LLM Prompting2026.03 | 39.2 | 32.1 |