Deductive logical reasoning on ProntoQA (test)
2.8Error RateICL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ICLModel=Qwen3-4B-Instruct-2507, Inference Type=In-context learning (5-shots)2026.01 | 2.8 | 1.2 | |
| ICLModel=Qwen2.5-3B-Instruct, Inference Type=In-context learning (5-shots)2026.01 | 4.4 | 0.4 | |
| ICLModel=Gemma-3-4B-Instruct, Inference Type=In-context learning (5-shots)2026.01 | 5.6 | 0.8 | |
| ICLModel=Phi-4-mini-Instruct, Inference Type=In-context learning (5-shots)2026.01 | 6.4 | 2.6 | |
| SFT+Model=Phi-4-mini-Instruct, Inference Type=SFT on symbolic ProofWriter and FOLIO combined2026.01 | 44.6 | 30.6 | |
| IncrementModel=Gemma-3-4B-Instruct, Inference Type=Incremental inference2026.01 | 50.8 | 23.6 | |
| SFT+Model=Gemma-3-4B-Instruct, Inference Type=SFT on symbolic ProofWriter and FOLIO combined2026.01 | 51.8 | 23.2 | |
| IncrementModel=Phi-4-mini-Instruct, Inference Type=Incremental inference2026.01 | 52.4 | 34.8 | |
| IncrementModel=Qwen2.5-3B-Instruct, Inference Type=Incremental inference2026.01 | 54.4 | 19.2 | |
| SFT+Model=Qwen2.5-3B-Instruct, Inference Type=SFT on symbolic ProofWriter and FOLIO combined2026.01 | 55.6 | 21.6 | |
| SFT+Model=Qwen3-4B-Instruct-2507, Inference Type=SFT on symbolic ProofWriter and FOLIO combined2026.01 | 62.2 | 51.6 | |
| IncrementModel=Qwen3-4B-Instruct-2507, Inference Type=Incremental inference2026.01 | 80.6 | 64.4 |