Deductive logical reasoning on ProofWriter (test)
100ExcRateSFT+
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SFT+Model=Phi-4-mini-Instruct, Inference Type=SFT on symbolic ProofWriter and FOLIO combined2026.01 | 100 | 97.83 | |
| IncrementModel=Phi-4-mini-Instruct, Inference Type=Incremental inference2026.01 | 100 | 97.83 | |
| SFT+Model=Gemma-3-4B-Instruct, Inference Type=SFT on symbolic ProofWriter and FOLIO combined2026.01 | 98.83 | 96.83 | |
| IncrementModel=Gemma-3-4B-Instruct, Inference Type=Incremental inference2026.01 | 98.5 | 96.5 | |
| SFT+Model=Qwen2.5-3B-Instruct, Inference Type=SFT on symbolic ProofWriter and FOLIO combined2026.01 | 96.83 | 93.33 | |
| IncrementModel=Qwen2.5-3B-Instruct, Inference Type=Incremental inference2026.01 | 96.5 | 93 | |
| IncrementModel=Qwen3-4B-Instruct-2507, Inference Type=Incremental inference2026.01 | 96.5 | 94.17 | |
| SFT+Model=Qwen3-4B-Instruct-2507, Inference Type=SFT on symbolic ProofWriter and FOLIO combined2026.01 | 95 | 92.83 | |
| ICLModel=Phi-4-mini-Instruct, Inference Type=In-context learning (5-shots)2026.01 | 87 | 68.67 | |
| ICLModel=Gemma-3-4B-Instruct, Inference Type=In-context learning (5-shots)2026.01 | 74.5 | 58.33 | |
| ICLModel=Qwen3-4B-Instruct-2507, Inference Type=In-context learning (5-shots)2026.01 | 42.67 | 35.5 | |
| ICLModel=Qwen2.5-3B-Instruct, Inference Type=In-context learning (5-shots)2026.01 | 36.67 | 25.5 |