First-Order Logic translation on FOLIO (test)
66BLEUQwen3-1.7B-SGRPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-1.7B-SGRPOTraining=SGRPO2025.12 | 66 | — | 87.4 | |
| Qwen3-1.7B-SFTTraining=Supervised Fine-Tuning2025.12 | 61.2 | — | 85 | |
| ChatGPT-4o2025.12 | 38.4 | 82.6 | 80.9 | |
| LogicLLaMA-13Bconfiguration=RLHF Corre.2025.12 | 38.4 | 85.8 | — | |
| LogicLLaMA-7Bconfiguration=RLHF Corre.2025.12 | 37.8 | 84.1 | — | |
| DeepSeek-V32025.12 | 37.6 | 83 | 79.2 | |
| ChatGPT-3.52025.12 | 37 | 80.2 | 77.6 |