Logical Reasoning on PC-83K Hard
61AccuracyRL (PC-83K)
Evaluation Results
| Method | Links | |
|---|---|---|
| RL (PC-83K)Training Data=PC-83K, Base Model=Qwen2.5-7B-Instruct, Training Method=RL (GRPO)2025.08 | 61 | |
| RL (PC-83K+PC-SL-35K)Training Data=PC-83K + PC-SL-35K, Base Model=Qwen2.5-7B-Instruct, Training Method=RL2025.08 | 54.1 | |
| SFTTraining Data=PC-83K, Base Model=Qwen2.5-7B-Instruct2025.08 | 48 | |
| RL (PC-SL-35K)Training Data=PC-SL-35K, Base Model=Qwen2.5-7B-Instruct, Training Method=RL2025.08 | 14.3 | |
| Qwen2.5-7B-InstructBase Model=Qwen2.5-7B-Instruct2025.08 | 12.1 |