Logical Reasoning on PC-83K (Normal)
71AccuracyRL (PC-83K)
Evaluation Results
| Method | Links | |
|---|---|---|
| RL (PC-83K)Training Data=PC-83K, Base Model=Qwen2.5-7B-Instruct, Training Method=RL (GRPO)2025.08 | 71 | |
| RL (PC-83K+PC-SL-35K)Training Data=PC-83K + PC-SL-35K, Base Model=Qwen2.5-7B-Instruct, Training Method=RL2025.08 | 64.8 | |
| SFTTraining Data=PC-83K, Base Model=Qwen2.5-7B-Instruct2025.08 | 61.9 | |
| RL (PC-SL-35K)Training Data=PC-SL-35K, Base Model=Qwen2.5-7B-Instruct, Training Method=RL2025.08 | 22 | |
| Qwen2.5-7B-InstructBase Model=Qwen2.5-7B-Instruct2025.08 | 16.8 |