Logical Reasoning on BBEH mini
17AccuracyRL (PC-83K)
Evaluation Results
| Method | Links | |
|---|---|---|
| RL (PC-83K)Training Data=PC-83K, Base Model=Qwen2.5-7B-Instruct, Training Method=RL (GRPO)2025.08 | 17 | |
| RL (PC-83K+PC-SL-35K)Training Data=PC-83K + PC-SL-35K, Base Model=Qwen2.5-7B-Instruct, Training Method=RL2025.08 | 17 | |
| RL (PC-SL-35K)Training Data=PC-SL-35K, Base Model=Qwen2.5-7B-Instruct, Training Method=RL2025.08 | 16.5 | |
| Qwen2.5-7B-InstructBase Model=Qwen2.5-7B-Instruct2025.08 | 11.3 | |
| SFTTraining Data=PC-83K, Base Model=Qwen2.5-7B-Instruct2025.08 | 9.8 | |
| SynLogic-7BBase Model=7B Parameters2025.08 | 8 |