Legal Reasoning on PRBench-Legal 500 prompts
44.1GPT-5 Grader ScoreQwen3.5-4B RL on Agentic Self-Instruct data
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3.5-4B RL on Agentic Self-Instruct dataRL training data=Agentic Self-Instruct, Base model size=4B2026.06 | 44.1 | 39.3 | |
| Qwen3.5-397BRL training data=no additional RL, Base model size=397B2026.06 | 40.4 | 35.8 | |
| Qwen3.5-4B RL on CoT Self-Instruct dataRL training data=CoT Self-Instruct, Base model size=4B2026.06 | 37.7 | 34.3 | |
| Qwen3.5-4BRL training data=no additional RL, Base model size=4B2026.06 | 28 | 24.5 |