Legal Reasoning on LegalBench (Comprehensive Metrics)
82.3AccuracyQwQ-32B
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| QwQ-32BModel Type=Natively-trained reasoning LM, Parameters=32B2026.06 | 82.3 | 62.4 | 76.6 | 73.7 | 80.9 | 74.7 | 77.2 | 71.2 | |
| DeepSeek-R1-8BModel Type=Distilled LRM, Parameters=8B2026.06 | 76.2 | 66.6 | 69.9 | 67.4 | 74.6 | 77.9 | 79.3 | 67.8 |