SG-DT parsing on NormBench (Random split)
100Done RateHeuristic Attach
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Heuristic AttachModel Category=Trainable Baselines2026.06 | 100 | 17.39 | 8.74 | 0.8409 | 100 | 0 | 62.78 | 25.6 | 83.83 | 1.14 | |
| Neural Biaffine ParserModel Category=Trainable Baselines2026.06 | 100 | 19.26 | 9.37 | 0.8287 | 100 | 0 | 67.2 | 28.13 | 84.01 | 2.27 | |
| Constrained DecoderModel Category=Trainable Baselines2026.06 | 100 | 19.26 | 9.29 | 0.8294 | 100 | 0 | 67.17 | 28.02 | 84.01 | 2.27 | |
| GPT-5.2 (2025-12-11)Model Category=Closed Frontier (General)2026.06 | 84.11 | 42.04 | 21.5 | 0.6866 | 79.29 | 20.71 | 62.91 | 47.06 | 77.57 | 56.87 | |
| Gemini-3-Pro-previewModel Category=Closed Frontier (General)2026.06 | 82.81 | 45.19 | 24.62 | 0.6572 | 77.28 | 22.72 | 64.34 | 50.28 | 76.04 | 54.52 | |
| Qwen3-235B-A22B-thinkingModel Category=Open / Open-Weight (General)2026.06 | 81.78 | 40.9 | 20.16 | 0.7031 | 73.39 | 26.52 | 62.39 | 47.29 | 74.31 | 48.79 | |
| DeepSeek-V3.2Model Category=Open / Open-Weight (General)2026.06 | 78.62 | 42.34 | 23.01 | 0.6788 | 73.5 | 26.41 | 59.73 | 47.51 | 73.15 | 61.5 | |
| Claude-Opus-4.5Model Category=Closed Frontier (General)2026.06 | 76.67 | 42.85 | 22.83 | 0.6806 | 70.7 | 29.11 | 59.15 | 47.8 | 71.52 | 64 | |
| LawGPTModel Category=Legal-Domain LLMs2026.06 | 76.3 | 15.77 | 4.81 | 0.8965 | 60.56 | 37.21 | 41.22 | 20.74 | 59.36 | 7.36 | |
| GLM-4.7Model Category=Open / Open-Weight (General)2026.06 | 74.72 | 38.56 | 20.72 | 0.7087 | 68.71 | 31.2 | 56.47 | 42.51 | 67.58 | 47.35 | |
| ChatLawModel Category=Legal-Domain LLMs2026.06 | 72.49 | 18.04 | 4.65 | 0.8896 | 69.26 | 30.74 | 43.34 | 22.55 | 57.9 | 0.65 | |
| LegalOne-8BModel Category=Legal-Domain LLMs2026.06 | 65.15 | 15.14 | 2.66 | 0.9222 | 54.55 | 44.8 | 35.04 | 15.82 | 54.53 | 29.39 |