Legal Inquisitive Dialogue on U.S. Supreme Court Oral Argument dataset
4.01CS ScoreSaulLM-7B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| SaulLM-7BDescription=Specialized legal LLM2026.05 | 4.01 | 3.91 | 4.56 | 3.75 | 4.06 | |
| Dual-agent hierarchical RL (Ours)2026.05 | 4.01 | 3.98 | 4.89 | 4.47 | 4.34 | |
| VaRMIMechanism=Offline policy gradient2026.05 | 4 | 3.94 | 4.71 | 3.93 | 4.15 | |
| Vanilla Llama3Type=Prompt-only, Base Model=Llama3-8B-Instruct2026.05 | 3.99 | 3.94 | 4.7 | 3.92 | 4.14 | |
| HudečekApproach=Structured pipeline2026.05 | 3.99 | 3.97 | 4.77 | 3.63 | 4.09 | |
| SFT Llama3Type=Fine-tuned, Base Model=Llama3-8B-Instruct2026.05 | 3.98 | 3.81 | 4.45 | 3.38 | 3.91 | |
| ArCHerStructure=Hierarchical Actor-Critic, Variant=Offline2026.05 | 3.96 | 3.79 | 4.17 | 4.22 | 4.04 |