Terms of Penalty Prediction on JurisMM-Text (test)
40.9AccuracyJurisMMA
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| JurisMMABackbone=GPT-4o, Knowledge Base (KB)=true, Multi-agent collaboration (MA)=true2026.01 | 40.9 | 52 | 39.3 | 36.1 | 0.848 | |
| JurisMMABackbone=GPT-4o, Knowledge Base (KB)=false, Multi-agent collaboration (MA)=true2026.01 | 37.8 | 45.4 | 37.1 | 33.5 | 0.848 | |
| TOPJUDGE2026.01 | 37.3 | 35.8 | 36.4 | 34.2 | — | |
| GPT-4o-202503262026.01 | 37.2 | 43.2 | 35.7 | 31.9 | 0.842 | |
| TextCNN2026.01 | 32.5 | 32.1 | 28.2 | 28.5 | — | |
| MPBFN2026.01 | 32.4 | 34.3 | 26.7 | 26.3 | — | |
| JurisMMABackbone=GPT-4o, Knowledge Base (KB)=true, Multi-agent collaboration (MA)=false2026.01 | 30.3 | 47.5 | 30.4 | 27.2 | 0.803 | |
| Qwen2.5-VL-7BSFT=true2026.01 | 25.2 | 12.3 | 9.6 | 4 | 0.818 | |
| Qwen2.5-7B-Instruct2026.01 | 23.7 | 22.7 | 14.5 | 12.4 | 0.826 | |
| JurisMMABackbone=Qwen2.5-VL-7B, Knowledge Base (KB)=false, Multi-agent collaboration (MA)=true2026.01 | 21.6 | 25.8 | 14.8 | 13.2 | 0.782 | |
| JurisMMABackbone=Qwen2.5-VL-7B, Knowledge Base (KB)=true, Multi-agent collaboration (MA)=false2026.01 | 20.7 | 25.5 | 14.9 | 12.5 | 0.789 | |
| GLM-4V-9B2026.01 | 19.9 | 32 | 19.7 | 16.3 | 0.668 | |
| Qwen2.5-VL-3B-Instruct2026.01 | 19.3 | 16.9 | 12.2 | 10 | 0.787 | |
| JurisMMABackbone=Qwen2.5-VL-7B, Knowledge Base (KB)=true, Multi-agent collaboration (MA)=true2026.01 | 18.9 | 17.5 | 12.3 | 10.3 | 0.784 | |
| Qwen2.5-VL-7B-Instruct2026.01 | 18.8 | 13.5 | 11.9 | 9.6 | 0.799 | |
| mPLUG-7B2026.01 | 10 | 25 | 16.8 | 6.7 | 0.507 |