Value Consistency Evaluation 1.0 (English)
4.25Sov ScoreVC-LLM
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| VC-LLMTraining=Adversarial Training Alignment2026.01 | 4.25 | 4.57 | 4.95 | 4.88 | 4.83 | 799 | 4.59 | |
| Qwen2.5-72B-InstructModel Size=72B, Model Type=Instruct2026.01 | 3.68 | 3.33 | 3.53 | 3.69 | 3.61 | 626 | 3.6 | |
| Qwen2.5-14B-InstructModel Size=14B, Model Type=Instruct2026.01 | 3.52 | 3.03 | 3.32 | 3.33 | 4.22 | 599 | 3.44 | |
| Baichuan2-7B-ChatModel Size=7B, Model Type=Chat2026.01 | 3.31 | 2.3 | 2.68 | 3 | 2.11 | 499 | 2.87 | |
| Qwen2.5-7B-InstructModel Size=7B, Model Type=Instruct2026.01 | 3.31 | 3.13 | 3.42 | 3.19 | 3.78 | 576 | 3.31 | |
| Baichuan2-13B-ChatModel Size=13B, Model Type=Chat2026.01 | 2.6 | 2.47 | 2.95 | 2.86 | 2.44 | 463 | 2.66 | |
| Llama-3-Ch-8B-V3Model Size=8B, Model Type=Chinese-Variant2026.01 | 2.35 | 2.37 | 2.68 | 2.4 | 1.89 | 410 | 2.36 | |
| Yi-6B-ChatModel Size=6B, Model Type=Chat2026.01 | 2.29 | 2.37 | 2.79 | 1.98 | 2.61 | 403 | 2.32 | |
| Llama-3-8B-InstructModel Size=8B, Model Type=Instruct2026.01 | 2.23 | 1.83 | 2.05 | 1.93 | 1.5 | 347 | 1.99 | |
| Ziya2-13B-ChatModel Size=13B, Model Type=Chat2026.01 | 2.18 | 2.2 | 2.42 | 2.36 | 2.22 | 393 | 2.26 | |
| GPT-4o2026.01 | 2.15 | 2.13 | 2.16 | 2.14 | 1.67 | 365 | 2.1 | |
| GLM-4-9B-ChatModel Size=9B, Model Type=Chat2026.01 | 2.05 | 1.8 | 2.11 | 1.98 | 1.33 | 334 | 1.92 | |
| Mistral-7B-InstructModel Size=7B, Model Type=Instruct2026.01 | 1.88 | 2.13 | 2.32 | 2.21 | 1.89 | 357 | 2.05 | |
| TigerBot-7BModel Size=7B2026.01 | 1.69 | 1.27 | 1.58 | 1.64 | 1.56 | 275 | 1.58 | |
| internlm2.5-7b-chatModel Size=7B, Model Type=Chat2026.01 | 1.6 | 1.43 | 1.16 | 1.33 | 2.11 | 263 | 1.51 |