Steerability Evaluation on triage
22EffectDeepSeek V3.2
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| DeepSeek V3.2Reasoning=low2026.02 | 22 | 1.46 | 1.45 | 0.88 | 46 | 77.1 | 0 | |
| DeepSeek V3.2Reasoning=off2026.02 | 22 | 1.42 | 1.38 | 0.82 | 38 | 92.9 | 3.1 | |
| Qwen3-235BReasoning=low2026.02 | 18 | 1.11 | 1.06 | 0.88 | 64 | 67.1 | 6.4 | |
| Qwen3-235BReasoning=off2026.02 | 18 | 1.56 | 1.4 | 1.3 | 56 | 85.7 | 10 | |
| Grok 4.1 FastReasoning=off2026.02 | 16 | 1.22 | 0.73 | 1.27 | 63 | 80 | 30.4 | |
| Llama 3.3 70BReasoning=none2026.02 | 16 | 1.88 | 1.22 | 2.64 | 66 | 85.7 | 25 | |
| Grok 4.1 FastReasoning=low2026.02 | 12 | 0.85 | 0.78 | 0.44 | 53 | 37.1 | 11.5 | |
| Llama 3.3 70BReasoning=before2026.02 | 12 | 0.57 | 0.51 | 0.72 | 72 | 58.6 | 7.3 | |
| GPT-5.2Reasoning=low2026.02 | 7 | 0.34 | 0.18 | 0.42 | 66 | 42.9 | 26.7 | |
| GPT-5.2Reasoning=off2026.02 | 7 | 0.4 | -0.16 | 0.61 | 64 | 51.4 | 75 |