Length-controlled text generation on LIFEBench
44Equal To Length DeviationQwen2.5-7B-Instruct
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen2.5-7B-InstructLenVM=1.5B2026.04 | 44 | 64.8 | 96.1 | 99.5 | |
| Qwen2.5-3B-InstructLenVM=1.5B2026.04 | 56 | 62.6 | 93 | 93.1 | |
| Qwen2.5-32B-InstructLenVM=1.7B2026.04 | 57 | 67.2 | 99.4 | 99.8 | |
| Claude-3-OpusModel Category=Closed-source frontier models2026.04 | 66 | 35.5 | 51.5 | 100 | |
| Qwen2.5-7B-InstructLenVM=None2026.04 | 71 | 30.9 | 98.5 | 89.1 | |
| GPT-4oModel Category=Closed-source frontier models2026.04 | 74 | 35.5 | 77.9 | 98.5 | |
| Qwen2.5-3B-InstructLenVM=None2026.04 | 83 | 25.6 | 92.1 | 94.6 | |
| Claude-3-Opus-thinkingModel Category=Closed-source frontier models2026.04 | 87 | 53.2 | 67.4 | 100 | |
| Qwen2.5-32B-InstructLenVM=None2026.04 | 90 | 36.8 | 87 | 99.3 | |
| Gemini-1.5-Pro-PreviewModel Category=Closed-source frontier models2026.04 | 91 | 49.3 | 70.7 | 100 | |
| Claude-3.5-SonnetModel Category=Closed-source frontier models2026.04 | 105 | 34.1 | 62.9 | 100 | |
| Gemini-1.5-Flash-PreviewModel Category=Closed-source frontier models2026.04 | 123 | 40.3 | 57.3 | 99.6 | |
| Claude-3.5-Sonnet-thinkingModel Category=Closed-source frontier models2026.04 | 124 | 51.3 | 69.3 | 100 | |
| GPT-4o-mini-thinkingModel Category=Closed-source frontier models2026.04 | 131 | 47.8 | 72.7 | 98.9 | |
| GPT-4o-miniModel Category=Closed-source frontier models2026.04 | 135 | 37.4 | 65.4 | 98.9 |