Task-solving on CL-bench (test)
25.8Overall Score (%)Ctx2Skill
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Ctx2SkillBackbone=GPT-5.1, Inference-time Augmentation=true2026.04 | 25.8 | 27.9 | 24.9 | 26.9 | 19.1 | |
| AutoSkill4DocBackbone=GPT-5.1, Inference-time Augmentation=true2026.04 | 22.7 | 25.3 | 21.5 | 23.1 | 16 | |
| PromptingBackbone=GPT-5.1, Inference-time Augmentation=true2026.04 | 22.1 | 24.7 | 21.1 | 22.4 | 15.5 | |
| Ctx2SkillBackbone=GPT-5.2, Inference-time Augmentation=true2026.04 | 21.4 | 22.2 | 20.4 | 25.4 | 12.6 | |
| GPT-5.1Inference Strategy=Base, Inference-time Augmentation=false2026.04 | 21.1 | 22.4 | 21 | 22.8 | 13.6 | |
| Claude Opus 4.5Inference Strategy=Base, Inference-time Augmentation=false2026.04 | 21 | 23.7 | 19 | 22.6 | 15.1 | |
| AutoSkill4DocBackbone=GPT-5.2, Inference-time Augmentation=true2026.04 | 19.7 | 20.5 | 18.8 | 23 | 11.6 | |
| Kimi K2.5Inference Strategy=Base, Inference-time Augmentation=false2026.04 | 19.2 | 19.1 | 19.4 | 21.3 | 14.4 | |
| PromptingBackbone=GPT-5.2, Inference-time Augmentation=true2026.04 | 19.1 | 19.6 | 18.3 | 22.6 | 11.1 | |
| GPT-5.2Inference Strategy=Base, Inference-time Augmentation=false2026.04 | 18.2 | 19.5 | 18 | 19.1 | 12.1 | |
| Ctx2SkillBackbone=GPT-4.1, Inference-time Augmentation=true2026.04 | 16.5 | 16.8 | 17.6 | 17.6 | 9.7 | |
| Gemini 3 ProInference Strategy=Base, Inference-time Augmentation=false2026.04 | 15.8 | 15.5 | 17.7 | 16.4 | 10.1 | |
| DeepSeek V3.2 ThinkingInference Strategy=Base, Inference-time Augmentation=false2026.04 | 13.2 | 13.6 | 13.8 | 14.2 | 8 | |
| AutoSkill4DocBackbone=GPT-4.1, Inference-time Augmentation=true2026.04 | 13.2 | 13.3 | 13.1 | 15 | 8.7 | |
| PromptingBackbone=GPT-4.1, Inference-time Augmentation=true2026.04 | 12.3 | 12.4 | 12.3 | 13.9 | 8.2 | |
| GPT-4.1Inference Strategy=Base, Inference-time Augmentation=false2026.04 | 11.1 | 10.6 | 14.8 | 10.4 | 4.6 |