Long-CoT Question Generation on CoDiQ-Bench 1.0 (test)
4.2Dialogue RoundsQwen3-4B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Qwen3-4BPrompting Strategy=Direct Prompt2026.02 | 4.2 | 1,419.7 | 36.8 | 40.4 | 38.6 | |
| Qwen3-8BPrompting Strategy=Direct Prompt2026.02 | 3.4 | 1,130.5 | 39.2 | 59.6 | 49.4 | |
| CoDiQ-Gen-8BPrompting Strategy=CoDiQ Generator2026.02 | 3.4 | 7,499.6 | 58.9 | 58.1 | 58.5 | |
| Qwen3-1.7BPrompting Strategy=Direct Prompt2026.02 | 3.3 | 844.5 | 25.6 | 37.1 | 31.4 | |
| Qwen3-14BPrompting Strategy=Direct Prompt2026.02 | 3.1 | 2,076.4 | 45.9 | 44.4 | 45.2 | |
| GPT-OSS-20BPrompting Strategy=Direct Prompt2026.02 | 2.9 | 5,528.2 | 68.5 | 74.4 | 71.5 | |
| GLM-4.6Prompting Strategy=Direct Prompt2026.02 | 2.8 | 3,385.8 | 71.2 | 65.8 | 68.5 | |
| Qwen3-4BPrompting Strategy=CoDiQ Prompt2026.02 | 2.8 | 4,422.3 | 49.1 | 42.7 | 45.9 | |
| GLM-Z1-9B-0414Prompting Strategy=Direct Prompt2026.02 | 2.7 | 1,229.8 | 48.8 | 43.7 | 46.3 | |
| GLM-4.6Prompting Strategy=CoDiQ Prompt2026.02 | 2.7 | 7,143.8 | 73.2 | 83.3 | 78.3 | |
| Qwen3-14BPrompting Strategy=CoDiQ Prompt2026.02 | 2.6 | 5,281.9 | 53.9 | 44.2 | 49.1 | |
| Qwen3-0.6BPrompting Strategy=Direct Prompt2026.02 | 2.4 | 314.3 | 17.2 | 35 | 26.1 | |
| Qwen3-8BPrompting Strategy=CoDiQ Prompt2026.02 | 2.4 | 4,155.6 | 49.8 | 41.9 | 45.8 | |
| Qwen3-32BPrompting Strategy=Direct Prompt2026.02 | 2.3 | 1,239.3 | 50.6 | 54.8 | 52.7 | |
| Qwen3-32BPrompting Strategy=CoDiQ Prompt2026.02 | 2.2 | 4,893.6 | 63 | 46.5 | 54.8 | |
| GPT-OSS-20BPrompting Strategy=CoDiQ Prompt2026.02 | 2.1 | 8,057.3 | 63.8 | 61.5 | 62.7 | |
| GLM-Z1-9B-0414Prompting Strategy=CoDiQ Prompt2026.02 | 1.7 | 3,638.3 | 54.7 | 30 | 42.4 | |
| Qwen3-1.7BPrompting Strategy=CoDiQ Prompt2026.02 | 1.4 | 2,975.7 | 32.3 | 37.3 | 34.8 | |
| Qwen3-0.6BPrompting Strategy=CoDiQ Prompt2026.02 | 1 | 2,052.7 | 22.4 | 29.2 | 25.8 |