Reasoning on GSM8K
1AccuracyGPT-5.2
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5.2Category=MLLM, Note=direct2026.02 | 1 | — | — | |
| Claude-Sonnet 4.5Category=MLLM, Note=direct2026.02 | 1 | — | — | |
| TwC (Ours) - Only ImageCategory=Think with Comic, Note=direct2026.02 | 1 | — | — | |
| Gemini-3-ProCategory=MLLM, Note=direct2026.02 | 0.99 | — | — | |
| DeepSeek-R1Category=Reasoning LLM, Note=CoT2026.02 | 0.961 | — | — | |
| TwC (Ours) - Img & TxtCategory=Think with Comic, Note=G-t-R2026.02 | 0.954 | — | — | |
| Qwen3-235B-A22BCategory=Reasoning LLM, Note=CoT2026.02 | 0.943 | — | — | |
| eMoTModels=Qwen-32B2026.06 | 0.934 | — | — | |
| Qwen3-8BLoss function=base+aux2026.02 | 0.9333 | — | — | |
| BoTModels=GPT-42026.06 | 0.933 | — | — | |
| Qwen3-8BLoss function=base2026.02 | 0.9242 | — | — | |
| Qwen3-4BLoss function=base2026.02 | 0.9234 | — | — | |
| Qwen3-4BLoss function=base+aux2026.02 | 0.9181 | — | — | |
| RM-RegenBase Model=Llama 3.1-8B2026.03 | 0.872 | — | — | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=32026.03 | 0.87 | — | — | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=22026.03 | 0.866 | — | — | |
| Persona SwitchModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 0.8575 | — | — | |
| BaselineTopology=Complete2025.11 | 0.8541 | — | — | |
| Random SelectionModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 0.8471 | — | — | |
| Low-gap SelectionModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 0.837 | — | — | |
| GreedyModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 0.8362 | — | — | |
| RM-RegenBase Model=GPT-3.52026.03 | 0.836 | — | — | |
| ProCoBase Model=Llama 3.1-8B, Iterations=22026.03 | 0.834 | — | — | |
| ProCoBase Model=Llama 3.1-8B, Iterations=32026.03 | 0.834 | — | — | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=52026.03 | 0.832 | — | — | |
| MultinomialModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 0.8317 | — | — | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=42026.03 | 0.83 | — | — | |
| ProCoBase Model=Llama 3.1-8B, Iterations=52026.03 | 0.826 | — | — | |
| ST CoTBase Model=GPT-3.5, Iterations=22026.03 | 0.822 | — | — | |
| ProCoBase Model=GPT-3.5, Iterations=42026.03 | 0.818 | — | — | |
| MAS-ShieldTopology=Complete2025.11 | 0.8179 | 16.71 | — | |
| BaselineTopology=Chain2025.11 | 0.8173 | — | — | |
| MAS-ShieldTopology=Chain2025.11 | 0.8146 | 25.72 | — | |
| LLaMA3.1-8BLoss function=base+aux2026.02 | 0.8118 | — | — | |
| s1.1-7B2026.02 | 0.81 | — | — | |
| ProCoBase Model=GPT-3.5, Iterations=32026.03 | 0.808 | — | — | |
| ST CoTBase Model=GPT-3.5, Iterations=42026.03 | 0.806 | — | — | |
| ProCoBase Model=Llama 3.1-8B, Iterations=42026.03 | 0.804 | — | — | |
| ST CoTBase Model=GPT-3.5, Iterations=32026.03 | 0.802 | — | — | |
| BaselineTopology=Tree2025.11 | 0.8007 | — | — | |
| ProCoBase Model=GPT-3.5, Iterations=22026.03 | 0.8006 | — | — | |
| Top-pModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 0.7991 | — | — | |
| Top-kModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 0.7983 | — | — | |
| LLaMA3.1-8BLoss function=base2026.02 | 0.7885 | — | — | |
| MAS-ShieldTopology=Tree2025.11 | 0.7792 | 16.19 | — | |
| Role-Play PromptingModel=LLaMA-3.1 (8B), Zero-shot=true2026.01 | 0.7703 | — | — | |
| Persona SwitchModel=LLaMA-3.2 (3B), Zero-shot=true2026.01 | 0.7665 | — | — | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=22026.03 | 0.758 | — | — | |
| Sora 2Category=Think with Video, Note=V-o-T, Citation Source=Tong et al., 20252026.02 | 0.757 | — | — | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=32026.03 | 0.756 | — | — | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=42026.03 | 0.756 | — | — | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=52026.03 | 0.754 | — | — | |
| Random SelectionModel=LLaMA-3.2 (3B), Zero-shot=true2026.01 | 0.7473 | — | — | |
| BaselineTopology=Circle2025.11 | 0.7386 | — | — | |
| Self-RefineBase Model=GPT-3.5, Iterations=22026.03 | 0.734 | — | — | |
| GreedyModel=LLaMA-3.2 (3B), Zero-shot=true2026.01 | 0.7309 | — | — | |
| MultinomialModel=LLaMA-3.2 (3B), Zero-shot=true2026.01 | 0.7301 | — | — | |
| Low-gap SelectionModel=LLaMA-3.2 (3B), Zero-shot=true2026.01 | 0.7271 | — | — | |
| BaselineTopology=Star2025.11 | 0.7267 | — | — | |
| MAS-ShieldTopology=Circle2025.11 | 0.7217 | 10.42 | — | |
| PaLModels=Codex (175B)2026.06 | 0.72 | — | — | |
| Top-pModel=LLaMA-3.2 (3B), Zero-shot=true2026.01 | 0.7172 | — | — | |
| Role-Play PromptingModel=LLaMA-3.2 (3B), Zero-shot=true2026.01 | 0.7149 | — | — | |
| DGR SFTAlignment Dataset=DirectRefusal, Backbone=s1.1-7B2026.02 | 0.711 | — | — | |
| Self-RefineBase Model=GPT-3.5, Iterations=42026.03 | 0.706 | — | — | |
| Self-RefineBase Model=GPT-3.5, Iterations=32026.03 | 0.704 | — | — | |
| MAS-ShieldTopology=Star2025.11 | 0.7031 | 11.81 | — | |
| TASOModel=Qwen2.5 3B, #Param.=2.06 M, Evaluation Protocol=Zero-shot2025.09 | 0.7004 | — | — | |
| DGR SFTAlignment Dataset=STAR-1, Backbone=s1.1-7B2026.02 | 0.697 | — | — | |
| TWI-1-Generated PhotoCategory=Think with Image, Note=G-t-R2026.02 | 0.694 | — | — | |
| LoRA-XSModel=Qwen2.5 3B, #Param.=4.13 M, rank=128, Evaluation Protocol=Zero-shot2025.09 | 0.6909 | — | — | |
| VERAModel=Qwen2.5 3B, #Param.=1.42 M, rank=1024, Evaluation Protocol=Zero-shot2025.09 | 0.689 | — | — | |
| LoRAModel=Qwen2.5 3B, #Param.=14.97 M, rank=8, Evaluation Protocol=Zero-shot2025.09 | 0.6883 | — | — | |
| AdaLoRAModel=Qwen2.5 3B, #Param.=61.32 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 0.6872 | — | — | |
| DoRAModel=Qwen2.5 3B, #Param.=65.98 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 0.6853 | — | — | |
| LoRAModel=Qwen2.5 3B, #Param.=59.87 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 0.6834 | — | — | |
| Top-kModel=LLaMA-3.2 (3B), Zero-shot=true2026.01 | 0.6755 | — | — | |
| Fine-tuneModel=Qwen2.5 3B, #Param.=3151.91 M, Evaluation Protocol=Zero-shot2025.09 | 0.672 | — | — | |
| AdapterModel=Qwen2.5 3B, #Param.=67.10 M, Evaluation Protocol=Zero-shot2025.09 | 0.6644 | — | — | |
| Vanilla SFTAlignment Dataset=STAR-1, Backbone=s1.1-7B2026.02 | 0.663 | — | — | |
| AttackTopology=Complete2025.11 | 0.6508 | — | — | |
| IA3Model=Qwen2.5 3B, #Param.=1.35 M, Evaluation Protocol=Zero-shot2025.09 | 0.6464 | — | — | |
| Persona SwitchModel=Gemma-2 (2B), Zero-shot=true2026.01 | 0.6414 | — | — | |
| Role-Play PromptingModel=Gemma-2 (2B), Zero-shot=true2026.01 | 0.6361 | — | — | |
| Random SelectionModel=Gemma-2 (2B), Zero-shot=true2026.01 | 0.6325 | — | — | |
| GreedyModel=Gemma-2 (2B), Zero-shot=true2026.01 | 0.6308 | — | — | |
| Low-gap SelectionModel=Gemma-2 (2B), Zero-shot=true2026.01 | 0.6186 | — | — | |
| AttackTopology=Circle2025.11 | 0.6175 | — | — | |
| AttackTopology=Tree2025.11 | 0.6173 | — | — | |
| Top-pModel=Gemma-2 (2B), Zero-shot=true2026.01 | 0.6118 | — | — | |
| Qwen2.5-7B-Instruct2026.02 | 0.604 | — | — | |
| DGR SFTAlignment Dataset=R1-ACT, Backbone=s1.1-7B2026.02 | 0.6 | — | — | |
| Top-kModel=Gemma-2 (2B), Zero-shot=true2026.01 | 0.5959 | — | — | |
| MultinomialModel=Gemma-2 (2B), Zero-shot=true2026.01 | 0.5921 | — | — | |
| AttackTopology=Star2025.11 | 0.585 | — | — | |
| AttackTopology=Chain2025.11 | 0.5573 | — | — | |
| Fine-tuneModel=LLaMA3.2 3B, #Param.=3266.58 M, Evaluation Protocol=Zero-shot2025.09 | 0.4511 | — | — | |
| AdaLoRAModel=LLaMA3.2 3B, #Param.=49.51 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 0.4106 | — | — | |
| DoRAModel=LLaMA3.2 3B, #Param.=53.73 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 0.4 | — | — | |
| TASOModel=LLaMA3.2 3B, #Param.=1.67 M, Evaluation Protocol=Zero-shot2025.09 | 0.3981 | — | — |