Discriminator-based DCS on Board Game Problems
51.5DCSSoft Thinking (no noise)
Evaluation Results
| Method | Links | |
|---|---|---|
| Soft Thinking (no noise)LLM=Llama-3.1 8B, Sequence Length=32, threshold=τ=0.92026.05 | 51.5 | |
| Output Emb.LLM=Llama-3.3 70B, Aggregation=Exc, threshold=τ=0.92026.05 | 51.4 | |
| Soft Thinking (no noise)LLM=Llama-3.3 70B, Sequence Length=1, threshold=τ=0.92026.05 | 50.9 | |
| Soft Thinking (no noise)LLM=Llama-3.1 8B, Sequence Length=64, threshold=τ=0.92026.05 | 50.6 | |
| Output Emb.LLM=Llama-3.3 70B, Aggregation=Pool, threshold=τ=0.92026.05 | 50.6 | |
| Soft Thinking (Gumbel)LLM=Llama-3.3 70B, Sequence Length=1, threshold=τ=0.92026.05 | 50.6 | |
| Last Input Tok.LLM=Llama-3.1 8B, Aggregation=All, threshold=τ=0.92026.05 | 50.5 | |
| Latent ThinkingLLM=Llama-3.1 8B, Sequence Length=1, threshold=τ=0.92026.05 | 50.5 | |
| Soft Thinking (no noise)LLM=Llama-3.1 8B, Sequence Length=16, threshold=τ=0.92026.05 | 50.4 | |
| Soft Thinking (Gumbel)LLM=Llama-3.1 8B, Sequence Length=32, threshold=τ=0.92026.05 | 50.4 | |
| Soft Thinking (Gumbel)LLM=Llama-3.1 8B, Sequence Length=16, threshold=τ=0.92026.05 | 50.3 | |
| Soft Thinking (Gumbel)LLM=Llama-3.3 70B, Sequence Length=16, threshold=τ=0.92026.05 | 50.3 | |
| Soft Thinking (no noise)LLM=Llama-3.1 8B, Sequence Length=1, threshold=τ=0.92026.05 | 50.2 | |
| Soft Thinking (no noise)LLM=Llama-3.1 8B, Sequence Length=128, threshold=τ=0.92026.05 | 50.2 | |
| Soft Thinking (Gumbel)LLM=Llama-3.1 8B, Sequence Length=1, threshold=τ=0.92026.05 | 50.2 | |
| Soft Thinking (no noise)LLM=Llama-3.3 70B, Sequence Length=16, threshold=τ=0.92026.05 | 50.2 | |
| Last Input Tok.LLM=DS-R1-Qwen 32B, Aggregation=All, threshold=τ=0.92026.05 | 50.2 | |
| Soft Thinking (Gumbel)LLM=Llama-3.1 8B, Sequence Length=64, threshold=τ=0.92026.05 | 50 | |
| Last Input Tok.LLM=Llama-3.3 70B, Aggregation=All, threshold=τ=0.92026.05 | 50 | |
| Last Input Tok.LLM=Llama-3.1 8B, Aggregation=Final, threshold=τ=0.92026.05 | 49.9 | |
| Soft Thinking (Gumbel)LLM=Llama-3.1 8B, Sequence Length=128, threshold=τ=0.92026.05 | 49.9 | |
| Latent ThinkingLLM=Llama-3.1 8B, Sequence Length=16, threshold=τ=0.92026.05 | 49.9 | |
| RVLLM=Llama-3.1 8B, Type=Baseline, threshold=τ=0.92026.05 | 49.9 | |
| Soft Thinking (no noise)LLM=Llama-3.3 70B, Sequence Length=32, threshold=τ=0.92026.05 | 49.9 | |
| IELLM=Llama-3.3 70B, Type=Baseline, threshold=τ=0.92026.05 | 49.9 | |
| Soft Thinking (Gumbel)LLM=DS-R1-Qwen 32B, Sequence Length=1, threshold=τ=0.92026.05 | 49.9 | |
| Latent ThinkingLLM=Llama-3.1 8B, Sequence Length=64, threshold=τ=0.92026.05 | 49.8 | |
| IELLM=Llama-3.1 8B, Type=Baseline, threshold=τ=0.92026.05 | 49.8 | |
| Last Input Tok.LLM=Llama-3.3 70B, Aggregation=Final, threshold=τ=0.92026.05 | 49.8 | |
| Soft Thinking (no noise)LLM=Llama-3.3 70B, Sequence Length=128, threshold=τ=0.92026.05 | 49.8 | |
| Soft Thinking (Gumbel)LLM=Llama-3.3 70B, Sequence Length=64, threshold=τ=0.92026.05 | 49.8 | |
| Latent ThinkingLLM=Llama-3.3 70B, Sequence Length=1, threshold=τ=0.92026.05 | 49.8 | |
| Latent ThinkingLLM=Llama-3.3 70B, Sequence Length=16, threshold=τ=0.92026.05 | 49.8 | |
| Latent ThinkingLLM=Llama-3.3 70B, Sequence Length=32, threshold=τ=0.92026.05 | 49.8 | |
| Latent ThinkingLLM=Llama-3.3 70B, Sequence Length=64, threshold=τ=0.92026.05 | 49.8 | |
| Latent ThinkingLLM=Llama-3.3 70B, Sequence Length=128, threshold=τ=0.92026.05 | 49.8 | |
| RVLLM=Llama-3.3 70B, Type=Baseline, threshold=τ=0.92026.05 | 49.8 | |
| Soft Thinking (no noise)LLM=DS-R1-Qwen 32B, Sequence Length=1, threshold=τ=0.92026.05 | 49.8 | |
| Soft Thinking (Gumbel)LLM=DS-R1-Qwen 32B, Sequence Length=16, threshold=τ=0.92026.05 | 49.8 | |
| Latent ThinkingLLM=Llama-3.1 8B, Sequence Length=128, threshold=τ=0.92026.05 | 49.7 | |
| Soft Thinking (no noise)LLM=Llama-3.3 70B, Sequence Length=64, threshold=τ=0.92026.05 | 49.7 | |
| Soft Thinking (Gumbel)LLM=Llama-3.3 70B, Sequence Length=128, threshold=τ=0.92026.05 | 49.7 | |
| Latent ThinkingLLM=DS-R1-Qwen 32B, Sequence Length=16, threshold=τ=0.92026.05 | 49.7 | |
| Latent ThinkingLLM=DS-R1-Qwen 32B, Sequence Length=128, threshold=τ=0.92026.05 | 49.7 | |
| RVLLM=DS-R1-Qwen 32B, Type=Baseline, threshold=τ=0.92026.05 | 49.7 | |
| Last Input Tok.LLM=DS-R1-Qwen 32B, Aggregation=Final, threshold=τ=0.92026.05 | 49.6 | |
| Soft Thinking (no noise)LLM=DS-R1-Qwen 32B, Sequence Length=128, threshold=τ=0.92026.05 | 49.6 | |
| Latent ThinkingLLM=DS-R1-Qwen 32B, Sequence Length=1, threshold=τ=0.92026.05 | 49.6 | |
| Latent ThinkingLLM=DS-R1-Qwen 32B, Sequence Length=64, threshold=τ=0.92026.05 | 49.6 | |
| Latent ThinkingLLM=Llama-3.1 8B, Sequence Length=32, threshold=τ=0.92026.05 | 49.5 | |
| Soft Thinking (Gumbel)LLM=Llama-3.3 70B, Sequence Length=32, threshold=τ=0.92026.05 | 49.5 | |
| Soft Thinking (Gumbel)LLM=DS-R1-Qwen 32B, Sequence Length=64, threshold=τ=0.92026.05 | 49.5 | |
| Soft Thinking (Gumbel)LLM=DS-R1-Qwen 32B, Sequence Length=128, threshold=τ=0.92026.05 | 49.5 | |
| Soft Thinking (no noise)LLM=DS-R1-Qwen 32B, Sequence Length=64, threshold=τ=0.92026.05 | 49.4 | |
| Soft Thinking (Gumbel)LLM=DS-R1-Qwen 32B, Sequence Length=32, threshold=τ=0.92026.05 | 49.3 | |
| Soft Thinking (no noise)LLM=DS-R1-Qwen 32B, Sequence Length=16, threshold=τ=0.92026.05 | 49.2 | |
| Latent ThinkingLLM=DS-R1-Qwen 32B, Sequence Length=32, threshold=τ=0.92026.05 | 49.1 | |
| Soft Thinking (no noise)LLM=DS-R1-Qwen 32B, Sequence Length=32, threshold=τ=0.92026.05 | 48.2 | |
| IELLM=DS-R1-Qwen 32B, Type=Baseline, threshold=τ=0.92026.05 | 47.9 | |
| Output Emb.LLM=Llama-3.1 8B, Aggregation=Pool, threshold=τ=0.92026.05 | 43.6 | |
| Output Emb.LLM=DS-R1-Qwen 32B, Aggregation=Pool, threshold=τ=0.92026.05 | 40.4 | |
| Output Emb.LLM=DS-R1-Qwen 32B, Aggregation=Exc, threshold=τ=0.92026.05 | 39.9 | |
| Output Emb.LLM=Llama-3.1 8B, Aggregation=Exc, threshold=τ=0.92026.05 | 39.3 |