Reasoning on DetectBench
70.3AccuracyGemini2.5-Flash
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini2.5-FlashPrompt Instruction=CoQ Prompt2025.10 | 70.3 | 601.4 | |
| Gemini2.5-FlashPrompt Instruction=CoT Prompt2025.10 | 68.6 | 430.1 | |
| Gemini2.5-FlashPrompt Instruction=Base2025.10 | 64.1 | 285.4 | |
| Qwen2.5-7BPrompt Instruction=Curious CoQ2025.10 | 64 | 632.5 | |
| Gemma3-12BPrompt Instruction=Curious CoQ2025.10 | 61.9 | 715.8 | |
| Llama3-8BPrompt Instruction=Curious CoQ2025.10 | 58.5 | 618.3 | |
| Qwen2.5-7BPrompt Instruction=Refined CoT2025.10 | 56.4 | 554.6 | |
| Gemma3-12BPrompt Instruction=Vanilla CoT2025.10 | 56.2 | 548.7 | |
| Gemma3-12BPrompt Instruction=Refined CoT2025.10 | 55.8 | 583.2 | |
| Qwen2.5-7BPrompt Instruction=Vanilla CoT2025.10 | 54.5 | 511.2 | |
| Llama3-8BPrompt Instruction=Refined CoT2025.10 | 52.7 | 489.5 | |
| Gemma3-12BPrompt Instruction=Base2025.10 | 48.7 | 324.6 | |
| Llama3-8BPrompt Instruction=Vanilla CoT2025.10 | 48.3 | 387.2 | |
| GPT-4o-miniPrompt Instruction=CoQ Prompt2025.10 | 45.2 | 475.9 | |
| GPT-4o-miniPrompt Instruction=CoT Prompt2025.10 | 42.8 | 242.5 | |
| Qwen2.5-7BPrompt Instruction=Base2025.10 | 40.3 | 267.4 | |
| GPT-4o-miniPrompt Instruction=Base2025.10 | 38.2 | 156.3 | |
| Llama3-8BPrompt Instruction=Base2025.10 | 25 | 114.8 |