Reasoning on ARC-AGI public evaluation set V2
97.9AccuracyConfluence Lab
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Confluence LabStrategy=Code-execution methods, Model 1=—, Model 2=—, Code execution=True, N=4, K=22026.04 | 97.9 | 11.77 | — | |
| SQUEEZE EVOLVEStrategy=Full pipeline, Model 1=—, Model 2=Gemini 3.1 Pro, T (Iterations)=10, N=4, K=22026.04 | 97.5 | 7.74 | 3.7 | |
| SQUEEZE EVOLVEStrategy=Single recombination, Model 1=Gemini 3.0 Flash, Model 2=Gemini 3.1 Pro, T (Iterations)=2, N=4, K=22026.04 | 97.5 | 5.93 | 4.9 | |
| ImbueStrategy=Code-execution methods, Model 1=—, Model 2=Gemini 3.1 Pro, Code execution=True, N=4, K=22026.04 | 95.1 | 8.71 | — | |
| SQUEEZE EVOLVEStrategy=Single recombination, Model 1=—, Model 2=Gemini 3.1 Pro, T (Iterations)=2, N=4, K=22026.04 | 94.2 | 5.62 | 5.1 | |
| RSAStrategy=Full pipeline, Model 1=—, Model 2=Gemini 3.1 Pro, T (Iterations)=10, N=4, K=22026.04 | 93.3 | 28.85 | 1 |