Long-Context Reasoning on BrowsCompLong
88.07AccuracyGemini-3.0-pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-3.0-pro2026.03 | 88.07 | |
| Deepseek-R1-Distill-Qwen-14B + TableLong2026.03 | 74.31 | |
| Deepseek-R1-Distill-Qwen-32B + TableLong2026.03 | 74.31 | |
| Qwen-Long-L12026.03 | 69.93 | |
| Qwen3-32B + TableLong2026.03 | 65.44 | |
| Deepseek-R1-Distill-Qwen-32B2026.03 | 64.22 | |
| Qwen2.5-32B-Instruct + TableLong2026.03 | 60.55 | |
| Qwen3-32B2026.03 | 59.33 | |
| Deepseek-v3.12026.03 | 56.27 | |
| Qwen2.5-32B-Instruct2026.03 | 52.56 | |
| Deepseek-R1-Distill-Qwen-14B2026.03 | 51.38 |