Reasoning over Large Structured Context on Canvas
4.96ReasoningJudge ScoreGPT-5
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5Optimization=Standard JSON serialization2026.04 | 4.96 | 122.8 | 10.45 | |
| GPT-5 + HYVEOptimization=HYVE pipeline2026.04 | 4.96 | 38.2 | 8.99 | |
| GPT-4.1 + HYVEOptimization=HYVE pipeline2026.04 | 4.95 | 35.1 | 3 | |
| GPT-4.1Optimization=Standard JSON serialization2026.04 | 4.94 | 123.4 | 3.22 |