Reasoning over Large Structured Context on Anom
4.03ReasoningJudge ScoreGPT-4.1 + HYVE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4.1 + HYVEOptimization=HYVE pipeline2026.04 | 4.03 | 8.9 | 4.38 | |
| GPT-5 + HYVEOptimization=HYVE pipeline2026.04 | 3.77 | 10.9 | 20.71 | |
| GPT-5Optimization=Standard JSON serialization2026.04 | 3.28 | 18.5 | 28.73 | |
| GPT-4.1Optimization=Standard JSON serialization2026.04 | 3.22 | 15.8 | 7.49 |