Long-horizon Agent Performance on CMS Medicare dataset 15-turn
75Score (out of 300)ContextForge
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| ContextForgeLLM Backbone=GPT-5.4, Number of cycles=2, Turn Count=152026.05 | 75 | 7.6 | 25,666 | 93.3333 | 8 | 13.4 | |
| Fabric Agt.LLM Backbone=GPT-5.4, Number of cycles=2, Turn Count=15, Component=Azure AI Foundry agent with Fabric Data Agent tool2026.05 | 64.6667 | 60.5 | 345,112 | 90 | — | — |