Earth Observation Agent Performance on OpenEarth-Bench
85.4Data Preparation AccuracyOpenEarth-Agent
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| OpenEarth-AgentLLM Backbone=GPT-5, Evaluation Protocol=Stage-Wise2026.03 | 85.4 | 85.27 | 76.66 | |
| OpenEarth-AgentLLM Backbone=GPT-5, Evaluation Protocol=End-to-End2026.03 | 82.38 | 74.83 | 58.72 | |
| OpenEarth-AgentLLM Backbone=Seed-1.6, Evaluation Protocol=Stage-Wise2026.03 | 79.36 | 82.38 | 71.81 | |
| OpenEarth-AgentLLM Backbone=Kimik2, Evaluation Protocol=Stage-Wise2026.03 | 77.85 | 79.19 | 66.11 | |
| OpenEarth-AgentLLM Backbone=Seed-1.6, Evaluation Protocol=End-to-End2026.03 | 77.18 | 68.12 | 51.84 | |
| OpenEarth-AgentLLM Backbone=Gemini-2.5-Flash, Evaluation Protocol=Stage-Wise2026.03 | 76.17 | 77.52 | 64.76 | |
| OpenEarth-AgentLLM Backbone=Kimik2, Evaluation Protocol=End-to-End2026.03 | 75.84 | 64.09 | 47.81 | |
| OpenEarth-AgentLLM Backbone=DeepSeek-V3.1, Evaluation Protocol=Stage-Wise2026.03 | 75.17 | 78.02 | 63.08 | |
| OpenEarth-AgentLLM Backbone=Gemini-2.5-Flash, Evaluation Protocol=End-to-End2026.03 | 74.16 | 61.58 | 45.47 | |
| OpenEarth-AgentLLM Backbone=DeepSeek-V3.1, Evaluation Protocol=End-to-End2026.03 | 72.82 | 60.74 | 43.12 | |
| OpenEarth-AgentLLM Backbone=Qwen3-Max, Evaluation Protocol=Stage-Wise2026.03 | 71.31 | 73.82 | 60.57 | |
| OpenEarth-AgentLLM Backbone=Qwen3-Max, Evaluation Protocol=End-to-End2026.03 | 69.79 | 56.38 | 39.26 |