Scene Classification on AID
98.33Top-1 AccGeoSolver
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GeoSolverModel Category=Ours2026.03 | 98.33 | — | — | |
| RS Agent2024.06 | 96.88 | — | — | |
| LWGANet L2Backbone Type=Hybrid, Params. (M)=13.0, FLOPs (G)=1.87, GPU Speed (FPS)=3308, CPU Speed (FPS)=16.18, ARM Speed (FPS)=274.3, Training image size=224x2242025.01 | 95.45 | — | — | |
| MobileViT SBackbone Type=Hybrid, Params. (M)=5.03, FLOPs (G)=1.75, GPU Speed (FPS)=2681, CPU Speed (FPS)=10.20, ARM Speed (FPS)=152.7, Training image size=256x2562025.01 | 95.25 | — | — | |
| MobileViT XSBackbone Type=Hybrid, Params. (M)=2.02, FLOPs (G)=0.900, GPU Speed (FPS)=3300, CPU Speed (FPS)=12.99, ARM Speed (FPS)=306.3, Training image size=256x2562025.01 | 95.2 | — | — | |
| LWGANet L1Backbone Type=Hybrid, Params. (M)=5.90, FLOPs (G)=0.709, GPU Speed (FPS)=6418, CPU Speed (FPS)=34.08, ARM Speed (FPS)=375.8, Training image size=224x2242025.01 | 94.85 | — | — | |
| LWGANet L0Backbone Type=Hybrid, Params. (M)=1.72, FLOPs (G)=0.186, GPU Speed (FPS)=13234, CPU Speed (FPS)=80.00, ARM Speed (FPS)=687.8, Training image size=224x2242025.01 | 94.6 | — | — | |
| MF-RSVLMLLM=Vicuna-1.5-7B2025.12 | 94.37 | — | — | |
| FUSE-RSVLMPublication=arXiv’252026.04 | 94.37 | — | — | |
| Efficientformer V2 S2Backbone Type=Transformer, Params. (M)=12.3, FLOPs (G)=1.26, GPU Speed (FPS)=642, CPU Speed (FPS)=24.74, ARM Speed (FPS)=123.8, Training image size=224x2242025.01 | 94.2 | — | — | |
| Efficientformer V2 S1Backbone Type=Transformer, Params. (M)=5.87, FLOPs (G)=0.661, GPU Speed (FPS)=1211, CPU Speed (FPS)=36.96, ARM Speed (FPS)=204.5, Training image size=224x2242025.01 | 93.95 | — | — | |
| MobileNet V2 2.0×Backbone Type=CNN, Params. (M)=8.81, FLOPs (G)=1.17, GPU Speed (FPS)=5567, CPU Speed (FPS)=17.03, ARM Speed (FPS)=372.3, Training image size=224x2242025.01 | 93.85 | — | — | |
| Efficientformer V2 S0Backbone Type=Transformer, Params. (M)=3.36, FLOPs (G)=0.396, GPU Speed (FPS)=1299, CPU Speed (FPS)=54.00, ARM Speed (FPS)=272.0, Training image size=224x2242025.01 | 93.8 | — | — | |
| MobileNet V2 1.0×Backbone Type=CNN, Params. (M)=2.28, FLOPs (G)=0.319, GPU Speed (FPS)=11301, CPU Speed (FPS)=49.11, ARM Speed (FPS)=785.4, Training image size=224x2242025.01 | 93.65 | — | — | |
| FasterNet T2Backbone Type=CNN, Params. (M)=13.8, FLOPs (G)=1.91, GPU Speed (FPS)=6852, CPU Speed (FPS)=26.40, ARM Speed (FPS)=669.8, Training image size=224x2242025.01 | 93.6 | — | — | |
| PVT V2 B1Backbone Type=Transformer, Params. (M)=13.5, FLOPs (G)=2.04, GPU Speed (FPS)=2369, CPU Speed (FPS)=15.96, ARM Speed (FPS)=145.3, Training image size=224x2242025.01 | 93.45 | — | — | |
| MobileViT XXSBackbone Type=Hybrid, Params. (M)=1.03, FLOPs (G)=0.333, GPU Speed (FPS)=4811, CPU Speed (FPS)=31.87, ARM Speed (FPS)=453.6, Training image size=256x2562025.01 | 93.3 | — | — | |
| FasterNet T1Backbone Type=CNN, Params. (M)=6.37, FLOPs (G)=0.855, GPU Speed (FPS)=11876, CPU Speed (FPS)=51.42, ARM Speed (FPS)=650.1, Training image size=224x2242025.01 | 93.2 | — | — | |
| EdgeViT XXSBackbone Type=Transformer, Params. (M)=3.79, FLOPs (G)=0.546, GPU Speed (FPS)=4153, CPU Speed (FPS)=46.36, ARM Speed (FPS)=259.9, Training image size=224x2242025.01 | 93.1 | — | — | |
| PVT V2 B0Backbone Type=Transformer, Params. (M)=3.42, FLOPs (G)=0.533, GPU Speed (FPS)=3843, CPU Speed (FPS)=40.26, ARM Speed (FPS)=243.4, Training image size=224x2242025.01 | 93.1 | — | — | |
| FasterNet T0Backbone Type=CNN, Params. (M)=2.68, FLOPs (G)=0.338, GPU Speed (FPS)=18276, CPU Speed (FPS)=106.4, ARM Speed (FPS)=839.7, Training image size=224x2242025.01 | 92.85 | — | — | |
| GeoZero2025.11 | 92.55 | — | — | |
| VHM2025.11 | 91.7 | — | — | |
| VHMLLM=Vicuna-1.5-7B2025.12 | 91.7 | — | — | |
| VHMPublication=AAAI’252026.04 | 91.7 | — | — | |
| RingMo-Agent2025.11 | 91.67 | — | — | |
| RemoteAgentPublication=-2026.04 | 91.34 | — | — | |
| LHRS-Bot2024.06 | 91.29 | — | — | |
| LHRS-BotLLM=LLaMA-2-7B2025.12 | 91.26 | — | — | |
| LHRS-BotPublication=ECCV’242026.04 | 91.26 | — | — | |
| TinyRS-R12025.11 | 90.2 | — | — | |
| EarthDialLLM=Phi-3-mini2025.12 | 87.57 | — | — | |
| EarthDialPublication=CVPR’252026.04 | 87.57 | — | — | |
| ScoreRS2025.11 | 85.9 | — | — | |
| GeoMagPublication=MM’252026.04 | 83.03 | — | — | |
| VHMModel Category=Open-source Remote Sensing Vision-Language Models2026.03 | 79 | — | — | |
| VLM2GeoVecEvaluation Protocol=Zero-shot, Prompting=20-prompt ensemble2025.12 | 77.25 | 2.3 | 1 | |
| Gemini-2.0-flashModel Category=Closed-source Commercial Vision-Language Models2026.03 | 76 | — | — | |
| RS-EoTModel Category=Open-source Reasoning Vision-Language Models2026.03 | 75.83 | — | — | |
| ChatGPT-5Model Category=Closed-source Commercial Vision-Language Models2026.03 | 75.5 | — | — | |
| SkySenseGPTModel Category=Open-source Remote Sensing Vision-Language Models2026.03 | 75.5 | — | — | |
| RemoteCLIPEvaluation Protocol=Zero-shot, Prompting=20-prompt ensemble2025.12 | 75.35 | 4.2 | 4 | |
| InternVL3.5LLM=InternVL3.52025.12 | 73.8 | — | — | |
| InternVL3.5Publication=arXiv’252026.04 | 73.8 | — | — | |
| GeoChatEvaluation Protocol=Zero-shot2025.12 | 73.55 | 4.3 | 5 | |
| GeoChatLLM=Vicuna-1.5-7B2025.12 | 73.17 | — | — | |
| GeoChatPublication=CVPR’242026.04 | 73.17 | — | — | |
| GeoRSCLIPEvaluation Protocol=Zero-shot, Prompting=20-prompt ensemble2025.12 | 72.85 | 3 | 3 | |
| GeoChat2025.11 | 72.03 | — | — | |
| GeoChat2024.06 | 72.03 | — | — | |
| SkyCLIPEvaluation Protocol=Zero-shot, Prompting=20-prompt ensemble2025.12 | 71.75 | 2.5 | 2 | |
| Qwen2.5-VLModel Category=Open-source Vision-Language Models2026.03 | 71.67 | — | — | |
| Qwen3-VL-8B-Instruct2025.11 | 71.4 | — | — | |
| Kimi-VL-ThinkingModel Category=Open-source Reasoning Vision-Language Models2026.03 | 70.5 | — | — | |
| CLIPEvaluation Protocol=Zero-shot, Prompting=20-prompt ensemble2025.12 | 70.1 | 4.8 | 6 | |
| GLM-4.1V-ThinkingModel Category=Open-source Reasoning Vision-Language Models2026.03 | 69.67 | — | — | |
| EarthDialModel Category=Open-source Remote Sensing Vision-Language Models2026.03 | 67.33 | — | — | |
| MiniCPM-V-2.6LLM=Qwen2-7B2025.12 | 65.07 | — | — | |
| VLM2VecEvaluation Protocol=Zero-shot, Prompting=20-prompt ensemble2025.12 | 64.25 | 6.5 | 7 | |
| InternVL2.5LLM=InternLM-2.52025.12 | 64 | — | — | |
| Qwen2.5-VLLLM=Qwen2.5-7B2025.12 | 63.07 | — | — | |
| Qwen2.5-VLPublication=arXiv’252026.04 | 63.07 | — | — | |
| Claude-sonnet-4Model Category=Closed-source Commercial Vision-Language Models2026.03 | 60.33 | — | — | |
| Phi-3.5-VisionLLM=Phi-3.52025.12 | 56.57 | — | — | |
| Phi3.5-VisionPublication=arXiv’242026.04 | 56.57 | — | — | |
| Qwen-VLLLM=Qwen-7B2025.12 | 55.3 | — | — | |
| MiniGPTv22024.06 | 52.6 | — | — | |
| Qwen-VL-Chat2024.06 | 52.6 | — | — | |
| InternLM-XComposeLLM=InternLM-7B2025.12 | 51.61 | — | — | |
| LLaVA-1.52024.06 | 51 | — | — | |
| mPLUG-OWL2LLM=LLaMA-2-7B2025.12 | 48.79 | — | — | |
| MiniGPTv2LLM=LLaMA-2-7B2025.12 | 32.96 | — | — | |
| LLaVA-1.5LLM=Vicuna-1.5-7B2025.12 | 31.1 | — | — | |
| InstructBLIPLLM=Vicuna-7B2025.12 | 29.5 | — | — | |
| MiniGPT-v2Model Category=Open-source Vision-Language Models2026.03 | 27.17 | — | — |