LLM behavior monitoring on SHADE-Arena (test)
75.2AccuracyAgent-ToM (external memory, Stage 1–2)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Agent-ToM (external memory, Stage 1–2)Learning Strategy=external memory, Stage Configuration=Stage 1–2, Inference Model=GPT-5-mini2026.05 | 75.2 | 78.2 | 74.8 | 81.8 | |
| Sequential EnsembleMonitoring Architecture=Ensemble-based Monitoring, Inference Model=Claude 3.6 Sonnet2026.05 | 72.5 | 72.7 | 61.8 | 88.4 | |
| Agent-ToM (external memory, Stage 1)Learning Strategy=external memory, Stage Configuration=Stage 1, Inference Model=GPT-5-mini2026.05 | 69.3 | 71.7 | 65.5 | 79.2 | |
| Agent-ToM (Prompt-level learning)Learning Strategy=Prompt-level learning, Inference Model=GPT-5-mini2026.05 | 69 | 68.6 | 57 | 85.9 | |
| Agent-Auditor (nearest neighbor retrieval)Evaluation Protocol=nearest neighbor retrieval, Monitoring Architecture=Retrieval-Based Monitoring, Inference Model=GPT-5-mini2026.05 | 66.3 | 69.7 | 65.3 | 74.8 | |
| ToM (zeroshot)Evaluation Protocol=zeroshot, Monitoring Architecture=Structured Multi-Stage Monitoring, Inference Model=GPT-5-mini2026.05 | 65.9 | 61.5 | 45.7 | 93.7 | |
| LLM-Judge (zeroshot)Evaluation Protocol=zeroshot, Monitoring Architecture=Single-Pass Monitoring, Inference Model=GPT-5-mini2026.05 | 63.5 | 59.6 | 45.4 | 86.9 | |
| Async Ensemble (Majority Vote)Monitoring Architecture=Ensemble-based Monitoring, Aggregation Strategy=Majority Vote, Inference Model=GPT-5-mini2026.05 | 62.5 | 76.1 | 99 | 61.7 |