LLM Judge Policy Invariance Evaluation on Judge Card LLM Safety Judge Stress-test
70PISGPT-4o-mini
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT-4o-miniInterpretation=Moderate2026.05 | 70 | 1.1 | 99 | 18 | |
| Claude-HaikuInterpretation=Fragile2026.05 | 47 | 3.6 | 100 | 31 | |
| DeepSeek-V3.2Interpretation=Unreliable2026.05 | 28 | 3.5 | 100 | 43 | |
| Gemini-FlashInterpretation=Unreliable / Fragile2026.05 | 3 | 7.6 | 100 | 29 |