Political Consistency Evaluation on Polarized Contrastive Pairs
61.5Sentiment ConsistencyQwen3-14B + PCT
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-14B + PCTTraining Judge=Gemini 3.1 Pro, Evaluation Judge=GPT-5.5, Training Algorithm=GRPO, Adaptation=LoRA2026.05 | 61.5 | 95.1 | 78.3 | |
| Grok 4.1 FastEvaluation Judge=GPT-5.52026.05 | 47.4 | 87.6 | 67.5 | |
| Gemini 3.1 ProEvaluation Judge=GPT-5.52026.05 | 40.5 | 72.8 | 56.6 | |
| Claude Opus 4.7Evaluation Judge=GPT-5.52026.05 | 39.3 | 64.3 | 51.8 | |
| GPT-5.5Evaluation Judge=GPT-5.52026.05 | 38 | 76.3 | 57.2 | |
| DeepSeek V4 ProEvaluation Judge=GPT-5.52026.05 | 33.2 | 78.8 | 56 | |
| Mistral Medium 3.5Evaluation Judge=GPT-5.52026.05 | 31.1 | 82.9 | 57 | |
| Grok 4.3Evaluation Judge=GPT-5.52026.05 | 25.2 | 71.5 | 48.4 | |
| Qwen3-14BRole=Baseline, Evaluation Judge=GPT-5.52026.05 | 20.9 | 51.6 | 36.3 |