LLM-as-Judge Response Evaluation on Overton Pluralistic 5 Perspectives
4.723HelpfulnessImplicit OP-GRPO
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Implicit OP-GRPOBackbone=Llama3.2-3B-Instruct, Judge=GPT-4.12026.02 | 4.723 | 4.917 | 4.83 | 4.773 | 4.74 | 4.797 | 4.725 | |
| Implicit OP-GRPOBackbone=Qwen2.5-3B-Instruct, Judge=GPT-4.12026.02 | 4.693 | 4.77 | 4.93 | 4.83 | 3.957 | 4.636 | 4.608 | |
| Qwen3-8B (Explicit OP)Judge=GPT-4.12026.02 | 4.607 | 4.337 | 4.83 | 4.467 | 3.62 | 4.372 | 4.38 | |
| Modular Pluralism (Qwen3-14B)Judge=GPT-4.12026.02 | 4.58 | 4.62 | 4.89 | 4.313 | 3.567 | 4.394 | 4.394 | |
| Implicit OP-GRPOBackbone=Qwen2.5-1.5B-Instruct, Judge=GPT-4.12026.02 | 4.48 | 4.647 | 4.873 | 4.52 | 3.78 | 4.46 | 4.47 | |
| GPT-OSS-20B (Explicit OP)Judge=GPT-4.12026.02 | 4.123 | 3.88 | 4.915 | 3.247 | 4.089 | 4.051 | 4.129 | |
| Explicit PromptingBackbone=Qwen2.5-3B-Instruct, Judge=GPT-4.12026.02 | 3.913 | 4.59 | 4.793 | 4.283 | 3.333 | 4.183 | 4.204 | |
| Explicit PromptingBackbone=Llama3.2-3B-Instruct, Judge=GPT-4.12026.02 | 3.87 | 4.37 | 4.673 | 4.21 | 3.443 | 4.113 | 4.103 | |
| Explicit PromptingBackbone=Qwen2.5-1.5B-Instruct, Judge=GPT-4.12026.02 | 3.82 | 3.81 | 4.43 | 3.855 | 2.945 | 3.772 | 3.863 |