LLM-as-Judge Response Evaluation on Overton Pluralistic 10 Perspectives
4.787HelpfulnessImplicit OP-GRPO
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Implicit OP-GRPOBackbone=Llama3.2-3B-Instruct, Judge=GPT-4.12026.02 | 4.787 | 4.91 | 4.937 | 4.757 | 3.877 | 4.653 | 4.725 | |
| Implicit OP-GRPOBackbone=Qwen2.5-3B-Instruct, Judge=GPT-4.12026.02 | 4.697 | 4.81 | 4.903 | 4.683 | 3.807 | 4.58 | 4.608 | |
| Qwen3-8B (Explicit OP)Judge=GPT-4.12026.02 | 4.69 | 4.43 | 4.893 | 4.37 | 3.55 | 4.387 | 4.38 | |
| Modular Pluralism (Qwen3-14B)Judge=GPT-4.12026.02 | 4.663 | 4.783 | 4.94 | 4.083 | 3.497 | 4.393 | 4.394 | |
| Implicit OP-GRPOBackbone=Qwen2.5-1.5B-Instruct, Judge=GPT-4.12026.02 | 4.567 | 4.71 | 4.89 | 4.5 | 3.733 | 4.48 | 4.47 | |
| Explicit PromptingBackbone=Qwen2.5-3B-Instruct, Judge=GPT-4.12026.02 | 4.18 | 4.757 | 4.823 | 3.95 | 3.42 | 4.226 | 4.204 | |
| Explicit PromptingBackbone=Qwen2.5-1.5B-Instruct, Judge=GPT-4.12026.02 | 4.1 | 4.07 | 4.643 | 3.817 | 3.133 | 3.953 | 3.863 | |
| GPT-OSS-20B (Explicit OP)Judge=GPT-4.12026.02 | 4.089 | 4.912 | 3.653 | 4.201 | 4.181 | 4.207 | 4.129 | |
| Explicit PromptingBackbone=Llama3.2-3B-Instruct, Judge=GPT-4.12026.02 | 4.047 | 4.487 | 4.783 | 3.817 | 3.333 | 4.093 | 4.103 |