Preference Prediction on PRISM (test)
66.62AccuracyEXACT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| EXACTBase Model=Llama-3.1-8B2026.02 | 66.62 | 0.98 | |
| EXACTBase Model=Qwen2.5-7B-Instruct2026.02 | 66.41 | 0.64 | |
| EXACTBase Model=Gemma-2-9B-it2026.02 | 65.93 | 0.15 | |
| Finetuned Reward ModelEvaluation Protocol=Bradley-Terry Reward Model, Backbone=Llama 3.1 8B2025.06 | 64.29 | — | |
| SynthesizeMe (+ FT RM + Personas)Evaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.3 70B2025.06 | 64.03 | — | |
| Finetuned Reward ModelEvaluation Protocol=Bradley-Terry Reward Model, Backbone=Llama 3.3 70B2025.06 | 63.5 | — | |
| SynthesizeMe (+ FT RM + Personas + Demos)Evaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.3 70B2025.06 | 63.44 | — | |
| SynthesizeMe (+ FT RM + Personas)Evaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.1 8B2025.06 | 63.11 | — | |
| SynthesizeMe (+ FT RM + Personas + Demos)Evaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.1 8B2025.06 | 62.74 | — | |
| Finetuned Reward ModelEvaluation Protocol=Bradley-Terry Reward Model, Backbone=Llama 3.2 3B2025.06 | 61.66 | — | |
| SynthesizeMe (+ FT RM + Personas)Evaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.2 3B2025.06 | 61.53 | — | |
| DriftBase Model=Gemma-2-9B-it2026.02 | 61.31 | 0.54 | |
| SynthesizeMe (+ FT RM + Personas + Demos)Evaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.2 3B2025.06 | 61.24 | — | |
| BaseBase Model=Gemma-2-9B-it2026.02 | 59.27 | 0.16 | |
| DriftBase Model=Llama-3.1-8B2026.02 | 58.96 | 0.47 | |
| RewardBase Model=Gemma-2-9B-it2026.02 | 58.91 | 0.23 | |
| VPLEvaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.2 3B2025.06 | 58.26 | — | |
| VPLEvaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.1 8B2025.06 | 58.23 | — | |
| SynthesizeMe (Just Demos)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.3 70B2025.06 | 57.76 | — | |
| BaseBase Model=Llama-3.1-8B2026.02 | 57.22 | 0.21 | |
| SynthesizeMe (Personas + Demos)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.3 70B2025.06 | 56.99 | — | |
| PALEvaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.2 3B2025.06 | 56.81 | — | |
| GPOEvaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.1 8B2025.06 | 56.48 | — | |
| DriftBase Model=Qwen2.5-7B-Instruct2026.02 | 56.16 | 0.28 | |
| GPOEvaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.3 70B2025.06 | 55.65 | — | |
| GPOEvaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.2 3B2025.06 | 55.26 | — | |
| SynthesizeMe (Personas + Demos + Distill)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.1 8B2025.06 | 55.24 | — | |
| DemographicsEvaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.2 3B2025.06 | 54.95 | — | |
| SynthesizeMe (Personas + Distill)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.1 8B2025.06 | 54.95 | — | |
| SynthesizeMe (Just Demos)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.1 8B2025.06 | 54.93 | — | |
| BaseBase Model=Qwen2.5-7B-Instruct2026.02 | 54.83 | 0.11 | |
| RewardBase Model=Llama-3.1-8B2026.02 | 54.78 | 0.11 | |
| DefaultEvaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.3 70B2025.06 | 54.35 | — | |
| PALEvaluation Protocol=Finetuned Reward Model, Backbone=Llama 3.1 8B2025.06 | 54.23 | — | |
| MemoryEvaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.3 70B2025.06 | 54.2 | — | |
| MemoryEvaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.1 8B2025.06 | 54.17 | — | |
| DemographicsEvaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.1 8B2025.06 | 54.06 | — | |
| DemographicsEvaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.3 70B2025.06 | 53.89 | — | |
| SynthesizeMe (Just Personas)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.3 70B2025.06 | 53.84 | — | |
| SynthesizeMe (Just Personas)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.1 8B2025.06 | 53.66 | — | |
| SynthesizeMe (Personas + Demos)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.1 8B2025.06 | 53.3 | — | |
| DefaultEvaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.1 8B2025.06 | 52.8 | — | |
| SynthesizeMe (Personas + Distill)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.2 3B2025.06 | 52.21 | — | |
| SynthesizeMe (Personas + Demos + Distill)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.2 3B2025.06 | 52.09 | — | |
| SynthesizeMe (Just Demos)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.2 3B2025.06 | 51.7 | — | |
| DefaultEvaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.2 3B2025.06 | 51.65 | — | |
| SynthesizeMe (Personas + Demos)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.2 3B2025.06 | 51.52 | — | |
| SynthesizeMe (Just Personas)Evaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.2 3B2025.06 | 51.12 | — | |
| MemoryEvaluation Protocol=In-Context LLM as a Judge, Backbone=Llama 3.2 3B2025.06 | 50.86 | — | |
| RandomEvaluation Protocol=Baseline2025.06 | 50 | — | |
| RewardBase Model=Qwen2.5-7B-Instruct2026.02 | 49.58 | 1.02 |