Latent Preference Modeling on MPT Context-Free Average
58.5F1 ScorePREFINE
Evaluation Results
| Method | Links | |
|---|---|---|
| PREFINEBase LLM=Gemini-3-Flash2026.04 | 58.5 | |
| PREFINEBase LLM=GPT-5-mini2026.04 | 56.77 | |
| PREFINEBase LLM=GPT-52026.04 | 55.95 | |
| Gemini-3-FlashMethodology=Base Prompting2026.04 | 52.53 | |
| GPT-5-miniMethodology=Base Prompting2026.04 | 51.99 | |
| GPT-5Methodology=Base Prompting2026.04 | 49.83 | |
| PREFINEBase LLM=GPT-4o-mini2026.04 | 48.62 | |
| LangMemMethodology=Memory-Augmented, Backbone=Gemini-3-Flash2026.04 | 48.5 | |
| Mem0Methodology=Memory-Augmented, Backbone=Gemini-3-Flash2026.04 | 46.18 | |
| GPT-4o-miniMethodology=Base Prompting2026.04 | 45.26 | |
| RAG (Top-5)Methodology=Memory-Augmented, Backbone=Gemini-3-Flash2026.04 | 44.88 | |
| PREFINEBase LLM=Gemma-3-12B2026.04 | 44.01 | |
| PREFINEBase LLM=R1-Distill-Llama-8B2026.04 | 35.03 | |
| PREFINEBase LLM=CodeGemma-7B2026.04 | 33.56 | |
| Gemma-3-12BMethodology=Base Prompting2026.04 | 32.66 | |
| R1-Distill-Llama-8BMethodology=Base Prompting2026.04 | 30.96 | |
| PREFINEBase LLM=R1-Distill-Qwen-7B2026.04 | 30.65 | |
| CodeGemma-7BMethodology=Base Prompting2026.04 | 19.42 | |
| R1-Distill-Qwen-7BMethodology=Base Prompting2026.04 | 18.59 |