Preference-driven Tool Calling on MPT Context-Guided, Preference Recall
64.88P-EMPREFINE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PREFINEBase LLM=Gemini-3-Flash2026.04 | 64.88 | 72.76 | 74.75 | |
| LangMemMethodology=Memory-Augmented, Backbone=Gemini-3-Flash2026.04 | 64.4 | 64.54 | 67.83 | |
| Gemini-3-FlashMethodology=Base Prompting2026.04 | 62.65 | 72.73 | 74.25 | |
| PREFINEBase LLM=CodeGemma-7B2026.04 | 59.64 | 69.5 | 70.51 | |
| PREFINEBase LLM=GPT-52026.04 | 52.41 | 66.74 | 67.85 | |
| PREFINEBase LLM=GPT-5-mini2026.04 | 51.45 | 68.03 | 68.08 | |
| GPT-5Methodology=Base Prompting2026.04 | 51.2 | 62.33 | 64.77 | |
| RAG (Top-5)Methodology=Memory-Augmented, Backbone=Gemini-3-Flash2026.04 | 50.6 | 69.14 | 67.99 | |
| PREFINEBase LLM=GPT-4o-mini2026.04 | 49.88 | 72.65 | 68.71 | |
| GPT-5-miniMethodology=Base Prompting2026.04 | 47.59 | 65.38 | 66.69 | |
| PREFINEBase LLM=R1-Distill-Llama-8B2026.04 | 42.05 | 62.35 | 61.63 | |
| R1-Distill-Llama-8BMethodology=Base Prompting2026.04 | 34.94 | 65.12 | 61.03 | |
| GPT-4o-miniMethodology=Base Prompting2026.04 | 32.23 | 58.21 | 53.54 | |
| PREFINEBase LLM=R1-Distill-Qwen-7B2026.04 | 32.17 | 59.05 | 54.6 | |
| Mem0Methodology=Memory-Augmented, Backbone=Gemini-3-Flash2026.04 | 31.93 | 64.59 | 59.79 | |
| PREFINEBase LLM=Gemma-3-12B2026.04 | 20.48 | 79.28 | 69.27 | |
| CodeGemma-7BMethodology=Base Prompting2026.04 | 18.67 | 38.88 | 38.17 | |
| R1-Distill-Qwen-7BMethodology=Base Prompting2026.04 | 13.55 | 33.49 | 31.58 | |
| Gemma-3-12BMethodology=Base Prompting2026.04 | 7.23 | 60.36 | 49.49 |