Preference-driven Tool Calling on MPT Context-Guided Average
67.18OA-F1PREFINE
Evaluation Results
| Method | Links | |
|---|---|---|
| PREFINEBase LLM=Gemini-3-Flash2026.04 | 67.18 | |
| Gemini-3-FlashMethodology=Base Prompting2026.04 | 65.76 | |
| PREFINEBase LLM=Gemma-3-12B2026.04 | 65.35 | |
| PREFINEBase LLM=GPT-52026.04 | 63.98 | |
| PREFINEBase LLM=GPT-5-mini2026.04 | 63.9 | |
| PREFINEBase LLM=GPT-4o-mini2026.04 | 63.58 | |
| PREFINEBase LLM=CodeGemma-7B2026.04 | 61.83 | |
| RAG (Top-5)Methodology=Memory-Augmented, Backbone=Gemini-3-Flash2026.04 | 61.74 | |
| GPT-5Methodology=Base Prompting2026.04 | 61.42 | |
| GPT-5-miniMethodology=Base Prompting2026.04 | 60.24 | |
| LangMemMethodology=Memory-Augmented, Backbone=Gemini-3-Flash2026.04 | 59.4 | |
| Mem0Methodology=Memory-Augmented, Backbone=Gemini-3-Flash2026.04 | 58.9 | |
| R1-Distill-Llama-8BMethodology=Base Prompting2026.04 | 56.21 | |
| PREFINEBase LLM=R1-Distill-Llama-8B2026.04 | 54.27 | |
| GPT-4o-miniMethodology=Base Prompting2026.04 | 53.27 | |
| PREFINEBase LLM=R1-Distill-Qwen-7B2026.04 | 47.87 | |
| Gemma-3-12BMethodology=Base Prompting2026.04 | 46.95 | |
| CodeGemma-7BMethodology=Base Prompting2026.04 | 32.63 | |
| R1-Distill-Qwen-7BMethodology=Base Prompting2026.04 | 25.73 |