General Performance on AlpacaEval
98WinrateVanilla
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VanillaModel=GPT-42023.11 | 98 | 39 | |
| Goal PrioritizationModel=GPT-42023.11 | 98 | 38.4 | |
| VanillaModel=ChatGPT2023.11 | 97 | 37.8 | |
| Goal PrioritizationModel=ChatGPT2023.11 | 96 | 36.3 | |
| Self-ReminderModel=GPT-42023.11 | 96 | 37.4 | |
| Self-ReminderModel=ChatGPT2023.11 | 95 | 35.2 | |
| VanillaModel=Vicuna-33B2023.11 | 95 | 36.6 | |
| Goal PrioritizationModel=Vicuna-33B2023.11 | 92 | 33.8 | |
| VanillaModel=Llama2-13B-Chat2023.11 | 91 | 33.8 | |
| VanillaModel=Llama2-7B-Chat2023.11 | 88 | 34.9 | |
| Self-ReminderModel=Vicuna-33B2023.11 | 86 | 33.3 | |
| VanillaModel=Vicuna-13B2023.11 | 84 | 32.6 | |
| Goal PrioritizationModel=Vicuna-13B2023.11 | 84 | 31.1 | |
| Goal PrioritizationModel=Llama2-13B-Chat2023.11 | 81 | 29.6 | |
| BaselineTarget Model=Zephyr-7B-Beta2025.05 | 78.35 | — | |
| VanillaModel=Vicuna-7B2023.11 | 78 | 30.9 | |
| MTSA-T3Target Model=Zephyr-7B-Beta, Iteration=32025.05 | 77.45 | — | |
| Self-ReminderModel=Vicuna-13B2023.11 | 76 | 29.3 | |
| Self-ReminderModel=Llama2-7B-Chat2023.11 | 75 | 29.8 | |
| Goal PrioritizationModel=Llama2-7B-Chat2023.11 | 74 | 28.8 | |
| Self-ReminderModel=Llama2-13B-Chat2023.11 | 74 | 29.9 | |
| Self-ReminderModel=Vicuna-7B2023.11 | 72 | 29.1 | |
| BaselineTarget Model=Llama2-7B-Chat2025.05 | 71.39 | — | |
| MTSA-T3Target Model=Llama2-7B-Chat, Iteration=32025.05 | 70.21 | — | |
| Goal PrioritizationModel=Vicuna-7B2023.11 | 68 | 27.5 |