Word Sorting on BIG-bench Hard (test)
45Test AccuracyBest Local f-Ens.
Evaluation Results
| Method | Links | |
|---|---|---|
| Best Local f-Ens.Base Model=Llama2026.03 | 45 | |
| Best Global f-Ens.Base Model=Llama, SMC Type=Token SMC, M=252026.03 | 42.7 | |
| Best Global f-Ens.Base Model=Llama, SMC Type=Token SMC, M=102026.03 | 42.3 | |
| Local Prob. Avg.Base Model=Llama2026.03 | 41.2 | |
| Best Cross-Model f-Ens.Base Models=P+L, SMC Type=Byte SMC, M=10, Prompt=Prompt 12026.03 | 41.1 | |
| Prompt 1Base Model=Llama2026.03 | 38.6 | |
| Prompt 2Base Model=Llama2026.03 | 36.5 | |
| Best Global f-Ens.Base Model=Llama, SMC Type=Byte SMC, M=102026.03 | 34 | |
| Best Global f-Ens.Base Model=Phi, SMC Type=Token SMC, M=252026.03 | 28.6 | |
| Best Global f-Ens.Base Model=Phi, SMC Type=Token SMC, M=102026.03 | 28.2 | |
| PE2Final Prompt=Sort the given words alphabetically, but exclude 'List:' or similar formatting elements. Ensure every word is considered., Task Model=Mistral-7B-Instruct-v0.2, Prompt Proposal Model=gpt-4-turbo2023.11 | 28 | |
| Prompt 1Base Model=Phi2026.03 | 27.7 | |
| Best Local f-Ens.Base Model=Phi2026.03 | 27.2 | |
| Local Prob. Avg.Base Model=Phi2026.03 | 25.8 | |
| Prompt 2Base Model=Phi2026.03 | 24.3 | |
| Best Global f-Ens.Base Model=Phi, SMC Type=Byte SMC, M=102026.03 | 23.8 | |
| Iterative APEFinal Prompt=Let's tackle this systematically, advancing step by step., Task Model=Mistral-7B-Instruct-v0.2, Prompt Proposal Model=gpt-4-turbo2023.11 | 20 | |
| APOFinal Prompt=Alphabetically sort the words below, correcting any typos. Ignore capitalization, treat abbreviations and possessives normally. Exclude 'List:' from items., Task Model=Mistral-7B-Instruct-v0.2, Prompt Proposal Model=gpt-4-turbo2023.11 | 16 | |
| Best Global f-Ens.Base Model=Qwen, SMC Type=Byte SMC, M=102026.03 | 15.5 | |
| Best Global f-Ens.Base Model=Qwen, SMC Type=Token SMC, M=102026.03 | 15.3 | |
| Best Local f-Ens.Base Model=Qwen2026.03 | 15 | |
| Best Global f-Ens.Base Model=Qwen, SMC Type=Token SMC, M=252026.03 | 14.9 | |
| Prompt 1Base Model=Qwen2026.03 | 13.9 | |
| Local Prob. Avg.Base Model=Qwen2026.03 | 9.6 | |
| Prompt 2Base Model=Qwen2026.03 | 5.8 | |
| Zero-shot CoTFinal Prompt=Let's think step by step., Task Model=Mistral-7B-Instruct-v0.2, Prompt Proposal Model=gpt-4-turbo2023.11 | 4 |