Theory of Mind reasoning on BigTOM False Belief
99.4AccuracyMeTHanol
Evaluation Results
| Method | Links | |
|---|---|---|
| MeTHanolBase=Llama3-8B2024.09 | 99.4 | |
| gpt-4Prompting strategy=0-shot CoT2023.11 | 93.25 | |
| gpt-4Prompting strategy=SIMTOM2023.11 | 92 | |
| gpt-4Prompting strategy=0-Shot2023.11 | 89 | |
| SimTomBase=GPT-42024.09 | 87.8 | |
| SFTBase=Llama3-8B2024.09 | 77.7 | |
| COTBase=GPT-42024.09 | 74.4 | |
| directBase=Llama3-8B2024.09 | 71.3 | |
| Llama2-7b-chatPrompting strategy=SIMTOM2023.11 | 70.5 | |
| gpt-3.5-turboPrompting strategy=SIMTOM2023.11 | 70.5 | |
| directBase=GPT-42024.09 | 66.5 | |
| Llama2-13b-chatPrompting strategy=SIMTOM2023.11 | 61.75 | |
| gpt-3.5-turboPrompting strategy=0-shot CoT2023.11 | 56.25 | |
| Llama2-13b-chatPrompting strategy=0-shot CoT2023.11 | 52.25 | |
| Llama2-7b-chatPrompting strategy=0-Shot2023.11 | 47.5 | |
| Llama2-13b-chatPrompting strategy=0-Shot2023.11 | 41.25 | |
| gpt-3.5-turboPrompting strategy=0-Shot2023.11 | 41 | |
| Llama2-7b-chatPrompting strategy=0-shot CoT2023.11 | 31.5 |