Theory of Mind reasoning on ToMI False Belief
98.2AccuracyMeTHanol
Evaluation Results
| Method | Links | |
|---|---|---|
| MeTHanolBase=Llama3-8B2024.09 | 98.2 | |
| COTBase=GPT-42024.09 | 95.5 | |
| SimTomBase=GPT-42024.09 | 95 | |
| directBase=GPT-42024.09 | 92.5 | |
| gpt-4Prompting strategy=SIMTOM2023.11 | 87.75 | |
| gpt-3.5-turboPrompting strategy=SIMTOM2023.11 | 81 | |
| gpt-4Prompting strategy=0-shot CoT2023.11 | 74.25 | |
| gpt-3.5-turboPrompting strategy=0-Shot2023.11 | 67.25 | |
| SFTBase=Llama3-8B2024.09 | 43.2 | |
| Llama2-7b-chatPrompting strategy=SIMTOM2023.11 | 40 | |
| Llama2-13b-chatPrompting strategy=0-Shot2023.11 | 39.25 | |
| Llama2-13b-chatPrompting strategy=SIMTOM2023.11 | 35.5 | |
| gpt-3.5-turboPrompting strategy=0-shot CoT2023.11 | 34 | |
| Llama2-7b-chatPrompting strategy=0-Shot2023.11 | 28.25 | |
| gpt-4Prompting strategy=0-Shot2023.11 | 25.5 | |
| Llama2-7b-chatPrompting strategy=0-shot CoT2023.11 | 24 | |
| directBase=Llama3-8B2024.09 | 22.2 | |
| Llama2-13b-chatPrompting strategy=0-shot CoT2023.11 | 16.5 |