Theory of Mind Reasoning on COMMON-TOM 1.0 (test)
80Total AccuracyHuman Performance
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Human Performance2024.03 | 80 | 85 | 80 | 75 | |
| ReCoGBase model=FLAN-T52024.03 | 71 | 70.4 | 71.3 | 71.2 | |
| Mistral-7BEvaluation Protocol=Fine-Tune, Fine-tuning method=LoRA2024.03 | 64 | 64.8 | 63.9 | 63.2 | |
| gpt-4-0613Evaluation Protocol=Zero-Shot2024.03 | 63.4 | 65.5 | 62.5 | 62.1 | |
| Mistral-7B-InstructEvaluation Protocol=Zero-Shot2024.03 | 60.6 | 63.3 | 60.5 | 58 | |
| gpt-3.5-turbo-0613Evaluation Protocol=Zero-Shot2024.03 | 57 | 60.7 | 57.7 | 53 | |
| Random Baseline2024.03 | 50.4 | 50.3 | 50.5 | 50.4 |