Multiple Choice Question Answering on CareGiving-QA 1.0 (test)
72.1Accuracy (Remember)GPT-4o-Mini
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GPT-4o-Miniinference_mode=Zero-shot2026.01 | 72.1 | 77.1 | 77 | 82.8 | 77.2 | |
| Llama-3-70Binference_mode=Zero-shot2026.01 | 70.8 | 73.5 | 70.4 | 83.3 | 74.5 | |
| GPT-4oinference_mode=Zero-shot2026.01 | 70.6 | 73.4 | 72.5 | 84 | 75.1 | |
| Qwen-2.5-72Binference_mode=Zero-shot2026.01 | 70.5 | 75.4 | 71.2 | 81.1 | 74.6 | |
| Kimi-K2inference_mode=Zero-shot2026.01 | 67.5 | 71.9 | 70.4 | 83.5 | 73.3 | |
| Qwen-2-72B-Instructinference_mode=Zero-shot2026.01 | 64.4 | 70.9 | 69.7 | 80 | 71.2 | |
| Mixtral-8x7Binference_mode=Zero-shot2026.01 | 63.2 | 70 | 69 | 71 | 68.2 | |
| DeepSeek-V3inference_mode=Zero-shot2026.01 | 60.4 | 72.3 | 72.7 | 82.5 | 72 |