Binary Classification on Resistance behavior dataset 5-fold CV (test)
91.31Precision (Overall)PsyFIRE
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| PsyFIREBackbone=Llama-3.1-8B-Instruct, Strategy=Fine-Tuning (FPFT), Rationale=Included2026.01 | 91.31 | 91.21 | 91.25 | 91.41 | |
| GPT-4oPrompting Strategy=Few-Shot2026.01 | 84.34 | 79.96 | 80.65 | 81.91 | |
| Claude-3.5-SonnetPrompting Strategy=Zero-Shot2026.01 | 84.21 | 79.05 | 79.75 | 81.21 | |
| Claude-3.5-SonnetPrompting Strategy=Few-Shot2026.01 | 83.18 | 80.96 | 81.5 | 82.31 | |
| GPT-4oPrompting Strategy=Zero-Shot2026.01 | 82.56 | 78.28 | 78.91 | 80.3 | |
| Qwen2.5-7B-InstructPrompting Strategy=Few-Shot2026.01 | 79.67 | 70.54 | 70.28 | 73.88 | |
| Qwen2.5-7B-InstructPrompting Strategy=Zero-Shot2026.01 | 79.1 | 69.75 | 69.34 | 73.19 | |
| Qwen2.5-14B-InstructPrompting Strategy=Zero-Shot2026.01 | 78.47 | 72.74 | 72.58 | 74.91 | |
| Qwen2.5-32B-InstructPrompting Strategy=Zero-Shot2026.01 | 78.14 | 73.5 | 73.37 | 75.28 | |
| Qwen2.5-14B-InstructPrompting Strategy=Few-Shot2026.01 | 78.03 | 73.11 | 73.05 | 75.13 | |
| Llama-3.1-70B-InstructPrompting Strategy=Few-Shot2026.01 | 77.85 | 77.82 | 77.72 | 78.08 | |
| Qwen2.5-72B-InstructPrompting Strategy=Few-Shot2026.01 | 77.38 | 77.42 | 77.22 | 77.55 | |
| Llama-3.1-8B-InstructPrompting Strategy=Few-Shot2026.01 | 77.35 | 72.47 | 72.11 | 74.16 | |
| Llama-3.1-70B-InstructPrompting Strategy=Zero-Shot2026.01 | 77.13 | 76.3 | 76.56 | 77.27 | |
| Qwen2.5-72B-InstructPrompting Strategy=Zero-Shot2026.01 | 76.96 | 75.59 | 75.93 | 76.85 | |
| Qwen2.5-32B-InstructPrompting Strategy=Few-Shot2026.01 | 76.93 | 74.82 | 74.9 | 75.96 | |
| Llama-3.1-8B-InstructPrompting Strategy=Zero-Shot2026.01 | 73.79 | 72.81 | 70.96 | 71.17 |