Task Dimension Prediction on Working Alliance Client Self-Reports (test)
0.5Pearson CorrelationCARE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CARE2026.02 | 0.5 | 0.49 | 1.05 | |
| DeepSeek-R1mode=zero-shot2026.02 | 0.49 | 0.48 | 1.17 | |
| GPT-4omode=zero-shot2026.02 | 0.48 | 0.49 | 1.62 | |
| Llama-3.1-70B-Instructmode=zero-shot2026.02 | 0.45 | 0.43 | 1.32 | |
| Claude-3-Sonnetmode=zero-shot2026.02 | 0.43 | 0.41 | 1.47 | |
| Qwen2.5-14B-Instructmode=zero-shot2026.02 | 0.37 | 0.35 | 1.76 | |
| GPT-4o-minimode=zero-shot2026.02 | 0.36 | 0.34 | 1.11 | |
| Qwen2.5-32B-Instructmode=zero-shot2026.02 | 0.36 | 0.35 | 1.83 | |
| Llama-3.1-8B-Instructmode=zero-shot2026.02 | 0.36 | 0.36 | 2.33 | |
| Qwen2.5-72B-Instructmode=zero-shot2026.02 | 0.34 | 0.33 | 1.62 | |
| Human Counselor2026.02 | 0.3 | 0.28 | 1.61 | |
| GPT-3.5-Turbomode=zero-shot2026.02 | 0.3 | 0.29 | 1.16 | |
| Qwen2.5-7B-Instructmode=zero-shot2026.02 | 0.28 | 0.27 | 1.76 |