Cognitive State Tracking on User Fidelity Experiment Dataset
77.6E-A ScoreCogWM-14B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| CogWM-14BParameter Count=14B, Training Mode=Joint Optimization2026.06 | 77.6 | 77.3 | 70.5 | 45.8 | 29 | |
| GPT-5.52026.06 | 36.4 | 47.4 | 43.3 | 44.2 | 46 | |
| DS-V4-Pro2026.06 | 31 | 46.2 | 44.5 | 41.2 | 52 | |
| Dual-LLMBackbone=GPT-5.52026.06 | 30.9 | 50 | 44.6 | 44.5 | 47 | |
| Qwen3-14BParameter Count=14B2026.06 | 28.1 | 38.4 | 35.8 | 41.6 | 63 | |
| utt-only-14BParameter Count=14B, Training Mode=Utterance-only fine-tuning2026.06 | 12.2 | 28.5 | 34 | 26.4 | — |