Conversational Intervention on When2Speak
74Macro F1SFT Llama-3.1-8B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SFT Llama-3.1-8BBackbone=Llama-3.1-8B, Training Strategy=SFT, Status=baseline2026.05 | 74 | 47.9 | 4.4 | 52.1 | |
| RL with asymmetric reward shapingBackbone=Llama-3.1-8B, Training Strategy=Reinforcement Learning (RL), Lambda (λ)=0.50, Number of Seeds=32026.05 | 67.3 | 81.4 | 22.7 | 18.6 | |
| Llama-3.1-8B zero-shotBackbone=Llama-3.1-8B, Training Strategy=zero-shot2026.05 | 33.7 | 87.2 | 73.1 | 12.8 |