Prompt-based prediction ranking on Instructional Support Teacher Experience (test)
0Delta ScoreGround-truth DPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Ground-truth DPOBackbone=Qwen2.5-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0 | 2.93 | — | |
| Counterfactual DPOBackbone=Qwen2.5-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0 | 1.63 | -0.08 | |
| Ground-truth DPOBackbone=Qwen2.5-7B-Instruct, Evaluation Protocol=Avg@52026.04 | 0 | 2.93 | — | |
| Counterfactual DPOBackbone=Qwen2.5-7B-Instruct, Evaluation Protocol=Avg@52026.04 | 0 | 1.62 | — | |
| Ground-truth DPOBackbone=Llama-3.1-8B-Instruct, Evaluation Protocol=Avg@52026.04 | 0 | 2.93 | — | |
| Counterfactual DPOBackbone=Llama-3.1-8B-Instruct, Evaluation Protocol=Avg@52026.04 | 0 | 2.48 | — | |
| Debiasing-DPOBackbone=Qwen2.5-7B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.04 | 1.8 | 0.23 | |
| Debiasing-DPOBackbone=Llama-3.1-8B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.04 | 2.2 | 0.21 | |
| Debiasing-DPOBackbone=Qwen2.5-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.05 | 1.49 | 0.22 | |
| Counterfactual DPOBackbone=Llama-3.2-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.07 | 1.18 | -0.01 | |
| Ground-truth DPOBackbone=Llama-3.2-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.14 | 1.12 | 0.08 | |
| Debiasing-DPOBackbone=Llama-3.2-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.15 | 1.23 | 0.16 | |
| DefaultBackbone=Llama-3.1-8B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.23 | 2.46 | 0.08 | |
| DefaultBackbone=Qwen2.5-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.3 | 1.71 | 0.19 | |
| SFT with ground truthBackbone=Qwen2.5-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.32 | 1.67 | 0.09 | |
| SFT with ground truthBackbone=Qwen2.5-7B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.46 | 1.52 | 0.18 | |
| DefaultBackbone=Qwen2.5-7B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.47 | 1.51 | 0.21 | |
| SFT with ground truthBackbone=Llama-3.2-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.74 | 1.29 | 0.11 | |
| DefaultBackbone=Llama-3.2-3B-Instruct, Evaluation Protocol=Avg@52026.04 | 0.79 | 2.45 | 0.13 | |
| SFT with ground truthBackbone=Llama-3.1-8B-Instruct, Evaluation Protocol=Avg@52026.04 | 1.2 | 2.17 | 0.29 |