Feedback Evaluation Alignment on MT Bench
0.494Kendall's TauTRACT
Evaluation Results
| Method | Links | |
|---|---|---|
| TRACTCoT=true, Train=C-RAFT, Data=Self, Infer.=C-RAIL, Backbone=Mistral-7B-Instruct2025.03 | 0.494 | |
| RAFTCoT=false, Train=RAFT, Data=GPT-4, Infer.=RAIL, Backbone=Mistral-7B-Instruct2025.03 | 0.455 | |
| CECoT=false, Train=CE, Data=GPT-4, Infer.=Decode, Backbone=Mistral-7B-Instruct2025.03 | 0.429 | |
| RAILCoT=false, Train=None, Data=None, Infer.=RAIL, Backbone=Mistral-7B-Instruct2025.03 | 0.398 | |
| Prometheus-2-7BCoT=true, Training objective=Prometheus-2-7B, Inference method=Decode2025.03 | 0.392 | |
| TRACTCoT=true, Training objective=C-RAFT, Data source=Self, Inference method=C-RAIL, Backbone=Mistral-7B-Instruct2025.03 | 0.386 | |
| Mistral-7B-Instruct + CE (GPT-4 CoT)CoT=true, Training objective=CE, Data source=GPT-4, Inference method=Decode, Backbone=Mistral-7B-Instruct2025.03 | 0.38 | |
| CECoT=true, Train=CE, Data=GPT-4, Infer.=Decode, Backbone=Mistral-7B-Instruct2025.03 | 0.372 | |
| Mistral-7B-Instruct + RAFTCoT=false, Training objective=RAFT, Data source=GPT-4 (score only), Inference method=RAIL, Backbone=Mistral-7B-Instruct2025.03 | 0.342 | |
| Mistral-7B-Instruct + CE (GPT-4 Score)CoT=false, Training objective=CE, Data source=GPT-4 (score only), Inference method=Decode, Backbone=Mistral-7B-Instruct2025.03 | 0.211 | |
| Mistral-7B-Instruct + RAIL BaselineCoT=false, Inference method=RAIL, Backbone=Mistral-7B-Instruct2025.03 | 0.15 |