Feedback evaluation on Vicuna Bench (test)
0.468Kendall's TauTRACT
Evaluation Results
| Method | Links | |
|---|---|---|
| TRACTCoT=true, Train=C-RAFT, Data=Self, Infer.=C-RAIL, Backbone=Mistral-7B-Instruct2025.03 | 0.468 | |
| CECoT=true, Train=CE, Data=GPT-4, Infer.=Decode, Backbone=Mistral-7B-Instruct2025.03 | 0.406 | |
| RAFTCoT=false, Train=RAFT, Data=GPT-4, Infer.=RAIL, Backbone=Mistral-7B-Instruct2025.03 | 0.386 | |
| RAILCoT=false, Train=None, Data=None, Infer.=RAIL, Backbone=Mistral-7B-Instruct2025.03 | 0.36 | |
| CECoT=false, Train=CE, Data=GPT-4, Infer.=Decode, Backbone=Mistral-7B-Instruct2025.03 | 0.333 |