LLM-as-a-judge evaluation on Vicuna Bench
0.605Pearson Correlation (r)TRACT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TRACTBackbone=Llama3.1-8B, COT=true, Train=C-RAFT, Data=Self, Inf=C-RAIL2025.03 | 0.605 | 0.65 | |
| TRACTBackbone=Mistral-7B-Instruct, CoT=true, Train Objective=C-RAFT, Training Data=Self, Inference Method=C-RAIL2025.03 | 0.593 | 0.552 | |
| Mistral-7B-Instruct (RAFT on GPT-4)Backbone=Mistral-7B-Instruct, CoT=false, Train Objective=RAFT†, Training Data=GPT-4†, Inference Method=RAIL2025.03 | 0.567 | 0.519 | |
| TRACT (Ablation: Objective CE)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=CE, Training Data=Self, Inference Method=C-RAIL2025.03 | 0.562 | 0.526 | |
| TRACT (Ablation: Data GPT-4)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=C-RAFT, Training Data=GPT-4, Inference Method=C-RAIL2025.03 | 0.528 | 0.513 | |
| Baseline B.4 (RAFT)Backbone=Llama3.1-8B, COT=false, Train=RAFT, Data=GPT-4+, Inf=RAIL2025.03 | 0.509 | 0.541 | |
| TRACT (Ablation: Stage 1 Init)Backbone=Mistral-7B-Instruct, Note=Stage 2 initialized from ps, CoT=true2025.03 | 0.505 | 0.477 | |
| Prometheus-2-7BBackbone=Prometheus-2-7B, Inference Method=Decode2025.03 | 0.488 | 0.48 | |
| Baseline B.3 (RAIL)Backbone=Llama3.1-8B, COT=false, Inf=RAIL2025.03 | 0.485 | 0.487 | |
| Baseline B.2 (CE)Backbone=Llama3.1-8B, COT=true, Train=CE, Data=GPT-4, Inf=Decode2025.03 | 0.467 | 0.483 | |
| Mistral-7B-Instruct (CE on GPT-4 CoT)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=CE, Training Data=GPT-4, Inference Method=Decode2025.03 | 0.463 | 0.456 | |
| Mistral-7B-Instruct (CE on GPT-4 scores)Backbone=Mistral-7B-Instruct, CoT=false, Train Objective=CE, Training Data=GPT-4†, Inference Method=Decode2025.03 | 0.429 | 0.414 | |
| Mistral-7B-Instruct (CE on Self-gen CoT)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=CE, Training Data=Self, Inference Method=Decode2025.03 | 0.418 | 0.404 | |
| Baseline B.1 (CE)Backbone=Llama3.1-8B, COT=false, Train=CE, Data=GPT-4, Inf=Decode2025.03 | 0.4 | 0.423 | |
| Mistral-7B-Instruct (Zero-shot RAIL)Backbone=Mistral-7B-Instruct, CoT=false, Inference Method=RAIL2025.03 | 0.281 | 0.165 | |
| CLoud (Reward model)Backbone=Llama-3-8B, Inference Method=Decode*2025.03 | 0.229 | 0.311 |