Correlation Metrics on MT-Bench (LLM-as-a-Judge)
0.689Pearson's rQwen3-32B REAL (ours)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-32B REAL (ours)Training=RL, Inference=RAIL2026.03 | 0.689 | 0.691 | 0.552 | |
| TRACTBackbone=Llama3.1-8B, COT=true, Train=C-RAFT, Data=Self, Inf=C-RAIL2025.03 | 0.672 | 0.639 | — | |
| Baseline B.4 (RAFT)Backbone=Llama3.1-8B, COT=false, Train=RAFT, Data=GPT-4+, Inf=RAIL2025.03 | 0.618 | 0.614 | — | |
| Qwen3-8B REAL (ours)Training=RL, Inference=RAIL2026.03 | 0.617 | 0.608 | 0.471 | |
| Qwen3-32B RAFTTraining=SFT, Inference=RAIL2026.03 | 0.611 | 0.596 | 0.439 | |
| Mistral2-7B REAL (ours)Training=RL, Inference=RAIL2026.03 | 0.593 | 0.569 | 0.422 | |
| Qwen3-8B TRACTTraining=SFT, Inference=RAIL2026.03 | 0.558 | 0.586 | 0.436 | |
| TRACTBackbone=Mistral-7B-Instruct, CoT=true, Train Objective=C-RAFT, Training Data=Self, Inference Method=C-RAIL2025.03 | 0.555 | 0.529 | — | |
| Baseline B.3 (RAIL)Backbone=Llama3.1-8B, COT=false, Inf=RAIL2025.03 | 0.547 | 0.583 | — | |
| Baseline B.1 (CE)Backbone=Llama3.1-8B, COT=false, Train=CE, Data=GPT-4, Inf=Decode2025.03 | 0.541 | 0.556 | — | |
| Qwen3-8B RAFTTraining=SFT, Inference=RAIL2026.03 | 0.541 | 0.517 | 0.385 | |
| Prometheus-2-7BTraining=SFT, Inference=RAIL2026.03 | 0.538 | 0.511 | 0.389 | |
| Mistral2-7B Standard RLTraining=RL, Inference=RAIL2026.03 | 0.529 | 0.507 | 0.371 | |
| Mistral2-7B TRACTTraining=SFT, Inference=RAIL2026.03 | 0.521 | 0.501 | 0.366 | |
| Prometheus-2-7BBackbone=Prometheus-2-7B, Inference Method=Decode2025.03 | 0.519 | 0.483 | — | |
| Prometheus-2-7BTraining=SFT, Inference=Standard2026.03 | 0.519 | 0.483 | 0.392 | |
| TRACT (Ablation: Objective CE)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=CE, Training Data=Self, Inference Method=C-RAIL2025.03 | 0.517 | 0.503 | — | |
| CLoud (Reward model)Backbone=Llama-3-8B, Inference Method=Decode*2025.03 | 0.511 | 0.506 | — | |
| Qwen3-8B Standard RLTraining=RL, Inference=RAIL2026.03 | 0.495 | 0.535 | 0.397 | |
| Mistral-7B-Instruct (RAFT on GPT-4)Backbone=Mistral-7B-Instruct, CoT=false, Train Objective=RAFT†, Training Data=GPT-4†, Inference Method=RAIL2025.03 | 0.483 | 0.469 | — | |
| Mistral-7B-Instruct (CE on GPT-4 CoT)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=CE, Training Data=GPT-4, Inference Method=Decode2025.03 | 0.48 | 0.482 | — | |
| Prometheus-1-13BTraining=SFT, Inference=Standard2026.03 | 0.467 | 0.455 | 0.345 | |
| Baseline B.2 (CE)Backbone=Llama3.1-8B, COT=true, Train=CE, Data=GPT-4, Inf=Decode2025.03 | 0.466 | 0.494 | — | |
| Mistral-7B-Instruct (CE on Self-gen CoT)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=CE, Training Data=Self, Inference Method=Decode2025.03 | 0.435 | 0.426 | — | |
| TRACT (Ablation: Stage 1 Init)Backbone=Mistral-7B-Instruct, Note=Stage 2 initialized from ps, CoT=true2025.03 | 0.432 | 0.421 | — | |
| Qwen3-32B BaseTraining=None, Inference=RAIL2026.03 | 0.425 | 0.468 | 0.353 | |
| GPT-3.5-Turbo-0613Training=None, Inference=Standard2026.03 | 0.422 | 0.371 | 0.299 | |
| TRACT (Ablation: Data GPT-4)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=C-RAFT, Training Data=GPT-4, Inference Method=C-RAIL2025.03 | 0.399 | 0.418 | — | |
| Mistral2-7B RAFTTraining=SFT, Inference=RAIL2026.03 | 0.399 | 0.418 | 0.307 | |
| Prometheus-1-7BTraining=SFT, Inference=Standard2026.03 | 0.367 | 0.371 | 0.285 | |
| Qwen3-8B BaseTraining=None, Inference=RAIL2026.03 | 0.359 | 0.355 | 0.27 | |
| Qwen3-8B BaseTraining=None, Inference=Standard2026.03 | 0.355 | 0.327 | 0.265 | |
| Mistral2-7B Base (w/ warmup)Training=None, Inference=RAIL2026.03 | 0.32 | 0.299 | 0.214 | |
| Mistral-7B-Instruct (Zero-shot RAIL)Backbone=Mistral-7B-Instruct, CoT=false, Inference Method=RAIL2025.03 | 0.309 | 0.216 | — | |
| Mistral2-7B Base (w/ warmup)Training=None, Inference=Standard2026.03 | 0.309 | 0.318 | 0.253 | |
| Mistral-7B-Instruct (CE on GPT-4 scores)Backbone=Mistral-7B-Instruct, CoT=false, Train Objective=CE, Training Data=GPT-4†, Inference Method=Decode2025.03 | 0.279 | 0.268 | — |