LLM-as-a-Judge Evaluation on FLASK
0.589Pearson's rQwen3-32B REAL (ours)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-32B REAL (ours)Training=RL, Inference=RAIL2026.03 | 0.589 | 0.586 | 0.474 | |
| Mistral2-7B REAL (ours)Training=RL, Inference=RAIL2026.03 | 0.56 | 0.541 | 0.411 | |
| Qwen3-32B BaseTraining=None, Inference=RAIL2026.03 | 0.543 | 0.604 | 0.472 | |
| Qwen3-8B REAL (ours)Training=RL, Inference=RAIL2026.03 | 0.538 | 0.539 | 0.431 | |
| Prometheus-2-7BTraining=SFT, Inference=RAIL2026.03 | 0.525 | 0.514 | 0.392 | |
| Qwen3-8B Standard RLTraining=RL, Inference=RAIL2026.03 | 0.523 | 0.524 | 0.395 | |
| Qwen3-32B RAFTTraining=SFT, Inference=RAIL2026.03 | 0.521 | 0.529 | 0.399 | |
| TRACTBackbone=Mistral-7B-Instruct, CoT=true, Train Objective=C-RAFT, Training Data=Self, Inference Method=C-RAIL2025.03 | 0.518 | 0.501 | — | |
| Mistral2-7B Standard RLTraining=RL, Inference=RAIL2026.03 | 0.516 | 0.505 | 0.379 | |
| Qwen3-8B TRACTTraining=SFT, Inference=RAIL2026.03 | 0.514 | 0.513 | 0.385 | |
| Prometheus-2-7BBackbone=Prometheus-2-7B, Inference Method=Decode2025.03 | 0.512 | 0.493 | — | |
| Prometheus-2-7BTraining=SFT, Inference=Standard2026.03 | 0.512 | 0.493 | 0.405 | |
| Mistral-7B-Instruct (RAFT on GPT-4)Backbone=Mistral-7B-Instruct, CoT=false, Train Objective=RAFT†, Training Data=GPT-4†, Inference Method=RAIL2025.03 | 0.509 | 0.502 | — | |
| Mistral2-7B TRACTTraining=SFT, Inference=RAIL2026.03 | 0.507 | 0.5 | 0.372 | |
| Baseline B.4 (RAFT)Backbone=Llama3.1-8B, COT=false, Train=RAFT, Data=GPT-4+, Inf=RAIL2025.03 | 0.506 | 0.493 | — | |
| TRACTBackbone=Llama3.1-8B, COT=true, Train=C-RAFT, Data=Self, Inf=C-RAIL2025.03 | 0.5 | 0.493 | — | |
| Qwen3-8B RAFTTraining=SFT, Inference=RAIL2026.03 | 0.492 | 0.501 | 0.375 | |
| Baseline B.2 (CE)Backbone=Llama3.1-8B, COT=true, Train=CE, Data=GPT-4, Inf=Decode2025.03 | 0.475 | 0.484 | — | |
| TRACT (Ablation: Objective CE)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=CE, Training Data=Self, Inference Method=C-RAIL2025.03 | 0.468 | 0.436 | — | |
| Prometheus-1-13BTraining=SFT, Inference=Standard2026.03 | 0.466 | 0.429 | 0.346 | |
| Prometheus-1-7BTraining=SFT, Inference=Standard2026.03 | 0.457 | 0.457 | 0.365 | |
| Qwen3-8B BaseTraining=None, Inference=RAIL2026.03 | 0.45 | 0.483 | 0.385 | |
| TRACT (Ablation: Stage 1 Init)Backbone=Mistral-7B-Instruct, Note=Stage 2 initialized from ps, CoT=true2025.03 | 0.448 | 0.437 | — | |
| Qwen3-8B BaseTraining=None, Inference=Standard2026.03 | 0.448 | 0.48 | 0.401 | |
| Baseline B.1 (CE)Backbone=Llama3.1-8B, COT=false, Train=CE, Data=GPT-4, Inf=Decode2025.03 | 0.435 | 0.433 | — | |
| Mistral2-7B Base (w/ warmup)Training=None, Inference=RAIL2026.03 | 0.425 | 0.437 | 0.323 | |
| TRACT (Ablation: Data GPT-4)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=C-RAFT, Training Data=GPT-4, Inference Method=C-RAIL2025.03 | 0.418 | 0.419 | — | |
| Mistral2-7B RAFTTraining=SFT, Inference=RAIL2026.03 | 0.418 | 0.419 | 0.315 | |
| Mistral2-7B Base (w/ warmup)Training=None, Inference=Standard2026.03 | 0.415 | 0.419 | 0.341 | |
| Mistral-7B-Instruct (CE on GPT-4 CoT)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=CE, Training Data=GPT-4, Inference Method=Decode2025.03 | 0.413 | 0.407 | — | |
| Baseline B.3 (RAIL)Backbone=Llama3.1-8B, COT=false, Inf=RAIL2025.03 | 0.412 | 0.445 | — | |
| Mistral-7B-Instruct (CE on Self-gen CoT)Backbone=Mistral-7B-Instruct, CoT=true, Train Objective=CE, Training Data=Self, Inference Method=Decode2025.03 | 0.358 | 0.346 | — | |
| Mistral-7B-Instruct (CE on GPT-4 scores)Backbone=Mistral-7B-Instruct, CoT=false, Train Objective=CE, Training Data=GPT-4†, Inference Method=Decode2025.03 | 0.355 | 0.361 | — | |
| GPT-3.5-Turbo-0613Training=None, Inference=Standard2026.03 | 0.27 | 0.232 | 0.187 | |
| CLoud (Reward model)Backbone=Llama-3-8B, Inference Method=Decode*2025.03 | 0.228 | 0.168 | — | |
| Mistral-7B-Instruct (Zero-shot RAIL)Backbone=Mistral-7B-Instruct, CoT=false, Inference Method=RAIL2025.03 | 0.2 | 0.149 | — |