Truthfulness on TruthfulQA (Accuracy (%))
97.55Truthfulness AccuracyMetaCrit + Claude-3.5-Sonnet
Evaluation Results
| Method | Links | |
|---|---|---|
| MetaCrit + Claude-3.5-SonnetGroup=Ours, Base Model=Claude-3.5-Sonnet, Monitoring agent (phi_up)=included, Control agent (phi_down)=included2025.07 | 97.55 | |
| MetaCrit + GPT-4oGroup=Ours, Base Model=GPT-4o, Monitoring agent (phi_up)=included, Control agent (phi_down)=included2025.07 | 95.83 | |
| MetaCrit + GPT-3.5-TurboGroup=Ours, Base Model=GPT-3.5-Turbo, Monitoring agent (phi_up)=included, Control agent (phi_down)=included2025.07 | 94.12 | |
| MetaCrit + DeepSeek-v3Group=Ours, Base Model=DeepSeek-v3, Monitoring agent (phi_up)=included, Control agent (phi_down)=included2025.07 | 93.75 | |
| GPT-4oEvaluation Protocol=Closed-Source API, Emission (gCO2/q)=4.52†, Throughput (Tok/s)=55.22026.03 | 92.2 | |
| Gemini 2.5 ProEvaluation Protocol=Closed-Source API, Emission (gCO2/q)=3.90†, Throughput (Tok/s)=58.42026.03 | 90.1 | |
| Claude 3.5 SonnetEvaluation Protocol=Closed-Source API, Emission (gCO2/q)=3.85†, Throughput (Tok/s)=62.12026.03 | 89.5 | |
| MEPGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 89.35 | |
| EcoThinkEvaluation Protocol=Adaptive Inference, Emission (gCO2/q)=1.32, Throughput (Tok/s)=148.62026.03 | 88.7 | |
| Qwen-3-8B-InstructEvaluation Protocol=Standard CoT, Emission (gCO2/q)=2.12, Throughput (Tok/s)=98.52026.03 | 81.5 | |
| MADGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 80.67 | |
| ExpertPromptingGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 80.66 | |
| FrugalGPT (Cascade)Evaluation Protocol=Standard CoT, Emission (gCO2/q)=1.95, Throughput (Tok/s)=88.52026.03 | 80.2 | |
| Llama-3.1-8B-InstructEvaluation Protocol=Standard CoT, Emission (gCO2/q)=2.15, Throughput (Tok/s)=95.82026.03 | 78.4 | |
| Universal Self-consistencyGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 77.11 | |
| MetaCrit w/o ϕ↓Group=Ablation, Base Model=GPT-3.5-Turbo, Monitoring agent (phi_up)=included, Control agent (phi_down)=excluded2025.07 | 76.62 | |
| Self-refineGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 75.89 | |
| MetaCrit w/o ϕ↑Group=Ablation, Base Model=GPT-3.5-Turbo, Monitoring agent (phi_up)=excluded, Control agent (phi_down)=included2025.07 | 75.28 | |
| Zero-shot-CoTGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 70.38 | |
| Qwen-2.5-7B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 63.1 | |
| Gemma-2-9B-itDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 61.4 | |
| TI-DPOModel=LLaMA-3.2-3B2025.05 | 57 | |
| OLMo-2-7B-1124-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 56.3 | |
| Ministral-8B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 55.5 | |
| Mutual-TaughtBackbone=Llama-3-8B-Instruct, Fine-tuning method=Mutual-Taught2025.05 | 55.21 | |
| OLMo-v1.7-7B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 55.2 | |
| PerSynStudent Model=Qwen2.5-3B2025.10 | 55.14 | |
| Llama-3.1-8B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 55.1 | |
| Tulu 3 8BDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 55 | |
| IndividualBackbone=Qwen-14B2025.12 | 54.35 | |
| DTS-TBackbone=Qwen-14B2025.12 | 54.12 | |
| DTS-DBackbone=Qwen-14B2025.12 | 54.11 | |
| TPOModel=LLaMA-3.2-3B2025.05 | 54 | |
| DTS-T*Backbone=Qwen-14B, Variant=Optimized2025.12 | 53.99 | |
| DTS-D*Backbone=Qwen-14B, Variant=Optimized2025.12 | 53.99 | |
| CARStudent Model=Qwen2.5-3B2025.10 | 53.81 | |
| GRPOModel=LLaMA-3.2-3B2025.05 | 53.8 | |
| T-SwitchBackbone=Qwen-14B2025.12 | 53.72 | |
| Task-ArithmeticBackbone=Qwen-14B2025.12 | 53.52 | |
| CPOModel=LLaMA-3.2-3B2025.05 | 53.5 | |
| Family-StrongStudent Model=Qwen2.5-3B2025.10 | 53.37 | |
| KTOModel=LLaMA-3.2-3B2025.05 | 53 | |
| Iterative DPOBackbone=Llama-3-8B-Instruct, Fine-tuning method=Iterative DPO2025.05 | 52.91 | |
| EMR-MERGINGBackbone=Qwen-14B2025.12 | 52.91 | |
| DAREBackbone=Qwen-14B2025.12 | 52.79 | |
| Twin-MergingBackbone=Qwen-14B2025.12 | 52.78 | |
| PerSynStudent Model=Qwen2.5-1.5B2025.10 | 52.22 | |
| DPOModel=LLaMA-3.2-3B2025.05 | 52 | |
| TDPOModel=LLaMA-3.2-3B2025.05 | 52 | |
| MixStudent Model=Qwen2.5-3B2025.10 | 51.82 | |
| Llama-3-8B-InstructFine-tuning method=None (Base)2025.05 | 51.64 | |
| MAP-Neo-7B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 51.6 | |
| SIMPOModel=LLaMA-3.2-3B2025.05 | 51.5 | |
| Ties-MergingBackbone=Qwen-14B2025.12 | 51.46 | |
| StrongStudent Model=Qwen2.5-3B2025.10 | 51.21 | |
| SFTModel=LLaMA-3.2-3B2025.05 | 51 | |
| CARStudent Model=Qwen2.5-1.5B2025.10 | 50.98 | |
| Family-StrongStudent Model=Qwen2.5-1.5B2025.10 | 50.45 | |
| MixStudent Model=Qwen2.5-1.5B2025.10 | 49.73 | |
| OLMoE-1B-7B-0924-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 49.1 | |
| StrongStudent Model=Qwen2.5-1.5B2025.10 | 49.04 | |
| IPOModel=LLaMA-3.2-3B2025.05 | 49 | |
| OLMo-2-7B-SFTDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 47.8 | |
| EvaByte-SFTDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 46.3 | |
| PerSynStudent Model=Llama-3.2-3B2025.10 | 46.15 | |
| OLMo-7B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 44.5 | |
| CARStudent Model=Llama-3.2-3B2025.10 | 44.32 | |
| PerSynStudent Model=Gemma-2-2B2025.10 | 43.87 | |
| MixStudent Model=Llama-3.2-3B2025.10 | 43.56 | |
| PerSynStudent Model=Qwen2.5-0.5B2025.10 | 43.01 | |
| Family-StrongStudent Model=Llama-3.2-3B2025.10 | 42.63 | |
| StrongStudent Model=Llama-3.2-3B2025.10 | 42.31 | |
| CARStudent Model=Gemma-2-2B2025.10 | 42.28 | |
| CARStudent Model=Qwen2.5-0.5B2025.10 | 41.85 | |
| Family-StrongStudent Model=Gemma-2-2B2025.10 | 41.64 | |
| Family-StrongStudent Model=Qwen2.5-0.5B2025.10 | 41.43 | |
| MixStudent Model=Gemma-2-2B2025.10 | 40.83 | |
| MixStudent Model=Qwen2.5-0.5B2025.10 | 40.54 | |
| StrongStudent Model=Gemma-2-2B2025.10 | 40.17 | |
| StrongStudent Model=Qwen2.5-0.5B2025.10 | 39.89 | |
| R1 - 8B + UnsafeChain fullSize=8B, Alignment=UnsafeChain (full)2025.07 | 24 | |
| R1 - 8BSize=8B2025.07 | 23 | |
| R1 - 8B + STAR-1Size=8B, Alignment=STAR-12025.07 | 20.5 | |
| R1 - 8B + UnsafeChain randomSize=8B, Alignment=UnsafeChain (random subset)2025.07 | 20.5 | |
| R1 - 8B + UnsafeChain selectedSize=8B, Alignment=UnsafeChain (selected subset)2025.07 | 20.5 | |
| R1 - 8B + SafeChainSize=8B, Alignment=SafeChain2025.07 | 14 | |
| R1 - 7B + UnsafeChain fullSize=7B, Alignment=UnsafeChain (full)2025.07 | 14 | |
| R1 - 7B + STAR-1Size=7B, Alignment=STAR-12025.07 | 12 | |
| R1 - 7B + UnsafeChain selectedSize=7B, Alignment=UnsafeChain (selected subset)2025.07 | 11.5 | |
| R1 - 7BSize=7B2025.07 | 11 | |
| R1 - 7B + UnsafeChain randomSize=7B, Alignment=UnsafeChain (random subset)2025.07 | 7 | |
| R1 - 1.5B + SafeChainSize=1.5B, Alignment=SafeChain2025.07 | 7 | |
| R1 - 1.5B + UnsafeChain fullSize=1.5B, Alignment=UnsafeChain (full)2025.07 | 6.5 | |
| R1 - 1.5B + UnsafeChain randomSize=1.5B, Alignment=UnsafeChain (random subset)2025.07 | 2 | |
| R1 - 1.5B + UnsafeChain selectedSize=1.5B, Alignment=UnsafeChain (selected subset)2025.07 | 2 | |
| R1 - 1.5BSize=1.5B2025.07 | 1.5 | |
| R1 - 1.5B + STAR-1Size=1.5B, Alignment=STAR-12025.07 | 1.5 | |
| R1 - 7B + SafeChainSize=7B, Alignment=SafeChain2025.07 | 0.5 |