Language Understanding on MMLU (Accuracy and Forgetting)
87.56MMLU AccuracyQwen3-14B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-14BModel Variant=base2026.03 | 87.56 | — | |
| FinTool-Qwen3-14BModel Variant=domain-tuned2026.03 | 87.02 | — | |
| Qwen3-8BModel Variant=base2026.03 | 85.16 | — | |
| FinTool-Qwen3-8BModel Variant=domain-tuned2026.03 | 84.34 | — | |
| DALAAgent Category=Multi-Agent2025.11 | 84.32 | — | |
| AgentPrune-RAgent Category=Multi-Agent2025.11 | 83.94 | — | |
| ComplexCoTAgent Category=Single-Agent2025.11 | 83.78 | — | |
| LLM-DebateAgent Category=Multi-Agent2025.11 | 83.69 | — | |
| PHPAgent Category=Multi-Agent2025.11 | 83.45 | — | |
| SCAgent Category=Single-Agent2025.11 | 82.66 | — | |
| CoTAgent Category=Single-Agent2025.11 | 82.65 | — | |
| VanillaAgent Category=Single-Agent2025.11 | 82.14 | — | |
| DyLANAgent Category=Multi-Agent2025.11 | 80.16 | — | |
| Qwen2.5-14BPrecision=FP162026.04 | 77 | — | |
| Bit-by-BitBackbone=Qwen2.5-14B, Precision=w2a162026.04 | 75 | — | |
| Bit-by-BitBackbone=Qwen2.5-14B, Precision=w2a22026.04 | 75 | — | |
| Base2026.03 | 73.2 | — | |
| CoT MonitorBackbone=Qwen3-8B, Scenario=Sycophancy2026.03 | 73.12 | — | |
| SARBackbone=Qwen3-8B, Scenario=Sycophancy2026.03 | 73.06 | — | |
| PACED Forward KLWeighting=Beta2026.03 | 73 | 0.2 | |
| Qwen3-8BScenario=Strategic Deception2026.03 | 72.96 | — | |
| Qwen3-8BScenario=Sycophancy2026.03 | 72.96 | — | |
| GRPOBackbone=Qwen3-8B, Scenario=Sycophancy2026.03 | 72.94 | — | |
| GRPOBackbone=Qwen3-8B, Scenario=Strategic Deception2026.03 | 71.91 | — | |
| CoT MonitorBackbone=Qwen3-8B, Scenario=Strategic Deception2026.03 | 71.88 | — | |
| SARBackbone=Qwen3-8B, Scenario=Strategic Deception2026.03 | 71.82 | — | |
| Qwen2.5-7BPrecision=FP162026.04 | 71 | — | |
| Hard Filter Forward KLWeighting=Hard2026.03 | 70.7 | 2.5 | |
| BaseWeighting=-2026.03 | 70.6 | — | |
| Hard Filter Reverse KLWeighting=Hard2026.03 | 70.1 | 0.5 | |
| BaseAllowed Domains=None, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 70.1 | — | |
| PALETTEAllowed Domains=Violence | Hate, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 70.1 | — | |
| PACED Reverse KLWeighting=Beta2026.03 | 70 | 0.6 | |
| Bit-by-BitBackbone=Qwen2.5-7B, Precision=w2a162026.04 | 70 | — | |
| Bit-by-BitBackbone=Qwen2.5-7B, Precision=w2a22026.04 | 70 | — | |
| PALETTEAllowed Domains=Disinfo | Sexual, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 70 | — | |
| AKLWeighting=Token-level2026.03 | 69.8 | 0.8 | |
| PALETTEAllowed Domains=Sexual | Illegal, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 69.7 | — | |
| PALETTEAllowed Domains=Illegal | Violence, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 69.4 | — | |
| PALETTEAllowed Domains=Hate | Disinfo, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 69.2 | — | |
| AKLWeighting=Token-level2026.03 | 68.6 | 2.8 | |
| Reverse KL (unweighted)Weighting=None2026.03 | 68.4 | 2.2 | |
| CoT MonitorBackbone=Llama-3.1-8B, Scenario=Sycophancy2026.03 | 68.2 | — | |
| GRPOBackbone=Llama-3.1-8B, Scenario=Sycophancy2026.03 | 68.13 | — | |
| Llama-3.1-8BScenario=Strategic Deception2026.03 | 68.12 | — | |
| Llama-3.1-8BScenario=Sycophancy2026.03 | 68.12 | — | |
| SARBackbone=Llama-3.1-8B, Scenario=Sycophancy2026.03 | 68.02 | — | |
| Forward KL (unweighted)Weighting=None2026.03 | 66.4 | 6.8 | |
| CoT MonitorBackbone=Llama-3.1-8B, Scenario=Strategic Deception2026.03 | 66.19 | — | |
| CASTAllowed Domains=Disinfo | Sexual, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 65.8 | — | |
| CASTAllowed Domains=Illegal | Violence, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 65.5 | — | |
| GRPOBackbone=Llama-3.1-8B, Scenario=Strategic Deception2026.03 | 65.24 | — | |
| SARBackbone=Llama-3.1-8B, Scenario=Strategic Deception2026.03 | 64.91 | — | |
| CASTAllowed Domains=Sexual | Illegal, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 64.2 | — | |
| COVERCALModel=LLaMA-3-8B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 64.1 | — | |
| CASTAllowed Domains=Hate | Disinfo, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 64.1 | — | |
| CASTAllowed Domains=Violence | Hate, Backbone=QWEN2.5-7B-INSTRUCT2026.05 | 63.4 | — | |
| Max-ActVarModel=LLaMA-3-8B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 63.2 | — | |
| RandomModel=LLaMA-3-8B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 62.5 | — | |
| COVERCALModel=Mistral-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 61.6 | — | |
| LoRAModel=LLaMA-3.1 8B (4-bit), Rank=r=1, # Params=425K2026.04 | 60.9 | — | |
| SOLARModel=LLaMA-3.1 8B (4-bit), Configuration=SOLARr=1(1K→0.3K), # Params=40K (91% ↓)2026.04 | 60.9 | — | |
| Max-ActVarModel=Mistral-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 60.8 | — | |
| RandomModel=Mistral-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 60.2 | — | |
| NOLAModel=LLaMA-3.1 8B (4-bit), Configuration=1000 bases, # Params=128K2026.04 | 56.1 | — | |
| Avg BaselineExpert Combination=Code + Instruction2026.04 | 55.04 | — | |
| Precise FusionExpert Combination=Code + Instruction + Math2026.04 | 55.01 | — | |
| Precise FusionExpert Combination=Instruction + Math2026.04 | 54.98 | — | |
| WIDENExpert Combination=Code + Instruction2026.04 | 54.9 | — | |
| Precise FusionExpert Combination=Code + Instruction2026.04 | 54.66 | — | |
| WIDENExpert Combination=Code + Instruction + Math2026.04 | 54.58 | — | |
| Ties MergingExpert Combination=Instruction + Math2026.04 | 54.45 | — | |
| Avg BaselineExpert Combination=Instruction + Math2026.04 | 54.33 | — | |
| WIDENExpert Combination=Instruction + Math2026.04 | 54.2 | — | |
| Avg BaselineExpert Combination=Code + Instruction + Math2026.04 | 54.19 | — | |
| LoRAModel=LLaMA-3.2 3B, Rank=r=1, # Params=287K2026.04 | 54 | — | |
| SOLARModel=LLaMA-3.2 3B, Configuration=SOLARr=1(1K→0.1K), # Params=16K (94% ↓)2026.04 | 54 | — | |
| Ties MergingExpert Combination=Code + Instruction + Math2026.04 | 53.97 | — | |
| DARE Task ArithmeticExpert Combination=Code + Instruction2026.04 | 53.53 | — | |
| Task ArithmeticExpert Combination=Code + Instruction2026.04 | 53.44 | — | |
| Precise FusionExpert Combination=Code + Math2026.04 | 53.43 | — | |
| DARE TiesExpert Combination=Code + Instruction2026.04 | 53.4 | — | |
| WIDENExpert Combination=Code + Math2026.04 | 53.35 | — | |
| CodeModel Category=Base Models2026.04 | 52.96 | — | |
| NOLAModel=LLaMA-3.2 3B, Configuration=1000 bases, # Params=112K2026.04 | 52.7 | — | |
| Ties MergingExpert Combination=Code + Instruction2026.04 | 52.7 | — | |
| BaseModel Category=Base Models2026.04 | 52.34 | — | |
| Instruction TunedModel Category=Base Models2026.04 | 52.23 | — | |
| Avg BaselineExpert Combination=Code + Math2026.04 | 52.22 | — | |
| Task ArithmeticExpert Combination=Code + Math2026.04 | 52.1 | — | |
| Llama2-13BPrecision=FP162026.04 | 52 | — | |
| DARE Task ArithmeticExpert Combination=Code + Math2026.04 | 51.96 | — | |
| DARE TiesExpert Combination=Code + Math2026.04 | 51.92 | — | |
| Ties MergingExpert Combination=Code + Math2026.04 | 51.86 | — | |
| Task ArithmeticExpert Combination=Instruction + Math2026.04 | 51.5 | — | |
| DARE TiesExpert Combination=Instruction + Math2026.04 | 51.39 | — | |
| Task ArithmeticExpert Combination=Code + Instruction + Math2026.04 | 51.2 | — | |
| DARE TiesExpert Combination=Code + Instruction + Math2026.04 | 51.15 | — | |
| DARE Task ArithmeticExpert Combination=Instruction + Math2026.04 | 51.08 | — | |
| DARE Task ArithmeticExpert Combination=Code + Instruction + Math2026.04 | 50.89 | — |