Attribute Steering on Llama-2-7b-Chat-hf Open-Ended Generation
2.46Wealth ScoreBiPO
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| BiPOBackbone=Llama-2-7b-Chat-hf, Evaluation protocol=LLM-as-a-judge (GPT-4.1), Type=Preference-training baseline2026.02 | 2.46 | 2.02 | 3.04 | 2.42 | 2.8 | 3.4 | 1.54 | 2.92 | 2.12 | 2.82 | |
| SVFBackbone=Llama-2-7b-Chat-hf, Evaluation protocol=LLM-as-a-judge (GPT-4.1), Description=Steering Vector Fields2026.02 | 2.26 | 2.36 | 3.3 | 2.68 | 2.64 | 3.38 | 1.76 | 2.84 | 1.96 | 2.86 | |
| REDBackbone=Llama-2-7b-Chat-hf, Evaluation protocol=LLM-as-a-judge (GPT-4.1)2026.02 | 1.98 | 1.92 | 3.1 | 2.24 | 2.55 | 2.02 | 1.96 | 2.56 | 2.28 | 2.46 | |
| ICVBackbone=Llama-2-7b-Chat-hf, Evaluation protocol=LLM-as-a-judge (GPT-4.1)2026.02 | 1.84 | 1.72 | 2.86 | 2.3 | 2.38 | 2.56 | 2.12 | 2.39 | 2.18 | 2.42 | |
| CAA(m)Backbone=Llama-2-7b-Chat-hf, Evaluation protocol=LLM-as-a-judge (GPT-4.1), Variant=multi-layer2026.02 | 1.72 | 1.58 | 2.88 | 1.96 | 2.38 | 2.48 | 2.42 | 2.46 | 2.37 | 2.3 | |
| CAA(s)Backbone=Llama-2-7b-Chat-hf, Evaluation protocol=LLM-as-a-judge (GPT-4.1), Variant=single-layer2026.02 | 1.66 | 1.56 | 2.82 | 1.76 | 2.34 | 2.44 | 2.52 | 2.38 | 2.33 | 2.23 | |
| baseBackbone=Llama-2-7b-Chat-hf, Evaluation protocol=LLM-as-a-judge (GPT-4.1)2026.02 | 1.58 | 1.42 | 2.74 | 1.72 | 2.3 | 2.5 | 2.5 | 2.4 | 2.4 | 2.2 |