ResearchBenchmarksSafety Reasoning Evaluation on BeaverTails 5,000 prompts (subsampled)Follow4.68RelevanceAIDSAFE4.65924.66464.674.6754May 27, 2025Evaluation ResultsMethodMethodLinksRelevanceCoherenceCompletenessCoT Faithfulness (Policy)Response Faithfulness (Policy)Response Faithfulness (CoT)Delta (%)AIDSAFEAgentic Deliberation P...Agentic Deliberation Process=Yes, Base LLM=Mixtral 8x22B, Prompting Strategy=Multi-agent deliberation2025.054.684.964.924.274.915—LLM_ZSAgentic Deliberation P...Agentic Deliberation Process=No, Base LLM=Mixtral 8x22B, Prompting Strategy=Zero-shot2025.054.664.934.863.854.854.99—