ResearchBenchmarksKnowledge-intensive Question Answering on Average (CRAG, NQ, HotpotQA, MuSiQue)Follow36.8TruthfulnessGPT-521.51225.48129.4533.419Sep 30, 2025Evaluation ResultsMethodMethodLinksTruthfulnessHallucinationGPT-52025.0936.828.3OpenAI o32025.093432.7TruthRLBackbone=Llama3.3-70B-...Backbone=Llama3.3-70B-Instruct2025.0929.921PromptingBackbone=Llama3.3-70B-...Backbone=Llama3.3-70B-Instruct2025.0922.128.6