ResearchBenchmarksMulti-hop Retrieval on HotPotQA (Accuracy, F1, Precision, Recall)Follow82.7AccuracyCARE63.4668.45573.4578.445Apr 20, 2026Evaluation ResultsMethodMethodLinksAccuracyF1-ScoreRecallPrecisionCAREUnderlying LLM=GPT-4.1...Underlying LLM=GPT-4.1, Context list size (n)=10, Approach type=Base (Baseline)2026.0482.781.475.788DirectUnderlying LLM=GPT-4.1...Underlying LLM=GPT-4.1, Approach type=Direct2026.047265.85484.4IndirectUnderlying LLM=GPT-4.1...Underlying LLM=GPT-4.1, Approach type=Indirect2026.0464.247.432.289.8