ResearchBenchmarksHuman-Agent Interaction Assessment on Software Agent Case Study Claude-4 vs. GPT-5Follow-0.069Effect Size∆augment-0.07524-0.07362-0.072-0.07038Oct 10, 2025Evaluation ResultsMethodMethodLinksEffect SizeEffect Size (CI Lower)Effect Size (CI Upper)Significance Level (p-value)∆augmentCondition=augmentCondition=augment2025.10-0.069-0.116-0.021—∆naiveCondition=naiveCondition=naive2025.10-0.075-0.136-0.013—