ResearchBenchmarksDiagnostic Report Generation on Human-Eval 12-language (120 items)Follow7.41Human Mean ScoreMADE5.28845.83926.396.9408Jun 5, 2026Evaluation ResultsMethodMethodLinksHuman Mean ScoreAutomatic Mean ScoreHuman Win RateAutomatic Win RateMADE2026.067.418.187.996.2nanobotDescription=strongest...Description=strongest stable external agent baseline2026.065.935.3540.639.2cotDescription=strongest...Description=strongest single-LLM baseline2026.065.373.8521.514.6