ResearchTasksSelf-correctionFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedLlama-3.3-70B math n=30 (failure pool)L_MEMORY (Source-Conditioned Role Relabeling)86.7Correction Rate5Jun 5, 2026Qwen-72B math failure pool n=30L_MEMORY (Source-Conditioned Role Relabeling)80Correction Rate5Jun 5, 2026WikiText-2 and OpenWebTextOurs-0.6B82.2NLI Score3May 15, 2026
Llama-3.3-70B math n=30 (failure pool)L_MEMORY (Source-Conditioned Role Relabeling)86.7Correction Rate5Jun 5, 2026
Qwen-72B math failure pool n=30L_MEMORY (Source-Conditioned Role Relabeling)80Correction Rate5Jun 5, 2026