ResearchTasksSpecification AlignmentFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedSPECBENCH Average over scenariosGPT-5-chat93.73Safety Score33Jun 4, 2026Specification Alignment Evaluation Set (full (1500))Align376.4Safety Score16Jun 4, 2026SPECBENCHGPT-5.496.8Safety Score4Jun 4, 2026