ResearchTasksLLM GenerationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedRTX 3090 24GB (inference)Layer-Condensed KV Cache1,150Max Batch Size24Feb 26, 2026Llama-2-7BProtocol 2 (SIGMA)22.1Latency (LAN)12Feb 26, 2026A100 80GB (inference)Layer-Condensed KV Cache128Maximum Batch Size6Feb 26, 2026SyntheticH2O (20%)50.4Latency (s)6Feb 26, 2026SpecBenchEagle2116.95Tokens/s3Feb 26, 2026