Builds specialized training data, benchmarks, and evaluation environments that help frontier models and agents perform in high-stakes domains.
Contact Snorkel AI for current plans, volume pricing, and enterprise terms.
Trust profile
Snorkel AI announced the first group of projects funded by its $3 million Open Benchmarks Grants program. The initiative supports researchers and engineers building open-source datasets, benchmarks, and evaluation frameworks for AI systems. Supported projects include tools for evaluating agent performance in long-horizon workflows, coding agent reliability, and terminal-based environments. The program, which remains open for ongoing applications, was established with support from organizations including Hugging Face, Prime Intellect, Together AI, Factory, Harbor, and PyTorch.
Snorkel AI, in collaboration with researchers from Princeton University and the University of Wisconsin–Madison, released Senior SWE-bench, an open-source, Harbor-native benchmark. The tool evaluates the capabilities of AI agents in performing senior-level engineering tasks, such as investigating complex runtime bugs and implementing features based on underspecified requirements. The benchmark includes 100 tasks sourced from real-world open-source repositories and utilizes a validation agent and expert human review to assess code quality, fluency, and adherence to codebase conventions.