Builds specialized training data, benchmarks, and evaluation environments that help frontier models and agents perform in high-stakes domains.
Customer voice
Quotes published on Snorkel AI's official website.
). This allowed the joint team to obtain richer training data iteratively (e.g., adding geometric accent chairs as negative labels) to create better models. Snorkel’s experts also generated higher-quality datasets, achieving a 20+ point accuracy boost over vendor-supplied labels. Improved datasets led to improved model performance, enabling customers to find the right products at the right time. Since engaging Snorkel, Wayfair has also seen a 7% increase in view-read—times when a customer clicks through to a product presented to them from a search. This is a leading indicator for add-to-cart rate and increased revenue. Wyafair and Snorkel’s collaborative team overcame its challenges with: Programmatic labeling: By feeding pre-trained foundation models relevant prompts (e.g.,
Subratta ChatterjeePrincipal Data Scientist, MSKCCScaling clinical trial screening with document classification MSKCC, the world’s oldest and largest cancer center, sought to identify patients as candidates for clinical trial studies by classifying the presence of a relevant protein, HER-2. Reviewing patient records for HER-2 is onerous; clinicians and researchers must parse through complex, variable patient data. Snorkel’s experts, using our proprietary technology, collaborated with MSKCC’s team to accelerate the labeling of training data and enabled data-centric iteration. Working together, the team quickly developed a model with an overall accuracy of 93%. This now powers an AI-driven screening system to classify patient records at scale, speeding recruitment for clinical trials and advancing treatment research and development.
Retraining the model resolved this error. Notably, analysis revealed that many ground-truth labels—hand-labeled by domain experts—were incorrect. Correcting these labels removed a hidden constraint on model performance. With just a few rapid iterations, the team achieved an overall accuracy of 93% and an average F1 of 87% across all classes. The test set results were nearly as strong, with 92% accuracy and 87% F1. Programmatic labeling significantly reduced the time required to create high-quality training data for complex, domain-specific text. Model-guided error analysis identified data quality issues, including incorrect ground-truth labels, enabling rapid iteration. Labeling functions encoded the rationale behind each label, improving explainability and maintainability. The resulting document classification AI application now powers a clinical trial screening system that allows MSKCC to identify HER-2 in patient records without requiring human experts to manually review each one.
Customer outcomes
Selected customer stories from Snorkel AI.
Contact Snorkel AI for current plans, volume pricing, and enterprise terms.
Trust profile
Snorkel AI announced the first group of projects funded by its $3 million Open Benchmarks Grants program. The initiative supports researchers and engineers building open-source datasets, benchmarks, and evaluation frameworks for AI systems. Supported projects include tools for evaluating agent performance in long-horizon workflows, coding agent reliability, and terminal-based environments. The program, which remains open for ongoing applications, was established with support from organizations including Hugging Face, Prime Intellect, Together AI, Factory, Harbor, and PyTorch.
Snorkel AI, in collaboration with researchers from Princeton University and the University of Wisconsin–Madison, released Senior SWE-bench, an open-source, Harbor-native benchmark. The tool evaluates the capabilities of AI agents in performing senior-level engineering tasks, such as investigating complex runtime bugs and implementing features based on underspecified requirements. The benchmark includes 100 tasks sourced from real-world open-source repositories and utilizes a validation agent and expert human review to assess code quality, fluency, and adherence to codebase conventions.