Loading the SOTA2 catalog…
Compare ai evaluation & observability products for capabilities, deployment, pricing, and enterprise readiness. Ranked using verified practitioner feedback, product research, and enterprise trust signals.
| Company | Description | Matching products |
|---|---|---|
| Fabraix, Inc. | AI research lab building frontier red-teaming AI agents that continuously find security vulnerabilities in customer-facing AI. | Nyx 1 more |
| Physical Turing | Independent test and validation lab for humanoid policies and robots, conducting real-world trials for safety, dexterity, autonomy, and endurance. | |
| Middleware | Middleware is a full-stack observability platform that detects issues from infra, APM, RUM and resolves them using the AI SRE Agent. | |
| Centaur Labs |
| Centaur Labs (d/b/a Centaur.ai) provides data annotation and model evaluation services using a global network of expert labelers and collective intelligence aggregation. |
| Polymath AI Labs | Polymath is a data lab focused on increasing the reliability and autonomy of AI agents by building simulation environments for training and evaluating autonomous agents. |
| Datasaur | AI lab building secure, private LLMs and agentic workflows for regulated enterprises, deploying custom AI solutions behind customer firewalls. |
| Parea AI, Inc. | Experimentation and human annotation platform for teams building production-ready LLM applications. |
| Instance | Instance provides a verification layer for robot learning, delivering reliable success labels for training, evaluation, and reinforcement learning. |
| Scope | Agent Experience Platform that helps software companies make their products discoverable and usable by AI agents. |
| Oximy | AI adoption engine for every company. Oximy sees how your team uses AI, drives better usage, and proves the ROI. Detects 6,500+ tools with 60-second install and results in 90 days. |
| Appen Limited | Provider of high-quality data for the AI lifecycle, offering AI training data, annotation, labeling, collection, and managed services across text, image, audio, video, and geospatial data. |
| Nitrode | Provides high-quality spatial reasoning data that helps LLMs, agents, and world models understand dynamic environments, focused on advancing game development through frontier AI research. |
| TrustAI | Risk assessment platform for ERP AI agents. Assesses blast radius, correctness, hallucination/grounding, robustness at scale, and security/adversarial resistance before production deployment. |
| Arya AI | Enterprise-grade AI solutions for financial institutions and enterprises. |
| Qubrid AI | Open, inference-first AI platform for enterprise developers offering GPU compute, serverless inference, fine-tuning, and RAG on open-source models. |
| Bluejay | QA platform for AI agents that enables rigorous testing, monitoring, and improvement of voice and chat AI agents through real-world simulations and actionable observability. |
| Vibrant Labs | Applied research lab building novel data harvesting techniques, providing autoscaled RL data for frontier agents. | Ecom Bench 1 more |
| Solidroad | AI-powered QA and coaching platform for enterprise CX teams, reviewing 100% of customer interactions and generating agent training across human and AI agents. |
| Synth AI | The last mile for your agent's production readiness. |
| Modaic | Modaic provides an infrastructure layer for AI decision-making, targeting classification, ranking, judging, and extraction workflows. The platform scores model confidence, routes uncertain cases for human feedback, and iteratively improves underlying instructions to reduce manual oversight. |
| Moda | Continual learning layer for AI agents. Analyzes production conversations to surface user intent, agent failures, and frustration causes, then feeds validated improvements back into the agent harness. |
| Relari, Inc. | Relari, Inc. (dba Nuvi.dev) is an AI agent builder for Software 3.0, helping teams turn natural language specifications into reliable, testable agents with no code. |
| Tenyks | Vision AI Platform and MLOps tools that empower businesses to achieve workforce efficiency, operational excellence, and better ML models through visual data intelligence. |
| Raycaster | Builds benchmark data, evaluation harnesses, and reference agents for consequential professional work across documents, spreadsheets, and tools, with Workspace products that run the same harness. |
| Olam Labs | Research lab building multi-agent simulated environments and public arenas that measure social intelligence in AI models for safety, agent performance, and real-world use. |
| Voker | Analytics platform for AI agents that transforms interactions into structured insights. Track agent performance, identify knowledge gaps, and measure business impact. |
| Robocurve | A Public Benefit Corporation building open-source tools and independent benchmarks to help society understand the capabilities of AI and robots in the physical world. |
| Monte | Applied research lab building continual learning and post-training infrastructure for AI agents. Embeds with enterprise teams to ship agents that improve through experience. |
| Metoro | AI SRE and observability platform for Kubernetes that collects logs, metrics, traces, profiling, Kubernetes events, and deployment context with eBPF. AI agents detect production issues, investigate alerts, verify deployments, identify root causes, and open fix PRs. |
| Roark | Simulation testing and evals platform for voice and chat AI agents. Simulate every scenario before launch, score every production call on audio-native metrics. |
| Lucidic AI | The training and optimization platform for reliable AI agents. Lucidic AI builds the training and optimization layer for AI agents, enabling enterprises to measure, improve, and safely deploy agents through parameterization, targeted simulations, and automated optimization. |
| Experiential Labs | Enables intelligence layers to improve from experience by turning signals derived from production usage into realistic simulations that co-evolve with agents to continuously raise performance ceiling. |
| Rockfish Data | Rockfish Data operates a generative data platform that creates synthetic time-series data, enabling enterprises to test AI models and agents against rare events, edge cases, and scenarios before production deployment. |
| /dev/fast | An AI-native code forge that rebuilds code review, deployment infrastructure, and hosting for the AI era. |
| 2Digit Co., Ltd. | AI company developing financial AI for market prediction and industry trend analysis, founded in 2018. |
| Agentic Labs | Provides a context engine for AI teams, collecting, cleaning, and organizing context for agent deployments. |
| Aluna | Data foundry for life science AI. Evals and datasets for life science AI. |
| Ask Sage | Secure generative AI platform for government and regulated organizations; a BigBear.ai company founded by Nic Chaillan. |
| Autonomize AI | AI platform for extracting and using knowledge from healthcare and life-sciences data. |
| Chronicle Labs | AI agent testing and validation platform that turns production data into staging environments to catch failures before launch. |
| Clawvisor | Clawvisor provides an AI agent security gateway and control plane that secures agent interactions with tools and data through task-based authorization, credential vaulting, and audit logging. |
| Convexia | Convexia is an AI-maximalist pharma company that builds autonomous AI agents to scout, vet, and accelerate overlooked drug assets through an end-to-end sourcing and evaluation stack. |
| Coxwave Co., Ltd. | Operates Coxwave Align (also referred to as Align AI), an analytics engine for LLM-based conversational products enabling organizations to monitor, analyze, and evaluate AI chatbot interactions. |
| Dynamo AI | Dynamo AI offers end-to-end AI performance, security, and compliance solutions for delivering enterprise-grade generative AI. |
| Fits on Chips | Toolkit for LLM benchmarking that helps users achieve optimal performance across various deployment frameworks and devices by creating experiments, executing them, and gaining insights to optimize LLM deployment parameters. |
| HUD AI | The platform for building RL environments. Train specialized agents, evaluate models, and deliver frontier-grade post-training data to labs. |
| Koyal | Agentic AI Filmmaking Platform that turns scripts or audio into engaging films with consistent characters and storylines. |
| Maingen | AI-native software platform for solar operations and maintenance: unified monitoring, maintenance prioritization, and compliance. |
| Maitai | Enterprise AI platform for reliable, high-performance inference. Fine-tune, evaluate, and monitor production LLMs with confidence. |
| MangoDesk | MangoDesk evaluates and improves AI on meaningful use cases via production-grade RL environments, accelerating custom evals and post-training data creation. |