Snorkel AI’s Value Soars to $3.5B amid Exploding Demand for AI Training Data

Snorkel AI has clinched a $350 million Series E funding round, catapulting its valuation to $3.5 billion. This marks nearly a threefold increase from the $1.3 billion valuation secured during its Series D round about 17 months ago.

The round was co-led by Insight Partners and S32, with backers including Addition, Lightspeed, Greylock, GV, and Wells Fargo also participating.

Launched commercially in 2019 following several years of research at a Stanford AI lab, the startup initially focused on software tools to automate data labeling. However, over the past year it shifted to offering full datasets as a service—combining synthetic data generated by its models with domain experts to deliver complete, ready-to-use training sets.

Snorkel’s annualized run rate has jumped to $375 million, reflecting an eighteenfold growth in the past year. This surge is driven by AI labs and large enterprises demanding high-quality, scalable training data.

The company isn’t alone in this boom. Other AI-focused data businesses have reported similar explosive growth: an enterprise named Mercor is clocking in at around $2 billion in gross annualized revenue, Handshake recently hit $1 billion, and Micro1 has crossed the $500 million mark. However, it’s important to note that roughly 60–70% of top-line revenue in this sector is paid to domain experts (annotators, specialists), significantly reducing the net revenue after costs.

Why the Shift from Tools to Data-as-a-Service Matters

By moving away from purely providing automation tools for data labeling, Snorkel aims to deliver more value up front. Rather than just onboarding clients with software and letting them build their own pipelines, the startup now delivers fully processed datasets—streamlining the path from model development to deployment. This hybrid model emphasizes both machine-generated synthetic content and expert oversight.

In the eyes of its investors and clients, this shift answers a growing challenge in the AI industry: securing reliable, large-scale, domain-specific data without sacrificing quality. AI models are only as strong as the data feeding them, and demand for high-end datasets has surged dramatically. Snorkel’s revenue growth underscores the market’s appetite for solutions that simplify one of the most labor-intensive parts of model training.

The startup traces back to four years of research led by co-founder and CEO Alex Ratner at a Stanford lab. Since launching commercially, it has continuously evolved—first by helping clients label data more efficiently, and now by supplying full datasets ready for machine learning and reinforcement learning tasks.

Investors appear confident this strategy has staying power. With recurring contracts for complete training datasets, the company could hit more predictable revenue streams and scale with fewer dependencies on manual labor. That said, the high cost of paying domain experts—while necessary for quality—remains a pressure point for margins.

What to Watch Next:Can Snorkel sustain this rapid growth, especially as competition from other data-as-a-service providers intensifies? Will it be able to balance automation with expert input without eroding its margin? As AI development increasingly favors specialized, high-quality data, Snorkel could be well positioned—but execution will matter.