Training data is becoming its own gold rush, and Micro1 is one of the miners. The four-year-old startup’s gross annual run rate climbed from $100M to $500M in eight months, a person familiar with the company said, riding demand from labs and corporations desperate for unique material.
Micro1 hires domain experts, including doctors, lawyers, and scientists, on contract, and keeps roughly 60% to 70% of its gross revenue, putting net revenue in the hundreds of millions. Bigger rivals still tower above it. Mercor cleared $2B in gross annualized revenue this summer, and Handshake passed $1B earlier in the year.
Growth is expected to accelerate. Some researchers now hypothesize that AI spending on data could eventually rival spending on compute, and Micro1 reports contract sizes climbing at an increasing pace with margins expected to widen.
Synthetic data is a growing share of the work, including automated descriptions of video content that need no human involvement. Some datasets sell to multiple customers, lifting gross margins but drawing criticism from those who argue off-the-shelf data shipped to Chinese developers helps their models catch up to top US systems.
Micro1’s fivefold run-rate jump in under a year mirrors the wider labeling boom as labs compete for scarce, high-quality training material.