How Micro1’s $500M run rate reflects the AI data boom

Micro1 has expanded its gross annual run rate from $100 million to $500 million over the past eight months, according to a person familiar with the company. The growth points to strong demand for AI training data, even as off-the-shelf datasets and sales to foreign AI developers draw scrutiny.

How Micro1’s $500M run rate reflects the AI data boom

Demand for AI training data is turning data-labeling startups into some of the fastest-growing suppliers in the AI economy. Micro1, a four-year-old startup, is now one of the clearest examples of that shift.

According to a person familiar with the company, Micro1 grew its gross annual run rate from $100 million to $500 million over the past eight months. The company did not respond to a request for comment.

Why Micro1’s growth matters

Micro1 operates in a market where top labs and corporations need unique data to train and improve AI systems. That demand has created a boom for companies that can organize, generate, label, and evaluate data at scale.

The company’s gross annual run rate is not the same as what it keeps. Like similar startups that hire domain experts on a contract basis, including doctors, lawyers, and scientists, Micro1 retains roughly 60% to 70% of that figure. That places its net annual run rate between $150 million and $200 million.

Even with that adjustment, the pace of expansion is notable. Moving from $100 million to $500 million in gross annual run rate over the past eight months suggests that buyers are increasing the size and urgency of data contracts.

Micro1 is not the largest company in this group. Mercor hit $2 billion in gross annualized revenue this summer, while Handshake reached $1 billion earlier this year. But Micro1’s rise still shows that demand is broad enough to support multiple AI data providers rather than concentrating entirely around one or two companies.

The business behind AI training data

Data-labeling companies help turn raw material into something AI models can use. In Micro1’s case, the company works with experts who evaluate model outputs, a practice connected to reinforcement learning gyms.

The company’s work is not limited to expert review. Ali Ansari previously told TechCrunch that Micro1 is also building a robotics pre-training dataset by having hundreds of generalists record everyday object interactions in their homes.

That detail matters because it shows how varied AI training data has become. Model developers are not only looking for text or simple labels. They also need data tied to reasoning, video, physical interaction, and specialized professional judgment.

Some researchers are hypothesizing that future AI spending on data could rival spending on compute. If that view proves accurate, the companies supplying data could become increasingly important parts of the AI stack.

Margins, synthetic data, and repeatable datasets

Micro1 is also trying to improve the economics of the business. The startup is seeing contract sizes grow at an accelerated pace and expects its margins to expand over time.

One reason is synthetic data. Micro1 is increasingly generating synthetic data without human involvement, such as automated descriptions of video content. If less human labor is needed for some types of data creation, the company can potentially keep more of the revenue from each contract.

Another factor is reuse. Some data generated by Micro1 can be sold to multiple customers. A person familiar with the startup’s finances told TechCrunch that gross margins for this off-the-shelf data can be as high as 80% to 90%.

That model is different from purely custom data work. A custom project may depend heavily on people, time, and client-specific requirements. Off-the-shelf data can be created once and sold more than once, which changes the margin profile of the business.

The controversy around off-the-shelf data

Selling the same datasets to multiple clients has also created controversy. Critics argue that distributing off-the-shelf data to Chinese AI developers helps make their models as powerful as top U.S. models.

Ansari addressed the issue last month on X and said Micro1 does not sell its data to Chinese model makers, unlike some competitors.

“Some human data companies work with foreign adversaries. [A]nd the results show today in Kimi K3. We believe it’s shameful to claim American AI dominance desires while selling millions worth of data to countries that we are in adversarial competition with.”

The dispute shows that AI training data is no longer just an operational input. It has become part of a broader debate about competition, access, and which companies should be able to buy high-quality datasets.

From recruiting to data infrastructure

Micro1 did not begin as a data-labeling company. Like Mercor, it started as an AI recruiting startup.

The pivot came after Ansari noticed that data-labeling clients were using his AI platform to vet and recruit engineers for annotation. That customer behavior pushed the company toward the data-labeling business itself.

Micro1 raised its Series A at a $500 m illion valuation las t September. TechCrunch understands that the startup may have recently raised another round at a significantly higher valuation.

The company’s trajectory captures a larger change in the AI market. As model builders compete, the data used to train and refine those models is becoming a major commercial category of its own. Micro1’s growth does not prove how large the market will become, but it does show that demand for AI training data is already producing substantial revenue for newer entrants.