Why AI labs may spend $100 billion on training data

Andrew Ho left OpenAI after eight months and is building a startup focused on high-quality training data. His argument is that scaling large language models alone will not solve weak generalization, especially in specialized work such as bioinformatics and routine lab evaluation.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 0 ►

The story is mainly a business and capability update about improving AI training data, with only a mild lean toward more powerful systems.

Why AI labs may spend $100 billion on training data

Andrew Ho is making a direct bet on where the next major bottleneck in AI will appear: not only in bigger models, but in better training data. After leaving OpenAI after just eight months, he has launched a startup built around the idea that many valuable skills are still missing from the datasets used to train large language models.

His view is blunt. Large language models can look powerful in some settings, yet still perform unevenly when asked to handle contextual work that does not fit neatly into a test, benchmark, or automatically graded environment.

Why scaling alone may not be enough

Ho argues that the core problem is poor generalization. Even in an area such as programming, where major AI labs have invested heavily and models have improved, he says performance remains inconsistent.

The reason, in his view, is that most economically important work is not represented well enough in existing training data. Many tasks depend on context, judgment, and paths that are hard to label as clearly right or wrong.

Most work is highly contextual and not easily encoded into a gradable environment; even if we can observe a 'golden path' taken by a human which we believe to be good, it's challenging to understand whether alternate, counterfactual paths produce good or bad outcomes,

That problem matters because many AI training methods work best when success can be measured clearly. Code and complex math can offer stronger reward signals. Other forms of knowledge work may not.

Ho expects AI labs will eventually need to spend more than $100 billion on targeted data collection in the years ahead. The point is not simply to collect more material, but to collect data that captures the real shape of expert work.

The first targets: bioinformatics and lab work

Ho's startup is beginning with two product areas. The first is datasets for complex scientific analyses in bioinformatics, a topic he worked on at OpenAI.

The source article says that even current models like GPT-5.6 Sol hit only about a 30 percent success rate in that area. That figure is central to Ho's case: if advanced models still struggle in a scientific domain where better automation could be valuable, then better data may be the missing ingredient.

The second product area is routine lab work. One example is researchers submitting photos of experiments to AI models for evaluation. This kind of task is visual, procedural, and context-dependent, which makes it harder than simply answering a text prompt from a well-covered knowledge base.

Ho also plans to move into related areas, including:

  • Chemistry
  • Materials science
  • Healthcare
  • Broader knowledge work

Those planned areas all share the same basic challenge described in the source: AI systems need reliable skills in domains where existing datasets may not contain enough examples of what competent work actually looks like.

Specialization is replacing broad versatility

Cambridge researcher Adam Hunt shares Ho's skepticism. Hunt says his own view of large language models has moved from optimistic to increasingly pessimistic.

His argument is that the latest models are not becoming broadly more versatile. Instead, they are becoming sharper in some domains while stagnating or weakening in others.

Programming and complex math capabilities keep improving, according to the source. But areas such as language quality and simple logic are described as stagnating or even getting worse.

That pattern supports Ho's concern about uneven performance. If improvement depends heavily on whether a domain has clear reward signals and complete training data, then model progress may continue to cluster around tasks that are easy to evaluate.

Hunt's explanation is that reinforcement learning works well in code because the environment can provide clearer feedback. In other fields, the same kind of complete and reliable data is not available.

The generalization question remains open

The broader debate is still unresolved. The question is whether language models can develop capabilities that go beyond reproducing and recombining their training data, especially in areas where results cannot be automatically verified.

The source makes clear that scientists do not agree on the answer. Hunt puts his own confidence at only about 40 percent and acknowledges that future technical advances could prove him wrong.

One possible route he leaves open is that specialized AI models could be combined into something that resembles more general intelligence. That would not necessarily mean a single language model becomes broadly capable on its own. It could mean that different specialized systems work together in a larger architecture.

Google Deepmind's Tom Zahavy offers another explanation in a position paper titled LLMs can't jump. According to the source, Zahavy argues that language models are good at deduction and induction but fail at creative abduction.

In this context, creative abduction means the ability to invent a cause for which no linguistic precedent exists yet. That is a different kind of challenge from pattern completion or applying known reasoning steps.

Zahavy points to action-controllable world models as a possible fix because they allow counterfactual experiments. Large language models could still play a role in such systems, even if they reach limits when used alone.

What the $100 billion bet really signals

Ho's argument also has a business edge. He is skeptical of sky-high valuations at frontier labs like OpenAI or Anthropic, which he says are chronically unprofitable because they must keep spending growing sums on new models to stay ahead of cheaper rivals like Qwen or Kimi.

If he is right, the competitive focus may shift. Bigger models would still matter, but the decisive advantage could come from owning or producing targeted datasets for valuable domains.

That would make training data a strategic asset rather than a background input. In fields such as bioinformatics, lab work, chemistry, materials science, healthcare, and knowledge work, the hardest part may be capturing the judgment behind expert decisions.

The implication is simple: AI progress may depend less on scale by itself and more on whether systems can learn from data that reflects the real complexity of work.