Models & Research

Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won’t cut it

· July 30, 2026
Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won’t cut it

The business move

Andrew Ho, a former researcher at OpenAI, is leaving the company to launch a new startup focused on creating specialized training data for AI models. Alongside Cambridge researcher Adam Hunt, Ho argues that simply scaling large language models is no longer enough. Instead of becoming more versatile, these models are excelling in narrow tasks like coding and math but stagnating or even regressing in other areas. Ho predicts AI labs worldwide will need to invest over $100 billion on targeted, high-quality data to fix this imbalance and unlock broader AI capabilities.

Why it matters

The AI industry has long bet on bigger models with larger compute budgets as the key to progress. This move raises the bar on costs and energy without guaranteeing improvements across all tasks. Ho’s thesis shifts focus from raw scale to the quality and curation of training data—something many AI operators have underestimated. For founders, investors, and enterprises, this signals a likely surge in demand for specialized datasets and data-centric AI tools, pushing new business opportunities and raising operational expenses. It also challenges vendors relying solely on larger models, forcing them to rethink how they acquire and integrate domain-specific knowledge.

Who gains and who gets squeezed

Data providers and companies specializing in annotated, domain-focused datasets will gain more leverage and valuation as AI labs spend heavily on targeted data collection. Startups building tools to improve data quality, filtering, and curation stand to benefit from this shift. Meanwhile, generalist model providers and labs betting only on increasing model size without nuanced training data risk falling behind. Operators relying on out-of-the-box language models may face higher costs to adapt models for specialized enterprise or vertical use cases, squeezing margins and complicating deployments.

What to watch next

The effect on AI lab spending patterns could show up in contract announcements, new data partnerships, and acquisitions focused on specialty datasets. Watch for emerging startups funded to tackle data-centric AI problems and how major players adjust their roadmap between scaling models versus refining training data. Also, track developments in data labeling automation and synthetic data generation as alternatives to reduce the high cost of specialized data. How this $100 billion investment unfolds will influence AI pricing, capability gaps, and competitive dynamics over the next several years.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.