Jerry Tworek, CEO of Core Automation and a former OpenAI research leader, contends that the AI industry is measuring progress all wrong. The metric that matters, he argues, is not what models do best but what they consistently fail to do—gaps that determine whether AI can actually work in production.
Tworek made the case at the 2026 AI summit. His position reflects a widening crack in the industry consensus: the belief that throwing more data and computing power at transformer models will solve the fundamental problem. "The era of evals is done," Tworek said, meaning that static benchmarks no longer tell you whether an AI system can reason in the real world.
Core Automation, which Tworek founded after leaving OpenAI in January 2026, is built on this conviction. The company is developing new learning algorithms designed to move beyond pre-training and reinforcement learning, along with architectures that scale more efficiently than existing transformers. The stated goal is to build the "most automated AI lab in the world" by automating its own research processes.
Tworek's vision hinges on a structural shift: small teams using capable AI agents to do work that once required entire organizations. That demands solving what he sees as critical missing components in how models reason—problems that only emerge when you put systems into production loops, not when you run them against a test set.
He is not alone. A cohort of former OpenAI executives has launched what amounts to a coordinated challenge to the scaling thesis. Thinking Machines Lab is led by OpenAI's former CTO. Safe Superintelligence is led by its former chief scientist. All share the view that future progress depends on new paradigms, not bigger models or more training data.
Tworek's seven-year tenure at OpenAI ended, he said, because the company was not set up to pursue this kind of fundamental research. That departure—and the wider exodus—signals a shift in where AI R&D capital and talent are flowing.

