OpenAI's agentic coding tool Codex saw near-total adoption among its own employees in June — 98% were using it — but just 17% of organizational subscribers and less than 1% of individual subscribers had embraced the tool, according to a company-backed study reported by TechCrunch on August 24, 2026. The massive gap between internal and external usage highlights the challenge facing OpenAI as it tries to expand AI agents beyond software engineers to accountants, investors, doctors, and other white-collar professionals. The company released ChatGPT Work last month at its $20-per-month subscription tier, a modified version of Codex designed to let non-engineers access the same autonomous, multistep task completion that developers already enjoy.

The adoption chasm reflects both technical and user experience hurdles that OpenAI engineers acknowledge they're still solving. Andrew Ambrosino, lead engineer for OpenAI's desktop app, said the early Codex interface was "actively hostile" to non-engineering staff, displaying technical readouts like code diffs that meant nothing to communications or finance teams. His team has spent months making the tool more general-purpose. Meanwhile, casual use of the agent proved expensive: one journalist testing the product burned through more than 80 million tokens in four days, costing approximately $65 according to the model's own analysis — a subsidy of more than three times the monthly subscription price. OpenAI hasn't disclosed how many people use Work versus Codex, but the combined app has drawn just 20 million users compared to over a billion prompting ChatGPT online.

Thibault Sottiaux, who oversees OpenAI's core product work including Work, framed the mission as bringing agents to everyone: "In this new factor, ChatGPT can actually do entire, very complicated tasks for you all autonomously in a way that is delightful and safe." The report notes that reaching new professions matters commercially because agents working on longer tasks burn through more tokens, making them more profitable per user. Christian Catalini wrote on Andreessen Horowitz's blog that if labs can't quickly secure the complementary assets needed to scale AI in the market, "value will accrue elsewhere." Vertical-specific competitors like Harvey for law and Clay for sales have been pursuing those customers with model-agnostic approaches, plugging in whichever AI performs best at any given time.

The challenge of expanding beyond coding stems from fundamental differences in how work gets measured and evaluated, according to the report. Software either functions or it doesn't, creating clear benchmarks for AI performance. A strong presentation, business strategy, or sales pitch proves much harder to assess or trace. "One of the unique challenges with a product like this is just that it can really do anything," Ambrosino told TechCrunch. The report explains that most workflows outside engineering lack the digital traces — the step-by-step records of actions and outcomes — that let models learn what good performance looks like. Mario Zechner, creator of the open-source harness Pi, noted that management decisions often produce outcomes months later, making them impossible to capture in simple agent interactions. OpenAI says it uses its GDPval benchmark drawn from 44 occupations plus user feedback, but the report notes that much of the company's design insight comes from watching its own employees use the tools.

OpenAI's product stumbles also reflect competitive pressures. The company originally built Codex as a web app that bet on models being smart enough to handle tasks entirely autonomously with minimal user input — what engineers called being "a bit more AGI-pilled." Anthropic's Claude Code launched afterward with a conversational back-and-forth approach, surveying possibilities and offering users three or four options before proceeding, then checking back repeatedly. That approach proved more effective even though it demanded more work from users, and OpenAI eventually added more interaction opportunities. Download statistics show Claude Code led in demand until April 2026, though Codex has since gained a slight edge. The report notes that Sottiaux expects efficiency gains to drive down costs, pointing to a recent 80% price cut for OpenAI's Luna model, and predicting users should be able to accomplish the same tasks with less spending six months from now. For enterprise customers and everyday professionals alike, the question isn't whether AI agents can automate white-collar work — it's whether the interface, cost structure, and trust model can catch up to the technology's raw capability. The industry now faces a race to solve usability faster than competitors can carve out individual professions with purpose-built tools.