Can ChatGPT Work make AI agents useful beyond coding?

OpenAI is trying to move AI agents from coding into everyday office work with ChatGPT Work. The challenge is not only model capability, but trust, permissions, usability, and whether non-engineers will let an AI act across their digital tools.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

The story mildly leans toward Terminator because it centers on AI agents gaining access and permission to act across workplace systems, though the risks are mostly practical adoption concerns.

Can ChatGPT Work make AI agents useful beyond coding?

OpenAI is trying to turn ChatGPT from a place where people ask questions into a system that can carry out work across the tools people already use. Its biggest test is ChatGPT Work, a product aimed at white-collar workflows that reach well beyond software engineering.

The promise is simple to describe and hard to deliver: connect an LLM to email, Slack, calendars, documents, spreadsheets, SaaS platforms, and other work systems, then let it handle multi-step tasks. The central question is whether people will give an AI agent enough access to make that useful.

OpenAI wants agents to leave the coding niche

ChatGPT Work was released last month and is available on OpenAI’s lowest subscription tier, for $20 a month. It is based on a modified version of Codex, the company’s coding agent, but it is intended for a much broader group of workers.

For developers, AI agents already have a clearer role. Coding tasks can be broken into files, tests, diffs, and working or failing software. OpenAI now wants similar agentic behavior to help accountants, investors, doctors, and other workers whose daily output depends on digital systems.

That shift matters commercially. Agents that operate for longer stretches use more tokens, which can make them more valuable per user for OpenAI. It also matters strategically because software engineering is only one slice of the professional work AI companies need to reach.

Competitors focused on specific work categories, including Harvey for law and Clay for sales, are already pursuing customers with model-agnostic products. The source article also cites Christian Catalini on a16z’s Time to Build blog warning that if AI labs do not quickly gain the complementary assets needed to scale AI in the market, value may move elsewhere.

The adoption gap is still wide

Inside OpenAI, agent use looks very different from use outside the company. An OpenAI-backed study found that in June, 98% of OpenAI employees were using Codex. Among organizational subscribers, the figure was 17%, and among individual subscribers it was less than 1%.

That contrast frames the opportunity and the risk. OpenAI employees are close to the product, have incentives to experiment, and can tolerate rough edges. Most users do not have the same patience or context.

OpenAI also would not say how many people use Work versus Codex. The combined app is used by just 20 million people, while the company says more than a billion users are prompting ChatGPT online.

That suggests the familiar chat interface has reached mass behavior, while the agentic version of AI work still needs to prove itself. Asking a model for an answer is easy. Allowing it to take action across private systems is a different kind of decision.

Useful agents need access, and access creates tension

Andrew Ambrosino, the lead engineer for OpenAI’s desktop app, has given the app access to and control over his inbox, Slack account, phone, Notion, Figma, and more. He described that as part of testing the future, while acknowledging the privacy risk.

“If I’m asking it to write a document, is there a possibility that it’s going to pull from a private DM on that subject and not know that it’s not supposed to share some info? Yes,”

That tension sits at the center of ChatGPT Work. The product becomes more useful when it can see the information needed to complete a task. But the more it can see, the more users must trust it to handle sensitive context correctly.

OpenAI employees are already using the system for routine, data-heavy work, including weekly metrics reports and turning spreadsheets into planning tools. The article also describes venture capitalists using agents to assemble communications and analysis into investment memos, operations teams creating dashboards and data visualizations, and Sam Altman using it to plan his vacations.

In one example from the article, ChatGPT Work extracted a preschool calendar from email and put it into Google Calendar. In another, it created an auto-updating dashboard of metrics for publicly traded companies, made a queryable database of space launches, and sent a weekly email about new AI research posted at academic clearinghouses.

These examples point to the product’s practical lane: messy coordination work where information already exists but is scattered across formats, inboxes, and apps.

The interface may matter as much as the model

OpenAI engineers describe the software around the model as a harness. That harness determines what information the model can see, which tools it can use, and how it returns results. For an agent, the harness also supplies instructions and access for longer tasks.

Developers were able to benefit from command-line tools because that already matched how many of them work. Most office workers do not live in a command line. That is why OpenAI is trying to make agent workflows feel closer to prompting than to operating technical tooling.

ChatGPT Work adds buttons for selecting projects and plugins while still leaning on the familiar OpenAI interface. Ambrosino compared this transition to skeuomorphism, where digital tools borrowed cues from physical objects to help users understand them.

There are still practical usability problems. The article describes confusing permission setup for cloud drive access, settings split between web and mobile, and cases where ChatGPT Work can create Google Calendar events but not new calendars. It also notes that low effort settings can produce poor results, while higher effort may be necessary for useful work.

Why non-coding work is harder to judge

Software has clearer evaluation signals than many business tasks. Code can run or fail, even if quality still involves judgment. A presentation, strategy, sales pitch, or planning document is harder to score automatically.

OpenAI says it uses GDPVal, a benchmark drawn from 44 occupations and hundreds of knowledge work tests, along with user feedback. But Ambrosino also raised a more informal concern: whether OpenAI’s own internal workflows reflect what everyone else will do, or whether the company is unusual.

That question may decide how broadly ChatGPT Work succeeds. AI agents for office work need more than technical autonomy. They need clear controls, understandable permissions, useful feedback loops, and enough reliability that users keep trying after the first rough experience.

For now, ChatGPT Work represents OpenAI’s attempt to carry the agent idea out of software development and into everyday digital work. The product’s future depends on whether workers see it as a capable assistant, not just a powerful model with too many keys.