In Brief: Fresh OpenAI data show that in work contexts, people are more than twice as likely to ask ChatGPT to create a result or complete a task than they are outside of work. But that is not proof of autonomy or ROI: the sample covers individual Free, Go, Plus, and Pro plans, excludes Enterprise and Codex, and classifies messages automatically. For businesses, the main signal is this: adoption should be measured not by the number of chats, but by the share of verified outputs embedded in the real workflow.
This article is intended for business owners, functional leaders, CTOs, and AI transformation leaders. Below are the confirmed figures, the study’s limitations, and a practical plan; market forecasts and promises of specific ROI are not discussed.
Contents
- What OpenAI Found
- How to Read the Data Without Drawing the Wrong Conclusions
- Why Work Is Shifting from Questions to Results
- Which Use Cases Are Growing Fastest
- What Is Changing in Employee Roles
- The “Answer → Result → Process” Matrix
- 30-Day Implementation Plan
- What Metrics to Track
- Frequently Asked Questions
- Conclusion
What OpenAI Found
On August 6, 2026, OpenAI published its first country-level snapshot of how people use ChatGPT. The main result: in messages classified as work-related, users are more than twice as likely to shift into Doing mode — asking ChatGPT to edit, write code, perform analysis, or create another output — than they are outside of work.
Four more signals from the release:
- ChatGPT is used by more than 1 billion people;
- multimedia reached 7.8% of messages and became the fastest-growing major use case;
- in Brazil and Colombia, multimedia accounts for more than one in ten messages;
- the share of messages from users age 35+ rose over the past year in nearly every country; the global increase was 5 percentage points, and in France and the Czech Republic it was more than 10 points.
The Q2 2026 country rankings covered 144 countries. Usable estimates for multimedia were published for 126 countries, and the age time series with full estimates covered 111 countries. These are vendor-reported Measured OpenAI data, not an independent audit of the entire generative AI market.
| Signal | What It Means for Business | What It Does Not Prove |
|---|---|---|
| Doing at work is more than 2x more likely | employees use AI to produce artifacts | that the output is correct or accepted in production |
| Multimedia — 7.8% of messages | the AI interface is moving beyond text | that every company needs image generation |
| a 5 pp increase in the share of users age 35+ | the audience is broader than early enthusiasts | that age determines quality of use |
| more than 1 billion users | the interaction habit is already mainstream | that enterprise adoption is complete |
How to Read the Data Without Drawing the Wrong Conclusions
The OpenAI Signals v2.0 methodology is based on a sample of 300,000 messages per month from July 2024 through June 2026. It includes messages from adult users with reported ages. The texts are stripped of personal data, classified by models, and the public aggregates are protected with differential privacy.
There are four critical limitations:
- The sample does not include ChatGPT Enterprise or Codex messages. So, according to the authors, it likely understates work usage.
- Work, Doing, and topics are classifier outputs, not responses people gave in a survey.
- Users who deleted their accounts or opted out of data use for training were excluded.
- The share of messages does not equal the number of completed processes, hours saved, or financial impact.
So the correct takeaway is this: in individual ChatGPT, work messages are clearly more likely to be output-oriented. The incorrect takeaway is, “ChatGPT is already autonomously doing most office work.”
Why Work Is Shifting from Questions to Results
Outside work, people often look for explanations: how to choose a product, cook a dish, or understand a concept. At work, value appears when the output is an artifact: a spreadsheet, email, code, analysis, memo, presentation, or a completed CRM record.
That changes the smallest unit of adoption. It is no longer access to chat or a prompt library, but a chain of:
input data → model action → verification → handoff of the result into the record system.
If the last step is manual and recorded nowhere, the company sees activity but not productivity. If verification is missing, speed rises along with the risk of errors. The practical architectural principles for these longer workflows are covered in the piece “Asynchronous AI Agents: Architecture, Memory, and Verification”.
Which Use Cases Are Growing the Fastest
OpenAI classifies multimedia as image creation and analysis, as well as the generation or retrieval of other media. After the release of ChatGPT Images 2.0 in April 2026, this category’s share rose to 7.8% of messages. However, practical guidance, writing, and information seeking are still larger.
For a company, that is a reason to look for multimodal inputs wherever employees already work with more than just text:
| Function | Input | Output | Required Review |
|---|---|---|---|
| sales | call recording, correspondence | summary, next step, CRM fields | manager confirms facts and promises |
| support | screenshot, ticket, knowledge base | classification and draft reply | policy, personal data, escalation |
| marketing | brief, brand assets | creative and copy variations | rights, brand guidelines, factual claims |
| operations | document photo, spreadsheet | extracted fields and exceptions | reconciliation of totals and sample checks |
A use case’s popularity should not determine priority. Priority is set by a combination of frequency, manual effort, verifiability, and the cost of errors.
What Changes in Employee Roles
A related OpenAI report from July 27 analyzes more than 800,000 work messages from U.S. users. In it, 16.8% of all work messages and 43.5% of messages about profession-specific tasks relate to work typically associated with a different profession.
After excluding universal tasks, crossover is especially high in customer experience (77%), designers (75%), HR (69%), lawyers (56%) and marketers (53%). Among moderately active users, the share of tasks outside their own profession falls from 18.9% in teams of 2–5 people to 16.3% in organizations with more than 100 people.
This does not mean departments disappear. Rather, the number of small handoffs shrinks: the marketer does the initial analysis, the manager writes SQL, the designer checks the frontend prototype. The specialist in the adjacent function remains the owner of the standard and the exceptions.
The “Response → Result → Process” Matrix
AI Dawn suggests assessing implementation maturity across three levels. This is our analytical framework, not an OpenAI term.
| Level | What the AI Does | Control | Metric |
|---|---|---|---|
| 1. Response | explains, searches, suggests options | a person judges usefulness | time to decision |
| 2. Result | creates an email, code, spreadsheet, or design | the task owner accepts or edits | acceptance rate and edit time |
| 3. Process | receives data, acts through tools, records the outcome | access rules, evals, log, rollback | success rate, cost per task, incident rate |
Moving straight from Level 1 to Level 3 is risky. First, collect examples of accepted and rejected outputs at Level 2. They will become the test set for process automation.
The depth of control is determined by risk. A draft internal email can be reviewed after generation. A payment, an access-rights change, or a legal commitment should require explicit confirmation before action. For integrations with accounting systems, a step-by-step approach from the guide “How to Connect AI Agents to 1C” is useful.
30-Day Implementation Plan
- Days 1–3: choose one frequent process with 50–200 repetitions per month, measurable output, and reversible errors.
- Days 4–7: gather 30–50 real anonymized examples and benchmark outputs.
- Days 8–14: launch the “Result” mode without autonomous writing to the target system; an employee makes every decision.
- Days 15–21: calculate acceptance rate, median edit time, cost per accepted result, and error types.
- Days 22–26: Add retrieval, templates, tool constraints, and logging for the two most common errors.
- Days 27–30: Compare against the baseline and decide: stop, continue human-in-the-loop, or automate the low-risk part.
The numbers 30–50 and 50–200 — Estimated practical pilot thresholds, not OpenAI values. They provide enough repetition for an initial diagnosis, but they do not replace a statistically sound experiment.
What metrics to track
| Metric | Formula | Why it matters |
|---|---|---|
| Acceptance rate | accepted results / all results | separates production from a demo |
| Edit time | median minutes to acceptance | shows the hidden cost of review |
| Task success rate | tasks meeting the criteria / all runs | measures the process end to end |
| Cost per accepted task | model + infrastructure + review / accepted tasks | makes it possible to compare against the baseline |
| Incident rate | incidents / 1,000 runs | keeps quality under control as you scale |
| Escalation rate | handoffs to a specialist / all tasks | shows the system’s competency boundary |
Message volume and active users are useful for adoption, but they are not enough for ROI. An employee may send more messages because the answers are not good enough. Tie usage to accepted outcomes, cycle time, and errors.
Frequently asked questions
How is ChatGPT used in business in 2026?
In work messages, people are more than twice as likely to ask for a finished output or task execution as they are outside of work. Writing, coding, analysis, and multimedia use cases are common, but the exact profile depends on the role.
Does the report prove productivity growth?
No. It measures message structure and adoption, not hours saved, quality, or financial results. ROI must be measured within a specific process.
Does OpenAI Signals include enterprise accounts?
No. The dataset includes ChatGPT Free, Go, Plus, and Pro, and excludes Enterprise and Codex. As a result, activity at large companies is only partially reflected.
What does Doing mean?
It is an automatically identified request to perform an action or produce an output—for example, editing, coding, or analysis. Doing does not mean the output was accepted or that the action was completed autonomously.
What use case should we start with?
Start with a frequent, measurable, and reversible process where the output is easy to verify: support ticket classification, response drafts, field extraction, or preparing a recurring report.
What data should not be sent in a pilot?
Do not use personal, financial, medical, or commercially sensitive data without approved agreements, retention settings, access controls, and your company’s security policy.
Conclusion
OpenAI Signals shows an important shift: at work, ChatGPT is used primarily not as a reference tool, but as a tool for producing outputs. At the same time, the report does not measure quality, process completion, or return on investment, and its sample does not include Enterprise or Codex.
So the right next step for a business is not to buy licenses for everyone, but to choose one workflow and follow the path Response → Output → Process. First measure acceptance rate and the cost of review, then integrate data, and only after that delegate low-risk actions to the system. That is how mass AI adoption turns into manageable productivity.