Does More Agentic AI Usage Mean More Business Value?

Published by Industry AI Decision

Agent activity is a useful signal, but value appears only when bounded authority and accountable business outcomes mature with usage.

A team of professionals collaborates around a high-tech table displaying a digital interface with interconnected industrial machinery and robotic arms in a factory setting.

Activity becomes value only through controlled authority and accountable outcomes.

Standfirst: usage is a signal, not a value equation

OpenAI’s Enterprise Signals, updated August 12, 2026, captures a meaningful change in enterprise AI: organizations are asking models to do more than answer questions. As of June, Codex produced 64% of combined Codex and ChatGPT output tokens among enterprise customers. Firms in the top 10% of monthly usage generated 8.3 times as many output tokens per active user as firms around the middle of the distribution, up from a 2.6-times gap in January.

Those figures show intensity and divergence inside OpenAI’s enterprise customer base. They do not show that one group receives 8.3 times the economic value, that Codex produces 64% of enterprise outcomes, or that more tokens cause better performance. OpenAI itself says tokens are an imperfect measure of business value: a short output may be highly valuable, while a long one may add little.

My thesis is straightforward: agentic maturity requires alignment across Activity, Authority, and Accountability. Activity shows how extensively AI is used. Authority shows what systems and decisions an agent can affect. Accountability shows whether outcomes, controls, owners, and escalation paths are explicit. More activity can be a useful leading indicator, but business value emerges only when all three move together.

What changed: enterprise AI is moving from assistance toward delegation

The report defines agentic use for its headline comparison as Codex output tokens. It argues that agentic systems use tools to find information, edit files, and complete multi-step tasks autonomously or under supervision. Complex, long-running work can require more computation, which helps explain why agentic token volume can grow faster than ordinary conversational use.

The gap is not only about tokens. Among weekly active users, 21% at frontier firms use plugins and 19% use skills, versus 9% and 3% at typical firms. OpenAI also says organizations must govern where agents operate, what they access, when they act, and how higher-risk decisions are reviewed.

The strategic change is a shift from employees using chatbots to workflows delegating bounded work. Permission design, integration, verification, and process ownership therefore enter the value equation.

Reading the manufacturing signal correctly

The displayed industry table offers a useful warning against one-dimensional rankings. Manufacturing ranks sixth in ChatGPT adoption, first in ChatGPT intensity, fourth in Codex adoption, and seventh in API intensity. It therefore appears to have intensive ChatGPT use among adopters without leading every channel.

That pattern does not prove that manufacturing is more or less mature than another sector. ChatGPT intensity may reflect engineering analysis, documentation, procurement, quality work, or other knowledge tasks. Codex adoption may reflect internal software capacity. API intensity may reflect embedded products or automated workflows. Without task mix, cost, quality, business outcomes, and risk exposure, the ranks describe use—not ROI.

One source inconsistency matters: nearby prose calls manufacturing last in Codex and API intensity, while the table shows ranks four and seven. I use the displayed table and treat its ranks as a snapshot, not a maturity score.

The mechanism: the 3A Agentic Maturity framework

A1 — Activity

Activity measures exposure and depth: active users, workflows, delegated tasks, tool calls, tokens, repeat usage, and functional coverage. It shows whether AI remains an experiment or enters routine work. Yet tokens may reflect substantive work, verbosity, retries, or task difficulty. Activity must be segmented by workflow and read with cost and quality.

A2 — Authority

Authority measures what an agent may see, decide, and change: read a policy, query inventory, draft a plan, create a requisition, release a work order, or alter a parameter. It includes scope, reversibility, limits, segregation of duties, and approvals. Equal token volume can carry unequal consequences when one agent drafts and another executes.

A3 — Accountability

Accountability connects the agent to an owner, outcome, and control system: KPI, validation, audit trail, exceptions, incident response, override rights, and residual-risk acceptance. Post-deployment monitoring matters because models, data, integrations, and operating conditions change.

This logic is consistent with the NIST AI Risk Management Framework, which organizes risk work around Govern, Map, Measure, and Manage and treats governance as cross-cutting. It also aligns with ISO/IEC 42001, which frames responsible AI as a management system with policies, processes, risk treatment, traceability, and continual improvement.

The 3A framework is a diagnostic, not a certification. Activity without meaningful Authority can become demo theatre; Authority without Accountability is unmanaged autonomy; Accountability without Activity is paper governance. Maturity balances all three around an outcome.

A diagram illustrating a layered approach to data security and management, featuring icons for user profiles, security keys, analytics, and verification processes, interconnected with arrows and nodes.

The 3A framework prevents usage metrics from being mistaken for business value.

A manufacturing example

Consider a hypothetical agent supporting short-interval production planning when a material constraint threatens the schedule. The agent can gather demand, WIP, equipment availability, labor constraints, and approved routing rules, then propose alternatives. No result is assumed; this example defines how a pilot could be governed and measured.

Its authority ladder could be:

1. Observe: read approved data and explain the constraint.

2. Recommend: rank feasible schedule alternatives with assumptions.

3. Reversible workflow action: create a draft schedule, simulation, or approval task that a planner can reject.

4. Production-impacting action: release or revise a schedule in the execution system, only within explicit limits and approval policy.

Activity metrics could include eligible scenarios, completed analyses, planners, and cycle time. Authority records data scope, action level, approvals, rollback, and boundary violations. Accountability links decisions and owners to schedule adherence, lead time, WIP, changeovers, expedites, service, overrides, safety, and quality.

The question is not how many tokens the planner consumed, but whether the governed workflow improved its outcome without shifting cost or risk. A faster local decision could still increase WIP, changeovers, or downstream shortages. Manufacturing value is a system property.

My perspective: measure the conversion from intelligence to accountable action

OpenAI’s report offers useful adoption data, but a leaderboard can shift the objective from business improvement to usage maximization—the wrong target.

Tokens resemble machine runtime more than finished-goods value. Runtime indicates utilization, but a plant manager would not infer profit from spindle hours without yield, throughput, quality, energy, maintenance, and demand. Agentic AI needs the same discipline.

The better metric is value conversion: the share of agent-supported work producing a verified outcome within approved authority after review, rework, errors, risk, and cost. Measures vary: maintenance may use diagnostic lead time; quality, escape rate and false alarms; sourcing, cycle time, compliance, and landed cost.

A Stanford Digital Economy Lab field study of 5,172 customer-support agents found a 15% average increase in issues resolved per hour, with heterogeneous effects. It cannot be transferred to a factory, but its task-level outcome design is more informative than raw usage.

Three implications for manufacturing leaders

1. Benchmarking should trigger questions, not targets

If a plant has low usage, leaders should ask whether access, training, workflow design, or data readiness is blocking useful adoption. They should not mandate token quotas. If usage is high, they should ask which outcomes improved and whether review cost or risk rose.

2. Controls must scale with authority

Read-only analysis may need approved data and source traceability. Recommendations need validation criteria. Reversible transactions need logs, limits, and rollback. Production-impacting actions need stronger testing, segregation of duties, real-time monitoring, and stop conditions. Governance intensity should follow consequence.

3. Complementary capabilities determine whether activity converts to value

The TCS–AWS Future-Ready Manufacturing Study, based on 216 executives, reports that 74% expect agents to manage 11–50% of routine production decisions within three years, but only 21% consider their organizations fully AI-ready. These are expectations and self-assessments, not realized ROI. Still, the contrast highlights the importance of data integration, process redesign, workforce skills, and system readiness.

Limitations and counterargument

Enterprise Signals covers aggregated, de-identified OpenAI enterprise usage, not a representative census. Automated systems classified content, and the ranking lacks workflow-level revenue, cost, quality, or safety outcomes. Because “frontier” means top-decile usage, the 8.3-times gap is a usage comparison, not financial-performance evidence.

Codex tokens are product-specific. They omit some agentic work, input quality, review time, cost, correction, and displaced value. Industry samples may mix office and plant work. Correlation cannot show that deeper use caused capability; digitally stronger firms may both use more AI and perform better.

The counterargument is that usage still matters. Repeated exposure can reveal use cases, build skills, and create organizational learning; very low activity may make value impossible. I agree. Activity is a leading indicator and an experimental input. The mistake is treating it as the outcome.

Five actions leaders can take now

1. Build a workflow-level baseline. Define the current time, quality, cost, risk, and service measures before adding an agent.

2. Assign an authority level. Document what the agent can read, recommend, draft, transact, or change; include limits, approvals, and rollback.

3. Name the accountable owner. One role must own the business outcome, validation rule, exception process, and residual-risk decision.

4. Test value conversion. Compare governed agent-supported work with a credible baseline; include human review, rework, errors, and operating cost.

5. Scale with a 3A portfolio review. Expand workflows only when Activity is sustained, Authority is controlled, and Accountability demonstrates acceptable outcomes.

Conclusion

The rise in agentic AI usage is a real enterprise signal. It suggests that organizations are delegating more complex work and that adoption patterns are diverging. But it is not a business-value scoreboard.

For manufacturers, the disciplined response is neither to dismiss usage data nor to chase it. Use Activity to understand diffusion, Authority to define operational exposure, and Accountability to prove outcomes and control risk. The winning organization will not necessarily generate the most tokens. It will convert the right work into verified, governed results.

Frequently Asked Questions

1. Are output tokens a valid KPI for AI ROI?

They are a useful activity indicator, but not an ROI measure. ROI requires workflow-level benefits, costs, quality, risk, and a credible baseline.

2. Does the 8.3-times frontier gap mean frontier firms create 8.3 times more value?

No. The figure compares output tokens per active user between usage-defined groups. It does not compare profit, productivity, quality, or risk-adjusted returns.

3. What do the manufacturing ranks show?

In the displayed table, manufacturing ranks first in ChatGPT intensity, fourth in Codex adoption, and seventh in API intensity. These are channel-specific usage ranks, not an industry maturity ranking.

4. Which of the 3As should a company improve first?

Start with the weakest constraint for a defined workflow. Low Activity calls for access and task design; excessive Authority calls for limits; weak Accountability calls for owners, metrics, validation, and escalation before scaling.

PUT THE IDEAS TO WORK

Assess a workflow from your own operation.

Use the AI Readiness Assessment to review preparation, identify evidence gaps and save a working record.

KEEP READING

Related guides & perspectives.

Follow the wider topic with another useful question.

RECEIVE NEW ARTICLES

Read the next perspective.

New analysis and learning articles on manufacturing AI, business value and accountable decisions.

Manage delivery preferences or unsubscribe at any time. Privacy policy

Response

  1. […] Agentic AI Business Value: The 3A Maturity Framework […]

Leave a Reply

Discover more from Industry AI Decision | Agentic Manufacturing & Decision Intelligence

Subscribe now to keep reading and get access to the full archive.

Continue reading