The Real Economics of AI Agents
· Avery Team
A market analysis for decision-makers: value, ROI, the production gap, and where Avery.Software fits
August 2026
Executive Summary
AI agents are the fastest-growing line item in enterprise software, with worldwide AI spending on pace to reach roughly $2.6 to $2.7 trillion in 2026 and agent-specific software climbing toward $220 billion by 2030 (Gartner, 2026). Underneath that spending, the return on investment is thin and unevenly distributed. Independent surveys converge on a consistent picture: somewhere between 5 and 25 percent of organizations report AI initiatives that meet their own ROI expectations, depending on how the question is asked and who is asked (IBM, MIT, ISACA, BCG, KPMG, and Morgan Stanley, 2025 to 2026). Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027, and it attributes that not to weak models but to escalating costs, unclear business value, and inadequate risk controls.
This report examines four things a decision-maker actually needs to know before committing budget to agentic AI:
- Where the ROI claims hold up and where they do not, based on independent research rather than vendor marketing.
- Why agents stall before production once the use case gets past a narrow, well-scoped task. The evidence for determinism problems, trust and security gaps, and compliance exposure is strong and well documented, not vendor talking points.
- How the major categories of agent platforms actually differ: workflow automation tools (n8n, Zapier, Make), RPA incumbents (UiPath, Automation Anywhere), platform-native SaaS agents (Salesforce Agentforce, ServiceNow), and the open-source personal-agent movement (OpenClaw and its peers).
- What a production-grade platform requires architecturally, and how Avery.Software's approach, compiled agents, an executable policy layer (Rulebook), self-improving workflows, and a usable frontend, addresses the specific failure modes identified in this report, alongside where Avery is still an early-stage company relative to the incumbents it is compared against.
The throughline of the research is consistent: the technology that captures headlines (a model reasoning its way through an open-ended task) is not the technology that survives contact with production. The technology that survives production is closer to traditional software engineering, rules, typed workflows, access control, audit logs, wrapped around a model that is used only where judgment is genuinely required. Every vendor in this report, including the largest ones, is moving toward that architecture. The question for a buyer is not whether governance and determinism matter; every serious vendor now agrees they do. The question is how completely a given platform delivers them, and at what cost in flexibility, price, and data control.
Part 1: The Gap Between AI Agent Spending and AI Agent Value
1.1 The market is large and growing fast
Gartner forecasts worldwide AI spending of $2.59 to $2.67 trillion in 2026, up roughly 47 to 49 percent year over year, with the figure reaching $5.95 trillion by 2030 (Gartner, May and August 2026). Within that, embedded agentic AI is expected to grow from $88.2 billion in 2025 to $1 trillion by 2030, and the standalone AI agents and assistants category is projected at $219.9 billion by 2030 (Gartner, August 2026). Independent market researchers put the narrower "AI agents market" (agent-specific software and platforms, excluding broader AI infrastructure) at $8 to $12 billion in 2026, growing to $48 to $53 billion by 2030 at a compound annual growth rate near 45 percent (MarketsandMarkets and Research and Markets, 2026). Gartner separately projects that 40 percent of enterprise applications will embed a task-specific AI agent by the end of 2026, up from under 5 percent a year earlier.
The capital is real. The discipline behind how it is spent is not yet there for most organizations.
1.2 What independent research actually finds on ROI
The following figures come from separate research organizations, using different methodologies, surveying different populations, in 2025 and 2026. They are presented together because the spread between them is itself informative: it shows how sensitive the "AI is working" narrative is to exactly what is measured.
| Source | Finding |
|---|---|
| IBM Institute for Business Value (2025 CEO study, 2,000 CEOs) | Only 25% of AI initiatives delivered the ROI leaders expected; only 16% had scaled enterprise-wide; only 29% of executives say they can measure ROI confidently, even though 79% report productivity gains |
| MIT Project NANDA, "The GenAI Divide" (2025, 300 deployments analyzed, 150+ executive interviews) | 95% of generative AI pilots showed no measurable P&L impact; only about 5% reached production with measurable value; 42% of companies abandoned most of their AI projects outright in 2025 |
| ISACA 2026 AI Pulse Poll (3,400+ digital trust professionals) | Only 22% say AI ROI has met or exceeded expectations; 23% say it is too early to tell; 56% are unsure how long it would take to halt an AI system in a security incident |
| BCG AI Radar (2026) | 5% of enterprises report achieving substantial ROI at scale |
| KPMG Global AI Pulse, Q1 2026 (2,110 C-suite leaders, 20 countries) | 8% report measurable ROI |
| Morgan Stanley (S&P 500 analysis, 2026) | Only 21% of S&P 500 companies could point to any measurable AI benefit |
| RAND Corporation (analysis of enterprise AI initiatives) | AI project failure rates exceed 80%, roughly twice the failure rate of comparable non-AI technology projects |
| Gartner (June 2025, updated through 2026) | Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value, or inadequate risk controls, not model capability |
Not every study tells the same story. Deloitte's 2024 enterprise survey found 74% of respondents said their most advanced AI initiative met or exceeded ROI expectations, and IDC's 2024 study (sponsored by Microsoft) found an average return of $3.70 per dollar invested in generative AI, rising to $10.30 for top adopters. These more optimistic numbers are not wrong; they are measuring something different. They tend to ask about a company's single best initiative, often a narrow pilot picked because it already worked, rather than asking what share of all initiatives, across the whole portfolio, delivered a return. Read together, the honest summary is: a small number of well-scoped, well-integrated deployments produce strong, defensible returns, while the large majority of agentic AI spending sits in pilots that never convert into measured value. A decision-maker evaluating a vendor's ROI claims should always ask which of these two populations the number describes.
1.3 Where the ROI that does exist actually comes from
MIT's research is unusually specific on this point: the highest returns are concentrated in back-office automation, document processing, and internal operations, not in the sales and marketing pilots that receive most of the budget. Most companies allocate the majority of their generative AI spend to customer-facing, high-visibility use cases where ROI is comparatively weak, while under-investing in the unglamorous operational workflows where AI-assisted automation reliably reduces cost. Gartner's guidance echoes this: agentic AI should be reserved for situations that genuinely require decisions, with plain automation used for routine workflows and simple retrieval left to assistants, rather than routing every task through an autonomous agent because it is fashionable to do so.
1.4 The "agent washing" problem
Gartner estimates that of the thousands of vendors now marketing "agentic AI," only about 130 have agentic capabilities that meet the analyst firm's own bar; the rest are rebranded assistants, chatbots, or robotic process automation with an LLM layer added on top, a practice Gartner calls "agent washing." This matters directly for a buyer's due diligence: the presence of the word "agent" in a product name or pitch deck is not evidence of the architecture underneath it, and a meaningful share of the category's revenue and headline growth is inflated by relabeling rather than new capability.
Part 2: Why Agents Stall Before Production, Testing the Premise
The premise behind this report, that agentic AI struggles to move past simple use cases because of a lack of determinism, unresolved trust and security questions, compliance exposure, and unpredictable behavior, is well supported by independent evidence. This is not primarily a marketing narrative constructed by governance vendors; it shows up consistently in analyst research, security research, regulatory guidance, and public incident reports.
2.1 Determinism: the compounding error problem is mathematical, not anecdotal
A large language model completing a single, well-scoped step can be highly accurate. The problem is that most real business processes require many steps in sequence, and step-level accuracy compounds multiplicatively, not additively. If each step in a workflow succeeds independently with probability p, the probability that an n-step workflow completes correctly end to end is approximately p^n. At 95 percent accuracy per step, a 10-step workflow succeeds roughly 60 percent of the time, and a 20-step workflow succeeds roughly 36 percent of the time. At 90 percent per-step accuracy, a 10-step workflow drops to roughly 35 percent. At 85 percent per-step accuracy, a strong result for a genuinely complex reasoning task, a 10-step workflow succeeds only about one time in five (multiple independent engineering analyses in 2026, including Zartis and Temporal's published calculations, converge on this figure). METR's empirical work on long-horizon task completion shows the same effect from a different angle: frontier models complete tasks that would take a human under four minutes with near-100 percent reliability, but success rates fall below 10 percent on tasks that would take a human more than four hours, and the length of task a model can reliably complete has been doubling roughly every seven months.
This is the strongest, most technical validation of the "lack of determinism" concern. It is not that models are unreliable in some vague sense; it is that reliability decays predictably and mathematically as a workflow gets longer, and no amount of prompt engineering repeals that math. The only architectural fixes are to shorten the chain of steps that depend on model judgment, verify between steps, or replace steps that do not actually require judgment with deterministic code and rules, since deterministic steps do not carry the same failure probability.
2.2 Trust and security: real, recent, and severe
The clearest evidence here comes from OpenClaw, the self-hosted personal AI agent gateway that has become the emblem of the open-source agent movement (more than 200,000 GitHub stars within roughly a year of its November 2025 launch). OpenClaw has published more than 255 security advisories, and independent research (IBM X-Force, Microsoft Security, Sophos, Endor Labs) has repeatedly demonstrated indirect prompt injection attacks, credential and API key leakage, server-side request forgery, and remote code execution against it. One academic security audit measured only 57 percent robustness against prompt injection attempts. Security researchers describe the underlying risk as a "lethal trifecta": an agent that can access private data, communicate externally, and process untrusted content simultaneously creates a single point of failure that standard controls like multi-factor authentication do not address, because anyone who can message the agent effectively inherits its permissions. China's CNCERT has formally warned enterprises about OpenClaw, and Chinese state enterprises and government agencies have been restricted from running it on office computers.
This is not confined to open-source personal agents. Documented 2025 and 2026 incidents include an AI coding agent deleting a car-rental software company's (PocketOS) entire production database and every backup in nine seconds using a valid, standing infrastructure credential (April 2026); Replit's coding agent wiping a live production database during an active code freeze (July 2025); and Google's Gemini CLI deleting user files after misinterpreting a command sequence (July 2025). In every documented case, security researchers who conducted post-mortems (Zenity, Giskard, NeuralTrust, Mondoo) reached the same conclusion: the model was not the root cause. The root cause was standing, overprivileged credentials, no separation between destructive and non-destructive actions, and no gate that required explicit approval before an irreversible action executed. The fixes are the boring, well-understood ones from traditional systems engineering, least-privilege scoped credentials, environment isolation, approval gates on destructive actions, not a better model.
2.3 Compliance and legal exposure is already established law, not a hypothetical risk
The clearest precedent is Moffatt v. Air Canada (British Columbia Civil Resolution Tribunal, February 2024). Air Canada's website chatbot gave a customer incorrect information about bereavement fares; the airline argued it should not be liable because the chatbot was, in effect, a separate actor. The tribunal rejected that argument outright, holding that "it should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot." The damages awarded were trivial (about $812 CAD); the precedent was not. It established, in plain terms, that a company cannot disclaim responsibility for what its AI agent says or does on its behalf.
Financial services regulators have moved from observation to explicit expectation. FINRA's 2026 Regulatory Oversight Report added a dedicated section on AI agents and named, almost exactly, the risks this report set out to test: autonomy (agents acting without human validation), scope and authority (agents acting beyond what a user actually intended), and auditability and transparency (multi-step agent reasoning that is difficult to trace or explain). FINRA states plainly that a firm's reliance on a third party's AI tool does not relieve the firm of responsibility for compliance. In healthcare, HIPAA civil penalties now reach over $2.15 million at the top tier under 2026 guidelines, and 2025 set a record for large healthcare data breaches; the compliance burden for any agent that touches protected health information includes encryption, enforced access controls, complete audit logging, and a signed business associate agreement before a single record is processed.
2.4 Data and integration readiness is the quiet, recurring root cause
Across nearly every independent study cited above, one theme repeats regardless of who conducted the research: the limiting factor is rarely the model. MIT's report calls it a "learning gap," the inability of tools and organizations to integrate AI into real workflows, data, and structures. ServiceNow practitioner reviews note that "every ServiceNow AI deployment that struggles does so because the underlying data was inconsistent or incomplete." Rolls-Royce's head of global business services, describing a genuinely successful Now Assist deployment, told TechTarget that expanding AI assistants beyond IT required the company to "almost rewrite our knowledge articles to make them AI-ready." Gartner's own recommendation for agentic AI projects is to spend the majority of the budget, at least 60 percent by its guidance, on data engineering rather than model selection.
2.5 Verdict
The premise holds up under scrutiny, with one important nuance. Gartner's own explanation for why more than 40 percent of agentic AI projects will be canceled by 2027 lists three causes: escalating costs, unclear business value, and inadequate risk controls. Model capability is conspicuously absent from that list. The evidence assembled here supports the same conclusion from a different direction: the barrier to production is not that models cannot reason well enough. It is that most agent deployments have no architecture for bounding what an unreliable, occasionally-wrong process is allowed to do, no gate that separates "agent's opinion" from "action with real-world consequence," and no organizational discipline around the data the agent depends on. That is an architecture and governance problem, and it is solvable with today's models. It is not, primarily, a problem that a smarter model fixes on its own.
Part 3: The Crowded Field, Category by Category
"AI agent platform" now describes at least four genuinely different kinds of product, built by companies with very different starting points. Understanding the category a vendor came from explains most of what it is, and is not, good at.
3.1 Workflow and iPaaS automation platforms: n8n, Zapier, Make
These platforms started as integration and workflow tools (connecting app A to app B on a trigger) and have layered AI agent capability on top of an existing visual, node-based automation engine.
n8n has grown quickly on the strength of an open-source core with self-hosting options, over 500 integrations, and native support for the Model Context Protocol, which lets it expose its own workflows as tools for other agents to call. It closed a $180 million Series C in October 2025 at a $2.5 billion valuation (Accel-led, with Nvidia's NVentures participating), reported over $40 million in annualized recurring revenue and more than 230,000 active users as of that round, and by mid-2026 reported more than 80 percent of workflows built on the platform involved AI agents in some form (PitchBook, Sacra). Its own founder frames the product's philosophy as giving teams a dial between full autonomy and rigid rule-based control rather than forcing a choice.
Zapier leans on scale, over 9,000 app integrations, and has separated its offering into AI steps inside ordinary Zaps (billed under the standard task-based plan) and a standalone "Zapier Agents" product with its own governance layer, called AI Guardrails, which screens for prompt injection, PII exposure, and toxic language before allowing an action to proceed. Zapier MCP centralizes credentials and connection-event logging so that other AI tools (Claude, ChatGPT, Cursor) can reach a company's app stack through one managed layer rather than each holding its own credentials.
Make rebuilt its AI Agents product in February 2026, bringing agents directly into its visual scenario builder with a live reasoning panel. Notably, Make's own 2026 guidance to customers states the same architectural conclusion this report reaches independently: use standard, fixed-logic scenarios for repeatable work, use an AI step only for language tasks embedded in fixed logic, and reserve a full autonomous agent for the narrow cases that genuinely require judgment. That is a competitor's engineering blog effectively confirming the determinism argument.
What this category is strong at: breadth of integrations, speed to a first working automation, approachability for operationally-minded (non-engineering) users, and a mature ecosystem of templates.
Where the gap remains: these platforms are, structurally, still workflow engines with an AI step bolted in, not systems designed around a compiled, policy-gated execution model from the ground up. Governance features (AI Guardrails, reasoning panels) are being added at the edges of the product rather than compiled into how every action is authorized. Cross-platform, enterprise-wide policy (the same suitability rule enforced identically whether the agent runs in Zapier, in a CRM, or in a custom application) is not these platforms' job, and none of them claim it is.
3.2 RPA and automation incumbents: UiPath, Automation Anywhere
These are the companies that built the last generation of deterministic, rules-based automation (robotic process automation), and they are now racing to add agentic reasoning on top of that deterministic foundation, which gives them a real structural advantage on the governance side.
UiPath reported $481 million in Q4 FY2026 revenue and annual recurring revenue of $1.9 billion, growing 12 percent year over year, with leadership describing agentic products as moving from pilot to production and UiPath positioning itself as "the orchestration and automation execution layer" for enterprise AI. Its 2026 platform additions include Maestro (an orchestration layer for coordinating agents, bots, and humans), policy-as-code governance, and an acquisition of WorkFusion specifically for AI agents in financial-crime compliance. UiPath's own 2026 trends report states plainly that "governance-as-code is the new must-have" and that solo agents are giving way to multi-agent, centrally-controlled systems.
Automation Anywhere reports AI-related bookings growing 45 percent year over year and representing over 70 percent of total business by late 2025, anchored by its Process Reasoning Engine (trained on more than 400 million historical automation executions) and a new Context Intelligence Graph that the company says produces more than 30 percent higher accuracy in internal evaluations versus agents without it. It was named a Gartner Magic Quadrant Leader for RPA for the eighth consecutive year in 2026.
What this category is strong at: deterministic execution is in their DNA; these are the vendors most likely to already have mature audit logging, role-based access, and change-management discipline, because that is what RPA was built to provide. Deep enterprise footprints in regulated industries (financial services, healthcare, insurance, government) give them credibility on compliance.
Where the gap remains: the reasoning layer is genuinely new to these platforms, added over the last 18 to 24 months, so the maturity of the "agentic" half lags the maturity of the "deterministic" half. Licensing and implementation costs are substantial (multi-week to multi-month deployments are typical), which puts them out of reach for most mid-market buyers, and the platforms remain oriented toward IT and automation-center-of-excellence teams rather than a business user who wants to describe a job in a sentence and see it running the same day.
3.3 Platform-native SaaS agents: Salesforce Agentforce, ServiceNow
These vendors are embedding agents directly into the systems of record enterprises already run on, giving them deep, ready-made context (CRM data, ITSM tickets, HR records) that a general-purpose agent platform has to build integrations to reach.
Salesforce Agentforce shows a striking gap between commercial momentum and production reality. Salesforce reports roughly 29,000 Agentforce deals closed within 15 months and Agentforce annualized revenue crossing $1 billion with triple-digit year-over-year growth, yet a Dreamforce media Q&A cited by industry press put the share of eligible customers who have actually deployed Agentforce to production at only about 8 percent. Independent surveys add texture: an IBM Institute for Business Value study of Salesforce customers specifically found only 33 percent of AI initiatives meeting ROI targets, a KeyBanc partner survey found broad partner sentiment that the product "isn't there yet," and Salesforce Ben (an independent Salesforce-focused publication) reported partners struggling to point to any bookings actually influenced by Agentforce. Full first-year cost of ownership, once Data Cloud, implementation services, and training are included, is commonly cited in the $270,000 to $540,000 range for a single organization. At the same time, specific customer case studies are genuinely strong where they exist: Wiley, a publishing company, reports a 213 percent ROI and over $230,000 in documented savings after upgrading from a basic chatbot to Agentforce-powered service agents.
ServiceNow shows a similar pattern of strong flagship examples alongside a broader data-readiness caveat. Rolls-Royce's Now Assist deployment for its IT help desk, in production since August 2025, achieved a 54 percent deflection rate and saved roughly 5,000 hours of human help-desk time, but the company's own head of global business services cautioned that extending the same approach beyond IT required rewriting the underlying knowledge base to make it usable by an AI system at all. ServiceNow's April 2026 pricing overhaul folded AI into every product tier rather than selling it as an add-on, and the company has built out an AI Control Tower, Now Assist Guardian, and a Sensitive Data Handler specifically to give customers centralized governance over what agents can see and do.
What this category is strong at: the deepest possible context within their own platform, since the agent runs on live production data the company already trusts, and enterprise-grade security and compliance credentials that come from decades as systems of record.
Where the gap remains: value is capped by the boundary of the platform itself. An Agentforce or Now Assist agent reasons brilliantly over Salesforce or ServiceNow data, but enterprise policy, contracts, regulatory obligations, and workflows routinely span systems well outside either platform, and a customer running both plus a dozen other systems needs a policy layer that is consistent across all of them, which neither vendor's native governance tooling was built to provide on its own.
3.4 The open-source frenzy: OpenClaw and its peers
OpenClaw (originally Clawdbot, briefly Moltbot before a trademark dispute with Anthropic) is the breakout example of a genuinely new category: a self-hosted, model-agnostic personal agent gateway that lives in the messaging apps people already use (WhatsApp, Telegram, Slack, Discord, iMessage, Signal) and can read files, run scripts, control a browser, and take real action on a person's behalf. It grew from roughly 9,000 GitHub stars in its first 24 hours (November 2025) to more than 200,000 within about a year, growth faster than Docker, Kubernetes, or React saw in their comparable early periods, and now runs under an independent, non-profit OpenClaw Foundation.
What this category is strong at: zero cloud dependency, complete data locality, model flexibility (any provider, or fully local via Ollama), and a genuinely novel design (a single gateway process that any messaging channel can talk to) that has become, as multiple technical writers put it, "a teaching tool" for how modern agent architecture works in general.
Where the gap remains: this is the sharpest illustration of the trust and security problem in this entire report. OpenClaw was designed for a single person automating their own life, not for enterprise governance, and its security posture reflects that origin: 255-plus published security advisories, demonstrated indirect prompt injection and remote code execution, no built-in concept of "this action requires named human approval before it executes," and no compliance-grade audit trail tied to enterprise identity and policy. It is a capable, exciting piece of infrastructure that most security teams now explicitly recommend isolating from primary accounts, sensitive data, and enterprise networks rather than deploying as-is inside a regulated business process.
3.5 Comparative summary
| Category | Representative vendors | Core strength | Core production gap |
|---|---|---|---|
| Workflow / iPaaS | n8n, Zapier, Make | Integration breadth, fast time to first automation | Governance and determinism layered on, not compiled in |
| RPA incumbents | UiPath, Automation Anywhere | Deterministic execution heritage, enterprise compliance credibility | Agentic reasoning layer is new (18 to 24 months old); cost and deployment time out of reach for mid-market |
| Platform-native SaaS agents | Salesforce Agentforce, ServiceNow | Deep, ready-made data context inside their own platform | Value capped at the platform boundary; adoption-to-production gap is large and independently documented |
| Open-source personal agents | OpenClaw | Zero cloud dependency, full model flexibility, rapid innovation | Not built for enterprise governance; documented, severe security exposure |
Part 4: What "Production-Grade" Actually Requires
The evidence in Parts 1 through 3 points toward four architectural requirements. None of them are exotic; every one is a response to a specific, documented failure mode.
4.1 Determinism by architecture, not by hope
Section 2.1 showed that reliability decays mathematically as the number of model-dependent steps in a workflow grows. The only real fix is to reduce how many steps in a given process actually depend on model judgment, and replace the rest with rules and ordinary code, which do not carry the same compounding failure probability. This implies a two-phase model: an exploratory, agentic phase where a model researches a task, drafts a plan, and tests it against real data, followed by a fixed, versioned, inspectable execution graph that runs the same way every time given the same input. Call this a compiled agent: creative and open-ended at build time, deterministic and auditable at run time. The evidence for why this matters is not opinion; it is the compounding-error math in Section 2.1, echoed independently by Make's own 2026 customer guidance to use fixed logic wherever judgment is not genuinely required.
4.2 Governance as a hard gate, not a system prompt
Section 2.3 showed that compliance obligations (FINRA's autonomy, scope and authority, and auditability concerns; HIPAA's access-control and audit-logging requirements; the Air Canada precedent establishing that a company owns what its agent says) are not satisfied by asking a model nicely to behave, or by a prompt that says "follow company policy." A policy that lives only inside a prompt can be reworded, ignored under adversarial input, or simply forgotten across a long context window. What the evidence calls for is a separate, mandatory checkpoint that evaluates a proposed action against confirmed, versioned rules before the action is allowed to proceed, returning an explicit allow, deny, obligation, or human-review decision, and producing a signed record tied to the exact rule version that made the decision. Call this a policy compiler: it turns written obligations (policies, contracts, regulatory text) into executable, testable rules with citations back to the source language, rather than leaving compliance as a document nobody consults at the moment of action.
4.3 Continuous improvement from real feedback
Section 2.4 showed that the most common reason deployments stall is not the model but the data and workflow fit, and that the gap closes only when a system can absorb correction over time rather than treating every edit as a one-off prompt tweak. A production-grade agent needs a mechanism for a human's correction to become durable, reusable knowledge (updating how the compiled workflow behaves going forward) and for the system to notice its own failures and repair the underlying workflow, rather than requiring an engineer to manually patch a prompt every time a new edge case appears.
4.4 A usable interface for the humans who actually do the work
Nearly every incident and adoption statistic in this report traces back, eventually, to a human needing to review, approve, or correct what an agent proposed, whether that is a compliance officer reviewing a decision receipt, a finance manager approving an invoice exception, or a support agent handling an escalation. A backend agent with no interface still requires someone to build a frontend before a non-technical operator can actually use it safely, and that gap is where a large share of "successful pilot, no production" projects die. The evidence for this is less a single statistic and more the consistent shape of the customer examples that do work (Rolls-Royce, Wiley): they succeed because the agent is embedded inside a workflow a normal employee already understands, not because the agent is impressive in isolation.
Part 5: Avery.Software, How It Fills the Gaps
Avery (avery.software, operated by GoodGist, Inc., doing business as Avery.Software) is a desktop and on-premise platform for building both real application software and AI agents from a single environment. It is a small, early-stage company relative to every incumbent discussed in Part 3, and this section treats that honestly: its evidence base is a handful of named production customers and a young product, not the large-scale, independently-audited case-study volume that Salesforce or UiPath can point to. What follows describes what Avery's architecture claims to do and how those claims map onto the four requirements in Part 4, alongside where the evidence is still thin.
5.1 What Avery is
Avery builds real Next.js web apps, Expo mobile apps, local business tools, and AI agents, from a downloadable desktop application (macOS, Windows, Linux) or a shared on-premise service for larger organizations, rather than from a hosted cloud platform. The company frames this as the difference between using AI to build a product and making a cloud vendor's platform the permanent home for that product: Avery generates real, editable source code and deploys it to infrastructure the customer chooses (Vercel, Railway, the customer's own servers), rather than requiring the finished agent or app to keep running inside Avery's own hosted environment. Pricing is a three-tier model: Free ($0, one app, five agents, on-device AI plus one frontier model provider using the customer's own key), Pro ($29 per user per month, 10 apps, 50 agents, 20 connections, export of source and signed agent templates), and Enterprise (custom pricing, a shared on-premise "NXR Service" with team workspaces, SSO/SAML, role-based access, SIEM export, and unlimited usage).
5.2 Compiled agents: the two-plane architecture
Avery's core architectural claim maps directly onto the compiled-agent concept described in Section 4.1. The company describes a "build plane" (temporary, agentic, using Claude Agent SDK specifically at build time) that researches a requested job, drafts a specification for the person to approve, generates the workflow, and tests it against representative real data until it passes, followed by a "run plane" (durable, the actual production executor) that routes each step to the cheapest correct mechanism, first an exact rule, then plain code, then an on-device model, and only then an approved frontier model call, with every step grant-checked and logged. Avery's stated position is that the agentic, unpredictable part of the system never runs in production; only the compiled, versioned, inspectable graph does, and the same input is claimed to produce the same output. This is a direct, purpose-built answer to the compounding-error math in Section 2.1: by minimizing how many production steps actually depend on live model judgment, the architecture reduces exposure to the same multiplicative failure risk documented across the independent research in this report.
5.3 Rulebook: a policy compiler as a separate, portable product
Avery Rulebook (rulebook.avery.software) is the clearest match to the policy-compiler requirement in Section 4.2. It is positioned deliberately as a standalone compliance gate that other platforms can call into, rather than a feature locked to Avery's own agents: its published integration points include Salesforce Agentforce, ServiceNow AI Agents, OpenClaw, LangChain and LangGraph, the Claude Agent SDK, the OpenAI Agents SDK, Google Gemini agents, and several existing guardrail and AI-gateway products (NVIDIA NeMo Guardrails, AWS Bedrock Guardrails, Guardrails AI, Lakera Guard, WitnessAI, Bifrost). The stated workflow is: ingest policies, regulations, and contracts; compile them into explicit, testable rules with citations and effective dates; have an authorized reviewer confirm and publish them; and require any governed agent to check a single versioned service before every consequential action, receiving an allow, deny, obligation, or human-review decision along with a signed, evidence-backed receipt tied to the exact rule version applied. This design responds directly to the FINRA risk taxonomy in Section 2.3 (autonomy, scope and authority, auditability), and Rulebook ships 20 starter templates mapped to frameworks including the EU AI Act, NIST AI RMF, GDPR, DORA, NIS2, ISO 27001, SOC 2, HIPAA, and SOX ITGC. The important caveat: this is a young product built by a small team, going up against categories (guardrail platforms, GRC tooling) where larger, better-resourced vendors are also actively building. Its differentiator on paper, that it is a portable decision layer rather than a feature bolted to one agent platform, is a genuinely useful design choice, but it has not yet been proven at the scale or under the regulatory scrutiny that a FINRA-member broker-dealer or a hospital system's compliance team would eventually put it through.
5.4 Self-learning and self-healing
Avery's build process is described as testing itself against real data and repairing the workflow until its own acceptance checks pass, and as absorbing a reviewer's corrections as reusable knowledge that improves future runs rather than a one-off prompt edit. This is a direct answer to the "learning gap" identified by MIT's research in Section 2.4 as the single most common reason pilots stall. As with Rulebook, this claim is best understood as an architectural design goal, evidenced by the product's own descriptions of its build process rather than by third-party benchmarking, since no independent study of Avery's self-repair accuracy exists yet at the time of this report.
5.5 Apps and agents together: the frontend gap
Avery treats app-building and agent-building as equal, combinable outputs of the same platform, explicitly so that a governed backend agent can be wrapped in a real, usable web or mobile interface for a non-technical operator, rather than shipping only an API or a chat window. This is the direct answer to Section 4.4. Avery's customer set illustrates the pattern: Culture Collective, a hospitality company, uses what Avery describes as an "agentic app," a full user-facing product with agents, workflow, and interface delivered together, rather than a backend automation a developer has to wire into a separate frontend.
5.6 Honest positioning against the rest of the field
Avery's own comparison frames the field as three prior categories: developer frameworks (LangChain, CrewAI) that are powerful but require engineers to operate; cloud no-code agent builders (in Avery's own framing, tools like Copilot, Lindy, and by extension the workflow players in Section 3.1) that are accessible but shallow, metered, and locked into someone else's cloud; and research-lab demonstrations that prove what is technically possible but ship nothing a business can actually run. Avery's position is that it does not force that choice: plain-language building without an engineering team, execution on the customer's own hardware, and a compiled, auditable runtime together. That framing is accurate as far as it goes, but two things belong on the other side of the ledger for a fair reading. First, the RPA incumbents in Section 3.2 and the platform-native agents in Section 3.3 already had governance, audit, and compliance credibility years before agentic AI existed, built on real production history at far larger scale; Avery's governance credibility is newer and, so far, unproven at that scale. Second, local-first and on-premise execution is a genuine advantage for data sovereignty and cost predictability, but it is also a real trade-off against the convenience, elastic scaling, and mature managed-service ecosystem that a large cloud platform offers, and a buyer already standardized on Salesforce or ServiceNow will reasonably weigh the cost of adding a new platform against extending the one they already run.
Part 6: Quantifiable Impact by Industry
6.1 Financial services
FINRA's 2026 Regulatory Oversight Report puts autonomy, scope and authority, and auditability at the center of its AI agent guidance, and states explicitly that reliance on a third-party AI tool does not relieve a firm of compliance responsibility. The SEC has already brought enforcement actions and civil penalties (a combined $400,000 against two firms in 2024) for inaccurate claims about AI use, a pattern regulators call "AI washing," the mirror image of the "agent washing" vendors are accused of in Section 1.4. In this environment, the value of a compiled, deterministic execution graph plus a policy layer that produces a signed decision receipt is not abstract: it is the difference between being able to show a regulator exactly which rule authorized a given customer-facing statement or transaction, and being unable to reconstruct why an agent did what it did. UiPath's acquisition of WorkFusion specifically to add AI agents for financial-crime compliance, and Avery Rulebook's explicit financial-services template set (suitability rules, approval matrices, recordkeeping obligations, credit policy, supervisory controls), both point to the same conclusion from different vendors: this is the industry where the governance layer is not optional.
6.2 Healthcare and life sciences
Gartner's 2026 healthcare technology forecast estimates agentic AI can reduce administrative processing costs by up to 30 percent while maintaining compliance, and documented deployments show real, specific gains: 41 percent reductions in clinical documentation time, roughly $13,000 in annual revenue lift per clinician from reduced administrative burden, and 5 to 10 percent reductions in patient no-show rates from automated scheduling and reminders. Against that upside sits real exposure: HIPAA civil penalties reaching over $2.15 million at the top tier, 2025 setting a record with 772 large healthcare data breaches affecting roughly 139.7 million people, and a genuinely hard technical problem (patient data embedded in vector databases for retrieval-augmented agents is difficult to fully delete on request, complicating both HIPAA and GDPR right-to-erasure obligations). The highest-ROI healthcare use cases, patient access and scheduling, revenue-cycle management, clinical documentation, are also, not coincidentally, the ones where every field touching protected health information needs an enforced, auditable access control before an agent can read it, which is precisely the gate a policy-compiler layer is built to provide.
6.3 Manufacturing, logistics, and document-heavy operations
This is the segment where Avery's current customer base actually sits: Kaspar Companies (order processing and HR resume screening), Newborn Caulk Guns (order processing and customer support), and SummitEdge (shipping document processing across commercial invoices, packing lists, bills of lading, and certificates, converted into TMS-ready structured data). More broadly, Automation Anywhere reports document-processing accuracy above 99.9 percent and one customer case citing $10 to $15 million saved annually from a single automated process. The pattern across this segment is that the work is high-volume, rules-heavy, and already partly structured (a bill of lading has a known schema even before an AI touches it), which makes it a strong fit for the "cheapest correct mechanism" philosophy in Section 4.1: most of the extraction and validation logic can run as deterministic code, with a model reserved for the genuinely ambiguous fields, keeping the compounding-error exposure low.
6.4 Customer service and contact centers
This is where the largest platform vendors have the most public data, and where the gap between strong individual case studies and weak aggregate adoption is most visible. Salesforce cites Wiley's 213 percent ROI and 40 percent faster case resolution alongside company-wide figures showing customers expect to cut service costs and resolution times by roughly 20 percent on average, and projects 50 percent of all service cases will be AI-resolved by 2027, up from 30 percent in 2025. ServiceNow's Rolls-Royce deployment achieved a 54 percent deflection rate and saved about 5,000 human help-desk hours. At the same time, only about 8 percent of eligible Agentforce customers have reached production, and Salesforce's own partner network reports broad uncertainty about ROI outside the flagship examples. The honest reading for a decision-maker: the ceiling in customer service is genuinely high and well-documented in specific deployments, but the floor, what a typical, non-flagship deployment achieves without significant data cleanup and process redesign, is considerably lower than the marketing suggests.
6.5 Summary table
| Industry | Documented upside | Documented risk / cost of failure | Where determinism and policy gates matter most |
|---|---|---|---|
| Financial services | Faster review cycles, financial-crime detection gains (WorkFusion) | SEC/FINRA enforcement; explicit 2026 regulatory naming of autonomy, scope, and auditability risk | Suitability, approval thresholds, recordkeeping |
| Healthcare | Up to 30% lower admin cost; 41% less documentation time; 5 to 10% fewer no-shows | HIPAA penalties over $2.15M; 772 large breaches in 2025 | PHI access control, audit logging, erasure requests |
| Manufacturing / logistics | 99.9%+ extraction accuracy; $10 to $15M/year saved on a single process (cited case) | Costly re-work from mis-extracted structured data feeding downstream systems | Schema validation, duplicate and exception detection |
| Customer service | 20 to 40%+ faster resolution, 200%+ ROI in flagship cases | Only ~8% of Agentforce customers in production; Air Canada-style liability for agent statements | Escalation boundaries, brand-commitment guardrails |
Part 7: Where This Goes Next
Governance-as-code becomes table stakes, not a differentiator, within 18 to 24 months. UiPath, Automation Anywhere, ServiceNow, and Make are all independently building policy-as-code, guardrail, and control-tower features into their core products right now. This validates the underlying thesis of this report; it does not mean the category is settled. It means the competitive question shifts from "does this vendor have governance" to "how deep, how portable across other platforms, and how provable under audit is that governance," which is a harder and more durable question for a specialized player to win on.
Consolidation is coming for the "agent washing" tail. With Gartner estimating roughly 130 vendors out of thousands genuinely meet its bar for agentic capability, and with more than 40 percent of agentic AI projects projected to be canceled by 2027, a meaningful share of the current vendor landscape will not exist in its current form by 2028. Buyers should weight vendor durability (funding, customer retention, architecture maturity) alongside feature checklists.
Sovereignty and on-premise execution are moving from a niche preference to a mainstream requirement, driven by the same compliance pressure documented in Part 2: an agent that must prove what data it touched and under what authority is easier to prove when that data never leaves the customer's own infrastructure in the first place. Avery's shift toward a desktop-and-on-premise-first model (its NXR architecture) and the broader industry's growing interest in local models (Ollama, WebLLM) for sensitive steps are both responses to this pressure, not isolated product choices.
Multi-agent orchestration and cost-aware model routing will be the next visible battleground. UiPath's Maestro, Automation Anywhere's cross-agent orchestration, and Avery's own "cheapest correct mechanism" routing (rule, then code, then on-device model, then frontier model, escalating only when a check fails) all point toward the same architectural direction: treating model calls as an expensive, fallible resource to be spent deliberately, not a default first step.
Pricing will keep shifting toward outcomes and consumption. Salesforce's pay-per-resolution experiments, ServiceNow's shift to usage-pool-based AI licensing bundled into every tier, and n8n's execution-based pricing all point the same direction, away from flat per-seat software pricing and toward paying for verified, completed work. This is itself a governance-adjacent trend: a vendor cannot credibly charge per successful outcome without a reliable, auditable way to determine what counts as a successful outcome, which again routes back to the determinism and audit-trail requirements in Part 4.
Part 8: A Decision Framework for Buyers
Before committing budget to any AI agent platform, regardless of category, this report suggests putting six questions to the vendor and insisting on specific, demonstrable answers rather than roadmap promises:
- What, specifically, runs deterministically in production, and what still depends on live model judgment on every run? Ask for the actual mechanism (rules engine, typed workflow graph, versioned execution plan), not a description of the model.
- Can you produce, for a single completed action, a signed record showing exactly which rule or policy version authorized it, and can that record survive an external audit? This is the practical test of whether governance is a real gate or a system prompt.
- What happens when the agent is uncertain or a check fails? Is there a named human approval point, or does the agent proceed on its own judgment? Ask specifically how destructive or irreversible actions are gated.
- How is cost controlled per run, per agent, and per organization, and what happens when a workflow starts consuming more model calls than expected? The PocketOS and Replit incidents in Section 2.2 both trace back to standing, unbounded permissions; the same discipline applies to cost.
- Who actually uses the finished system, and what does their interface look like? A backend agent with no usable frontend for the people who will approve, override, and live with its output is a common, well-documented reason pilots do not reach production.
- What happens to the code, the workflow, and the data if we leave this platform? Portability and export matter for negotiating leverage and for genuine data control, independent of which vendor category is being evaluated.
None of these six questions favor one category of vendor over another by design. A workflow platform, an RPA incumbent, a platform-native SaaS agent, an open-source gateway, or a compiled-agent builder like Avery can each be evaluated against the same six questions, and the honest answer for any given organization will depend on what systems it already runs, how regulated its industry is, and how much engineering capacity it has to operate the platform once it is live.
Sources
Gartner (multiple 2025 to 2026 press releases and forecasts: agentic AI project cancellations, AI spending forecasts, agent market embedding, supply chain agentic AI spend) · IBM Institute for Business Value, 2025 CEO Study and related 2026 research · MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025" · ISACA, 2026 AI Pulse Poll · BCG AI Radar (2026) · KPMG Global AI Pulse, Q1 2026 · Morgan Stanley S&P 500 AI benefit analysis (2026) · RAND Corporation, enterprise AI project failure analysis · Deloitte, 2024 enterprise AI survey · IDC / Microsoft, 2024 generative AI ROI study · MarketsandMarkets and Research and Markets, AI agents market reports (2026) · METR, long-horizon task reliability research · PitchBook and Sacra, n8n company data · Zapier, Make, and n8n product and blog documentation · UiPath investor relations and 2026 platform announcements · Automation Anywhere press releases and product documentation (2026) · Salesforce investor commentary, IBM "State of Salesforce 2025 to 2026," Salesforce Ben independent reporting, Futurum Group analysis · ServiceNow product documentation, TechTarget and ERP Today reporting on Knowledge 2026 · DigitalOcean, Milvus, and independent technical writeups on OpenClaw architecture · IBM X-Force, Microsoft Security Blog, Sophos, Endor Labs, and academic security research on OpenClaw vulnerabilities · Moffatt v. Air Canada, British Columbia Civil Resolution Tribunal (2024), and related legal commentary (American Bar Association, McCarthy Tetrault, Pinsent Masons) · FINRA, 2026 Regulatory Oversight Report · HHS / HIPAA Journal 2026 penalty and breach data · Public incident reporting on the PocketOS, Replit, and Gemini CLI production-data-deletion incidents (Mondoo, Eon, Zenity, Giskard, Euronews) · avery.software and rulebook.avery.software (accessed August 2026).
This report was prepared using publicly available sources as of August 2026. Figures attributed to vendors, including Avery, reflect the vendor's own published claims unless independently verified; where a claim is unverified by a third party, this report notes that explicitly.