"Agentic work units" measure activity. They do not measure whether the refund cleared, the quote went out, or the case closed correctly. The Completion Index measures that, on the benchmark Salesforce built for exactly this purpose.
Published by Salesforce AI Research in June 2025 and accepted by TMLR in 2026, CRMArena-Pro is the only expert-validated public benchmark for agents on realistic CRM business flows: 19 tasks across sales, service, and CPQ, in B2B and B2C environments, spanning database querying, text reasoning, workflow execution, and policy compliance.
| Condition | Best agent accuracy | What it means |
|---|---|---|
| Single-turn tasks | ~58% | Even simple, one-shot business tasks fail about 4 times in 10 |
| Multi-turn tasks | ~35% | On realistic back-and-forth flows, roughly 2 of 3 fail to complete correctly |
| Confidentiality awareness | Near zero without prompting | Agents volunteered sensitive data unless explicitly instructed not to |
Source: Huang et al., CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions, Salesforce AI Research, arXiv:2505.18878.
Every flow that doesn't finish gets redone by a human, escalated, or quietly dropped. At the published 35% completion rate, that is most of them. Put in your numbers.
Illustrative estimate. Baseline completion rate reflects Salesforce AI Research's published CRMArena-Pro multi-turn results (~35%); your workload will differ. The calculator makes no claim about MOTA's performance. Request our methodology and internal results via early access.
In our internal runs on CRMArena-Pro task environments, a frontier LLM paired with MOTA's planning layer substantially closes the multi-turn gap. These are our own evaluations rather than Salesforce's figures, and we share the run logs so you can check us.
The steepest published drop, 58% to 35%, is where deterministic planning helps most. Our internal runs show the largest gains here.
The same intent produces the same plan and the same outcome. Variance between runs is a design target we hold ourselves to.
Every benchmark run produces the same human-readable audit trail a production deployment would.
Request full results & methodology Shared with early-access applicants · reproducible
The market has converged on the three guarantees MOTA is built around: reliably, durably, and provably complete. Salesforce itself now agrees; its Agent Script beta adds deterministic control to Agentforce after customers struggled in critical workflows.
Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027, citing costs, unclear value, and inadequate risk controls. The enterprise bar remains the straight-through processing standard: flows that finish without intervention.
Forrester expects fewer than 15% of firms to enable agentic features in 2026, and notes that deterministic automation remains "the backbone of reliability and compliance". Durable, resumable execution is becoming the price of admission.
Gartner's AI TRiSM framework, extended in 2026 with its first Market Guide for Guardian Agents, makes runtime traceability and policy enforcement a formal requirement for trustworthy agents, and the EU AI Act's high-risk obligations arrive from August 2026.
Sources: Gartner press release, June 2025 · Forrester "Predictions 2026: Automation at the Crossroads," Nov 2025 · Gartner AI TRiSM / Guardian Agents Market Guide, Feb 2026 · CIO.com on the Agentforce recalibration, Jan 2026. Full citations in Research and the trade press table.