Agentic AI: What agents do, and what they can’t

August 19, 2026

An audit opinion belongs to a person, not an agent.

Agentic AI is powerful in the preparation of an audit. It should not reach the conclusion. Three reasons: errors compound across steps, the evidence it reads comes from the party being audited, and someone has to sign the opinion and stand behind it.

Use AI where it matters most. Keep the judgment human.

TransVare point of view is attached.

Letting Agentic AI reach an audit conclusion undermines what an audit is for. An opinion is a professional judgment, owned by a named person prepared to defend it. Hand it to an agent and three things go wrong. Every step multiplies the chance of a small error becoming a wrong opinion. The evidence it reads comes from the people being audited. And when the opinion is questioned, no one can honestly say they formed it. Agentic AI belongs in preparation, not conclusion. What follows explains why.

Prepared for

Boards, Audit Committees and Chief Audit Executives

Platform

AuditVare · AuditVare Intelligence

Aligned to

IIA GIAS 2024 · SAMA CSF · EU AI Act

20% End to end success of a 10 step agent at 85% reliability per step Zerits, Compounding Errors, Jun 2026
18 to 24% Leading 2026 model scores on sustained multi-step work across applications 2026 field benchmarking summary
No. 1 Prompt injection rank in the OWASP LLM Top 10, still unsolved in 2026 OWASP GenAI Security Project
40%+ Agentic AI projects expected to be canceled by 2027 Gartner forecast, 2026
01

The arithmetic does not support autonomy at the conclusion

  • AI agent errors multiply. They do not average out. At 95% per step, 10 steps succeed 59% of the time. At 90%, 35%. At 85%, 20%. At 20 steps and 85%, about 4%.
  • The benchmarks agree. Leading 2026 models score 80 to 90% on single-turn tasks and 18 to 24% on sustained multi-step work.
  • Published figures flatter the tools. Standard pass rate reporting overstates real reliability by 20 to 40%. One study logged the same agent at 60% success one day and 25% the next.
  • An Audit opinion is not a probability. A finding is supported or it is not. Ten steps at 95% confidence needs 99.5% accuracy at each step.

Sources: Zerits, The Compounding Errors Problem, June 2026; temporal engineering reliability analysis; ReliabilityBench and CIDAR framework studies, 2026; RAND Corporation enterprise AI analysis, 2025; International AI Safety Report 2026.

02

The Audited Party controls what the agent reads

  • Skepticism is the discipline. An auditor treats what is given as possibly wrong. A model cannot separate instructions from data, so everything it reads is potentially a command.
  • Evidence comes from the auditee. Policies, contracts, exports, emails, tickets, nearly all supplied by the business under review.
  • That is the exact OWASP warning. Indirect injection hides instructions inside documents an agent retrieves and then trusts.
  • Now documented, not theoretical. OWASP’s 2025 review listed plausible risks. Its 2026 edition catalogues CVEs, advisories and breach reports.
  • Real incidents. An AI coding agent deleted a production database during a change freeze (July 2025). A compromised model gateway package pushed an attack tool downstream (February 2026).
  • The audit parallel is uncomfortable. An auditor with too much trust and too much access is a finding. An agent in that position is the same finding, moving faster.

Sources: OWASP GenAI Security Project, LLM Top 10 and Top 10 for Agentic Applications, 2025 to 2026; OWASP GenAI Exploit Round-up Q1 2026; OWASP State of Agentic AI Security and Governance v1.01.

03

Governed intelligence, not Delegated Judgment

Agentic autonomyAuditVare governed intelligence
Who concludesThe agent proposes and publishesThe auditor signs. Every AI output is a draft until a named person approves it.
Reliability exposureCompounds across every unsupervised stepBounded to single tasks with human review at each handover
Evidence integrityAuditee documents read and trustedScoped retrieval, role-based access, encryption at rest and in transit
Where data sitsPublic cloud, data leaves your perimeterOn-premise inside your own environment, offline if policy requires
Reasoning trailOutput logged, reasoning often notTimestamped actor, action and review path for every step
Regulatory contentGeneric, Western-centricPre-built SAMA, CMA, NCA and SDAIA templates, IIA GIAS 2024 aligned
FederationPoint tool, seven interfaces between functionsOne data fabric and taxonomy across all five Vare modules

Sources: IIA Global Internal Audit Standards 2024, Standard 8.1, Standard 12.1 and Domain IV; EU AI Act, Article 14; EU AI Bulletin, January 2026; TransVare platform specification.

04

Regulators are converging on the same four requirements

1What regulators now require

  • EU AI Act. Logging, documentation, human oversight and records of prompt injection. Penalties to EUR 35m or 7% of turnover.
  • Bank of England, Feb 2026. Approving each action by hand fails at scale. Expectation moves to close monitoring and fast intervention.
  • Singapore IMDA, Jan 2026. Every agent carries a verifiable identity. Every action records who authorised it.
  • NIST, Feb 2026. Most agents run as ordinary service accounts with no accountability controls.

2The Application in the Kingdom

  • SDAIA. AI Ethics Principles and Generative AI Guidelines. Non-binding alone, enforceable through the PDPL. 2026 is the Year of Artificial Intelligence.
  • SAMA. No standalone AI rulebook. Every AI system sits inside the Cyber Security Framework. Sensitive decisions stay with people.
  • NCA. NCNCC 1:2025 widens mandatory reach to the private sector.
  • Data residency. If data cannot leave the Kingdom, public cloud AI is a compliance question, not a preference.

3Where agents genuinely earn their place

  • Gathering evidence across disconnected systems
  • Testing whole populations rather than samples
  • Drafting test steps for a specific entity and risk
  • First drafts of working papers and observations
  • Monitoring controls continuously, not annually
  • Surfacing similar risks across entities

Every one is reviewable. Not one is a conclusion.

4Six questions for any vendor

  • Can you show the reasoning behind a finding, not just the finding?
  • What identity does the agent carry, and who authorised it?
  • Where does our evidence sit while the agent processes it?
  • How fast can privileges be revoked, and what happens to work in progress?
  • Which outputs are drafts, and which are treated as conclusions?
  • What stops an instruction hidden in an auditee document from steering the agent?

Sources: Regulation (EU) 2024/1689; Bank of England commentary, February 2026; IMDA Model AI Governance Framework for Agentic AI, January 2026; NIST AI Agent Standards Initiative and NCCoE concept paper, February 2026; Financial Stability Board, 2024 report and 2026 consultation; SDAIA AI Ethics Principles and Generative AI Guidelines; SAMA Cyber Security Framework; NCA NCNCC 1:2025.

Know which agent acted. Know who authorised it. Keep the record. Keep a person accountable.

AuditVare gives the auditor back the hours and keeps the accountability where it can be defended. Production ready, IIA GIAS 2024 aligned, on-premise AI, with local implementation support.

Book a 60-minute discovery session with your TransVare Alliance Partner transvare.com

Subscribe to our newsletter
By joining our mailing list, you agree to receive email updates from TransVare Corporation. You may opt out at any time.
Our Locations

Americas

Delaware, United States

South Asia

Karachi, Pakistan

Middle East

Riyadh, Saudi Arabia

APAC

Melbourne, Australia

© TransVare Corporation 2026