Governing AI execution as institutional infrastructure
Key takeaways
AI governance is operational governance: As agentic tools move into workflow execution, control is shifting from model approval to authority, supervision and recovery within live processes
Inventory must become more dynamic: AI inventories must move from periodic registers to live visibility to support firms in detecting use, classifying risk and mapping changes over time to fully capture co-pilots, local agents and user-built workflows
Accountability must follow the chain of action: Ownership and supervision of agents is key for responsibility and accountability, especially when workflows cross business, risk, compliance and technology silos
Controls must operate at run time: Access rights, telemetry, policy enforcement, prompt and output testing, source attribution and audit records must sit in the workflow, not only in committee materials
Do not outsource accountability: Leaders across the AI, compliance and risk functions foresee an emerging operating model that places greater value on human skills to interrogate outputs, understand processes and control automation
Artificial intelligence adoption across financial institutions is no longer an experimental question, nor is it limited by whether tools can produce useful outputs. The barriers that matter now are not technical. They are about control, accountability and the defensibility of outcomes.
Panellists on the Built to scale, built to fail? Governing AI execution roundtable emphasised that the true constraint now is whether an institution can see, govern and defend the way those outputs are used once AI is embedded into operating processes. “Control is actually the biggest issue … and the binding constraint with AI,” a senior risk and compliance research leader highlighted.
From model oversight to execution control
Agentic AI is driving a shift from overseeing how models generate outputs to controlling how autonomous systems act on those outputs. Unlike earlier analytical models, agentic systems can call tools, route work, escalate exceptions and trigger actions with less visible human intervention. That makes the governance question less about isolated model performance and more about the quality of the institutional operating layer around AI.
“Lack of visibility is the biggest risk,” one participant noted. It is key to map who authorised the activity, what data was used, what authority was exercised and what evidence was preserved, but most importantly what happens when the workflow fails.
As regulatory guidelines and standards for AI governance continue to evolve – under the US Federal Reserve’s SR 26-2 (formerly SR 11-7), the Office of the Comptroller of the Currency Heightened Standards, the Digital Operational Resilience Act and the European Union AI Act – institutions will be asked to explain how an automated decision was formed, what authority was exercised and whether it can be defended after the fact. The supervisory bar is shifting from evidence of controls to evidence of decision quality.
Inventory as operating risk
While existing model risk frameworks remain relevant, AI governance and risk practitioners questioned whether a traditional model inventory could fully capture outputs across co-pilots, local agents and user-configured workflows. The concern was not only if a formal use case had been approved, but whether second and third lines had access and visibility into what is being performed within organisational and functional silos.
One senior AI governance leader described a widening gap between formal governance and “what people are actually doing [behind] their desks when they get access to these tools”.
That distinction matters. A tool may begin as research support, then become a recurring workflow, then shape an operational decision. At that point, practitioners noted, the boundary between sandbox, end-user computing and production process becomes less tidy than the governance taxonomy assumes.
The emerging answer was not a single inventory but a tiered, risk-based view. General prompts, research and low-risk productivity tasks may need one level of control. Customer interaction, regulated reporting, trading support, financial crime disposition or any workflow with execution rights requires a different standard. Here, identity, ownership, application IDs, change controls, usage logs and escalation paths become vital. The inventory must transform from merely being a register to more of an operating control mechanism.
Control is actually the biggest issue … and the binding constraint with AI
Accountability must map to execution
The accountability debate among attendees exposed a deeper organisational issue. Senior risk and AI leaders argued that AI should not be governed simply as an application-based software. It must be considered a decision actor inside a process, especially once agents receive entitlements, call tools or act across applications. One attendee asked whether an agent should be viewed as a “digital person”, and onboarded, permissioned, monitored and mapped to an owner.
While a human employee has incentives, training and judgement, and faces consequences, an agent has instructions, data access, tool permissions and logs. Treating agents like digital employees may help firms assign ownership and map application access, but it does not solve the harder problem of cross-functional workflows, attendees highlighted. A process that touches know‑your-customer data, fraud and anti–money laundering signals, product information and client communication may not fail inside a single application or under a single accountable, executive employee.
One senior practitioner emphasised that “you cannot outsource accountability”. It requires an all‑encompassing approach. AI process builders, coders, process owners, risk functions and technology teams may all need defined responsibilities. Moreover, forcing accountability onto a single owner often places it on those least able to control the outcome. The question is not whether the buck stops somewhere, attendees noted. It is whether the chain of responsibility is credible enough before an incident occurs.
Controls must move to the point of decision
Attendees underscored the limits of policy-based governance. Policy manuals, use case approvals and committee reviews remain necessary, but they do not control an agent at the point it retrieves data, uses a tool, generates a recommendation or routes an exception. For AI execution, senior practitioners emphasised, controls need to live inside the workflow.
This transforms the AI control stack. Participants pointed to data access restrictions, prompt-injection testing and accuracy testing, source attribution, deterministic verification layers, output review, tool‑calling limits and logs that can reconstruct the basis for a decision.
One attendee drew the comparison with end‑user computing, in which poorly written spreadsheet macro could create risk, but was largely deterministic. Generative systems introduce stochastic behaviour into everyday work, making verification and evidence essential. Designing verifiable processes up front rather than measuring outcomes after the fact is key.
Participants drew a clear distinction between AI used for support and AI used for execution. Tools that help reduce false positives in regulatory change monitoring were viewed as more manageable when a human remained in the loop.
Risk appetite, data and telemetry
Against this backdrop, participants also highlighted that, in high-risk workflows, the right answer might be to make the process human-driven even if AI can technically perform the task. However, AI leaders in attendance also challenged the premise that AI only adds risk. It may transmute or reduce risk versus existing residual human error.
Data sits at the centre of this discussion. Senior practitioners laid focus not only on training data but on machine data, such as logs, access records, prompts, outputs, connector activity, entitlement changes and the telemetry needed to identify drift in behaviour. But institutions are lacking the tooling to operationalise this today. Without this layer, governance committees may know that an AI programme exists, but remain unable to evidence how it is behaving.
The same issue applies to cost and capacity. Token usage, model selection and infrastructure spend are not purely technology metrics; they are indicators of scale, cost dynamics, adoption and emerging concentration, attendees noted. Undoubtedly, firms with a mature telemetry layer will be better placed to distinguish experimentation from embedded operational resilience.
The human in the AI governance loop
If an agentic workflow fails mid-execution, a firm needs to know whether the system stops, rolls back, escalates, reroutes to a human or allows another workflow to intervene. The answer must be tested, not assumed. Senior practitioners emphasised that the audit record must show not just that a control existed, but how the decision was formed.
One attendee compared the cultural shift to cyber awareness, arguing that employees need to be trained to “think before they automate”.
Another said that the most valuable users might be those who can own their work product while using AI to challenge, accelerate or deepen it.
Overall, practitioners noted that senior management may encourage adoption while risk teams urge restraint, creating structural misalignment. However, the solution is not to oppose adoption, but to make approved tools easier to use and to create review points where repeated tasks become governed processes.
Make AI execution observable, defensible and recoverable
AI, risk and governance practitioners are clear that AI execution will not be governed by adding another approval layer to existing processes. It requires a clearer operating model for inventory, data, authority, telemetry, ownership and recovery. Model governance remains part of that architecture but it is no longer sufficient on its own.
The firms most likely to scale AI safely will be those that can answer the simple yet essential questions: What is running? What is it allowed to do? Whose authority is it exercising? What data did it use? What evidence exists? Who owns the process and what happens when it fails?
The issue is not whether AI should be used. It is whether the institution can make AI execution observable, defensible and recoverable before scale turns a local weakness into an enterprise control failure.
This report is drawn from discussions at an exclusive Risk.net AI governance roundtable, Built to scale, built to fail? Governing AI execution, convened in collaboration with Cisco and held in New York in June 2026.