How banks can better measure AI ROI and business impact

Banks face three interconnected challenges as they move from experimenting with AI to integrating it into their businesses: making AI adoption measurable, scalable and accountable. The measurement challenge is coming to the fore as rising costs force executives to explain what their growing portfolio of AI initiatives actually produces.

Processing Content

The early objective of generative AI adoption was to encourage experimentation. Banks tracked how many employees had access, how often they used the tools and how many use cases entered the pipeline. Those numbers showed that adoption was taking place. They did not necessarily show that it was creating value.

PwC’s 2026 Financial Services Workforce AI Survey released last week illustrates the tension. Seventy-seven percent of the 1,004 executives at U.S. financial institutions said most AI investments are not delivering measurable ROI. Yet firms have been actively pursuing benefits of AI. Over the past year, 49% focused their AI workforce efforts on improving productivity, 48% on reducing time spent on routine work and 46% on integrating AI into day-to-day workflows.  

These findings should not be read to imply that productivity and workflow improvements do not matter. They suggest something more important: Banks need to distinguish progress toward value from financial value itself.

That is the measurement challenge banks now face. The problem is banks often make one of two mistakes:

The first is treating activity as value. Employee access, prompts, training completion and the number of pilots all indicate that people are experimenting with AI. They do not establish that work has improved, business outcomes have changed or the banks have earned a financial return.

The second mistake is demanding a fully attributable financial return too early. A pilot designed to test whether an AI tool can complete a task safely cannot yet be judged by incremental revenue or enterprise cost reduction. Applying mature ROI tests to an immature project may cause management to abandon promising initiatives before they have had an opportunity to progress.

Banks should neither confuse activity with value nor demand mature financial outcomes from immature projects. Instead, they should measure whether each initiative is producing the results appropriate to its stage of development. (Of course they will also need to prove that these initiatives are promising candidates that will scale and that they can translate that into financial return eventually. We’ll discuss these in more detail in the second part of this series.)

What should banks measure?

AI adoption initiatives generally progress through four stages. Each stage has a different objective and therefore requires different evidence.

Experimentation Activity measures are relevant at this stage. A bank needs to know whether employees have access, are using the application and have received appropriate training. But access and usage are inputs, not results. The more important question is whether the application can perform its intended task effectively and within the bank’s risk parameters. A heavily used pilot that produces unreliable output is not progressing. Neither is one that performs well but fails security, privacy or compliance requirements. Experimentation should establish technical feasibility, user acceptance, output quality and control effectiveness. 

Workflow improvement Once an application enters production, measurement should shift from the user to the workflow. Does it reduce task time, increase productivity or improve accuracy? Does it eliminate work or simply shift it elsewhere in the workflow? “Time saved” is useful but incomplete. Time creates value only when it can be redeployed to increase capacity, improve service or reduce required resources. The correct baseline is the process the application replaced, not an idealized version of what AI might achieve.

Business outcomes Workflow gains matter only if they translate into outcomes the business cares about. In customer service, that may mean higher resolution rates, shorter wait times or improved satisfaction. In fraud detection, it may mean fewer losses without excessive false positives. In sales, it may mean higher conversion rates or stronger customer relationships. This stage connects operational change to business value and begins to test scalability. A tool that improves efficiency in isolation and does not move customer, revenue or risk metrics may have limited strategic value.

Financial outcomes Financial measurement becomes more credible once the earlier links have been established. The question is whether improved business outcomes generate incremental revenue, reduced costs or enhanced capital efficiency. Those benefits must be assessed against the full cost of the adoption initiative in question. 

That last requirement is becoming harder to ignore. EY recently found that 82% of senior leaders at organizations investing in AI are concerned about token usage and related costs. Yet only 64% said their organizations actively monitor token usage and have clear budget guardrails. 

This is why measuring AI value is more complicated than assigning a dollar amount to every hour reportedly saved, a challenge we’ll examine in greater detail in the third article in this series.

Why does this approach improve decision-making?

Matching the measure to the stage gives management a more useful basis for decision-making. It helps executives determine which initiatives should continue experimenting, move into production, scale across the business, be redesigned or stopped.

Management needs evidence that distinguishes initiatives that should be funded and expanded from those that should be changed or discontinued. A collection of usage statistics and estimates of hours saved alone is not enough as an AI initiative matures. They need evidence that these initial improvements can be turned into business and financial outcomes. 

Shareholders also need to know whether rapidly growing AI expenditures are becoming productive investments or simply another layer of technology spending. AI may transform banking over time, but that does not mean every initiative deserves continued funding.

The purpose of measurement is therefore not only to prove that past spending was justified but also to help management to decide what to do next. 

How can banks better connect the dots?

The four stages are not four disconnected score cards. Together they form an evidence chain that banks should build: Usage leads to effective task performance. Effective task performance improves a redesigned workflow. The improved workflow enhances a business outcome. The business outcome creates financial value.

Banks will not always be able to calculate a precise return when an AI initiative begins. They should still be able to establish what success means at the current stage, what evidence would validate progress to the next stage and what conditions would cause the initiative to be reconsidered.

Not every initiative will complete that journey. Experimentation is supposed to identify failures as well as success. What matters is that management can distinguish between the two and allocate resources accordingly. 

Banks do not need to stop experimenting until every AI investment can be reduced to an immediate ROI calculation. They do not need to stop presenting activity as evidence that value has been created. 

The real objective is to determine whether each initiative is progressing credibly toward financial value and whether the bank is learning enough at every stage to decide what should happen next.  

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *