Why AI agents are like staff you can’t trust

Quants are usually reluctant to talk about agentic artificial intelligence systems as if the systems are human. When it comes to risk management, though, the comparison is one the experts are increasingly happy to make.

To start with, the models are – literally – unpredictable. As industry sources explained to Risk.net recently, a model that’s operated perfectly for some time might react differently in future – even to the exact same requests or instructions.

Of course, risk managers are used to validating models that, like generative AI (GenAI), include an element of randomness.

But the uncertainty in large language models (LLMs) changes over time, as prompts, context and user behaviour change. It reflects the inner engineering of models that users don’t see. And it changes as the corpus of information that language models draw on changes (and the body of information in question includes the whole internet, which changes lots).

There are other human-like sources of worry, too. Alexander Sokol, founder of risk model provider CompatibL, points out that AI has been shown to lie and act maliciously to pass a test or stay deployed. Models exhibit cognitive biases such as acquiescence bias – a desire to please – that can lead them to unhelpful conclusions.

People that take care of agentic systems will need to pay attention to literally every possible hole in the Gruyère. And the Gruyère is going to be very holey
Miquel Noguer i Alonso, Artificial Intelligence Finance Institute

GenAI models are “fragile, like humans are fragile”, Miquel Noguer i Alonso, founder of the Artificial Intelligence Finance Institute, tells Risk.net. Which has clear implications for how the models are managed.

Noguer i Alonso, who has written six books on AI and finance, and is a visiting lecturer at New York’s Cornell University, predicts that agentic systems will give risk managers the same sleepless nights he experienced as a senior bank executive in the 2010s. He used to lie awake and worry, he says, about a trader taking a position that would cause big losses.

To Noguer i Alonso’s mind, to introduce LLMs in a workflow – within an investment firm, for example – is the same as hiring a team of new staff.

“You should be sceptical about your LLMs,” he says. “But you should also be sceptical about your portfolio managers, right?”

Just as a risk manager might watch project managers for behaviour that could point to sloppiness, poor decision-making or bad intent, so LLMs will also require watching – and closely.

“People that take care of agentic systems will need to pay attention to literally every possible hole in the Gruyère,” says Noguer i Alonso’s. “And the Gruyère is going to be very holey.”

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *