The Limitations and Agentic Finance

Harrison Mann
,
Head of Growth

Over the past few weeks we’ve talked a bit about Embedded FX, and what we think it means. While we are obviously big believers in automation, our point of view is somewhat different from what is often pitched in market. A lot of times when people talk about the future of payments, they gesture at fully-agentic treasury operations, systems that transact and adapt with little to no human interaction.
That seems like a kind of logical end point of the current technological curve, until you zoom the camera a little.
In a system designed this way, who makes decisions, and who takes responsibility for them? The machine? It seems unlikely that regulators are going to accept that. It seems unlikely your clients will if something goes wrong. No, you’re responsible, and for that reason you really don’t want agents running around everywhere making decisions. What you want are systems that are robust enough to stand up to uncertainty, while remaining steerable enough to stand up to audit.
Before we dig into that, here is a quick reminder of how we think about the different types of agentic systems, you can read the full article here.
The Four Basic Levels of Agentic Taxonomy
Rules-based automation. These are just rote decision trees: Condition met, action executed. They’ve existed for decades. The easiest example to reach for is software programmed to “issue payroll on the 5th of the month.”
Rules-based with exception handling. This is where a big portion of “intelligent” payments infrastructure currently lives. You’re still dealing with rules-driven software: If this, then that. The difference is, you’ve got a more elaborate decision tree. To return to the example of the payroll software, let’s say the 5th of the month falls on a Sunday. Software from #1 might stall, but a rules-based software with exception handling will roll the payment over to the next business day.
ML-driven systems with oversight. Machine learning-driven systems are designed to observe patterns, make recommendations and improve over time. Maybe it notices that after every payroll, payments through a certain corridor don’t land or get delayed if the 5th happens to land on a Friday. The system might propose an exceptional process for these payments when that happens—maybe releasing them the day before. This is a genuine game changer for treasury operations that could dramatically improve the efficacy of a payments system. Does that make it agentic? Debatable.
Autonomous end-to-end systems. Now we’re talking real agentics: A system that observes conditions, forms viewpoints, decides on a course of action and makes it, without a person in the loop at each step. It optimizes your global payroll, pays freelancers last (and on the last day of their invoice validity), and initiates cross-border transactions within the most statistically favourable conditions. If a problem arises, it might let you know, assuming someone is on the other side to read it.
Nobody actually wants reality #4 in finance at any level, for reasons we’ll get into. But also, let’s step off our soap box of taxonomy and be real: When people talk about agentic payments, what they have in mind is something like a mix of points 2 and 3, where you’ve got a largely rules-based system and then LLM-driven “agents” interacting with it to make things run more smoothly before intervention is needed.
What stops most people is that it’s hard to pinpoint where these agents should be intervening, and at what level of autonomy.
Why Full-Agentic Treasuries Aren’t the Goal
Common industry marketing presents automated cross-border payments as a maturity curve - as in, every treasury operation will eventually rise to full autonomy, it’s just a matter of when.
This is silly.
In “How Agentic AI Will Reshape Payments,” Sonja Davidovic and Hervé Tourpe observe that treasuries seeking to incorporate agentic AI “must reconcile two fundamentally different design logics: the adaptive, probabilistic nature of agentic AI systems and the deterministic requirements of financial market infrastructures.”
The design friction they describe manifests most visibly in the areas of traceability and information management.
The traceability issue is straightforward. It neatly sums up why assuming that all treasury operations will someday be automated and agentically driven is naive and foolhardy. As a rule, payment regimes require payment orders to be traceable to an authorized instruction. This tiny detail is what every audit sits on, and transactions performed by agents may not always correspond to explicit, transaction-level instructions. Their introduction thus requires really clear boundaries of intervention.
Eroding the human-authorization aspect of traceability doesn’t just erode the defensibility upon which the financial system operates, it invites the kinds of issues Davidovic and Tourpe worry about at scale: “without appropriate safeguards, delegating payment initiation to autonomous agents could introduce new operational, legal, and systemic risks, including misaligned incentives, model errors, and highly correlated automated behaviors across markets.”
The second detail relates to the treatment of information. Information security in finance is mainly concerned with the protection of two crucial assets: Identity information and the financial assets themselves. (Money is also information.)
AI-driven agents can only improve with the regular consumption of information. This has led to a surprising characteristic, according to Blue Bridge Group AI: It causes them to pursue it, even when that information is confidential and not necessary to their operations. This emerging problem, flagged in an AI security session at VivaTech in Paris this year, has resulted in the need for companies with agentic systems to create additional agents to mitigate this problem—black-hats to exploit system failures or data weaknesses (including employees); white-hats to patch these finds, and purple-hats to learn and adapt from both.
There are good reasons to introduce a model this unsteerable into your stack? As a means of detecting and mitigating novel kinds of fraud, maybe, but in many other domains purely agentic strategies introduce new risks with limited auditability, risk that has an enormous chance of blowing up in your face.
Now that I’ve sponged all the fun out, let’s talk about where we think autonomous cross-border payments earns its keep.
Realistic Expectations for Automated Cross-Border Payments
Agentic payments make the most sense when transactions are:
Repeatable
Bounded in value and risk
Auditable from end to end
Reversible or escalatable when they fail
Think about these characteristics as preconditions for autonomy to exist anywhere in your stack at all. These details may seem to dilute an agent’s full potential and scope, but we’re actually scoping for accountability.
Notably, this criteria doesn’t control for transaction size or volume. What they control for is the size of the blast radius in instances of error.
Here’s an example of where agentic payments make sense:
You pay a supplier every month for two years through the same corridor. This month, the invoice reference doesn’t match what’s on file; there’s an incorrect digit, probably a typo. A rules-based system might stall around this invoice. But an agent with the transaction history can infer, based on other details, that the vendor and amount are the same. It fits an established pattern, with one minor exception. The agent clears it, but flags it for non-urgent post-execution review.
If its decision turns out to be wrong, the issue is contained and reversible, and both the client and the corridor are well-known to your operations. A phone call would probably fix the issue if one emerges.
Here’s a situation you shouldn’t automate:
A new counterparty sends a payment instruction with a minor error—an incorrect reference number. The error isn’t bigger than the scenario in the first example, but in this case, no payment history exists to compare this document to, and the amount is large. The corridor in question has capital controls, so if you get this wrong and need to call the money back, you might have to produce a regulatory filing.
From an agent’s point of view, these two scenarios are meaningfully similar. That’s what you should take away: Agentic suitability isn’t about how hard something is; it’s about the size of the undo.
Realistically, most real cases you encounter won’t split as neatly as these examples show. You might encounter a repeatable, bounded transaction with a shaky audit trail, or an easily-reversible transaction that doesn’t repeat. The criteria we’ve shared are a starting point from which to safely expand as you get acquainted with how your operations metabolize agentic systems.
Share article
Read other articles
Stay informed with our latest articles on currency launches, institutional FX trends, and global liquidity.






