Can AI trust us?

There is a loud argument about whether AI can be trusted with your numbers. In a finance context, the unreliable component is almost never the machine.

Cover image for “Can AI trust us?”
In short

AI fails in finance functions because of missing context, not model accuracy. A business that defines revenue four different ways cannot blame a model for picking one. Making AI trustworthy is four pieces of work, none of which involve the model: write the context down once, engineer repeatability instead of prompting for it, close the correction loop, and staff the workflow the way you would staff a team.

A few weeks ago, I interviewed three people from the same finance team, asking them the same question: how does this business define a booking?

I got three answers. Four, actually, because one of them gave me two definitions and then explained that it depends on the deal.

Now, hold that thought in your head, because there’s a loud conversation online right now about whether AI can be trusted with your numbers. People talk about accuracy, hallucinations, models that change so often that anything you build on one is already stale by the next quarter. All of this gets debated as though the unreliable component in the equation is obviously the machine.

In a finance context, that is almost never my observation. Can a business that defines revenue four different ways really blame a model for picking one? And it picked without being told which definition you meant, because nobody in the organization ever recorded the ‘right’ definition to use.

We’ve spent two years running a trust test on the machine. In reality, the machine has been quietly running one on us.

The three objections, taken seriously

The objections to trusting AI are real and they deserve more than a dismissive wave, so let’s take them at face value.

  • It makes mistakes. True. So does every analyst you’ve ever hired, usually a lot more in month one than in month twelve.
  • It doesn’t understand our business. Also true. Neither did the graduate you hired in March. The difference is that you gave her six months of context, and you gave the model a prompt.
  • It changes too fast to build on. True again, and the least examined of the three. Models change. What you build around them doesn’t have to.

Every one of those objections describes the relationship, not the model. And the relationship is the part you control.

Every one of these is a correct answer to a question that didnt provide context
Figure 1. Every one of these is a correct answer to a question that didn’t provide context.

Which leads to the more awkward question. If the machine can’t be trusted with your numbers, what exactly are your people doing differently? Mostly, they’re compensating. They’re filling in the gaps from memory and learned context.

What the compensating costs you

Finance work sits on four rungs.

  • Report. Assemble the data and present it. Answers what happened.
  • Explain. Isolate the drivers underneath it. Answers why it happened.
  • Advise. Turn the explanation into a forward view and name the measures that change the outcome. Answers what we do next.
  • Partner. Be a part of the decision and own what follows. Answers what we do differently, and whether it worked.

Value climbs steeply up that ladder. And so does the judgment required, which is exactly why the top two rungs are likely to stay human for a long time yet.

In most organizations I see, the bottom two rungs take up the majority of the function’s capacity — and it isn’t about the seniority of the people. The definitional mess I described earlier means every cycle carries a round of reconciliation, clarification, and some negotiation over which number everyone is using, and all of it gets re-run from scratch the following month.

AI is genuinely good at the bottom two rungs. So, why isn’t it taking more off your team’s hands? The answer has very little to do with accuracy, and a lot to do with the fact that nobody ever provided it with the context to be accurate.

Value climbs. Capacity sits at the bottom. That gap is the whole opportunity
Figure 2. Value climbs. Capacity sits at the bottom. That gap is the whole opportunity.

So, the question is worth restating — It’s not whether you trust AI. It’s whether you’ve built the conditions that would make it trustworthy. That’s four pieces of work, and none have anything to do with the model:

One: write the context down, once

Start with the job nobody wants.

This isn’t a data dictionary sitting in a folder that nobody opens. It’s a working context layer:

  • How revenue is recognized and where the judgment sits.
  • Which adjustments are standing and which are one-time.
  • What the reporting calendar actually is on the occasions it deviates from the one you published.
  • Which entities roll up where.
  • Which three customer names are really one customer.

Every hour a new analyst spends absorbing this information in their first six months is an hour of context that currently exists nowhere except in people’s heads. That is the asset. AI can’t inherit what was never written.

When I ran finance for Cisco’s MEA region, the most useful document my team had wasn’t a model. It was a data stack of internal notes on how the region’s numbers actually behaved: which quarter-ends distorted what, which markets needed a second look before taking action, and which reporting quirks were real versus habit. We wrote it because people kept moving on and taking the knowledge with them. As it turns out, we’d built precisely the thing an AI system needs, years before there was one to give it to.

Write yours. It will be shorter than you expect and more revealing than you’d think.

Two: engineer repeatability instead of hoping for it

The instinct with AI in finance is to ask it a question and then judge the answer. The discipline is to build the framework that produces the answer, and then run that.

This is what AI-directed engineering means in practice. A dashboard that somebody rebuilds by prompt each month will be slightly different every time, and those differences stay invisible until one of them matters. A dashboard built once inside a harness (a fixed set of inputs, transformations, checks and output formats that the model populates rather than reinvents) gives you the same answer on the hundredth run as it did on the first.

Stop asking the model to remember. Make the system remember, and let the model do the reasoning inside it.

That single shift removes most of the sting from the “models change too fast” objection. When the next model lands, you swap the engine and keep the chassis.

Three: close the loop

Almost every AI deployment I see in finance runs open loop. Input goes in, output comes out, a human reads it, and when it’s wrong the human fixes their own copy and gets on with the day. The correction dies on that desk. Next month the system makes the identical mistake because nothing ever told it otherwise.

A closed loop writes the correction back. When a reviewer overrides a classification, the override lands in the context layer rather than in the deck. When a variance explanation gets rejected, the reason is captured. The system arrives at the next close knowing something it didn’t know at the last one.

That is the difference between a tool that is exactly as good in month twelve as it was in month one, and a system that gets just a little better every cycle. Trust doesn’t get granted on the first run. It accrues, and only if you’ve built the mechanism that lets it accrue.

The flowcharts are nearly identical. One arrow is the entire difference between a tool and a system
Figure 3. The flowcharts are nearly identical. One arrow is the entire difference between a tool and a system.

Which leaves the question of what you’re actually building, because a dashboard was only ever the beginning.

Four: graduate from a dashboard to a system

Nearly every finance team’s AI journey starts with a dashboard, and that’s the right place to start — it’s an early win: It’s small, it’s visible, and it either works or it doesn’t.

The mistake is stopping there. A dashboard is an artifact. A system is a set of roles.

Most finance teams are at stage one and calling it an AI strategy. The value is two steps to the right
Figure 4. Most finance teams are at stage one and calling it an AI strategy. The value is two steps to the right.

Architect a finance process properly for AI and it starts to resemble a real org chart, because that is exactly what it is. One agent retrieves and prepares the data. One analyzes it and flags what moved. A separate one checks the first two against the rules — segregation of roles matters here because a reviewer who also did the work has never been a reviewer. One drafts the narrative. A named human owns the output and signs off on it.

Segregation of duties. Finance invented it. We just never thought to apply it to the machines.

Staff the workflow the way youd staff a team. If you wouldnt let one analyst prepare, review and sign the same pack, dont let one agent do it either
Figure 5. Staff the workflow the way you’d staff a team. If you wouldn’t let one analyst prepare, review and sign the same pack, don’t let one agent do it either.

Which brings me to what the whole climb is for, because it was never about cost.

What the climb is actually for

There’s a version of this argument that ends with a warning about being replaced. I don’t find it useful and I don’t think it’s true.

Automating the bottom of the ladder buys you exactly one thing: the top of it. Trust me, nobody in your business is waiting for a faster variance report. They’re waiting for someone who can walk into a decision before it’s taken and bring a view worth arguing with. That seat has been open in most companies for years, and the reason finance hasn’t taken it is rarely ambition or capability. It’s about scale. There aren’t enough hours left after the close.

Everyone’s going to end up running much the same models. The teams that pull ahead will be the ones that bothered to write down what their own numbers mean. So I invite you to run the test with your finance team. Three people, one definition, no discussion. Whatever comes back is the honest state of your AI readiness, and it has nothing to do with which vendor is on your invoice.

The machine will be exactly as reliable as the context you can hand it. In most finance functions, that context lives in six people’s heads and nowhere else. Fixing that takes a Tuesday afternoon, a blank document, and some determination. Which is probably why so few have done it.

Kamran Habibollah is the founder of Third Horizon Capital Advisory, which provides CFO-grade strategic finance insights to founders, CEOs and boards. He spent 20+ years in global technology enterprises running multi-billion-dollar P&Ls.

Kamran Habibollah

Kamran Habibollah

Founder & Principal, Third Horizon

Twenty years in technology and telecommunications finance, including senior finance leadership at Cisco, across the Middle East, Africa and Europe. He advises founders, CEOs and boards on capital strategy, transactions and investor relations from Dubai.

More about Kamran →

start/

Is your context
written down?

If the answer is “it’s in people’s heads”, that is the first conversation — and it is a short one.