Skip to content

Why AI Reporting Tools Give You Wrong Numbers (And How to Tell)

The dangerous failure in AI reporting is not an error message. It is a clean, confident number that happens to be wrong - and here is why it happens so often.

By Karani Geoffrey, Founder & CEO, Upeosoft
In short

AI reporting tools get numbers wrong because they translate your question into a database query without knowing what your tables actually mean. Accounting systems store revenue, stock value and margin in non-obvious places, so a plausible query returns a plausible wrong figure. The fix is a curated set of business measures defined by a human, which the AI selects from rather than inventing.

Key takeaways
  • The risk is not that AI fails loudly - it is that it succeeds confidently and incorrectly.
  • Accounting databases do not store revenue and margin where their names suggest.
  • A tool writing its own queries will eventually miss credit notes, drafts or cost of goods.
  • The safe pattern is a fixed catalog of business measures, each checked once by a person.
  • Any tool should be able to show you where a number came from.
  • Prefer a tool that admits uncertainty over one that is never uncertain.

The demo that never gets checked

You have probably seen the demonstration. Someone types "show me last month's revenue" into a chat box, the AI writes a query, a chart appears, and everyone is impressed.

What the demonstration never includes is somebody verifying the number against the books.

When you do verify, on real business data, the failure rate is high enough to matter. Not because the technology is fake, but because writing a query and understanding a business are different skills, and only one of them can be inferred from a database schema.

Why your database lies about itself

Take any serious accounting or ERP system. Its structure reflects double-entry bookkeeping, not plain English, and that mismatch is where the errors live.

Ask a naive AI for last month's revenue and it will find something that looks like a sales invoice table, add up a column that looks like a total, and hand you a figure. That figure will not account for credit notes that reverse part of the sale. It will not exclude invoices still sitting in draft. In many systems, properly reconciled revenue lives in the general ledger rather than the invoice table at all.

Stock valuation is stored somewhere non-obvious too. And gross margin usually is not stored anywhere - it has to be reconstructed from cost of goods sold. None of this is exotic. It is simply how the system works, and a model that has never seen your configuration cannot know it. So it does the reasonable thing available to it: it guesses, fluently.

Why a confident wrong number is worse than none

Software that fails loudly costs you time. An error message, a blank screen, a call to support - irritating, but contained.

Software that fails quietly costs you decisions. You price against the wrong margin. You reorder against the wrong movement figure. You conclude a branch is doing well when it is not. And because the output looked exactly like every correct output, nothing prompts you to check.

This is why "the AI is right most of the time" is not a reassuring statistic in reporting. If you cannot tell which answers are in the wrong minority, an eighty-percent accuracy rate is not a tool - it is a trap with good manners.

The pattern that actually works

The fix is unglamorous and it works. Instead of letting the AI invent its own calculations, you give it a curated catalog of business measures - and a human defines each one properly, once.

  • Revenue and gross margin, by item and category, over any period.
  • Stock turnover and days of cover remaining.
  • Dead and slow stock, with capital tied up.
  • Best and worst movers, ranked on profit.
  • Supplier price variance for the same item.
  • Receivables aging and current till position.

Catalog first, and mean it

A catalog only protects you if the tool is actually required to use it. The rule that matters is that the AI must check the defined measures before anything else, and may only write its own query when it can state that no defined measure answers the question.

When it does improvise, three guardrails should apply. The query must be read-only, so nothing can be changed. It should be fenced with a row limit, a timeout and a hard scope to your company's data. And it should show its working - the query it ran and the assumptions it made - so you or your support partner can sanity-check it.

That last point costs the vendor something. It openly admits that this answer is less certain than the others. It is the right trade, because the alternative is a system where every answer looks equally confident and some of them are not.

How to test a tool before you trust it

You do not need technical skill to run a sensible evaluation. You need a figure you already know cold.

  • Pick a month you have closed and verified. Ask the tool for its revenue and compare.
  • Ask the same question two different ways and check the answers match each other.
  • Ask for a figure involving cost - gross margin is the classic - since that is where naive tools break.
  • Ask something deliberately awkward and see whether it admits it cannot answer.
  • Ask it to show its working, and see whether that option exists at all.
  • Repeat the spot-checks monthly for a season, not just during the trial.

Where Upeosoft stands on this

We build business systems for Kenyan companies, and we add AI to them where it earns its place. Our position on reporting is firm: an AI must reason over defined, verified business measures, not improvise against a live database.

That is slower to build. It means somebody sits down and works out exactly how gross margin should be computed in your setup, and writes it down, and tests it. It is also the only version of this that we are willing to let a business make pricing and purchasing decisions on.

If you are evaluating an AI reporting tool, or you have one and something about the numbers feels off, talk to us. Half an hour of checking is cheaper than a quarter of decisions built on a confident wrong figure.

Frequently asked questions

Is the AI just not clever enough to get this right?

No, and that is the uncomfortable part. Modern models write good SQL. The problem is that good SQL against a schema you do not understand is still the wrong answer. Your accounting system encodes decades of bookkeeping convention, and no model can infer your particular setup - your credit note handling, your draft invoices, how your costs are posted - from table names alone.

How would I even notice a wrong number?

Usually you do not, which is the whole problem. It surfaces months later when your accountant's figures disagree with the ones you have been making decisions on. The practical defence is to spot-check any new tool against a figure you already know cold - last month's total from a source you trust - and to keep checking for the first few months rather than just the first week.

What is a metric catalog in plain terms?

It is a list of the business questions that matter - revenue, gross margin, stock turnover, dead stock, receivables - where somebody has worked out, once and carefully, exactly how each one should be calculated for your system. The AI then picks from that list instead of improvising. Choosing the right question from a menu is a far smaller and safer job than inventing the arithmetic.

Does this mean AI reporting is not worth it?

Not at all. It means the value sits in the preparation, not the chat box. A tool built on properly defined measures gives you genuine speed - answers in seconds to questions that used to need somebody to build a report. The warning is only about tools that skip that preparation and point a model straight at your database.

What should I ask a vendor about this?

Ask one question: where does this number come from? If the answer is that the model generates a query each time, you know what you are buying and can price the risk accordingly. If the answer is that it maps to defined measures somebody has verified, ask to see how one of them is defined. A vendor who can show you that is a vendor who has done the work.

Karani Geoffrey
Karani Geoffrey
Founder & CEO, Upeosoft

Karani Geoffrey is the Founder & CEO of Upeosoft, a software and automation company rooted in Kenya. He builds custom software, AI systems, and production-grade ERPNext for businesses across East Africa, and writes about the Kenyan realities - eTIMS, M-Pesa, SHIF, unreliable internet and power - that make or break real systems.

Next step

Want this working in your business?

Upeosoft builds and hardens the systems behind this article - for real Kenyan operations, with eTIMS, M-Pesa and offline realities handled.

Keep reading

AI and Automation for BusinessBuyer's guide

An AI Business Analyst for Your Shop: What It Can Actually Tell You

Your shop already produces the data that would answer your hardest questions. An AI analyst is the thing that finally lets you ask them in plain language and get a real answer.

7 min readRead article →
ERP and Business Systems

Understanding Business Dashboards: Which Numbers Actually Matter

A good dashboard answers the few questions that actually change what you do. Here is how to tell the numbers that matter from the vanity metrics that just look busy.

5 min readRead article →
AI and Automation for Business

How to Spot AI Snake Oil: Real Tools vs Hype

A practical guide to telling genuine AI tools from expensive hype, with the warning signs and the honest questions that expose snake oil before you pay.

7 min readRead article →