expenieexpenie

Guides

How AI expense categorization works, and why it guesses wrong

Categorization is inference from a merchant name and context. Understanding what the model can and cannot see explains most of its mistakes.

expenie guide cover: how AI expense categorization works and why it guesses wrong

In short

The short answer

AI categorization infers a category from the merchant name, amount, and any text you provide. It has no access to what you actually bought, so a supermarket that sells electronics and a shop that sells everything are inherently ambiguous. Adding a few words of context improves suggestions more than anything else.

Inference, not lookup

The intuitive model is that the app has a database mapping merchants to categories. That is how older rules-based systems worked, and it explains neither the strengths nor the failures of the current generation.

A language model categorising a spend is reasoning from what it knows about the world. Given "Riverside Dental — 180", it infers a healthcare-type category from the name, without ever having seen that specific business.

This generalises far better than a lookup table. A new coffee shop with an unusual name is usually categorised correctly on the first encounter, which no merchant database could do.

It also fails differently. A lookup miss produces no suggestion — obviously unhelpful, and obviously something you must fill in. An inference miss produces a confident, plausible, wrong suggestion, which is much easier to accept without noticing.

What the model can and cannot see

Most categorisation errors are explained by the information simply not being available.

It can see the merchant name as printed, the amount, the date, whatever text you wrote, any items listed on a receipt, and the set of categories you actually use.

It cannot see what you bought when the receipt does not itemise, why you bought it, whether it was for you or for work, whether it is reimbursable, or how you personally think about a borderline case.

That last one is underrated. Whether a work lunch is Dining Out or Business is a decision about your own categorisation scheme, not a fact about the world. No amount of model capability resolves it, because the answer lives in your head.

The cases it reliably gets wrong

Predictable failure patterns, all following from the visibility limits above:

  • General merchants. A large supermarket sells food, clothing, electronics, and garden furniture. The merchant name genuinely does not determine the category.
  • Online marketplaces. A charge from a large platform tells you nothing about what was purchased.
  • Payment intermediaries. A charge showing only a processor's name hides the actual merchant entirely.
  • Abbreviated or coded merchant names, which appear as strings that mean nothing without context.
  • Personal conventions. If you file petrol under Car rather than Transport, the model has no way to know until it has seen your categories.
  • Business versus personal on an identical purchase. Same merchant, same amount, different answer depending on purpose.

Notice that none of these are model failures in a meaningful sense. In every case the information required to answer correctly was not present.

How to get better suggestions

The highest-leverage fix is providing the context the model lacks — and it takes a few words.

"Supermarket 84" is ambiguous. "Supermarket 84, groceries and a phone charger" is not. That is four extra words and it removes the guesswork entirely.

Other things that help:

  1. Use category names that describe the thing rather than the context. "Groceries" is clearer to a model than "Weekly shop".
  2. Keep the list flat and reasonably small. Twelve well-named categories give a model a much clearer target than forty overlapping ones.
  3. Say the purpose when it is not obvious. "Client dinner" versus "dinner" is the whole difference for a business split.
  4. Mention the account when it matters. "On the business card" resolves an ambiguity nothing else can.

In expenie, capture requests include a live snapshot of your workspace — your actual accounts and categories — so suggestions are drawn from the set you use rather than from generic labels. It cannot infer your personal conventions, but it will not invent a category that does not exist.

Why you should still check

Category errors are less damaging than amount errors, and they are not harmless.

A misfiled expense corrupts the one thing categories exist for: budget totals. If a third of your dining spend lands in Groceries, both categories are wrong, your dining budget looks healthy while you overspend, and the data cannot tell you what is happening.

Worse, the error is systematic rather than random. A model that miscategorises a particular merchant will do it consistently, so the distortion accumulates in the same direction month after month.

This is why the confirm step matters even for a field that feels minor. In expenie a draft is not a transaction — you accept, edit, or skip each one, and correcting a category at that moment costs a second.

What good use looks like

The realistic pattern for someone using this well:

You describe a spend or photograph a receipt. A draft appears with amount, date, payee, and category filled in. You glance at the amount, glance at the category, and either accept or change one field. The whole thing takes a couple of seconds and saved you typing four fields.

Over time you learn which merchants it gets wrong for you specifically, and you add two words of context for those. That is a small, permanent improvement that costs nothing.

What good use does not look like: accepting drafts without reading them because it is usually right. Usually right is a very different thing from right, and the difference accumulates in your budgets.

FAQ

How does AI decide what category an expense belongs to?
It infers from the merchant name, amount, and any text you provide, reasoning about the world rather than looking up a database. That is why it handles unfamiliar merchants well and ambiguous ones badly.
Why does it keep miscategorising my supermarket spending?
Because a large supermarket genuinely sells across categories, and the merchant name does not determine what you bought. Adding a few words like "groceries and a charger" removes the ambiguity.
Does it learn my personal categorisation habits?
It works from the categories you actually use, so it will not invent ones you do not have. It cannot infer personal conventions like filing petrol under Car rather than Transport — say so in the text instead.
Do category errors actually matter?
Yes, because categories exist to drive budgets. The errors are systematic rather than random — a merchant miscategorised once is miscategorised every time — so the distortion accumulates in one direction.

Try expenie

Solo private ledger. Manual entry. Statement-first month. 14-day full Pro trial, then subscribe.

Try AI capture on Pro