Pull a receipt out of your back pocket after a long day and look at it. The ink is faded. There's a crease straight through the total. Half the top is a logo, the middle is a list of items in a cramped font, and the tax is tucked in near the bottom next to three numbers that all look like they could be the total.
A human reads that in a couple seconds. Getting software to do the same is genuinely hard, and the fact that it works most of the time now is one of those quiet advances that snuck up on everyone.
First, the software has to see the letters
The starting point is OCR, which just means optical character recognition, turning a picture of text into actual text characters. This part has been around for decades. Scan a clean printed page and it's nearly perfect.
Receipts break the easy version. The paper is curved, the lighting in your kitchen is bad, thermal ink fades unevenly, and a crease can wipe out a whole line. So before it reads anything, a good scanner tries to flatten and clean the image, straightening the angle you shot it at and boosting the contrast so faint numbers come back into view.
Reading isn't the same as understanding
Here's the part people underestimate. Even with perfect character recognition, you're left with a jumble of text. The word "TOTAL" might appear three times: subtotal, total, and "total savings." There are five dollar amounts on the page. Which one goes in your expense record?
This is where the AI layer does the real work. Instead of hunting for a fixed spot on the page, it reads the whole receipt in context, the way you would, and reasons about what each number means.
The clever bit isn't reading the numbers. It's knowing which number you actually care about.
It learns the patterns that hold across thousands of different receipt layouts. The total usually sits below the subtotal and tax. Tax often sits right above it, sometimes with a percentage next to it. The merchant name is usually up top, larger, often near a phone number or address. None of these rules are absolute, so the model weighs them together rather than trusting any single one.
What good extraction actually pulls out
When it works, a single photo turns into a tidy little record without you typing anything:
- The merchant, cleaned up from the messy header text.
- The date, converted to a real date even when the receipt writes it oddly.
- The total, chosen from among all the dollar amounts on the page.
- The tax, separated out so it's ready for bookkeeping.
That structured data is the whole point. A photo you can't search is just a photo. Once the app behind my receipt scanning setup has pulled those fields, I can search by store, sort by date, or total up a category without ever opening the images.
Where it still slips, and why that's fine
It's not magic. A receipt that's genuinely destroyed, torn across the total, ink completely gone, will stump it, the same way it would stump you. Handwritten receipts from a small shop are harder than printed ones. Very unusual layouts occasionally put the wrong number in a field.
The healthy way to use it is to let the AI do the ninety-odd percent that's tedious, then glance at the total for a half-second before saving. That quick check takes the pressure off the software being perfect, and it almost always is right anyway. What used to be manual data entry becomes a quick nod of approval, and that's a trade I'll take every time.