The full reel breakdown — the last of five

Four agents for the paperwork nobody talks about.

In the reel I said I'd write the whole system out and send it. Here it is — four agents, the rulebook that decides what counts as a mismatch and what should be ignored, an honest section on where this breaks, and a sheet you can copy. This is also the last reel in the series, so the end of the page links to the other four.

Here's what you actually get:Every bill and invoice read the day it arrives instead of at month end. Every one checked against something — the order, the bank, the shelf. Whatever matches gets entered. Whatever doesn't reaches the one person who can fix it, the same day, while they still remember the transaction. And whatever the system isn't sure about goes to a person before it ever touches your books.
Problem

The one person who only does paperwork

In every business there's one person who only handles paperwork. Not sales, not delivery — invoices, entries, mismatches, reports. Some days that is the entire day.

This work is invisible, and that's the whole problem. Nothing breaks when it's done. Nothing visibly breaks when it isn't, either — it breaks three weeks later at month end, when the numbers don't tie out and nobody can remember which bill it was or which delivery was short.

So it never gets automated. Everyone automates the loud things: leads, calls, messages. Nobody automates the thing quietly eating four hours a day, because nobody is complaining about it.

But it happens every single day. And work that happens every day gets paid for every month.

System

Four agents, four separate jobs

Each agent does one thing. One. That is exactly why you can tell which part is working and which isn't.

  1. Agent 1

    Extracts the data

    Invoices, bills, PDFs, photos of documents. It pulls out the fields that matter — date, party name, GSTIN, invoice number, line items, quantities, rates, tax, total. It judges nothing at this stage and compares nothing. It only reads.

    It also writes a confidence number against every field it read. That number is the single most important output of this agent, and everything downstream runs off it. There's a whole section below on what happens when it's low.

  2. Agent 2

    Checks whether it reconciles

    It takes what Agent 1 read and puts it next to something else — a purchase order, a bank line, a physical stock count. Three comparisons, written out in full in the rulebook below. It fixes nothing and enters nothing. It produces one verdict per document: matches, doesn't match, or inside tolerance and not worth anyone's time.

    This is the agent that decides whether a human being gets disturbed. Most of the value in the whole system sits in one thing it does — deciding what not to flag.

  3. Agent 3

    Enters it

    Into the sheet, or into Tally, Zoho, Vyapar — whatever you already run. One row per document, with the source file linked against the row, so anyone auditing the entry six months later can open the bill it came from in one tap.

    It only enters what matched. Anything Agent 2 flagged, and anything below the confidence threshold, does not get entered — it waits in the review queue. Wrong data sitting in your books is far more expensive than no data, because you'll trust it.

  4. Agent 4

    Alerts, wherever there's a mismatch

    Immediately — not in a daily digest, not in a weekly report. A WhatsApp message to the one person who can act on it, carrying the amount of the gap, the document, and what it was compared against, so nobody has to go hunting before they can answer.

    Immediately, but only for what Agent 2 marked worth alerting on. That restriction is the whole design, not a limitation of it — read the Ignore line in every rulebook block below.

Rulebook

Agent 2's reconciliation rulebook

Three comparisons cover almost everything a small business needs checked. Every block answers the same four questions in the same order — what gets compared, what's a real mismatch, what must be ignored, and who gets the alert.

Jump straight to the one you need:

Same four things in every block — Compare (what Agent 2 puts side by side), Flag (what's a genuine mismatch), Ignore (what must never raise an alert), Alert (who it goes to). The Ignore line is highlighted in all three blocks on purpose. It's the one that decides whether this system is still switched on in week three.

Invoice — the supplier's bill against what you ordered and what actually turned up.

Invoice — three documents, and the third one is the one everybody skips

Most people check the supplier's invoice against the purchase order. Almost nobody checks it against what physically arrived. That's where the money actually leaks, and it usually isn't the rate — it's the quantity.

  • Compare What Agent 2 compares

    Three documents, not two. The invoice (what they're charging), the purchase order (what you agreed), and the delivery challan or goods received note (what actually landed). Line by line: item, quantity, rate, taxable value, GST rate, line total. And if there's no GRN for that delivery, the agent says so on the row instead of quietly calling a two-way check a three-way one.

  • Flag What counts as a mismatch

    Rate charged above the rate on the PO — any amount. Quantity billed above quantity received — any amount. A GST rate different from what that supplier normally charges on that item. An invoice number you already have a row for, because duplicate billing is real and it's usually an accident, not a fraud. And a total that doesn't equal the sum of the invoice's own lines, which happens more often than you'd think on manually made bills.

  • Ignore What gets ignored

    Rounding, first and foremost. Every accounting package rounds totals to the rupee and tax lines to two decimals, so a ₹0.50 or ₹2 gap between your arithmetic and theirs is arithmetic, not a dispute. Set a tolerance — ₹5 or 0.5% of the invoice value, whichever is smaller — and anything inside it gets entered silently, no alert, no queue. Also ignored: the party name written differently ('Sharma Traders' against 'Sharma Trading Co.') when the GSTIN matches, and the challan date being a day off the invoice date. None of those are errors. And if Agent 4 pings you about them, you will start ignoring Agent 4 — and then you'll scroll past the ₹6,000 rate difference sitting three alerts below.

  • Alert Who gets the alert

    Whoever raised the purchase order — not accounts. Accounts can only tell you the bill doesn't tie out. The person who placed the order is the only one who knows whether the rate got renegotiated on a phone call last Tuesday. The alert carries the supplier, the invoice number, the size of the gap, and which of the three documents disagrees. The entry is not made until they answer.

The most common real mismatch in Indian SMEs isn't the rate, it's the quantity — billed for 100, 96 arrive, and nobody catches it because nobody counts against the challan at the gate. The system can only catch that if a GRN exists in some form. If your receiving process is 'someone signs the challan and puts it in a drawer', this check has nothing to compare against, and no amount of AI fixes that. Fix the drawer first, then automate. That is a cheaper project and it finds more money.

Read the whole rulebook from the top

Payment — the bank line against the invoice against what the books already say.

Payment — the bank statement is the only document that tells the truth

One rule sits under this entire block: the bank is the fact. An invoice is a claim. A screenshot is a claim. A ledger entry is somebody's memory of a claim. Only the bank statement says money actually moved.

  • Compare What Agent 2 compares

    Every bank credit and debit against the invoice or bill it's meant to settle, and both of those against what the ledger already shows. Three fields do the matching — amount, a date window, and the reference (UTR, cheque number, or the narration text). A payment that matches no open invoice at all is as much a finding as one that matches badly: unallocated money sitting in the account means somebody's invoice is still showing unpaid in your books.

  • Flag What counts as a mismatch

    A short payment against an invoice with no recorded reason. Money received that matches no open invoice. The same invoice settled twice. A bank debit with no bill behind it. And an invoice past its credit period with nothing against it — that last one isn't strictly a mismatch, it's a report, and it's the single most valuable thing this agent produces, because it's the list of people who owe you money that nobody has looked at this month.

  • Ignore What gets ignored

    Bank charges and UPI/NEFT fees shaving a few rupees off the credited amount. TDS deducted by the customer — a ₹27,500 invoice settled at ₹25,000 is the standard 10%, and it is a known deduction, not a short payment; once you've told the system that this customer deducts TDS, it becomes an expected line and stops being an alert forever. Payment dates a day or two either side of the invoice date. And any per-transaction gap inside the same ₹5-or-0.5% tolerance. Every one of these is a completely normal Tuesday. Alert on them and within a week nobody opens the alerts, which means the duplicate payment goes through unread.

  • Alert Who gets the alert

    Accounts, for anything about an amount. The owner directly, for exactly two things: money in that matches no invoice, and money out that matches no bill. Those two are the only ones that are ever fraud or a serious error, and they are also the two most likely to get quietly sorted out by whoever caused them if the alert lands on the wrong desk.

One thing to be strict about: a payment screenshot sent by a customer is not a bank entry, and Agent 2 must never reconcile against it. It goes into the row as a claim, with the word 'claimed' next to the amount, and it clears only when the bank statement shows it. Pending transactions screenshot identically to successful ones, UPI reversals happen, and a screenshot can simply be made. If you read the WhatsApp desk page, this is the same rule seen from the other end — that system writes the row as unverified, and this is the system that clears it.

Read the whole rulebook from the top

Inventory — what the books say you have against what's actually on the shelf.

Inventory — stock on paper against stock on the shelf

This is the check done once a year, in a panic, in March. It should run on a rolling basis for the ten items that actually matter — and counting ten items is a small enough job to do every week.

  • Compare What Agent 2 compares

    Closing stock in the software against a physical count, item by item, for a defined set of items over a defined period. And the movement in between: opening stock, plus purchases, minus sales, should equal what's on the shelf. When it doesn't, the gap is one of three things — a sale that wasn't entered, a purchase or return that wasn't entered, or stock that left without a document.

  • Flag What counts as a mismatch

    Any gap at all on a high-value item — one missing unit of something worth ₹40,000 is not a rounding problem. A gap moving in the same direction on the same item, count after count, which is the actual signature of a leak as opposed to a bad count. Negative stock in the software, which is always a data error and always means an entry is missing. And stock that hasn't moved in 90 days, which isn't a mismatch either — it's money sitting on a shelf, and it's worth a separate line in the report.

  • Ignore What gets ignored

    Small variance on loose or weighed goods. If you sell rice, oil, cloth, or hardware by weight or cut length, a 1-2% variance is physics — moisture, spillage, cutting loss — not theft. Set that tolerance per category, never one number for the whole store: a tolerance that's sensible for loose grain is absurd for mobile phones. Also ignore counts taken mid-delivery, when stock is physically in the shop but not yet entered — the system knows a GRN is pending on that item and simply doesn't count it that day. Flag either of these and you produce a steady stream of alerts about nothing, and the fastest result of that is a store person who stops counting properly.

  • Alert Who gets the alert

    Whoever holds the stock, for anything in a normal range — it's usually theirs to correct and it's usually a missing entry, so it should reach them first and quietly. The owner directly, for two things only: a high-value item short, and a gap that has grown three counts running. That second one is deliberately slow. One bad count is a bad count. Three in the same direction is something else, and it needs to reach the person who can act on it without passing through the person it might be about.

Be honest about what this check is. It does not find theft. It finds a gap, and a gap has four ordinary explanations before it has a dishonest one — an unentered sale, an unentered purchase return, a unit given away as a sample, and a wrong unit of measure (a box entered as a piece, which alone accounts for a startling share of them). So the alert has to read as 'go and look', never as an accusation. Send an accusation to the wrong person once and you will never get an honest count out of that shop again — and after that the system is worse than useless, because now the numbers are being managed for you.

Read the whole rulebook from the top
Limits

Where this breaks

This is the part I'd want to read before paying anybody for this, so here it is straight. Extraction accuracy on Indian small-business paperwork is not 100%, and anyone telling you otherwise has not run it on a real pile of bills.

Look at what the pile actually contains. A clean GST invoice, emailed as a PDF. A photo of a printed bill taken at an angle, in shop light, with a thumb in the corner. A faded thermal print that was already grey when it was handed over. A dot-matrix invoice on carbon paper. And a handwritten kachcha bill in a mix of English and Hindi, with the total scratched out and rewritten. Those are five completely different problems, and only the first one is easy.

Here's the rough shape of it — and where these numbers come from matters, so: these are the ranges this class of document typically falls into, what extraction generally manages on paperwork like this. They are not measurements from your pile, or from any one business's pile. A typed PDF invoice sits around 98% field accuracy and is effectively a solved problem. A clear photo of a printed bill, somewhere around 90-95%, with the failures usually in the line items rather than the total. Faded thermal, dot-matrix and carbon copies, closer to 75-85%. Handwritten bills, roughly 50-70% — and worse than that number sounds, because you cannot tell which half it got wrong.

Across a mixed pile that typically works out to something like 10% to 20% of documents needing a human touch in the first month. It improves from there as you add layout templates for your regular suppliers, because the same six suppliers usually account for most of your volume and each of them sends the same shaped bill every time.

So here's what you do about it, and it's the whole answer: every field comes out with a confidence score, and anything below the threshold does not get entered, does not get reconciled, and does not raise an alert. 0.85 is a sane place to start. It goes to a review queue where a person sees the document image and the extracted numbers side by side and either confirms or corrects. Confirming is a ten-second job. Finding a wrong number in your books three months later is a two-hour job, and it costs you the trust you had in every other number next to it.

Which brings the fair objection: 'so I still need a person'. Yes. The question was never whether a person is involved — it's how much of their day comes back. That's answerable with a number, so here's the arithmetic on a business doing 40 documents a day.

Entering one document by hand is 2 to 4 minutes end to end — finding it, reading it, keying it in, tying it out. Forty documents is about two hours a day, and those two hours are what you're paying for right now. Clearing one review-queue item is 20 to 40 seconds, because the document image and the extracted fields are already side by side and the job is to confirm or correct one number, not type nine.

So, mostly emailed PDFs and clear photos, at a typical 10% in the queue: that's 4 documents a day and roughly two minutes of clearing — plus the two or three genuine mismatches that need an actual decision from you. Call it 15 minutes a day against two hours. A mixed pile with faded thermal and carbon copies in it, at around 20%: 8 documents, still under ten minutes of queue.

And the honest end of it. Where most of the bills are handwritten, the queue typically runs at 40-50% of documents, and each one is slower to clear, because now the reviewer is reading handwriting rather than checking a machine's reading. Forty documents a day, half of them queued at a minute each, is about 30 minutes a day. That's still a quarter of the two hours, and it's worth paying for — but say what it actually is: you haven't removed the paperwork person, you've changed the job from entering to checking. Anyone selling you 'fully automated' on a handwritten pile has not run it on one.

Every number on this page is the industry-typical shape of the problem, not a measurement of your business — so measure your own queue rate before you believe any of them, mine included. It's the first thing the system tells you, in week one, and it's the right number to judge the whole decision on. If it lands near the top of that range, the cheapest fix isn't a better model — it's the input. Get your regular suppliers to email PDFs, and stop accepting handwritten bills from the ones where you have the leverage to insist. Every document you move from handwritten to PDF leaves the queue permanently.

And one thing it will never do: it cannot catch a bill that never arrived. A supplier who doesn't send an invoice, a sale made off the books, cash that never touched the bank — none of that appears in any document, so none of it appears in any check. This system reconciles what exists. It has nothing to say about what doesn't.

Accuracy isn't the number to negotiate. The number to negotiate is what happens to the 10% it isn't sure about — and the only acceptable answer is that a person sees it before it reaches your books.

Why

Four narrow agents instead of one big automation

Fair question — why not one automation that reads the bill and puts it in the books?

Because reading and deciding are different jobs, and one thing doing both hides which one failed. The entry comes out wrong and you cannot tell whether it read ₹4,800 as ₹4,300, or read it perfectly and compared it against the wrong purchase order. Two completely different fixes, and no way to tell them apart from the outside.

Second, the confidence number has to survive the handoff. One big automation has one output — the entry. It has no way of saying 'I've entered this and I'm 60% sure of the amount'. Split into four, the confidence Agent 1 produces becomes a thing Agent 3 is allowed to refuse to act on. That refusal is the safety mechanism of the entire system, and it only exists because the reader and the writer are separate.

Third, and this is the one people miss: the four jobs need opposite temperaments. Extraction should be generous — attempt everything, read every smudged line. Checking should be strict — rules, tolerances, no benefit of the doubt. Entry should be conservative — write nothing that isn't confirmed. Alerting should be quiet — say almost nothing, and only to one person. You cannot tune one system to be generous and strict and conservative and quiet at the same time. Four agents get one temperament each.

And the practical one, as always: switch Agent 3 off and keep the other three. You get documents read, mismatches caught and alerts sent, with a person making the actual entries. Most businesses should start exactly there. Entry is the last of the four to automate, not the first — it's the only one that writes into something you can't easily undo.

One automation gives you an entry and no way to audit it. Four give you a number, a verdict, an entry and an alert — four places to check, and four things you can switch off one at a time. The entry should be the last one you switch on.

The series

Five reels, five systems, and the same shape underneath

This is the last one in the series. Five reels, five systems, and it's worth saying out loud what all five had in common — because it isn't the tools and it definitely isn't the prompts.

Not one of them was a single AI doing everything. Every one was the work broken into pieces, with each piece handed to an agent that does one thing. Five agents for the lead pipeline. Five for the front desk. Four for the comeback list. Five for the WhatsApp inbox. Four here.

The reason is the same every time. A narrow job can be described exactly, which means it can be tested, which means when it goes wrong you know which piece went wrong. And you can switch one piece off without switching the system off. That last part is what makes any of this safe enough to run on a business that has to keep operating tomorrow.

For the freelancers and agency people reading — this is the part worth internalising. Nobody is paying you for a prompt. A prompt is something the client can get for free, from the same model you're using, in about four minutes, and increasingly they know that. What they cannot do is break their own business into steps, work out which step must never be automated, and decide what the system does when it isn't sure. That is the work. That is what the invoice is for.

It's also why the boring systems sell best. Nobody gets excited about invoice reconciliation. But it happens every single day, it is costing a real salary right now, and work that happens every day gets paid for every month.

The structure gets paid for, not the prompt. Five reels, and that's the whole argument.

All five

The other four systems

Each one is written out the same way as this page — the agents, the rulebook, and a sheet to copy. Read whichever one is closest to your problem; you don't need them in order.

  1. One

    The 5-agent lead system

    What happens to a lead between arriving and being called. Five agents, and a scoring rulebook that's written separately for property, insurance and coaching — because a good lead means something different in each.

    sharanjeetdigital.in/r/system — includes a scoring sheet per industry.

  2. Two

    The AI receptionist

    The calls nobody picks up. Five agents that answer, take the details, put them in the CRM, follow up, and send a weekly report. It has a live calculator on it that puts a rupee number on your own missed calls.

    sharanjeetdigital.in/r/receptionist — the calculator is the fastest thing to try on this whole series.

  3. Three

    The comeback system

    The dead list every business is sitting on — old customers and old enquiries. Four agents, a tier rulebook that decides who's worth a message, and the actual WhatsApp messages that get replies.

    sharanjeetdigital.in/r/comeback — the block on who to never contact again is the important one.

  4. Four

    The WhatsApp desk

    One inbox with orders, questions, complaints and payments all landing in the same list. Five agents and a triage rulebook for all four categories, including where the bot is forbidden from answering.

    sharanjeetdigital.in/r/whatsapp — its payment rules and this page's payment block are two halves of the same job.

  5. Five

    The back-office agent — you're on it

    The paperwork. Four agents, the reconciliation rulebook above, the honest section on where extraction breaks, and the log sheet below.

    The closer, and the least glamorous of the five. Also the one with the clearest monthly cost sitting next to it.

Sheet

The reconciliation log — make a copy

Ten columns — date, doc_type, source, extracted_amount, expected_amount, variance, confidence, flag, assigned_to, status. Agent 1 fills the extraction columns, Agent 2 fills variance and flag, Agent 3 works off the matched rows, and Agent 4 sends whatever has a name in assigned_to.

Two of those ten columns are doing the real work. confidence is what stops a badly-read number reaching your books — everything under the threshold gets held. flag is what stops the alerts becoming noise, and it has four values, not two: MATCH, IGNORE, MISMATCH, REVIEW. Most log sheets people build only have match and mismatch, and that missing IGNORE value is exactly why their alerts get muted.

Seven example rows inside, across all three reconciliation types. Including a handwritten bill at 0.41 confidence sitting in the human review queue, and a ₹0.50 rounding difference marked IGNORE and closed with nobody told — because the format has to teach the restraint too, not just the checking.

It's worth running by hand before any agent touches it. Take one week of bills, fill this in yourself, and count how many rows end up MISMATCH. Most businesses find it's two or three a week and each one is worth more than they expected. That count is the number that tells you whether to automate this at all.

Make a copy of the reconciliation log

The link opens on 'Make a copy' — you get your own copy, mine isn't touched.

Next

If you want it set up on your own paperwork

Take the rulebook and the sheet as they are — the tolerances and the three checks work for most businesses with the numbers adjusted. If you want it set up on your own paperwork, message me and tell me two things: roughly how many documents come in a day, and what shape they arrive in — emailed PDFs, photos, or handwritten. That second answer decides almost everything, and if most of yours are handwritten I'll tell you straight that the input needs fixing before any of this is worth paying for.

Message me on WhatsApp

Straight to WhatsApp. No forms, no email.