Article
AI-Generated Fake Invoices Are Now Passing AP Review. Here's What Catches Them
AI-generated fake invoices pass accounts payable (AP) review because most review checks whether a document looks right, not whether it matches an agreement. A fabricated invoice can be typeset perfectly, use a real vendor's logo, and still fail the moment it is checked against the contract, statement of work (SOW), or rate card behind it - because it cites a purchase order (PO) that doesn't exist, bills a rate the agreement never allowed, or asks for payment to an account the vendor never used.
That distinction matters more this year than it did two years ago. Generative AI tools have made it cheap to produce a convincing PDF, a matching email thread, and even a voice on a phone call confirming "yes, please update our bank details." None of that changes what an invoice has to be true against: the agreement the two parties actually signed. This piece walks through why a visually flawless invoice still fails an agreement check, which fraud signals survive AI polish, what manual review typically misses, and a checklist for the next invoice that looks a little too clean.
What changed: are AI-generated invoices actually reaching AP teams now?
The scale of the underlying problem is not new - fraud against accounts payable has always been a real cost. The FBI's Internet Crime Complaint Center (IC3) reported that business email compromise (BEC), the scam category that includes fake-vendor and fake-invoice schemes, caused close to $2.77 billion in reported losses in the United States in 2024 alone, across more than 21,000 complaints. That is the scale AP teams have been defending against for years, before generative AI entered the picture.
What has changed is the cost of producing a convincing fake. A forged invoice used to carry small tells: a slightly wrong logo, stilted English in the cover email, a font that didn't match the vendor's usual template. Generative AI tools remove most of those tells. A fabricated invoice, a supporting purchase order, and a plausible email thread can now be produced in minutes, in a vendor's usual tone, with a template that looks identical to the real one. That is a genuine shift in AP risk, even without a single new confirmed statistic to name it: the defense that used to work (a sharp-eyed reviewer noticing a document "looks off") stops working when nothing looks off.
Why does a flawless-looking fake invoice still fail an agreement check?
A polished invoice only has to survive whatever check is applied to it. If the check is "does this look like a real invoice from this vendor," AI-generated fraud is built to pass exactly that test. If the check is "does every line on this invoice match the contract, SOW, or rate card we actually signed with this vendor," the invoice has to survive a much harder bar, and most fabricated invoices cannot, because the fraud is built from the outside in - it mimics what an invoice looks like, not what the underlying agreement says.
Concretely, agreement-level verification checks things a visual review cannot see just by looking at the PDF:
- Does the invoice cite a PO number that exists in your system, for this vendor, with unspent balance left on it?
- Do the billed rates match the exact rate card in the SOW, rather than a rate that merely looks reasonable?
- Is the scope of work billed inside what the SOW actually covers, or has it quietly expanded?
- Is the bank account on file the one this vendor has used before, or did it change without the usual verification?
A fabricated invoice can nail the first three of these by copying an old, real invoice closely enough to fool a human skimming it. It is far harder to fabricate a PO number that happens to exist and have unspent balance, or a rate that happens to match a rate card the fraudster has never seen. That is the practical reason agreement-level checks outperform look-and-feel review: they are checking against a document the fraudster does not have, not against a template they can copy.
Which fraud signals survive AI polish?
Four signals keep showing up in AP fraud regardless of how good the fake document looks, because they are structural, not cosmetic:
- A changed bank account. AI can write a perfectly plausible "we've updated our banking details" email. It cannot make the new account match the vendor's payment history without a verification call to a known contact, not the number on the email.
- A payee mismatch. The name on the invoice, the name on the contract, and the name on the bank account don't line up. This is a paperwork problem a generated PDF does not fix.
- Threshold-parking. An amount set just under whatever dollar figure triggers a second approver. This pattern is about the number, not the document, so no amount of visual polish disguises it.
- A PO that doesn't exist. The invoice cites a purchase order number, but no such PO exists in your system, or it exists for a different vendor or a different, already-exhausted amount.
None of these require judging whether a document "looks real." They require checking the invoice against records the fraudster does not control. That is the same logic behind matching invoices at the contract and rate-card level, not just the PO, and it is worth reading how a real fake-invoice incident played out end to end in what a $67,000 fake invoice scam should teach every AP team.
What does manual AP review miss that AI-generated fraud exploits?
Manual review is built around pattern recognition: a reviewer has seen hundreds of real invoices from this vendor, and something about a fake one usually feels wrong. Standard optical character recognition (OCR) tools used in many AP workflows extract the text and numbers off a document, then check format and math, not whether the deal terms are real. Both approaches share the same blind spot: they evaluate the invoice as a document, not as a claim against an agreement.
That blind spot is exactly what AI-generated fraud is built to exploit. A generated invoice is optimized to be internally consistent and visually normal - correct math, correct formatting, a vendor name spelled right. It is not optimized to match a specific SOW's rate card, because the fraudster usually does not have that document. So the invoice sails through a check built for "does this look right" and stalls the moment the check becomes "does this match what we agreed to." Occupational fraud broadly stays a large, ongoing cost for exactly this reason: the Association of Certified Fraud Examiners' 2024 Report to the Nations puts the typical organization's annual fraud loss at roughly 5% of revenue, a figure that predates the current wave of AI-generated documents and shows how much room ordinary review already leaves open.
It is also worth remembering that fraud is not the only source of leakage a document-only review misses. Duplicate and erroneous payments, with no fraudulent intent at all, run 0.8% to as high as 1.5-2% of total annual disbursements even at well-run AP shops, according to APQC's Open Standards Benchmarking research. A review process that only asks "does this invoice look legitimate" will not catch those either, because they are legitimate-looking invoices with a legitimate vendor name attached to the wrong amount, or a repeat of a bill already paid.
What should you check before approving an invoice that looks a little too clean?
Use this order when an invoice raises even a small doubt, before it gets approved:
- Pull the PO it cites. Confirm the PO exists, belongs to this vendor, and has enough unspent balance to cover the invoice. If the PO number does not resolve to a real record, stop there.
- Check the rate against the SOW or rate card. Not "is this rate plausible," but "is this the exact rate this contract specifies for this role or deliverable." Checking contractor hours against the SOW before you pay is the same discipline applied to time-based billing.
- Confirm the bank account independently. Call a known contact at the vendor, using a number you already have on file, not one from the invoice or the email that flagged the change.
- Compare the scope billed to the scope contracted. A SOW that covers a fixed deliverable should not quietly grow into open-ended hours without a change order behind it.
- Look for threshold-parking. If the amount sits suspiciously close to your next approval tier, treat that as a reason to slow down, not a reason it "must be fine since it's small."
Comparison: what visual AP review checks versus what agreement-level verification checks
| Check | Visual / OCR-based AP review | Agreement-level verification |
|---|---|---|
| Vendor name and logo | Confirms it looks right | Confirms it matches the vendor on file for the cited contract |
| Line-item math | Confirms totals add up | Confirms totals add up |
| Rate charged | Not checked against contract terms | Checked against the exact rate card in the SOW |
| PO reference | Confirmed present on the document | Confirmed to exist, belong to this vendor, and have unspent balance |
| Scope of work billed | Not checked against contract scope | Checked against what the SOW actually covers |
| Bank account | Confirmed present, not confirmed correct | Flagged if it differs from the vendor's payment history |
| Approval-threshold patterns | Not typically monitored | Flagged when an amount sits just under an approval tier |
FAQ
Can AI-generated invoices fool standard OCR-based AP tools?
Yes. Standard OCR-based AP tools extract text, numbers, and formatting from a document and check that the math is internally consistent. They were not built to verify that an invoice's rate, scope, or PO reference matches a specific contract, SOW, or rate card, so a fabricated invoice that is internally consistent can pass an OCR check even though it fails an agreement-level check.
What is the single biggest tell that an invoice is fabricated, even if it looks perfect?
A PO reference that does not resolve to a real, matching record in your system - either the PO number doesn't exist, belongs to a different vendor, or has no unspent balance left to cover the amount billed. This tell survives AI polish because the fraudster typically does not have access to your PO records.
Should a changed bank account on an invoice ever be trusted without a separate call?
No. A bank account change communicated only through an invoice or an email should always be confirmed by calling a known contact at the vendor, using a phone number already on file, not one provided in the message that requested the change. This is the standard control against payee-mismatch and account-takeover fraud.
Does agreement-level verification replace the AP team's judgment?
No. It surfaces the specific mismatches (a PO that doesn't exist, a rate outside the contract, a scope that has grown) so the reviewer can make the call faster, with the clause behind each flag already identified, rather than replacing the human decision to approve or reject a payment.
One next step
If your team is fielding more invoices that look right but nag at you, the fastest way to find out how much agreement-level checks would catch is to run them against your own real invoices and agreements, not a demo dataset. Quittance checks every line against the contract, SOW, and rate card behind it, flags the exact clause a mismatch violates, and posts clean invoices to Xero as a draft bill for a person to approve - it never pays anything on its own. A pilot on your real spend data will show which of this month's invoices would have failed an agreement check, and how many are still sitting in the categories the checklist above describes.