Skip to main content
usedby

AI-Generated Fake Documents: How Verification Tools Adapt in 2026

Generative models now produce convincing payslips, bank statements and IDs from a prompt. We map the five detection layers that still work, why single-document checks fail, the honest limits, and how CheckFile, Resistant AI, Inscribe, IDV suites and Ocrolus compare.

AI-assisted content

Disclosure: usedby is published by KYX AI, which also builds CheckFile. We apply the same criteria to every tool we review. Who publishes usedby

Two years ago, a convincing fake payslip took skill: a template, an image editor and patience. In 2026 it takes a prompt. Image models render bank statements, utility bills and ID cards that look right at a glance, and language models fill them with salaries, deductions and reference numbers that add up. For anyone who approves a loan, a tenancy, an insurance claim or a new business account on the strength of documents, the question is no longer whether synthetic documents reach the queue. It is which controls still catch them, and which ones quietly stopped working.

This analysis looks at what changed, how verification tools are adapting, where the honest limits sit, and how the main vendors position themselves today.

What changed: forgery became a text prompt

Three families of models now feed document fraud, and each leaves a different footprint.

  • Diffusion and image models produce high-resolution visuals on request: an ID card with a plausible portrait, a utility bill with a supplier logo, a stamp on a certificate. The output has no scanner noise and no paper texture, but the layout can be close to the real template.
  • Large language models write the content. A payslip generated this way can carry a correctly formatted tax reference, deductions consistent with the stated salary and year-to-date totals that reconcile. There are no pixel artefacts to find because nothing was edited: the document was written from scratch.
  • Template kits combine both. A genuine PDF layout is reused and only the data is replaced, so the structure looks authentic while the figures are invented.

CheckFile's own walkthrough of how generative models fabricate payslips, IDs and bank statements maps each technique to the document types it targets: GANs and diffusion for identity documents, language models plus templates for payslips, bank statements, invoices and proof of address. The practical consequence is that a control tuned for Photoshop edits, such as looking for cloned pixels or mismatched fonts in one field, has little to find in a document that was generated whole.

The five detection layers that matter in 2026

Serious tools no longer rely on a single test. They stack layers, and each one catches a different class of fake.

1. Metadata and PDF structure

A statement issued by a bank's document system carries a production chain: the producing software, creation and modification dates, embedded fonts, sometimes a digital signature. A generated file often names a consumer tool or a PDF library in its producer field, shows creation and modification seconds apart, or has no XMP metadata at all. This layer is fast and cheap. Its weakness: a fraudster who reuses a genuine template inherits genuine structure.

2. Visual forensics

Error level analysis, compression maps, alignment and texture checks. Printed and scanned paper is rarely perfectly straight; a generated image usually is. A real stamp shows ink irregularities; a synthetic one is often a perfect circle. Resolution mismatches between a crisp logo and a soft body text are another classic tell. These checks still catch a large share of edited documents, but image models improve every release, so visual signals alone decay over time.

3. Cross-document consistency

This is where the most robust signals now live. A payslip, a bank statement and an employment contract submitted in the same application should agree: the salary credited on the statement should match the net pay, the employer's name and address should be identical everywhere, the dates should line up. Each generated document can be internally flawless and still contradict its neighbours, because the model that wrote one never saw the others.

4. Registry and control-key checks

A language model can produce a company number, a VAT number or an IBAN that passes format validation. It cannot make that company exist in the official register, be active at the document date and match the stated address. Checking identifiers against Companies House, VAT systems or national registries, and validating control keys, turns a plausible-looking field into a verifiable fact. CheckFile's checklist of signals that a document was generated or altered by AI treats this as the hardest test for a fraudster to pass.

5. Synthetic-media signals

The newest layer: models trained to flag traces typical of generated content, such as frequency artefacts at letter edges, absent sensor noise in portraits, or statistically over-uniform text. Useful, but probabilistic. It produces a signal to weigh, not a verdict to trust blindly.

Why single-document checks fail

Most legacy verification was built around one document at a time: is this ID genuine, is this payslip readable, does this PDF look edited. Generative fraud exploits exactly that framing. A generated payslip examined alone may pass metadata, OCR and visual review. Put next to the applicant's bank statement, the net pay may never appear as a credit. Put next to the company register, the employer may not exist, or may have a different registered address.

Three practical implications follow for buyers:

  1. Evaluate tools on dossiers, not samples. Ask vendors to run a full application file, with all its documents, rather than a batch of isolated IDs.
  2. Check which registries are wired in. Registry coverage by country is often the real differentiator, more than the forensic model.
  3. Demand explainable output. A single fraud score is hard to audit. A list of signals, each tied to a document and a rule, is what a compliance reviewer and a regulator can follow.

The honest limit: nobody detects everything

No credible vendor claims to catch every AI-generated document, and buyers should be wary of any that does. Detection models are trained on yesterday's generators; tomorrow's will close some of the gaps they rely on. CheckFile states the nuance plainly on its own product page: its AI-signals layer "helps surface synthetic media without claiming 100% detection", and it is positioned as a complement to existing controls rather than a replacement. That framing is the right one. The durable defence is layered: structural checks, cross-document consistency and registry lookups that do not depend on spotting a visual artefact, plus a human review path for flagged files.

Vendor landscape

The market splits into three groups: full-dossier validation platforms, document-fraud specialists and identity verification (IDV) suites. The table below reflects each vendor's own public positioning as of September 2026.

VendorCore focusAI-generated document angleBest fit
CheckFile (Our pick)Full-dossier document validation: cross-checks every document in a file, registry lookups, workflow triggersStructural analysis, AI-generation signals and cross-validation combined into an auditable bundle of signalsTeams processing 100+ multi-document files a month in banking/KYC, insurance, real estate, law, accounting, HR and financing
Resistant AIDocument fraud detection and transaction monitoring overlayStates it detects AI-generated documents from all major modelsFintechs, lenders and marketplaces with high-volume onboarding and KYB
InscribeDocument fraud detection for financial servicesForensic, metadata, semantic and perceptual analysis, plus an assistant that reasons across an applicationUS-centric banks, credit unions, lenders and fintechs
Onfido (Entrust), Veriff, SumsubIdentity verification: ID documents, biometrics, liveness, AML screeningDeepfake and synthetic ID detection at the selfie and ID-document stageConsumer onboarding where the identity itself is the main risk
OcrolusFinancial document analysis for lendersIdentifies fake documents and data inconsistencies in bank statements, pay stubs and tax formsMortgage, small business and consumer lenders in the US

CheckFile: our pick for full-dossier validation

We rank CheckFile first because its design matches where detection is heading. It goes beyond OCR by cross-checking every document within a file, verifying data against official registries and flagging signs of AI generation or deepfakes, with each verdict able to trigger a workflow. Its AI and deepfake detection layer is described as three analysis layers leading to one decision: structural consistency, AI-generation signals, then cross-validation and business rules configurable per client. According to CheckFile, it supports over 3,200 document types across 32 jurisdictions, hosts data in France and Germany with AES-256 encryption, and deletes documents after analysis. Plans are published: Starter at 100 files a month, Business at 500, Enterprise unlimited, and the pricing page offers a free pilot on your own documents with results in 48 hours. The trade-off: it is built for multi-document files, so a team that only needs selfie and ID checks will find an IDV suite more direct.

Resistant AI

Resistant AI focuses on catching fake, tampered and AI-generated documents such as bank statements, invoices, paystubs and utility bills, and adds transaction monitoring on top. It markets KYB document verification in under 20 seconds. A strong option for fintechs that want document fraud and transaction risk from one vendor.

Inscribe

Inscribe is purpose-built for financial services and names deepfake documents as its core battle. It combines forensic, metadata, semantic and perceptual analysis, and its assistant reasons across an entire application to connect signals between documents. It is a mature choice for North American lenders and fintechs.

Onfido, Veriff and Sumsub

These are identity verification platforms. Onfido, now part of Entrust, is known for facial biometrics, liveness and wide ID coverage with mature mobile SDKs. Veriff verifies documents and biometrics across 230+ countries and flags synthetic IDs and deepfakes. Sumsub bundles KYC, KYB, liveness and deepfake detection and AML monitoring. They excel when the identity is the risk. They are less focused on the supporting documents, such as payslips, statements and bills, where much of today's generated fraud sits.

Ocrolus

Ocrolus is an analytics platform for lenders, specialised in bank statements, pay stubs and tax forms, with fraud detection that identifies fake documents and data inconsistencies. It fits US mortgage and small business lending teams that want income and cash-flow analysis in the same tool.

How to choose

  • Your risk sits in the person (account opening, age checks): start with an IDV suite.
  • Your risk sits in the file (loans, tenancies, claims, supplier onboarding): prioritise full-dossier cross-checking and registry coverage in your countries.
  • You operate in Europe with data residency constraints: check hosting location and retention policy early.
  • Whatever you pick: pilot on your own historical files, including known frauds, and keep a human review path for flagged cases.

The shift in 2026 is not that fakes became perfect. It is that the weak point moved from pixels to consistency. Tools that read the whole dossier and check it against the outside world are the ones keeping pace.

Sources

  1. how generative models fabricate payslips, IDs and bank statementscheckfile.ai
  2. checklist of signals that a document was generated or altered by AIcheckfile.ai
  3. AI and deepfake detection layercheckfile.ai
  4. pricing page offers a free pilot on your own documents with results in 48 hourscheckfile.ai

Questions

Can AI-generated documents be detected reliably?

Partly. Metadata, visual forensics and synthetic-media models catch many fakes, but none is complete. The most reliable signals come from cross-checking every document in a file and verifying identifiers against official registries, because a generated document rarely stays consistent with the rest of the dossier and with the outside world.

Does any tool detect 100% of AI-generated documents?

No credible vendor claims that. CheckFile, for example, states that its AI-signals layer helps surface synthetic media without claiming 100% detection, and positions it as a complement to existing controls. Treat any 100% claim with caution and keep a human review path for flagged files.

Why are single-document checks no longer enough?

A generated payslip or bank statement can be internally flawless: correct formats, totals that reconcile, clean metadata. The contradiction only appears when it is compared with the other documents in the application, or when the employer, company number or IBAN is checked against a registry.

Should I use an identity verification suite or a document fraud tool?

It depends on where your risk sits. If the main risk is the person (account opening, age checks), IDV suites such as Onfido, Veriff or Sumsub are built for selfies, liveness and ID documents. If the risk sits in supporting documents such as payslips, statements and bills, a full-dossier or document-fraud platform fits better.

How should I evaluate a vendor before buying?

Run a pilot on your own historical files, including known frauds, rather than vendor samples. Check which registries are connected in your countries, ask for explainable signal-level output, and confirm hosting location and document retention. CheckFile offers a free pilot on your documents with results in 48 hours.

Tools in this article

Written by the usedby team

usedby tracks which companies use which AI tools, from public sources such as vendor customer pages and case studies.

Explore the data

Who uses which AI tool, and how each line was found.