Back

PDF comparison

Document tools

PDF workspace

Loading

Loading tool

The tool is loaded only when you open it.

All processing for this tool happens in your browser. Your input is not sent to a server.

About this tool

Compare a reference and candidate PDF on a shared physical page scale. Review both pages, the half-opacity overlay and magenta differences. Optional extracted-text changes and a JSON report complement the visual comparison.

Common uses

  • Review changes between two versions of a PDF.
  • Align pages after an inserted or removed page and inspect unmatched pages.

How to use it

  1. 1.Choose one reference and one candidate PDF, or load the local example. Set the candidate page offset and pixel tolerance; optionally enable extracted-text comparison.
  2. 2.Compare. Every physical page appears once, including unmatched pages. Review the four actual images for each pair, dimension warnings and optional text changes.
  3. 3.Download the JSON report after all pairs succeed. Files, settings, reset, language, clear and cancel invalidate old results.

Executed local examples

Aligned local versions

Load the 3-page reference/candidate example; offset 0, tolerance 12, extracted text enabled.
3 pairs: pair 1 has zero changed pixels; pair 2 moves the colored rectangle; pair 3 changes Original paragraph to Edited paragraph.

Keep unmatched pages

Use the same 3-page example with offset 1.
4 pairs: unmatched candidate page 1; reference 1 with candidate 2; reference 2 with candidate 3; unmatched reference page 3.

Limits and notes

  • Two PDFs, each 10 MiB and 1–20 pages; up to 40 pairs. Integer offset -19–19 and RGB tolerance 0–64. Preview edge 512px, at most 262144 pixels per pair, each PNG 1 MiB and all previews 24 MiB. JSON report 2 MiB. Reading, inspection and rendering have cancellable 30-second phase deadlines. PDF graph: 10000 indirect objects, 50000 visits, 32 direct levels; decoder memory is not fully bounded by raster caps.
  • Candidate page = reference page + offset. Positive offset aligns reference page 1 with a later candidate page. Unmatched physical pages are compared against white; positions with neither page are omitted. Each pair uses one scale derived from its largest page edge, with top-left alignment and white padding. CropBox and quarter-turn rotation are respected; unequal dimensions or rotations are reported separately.
  • A pixel differs when the largest absolute RGB channel difference is strictly greater than the tolerance, after compositing transparency on white. The overlay averages the two colors equally; differences are magenta on white. Antialiasing, fonts, readers and compressed resources can produce differences. Low-resolution matching does not prove identical files, semantic equivalence, preservation or full-page-detail equality.
  • Optional PDF.js text extraction is not OCR. Text items are concatenated in extraction order with hasEOL line breaks; normalization is disabled and spatial word gaps are not inferred. Each page: 10000 items, 4096 UTF-16 units and 200 lines; both inputs total 100000 units. Exact line LCS shows same, removed and added lines. Reading order, ligatures, hidden text and font mapping can mislead. Controls and bidi controls are visibly escaped in text previews; JSON escapes bidi controls and contains extracted text when enabled.
  • Encrypted PDFs, all AcroForm/XFA forms, signatures/permissions, JavaScript and unsafe document/page/additional actions reject. UserUnit and unsupported geometry reject. Ordinary URI/GoTo links may render without fetching or executing their targets. Only installed bundled PDF.js fonts/CMaps can be fetched. Original PDFs are not modified; metadata, attachments and hidden content are not compared or sanitized. No upload or input persistence.

Frequently asked questions

Does zero changed pixels prove the PDFs are identical?

No. It means the paired rendered pixels match at this resolution and tolerance. Metadata, attachments, hidden text and fine details may differ; font rendering can also cause false positives.

How does the page offset work?

Candidate page = reference page + offset. An offset of 1 aligns reference page 1 with candidate page 2. Candidate page 1 remains visible as an unmatched pair; all physical pages are included.

Does extracted-text comparison read scanned documents?

No OCR is performed. Only the PDF text layer is extracted. Extraction order, ligatures and hidden text can differ from the visible page.

Related tools