Skip to content

Latest commit

 

History

History

README.md

Find and replace text in a PDF with DsPdfJS

This sample searches a PDF for a term and rewrites every occurrence with Document Solutions for PDF JS (DsPdfJS).

Unlike drawing a box or new text on top of a page, PdfDocument.replaceText() edits the document's actual content stream. The change is therefore visible both in the rendered page and in the text you extract afterwards - the sample confirms this by counting occurrences of the old and new text once the replacement is done.

What it shows

  • Loading a PDF into a PdfDocument.
  • Counting occurrences with PdfDocument.findText().
  • Replacing every occurrence with PdfDocument.replaceText(), optionally overriding the font size.
  • Matching case-sensitively or case-insensitively.
  • Rendering before/after page previews with PdfPage.saveAsPng().
  • Verifying the edit by re-searching the document for the old and new text.
  • Saving the edited PDF entirely in the browser.

Run the sample

npm install
npm run dev

Open the local URL printed by Vite. The sample loads public/text-replace-sample.pdf automatically; you can also upload another PDF.

The predev script copies the DsPdf.wasm file shipped by @mescius/ds-pdf into public (the copied file is ignored by Git). The same step runs automatically before a production build:

npm run build

How the sample works

  1. Load the PDF and count how many times the search term appears with findText().
  2. Call doc.replaceText({ text, matchCase }, newText, undefined, null, fontSize) to rewrite every occurrence. Passing null for the font keeps the current font; passing null for the font size keeps the current size.
  3. Render each page to PNG for the "after" preview.
  4. Search again for the old and new text and report the counts, confirming the replacement.
  5. Offer the edited PDF for download.

By default the sample replaces Acme Corp with Globex Inc in the bundled quotation.

The bundled PDF is generated by tools/create-sample.mjs and contains only obvious fictional data. Regenerate it after editing that file with:

npm run make-document

A note on text recognition

replaceText() locates text using DsPdfJS's current recognition algorithm (PdfDocument.recognitionAlgorithm). On complex PDFs the result can depend on how the text is stored in the content stream, so a match that reads naturally to a human is not always a single replaceable run. The bundled sample document uses simple, plain text so the demonstration behaves predictably; when replacing text in arbitrary PDFs, review the output.

License key

Without a license key, DsPdfJS runs in trial mode and adds an evaluation notice to generated or rendered output. To use a key, call await DsPdfConfig.setLicenseKey("YOUR_KEY") before connectDsPdf() in src/Demos.ts.