Skip to content

About

Deterministic PDF-to-editable-PowerPoint reconstruction with reviewable OCR/AI inputs and IguanaTex-compatible math workflows

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

editable-slides

English | 한국어

Editable PDF-to-PowerPoint reconstruction with reviewable OCR/AI inputs and IguanaTex-compatible math workflows.

editable-slides rebuilds PDF slide decks as editable PowerPoint presentations. It preserves page geometry, text, raster images, and vector objects, and provides additional workflows for mathematical notation using native Office Math or IguanaTex-compatible LaTeX metadata.

The core converter is local and deterministic: it does not require an AI or OCR service. For difficult pages, its JSON catalogs are explicit integration points for human-reviewed suggestions from OCR, math-recognition, or AI reconstruction tools. This keeps recognition optional and makes the accepted result auditable.

Features

  • Extract portable PDF geometry manifests with image z-order information.
  • Reconstruct editable text, images, vectors, charts, diagrams, and code blocks.
  • Convert equation images to transparent PNG fallbacks.
  • Attach editable LaTeX source metadata compatible with IguanaTex workflows.
  • Accept reviewed catalog entries prepared manually or with external OCR/AI tools.
  • Support text-derived formulas with independent display geometry and fit rules.
  • Validate OOXML structure, equation coverage, transparency, and relationships.
  • Compare source and LibreOffice renders and detect catastrophic content loss.
  • Record a slide-by-slide raster audit for labels, arrows, formulas, charts, diagrams, and other semantic content embedded inside images.
  • Remove document properties, thumbnails, timestamps, and author identities.
  • Produce privacy-audited release packages with SHA-256 manifests.

Status and scope

PDF is a final-layout format, so reconstruction is necessarily heuristic. Text and basic geometry can often be recovered automatically. Equations, charts, complex diagrams, clipping masks, and unusual fonts may require reviewed catalog entries or manual correction. The tool does not bypass encryption or permissions on input documents.

Visual similarity is not proof of editability: a screenshot can match the PDF while keeping every label and arrow trapped in one image. Releases should use the object-level editability audit in addition to structural validation and full-slide rendering.

OCR and AI models are not bundled, invoked, or required by the core CLI. There is currently no one-command adapter for a particular model or hosted service; external suggestions must be reviewed and expressed in the documented catalogs.

Release gate

A green render comparison is necessary but insufficient. Release only when structural validation passes with zero unreviewed formula candidates, every raster is classified, every slide is reviewed, and representative labels, arrows, and formulas have been edited in Microsoft PowerPoint.

Formula coverage checks inspect transparent grayscale images and opaque PNGs on light canvases. Unreviewed candidates fail validation and include xref, asset path, page occurrence, instance ID, bbox, and pixel/component evidence in the structural report. Each candidate must become a reviewed formula or receive a specific documented raster exception. OCR and math recognition are proposals, not authoritative LaTeX sources.

The release workflow must render and compare the complete final deck, then review every slide and perform object-level PowerPoint checks. For IguanaTex, edit and regenerate a representative equation, save, and reopen the file. See Editability audit and Quality assurance.

Requirements

  • Python 3.11 or newer
  • LibreOffice for full-deck rendering
  • Poppler (pdftoppm) for PNG page renders
  • LaTeX and dvipng for text-derived equation rendering
  • IguanaTex in PowerPoint when editing embedded LaTeX interactively
  • Tesseract when discovering raster-diagram labels locally

LibreOffice, Poppler, LaTeX, dvipng, IguanaTex, and Tesseract are optional unless the associated workflow is used.

Install

python -m venv .venv
. .venv/bin/activate
python -m pip install -e .

The installed command is editable-slides. Commands can also be run from a checkout with PYTHONPATH=src python -m editable_slides.

Dependency licensing: PDF extraction uses PyMuPDF/MuPDF, which upstream offers under the GNU AGPL v3 or a commercial license. Running and distributing the complete installed application may therefore create obligations beyond this repository's Apache-2.0 license. Review the upstream terms and Third-party notices for your use case.

Quick start

editable-slides extract input.pdf \
  --manifest build/manifest.json \
  --images build/assets/images \
  --references build/assets/reference \
  --reference-scale 2

editable-slides build \
  --manifest build/manifest.json \
  --formulas formulas.json \
  --math-mode iguanatex \
  --output build/output.pptx

editable-slides sanitize build/output.pptx

editable-slides validate \
  --pptx build/output.pptx \
  --manifest build/manifest.json \
  --formulas formulas.json \
  --math-mode iguanatex \
  --privacy \
  --report build/structural.json

Synthetic example

The repository contains a generator for a small, original technical slide deck covering an image-derived integral and text-derived fraction, matrix, optimization, and multi-line equations. It exercises formula suppression, independent display geometry, fit, and anchor settings. It does not contain third-party course material, paper figures, or photographs.

make synthetic-example

For the full LibreOffice render and visual comparison:

make privacy-safe-release

Mathematical notation

Three math modes are supported:

  • image: transparent equation image without editable math metadata.
  • iguanatex: transparent fallback image plus embedded LaTeX source metadata.
  • native: fallback image plus native Office Math markup where supported.

Text-derived formula entries may separate the reviewed source region from the PowerPoint display geometry:

{
  "id": "gaussian-integral",
  "page": 3,
  "source_bbox": [70, 110, 430, 180],
  "display_bbox": [90, 115, 410, 175],
  "fit": "contain",
  "anchor": "left-center",
  "latex": "\\int_{-\\infty}^{\\infty} e^{-x^2}\\,dx = \\sqrt{\\pi}"
}

See Formula workflow and Catalog schema.

OCR, AI, and IguanaTex interoperability

OCR or AI-assisted reconstruction can be used upstream to propose text, equation LaTeX, bounding boxes, charts, diagrams, or code blocks. The reviewed results enter the deterministic build through versionable JSON catalogs rather than an opaque model call. In iguanatex mode, accepted LaTeX is embedded with an editable fallback object for subsequent interactive editing in PowerPoint.

Flat diagrams can use the auto-editable override handler to turn color regions into editable freeforms and labels into text boxes. Store reviewed labels in overrides.json for repeatable output, or omit them to run the optional local Tesseract adapter. Accepted label pixels are masked before the text boxes are created so duplicate raster text cannot remain underneath. labels: [] disables OCR deterministically; a missing Tesseract executable or a 45-second timeout is an explicit error rather than a silent raster fallback.

Transparent sources are composited safely and dominant light-neutral canvases are discarded instead of becoming black or light slide-covering shapes. The largest 150 vector regions are retained by default so raster specks cannot create thousands of unusable freeforms. Small arrowheads or marks can therefore require manual review. Vectorized arrows are independently selectable freeforms, not guaranteed semantic PowerPoint connector objects. See Catalog schema.

This division is intentional: recognition tools suggest semantic content; editable-slides records placement decisions, builds the PPTX, sanitizes its metadata, and validates the result. See Architecture, Catalog schema, and Related work.

Privacy

sanitize removes OOXML document-property parts, package thumbnails, ZIP entry timestamps, and author/person attributes. privacy-audit reports remaining package-level identity metadata.

This does not redact visible slide content. Names, email addresses, speaker notes, and other information displayed on a slide must be reviewed separately. See Privacy.

Responsible use

Users are responsible for ensuring that they have permission to process and redistribute their input documents and generated presentations. This software does not grant rights to input PDFs, embedded media, fonts, or output content.

Do not process untrusted PDFs, LaTeX, or office documents in a privileged environment. The integration workflow invokes external applications.

Development

make test
make audit-public
python -m build

See Contributing and Security.

Third-party software

This project builds on independently maintained open-source libraries and can interoperate with optional external tools. Their roles, official project links, and license identifiers are recorded in Third-party notices. Those projects are not bundled in this source repository unless a package installer obtains them separately.

License

The source code and original documentation in this repository are licensed under the Apache License 2.0. Input documents and generated presentations remain subject to their respective rights and licenses. Dependencies retain their own licenses; in particular, the PyMuPDF/MuPDF runtime is AGPL-3.0-or-later or commercially licensed by its upstream vendor.

About

Deterministic PDF-to-editable-PowerPoint reconstruction with reviewable OCR/AI inputs and IguanaTex-compatible math workflows

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages