Skip to content

Repository files navigation

PaperScan

Turn phone photos of documents into proper scans.

Automatic page detection · perspective correction · white balance · background trimming Lossless PNG at native resolution · fully offline · cross-platform

CI License: MIT .NET 8

English · 简体中文


Before — a photo After — a scan
A phone photo of a certificate on a dark desk The rectified, white-balanced, trimmed scan

Both images are synthetic and are regenerated by the test fixtures. The photo is perspective-skewed, tinted, and has a drop shadow and sensor noise; the scan is what paperscan produced from it with no arguments at all.

Two ways to use it

Open it in a browser Nothing to install. Drop a photo on the page, get a scan back. Works on a phone, including the camera. The photo is decoded, rectified and re-encoded on your device — there is no server, nothing is uploaded, and it works offline after the first visit. Drag the corner handles if the detector mis-frames your page, and pick the right orientation from four thumbnails.
Install the CLI Same algorithm, scriptable and batchable, writes lossless PNGs at full resolution. For when you have fifty photos and a shell.

Both share the same five-stage pipeline and the same regression target; see web/README.md for why there are two implementations of it.

The web URL above is a placeholder — see Deploying the web app to fill it in.

What it does

  • Finds the page for you. No cropping, no clicking corners. It thresholds the photo down to "bright and slightly warm" pixels and reads the page's corners off the resulting blob — including documents with a closed printed border, which trip up the obvious "largest bright blob" approach.
  • Rectifies the perspective. A homography maps the page quad onto a rectangle. The output is sized from the page's own edge lengths, so it is a 1:1 resample at the original pixel density — nothing is upscaled, nothing is thrown away.
  • Balances the white. The 90th percentile of each channel is taken as "paper white" and stretched to 250. Indoor yellow cast disappears; the red of a wax seal or the colours in a logo stay put.
  • Trims the background. Probes inwards from all four edges, finds where the paper actually starts, and crops. The result has no dark border and no white border.
  • Writes a lossless PNG. 24-bit RGB, no alpha, no JPEG re-encoding, no resampling beyond the corner-detection thumbnail.

Install

Download a release

Grab the archive for your platform from Releases, unpack it, and run paperscan. The self-contained builds need no .NET installation.

As a .NET tool

dotnet tool install --global PaperScan.Cli

From source

git clone https://github.com/YangRainmorning/PaperScan.git
cd PaperScan
./scripts/build.ps1 -Publish -SelfContained

That produces:

  • dist/paperscan.exe on Windows
  • dist/paperscan on Linux and macOS

Windows drag-and-drop

Once built, drag any number of photos onto scripts/paperscan.cmd. If no build exists yet it will create one first.

Usage

paperscan photo.jpg

That is the whole thing. Three files appear next to the photo:

photo-scan.png      the scan
photo-verify.png    detected corners drawn on the photo, for checking
photo-preview.png   downscaled copy, for looking at quickly

Batch mode works too:

paperscan *.jpg --quiet

Options

Option Default Meaning
-o, --out <file> <photo>-scan.png Output path. Single input only.
-c, --corners <quad> detected Manual corners, x,y;x,y;x,y;x,y in reading order: top-left;top-right;bottom-right;bottom-left
-r, --reading-edge <edge> top Which edge of the photo the page's readable "up" points at: top, right, bottom, left
-m, --margin <ratio> 0 White border, as a fraction of the page size
--no-trim off Keep the raw rectified canvas, skip background trimming
--no-verify off Do not write the corner overlay
--no-preview off Do not write the preview
--preview-width <px> 1600 Preview width
--detect-long-edge <px> 1024 Long edge of the thumbnail used for detection
--png <effort> fast PNG encoder effort: fast, balanced, small
-l, --lang <lang> auto UI language: auto, en, zh
-q, --quiet off Only report errors
-h, --help Show help
-V, --version Show version
Exit code Meaning
0 Success
1 At least one image failed
2 Bad usage

Getting the corners right

The page came out sideways. The photo was taken in portrait with the page lying across it. Tell the tool which photo edge the page's readable "up" points at:

paperscan photo.jpg --reading-edge right

Detection picked the wrong thing. Open photo-verify.png: the green quad shows what was detected. If it does not hug the page, read the pixel coordinates of the four corners off the photo (any image editor will do) and pass them in:

paperscan photo.jpg --reading-edge right \
  --corners "5834,456;5789,7975;489,7975;424,520"

The order is always reading order — top-left, top-right, bottom-right, bottom-left as you would name them looking at the page the right way up.

A white border is wanted. --margin 0.03 adds 3% of the page size on each side.

How it works

  1. Corner detection — downscale to a 1024px thumbnail, mark "bright, warm" pixels, group them into connected blobs, seed on the blob with the largest bounding box, merge in everything that overlaps it, then read the corners off the extremes of x+y and x−y.
  2. Rectification — solve an 8×8 linear system for the homography taking the unit square to the page quad, then walk the output canvas and inverse-map every pixel back to a bilinear sample of the source.
  3. White balance — build per-channel histograms on the canvas, take the 90th percentile as paper white, scale each channel so it lands on 250.
  4. Framing — draw the page polygon on a white canvas so the page quad's edges are cleanly cut (this is where the optional margin comes from).
  5. Trimming — fire 120 probes per edge inwards, find the first paper-bright pixel confirmed by a second sample, and crop to the deepest result.

docs/algorithm.md has the details, including why each parameter is what it is.

Performance

Measured on a 6144×8192 (50 MP) phone photo of a certificate, producing a 7450×5290 PNG:

Stage Time
Decode 0.6 s
Rectify + white balance 0.3 s
Corner detection < 0.1 s
Background trim 0.1 s
PNG encode (--png fast) 7 s
Total ~10 s

The rectification and white-balance loops are parallelised over rows. Results are bit-identical to a serial run, because the per-pixel work is independent and the histogram accumulators are integers.

--png balanced trades about 30 s for roughly 10% smaller files. --png fast is the default.

Limitations

  • A closed printed border used to break detection; it no longer does. If a document has a border, a large logo, or heavy decoration that the detector still mis-reads, use --corners.
  • All four page corners must be in frame. A cropped corner cannot be recovered.
  • Even lighting matters. A hard shadow across the page can pull the detector inwards; --corners is the escape hatch.
  • The input is the quality ceiling. The tool works on the pixels the camera produced and adds no loss of its own, but it cannot invent detail that is not there.
  • Output files are large. A 40 MP scan of a mostly-white page is 50–55 MB of PNG. That is the price of keeping every original pixel; re-encode to JPEG yourself if you need something smaller.
  • Handwriting and stamps are not cleaned up. No despeckling, no shadow removal, no OCR. It is a rectifier, not a document restorer.

Development

Requires the .NET 8 SDK (or newer — the projects target net8.0 with RollForward=LatestMajor, so a newer runtime is fine) and, for the web app, Node 24.

dotnet test                          # 44 tests
./scripts/build.ps1                  # restore, build, test
./scripts/build.ps1 -Publish -SelfContained
./scripts/build.ps1 -Publish -Runtime linux-x64 -OutputDirectory dist/linux

cd web
npm ci && npm test && npm run build  # 45 tests, then a ~34 KB site in web/dist
npm run dev                          # http://localhost:5173

The repository layout:

src/PaperScan.Core/     the engine — detection, homography, warping, trimming
src/PaperScan.Cli/      the paperscan command-line tool
tests/PaperScan.Tests/  xunit suite, including synthetic-photo end-to-end tests
web/                    the browser app — a dependency-free TypeScript port of Core
scripts/                build and drag-and-drop helpers
legacy/                 the frozen v0 PowerShell implementation this was ported from
docs/                   algorithm notes and README images

web/ has its own README covering the architecture, the deliberate deviation from the C# defaults, and how the browser path is smoke-tested.

See CONTRIBUTING.md.

Deploying the web app

web/dist is a static folder, so anything that serves files will do. Cloudflare Pages is the documented default because github.io is often slow or unreachable from mainland China:

  1. Cloudflare dashboard → Compute → Workers & Pages → Create → Pages → Connect to Git. Cloudflare moved this entry under Compute; on older dashboards Workers & Pages is a top-level sidebar item.
  2. Pick this repository.
  3. Root directory web, build command npm ci && npm run build, output directory dist.
  4. Environment variable NODE_VERSION = 24.
  5. Deploy. Pushes to main redeploy automatically.

Then replace paperscan in this README (and in README.zh-CN.md) with the assigned *.pages.dev hostname. web/public/_headers is picked up as-is and sets a strict CSP and long-lived caching for the hashed assets.

For GitHub Pages instead, set base to '/PaperScan/' in web/vite.config.ts and publish web/dist.

A note on the portable reference implementation

legacy/ contains the original single-file PowerShell tool. It is frozen — no fixes, no features — but it still runs on a bare Windows box with nothing installed, and it is the implementation the C# port was validated against. See legacy/README.md.

License

MIT — see LICENSE.

PaperScan depends on SixLabors.ImageSharp, which is licensed under the Six Labors Split License. It is free for open-source projects; commercial use may require a paid licence from Six Labors. If that is a problem for your use case, the code only needs an image decoder and encoder, and swapping in a permissively licensed one is a contained change.

The web app is unaffected: it uses the browser's own decoders and encoders, and web/src/core/ has no third-party dependencies at all.

About

Turn phone photos of documents into clean, rectified, white-balanced scans. .NET 8 CLI, fully offline, cross-platform.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages