Scanned PDF Cleaner
Straighten crooked pages, clear scanner borders, remove specks and blank sheets, cut out anything that should not be there, and make the result searchable. Everything happens in your browser, so the document is never uploaded.
Why Use Our PDF Cleaner?
It Measures By Actually Doing It
Every defect is weighed by performing the fix on a copy and seeing what came off, not by guessing from how dark an edge looks. That distinction is not academic: a black bar across the top of a page scored nine percent on the guess and was left alone, while the honest measure put a quarter of that page’s ink in it.
Untouched Pages Stay Untouched
A page you did not ask to change is copied across byte for byte, keeping its original image data and its text layer. Re-encoding an already-efficient scan makes it 70% to 160% larger, so the pages that need nothing are given nothing.
Searchable When It Is Finished
A straightened or trimmed page has its existing text moved to match its new geometry. A page that never had a text layer is read with Tesseract, on your own machine, and given an invisible one — so the finished file can be searched and copied from.
You See The Reason For Everything
Each proposal comes with the measurement behind it, in plain words, and every one can be overridden per page. Paint out anything that should not be in the document, draw a crop, or put a page back exactly as it arrived.
Nothing Leaves Your Machine
There is no upload step and no server. Even the recognition engine and its language model are served from this site rather than a third-party CDN, so no request carrying any part of your document is ever made.
Drop a scanned PDF here
It is read inside this browser tab. Nothing is uploaded, and no copy of it exists anywhere but on your own machine.
A Clean Scan in Four Steps
From a crooked, speckled photocopy to something worth archiving
Open the Scan
Drop in the PDF your scanner produced. It is read inside this tab — there is no upload, so a contract or a medical record never leaves the machine it is sitting on.
See What It Found
Every page is measured and cleaned automatically: black bars and edge lines cleared, scuffs and shadows wiped out of the margins, dirt removed, crooked pages straightened. You see the cleaned page, and can hold a button to compare it against the original.
Change Anything You Disagree With
Each page says what was found and what is being done about it, in plain words, and every switch can be turned off. Paint out anything that should not be there, draw a crop, or put a page back exactly as it arrived.
Download the Clean File
Pages you changed are rebuilt and made searchable. Pages you did not are copied across byte for byte, which is the only way to be certain nothing was quietly degraded.
Questions About Cleaning Scans
What the tool does to a document, and what it refuses to do
Is my document uploaded anywhere?
No. The PDF is read from your disk into the browser tab and worked on there. The decoders, the recognition engine and the language model are all served from this site rather than a third-party CDN, so no request carrying any part of your document is ever made — you can watch the network panel and confirm it.
What does it clean without being asked?
Black bars, edge lines and the sloping shadow a document feeder leaves; scuffs, smears and specks stranded out in the margins; blank pages; and pages that are noticeably crooked. All of it found and fixed automatically, with each page saying what was done to it. The one thing it will not do on its own is crop, because that changes the page size.
How does it know a scanner line from a line that belongs there?
By what is next to it. An underline has the word it underlines directly above; a table rule has the table. A feeder shadow has clear paper on every side of it for a centimetre, and so does a scuff, a smear or a fleck of dirt. Nothing is removed that has writing beside it, which is also why a page number sitting alone in an empty margin survives: it lines up with the column of text above it.
My scan is dirty. How does it tell dirt from a full stop?
By what the mark belongs to. Counting across a clean 300 dpi scan, 74 of the 85 small marks on the page touch a letter outright; on a dirty 150 dpi scan of the same kind of document there are 1,310 small marks and 636 of them sit three pixels or more from anything. Typography does not leave punctuation stranded — a full stop is set hard against the word in front of it — so a small mark with no letter near it is dirt. Size cannot be used for this: at 150 dpi a speck and a full stop are the same few pixels. The reach is set generously, because how far a full stop sits from its word differs between documents, and a value tuned on one scan removed the numbering dots on another. If your scan is dirty enough that you would rather have the cleaner page, there is a switch that shortens the reach, and it says what it costs.
Why was one of my pages left alone?
Because nothing on it was worth the cost of rebuilding it. Re-encoding a scan is not free — on a page already stored as CCITT or JBIG2 it can cost several times the file size — so a defect has to be visible before it justifies that. A page that needed only a fifth of a degree of straightening is copied across exactly as it was, and the card says so.
Will cleaning make my scan worse?
Not on a page you leave alone. A page nobody chose to change is copied across untouched, with its original image data, bookmarks and text layer intact, so it cannot be degraded at all. Only pages you actually asked to change are re-encoded, and the tool says which is which before you commit.
Why did the file get bigger instead of smaller?
That is what re-encoding an already-efficient scan does. A page stored as CCITT G4 or JBIG2 is close to optimal, and rewriting it typically costs 70% to 160% more space. It is precisely why unchanged pages are copied rather than rebuilt, and why the tool measures each page before deciding.
Does the text stay searchable?
Yes. A page copied unchanged keeps whatever text layer it arrived with. A page that is straightened or trimmed has its existing text moved to match the new geometry, and a page that never had one is read with Tesseract and given an invisible text layer, so you can search and copy from the finished document.
My scan is faded and lit unevenly. Will the text survive?
Yes. Flattening judges every pixel against its own neighbourhood rather than against one level for the whole page, so a pale corner is measured against pale paper instead of against the darkest part of the sheet. That is what keeps serifs and hairlines on a faded photocopy, where a single threshold erases them. Out in the margins, where there is no stroke to rescue, a stricter rule applies so a grey smudge is not promoted to solid black.
How does it decide to flatten a page to black and white?
By measuring the tones on it. Text on paper is bimodal — nearly every pixel is ink or paper with very little in between — and flattening such a page typically removes 85% to 92% of its size with no visible loss. A photograph or a halftone lives in the midtones, scores high on that measure, and is left with its greys.
What does painting out an area actually do?
It replaces those pixels with paper white in the image that gets written to the new PDF. The original content is not hidden behind a black box that can be lifted off later; it is not in the output file at all. The page has to be rebuilt for that, which is why the tool marks it as changed.
Can it fix a page that claims to be four feet tall?
Yes. Some scanning software writes the pixel count into the page box, so a 300 dpi letter page declares itself 2550 by 3510 points. Printers take that literally. Pages whose declared size implies an impossible resolution are flagged and rebuilt at the size the paper actually was.
How big a document can it handle?
It is limited by your machine rather than by a server quota, and a few hundred pages is comfortable on an ordinary laptop. Recognition is the slow part — roughly a second or two per page — so a long document with every page being read takes a while, and the progress bar tells you where it is.