Extract text from scans and photos
A scanned document is a picture of text, so copying, searching and selecting do not work on it. OCR reads the characters back out. You get either the plain text, or the same document with an invisible text layer added so it behaves like a normal PDF.
- Scanned PDFs and photos, up to 200 pages
- Seven languages including Russian, Japanese and Chinese
- Searchable PDF keeps the pages looking exactly as they were
- Files are deleted automatically — 24 hours free, 7 days on paid plans
PDF, PNG, JPG, TIFF, BMP, GIF, WebP or HEIC.
Pick the language actually on the page. The wrong one does not fail — it returns nonsense.
How to extract text from a scan
- 01
Upload the scan or photo
Choose a scanned PDF or an image of the page. Multi-page PDFs are read page by page, in order.
- 02
Choose the language
Select the language the document is actually written in. This matters more than it sounds: the wrong language does not produce an error, it produces confident nonsense.
- 03
Choose plain text or searchable PDF
Plain text gives you the characters to copy. Searchable PDF returns the original pages with a hidden text layer, so the file still looks the same but can be searched and selected.
- 04
Read or download the result
Text appears on the page ready to copy; a searchable PDF downloads straight away.
What affects accuracy
Resolution. Pages are rendered at 300 DPI before reading, which is the level Tesseract's own documentation recommends. A photo taken far away, or a scan made at 100 DPI, has less for the engine to work with.
Straightness and focus. A page photographed at an angle, or one where the text curves near the spine of a book, reads noticeably worse than a flat scan.
Language. Each language has its own trained model. Running Cyrillic or Japanese text through the English model returns fluent-looking rubbish rather than an error, which is why the picker only offers languages the server can genuinely load.
What OCR does not do. Layout is not preserved: columns, tables and headers come back as running text in reading order. If the original PDF already has a text layer, converting it to text uses that instead — exact, instant, and better than any OCR.
What people use this for
Contracts and invoices
A signed PDF that arrived as a scan, turned into a searchable file so a reference number can be found later.
Receipts photographed on a phone
Pull the amounts and dates out of a photo instead of typing them in again.
Old documents and archives
Paper records scanned years ago, made searchable without changing how the pages look.
Screenshots and slides
Recover the text from an image where the original file is long gone.
Frequently asked questions
What is a searchable PDF?
A PDF that looks exactly like the scan you uploaded, with an invisible layer of recognised text positioned over the words. The pages are unchanged visually, but you can now select, copy and search them, and so can other software.
Which languages are supported?
English, Russian, French, German, Spanish, Japanese and Simplified Chinese. The picker lists exactly what the server can load — it never offers a language that would silently produce nonsense.
How many pages can it handle?
Up to 200 pages per file. Each page is rendered and read separately, so a long document takes proportionally longer.
Is OCR accurate enough to trust?
For a clean, straight scan of printed text it is very accurate. Handwriting, low-resolution photos, angled pages and unusual fonts are all considerably worse. Always check figures that matter — OCR failures look like plausible text, not like errors.
Do you keep my document?
No longer than the retention on your plan: 24 hours on Free and for visitors without an account, 7 days on paid plans. The file is transferred over TLS and is not used for anything besides producing your result.
My PDF already has selectable text. Do I need OCR?
No. If a PDF already has a text layer, converting it to text uses that layer directly, which is exact and instant. OCR only runs when there is nothing to extract — that is, when the page really is just an image.