Scanned PDF to text
In a scan, each page is a picture of the paper, so there’s no text to copy. Text recognition (OCR) reads the letters on your device and gives you the text in a .txt file, plus a PDF you can now search.
Processed on your deviceNothing is uploaded to servers
Drag one or more PDFs here
PDF files
Files are opened on your device and never uploaded.
How it works
Choose the scan
Drag the scanned PDF (or a photo of a document saved as a PDF) into the workspace, or use the button.
Check the language
The page’s language is already selected; tick any others the document uses. The first time, MovaDoc downloads the recognition data: about 3 MB per language.
Download the text
You’ll see the start of the recognized text on screen. Click “Download the text (.txt)” to get all of it, or download the searchable PDF.
Frequently asked questions
Why can’t I copy the text from my PDF?
Because it’s a scan: each page is a picture of the paper, with no letters stored, and your PDF reader only sees a photo. OCR recognizes the letters in that picture and turns them into text you can copy, edit and search.
Does the text keep the document’s formatting?
No. The .txt file has just the text, line by line, with a blank line between paragraphs: no fonts, bold, images or table grid. For a Word document you can edit with headings, lists and tables, use “PDF to Word”, which also recognizes scanned pages.
Does it work with photos of documents and with handwriting?
With photos saved as a PDF, yes, if they’re sharp and straight: the better the picture, the fewer the mistakes. If the photo is a JPG or PNG, put it in a PDF first with “Images to PDF”. Handwriting isn’t recognized: the recognition data only works for printed text.
Can I get the text from several PDFs at once?
Yes, in two steps. With several PDFs, you get a ZIP with the searchable PDFs (without the .txt files). Then “Continue with…” → “PDF to text”, which also handles several files at once, gives you one .txt per PDF.
Is the document sent anywhere to be read?
No. Recognition runs in your browser, with the Tesseract engine. The first time, MovaDoc downloads the engine and the language data from its own site; after that, it even works offline. Your PDF and the recognized text never leave your device.