Digital PDF vs scanned PDF
A digital PDF contains selectable text. A scanned PDF is usually a page image and needs OCR before text can be copied.
Drop a PDF here
or click to choose a file
If the output is blank, the PDF may be scanned. Turn on browser OCR and extract again.
This tool uses pdf.js to read selectable text from the PDF structure inside your browser. For scanned or image-only pages, turn on browser OCR. OCR renders the page image locally and recognizes characters, so it is slower than normal text extraction.
A searchable invoice, research paper, resume, form, or report usually contains selectable text that can be extracted into plain text. A scanned document may need OCR first.
Invoices and forms may need cleanup around tables. Research papers can extract columns out of reading order. Reports and resumes usually work best when the PDF text is selectable.
Choose or drop a PDF. Extraction starts automatically for valid PDFs. Review the text, then use Download TXT.
The PDF may be scanned, image-only, encrypted, or encoded in a way the browser cannot read. Scanned documents need OCR.
Only after it is unlocked. Open the PDF in a reader with the password, save an unlocked copy, and try that file.
Yes. Enter a range such as 1, 3-5, 10 in the page range box before extracting.
The output is plain text. Simple tables may be readable, but complex tables and multi-column layouts often need manual cleanup.
No. The PDF is processed locally in your browser with pdf.js. OCR mode also runs in the browser after the OCR library loads.
PDF text extraction reads text already stored inside the file. OCR analyzes page images to recognize characters in scanned documents.
A digital PDF contains selectable text. A scanned PDF is usually a page image and needs OCR before text can be copied.
PDFs store positioned text objects, not paragraphs. Columns, sidebars, tables, and captions can extract in drawing order.
Use page ranges for long files, keep page numbers on while reviewing, and try Merge lines into paragraphs for prose documents.
Turn on browser OCR and extract again. If OCR still fails, re-export the PDF or open it in a PDF reader to check whether the page is an image.
Unlock protected PDFs before using the tool. For corrupted files, re-export from the source app or save a fresh copy from a PDF reader.