How to get the text out of a PDF
- Choose your PDF, or drag it onto the box.
- Pick the pages you want, such as 1-3 or 5, or leave it empty for every page.
- Press "Get the text". Copy it straight from the box, or download it as a .txt file.
How it works
A PDF stores each piece of text with its position on the page. Your browser reads those pieces with pdf.js and puts them back in order: a new line where the text moves down the page, a blank line where there's a bigger gap (usually a new paragraph), and a space where there's a gap between words on the same line.
For example, two pieces of 12-point text on the same line, "Invoice" ending at 100 points across and "total" starting at 103, get a space between them, because the 3-point gap is more than 15% of the text size (1.8 points). A piece 14 points lower starts a new line, because that's more than half a line down.
What it can't do
- Scanned pages: a scan or photo of a document is a picture, so there's no text to copy. You'd need OCR (text recognition) software for that.
- Layout: columns, tables and text boxes may come out in a slightly different order, and fonts, colours and pictures aren't kept.
- Protected PDFs: if a PDF needs a password to open, remove the password in the app that made it first.
Frequently asked questions
Is my PDF uploaded?
No. The text is read by your own browser. The PDF never leaves your device, and nothing is stored after you close the page.
Why is some text missing or in the wrong order?
Some PDFs store text in an unusual order, or draw letters as shapes rather than text. The tool can only copy text the PDF really contains. If a page says it had no text, it's probably scanned.
How do I count the words?
Paste the text into our word counter to see words, characters and reading time.