PDF to Text — Extract Text from a PDF
Browser tool that extracts the text of a PDF page by page for copying or saving as TXT
Runs in your browser · nothing is uploaded
Drop the PDF to extract text from here
Add one PDF (up to 100 MB). Adding another file replaces the current one. Files never leave your browser.
What it does
It turns the text inside a PDF into plain text. Use it to move a report into a text editor, quote part of a paper, or search a long contract for a phrase. You can open each page on its own and copy it, copy everything at once, or download a TXT file. When letters come out broken because of the fonts, or when a scan has no text at all, the page explains why and what to do. Everything runs in your browser.
How to use
- Drop one PDF on the dashed area or press Choose PDF.
- If you only need part of it, type a page range such as
1-3, 5. Empty means all pages. - Choose whether to put a separator such as
--- Page 1 ---between pages. - Press Extract text to get the full text and the text of each page. On long files, Stop ends early and keeps what was read.
- Copy all, copy one page, or download the TXT file.
How it works
- Extraction: the text pieces PDF.js reads are joined in stored order. A new line starts at an end-of-line mark or where the baseline moves by more than half a line; a space is added between pieces on the same line that are visibly apart. Trailing spaces are removed and blank lines squeezed to one.
- Garbled-text warning: shown when at least 3 characters, and at least 2% of the visible ones, are replacement characters (�), private-use characters or control characters.
- Pages without text: their numbers are listed; if every page is empty the file is treated as a scan and OCR is suggested.
- TXT file: UTF-8 with a byte-order mark and Windows line endings (CRLF), so Korean, Chinese and emoji open correctly even in older Notepad.
- Limits: files up to 100 MB. No OCR.
- Libraries: PDF.js (Apache-2.0). Character maps for Chinese, Japanese and Korean fonts are served from this site and loaded only when needed.
Examples
| File | Result |
|---|---|
| Report exported from Word | Paragraphs and lines come out as written |
| Two-column paper | Usually left column then right; table cells may run together |
| Scanned receipt | “No text was found” with an OCR hint |
| PDF whose font lacks a Unicode map | Garbled-text warning with the fix |
For the photos inside a PDF use Extract PDF Images; to keep the look of whole pages use PDF to JPG.
FAQ
Is my PDF uploaded?
No. Extraction runs in your browser only and the file is never uploaded. The text exists only on this page and in the TXT file you download.
The extracted text is garbled. Why?
When a PDF embeds a font without a Unicode map, the letters show on screen but come out as blanks or odd symbols when read as text. If enough of those characters appear, the tool warns you. Exporting the PDF again from the original program, or using text recognition (OCR), is the fix.
Why do I get nothing from a scanned PDF?
In a scanned or photographed PDF every page is a picture, so there is no text to extract, and this tool does not do OCR. When no text is found at all you are told so; run the file through a program or scanner app with text recognition first.
Are line breaks and order the same as in the original?
Text follows the order it is stored in the PDF, and lines are split at end-of-line marks and where the baseline jumps. Most documents come out in reading order, but multi-column layouts, tables and text boxes can be mixed up or run together.
Does it work with protected PDFs?
PDFs that ask for a password to open are not supported yet. PDFs that open without a password but restrict printing or editing do open, and their text can be read. Check the document's terms of use yourself.
Related tools
Extract PDF Images
Browser tool that saves the images embedded in a PDF as PNG or JPG, singly or as a ZIP
PDF to JPG
Render PDF pages to JPG or PNG at a chosen DPI and download a ZIP, in the browser
Split PDF
Split a PDF by page ranges, every N pages or per page, with ZIP download, in the browser
Character Counter
A character and word counter with bytes and platform length presets