PDF to text
Copy text from a PDF – without broken line breaks
A paragraph from a PDF needs to go into an e-mail, a Word document or a translation tool. After pasting, every line ends with a break, words are torn apart by “hyphen- ation” – or nothing can be selected at all. Here's how to get clean text.
Get started: open the PDF, copy the text as clean running text or save it as TXT – even from scans. Free, no upload.
Get text from PDFWhy copied PDF text is broken
A PDF is made for printing. It doesn't store “a paragraph with this text” but “these letters at exactly this position”. Lines, paragraphs and columns only follow from the layout. When copying, the app guesses where a line ends – and inserts a hard break there.
| Problem | Cause |
|---|---|
| A break after every line | Each line is a separate text piece in the PDF. |
| “implemen- tation” | The layout's hyphenation is copied as a real hyphen. |
| Columns mixed up | Lines from the left and right column are read alternately. |
| Strange characters instead of text | The font is embedded without a character map – programs can't read the text. |
| Nothing can be selected | The PDF is a scan (just an image) or copying is locked by a permission setting. |
Clean text in three steps
- Open the PDF. Drop the file into PDF to Text. The text of all pages appears in the text box.
- Check the settings. “Remove line breaks inside paragraphs” and “Join hyphenated words” are on by default. Bullets and numbered lists are kept. Only some pages? Enter e.g.
3-5, 8under “Pages”. - Copy or save. “Copy text” puts everything on the clipboard, “Save as TXT” downloads a text file.
The PDF never leaves your computer. Contracts, reports or medical letters are safe to convert – the text is read in your browser.
Running text instead of line salad – drop the PDF in, copy, done.
When nothing can be selected
- Scan or photo: the text is just an image. The tool detects such pages automatically and offers text recognition (see below).
- Copy protection: some PDFs forbid copying via a permission setting. Respect that for other people's work; for your own documents, the author or the original file helps.
- Password to open: without the password the PDF can't be read at all – not by tools either.
- Gibberish instead of text: the font lacks a character map. Text recognition helps here too, because it reads the letters from the image.
Scanned PDFs: text recognition
If the tool finds pages without text, a hint appears with the button “Recognize text (OCR)”. Recognition (Tesseract) also runs in your browser and reads German and English. The first time it downloads about 20 MB of recognition data; after that it's faster.
- The cleaner the scan, the better: straight, high-contrast, at least 200 dpi.
- Handwriting is hardly recognised; tables come out as running text.
- Always proofread – typical errors are
rnformor0forO.
Alternatives: Word, Acrobat, pdftotext
Microsoft Word
Word opens PDFs directly and converts them into an editable document. Fine for simple layouts; complex pages end up with many text boxes and breaks.
Adobe Acrobat
“Save as text” and OCR are part of the paid versions. The free Reader only lets you select and copy.
Google Docs
Upload the PDF to Google Drive and open it with Google Docs – Drive runs OCR in the process. The file then lives at Google, though.
Command line
pdftotext from Poppler extracts text from many PDFs in seconds, ocrmypdf makes scans searchable.
pdftotext -layout document.pdf document.txt
ocrmypdf -l eng scan.pdf scan-searchable.pdf
Get the text out of your PDF
Clean running text without line salad, copy or save as TXT – even from scans via OCR. Free, no upload.
Frequently asked questions
Why does copied PDF text have so many line breaks?
PDFs store text line by line at fixed positions. When copying, every line end becomes a hard break.
How do I remove the line breaks?
With PDF to Text: the option “Remove line breaks inside paragraphs” joins lines back into paragraphs while keeping lists.
Can I copy text from a scanned PDF?
Yes, with OCR. The tool detects pages without text and reads them with OCR on request – right in your browser.
Is my PDF uploaded?
No. OCR runs locally too. Only the recognition data is downloaded once.
Is formatting kept?
No, the result is plain text without bold, font sizes or table lines.
What about multi-column layouts?
They're usually read in the right order, but complex layouts can mix up columns.
Can I convert only some pages?
Yes. Enter e.g. 1-3, 5 under “Pages”.
Sources
- Tesseract OCR and Tesseract.js – the text recognition the tool uses in your browser.
- Poppler (pdftotext) and OCRmyPDF.
- PDF.js – the open-source library the tool uses to read the text.