PDF in Text umwandeln
Den Text aus einem PDF herausholen – sauber, ohne Zeilenumbrüche mitten im Satz. Zum Kopieren oder als TXT-Datei. Scans werden per Texterkennung gelesen.
Warum der kopierte Text oft kaputt ist
Ein PDF speichert Text Zeile für Zeile an festen Positionen. Kopiert man ihn direkt, steht hinter jeder Zeile ein Umbruch und getrennte Wörter bleiben getrennt.
Dieses Tool erkennt Absätze, verbindet die Zeilen wieder und fügt „Silben-
trennung“ zusammen. Listen bleiben Listen.
Gescannte PDFs
Ein Scan ist nur ein Bild – es gibt keinen Text zum Herauskopieren. Für solche Seiten bietet das Tool eine Texterkennung (OCR) für Deutsch und Englisch an.
Sie läuft ebenfalls in deinem Browser. Beim ersten Mal werden dafür einmalig etwa 20 MB Erkennungsdaten geladen.
Grenzen
- Mehrspaltige Layouts und Tabellen kommen als Fließtext heraus.
- Formatierung (fett, Schriftgröße) geht verloren – es ist reiner Text.
- OCR ist bei schlechten Scans und Handschrift ungenau.
Why copied text is often broken
A PDF stores text line by line at fixed positions. Copy it directly and every line ends with a break, and hyphenated words stay split.
This tool detects paragraphs, joins the lines again and repairs hyphen-
ation. Lists stay lists.
Scanned PDFs
A scan is just an image – there's no text to copy. For such pages the tool offers optical character recognition (OCR) for German and English.
It also runs in your browser. The first time, about 20 MB of recognition data is downloaded once.
Limits
- Multi-column layouts and tables come out as running text.
- Formatting (bold, font size) is lost – it's plain text.
- OCR is inaccurate for poor scans and handwriting.