Document
Turning a PDF into usable text: what extraction recovers, and what it loses
A PDF is designed so that a document looks the same everywhere, not so that its content can be reused. Copying and pasting from a PDF often gives broken lines, merged words or jumbled columns. Exporting the text to TXT or Markdown solves some of these problems, provided you understand what a PDF actually contains. This guide explains how extraction works in FileXvert, and how to get the most out of it.
What a PDF really contains
Unlike a Word document, a PDF contains no paragraphs or headings in the sense a word processor understands them. It contains drawing instructions: place these characters, in this font, at this position on the page. A paragraph is often just a series of text fragments positioned one below another, with no explicit link between them.
Recovering flowing text from a PDF therefore means rebuilding a reading order from those positions. For a simple, single-column document, the result is usually faithful. For a complex layout, it is a matter of interpretation, and every automatic extraction makes compromises.
Digital text or scanned image: the decisive test
Before converting anything, try this: open the PDF and try to select a word with the mouse. If you can, the document has a text layer and extraction will work. If the selection grabs the whole page like an image, or selects nothing, the PDF is a scan: a photo of a page, with no characters to extract.
FileXvert does not perform optical character recognition (OCR). A scanned PDF will therefore produce an empty or almost empty text file. This is not a conversion fault: there is literally no text in the file, only pixels. For this kind of document, OCR software is needed, and its output is what you can then convert.
How FileXvert rebuilds the text
Extraction relies on pdf.js, the PDF rendering engine developed by Mozilla and built into Firefox. It reads, page by page, the text fragments and their end-of-line markers. FileXvert joins them while respecting those line ends, removes superfluous spaces and limits consecutive blank lines, so that headings and paragraphs stay separate instead of merging into a single block.
Like everything else on the site, this runs entirely in your browser: the PDF is sent nowhere. That is particularly useful for documents you want to work with precisely without spreading them around: contracts, internal reports, statements.
What is kept, what is lost
Extraction recovers content, not presentation. In practice:
- Kept: the text, in page order, with line breaks and separations between blocks.
- Lost: formatting (bold, italics, sizes, colours), images, graphic headers and clickable links.
- Imperfect: tables, whose cells come out one after another, and multi-column layouts, whose lines may interleave.
- Worth checking: words hyphenated at the end of a line, which stay split, and repeated headers or footers, which reappear on every page.
TXT or Markdown?
The extracted content is the same in both cases: FileXvert does not automatically guess the headings or lists of a PDF, since those concepts do not exist in it. The difference lies in what you will do with the file.
Choose TXT if you simply want to read, search or archive the content: it opens everywhere, even in the most basic text editor. Choose Markdown (.md) if you plan to rework the text in a note-taking tool such as Obsidian, Notion or a documentation editor: just add a # in front of headings and a - in front of list items to get a structured document.
The recommended method
To get text that is genuinely reusable, work in four steps:
- Check that the PDF contains selectable text.
- Convert it to TXT or Markdown depending on the intended use.
- Do a clean-up pass: find and replace repeated headers, rejoin split words, reformat important tables.
- Compare key passages with the original PDF before quoting, publishing or relying on the text in a document that matters.
For a Word document, the same approach applies: FileXvert can also extract the text of a DOCX file to TXT or Markdown, or turn it into a simple PDF.