This extracts text, it does not run OCR#
A PDF built from a real document (exported from a word processor, generated from HTML, or produced by most modern software) has an actual text layer embedded in it — the words exist as data, not just pixels, and this tool reads that data directly. A PDF made by scanning a paper page, though, often contains nothing but an image of the page; there is no text data to extract, because none exists in the file. This tool does not attempt optical character recognition (reading the shapes of letters out of an image) — it is a read-the-existing-text tool, not a read-the-picture tool, and it says so clearly when a page has nothing extractable rather than returning blank output with no explanation.
Why extracted text sometimes reads oddly, even on a real text layer#
A PDF's text layer stores each piece of text with its own position on the page, not necessarily in the reading order a human would follow — a multi-column layout, a table, or text boxes placed for visual effect can all extract in an order that looks scrambled compared to how the page reads visually. This is a structural property of how PDF stores text, not a bug in the extraction; a PDF genuinely does not have a single "reading order" the way a plain text file does unless the document was specifically tagged with one.
Where this differs from PDF to JPG#
This tool gives you the words as editable, searchable plain text — useful for copying content, feeding it into another tool, or searching across a document. This site's PDF to JPG Converter goes the other direction: it produces a picture of each page, useful when the visual layout itself is what matters, not the words as text.