Extract the text from a PDF

Get the words out as plain text, Markdown or HTML โ€” with paragraphs and headings reconstructed, and an honest warning when there is no text to find.

Your file never leaves this device. There is no upload step โ€” the work happens in this browser tab.

Open the fileDrop the PDF on the box. You can remove pages you do not need first, so the output is only the part you care about.
Pick a shape for the outputPlain text for pasting somewhere else. Markdown if you want headings preserved as headings. HTML for something you can open and read straight away.
DownloadThe text is assembled in this tab. If the document turns out to be a scan, you are told before anything is downloaded rather than handed an empty file.

A PDF does not contain paragraphs

This is the thing that explains every frustration people have with PDF text extraction. A PDF stores instructions to place glyphs at coordinates โ€” put an "h" here, an "e" three points to the right โ€” and nothing about which of them form a word, a line, a sentence or a paragraph. The structure you see when you read it is something your eyes reconstruct from the geometry.

So extraction has to do the same reconstruction. Runs of glyphs sharing a baseline are gathered into a line. A gap wider than about a quarter of the type size is treated as a space, because the file often does not store spaces at all. A vertical jump bigger than a line and a half ends a paragraph. Lines set noticeably larger than the body are probably headings.

These are heuristics, and it is worth saying so plainly rather than pretending otherwise. They do well on the documents most people have: reports, articles, letters, contracts, invoices. They do badly on multi-column layouts and complex tables, because the reading order in those was never recorded in the file, and no amount of geometry recovers information that was never written down.

Which format to choose

Plain text gives you the words with paragraph breaks and nothing else. It is what you want if the text is going somewhere that will do its own formatting โ€” a search index, a spreadsheet cell, an email, another document.

Markdown keeps the structure the extraction could work out: lines set much larger than the body become headings, paragraphs stay paragraphs, and page breaks become horizontal rules. It is the right choice if a human is going to edit the result afterwards.

HTML is the same reconstruction wrapped in a small, readable page you can open in a browser immediately. Useful for reading a document on a phone without a PDF reader, or for pasting into a content management system.

When there is no text to extract

A scanned document is a picture of a page, not a page. There are no glyphs in it, only pixels arranged to look like glyphs, so there is nothing for extraction to find โ€” and a tool that returns an empty file without explanation leaves you assuming it broke.

So the file is sampled before any work starts. If none of the sampled pages contain text, you are told the document is almost certainly a scan and that optical character recognition is what turns it back into text. If only some pages have text โ€” a common case where a scanned exhibit has been appended to a typed contract โ€” the file is produced and you are told what proportion is missing.

That is a small piece of work that costs a second and saves the ten minutes you would otherwise spend wondering what went wrong.

What extraction loses

Layout, mostly. Columns are flattened into a single stream, which reads correctly for a single-column document and scrambles a newspaper. Tables lose their grid: the cells come out in reading order, which is often useful and never structured. Images are not extracted; footnotes usually land at the end of the page they belong to rather than attached to their reference.

What survives is the thing you usually wanted: the words, in order, in something you can search, edit, quote and paste. If the file is confidential, it is worth noting that it never left this device to get here โ€” which is not true of most extraction tools, and matters a great deal for exactly the documents people most often need extracted.

Questions

What people ask about extracting text

Why is my output empty?

Almost certainly because the PDF is a scan โ€” an image of a page rather than a page of characters. There is no text in it to extract. You will be told this before anything downloads, rather than handed an empty file. Optical character recognition is what converts a scan back into text.

Why is the text out of order?

Multi-column layouts and complex tables are the usual cause. A PDF records where each glyph sits, not the order a human should read them in, so a two-column page can interleave. Single-column documents come out in the right order.

Are tables preserved?

Not as tables. The cell contents come out in reading order but the grid is lost, because the file usually stores a table as text positioned inside drawn lines rather than as a structure.

Which format should I choose?

Plain text for pasting elsewhere, Markdown if a person will edit the result and you want headings kept, HTML if you want something you can open and read straight away.

Is my document uploaded to extract the text?

No. The parsing happens in this browser tab using pdf.js. You can watch the network panel while it runs. That matters here more than for most operations, because the documents people extract text from are often the confidential ones.