A PDF does not contain paragraphs
This is the thing that explains every frustration people have with PDF text extraction. A PDF stores instructions to place glyphs at coordinates โ put an "h" here, an "e" three points to the right โ and nothing about which of them form a word, a line, a sentence or a paragraph. The structure you see when you read it is something your eyes reconstruct from the geometry.
So extraction has to do the same reconstruction. Runs of glyphs sharing a baseline are gathered into a line. A gap wider than about a quarter of the type size is treated as a space, because the file often does not store spaces at all. A vertical jump bigger than a line and a half ends a paragraph. Lines set noticeably larger than the body are probably headings.
These are heuristics, and it is worth saying so plainly rather than pretending otherwise. They do well on the documents most people have: reports, articles, letters, contracts, invoices. They do badly on multi-column layouts and complex tables, because the reading order in those was never recorded in the file, and no amount of geometry recovers information that was never written down.
Which format to choose
Plain text gives you the words with paragraph breaks and nothing else. It is what you want if the text is going somewhere that will do its own formatting โ a search index, a spreadsheet cell, an email, another document.
Markdown keeps the structure the extraction could work out: lines set much larger than the body become headings, paragraphs stay paragraphs, and page breaks become horizontal rules. It is the right choice if a human is going to edit the result afterwards.
HTML is the same reconstruction wrapped in a small, readable page you can open in a browser immediately. Useful for reading a document on a phone without a PDF reader, or for pasting into a content management system.
When there is no text to extract
A scanned document is a picture of a page, not a page. There are no glyphs in it, only pixels arranged to look like glyphs, so there is nothing for extraction to find โ and a tool that returns an empty file without explanation leaves you assuming it broke.
So the file is sampled before any work starts. If none of the sampled pages contain text, you are told the document is almost certainly a scan and that optical character recognition is what turns it back into text. If only some pages have text โ a common case where a scanned exhibit has been appended to a typed contract โ the file is produced and you are told what proportion is missing.
That is a small piece of work that costs a second and saves the ten minutes you would otherwise spend wondering what went wrong.
What extraction loses
Layout, mostly. Columns are flattened into a single stream, which reads correctly for a single-column document and scrambles a newspaper. Tables lose their grid: the cells come out in reading order, which is often useful and never structured. Images are not extracted; footnotes usually land at the end of the page they belong to rather than attached to their reference.
What survives is the thing you usually wanted: the words, in order, in something you can search, edit, quote and paste. If the file is confidential, it is worth noting that it never left this device to get here โ which is not true of most extraction tools, and matters a great deal for exactly the documents people most often need extracted.