A PDF does not contain paragraphs, so they have to be rebuilt
This is the fact that explains every disappointing PDF-to-Word conversion you have ever had. A PDF is a description of marks on a page: put this glyph at this coordinate, in this font, at this size. It does not record that two lines belong to the same sentence, that a line is a heading, or that three lines are a bulleted list. The structure you see when you read one is something your eyes assemble from the geometry.
So a converter has to assemble it too, and the quality of a conversion is almost entirely the quality of that guessing. Four signals do the work here. Type size: a line set well above the body size is a heading. Vertical gap: a jump of more than about one and three-quarter line heights ends a block. Line fill: a line that stops short of the right margin — with room to spare for the first word of the next line — is the last line of its paragraph. And the left edge: a bullet glyph at a consistent indent is a list item, whose wrapped lines line up under the text rather than under the bullet.
The third signal is the one that matters most, because it is the only one that works on documents with no extra space between paragraphs. Get it wrong and you produce a .docx with one paragraph per visual line — technically a valid Word file, and useless to edit, because every sentence you touch re-wraps into a mess. That single behaviour is the difference between a converted document you can work in and one you rewrite from scratch.
Two conversions, because you cannot have both at once
A PDF is fixed: every glyph has a coordinate and the page never changes shape. A Word document is the opposite — it reflows, which is the entire reason to want one. Those two properties genuinely conflict, and a converter that claims to deliver both is choosing not to mention which one it quietly dropped. So this offers the choice explicitly.
Keep the layout pins every line to the position it occupied in the PDF, using Word's own paragraph frames. The page size and the margins are measured from the file rather than assumed, tables are placed where they were, and a line with a wide gap in the middle of it — a signature block, a form field, a right-aligned page number — is split so that both halves keep their positions instead of sliding together. What you give up is reflow: each line is its own frame, so typing into one does not push the next one down. It is the right choice for reading, sharing, printing, or lifting a figure out of a report.
Editable text does the older and harder thing: it works out which lines belong to the same paragraph and joins them back together, marks larger lines as real Word Heading 1, 2 and 3 styles so the navigation pane and an automatic table of contents both work, rebuilds bullet lists as genuine Word list items, and keeps the numbers a numbered list actually showed so a list that began at 4 still begins at 4. The layout is approximated rather than reproduced, and in exchange the document behaves like one you wrote.
Both modes carry bold and italic wherever the embedded font name says so, both set the type at the size the PDF used rather than flattening everything to one size, and both find tables, detect multi-column pages and extract embedded images.
Tables, found by alignment rather than by borders
A PDF stores no tables. It stores text at coordinates and, sometimes, lines drawn near that text. What a reader calls a table is something the eye assembles from alignment — a run of lines whose text starts at the same handful of horizontal positions.
So that is what is looked for here, rather than the borders. The obvious alternative is to hunt for the rectangles that make up a table's ruling, which is more precise when it works and fails completely on the very common table that has no borders at all: a price list, a specification sheet, an invoice with nothing but whitespace between its columns. Alignment is present in every table, because alignment is what makes it readable as one. Ruling is decoration that many tables do without.
The thresholds are deliberately cautious, because a false positive is far worse than a miss. Two aligned lines are not enough — a heading over a subheading, a two-line signature block, and a caption under a figure all align, and none of them is a table — so three rows are required. Two columns are required as well. And a block whose cells are long and fill their columns is rejected as prose, because that is what a two-column article looks like from the geometry, and rebuilding an article as a spreadsheet would be a worse result than leaving it as paragraphs.
What is not attempted: merged cells, which look identical from the geometry to a cell that simply has more text in it, so a wrong span is never guessed; nested tables; and a table split across a page break, which arrives as two tables because the file records nothing to say the two halves were one.
Images and columns, and what is left after them
Embedded images are extracted and placed — at the size and position they had in the PDF in layout mode, and in the right place in the flow in editable mode. Getting them out is not a call to a library: a PDF holds pixels in a resource dictionary and, somewhere in the content stream, an instruction to paint them through whatever the transformation matrix happens to be at that moment. Position and size are not properties of the image at all, so the content stream has to be walked the way a renderer walks it, tracking the graphics state, to find out where each picture landed.
Multi-column pages are detected, which matters more than it sounds. Sorted by vertical position — the only order the file itself suggests — a two-column page gives you the first line of the left column, the first line of the right column, the second line of the left column, and so on, and the result is unreadable. A column boundary leaves an unmistakable trace: a vertical strip that no line of text crosses, running most of the height of the block. Finding that strip and reading each side in turn is what turns an interleaved mess back into two columns of prose.
What is left: a chart drawn as line art is a few hundred drawing operators rather than a picture, so it is not extracted — rasterising a region of the page is a different feature with different trade-offs. Merged table cells are not reconstructed, because from the geometry a cell spanning two columns is indistinguishable from a cell with more text in it, and a wrong span is worse than two right cells.
Headers, footers, footnotes and page numbers arrive as ordinary text at the top or bottom of a page, because that is exactly what they are in the file. Text colour is not carried across. Exact fonts are not — the document is set in Calibri at the PDF's own sizes, because embedding the original typeface would mean shipping a copy of a font you may not have the right to redistribute.
That list is longer and blunter than the one you will find on converters promising to keep your formatting "100% intact". They cannot, and neither can this. What you are getting is either a document that looks like the original or one that edits like a document you wrote — and being told, in advance, which parts of the original did not make the trip.
A .docx is a ZIP of XML files, which is why this works in a tab
Rename a Word document to .zip and open it and you will find a folder structure: word/document.xml holds the content, word/styles.xml the styles, [Content_Types].xml a manifest, and a couple of small relationship files tie them together. That is the whole format for a text document. Writing one means writing that markup and zipping it, which is why no server and no conversion service is needed here.
It also means one detail decides whether the file opens at all: XML escaping. An ampersand, a less-than or a greater-than sign written straight into the markup makes it malformed, and Word's response to malformed XML is to declare the file corrupt without saying which character did it. Every character that leaves this converter is escaped, and the automated test for this page deliberately runs a document containing all three through the whole round trip.
The output is a plain, standards-compliant document with no macros, no tracking, no embedded objects and nothing phoning home. It opens in Microsoft Word, LibreOffice Writer, Google Docs, Apple Pages and anything else that reads OOXML. Because your file never leaves the device, a confidential contract stays confidential — which matters more here than for most operations, since the documents people most often need converted are the ones they least want to upload.
Scanned PDFs, and the conversion that is not possible
A scanned document is a photograph of a page. There are no characters in it, only pixels arranged to look like characters, so there is nothing for any converter to extract. A tool that hands you an empty Word file in that situation has not failed gracefully — it has left you assuming something is broken with your file, your browser or your download.
So the document is sampled before any work begins. If none of the sampled pages contain text, you are told plainly that the PDF is almost certainly a scan, that optical character recognition is what turns a scan back into text, and nothing is downloaded. If only some pages have text — very common when a scanned exhibit has been appended to a typed contract — the document is produced and you are told what proportion is missing.
The other case worth naming is a PDF exported from a design tool, where every line of text is its own positioned object and the reading order in the file bears no relation to the reading order on the page. Those convert badly everywhere, including here, and no amount of geometry recovers an order that was never recorded. If the result looks scrambled, that is what happened, and the fix is to go back to the source document rather than to try another converter.