Converting files

Why is converting a PDF back to Word or Excel never perfect?

A PDF stores positioned glyphs rather than paragraphs, tables or slides, so converting back is reconstruction from appearance. Text usually survives, structure is inferred, and anything the original expressed through layout alone is guesswork.

3 min read

It usually goes the same way: the text comes through, but the headings, tables and layout that made it a *document* don't. Once you know why, it is much easier to get a result you can use.

A PDF does not know what anything is

When a word processor exports a PDF, it turns meaning into appearance. "Heading 2, Calibri, 14pt, keep with next" becomes "draw these letters at these positions in this font at this size". The look is kept, but the file no longer knows it was a heading.

A PDF has no idea of:

  • a paragraph: only lines of letters that sit near each other
  • a table: only text lined up in columns, and sometimes lines drawn around it
  • a cell: the grid was never saved, only what it looked like
  • a slide layout: only shapes and text at positions
  • a heading: only text that is larger

Converting back means looking at the appearance and guessing the meaning. That is hard, and when it goes wrong, it goes wrong in the same few ways. No converter can bring back information the file never had, so a perfect round trip isn't possible.

What survives well, and what does not

Usually fine: the words themselves, in reading order; font sizes and weights; simple single-column text; images.

Usually close: where paragraphs end (a line ending in a full stop is probably the end of a paragraph, but not always); tables with visible lines, which give the converter something to work from.

Usually poor: tables with no lines, where the grid was only ever whitespace; multi-column layouts, where reading order has to be guessed; anything positioned by hand; footnotes; anything in a text box overlapping other content.

How to get a better result

Choose the closest target. A PDF of a spreadsheet should go to PDF to Excel, not to Word and then into Excel. A PDF of slides should go to PDF to PowerPoint. Each converter is looking for a different structure.

Check whether there is text at all. A scanned PDF has no text to convert, only pictures of it. Run OCR PDF first, or the output will be an empty document with a photo in it.

Expect to fix the tables. If the tables have no lines, plan to spend a few minutes putting columns right. The converter hasn't failed. The column edges are not in the file.

Ask whether you need to convert at all. If you only want to change a date or a name, editing the PDF directly with the PDF editor is quicker and keeps everything else exactly as it was. Sending a finished document through Word and back to change one line usually costs more than it saves.

Common questions

Why did my table come out as loose text?

Because it had no lines. Without them the grid is only empty space, and the converter can't tell where one cell ends and the next begins.

Which conversion is most reliable?

Simple single-column text to Word. The more the original used layout to show its structure, the more has to be guessed.

My converted file is empty. Why?

The PDF is almost certainly a scan, so there is no text in it to convert. OCR it first, then convert the result.