Converting files
How do I get a table out of a PDF and into Excel?
Convert the PDF to XLSX rather than copying and pasting. A converter reads the position of every number and rebuilds the columns, which pasting cannot do.
Copying a table out of a PDF and pasting it into Excel almost never works. You select what looks like a neat grid, paste, and everything lands in column A as one long strip of text.
That is not Excel being unhelpful. It is a consequence of what a PDF actually is, and once you know the reason, the right approach is obvious.
Why pasting collapses the columns
A PDF does not contain a table. It contains a list of instructions that say "draw the characters 1,240 at this exact point on the page", repeated a few hundred times. The visual alignment you see is a consequence of the coordinates, not a description of structure.
There is no row. There is no column. There is no cell. Those things exist in your eye, not in the file.
When you copy, the viewer hands over the text in the order it appears in the drawing instructions, with the geometry thrown away. Excel receives a stream of words and does the only sensible thing with it, which is to put the stream in one cell.
A converter does the work the viewer skips: it reads the coordinates, clusters values that share a horizontal band into a row, clusters values that share a vertical band into a column, and writes a real spreadsheet.
What to expect from your file
Not every PDF converts equally well, and the difference is almost entirely down to how it was made.
Made by exporting from a spreadsheet or a report tool. These convert well. The text is real text, the alignment is consistent, and the column boundaries are clean.
Made by scanning paper. There is no text in the file at all, only a picture of text. Nothing can rebuild the columns from that until the characters have been recognised. See the note below.
Made by a designer. Tables in brochures and annual reports often have merged cells, values spanning two columns, and figures placed by hand. These convert, but expect to fix a few cells.
Doing the conversion
PDF to Excel takes the file and gives you an XLSX. Open it, and check three things before you rely on it.
The header row. Converters have to guess whether the first row is a header or just the first record. Look at it: it takes a second and it affects every formula you write afterwards.
Number formatting. A figure that arrives as text sorts alphabetically, which puts 1,000 before 9. Select the column and check that Excel is right-aligning the values. If it is not, they are text, and Data → Text to Columns with no delimiter converts them in one step.
Merged cells. If the original had a heading spanning several columns, you may get a merged cell that breaks sorting. Unmerge it before doing anything else.
When the table runs over several pages
Long tables repeat their header on every page. A converter that does not notice will treat the repeated headers as data rows, and you end up with the word "Amount" scattered through your amounts column.
Two ways to handle it: convert, then sort by the first column and delete the repeated header rows in one block; or extract just the table pages with Split PDF and convert those. The second is more work but the result needs no cleaning.
If the PDF is a scan
A scanned page has no text to read, only pixels shaped like text. It has to be recognised first.
PDF to Text runs optical character recognition and gives you the words. The layout will not survive, so you get the values but not the grid, and you rebuild the columns yourself.
This is honest rather than pessimistic: recognising characters and reconstructing a table structure from a photograph are two different problems, and any tool claiming to do the second perfectly from a poor scan is overpromising. If the scan is good (straight, high contrast, 300 DPI) the recognition is usually excellent and the retyping is minimal. If it is a phone photo taken at an angle, Enhance it first, and it will genuinely help.
Going the other way
If your problem is the reverse (you have the data and want a PDF someone can read) do not print from Excel and hope. Excel to PDF scales the sheet properly, and if your data is already a CSV, CSV to PDF measures every column against its contents so nothing gets clipped.
The short version
Never paste. Convert. Check the header row, check the number formatting, check merged cells. If the page is a scan, recognise the text first and accept that you are rebuilding the grid by hand.
Cleaning up a converted sheet, in order
The sequence matters, because two of these steps get harder if you do them in the wrong order.
- Unmerge everything first. Select all, then Merge & Center to toggle it off. Merged cells break sorting and filtering, and you cannot see which cells are merged until you try to sort.
- Delete the repeated header rows. Sort by the first column and they gather in one block.
- Fix the number columns. Select each one and check the alignment. Right-aligned means Excel sees numbers; left-aligned means text.
- Then add your filters, formulas and formatting.
Doing step 4 before step 1 is the common mistake, and it means redoing it.
Two conversions that look the same and are not
Worth separating, because people ask for one when they want the other.
Converting the whole document gives you every page as a sheet, including the covering letter and the notes. Right when the PDF is essentially a spreadsheet that was printed.
Extracting one table from a report that happens to contain it is a different job. Split PDF pulls out just the pages the table lives on, and converting three pages rather than forty gives you a sheet with nothing to delete.
The second is faster and cleaner almost every time, and it takes about thirty seconds longer to set up.
Common questions
Why does pasting a PDF table into Excel put everything in one column?
Because a PDF has no table structure, only characters at coordinates. Copying keeps the text and discards the geometry, so Excel receives one stream of words. A converter reads the coordinates and rebuilds the rows and columns.
Will my scanned PDF convert to Excel?
Not directly. A scan contains a picture of text, not text. The characters have to be recognised first, and the column layout will not survive that step, so you get the values and rebuild the grid yourself.
How do I check whether my PDF has real text?
Try to select a number in your viewer. If a text cursor appears, the file has real text. If you get a selection rectangle over an image, it is a scan.
Why do numbers arrive as text after converting?
Because the converter cannot always tell a figure from a label. Select the column and use Data, then Text to Columns with no delimiter. That converts the whole column in one step.