Converting files

How to turn a CSV into a PDF table people can actually read

Use a converter that measures each column against its contents and detects the delimiter, rather than opening the CSV in a spreadsheet program and accepting its import defaults.

6 min read

A CSV is the least glamorous file format in regular use and one of the most important. Almost every system that holds data can produce one: accounting packages, databases, analytics dashboards, e-commerce platforms, payroll. When you need to get numbers out of one system and into another, or in front of somebody who just wants to read them, a CSV is usually what you get.

Turning that into a document someone can read, however, tends to go badly.

What a CSV is, and what it is missing

A CSV file is plain text. One line per row, values separated by a character. That is the entire format.

What it does not contain is everything that makes a table readable: column widths, fonts, alignment, borders, a marked header row, number formatting. A CSV holds 1250 and not £1,250.00, and it has no way to say that the first row is headings rather than data.

So converting a CSV is not really a conversion. It is a design job: deciding how wide each column should be, what the header should look like, how a long value should behave. Whoever does the converting is making those decisions, whether they think about it or not.

Why the usual route disappoints

The instinctive approach is to open the CSV in Excel or LibreOffice and export a PDF. This works, but you inherit that program's import defaults, and those defaults are the source of the familiar problems.

Every column arrives at the same default width. A column holding Yes and a column holding a two-line postal address are given identical space. One wastes it; the other clips, or shows ##### where a number would not fit.

The separator is assumed. Spreadsheet programs guess based on your locale. Files exported in much of Europe use semicolons, and opening one in a comma-configured Excel gives you a single column with everything crammed into it. One of the most common and most confusing spreadsheet problems there is.

The header is not a header. Row one is just row one. It looks like the data, and when the table runs onto page two there is nothing at the top telling you what the columns are.

Long values get truncated silently. This is the worst of them, because you may not notice. A cell that shows Northwind Trading Compa looks like data rather than a display problem.

What a good conversion does instead

Handled properly, every one of those is solvable.

Measure the columns. Look at what is actually in each column and give it the width it needs. A column of dates gets date-width; a notes column gets the rest.

Detect the delimiter properly. Not by locale, and not by counting characters. Counting picks the comma the moment a description field contains prose. The reliable test is which candidate produces a *consistent* column count across many lines. A semicolon file gives eight columns on every row when split by semicolons, and a chaotic number when split by commas. That is the signal.

Style the header and repeat it. On every page, not just the first. This is what decides whether page four of a table is readable at all.

Wrap rather than clip. If the table is genuinely wider than the page, narrow the widest columns first (the prose ones, which wrap gracefully) and leave the narrow ones alone. Dates and reference codes keep their values intact; the description column takes two lines. Nothing is ever cut off.

CSV to PDF does all four, and tells you afterwards which delimiter it used, so a wrong guess is visible rather than silent.

Quoted fields, and why they matter

The CSV format has one genuinely tricky rule: a value can be wrapped in double quotes, and inside those quotes the separator loses its meaning.

Ref,Client,Notes SR-1,"Acme, Inc.","Line one line two"

Here Acme, Inc. is one value even though it contains a comma, and the note spans two physical lines while remaining a single cell. A converter that splits naively on commas and newlines will mangle both. Addresses and comment fields hit this constantly.

The other half of the rule is the escaped quote: "" inside a quoted field means one literal quote character. That is how O""Brien Ltd in the file means O"Brien Ltd on the page.

Portrait or landscape?

Landscape is usually right for tabular data, and it is the sensible default. A table is nearly always wider than it is tall, and rotating the page buys about 40% more width before anything has to be scaled down.

Portrait makes sense for a narrow table (four or five columns) where landscape would leave a lake of white space to the right.

Practical points

  • Check the encoding. A file saved as UTF-8 shows accented characters and non-Latin scripts correctly. One saved in a legacy encoding may show them as mojibake. Files carrying a byte-order mark, and UTF-16 files, are handled automatically.
  • Ragged rows are fine. Exports sometimes produce rows with fewer values than the header. Short rows are padded rather than dropped; you keep the data and can see the gap.
  • Very large files. A file with tens of thousands of rows produces a very long PDF. Filter it to what you need first. A hundred-page table is rarely read by anyone.
  • Not really a table? If your file is a log or a report rather than a grid, TXT to PDF will treat it as text and preserve its alignment, which is often what you actually want.

Before you convert: two minutes that save the whole job

A CSV is written for a machine to read, and it usually needs one pass of human judgement before it makes a good page.

Drop the columns nobody reads. Exports carry internal IDs, timestamps to the millisecond, and status flags that mean nothing outside the system that produced them. Every one of them takes width away from the columns people actually look at. Deleting five columns does more for readability than any font setting.

Shorten the headers. "customer_account_reference_number" forces its whole column wide enough to hold it. "Account" says the same thing to a reader and lets the numbers below it breathe.

Decide about the totals row. A CSV export often has one, and it will be styled exactly like every other row, which makes it easy to misread as data. If the table is going in front of somebody who has to act on it, move the total out of the CSV and put it in the surrounding document instead.

Numbers, and the one thing a table renderer cannot guess

A CSV has no formatting. The value 1234.5 might be currency, a quantity, or a measurement, and nothing in the file says which.

That means whatever you see in the PDF is exactly what was in the file: no thousands separators, no currency symbols, no fixed decimal places, and no way for the converter to add them without guessing what the numbers mean.

If those things matter, format them before exporting the CSV, not after. In a spreadsheet, that means writing the formatted value into a new column with a text function, then exporting that column instead of the raw one. It is one extra step and it is the only way to get "£1,234.50" rather than "1234.5" on the page.

A 25-column CSV converted with every column kept and the header repeated.

Common questions

My CSV uses semicolons. Will that work?

Yes. The delimiter is detected by testing which candidate gives a consistent column count across the file, so semicolon, tab and pipe files all work. The result screen names the one it used, and you can override it.

Will long values be cut off?

No. Columns are measured against their contents, and if the whole table is still too wide the widest columns are narrowed first and their text wraps. Nothing is truncated.

Does the header repeat on later pages?

Yes, on every page, which is the difference between a long table being readable and being a wall of numbers.

What about commas and line breaks inside a value?

Handled properly. Quoted fields may contain the delimiter, escaped quotes and line breaks; a two-line address stays two lines inside its cell.