Working with scans

How do I get the text out of a screenshot or a photo?

Run the image through optical character recognition, which reads the shapes of the letters and gives you back real, selectable text.

5 min read

It might be an error message you need to search for, a page of a book, a receipt, or a slide someone photographed instead of sharing. The words are right there, but you can't select them.

To a computer, a picture of text is only a grid of coloured dots that look like letters to you. To get the words out, something has to recognise them, and that is what OCR does.

What OCR is doing

Optical character recognition works in steps, and knowing them explains the odd results you sometimes get.

First it finds the text: which parts of the image are writing and which are background, picture or noise. Then it splits out lines, then words, then single characters. Then it matches each character's shape against the letters it knows, and finally it checks the result against a dictionary. That is how "rn" gets corrected to "m" when the word around it makes that clear.

Each of those steps needs a clear enough image to get it right. That is why image quality matters far more than which tool you use.

What makes recognition accurate

Resolution. A character needs roughly 20 pixels of height to be recognised reliably. Text photographed from across a room does not have that, no matter how sharp the photo looks to you.

Contrast. Black on white is ideal. Grey on grey is hard. Text over a photo is the hardest case, because it is hard to tell the writing from the background.

Straightness. Every degree of tilt makes the lines harder to find. A page photographed at an angle, with the far edge smaller than the near one, is harder still, because the letters change size across the line.

Even lighting. A shadow across half the page, or a bright reflection from a window, can make one region unreadable while the rest is perfect.

Plain type. Ordinary printed fonts are recognised almost perfectly. Handwriting is a different and much harder problem. Decorative and script fonts sit in between and are unreliable.

Getting a good capture

If you are taking the photo yourself, four habits help accuracy more than anything else:

  1. Fill the frame with the text. Get close. Background is wasted pixels.
  2. Shoot straight down, not at an angle.
  3. Use even light. Near a window is good; direct sunlight and camera flash both create hot spots.
  4. Hold still. Motion blur destroys the letter shapes that recognition depends on.

A screenshot is better than a photo of a screen every time: no lens, no lighting, no angle, and perfectly sharp edges.

Doing it

Image to Text takes a JPG, PNG or WebP and gives you the recognised text. It handles several languages, and telling it which language the text is in improves accuracy, because the dictionary check then uses the right words.

If your source is a scanned PDF instead of an image, PDF to Text does the same job.

What to check in the result

OCR doesn't warn you when it goes wrong. It gives you text that looks right but isn't, so proofreading matters more here than usual.

The usual trouble spots:

  • 0 and O, 1 and l and I, 5 and S, 8 and B. Worst in reference numbers, where the dictionary check can't help because the string is not a word.
  • Decimal points and thousands separators. A misread separator changes a number by a factor of a thousand and looks perfectly normal.
  • Column boundaries. Two columns of text can be read straight across, interleaving two unrelated sentences.
  • Line breaks. Hyphenated words split across lines sometimes stay hyphenated.

If you are pulling out figures from an invoice, a statement or a table, check every number. If it is ordinary text, read it once. A wrong word in a paragraph stands out. A wrong digit in an amount doesn't.

When OCR is not the answer

The text is already selectable. Try selecting it first. If a text cursor appears, the words are real and you can copy them directly. Running OCR would only add errors to text that is already perfect.

The document exists somewhere as a file. Asking for the original is always better than recognising a picture of it. Ten seconds of asking beats proofreading.

It is handwriting. Printed-text OCR is not built for it, and the results will be poor enough to retype anyway.

Languages, and why telling it matters

Recognition uses a dictionary to decide between similar characters, and a dictionary in the wrong language makes things worse. It will "correct" a word that was read right into a similar-looking word from the language it expected.

So choosing the right language matters. With the right one, you get a few wrong letters. With the wrong one, you can get smooth-looking nonsense.

Two special cases:

Mixed-language documents. A page of Bengali with English technical terms is hard, because each language needs its own recognition. Running it twice, once per language, and keeping the good parts of each is clumsy, but it usually beats either run on its own.

Non-Latin scripts. Accuracy varies a lot by script and by how much the letter shapes depend on the letters around them. Connected scripts are harder than separated ones, and any script where a character changes shape depending on its neighbours is harder still. Expect to proofread more, not less.

What to do with the result

Recognised text comes as plain text with no formatting. The formatting was only part of the picture, so it can't be recovered.

If the source was a document with structure worth keeping, the practical route is to recognise the text, paste it into a word processor, apply headings and lists by hand, and produce a clean PDF with Word to PDF. You end up with a document that is searchable, selectable and a fraction of the size of the scan.

That is more work than a one-step converter. But it is the only way to get a correct result, because the structure has to come from a person who can read the page.

Common questions

Why is my OCR result full of mistakes?

Almost always image quality rather than the tool. Recognition needs roughly 20 pixels of character height, good contrast, a straight page and even lighting. Photographed-from-a-distance text fails all four.

Does OCR work on handwriting?

Not reliably. Printed-text recognition is a different problem from handwriting recognition, and the results are usually poor enough that retyping is faster.

Which characters get misread most often?

0 and O, 1 and l and I, 5 and S, 8 and B. They are worst inside reference numbers, where the dictionary check can't help because the string is not a word.

Should I enhance the image first?

Yes, if it is grey, speckled or low-contrast. Sharpening and a black-and-white conversion fix what the first step of recognition struggles with.