Skip to the tool

Extract text from images.

Words read out of images, right on your device — free OCR.

No uploads — 100% local No ads Free & open source

Drop images here

or browse your files

Files never leave your device. Everything runs in your browser, nothing touches a server — tools you've used even work offline.

Extract text from images entirely in your browser — screenshots, photos of documents, scans and whiteboards, recognized on your own device. Pick the document language, drop the images, and each one comes back as a .txt file. Nothing is ever uploaded.

Before / after

Scan of the opening page of "A Scandal in Bohemia" from the 1892 first edition of The Adventures of Sherlock Holmes — the page the text below was recognized from

Extracted text — verbatim tool output

Hoventure F
A SCANDAL IN BOHEMIA
I

RztOvKeN GSO Sherlock Holmes she is always the woman. I
Ae have seldom heard him mention her under any

W ati other name. In his eyes she eclipses and predom-
ONS inates the whole of her sex. It was not that he
felt any emotion akin to love for Irene Adler. All emotions,
and that one particularly, were abhorrent to his cold, precise,
but admirably balanced mind. He was, I take it, the most
perfect reasoning and observing machine that the world has
seen ; but, as a lover, he would have placed himself in a false
position. He never spoke of the softer passions, save with
a gibe and a sneer. They were admirable things for the
observer—excellent for drawing the veil from men’s motives
and actions. But for the trained reasoner to admit such in-
trusions into his own delicate and finely adjusted tempera-
ment was to introduce a distracting factor which might throw
a doubt upon all his mental results, Grit in a sensitive instru-
ment, or a crack in one of his own high-power lenses, would
not be more disturbing than a strong emotion in a nature
such as his. And yet there was but one woman to him, and
that woman was the late Irene Adler, of dubious and question-
able memory.

: I had seen little of Holmes lately. My marriage had drifted
us away from each other. My own complete happiness, and
the home-centred interests which rise up around the man who
first finds himself master of his own establishment, were suffi-

Words recognized

266

Language

English

Output

.txt — 1.5 KB

Real result, not a mock-up: the opening page of "A Scandal in Bohemia" — the 1892 first book edition of The Adventures of Sherlock Holmes — went through the Image to Text tool — Tesseract, the OCR engine that has read the world's paper for decades, compiled to WebAssembly — with the language set to English. The panel above is the verbatim output: 266 words of selectable, copyable text out of a flat scan, uncorrected. The Victorian body type comes out nearly flawless — "To Sherlock Holmes she is always the woman" and all; only the ornamental drop capital and the blackletter heading defeat the engine, which is the honest trade a 130-year-old page offers. Drop the same scan in yourself and you'll get the same text.

Book by Arthur Conan Doyle — public domain.

How it works

  1. Drop files anywhere on the page, click to browse, or paste with ⌘V.
  2. Pick a quality or preset — or set an exact target size and let the tool find it.
  3. Compress, compare before/after, and download — individually or as a ZIP.

What OCR reads well — and what it does not

Printed text is the sweet spot: books, receipts, screenshots, signs and scans recognize reliably when the text is sharp and reasonably horizontal. Handwriting is not — cursive and freeform notes come out garbled, and no browser OCR changes that honestly. Resolution matters too: if you can barely read it zoomed in, neither can the recognizer. Scanned PDFs have their own tool — OCR PDF keeps the pages and adds a searchable layer.

Eight languages, downloaded once

Recognition models for English, Slovenian, German, Italian, French, Spanish, Portuguese and Croatian ship with the site — the one you pick (1–2 MB) downloads on first use and stays cached, so later runs work offline. Choosing the right language matters: an English model reading German text will miss every umlaut.

Under the hood

Recognition runs on Tesseract — the open-source OCR engine that has read the world’s paper for decades — compiled to WebAssembly and running in a worker on your device. The page you are on only ever serves files; your images, the recognized text and the language models all live and die in your browser. That is also why there are no page limits and no queues: your hardware does the reading.

Frequently asked questions

Which languages are supported?

English, Slovenian, German, Italian, French, Spanish, Portuguese and Croatian. Pick the document’s language before running — recognition quality depends on it.

Does it read handwriting?

Not usefully — Tesseract is built for printed text. Clean print recognizes well; cursive and freeform handwriting come out garbled.

Why is my result poor?

Usually resolution or language: blurry, small or skewed text defeats any OCR, and the wrong language model misses accented characters. Use a sharper capture and double-check the language picker.

Are my images uploaded?

No — the pixels never leave your machine. Decoding and re-encoding both happen in your browser; there is no upload to wait for and no server-side copy to worry about afterwards. Want proof? Run one file through, switch your connection off, and run another — it still works.