Image to Text — Free OCR in Your Browser
Click to choose an image — or drag it here, or press Ctrl+V
JPG, PNG, WebP, GIF, BMP, AVIF · up to 25 pages at once · read in your browser, never uploaded
Drop in a screenshot, a photo of a page, or a scanned receipt and the text comes back as something you can select and copy. Recognition runs on a WebAssembly build of Tesseract inside the tab, so the picture itself never leaves your machine — which also means there is no page cap, no watermark, and no email box to fill in. Pick one or two of the 34 languages, choose how the page is laid out, and the cleanup switches turn the ragged line-per-line output into flowing paragraphs.
What accuracy to expect
Character accuracy depends almost entirely on how many pixels each letter has and how cleanly it separates from the background. These are the ranges Tesseract's LSTM engine typically lands in, and what each one costs you in proofreading.
| Source | Typical accuracy | What that means in practice |
|---|---|---|
| Flatbed scan, printed page, 300 DPI | 98–99% | A mistake every two or three lines |
| Sharp phone photo of a book, flat and lit | 95–98% | A few words per paragraph need a look |
| Screenshot at 2× zoom or on a retina display | 95–99% | Usually clean; punctuation is the weak point |
| Screenshot of 11 px UI text at 100% zoom | 70–90% | Fix with upscaling, or re-shoot it zoomed in |
| Thermal receipt, faded | 60–85% | Totals and dates have to be checked by hand |
| Photo at an angle, or in shadow | 50–80% | Re-shoot it square to the page instead |
| Cursive handwriting | near zero | Not what the model was trained on |
Which page layout to pick
This setting is Tesseract's page segmentation mode, and it is the single control that changes results the most. Automatic runs full layout analysis, which is the right call on a document and the wrong one on a screenshot with six words scattered across it — there, Tesseract looks for columns that do not exist and throws text away.
| Option | Tesseract PSM | Use it for |
|---|---|---|
| Automatic | 3 | Book pages, letters, articles, anything page-shaped |
| Sparse text | 11 | App screenshots, error dialogs, signs, memes, labels |
| Single block | 6 | A cropped paragraph, a quote, a code snippet |
| Single column | 4 | Receipts, invoices, a narrow column of mixed type sizes |
| Single line | 7 | A serial number, a licence plate, one heading |
If the output is wrong
- Nothing came back at all. Switch the layout to Sparse text. Automatic mode can decide a whole screenshot is a picture and return an empty string.
- Letters are right, accents are missing. The wrong language is loaded. English traineddata has no model for ä, ş or я, so it substitutes the closest Latin shape it knows.
- Random punctuation between letters. JPEG artefacts around the glyph edges. Turn on Boost contrast, and if you still have the original, save it as PNG instead of re-saving a JPEG.
- Dark-mode screenshot is unreadable. Turn on Invert. The tool guesses this from the first image, but it only guesses once.
- 0 and O, 1 and l are swapped. Those pairs are identical in many sans-serif fonts, and without a dictionary word around them the model has nothing to go on. Serial numbers and codes always need checking.
- Columns are interleaved. Crop the page to one column at a time and run them separately. Two-column layout analysis fails when the gutter is narrow or the text is skewed.
How it works
- 1 Add the image Click the drop area, drag files onto it, or just press Ctrl+V to paste a screenshot straight from the clipboard. JPG, PNG, WebP, GIF, BMP, and AVIF all work, and you can queue up to 25 pages at once.
- 2 Pick the language Choose the language of the text, not your interface language — this is what the recognition model is loaded for. If the page mixes two, add a second language; a Russian document with English brand names reads much better as rus+eng.
- 3 Match the layout Leave it on Automatic for a normal page. Switch to Sparse text for screenshots, UI, and labels scattered around an image, Single line for one line of large type, and Single column for a narrow column of body text.
- 4 Extract and clean up Press Extract text and watch the progress bar. When it finishes, the Join wrapped lines switch turns one-line-per-visual-line output into real paragraphs, then copy the result or download it as a .txt file.
Your images stay on your computer
The photo is decoded and read inside this tab — it is never uploaded, and no copy exists on our side to delete. One thing to be straight about: the first run downloads the recognition engine (about 3 MB of WebAssembly) and the language data (5–15 MB per language) from the jsDelivr CDN, and your browser caches both. So the first extraction needs a connection; after that the tool keeps working offline. The images and the text are never part of any request.
Frequently asked questions
- Is my image uploaded anywhere?
- No. The file is decoded by your browser and read by a WebAssembly build of Tesseract running in a worker thread in this tab. There is no upload request, no temporary file on a server, and nothing for us to delete afterwards. The only network traffic is the one-time download of the engine and the language data from jsDelivr — those are public static files and they carry no information about your image.
- How accurate is browser OCR?
- On a clean 300 DPI scan of printed text, Tesseract's LSTM engine lands around 98–99% of characters right, which is a typo every two or three lines. A sharp phone photo of a book page is usually 95–98%. A photo taken at an angle in dim light, or a screenshot of 11 px UI text, can fall below 85%, and that is the point where fixing the output takes longer than retyping it. The tool prints the mean confidence after every run so you know which case you are in before you start trusting the text.
- Can it read handwriting?
- Not reliably, no. Tesseract was trained on printed type, and the standard traineddata files have no handwriting model. Very neat block capitals sometimes come through; ordinary cursive comes back as noise. If you need handwriting, you need a cloud service trained on it — and that means uploading the image, which is exactly the trade-off this tool exists to avoid.
- Why is my screenshot coming out as gibberish?
- Almost always because the text is too small in pixels. Tesseract expects a capital letter around 30 px tall; a screenshot of a web page at 100% zoom gives it about 11. The tool already upscales anything under 1400 px on its short side, which fixes most of it, but you will get a better result by taking the screenshot zoomed in, at 200% browser zoom or on a retina display. Two other things to try: switch the layout to Sparse text, and turn on Invert if the text is light on a dark background.
- Why does the text come out with a line break after every line?
- Because that is literally what OCR sees — it reports one line of text per line on the page, with no idea which of those breaks were the author's and which were just the right margin. Turn on Join wrapped lines. It merges a line into the next one unless the line ended on a full stop, colon, or closing quote, or the next line opens a bullet or a number, so paragraphs flow again while lists keep their shape. Dehyphenate handles the other half of the problem: a word split as "recog-" / "nition" across two lines is put back together.
- Can it read a PDF?
- Not directly — this tool takes images. For a scanned PDF, render the pages to PNG first and drop them in here; you can queue 25 at a time and they come out as one text file with a header per page. If your PDF already has a text layer (anything exported from Word, a browser, or most invoicing software), you do not need OCR at all: select the text and copy it, or use a PDF text extractor, which is both exact and instant.
- Which languages can it recognise?
- 34, including English, Russian, Ukrainian, German, French, Spanish, Italian, Portuguese, Polish, Czech, Turkish, Greek, Arabic, Hebrew, Hindi, Thai, Vietnamese, Chinese (Simplified), Japanese, and Korean. You can combine any two in one pass. Combining costs accuracy on each, though — the model has more candidates to choose between at every character — so only add a second language if the page genuinely mixes scripts.
- Is there a page or file-size limit?
- 25 images per batch, and that is a limit on memory rather than on your account, because every page is held as decoded pixels in the tab. There is no daily quota and no file-size cap beyond what your browser can decode. Anything over 3500 px on its longest side is scaled down to that before recognition — past that point Tesseract gets slower without getting more accurate.
- What does the confidence number mean?
- It is Tesseract's own mean word confidence for the page, from 0 to 100, weighted here by how much text each page produced. Above 88 the output is usually safe to use with a skim. Between 75 and 88, check the digits and proper nouns specifically — those are where the model has the least language context to fall back on, so an invoice total or a surname is far more likely to be wrong than an ordinary word. Below 75, treat the result as a draft and read every line.
- Does it keep the table structure?
- Partly. Preserving inter-word spaces is switched on, so columns in a receipt or a simple table usually stay visually aligned in the output, and you can see where the rows were. What you do not get is a real table — no cells, no tab separators, no CSV. Tesseract can produce a hOCR layout with bounding boxes, but turning that back into rows and columns is a separate problem that no free tool solves well.