belun.app Blog

Image to Text — Free OCR in Your Browser

Click to choose an image — or drag it here, or press Ctrl+V

JPG, PNG, WebP, GIF, BMP, AVIF · up to 25 pages at once · read in your browser, never uploaded

Drop in a screenshot, a photo of a page, or a scanned receipt and the text comes back as something you can select and copy. Recognition runs on a WebAssembly build of Tesseract inside the tab, so the picture itself never leaves your machine — which also means there is no page cap, no watermark, and no email box to fill in. Pick one or two of the 34 languages, choose how the page is laid out, and the cleanup switches turn the ragged line-per-line output into flowing paragraphs.

What accuracy to expect

Character accuracy depends almost entirely on how many pixels each letter has and how cleanly it separates from the background. These are the ranges Tesseract's LSTM engine typically lands in, and what each one costs you in proofreading.

Source Typical accuracy What that means in practice
Flatbed scan, printed page, 300 DPI 98–99% A mistake every two or three lines
Sharp phone photo of a book, flat and lit 95–98% A few words per paragraph need a look
Screenshot at 2× zoom or on a retina display 95–99% Usually clean; punctuation is the weak point
Screenshot of 11 px UI text at 100% zoom 70–90% Fix with upscaling, or re-shoot it zoomed in
Thermal receipt, faded 60–85% Totals and dates have to be checked by hand
Photo at an angle, or in shadow 50–80% Re-shoot it square to the page instead
Cursive handwriting near zero Not what the model was trained on

Which page layout to pick

This setting is Tesseract's page segmentation mode, and it is the single control that changes results the most. Automatic runs full layout analysis, which is the right call on a document and the wrong one on a screenshot with six words scattered across it — there, Tesseract looks for columns that do not exist and throws text away.

Option Tesseract PSM Use it for
Automatic 3 Book pages, letters, articles, anything page-shaped
Sparse text 11 App screenshots, error dialogs, signs, memes, labels
Single block 6 A cropped paragraph, a quote, a code snippet
Single column 4 Receipts, invoices, a narrow column of mixed type sizes
Single line 7 A serial number, a licence plate, one heading

If the output is wrong

  • Nothing came back at all. Switch the layout to Sparse text. Automatic mode can decide a whole screenshot is a picture and return an empty string.
  • Letters are right, accents are missing. The wrong language is loaded. English traineddata has no model for ä, ş or я, so it substitutes the closest Latin shape it knows.
  • Random punctuation between letters. JPEG artefacts around the glyph edges. Turn on Boost contrast, and if you still have the original, save it as PNG instead of re-saving a JPEG.
  • Dark-mode screenshot is unreadable. Turn on Invert. The tool guesses this from the first image, but it only guesses once.
  • 0 and O, 1 and l are swapped. Those pairs are identical in many sans-serif fonts, and without a dictionary word around them the model has nothing to go on. Serial numbers and codes always need checking.
  • Columns are interleaved. Crop the page to one column at a time and run them separately. Two-column layout analysis fails when the gutter is narrow or the text is skewed.

How it works

  1. 1
    Add the image Click the drop area, drag files onto it, or just press Ctrl+V to paste a screenshot straight from the clipboard. JPG, PNG, WebP, GIF, BMP, and AVIF all work, and you can queue up to 25 pages at once.
  2. 2
    Pick the language Choose the language of the text, not your interface language — this is what the recognition model is loaded for. If the page mixes two, add a second language; a Russian document with English brand names reads much better as rus+eng.
  3. 3
    Match the layout Leave it on Automatic for a normal page. Switch to Sparse text for screenshots, UI, and labels scattered around an image, Single line for one line of large type, and Single column for a narrow column of body text.
  4. 4
    Extract and clean up Press Extract text and watch the progress bar. When it finishes, the Join wrapped lines switch turns one-line-per-visual-line output into real paragraphs, then copy the result or download it as a .txt file.

Your images stay on your computer

The photo is decoded and read inside this tab — it is never uploaded, and no copy exists on our side to delete. One thing to be straight about: the first run downloads the recognition engine (about 3 MB of WebAssembly) and the language data (5–15 MB per language) from the jsDelivr CDN, and your browser caches both. So the first extraction needs a connection; after that the tool keeps working offline. The images and the text are never part of any request.

Frequently asked questions

Is my image uploaded anywhere?
No. The file is decoded by your browser and read by a WebAssembly build of Tesseract running in a worker thread in this tab. There is no upload request, no temporary file on a server, and nothing for us to delete afterwards. The only network traffic is the one-time download of the engine and the language data from jsDelivr — those are public static files and they carry no information about your image.
How accurate is browser OCR?
On a clean 300 DPI scan of printed text, Tesseract's LSTM engine lands around 98–99% of characters right, which is a typo every two or three lines. A sharp phone photo of a book page is usually 95–98%. A photo taken at an angle in dim light, or a screenshot of 11 px UI text, can fall below 85%, and that is the point where fixing the output takes longer than retyping it. The tool prints the mean confidence after every run so you know which case you are in before you start trusting the text.
Can it read handwriting?
Not reliably, no. Tesseract was trained on printed type, and the standard traineddata files have no handwriting model. Very neat block capitals sometimes come through; ordinary cursive comes back as noise. If you need handwriting, you need a cloud service trained on it — and that means uploading the image, which is exactly the trade-off this tool exists to avoid.
Why is my screenshot coming out as gibberish?
Almost always because the text is too small in pixels. Tesseract expects a capital letter around 30 px tall; a screenshot of a web page at 100% zoom gives it about 11. The tool already upscales anything under 1400 px on its short side, which fixes most of it, but you will get a better result by taking the screenshot zoomed in, at 200% browser zoom or on a retina display. Two other things to try: switch the layout to Sparse text, and turn on Invert if the text is light on a dark background.
Why does the text come out with a line break after every line?
Because that is literally what OCR sees — it reports one line of text per line on the page, with no idea which of those breaks were the author's and which were just the right margin. Turn on Join wrapped lines. It merges a line into the next one unless the line ended on a full stop, colon, or closing quote, or the next line opens a bullet or a number, so paragraphs flow again while lists keep their shape. Dehyphenate handles the other half of the problem: a word split as "recog-" / "nition" across two lines is put back together.
Can it read a PDF?
Not directly — this tool takes images. For a scanned PDF, render the pages to PNG first and drop them in here; you can queue 25 at a time and they come out as one text file with a header per page. If your PDF already has a text layer (anything exported from Word, a browser, or most invoicing software), you do not need OCR at all: select the text and copy it, or use a PDF text extractor, which is both exact and instant.
Which languages can it recognise?
34, including English, Russian, Ukrainian, German, French, Spanish, Italian, Portuguese, Polish, Czech, Turkish, Greek, Arabic, Hebrew, Hindi, Thai, Vietnamese, Chinese (Simplified), Japanese, and Korean. You can combine any two in one pass. Combining costs accuracy on each, though — the model has more candidates to choose between at every character — so only add a second language if the page genuinely mixes scripts.
Is there a page or file-size limit?
25 images per batch, and that is a limit on memory rather than on your account, because every page is held as decoded pixels in the tab. There is no daily quota and no file-size cap beyond what your browser can decode. Anything over 3500 px on its longest side is scaled down to that before recognition — past that point Tesseract gets slower without getting more accurate.
What does the confidence number mean?
It is Tesseract's own mean word confidence for the page, from 0 to 100, weighted here by how much text each page produced. Above 88 the output is usually safe to use with a skim. Between 75 and 88, check the digits and proper nouns specifically — those are where the model has the least language context to fall back on, so an invoice total or a surname is far more likely to be wrong than an ordinary word. Below 75, treat the result as a draft and read every line.
Does it keep the table structure?
Partly. Preserving inter-word spaces is switched on, so columns in a receipt or a simple table usually stay visually aligned in the output, and you can see where the rows were. What you do not get is a real table — no cells, no tab separators, no CSV. Tesseract can produce a hOCR layout with bounding boxes, but turning that back into rows and columns is a separate problem that no free tool solves well.

From the blog

Why Your OCR Output Is Garbage (and How to Fix It) Resolution, contrast, page segmentation mode — the three settings that decide whether OCR gives you clean text or noise. Read the post →

Related tools