Half the context I want to give a model is locked inside something I can't
select from: a scanned book, a slide deck, a course viewer, a "PDF" that's
really page images. Copy-paste gets you nothing, and screenshotting 200 pages by
hand isn't a plan.<p>OCR It is a Chrome extension for that gap. You drag out a capture region once —
the text block of the reader, say. After that, one hotkey per page screenshots
that exact rectangle, OCRs it, and appends the result to a running transcript.
Or start an auto-run and it captures, turns the page, and repeats until the
document ends. Then Copy all, or Download .txt, and you have a file to paste
into Claude or drop into an agent's context.<p>Everything runs locally. Tesseract's wasm build and the language data (~10 MB)
are committed into the extension, so there are no network requests at all, no
API key, and no host permissions at install — single captures ride on activeTab.
The irony of an AI-adjacent tool that never talks to a server was not lost on
me, but the pages you're capturing are often exactly the ones you don't want to
ship to a third party.<p>Three things turned out more interesting than expected:<p>- MV3 service workers have no DOM and no Worker, so cropping and OCR live in an
offscreen document.<p>- The next-page control is stored as a <i>point</i>, not a CSS selector. A point
survives DOM re-renders and reaches into cross-origin iframes and shadow
roots, which nothing the top frame can express does. Routing it was the fiddly
part: window.screenX inside an iframe reports the browser window, not the
frame, so frames locate themselves by walking same-origin ancestors, and
across an origin boundary the parent hands the offset down by postMessage.<p>- The auto-run waits for each page's OCR before turning. That's what makes
end-of-document detection work; a timer-based loop sails past the last page
and fills your transcript with copies of it.<p>Limitations: Chrome's own PDF viewer can't be auto-advanced (it's a plugin no
extension can inject into, though capturing from it works fine); the region is a
fixed rectangle on screen, so resizing or zooming mid-run breaks it; and
accuracy tracks the source — crisp rendered text reads at 93-95% confidence,
scans need cleanup before they're worth feeding to anything.<p>Tests drive a real headless Chrome over CDP, which had its own surprises:
Chrome 137+ ignores --load-extension, and headless can't show the
optional-permission prompt, so the suite installs a copy with the grant baked in
plus a real toolbar click via Extensions.triggerAction to prove the ungranted
path still works.<p>MIT, no build step: <a href="https://github.com/thiagotigaz/ocr-it" rel="nofollow">https://github.com/thiagotigaz/ocr-it</a>