Extract Text (OCR)
Reads the real text layer where it exists, and falls back to English/Arabic OCR where it doesn't.
Pull text out of any PDF — typed or scanned.
Extract Text checks each page of a PDF for a real, selectable text layer first — that's instant and exact. Only pages with little or no extractable text (i.e. scans) fall back to on-device OCR, in English and Arabic. Everything runs in your browser: nothing uploads to a server, though the one-time OCR language pack is fetched from a CDN the first time you use it. Download the result as .txt or .docx, or copy it straight to your clipboard.
- Batch processing
- API access
- Encrypted transport
- No file retention
What people use Extract Text (OCR) for.
Pull the text out of a scanned contract so you can search or quote it.
One PDF with some typed pages and some scanned pages — both get handled automatically.
Grab a paragraph from a scan without retyping it.
Get clean text out before pasting into a translator, editor, or search index.
Three steps, no learning curve.
Built for people who care about their documents.
No need to know in advance which pages are scanned — each page is checked on its own.
The OCR fallback recognises English and Arabic text.
Download in whichever format you need, or just copy the text.
The PDF itself is never uploaded — only a one-time OCR language pack is fetched from a CDN.
Your documents belong to you.
Sable is designed around a single principle: your files are yours. We never train on them, never sell them, and never keep them longer than we need to.