Extract text from any PDF — digital, scanned, password-protected — in your browser. Urdu/Arabic OCR, Quran mode, entity extraction, DOCX export. No upload.
The PDF Text Extractor is a complete browser-side extraction pipeline: digital text parsing via Mozilla's pdf.js, OCR fallback through Tesseract.js for scanned pages, image preprocessing for low-quality scans, password-protected PDF support, Quran mode detection, entity extraction, and exports to TXT, Markdown, DOCX, HTML, and JSON. Every step happens in your browser — your documents never leave the device.
Most PDFs are mixed: some pages have a clean text layer while others are scanned images. Whole-document extractors fail on these documents. This tool classifies every page independently — if pdf.js returns garbled, missing, or invisible-only text, the page is re-rendered at high DPI and passed to OCR with the language you chose.
Non-Latin scripts are first-class. The OCR worker uses the matching Tesseract language pack (urd, ara, fas, heb), RTL pages are rendered at 3.5x scale for sharper glyph fidelity, and the output panel switches to right-to-left direction automatically. Quranic Unicode marks are auto-detected and preserved through NFC normalization.
No server, no API call, no upload, no tracking on document content. The Network tab shows zero outbound requests during extraction.
Pick the correct language before extraction — English OCR will produce garbage for Arabic, Urdu, Hindi, Chinese etc.
Language: setOcrLang(e.target.value)} disabled={isProcessing} className="px-3 py-2 rounded-md border border-border bg-background text-foreground text-sm min-w-[220px]" aria-label="OCR language" > Arabic (العربية) Urdu (اردو) Arabic + Urdu Urdu + English Arabic + English Persian / Farsi English Spanish French German Hindi Bengali Chinese (Simplified) Chinese (Traditional) Japanese Korean Russian Hebrew setShowSettings(s => !s)} className="text-sm text-primary underline-offset-4 hover:underline" aria-expanded={showSettings} > {showSettings ? 'Hide settings' : 'More settings'} {showSettings && ( setRemoveHF(e.target.checked)} disabled={isProcessing} /> Remove repeated headers / footers / page numbers { await cacheClear(); setCacheHit(false); toast({ title: 'Cache cleared' }); }} > Clear extraction cache (all PDFs) )} {/* PDF info panels */} {(metadata && (metadata.title || metadata.author || metadata.creator || metadata.pdfVersion)) && ({error}
)} {results.length > 0 && ( {/* Badges */} {isQuran && ( Quran PDF detected — special Unicode preserved )} {duplicates.length > 0 && ( Duplicate pages: {duplicates.map(g => g.join('=')).join(', ')} )} {cacheHit && ( Loaded from cache )} {/* Toolbar */} {results.length} page(s) · {results.reduce((s, r) => s + r.charCount, 0).toLocaleString()} chars · {results.filter(r => r.method === 'ocr').length} OCR · {results.filter(r => r.method === 'digital').length} digital setSearch(e.target.value)} placeholder="Search pages..." className="pl-7 h-8 w-44 text-sm" aria-label="Search extracted text" /> setExportRange(e.target.value)} placeholder="Pages: 1-3,5" className="h-8 w-32 text-sm" aria-label="Page range for export" /> Copy all handleDownload('txt')} className="gap-1"> TXT handleDownload('md')} className="gap-1"> MD handleDownload('docx')} className="gap-1"> DOCX handleDownload('html')} className="gap-1"> HTML handleDownload('pdf')} className="gap-1"> PDF handleDownload('json')} className="gap-1"> JSON {entities && (entities.emails.length + entities.phones.length + entities.urls.length + entities.dates.length + entities.cnics.length + entities.currencies.length > 0) && ( handleDownload('entities-csv')} className="gap-1"> Entities CSV )} {extractingImages ? 'Extracting…' : 'Images (ZIP)'} {extractingTables ? 'Detecting…' : 'Tables (CSV)'} setShowThumbs(s => !s)} className="gap-1"> {showThumbs ? 'Hide preview' : 'Split-view preview'} {/* Entities panel */} {entities && (entities.emails.length + entities.phones.length + entities.urls.length + entities.dates.length + entities.cnics.length + entities.currencies.length > 0) && (No pages match "{search}".
)} {filteredResults.map((r) => ( Page {r.page}It runs a hybrid per-page engine — pdf.js first, then Tesseract OCR fallback for any garbled or scanned page. Invisible text layers are dropped, hyphenation is repaired, repeated headers/footers/page numbers are stripped, and entities (emails, phones, dates, CNICs) are extracted automatically.
Yes. If the PDF asks for a password, the tool prompts you to enter it. The password stays in your browser — never sent anywhere.
Yes. The tool auto-detects Quranic Unicode marks (U+06D6–U+06ED, ﴾ ﴿, Bismillah ﷽) and shows a Quran badge. Tajweed marks and ayah numbers are preserved through NFC normalization.
Yes — DOCX export is built in. RTL pages are exported with proper right-to-left direction so Urdu and Arabic open correctly in Word.
No. Every step — parsing, OCR, entity extraction, export generation — happens in your browser. You can verify with DevTools Network tab.
Stylized fonts (especially Nastaliq calligraphy) push browser OCR to its physical limit. Re-OCR a page with a different language combination, or edit inline.
Browse all free tools · Guides and tutorials · PDF tools · Developer tools · Text tools · SEO tools