📜 Indic Reader

Gujarati · Hindi · Sanskrit

Source Input

📥

Drop files here or use the buttons below

Images · PDF · DOCX · TXT · RTF · HTML · XML

OCR & playback options are in ⚙️ Settings
0%
1.0×
🕉 Chant
Auto-scroll
Pause at skip
📖 Continuous
Skipped section
Continue where you left off?

Something went wrong

    📞 No Gujarati voice on this device

      🔐 Enable Google Drive sign-in

      One-time setup. After this you can open any Drive file you own — no link sharing needed. Your Google account stays with Google; this app only receives temporary read-only access during pick.
      📖 How to get the two values (≈ 5 minutes)
      1. Open Google Cloud Console and create a project (any name).
      2. APIs & Services → Library. Search for and enable both:
        • Google Picker API
        • Google Drive API
      3. APIs & Services → OAuth consent screen → User Type: External → fill App name, your email, save. Add your Google email under Test users.
      4. APIs & Services → Credentials → + Create Credentials → OAuth client ID → Application type: Web application → under "Authorized JavaScript origins" add the URL where this app runs. Click Create. Copy the Client ID (ends in .apps.googleusercontent.com).
      5. + Create Credentials → API key. Copy it (starts with AIzaSy). Optionally restrict it to the same HTTP referrer for security.
      6. Paste both values below and click Save.
      💡 The current page URL is: — paste this into "Authorized JavaScript origins" in step 4.
      🔒 Where to store these credentials

      ☁️ Cloud OCR Setup

      Cloud OCR significantly improves accuracy on poor-quality scans, photos, and books with underlines or marginalia. The local browser OCR remains a transparent fallback if cloud OCR is unreachable.
      ⚠ Privacy: When enabled, page images are sent to your chosen cloud provider for processing. Files no longer stay only on your device. Each provider has its own data-handling policy — review before enabling.
      💡 Saving here only stores your credentials. To actually use cloud OCR, switch on the ☁️ Cloud OCR toggle on the main screen.
      🔒 Device password manager — store the key/URL in your device's secure store (Google Password Manager, iCloud Keychain, or Windows), protected by your device lock.

      Privacy Policy & Terms of Use

      📚 Educational use only

      Indic Reader is a free, non-commercial tool for the educational study of Gujarati, Hindi, and Sanskrit scriptures and literature. By using this software you agree to use it solely for personal study, teaching, or non-commercial research. Commercial use, redistribution, or rehosting requires written permission from the author.

      © Copyright

      © 2026 The Indic Reader contributors. Software provided "as is" without warranty of any kind. Source code is released for educational and personal use. All rights reserved.

      Scripture and document content uploaded or pasted by you remains entirely your responsibility. Ensure you have the right to access, narrate, and translate any text you process. Where source materials are copyrighted (translations, modern commentaries, scanned editions), respect those rights.

      🔒 Data privacy

      What stays on your device (always):

      • Every file you open — PDFs, images, DOCX, pasted text
      • OCR results and extracted text
      • OCR cache — recognized text and word positions are cached per page (in this browser's IndexedDB), keyed by file + page + OCR settings, so reopening or re-reading a page is instant and avoids repeat cloud-OCR calls. Bounded in size (oldest entries auto-removed); cleared when you turn off the 🗄 OCR cache toggle.
      • Continue / reopen — the last document you opened is cached on-device (IndexedDB) so the "Continue" button can reopen it; your reading position (page, line, word) is remembered to resume where you left off.
      • Page bookmarks (including any notes you attach and the full line text), page selections, and current playback position
      • Theme preference, voice preference, chant mode, auto-scroll, language filter, and the Prefs toggles (Pause on call, Resume, OCR cache, Check OCR)
      • Drive Picker credentials (if you configured them) — stored only in this browser's localStorage
      • Device password manager (optional) — if you tap "Save to device" in Cloud OCR settings, your API key or Worker URL is handed to your browser's Credential Management API and held by your operating system's secure password store (Google Password Manager, iCloud Keychain, or Windows), protected by your device lock. The app and its (non-existent) server never see it again except when you tap "Restore". Availability depends on your browser/OS.
      • Translation cache — built up locally as you tap Show Meanings or Translate
      • Manual corrections — your per-word and per-line language overrides, line text corrections (from inline line edit or Re-OCR), manual verse references (for vedabase.io links), and saved sandhi splits

      What may go online (only with your action):

      • Show Meanings sends individual Sanskrit words to two third-party services for accurate dictionary lookup. Disabled by default; requires explicit consent. The Cologne University Monier-Williams API (api.c-salt.uni-koeln.de) is queried first as the authoritative source. Google Translate's public endpoint is used as a fallback and to translate English definitions into Gujarati. Each call exposes your IP address and the word being looked up to these services.
      • Cloud OCR (optional, disabled by default; requires explicit consent) sends the entire page image to one of two providers you can configure:
        • Google Cloud Vision via your own Cloudflare Worker — the page image traverses Cloudflare's edge (subject to their data-processing terms) and is sent to Google Cloud Vision. Google's privacy policy applies; per Google's documentation, Vision API requests are not used to improve their models without your explicit consent in the Google Cloud Console settings.
        • OCR.space direct — the page image is sent to ocr.space servers. Per their privacy policy, free-tier requests may be retained for service improvement.
        When Cloud OCR is enabled, files no longer stay only on your device. The processing provider sees your IP address and the page image contents. Recognition results return to your browser only — they are not stored on our side.
      • Split Sandhi (optional, disabled until you use it) sends only the Sanskrit verse text — never any file or image — to the University of Hyderabad Heritage segmenter (sanskrit.uohyd.ac.in), a public academic service, routed through your own Cloudflare Worker. The Worker caches results for 24 hours. If the service is unreachable, the app silently falls back to an on-device rule-based assistant; no text leaves your device in that case. The segmenter sees the verse text and (via the Worker) Cloudflare's edge sees the request. The overlay also shows a per-word meaning in your chosen language (Gujarati/Hindi/English) — this sends each individual word to Google Translate, and tapping a word queries the Cologne Monier-Williams dictionary, both subject to the same consent as Show Meanings below. Your chosen meaning-language is stored locally.
      • Re-OCR (region, single line, or continuous next page) (optional) sends only an image to your chosen OCR provider (Google Cloud Vision via your Worker, or OCR.space, or processed entirely on-device with local OCR): the cropped rectangle you select, the single line you choose to fix, or — in Continuous mode — the next page being prepared just before narration reaches it. You choose the language hint each time. Same provider data-handling as Cloud OCR above; results are cached on-device.
      • Translate verse / Translate page (optional; requires consent) sends text — a single Sanskrit verse, or the current page's Gujarati/Hindi lines — to Google Translate's public endpoint to render it in another language. Sanskrit verses are not sent when you use page translate (Gujarati↔Hindi). Your IP address and the submitted text are visible to Google; no file or image is sent. Results are cached on-device.
      • Google Drive Picker/link load — only when you click those buttons. Drive sees the URLs you fetch and (for Picker) your Google account identity.
      • First-launch CDN load — the PDF reader, OCR engine, and DOCX parser load from cdn.jsdelivr.net. After first load they're cached locally and re-used offline. The CDN sees your IP address on first load only.
      • Tesseract language data — Gujarati and Devanagari recognition data is downloaded from tessdata.projectnaptha.com on first OCR. Cached afterward.

      No cookies are set by this app. No analytics, no telemetry, no tracking pixels. We don't operate a server — the app runs entirely in your browser. If you host it elsewhere (e.g. GitHub Pages), that host's standard logs apply.

      🇪🇺 GDPR / 🇮🇳 DPDP rights

      Because this app collects no personal data on any server, the standard data-subject rights (access, rectification, erasure, portability, objection) are exercised by you directly:

      • Access / Portability: use the Copy Text button in the reader to export all extracted content.
      • Erasure: click Clear All to wipe in-memory state. To wipe stored preferences and saved corrections (language/text overrides, manual verse references, sandhi splits, translation cache, Drive credentials), use your browser's "Clear site data" (DevTools → Application → Storage in Chrome; Settings → Site Settings on Android).
      • Withdraw consent: click Reset privacy choices below to revoke previous consent and re-show the privacy banner.

      ⚠️ Disclaimers

      • OCR accuracy varies with scan quality. Always verify against the original.
      • Inline script-switching is a known OCR limitation: short Sanskrit/Devanagari phrases embedded mid-sentence within Gujarati or Hindi prose are sometimes misread as the surrounding script (e.g. भगवान् काव्यः may appear as માવાનું વાસ્થ્યઃ). This affects all OCR engines — Google, Azure, Tesseract — because text-region detection assigns one dominant script per visual line. Standalone Sanskrit verse lines and standalone Gujarati paragraphs OCR correctly. For inline mixed-script passages, verify against the original page image (Document View).
      • Sandhi-viccheda (compound word splitting) uses the University of Hyderabad Heritage morphological segmenter when online, with an on-device rule-based assistant as fallback. The segmenter gives the ranked best analysis but Sanskrit segmentation has inherent ambiguity — treat results as a study aid and use the built-in split/merge editing to verify. The fallback assistant is heuristic only.
      • Chhand (metre) detection: when a verse matches a known guru/laghu pattern, the exact metre name is shown (rigorous scansion); otherwise it falls back to an approximate syllable-count family. The 📐 Scan view marks each syllable heavy/light by fixed rules. All computed on-device; accuracy depends on clean text (OCR errors shift the marks). Treat it as a study aid.
      • Re-OCR a region and manual language/text overrides let you correct mistakes, but corrections are only as accurate as what you enter or accept. Verify against the original page image.
      • Auto-translation is statistical, not a Sanskrit dictionary. For canonical meanings consult Monier-Williams, Apte, or a trusted teacher.
      • Narration depends on the Text-to-Speech voices installed on your device. Quality varies by platform.

      📬 Contact

      Questions, takedown requests, or feedback about copyrighted content: open an issue on the project repository.

      ★ Bookmarks

      Browse