📚 Educational use only
Indic Reader is a free, non-commercial tool for the educational study of Gujarati, Hindi, and Sanskrit
scriptures and literature. By using this software you agree to use it solely for personal study,
teaching, or non-commercial research. Commercial use, redistribution, or rehosting requires
written permission from the author.
© Copyright
© 2026 The Indic Reader contributors. Software provided
"as is" without warranty of any kind. Source code is released for educational and personal use.
All rights reserved.
Scripture and document content uploaded or pasted by you remains entirely your responsibility.
Ensure you have the right to access, narrate, and translate any text you process. Where source
materials are copyrighted (translations, modern commentaries, scanned editions), respect those rights.
🔒 Data privacy
What stays on your device (always):
- Every file you open — PDFs, images, DOCX, pasted text
- OCR results and extracted text
- OCR cache — recognized text and word positions are cached per page (in this browser's IndexedDB), keyed by file + page + OCR settings, so reopening or re-reading a page is instant and avoids repeat cloud-OCR calls. Bounded in size (oldest entries auto-removed); cleared when you turn off the 🗄 OCR cache toggle.
- Continue / reopen — the last document you opened is cached on-device (IndexedDB) so the "Continue" button can reopen it; your reading position (page, line, word) is remembered to resume where you left off.
- Page bookmarks (including any notes you attach and the full line text), page selections, and current playback position
- Theme preference, voice preference, chant mode, auto-scroll, language filter, and the Prefs toggles (Pause on call, Resume, OCR cache, Check OCR)
- Drive Picker credentials (if you configured them) — stored only in this browser's
localStorage
- Device password manager (optional) — if you tap "Save to device" in Cloud OCR settings, your API key or Worker URL is handed to your browser's Credential Management API and held by your operating system's secure password store (Google Password Manager, iCloud Keychain, or Windows), protected by your device lock. The app and its (non-existent) server never see it again except when you tap "Restore". Availability depends on your browser/OS.
- Translation cache — built up locally as you tap Show Meanings or Translate
- Manual corrections — your per-word and per-line language overrides, line text corrections (from inline line edit or Re-OCR), manual verse references (for vedabase.io links), and saved sandhi splits
What may go online (only with your action):
- Show Meanings sends individual Sanskrit words to two third-party services for accurate dictionary lookup. Disabled by default; requires explicit consent. The Cologne University Monier-Williams API (api.c-salt.uni-koeln.de) is queried first as the authoritative source. Google Translate's public endpoint is used as a fallback and to translate English definitions into Gujarati. Each call exposes your IP address and the word being looked up to these services.
- Cloud OCR (optional, disabled by default; requires explicit consent) sends the entire page image to one of two providers you can configure:
- Google Cloud Vision via your own Cloudflare Worker — the page image traverses Cloudflare's edge (subject to their data-processing terms) and is sent to Google Cloud Vision. Google's privacy policy applies; per Google's documentation, Vision API requests are not used to improve their models without your explicit consent in the Google Cloud Console settings.
- OCR.space direct — the page image is sent to ocr.space servers. Per their privacy policy, free-tier requests may be retained for service improvement.
When Cloud OCR is enabled, files no longer stay only on your device. The processing provider sees your IP address and the page image contents. Recognition results return to your browser only — they are not stored on our side.
- Split Sandhi (optional, disabled until you use it) sends only the Sanskrit verse text — never any file or image — to the University of Hyderabad Heritage segmenter (sanskrit.uohyd.ac.in), a public academic service, routed through your own Cloudflare Worker. The Worker caches results for 24 hours. If the service is unreachable, the app silently falls back to an on-device rule-based assistant; no text leaves your device in that case. The segmenter sees the verse text and (via the Worker) Cloudflare's edge sees the request. The overlay also shows a per-word meaning in your chosen language (Gujarati/Hindi/English) — this sends each individual word to Google Translate, and tapping a word queries the Cologne Monier-Williams dictionary, both subject to the same consent as Show Meanings below. Your chosen meaning-language is stored locally.
- Re-OCR (region, single line, or continuous next page) (optional) sends only an image to your chosen OCR provider (Google Cloud Vision via your Worker, or OCR.space, or processed entirely on-device with local OCR): the cropped rectangle you select, the single line you choose to fix, or — in Continuous mode — the next page being prepared just before narration reaches it. You choose the language hint each time. Same provider data-handling as Cloud OCR above; results are cached on-device.
- Translate verse / Translate page (optional; requires consent) sends text — a single Sanskrit verse, or the current page's Gujarati/Hindi lines — to Google Translate's public endpoint to render it in another language. Sanskrit verses are not sent when you use page translate (Gujarati↔Hindi). Your IP address and the submitted text are visible to Google; no file or image is sent. Results are cached on-device.
- Google Drive Picker/link load — only when you click those buttons. Drive sees the URLs you fetch and (for Picker) your Google account identity.
- First-launch CDN load — the PDF reader, OCR engine, and DOCX parser load from
cdn.jsdelivr.net. After first load they're cached locally and re-used offline. The CDN sees your IP address on first load only.
- Tesseract language data — Gujarati and Devanagari recognition data is downloaded from
tessdata.projectnaptha.com on first OCR. Cached afterward.
No cookies are set by this app. No analytics, no telemetry, no tracking pixels. We don't
operate a server — the app runs entirely in your browser. If you host it elsewhere (e.g. GitHub
Pages), that host's standard logs apply.
🇪🇺 GDPR / 🇮🇳 DPDP rights
Because this app collects no personal data on any server, the standard data-subject rights
(access, rectification, erasure, portability, objection) are exercised by you directly:
- Access / Portability: use the Copy Text button in the reader to export all extracted content.
- Erasure: click Clear All to wipe in-memory state. To wipe stored preferences and saved corrections (language/text overrides, manual verse references, sandhi splits, translation cache, Drive credentials), use your browser's "Clear site data" (DevTools → Application → Storage in Chrome; Settings → Site Settings on Android).
- Withdraw consent: click Reset privacy choices below to revoke previous consent and re-show the privacy banner.
⚠️ Disclaimers
- OCR accuracy varies with scan quality. Always verify against the original.
- Inline script-switching is a known OCR limitation: short Sanskrit/Devanagari phrases embedded mid-sentence within Gujarati or Hindi prose are sometimes misread as the surrounding script (e.g.
भगवान् काव्यः may appear as માવાનું વાસ્થ્યઃ). This affects all OCR engines — Google, Azure, Tesseract — because text-region detection assigns one dominant script per visual line. Standalone Sanskrit verse lines and standalone Gujarati paragraphs OCR correctly. For inline mixed-script passages, verify against the original page image (Document View).
- Sandhi-viccheda (compound word splitting) uses the University of Hyderabad Heritage morphological segmenter when online, with an on-device rule-based assistant as fallback. The segmenter gives the ranked best analysis but Sanskrit segmentation has inherent ambiguity — treat results as a study aid and use the built-in split/merge editing to verify. The fallback assistant is heuristic only.
- Chhand (metre) detection: when a verse matches a known guru/laghu pattern, the exact metre name is shown (rigorous scansion); otherwise it falls back to an approximate syllable-count family. The 📐 Scan view marks each syllable heavy/light by fixed rules. All computed on-device; accuracy depends on clean text (OCR errors shift the marks). Treat it as a study aid.
- Re-OCR a region and manual language/text overrides let you correct mistakes, but corrections are only as accurate as what you enter or accept. Verify against the original page image.
- Auto-translation is statistical, not a Sanskrit dictionary. For canonical meanings consult Monier-Williams, Apte, or a trusted teacher.
- Narration depends on the Text-to-Speech voices installed on your device. Quality varies by platform.
📬 Contact
Questions, takedown requests, or feedback about copyrighted content: open an issue on the project repository.