Sanskrit · c. 600–1200
Read the Sanskrit that travelled east
Siddhaṃ is the hand in which Sanskrit reached China and Japan, and almost nobody who reads Devanagari can read it cold. Its conjunct consonants fuse two or three letters into a single shape, its vowels hang above and below the line, and on a palm leaf the string hole punches a gap through the middle of every line. Evidano transcribes the akṣaras as they stand and lets you check each line against the leaf.
Sample pages
Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Hōryū-ji, Nara, manuscript c. 6th–8th century; facsimile 1880 · Source · Public domain · Facsimile by Autotype, London, for Buddhist Texts from Japan (Anecdota Oxoniensia)
Transcription
The upper leaf, which holds the whole of the shorter Prajñāpāramitāhṛdayasūtra as edited from these very leaves by F. Max Müller and Bunyiu Nanjio in 1884. The standard Sanskrit text is given here in Devanagari from Wikisource, with the source’s line breaks; the leaf itself opens with namaḥ sarvajñāya, as this text does, but its readings are corrupt in places, dropping anusvāras and visargas and omitting a phrase, so the two should be compared rather than assumed identical.
॥ अथ प्रज्ञापारमिताहृदयसूत्रम् ॥ ॥ नमः सर्वज्ञाय ॥ आर्यावलोकितेश्वरो बोधिसत्त्वो गंभीरायां प्रज्ञापारमितायां चर्यां चरमाणो व्यवलोकयति स्म । पंचस्कन्धाः । तांश्च स्वभावशून्यान्पश्यति स्म । इह शारिपुत्र रूपं शून्यता शून्यतैव रूपं रूपान्न पृथक्शून्यता शून्यताया न पृथग्रूपं यद्रूपं सा शूयता या शून्यता तद्रूपं । एवमेव वेदनासंज्ञासंस्कारविज्ञानानि । इह शारिपुत्र सर्वधर्माः शून्यतालक्षणा अनुत्पन्ना अनिरुद्धा अमला न विमला नोना न परिपूर्णाः । तस्माच्छारिपुत्र शून्यतायां न रूपं न वेदना न संज्ञा न संस्कारा न विज्ञानानि । न चक्षुःश्रोत्रघ्राणजिह्वाकायमनांसी । न रूपशब्दगंधरसस्प्रष्टव्यधर्माः । न चक्षुर्धातुर्यावन्न मनोविज्ञानधातुः । न विद्या नाविद्या न विद्याक्षयो नाविद्याक्षयो यावन्न जरामरणं न जरामरणक्षयो न दुःखसमुदयनिरोधमार्गा न ज्ञानं न प्राप्तिः ॥ तस्मादप्राप्तित्वाद्बोधिसत्त्वाणां प्रज्ञापारमितामाश्रित्य विहरत्यचित्तावरणः । चित्तावरणनास्तित्वादत्रस्तो विपार्यासातिक्रान्तो निष्ठनिर्वाणः ॥ त्र्यध्वव्यवस्थिताः सर्वबुद्धाः प्रज्ञापारमितामाश्रित्यानुत्तरां सम्यक्सम्बोधिमभिसंबुद्धाः ॥ तस्माज्ज्ञातव्यं प्रज्ञापारमिता महामन्त्रो महाविद्यामन्त्रो ऽनुत्तरमन्त्रो ऽसमसममन्त्रः सर्वदुःखप्रशमनः । सत्यममिथ्यत्वात् । प्रज्ञपारमितायामुक्तो मन्त्रः । तद्यथा गते गते पारगते पारसंगते बोधि स्वाहा ॥ इति प्रज्ञापारमिताहृदयं समाप्तम् ॥
Transcription from हृदयसूत्रम् (Prajñāpāramitāhṛdayasūtra), via Sanskrit Wikisource (CC BY-SA 4.0).

Bibliothèque nationale de France, c. 8th–10th century · Source · Public domain · Bibliothèque nationale de France
Transcription being verified
The sheet carries the same shorter recension edited by Müller and Nanjio, The Ancient Palm-Leaves (Anecdota Oxoniensia, 1884); a line-by-line reading of this copy is being matched against it and against the Gallica scan.
What the model is told to watch for: Expect Siddham with a horizontal head line, script-specific conjunct consonants, dependent vowel signs, and akṣaras whose components can extend above or below the line. Distinguish independent from dependent vowels, virama or vowel-cancellation forms, conjunct clusters, anusvāra, visarga, jihvāmūlīya, and upadhmānīya. Note daṇḍa (।), double daṇḍa (॥), siddhaṃ signs, section marks, invocation symbols, and visual spacing that does not necessarily correspond to lexical boundaries after sandhi.

Bibliothèque nationale de France, Pelliot chinois 2778, c. 800–1000 · Source · Public domain · Bibliothèque nationale de France
Transcription being verified
The Chinese transliteration column corresponds to the Nīlakaṇṭha dhāraṇī as transmitted in the Chinese canon; the Siddhaṃ line is being read against the Gallica scan before publication.
What the model is told to watch for: Expect Siddham with a horizontal head line, script-specific conjunct consonants, dependent vowel signs, and akṣaras whose components can extend above or below the line. Distinguish independent from dependent vowels, virama or vowel-cancellation forms, conjunct clusters, anusvāra, visarga, jihvāmūlīya, and upadhmānīya. Note daṇḍa (।), double daṇḍa (॥), siddhaṃ signs, section marks, invocation symbols, and visual spacing that does not necessarily correspond to lexical boundaries after sandhi.

British Library, Or.8212/175, c. 8th–10th century · Source · Public domain · International Dunhuang Project
Transcription being verified
Catalogued by the International Dunhuang Project as a bilingual Sanskrit and Sogdian scroll of the Nīlakaṇṭha dhāraṇī; the Sanskrit lines shown here are being read against the IDP images.
What the model is told to watch for: Expect Siddham with a horizontal head line, script-specific conjunct consonants, dependent vowel signs, and akṣaras whose components can extend above or below the line. Distinguish independent from dependent vowels, virama or vowel-cancellation forms, conjunct clusters, anusvāra, visarga, jihvāmūlīya, and upadhmānīya. Note daṇḍa (।), double daṇḍa (॥), siddhaṃ signs, section marks, invocation symbols, and visual spacing that does not necessarily correspond to lexical boundaries after sandhi.

Source · Public domain · Photograph by WikiTaro, released to the public domain
Transcription
The opening of the shorter Sukhāvatīvyūha (Amitābha Sūtra), section 1, matching the five columns shown: the invocation, then the sutra from evaṃ mayā śrutam to the assembly of monks. The standard Sanskrit text is quoted in Devanagari from Wikisource; the print itself sets these words in Siddhaṃ with katakana readings, so the two run in parallel rather than glyph for glyph.
॥ नमः सर्वज्ञाय ॥ एवं मया श्रुतम् । एकस्मिन् समये भगवाञ्श्रावस्त्यां विहरति स्म जेतवनेऽनाथपिंडदस्यारामे महता भिक्षुसंघेन सार्धम् अर्धत्रयोदशभिर् भिक्षुशतैरभिज्ञानाभिज्ञातैः स्थविरैर्महाश्रावकैः सर्वैरर्हद्भिः ।
Transcription from सुखावतीव्यूहः (shorter Sukhāvatīvyūha), section 1, via Sanskrit Wikisource (CC BY-SA 4.0).
What a Siddhaṃ page looks like
Siddhaṃ, or siddhamātṛkā, is the north Indian book hand of roughly the sixth to the twelfth century, the stage between late Gupta writing and early Nāgarī. In India it kept evolving; carried east along the Silk Road with Buddhist scriptures it did not. Chinese and Japanese monks learned it as the proper script for mantras and dhāraṇīs, held its letterforms fixed for a thousand years, and are still taught it in the Shingon and Tendai schools as bonji. The name comes from the word siddhaṃ, "accomplished", written at the head of the syllabary a pupil copied out first.
The unit of writing is the akṣara, a consonant with its inherent a, not a letter. Change the vowel and you attach a sign that may sit on top (e), underneath (u), to the left (i) or to the right (ā), or wrap round the whole shape (o, au); write a consonant with no vowel at all and it either takes a virāma stroke or is stacked into the next one. Those stacks, the saṃyuktākṣara, are what make the script hard: kta, ṣṭra, ntva and their kin fuse two or three consonants into one compound sign whose parts are squeezed, halved or turned on their side. Add the anusvāra dot for a nasal, the two dots of visarga, and the ardhavisarga forms jihvāmūlīya and upadhmānīya that appear before k and p in careful Buddhist copies.
Layout depends on where the copy was made. An Indian or Central Asian pothi is a wide, shallow palm leaf written along its length, with a blank disc left for the binding string and folio numbers in the left margin; there is no word division, and sandhi welds the last sound of one word to the first of the next, so the gaps that do appear are not lexical. A Chinese or Japanese copy turns the same script through ninety degrees into vertical columns, often with a phonetic gloss beside every syllable. Section ends are marked by a daṇḍa or double daṇḍa, and each text opens with an auspicious mark frequently misread as oṃ.
Why ordinary OCR struggles here
Stacked consonants read as one letter
A generic recogniser is trained to find one character per glyph. In Siddhaṃ a single glyph can be three consonants plus a vowel sign, and the components are deformed: a subscript ra becomes a hook, a preceding r rides on the following letter as a hat, and ta under ka loses its crossbar. Getting kṣ, ṣṭ, ndh and hm out of those shapes needs a model that knows which clusters Sanskrit allows.
Vowels on four sides of the same sign
The same stroke means different things depending on which side of the akṣara it sits, and an i sign written to the left belongs to the consonant that follows it, not the one before. Ordinary line-based OCR flattens the whole column of marks into one reading order and produces syllables no Sanskrit word contains.
Scribes who did not know the language
Many East Asian copies were made by people who could draw the letters but not read them. Anusvāras and visargas drop out, whole phrases repeat, and dhāraṇī syllables mutate. The Hōryūji leaves are a famous example. A transcription has to record what is on the leaf rather than quietly restore the standard text, and mark where the two diverge.
Two scripts, two directions, one page
Dunhuang sheets alternate lines of Siddhaṃ with Sogdian; Chinese scrolls set the Sanskrit in a grid beside a character-by-character phonetic gloss; Japanese blockbooks add katakana readings and a Chinese rendering in the same column. Reading order has to be resolved before a single syllable is transcribed.
Who works with this material
Editors of Buddhist Sanskrit texts
Anyone collating a sutra or a dhāraṇī across Indian, Central Asian and East Asian witnesses needs each copy recorded as written, with its omissions and false readings intact, before the variants can be weighed. A diplomatic first pass makes the apparatus buildable.
Silk Road manuscript projects
Collections of Dunhuang, Turfan and Gilgit fragments hold thousands of unedited scraps in which a handful of legible akṣaras is enough to identify the text. Searchable transliteration turns a box of fragments into something that can be matched against a corpus.
Temple archives and bonji specialists
Japanese monasteries hold sutra copies, mandala inscriptions and calligraphic exemplars in Siddhaṃ that catalogue records describe only by title. Transcribing the syllables lets a cataloguer say which dhāraṇī a scroll actually carries.
Historians of script and encoding
Palaeographers tracing the road from Gupta to Nāgarī, and font and Unicode implementers testing conjunct behaviour, both need transcriptions tied to particular shapes on particular folios rather than to a normalised modern text.
Getting the best transcription
Choose the target alphabet before you start
Say whether you want IAST romanisation, Devanagari or Siddhaṃ in Unicode. IAST is easiest to check and to search; Devanagari suits comparison with printed editions; Siddhaṃ keeps the encoding closest to the page. Mixing them across a manuscript makes collation painful later.
Set a policy for anusvāra, visarga and sandhi
Decide whether a missing anusvāra should be supplied in brackets or left out, and whether the transcription should keep the manuscript’s continuous sandhi or split words. For dhāraṇīs, ask for syllable-by-syllable output: the strings are not words and any attempt to normalise them destroys evidence.
Describe the page furniture
Tell the model that the blank disc in a palm leaf is a string hole and not a gap in the text, that the number at the left edge is the folio number, and that a daṇḍa or double daṇḍa ends a section. On an East Asian copy, say whether the Chinese or katakana gloss should be transcribed alongside or ignored.
Test on a damaged folio, not a clean one
Start with a leaf where the ink has flaked and only the outline of the akṣaras survives. If uncertain syllables come back marked rather than invented there, the intact folios will look after themselves.
Further reading on this hand
The palaeography guides the in-app selector points to for this domain, if you want to check a transcription against the standard references.
- Proposal to Encode the Siddham Script in ISO/IEC 10646
This Unicode proposal was prepared for the International Organization for Standardization and International Electrotechnical Commission (ISO/IEC) encoding process. It documents Siddham's historical use for Sanskrit Buddhist texts and provides a character repertoire with representative manuscript forms. The proposal explains vowels, consonants, combining signs, virama behavior, conjuncts, punctuation, and script-specific symbols. It is useful for Siddham transcription because it links palaeographical shapes to distinct encodable characters and combining sequences.
Frequently asked questions
- Can I get the output in IAST, Devanagari or Siddhaṃ Unicode?
- All three. Ask for one in the prompt and it is applied consistently across the manuscript. IAST is the usual choice for editing and searching, Devanagari for checking against printed editions, and Siddhaṃ Unicode when the encoding itself is the point of the project.
- How are conjunct consonants handled?
- A stack such as kṣa, ṣṭra or ndha is read as the cluster it represents rather than as its top component. Where a subscript has flaked away or the stack is ambiguous, the reading is flagged instead of being guessed from context.
- What happens to the string hole in a palm leaf?
- It is treated as a hole, not as punctuation or a word break. The text either side of it belongs to the same line and is joined, and the gap can be noted in the output if you want a record of where it falls.
- Will the transcription correct a corrupt dhāraṇī?
- Not silently. Many East Asian copies were made by scribes with no Sanskrit and they drop nasals, repeat phrases and garble syllables. The default is to record what is on the page; if you want the standard reading supplied as well, ask for it in brackets so both survive.
- Can it read the Chinese or katakana gloss beside the Sanskrit?
- Yes. On a Dunhuang grid or a Japanese blockbook the columns can be transcribed together, keeping each syllable beside its phonetic gloss, or the Sanskrit alone can be extracted if the gloss is not wanted.
Related scripts and pages
Transcribe your Siddhaṃ folios
Upload a leaf or a scroll, choose Siddham Buddhist Manuscripts as the domain, and check the first page before you run the rest.
