Persian · c. 900–1200
Read Persian as it was first written down
The first Persian books were copied in an alphabet built for Arabic, and the scribes did not always add the marks that keep the two languages apart: p, ch, zh and g wear the coats of b, j, z and k, and no short vowel is written at all. Evidano reads these folios line by line, keeps an Arabic base text separate from the Persian glossing it, and puts every reading beside the page so you can check it.
Sample pages
Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Library of Congress, African and Middle Eastern Division, 1-90-154.150, 11th–13th century · Source · Public domain · Library of Congress, African and Middle East Division, Near East Section Persian Manuscript Collection
Transcription being verified
The Library of Congress record identifies the text — the heading of sūra 19 and its basmala on this side, the closing words of sūra 18:110 on the recto — but does not print the sub-linear Persian rendering, which is unedited.
What the model is told to watch for: Expect early Persian written in Qurʾanic-style Kufic or increasingly regular Naskh, often with Arabic passages and inconsistent marking of Persian sounds. Distinguish پ, چ, ژ, and گ where scribes adapt or modify Arabic letterforms, and watch for missing or displaced dots in ب ت ث ن ي and ج ح خ groups. Note vowel signs used selectively, Arabic ligatures, interlinear Persian glosses, catchwords, section markers, and scripts changing with language or textual function.

Arthur M. Sackler Gallery, Smithsonian Institution, S1997.95, 11th–12th century · Source · Public domain
Transcription
The nine lines of the folio, which begin in the middle of sūra 8:45 and break off in the middle of 8:48; the four verses are printed here whole, in the modern imlāʾī spelling, so the opening of 45 and the remainder of 48 are supplied and are not on the leaf. The manuscript carries only occasional vowels.
يَا أَيُّهَا الَّذِينَ آمَنُوا إِذَا لَقِيتُمْ فِئَةً فَاثْبُتُوا وَاذْكُرُوا اللَّهَ كَثِيرًا لَعَلَّكُمْ تُفْلِحُونَ وَأَطِيعُوا اللَّهَ وَرَسُولَهُ وَلَا تَنَازَعُوا فَتَفْشَلُوا وَتَذْهَبَ رِيحُكُمْ وَاصْبِرُوا إِنَّ اللَّهَ مَعَ الصَّابِرِينَ وَلَا تَكُونُوا كَالَّذِينَ خَرَجُوا مِنْ دِيَارِهِمْ بَطَرًا وَرِئَاءَ النَّاسِ وَيَصُدُّونَ عَنْ سَبِيلِ اللَّهِ وَاللَّهُ بِمَا يَعْمَلُونَ مُحِيطٌ وَإِذْ زَيَّنَ لَهُمُ الشَّيْطَانُ أَعْمَالَهُمْ وَقَالَ لَا غَالِبَ لَكُمُ الْيَوْمَ مِنَ النَّاسِ وَإِنِّي جَارٌ لَكُمْ فَلَمَّا تَرَاءَتِ الْفِئَتَانِ نَكَصَ عَلَى عَقِبَيْهِ وَقَالَ إِنِّي بَرِيءٌ مِنْكُمْ إِنِّي أَرَى مَا لَا تَرَوْنَ إِنِّي أَخَافُ اللَّهَ وَاللَّهُ شَدِيدُ الْعِقَابِ
Transcription from القرآن الكريم (بالرسم الإملائي)/سورة الأنفال, Arabic Wikisource (Public domain).

Library of Congress, African and Middle Eastern Division, 1-89-154.172, 14th century · Source · Public domain · Library of Congress, African and Middle East Division, Near East Section Persian Manuscript Collection
Transcription
The Arabic base text only, in the modern imlāʾī spelling: sūra 3:85–88, the range recorded for this side of the leaf. The first line begins in the middle of verse 85 and the last breaks off in verse 88, so the ends of both verses are supplied from the printed text. The Persian glosses under the words are not included; they have not been edited.
وَمَنْ يَبْتَغِ غَيْرَ الْإِسْلَامِ دِينًا فَلَنْ يُقْبَلَ مِنْهُ وَهُوَ فِي الْآخِرَةِ مِنَ الْخَاسِرِينَ كَيْفَ يَهْدِي اللَّهُ قَوْمًا كَفَرُوا بَعْدَ إِيمَانِهِمْ وَشَهِدُوا أَنَّ الرَّسُولَ حَقٌّ وَجَاءَهُمُ الْبَيِّنَاتُ وَاللَّهُ لَا يَهْدِي الْقَوْمَ الظَّالِمِينَ أُولَئِكَ جَزَاؤُهُمْ أَنَّ عَلَيْهِمْ لَعْنَةَ اللَّهِ وَالْمَلَائِكَةِ وَالنَّاسِ أَجْمَعِينَ خَالِدِينَ فِيهَا لَا يُخَفَّفُ عَنْهُمُ الْعَذَابُ وَلَا هُمْ يُنْظَرُونَ
Transcription from القرآن الكريم (بالرسم الإملائي)/سورة آل عمران, Arabic Wikisource (Public domain).

Library of Congress, African and Middle Eastern Division, 1-85-154.69, 13th century · Source · Public domain · Library of Congress, African and Middle East Division, Near East Section Persian Manuscript Collection
Transcription being verified
Balʿamī’s Persian text of this preface was translated into French by Hermann Zotenberg in his Chronique de Tabari (Paris, 1867–74), which is out of copyright, but no public-domain edition prints the Persian of this particular copy.
What the model is told to watch for: Expect early Persian written in Qurʾanic-style Kufic or increasingly regular Naskh, often with Arabic passages and inconsistent marking of Persian sounds. Distinguish پ, چ, ژ, and گ where scribes adapt or modify Arabic letterforms, and watch for missing or displaced dots in ب ت ث ن ي and ج ح خ groups. Note vowel signs used selectively, Arabic ligatures, interlinear Persian glosses, catchwords, section markers, and scripts changing with language or textual function.
What an early Persian page looks like
New Persian appears as a written language in Khurasan and Transoxiana in the ninth century, set down in Arabic letters on rag paper. Almost nothing survives in a copy made before c. 1200 — the oldest dated example is a pharmacological handbook written out in 1055 by the poet Asadī Ṭūsī and now in Vienna — so the period is studied from Qurʾans copied in Iran, from translations glossed between the lines of an Arabic text, and from thirteenth- and fourteenth-century copies of tenth- and eleventh-century works. Two book scripts carry them: eastern Kufic, the angular New Style, with tall leaning shafts and bowls swept far below the line, knotted with interlace in its plaited form and kept for gold headings long after the body text had moved on; and naskh, rounded and evenly spaced, which becomes the ordinary hand of Persian prose in the eleventh century and never leaves it.
The alphabet was a poor fit and the early scribes patched it unevenly. Persian needed four letters Arabic did not have, and پ, چ, ژ and گ are routinely written as plain ب, ج, ز and ک, so that گفت and کفت, پیش and بیش are the same shape on the page. Early orthography also keeps distinctions modern Persian has dropped: a ذ where later copies write د, and prefixes and suffixes such as بی، می، ها and تر set apart from their word or joined to it by no fixed rule. Short vowels are not written and neither is the eżāfe that links a noun to its adjective; the vowel signs that do appear were often added by a later reader in red, with hamzas, shaddas and corrections that belong to a different century from the text.
Layout is doing work as well. In an interlinear Qurʾan the Arabic runs large across the page and the Persian is squeezed under each word on a slant, so the eye moves horizontally through one text and diagonally through the other; in other copies the Persian is continuous and picked out in red. Gold rosettes divide verses, lamp-shaped medallions in the margin mark the thirtieth parts, marginal notes tell a reciter how to lighten a consonant, and a catchword at the foot of the verso gives the first word of the next leaf.
Why ordinary OCR struggles here
One skeleton, five or six letters
Strip the dots and ب، ت، ث، پ، ن and ی collapse into one shape, as do ج، چ، ح and خ. Early scribes place dots loosely, group them into a slanting dash, or leave them off where the word is obvious to them. An engine trained on printed Persian, which is always fully dotted, settles these by guessing the commonest word rather than by weighing the sentence.
The letters Persian added and the scribes omitted
Because پ، چ، ژ and گ are usually written as ب، ج، ز and ک, deciding between گور and کور, or پرده and برده, is a question about the sentence rather than about the ink. The transcription has to record what stands on the page and, if you want it, the modernised form beside it, so that a lexicographer can see which is which.
Two languages interleaved on the same lines
An interlinear Qurʾan is not one text but two, in different scripts, at different sizes and on different angles. Read straight through and the Arabic and the Persian fuse into nonsense; the reading order has to be rebuilt word by word, with base text and gloss kept in separate streams.
Spellings that no modern dictionary lists
Archaic forms, the ذ for د, detached prefixes and vocabulary that fell out of use by the fifteenth century all look like errors to a system tuned on modern Persian, which repairs them silently. Vowel points and corrections added centuries later sit in the same space and are easily taken for the scribe’s own.
Who works with this material
Editors of early Persian prose
Anyone establishing the text of a tenth- or eleventh-century work from later copies needs a reading that preserves each copy’s own spelling rather than a tidied modern one. Variants become visible only when the ذ, the detached prefix and the missing gāf stroke stand as the scribe left them.
Scholars of Qurʾan translation
The interlinear versions are the earliest continuous evidence for Persian prose, and their value lies in the pairing of each Arabic word with its Persian equivalent. What such a project needs is an aligned output — Arabic token, Persian token — not a flat block of text in which the two have been merged.
Cataloguers of detached folios
Museums and libraries hold thousands of single leaves cut from dispersed manuscripts. Identifying one means reading enough of it to locate the passage. A transcription of the visible lines turns an unidentified leaf into a searchable record.
Historical lexicography and corpus building
Dictionaries of early Persian are built from dated attestations. A corpus assembled from manuscript pages is only usable if archaic forms survive the transcription and if uncertain readings are flagged instead of quietly regularised.
Getting the best transcription
Say which language you want back
On a bilingual page, ask for the Arabic and the Persian as two labelled streams, and say whether the glosses should be aligned to the word above them or listed line by line. Settling this before the run saves untangling the output afterwards.
Fix a rule for the four Persian letters
Choose whether ب، ج، ز and ک are transcribed as written or normalised to پ، چ، ژ and گ where the word requires it. If you want both, ask for the manuscript form first and the modern form in brackets; set once, the rule holds across a whole codex.
Protect the old orthography
State that the ذ, separated prefixes and unfamiliar spellings must be kept, and that anything the model would ordinarily correct is to be marked rather than changed. Ask for later additions in red — vowels, hamzas, corrections — to be reported apart from the black text.
Test on a page with a heading
Start with a folio that mixes gold Kufic in a heading, black naskh in the body and a marginal note. If the three are kept apart there, the plain pages of the same codex will present nothing new.
Further reading on this hand
The palaeography guides the in-app selector points to for this domain, if you want to check a transcription against the standard references.
- Scripts
This University of Pennsylvania guide introduces major scripts used in Islamic manuscripts, including Kufic and Naskh. It contrasts angular early writing with the proportioned cursive book hands that became dominant in later codices. The guide discusses stroke form, joining, ligatures, script function, and relationships among regional manuscript traditions. It is useful for early New Persian material because Persian scribes adapted these Arabic-script systems while adding letters for Persian sounds.
Frequently asked questions
- Can it tell the Persian from the Arabic when both are on the page?
- Yes. The two are distinguished by script, size and position rather than by ink colour, and the transcription can label each stretch so that a bilingual folio comes back as an Arabic text with its Persian rendering attached, not as one continuous line.
- Does it add the dots for پ, چ, ژ and گ that the scribe left out?
- Only if you ask. The default keeps the letters as written, because for a text-critical edition the undotted form is the evidence. Ask for a normalised version and the modernised letters are supplied, with the manuscript form retained alongside if you want both.
- How are word-by-word interlinear glosses handled?
- They are read as a separate layer and can be returned aligned to the Arabic word each one translates. Because they are written on a slant and often overlap the line below, uncertain alignments are flagged rather than guessed at.
- Will archaic spellings be modernised without warning?
- No. Old forms are treated as readings rather than mistakes and are reproduced. A normalised text for searching is a second output; the first stays close to the page so the two can be compared.
- Is eastern Kufic legible to the model?
- The angular New Style is recognised as a distinct hand, including its plaited form in headings where interlace fills the space between the shafts. Gold-on-red headings and faded ink are the harder cases, and letters that cannot be read are marked instead of invented.
Related scripts and pages
Transcribe your early Persian folios
Upload a leaf, choose Early New Persian Kufic and Naskh Manuscripts as the domain, and settle the glossing rules on one page before you run the codex.
