Ottoman Turkish · c. 1500–1800
From a divan folio to a searchable text
The classical Ottoman book is written in two hands that behave quite differently: an upright naskh that keeps to its line, and a nastaʿliq whose words slide downhill and stack on top of one another. Evidano transcribes both, keeps the couplet structure of a divan and the red rubrics of a chronicle, and lets you compare each line with the folio before you accept it.
Sample pages
Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Free Library of Philadelphia, Rare Book Department, Lewis O 92, 1511 · Source · Public domain · Digitised by the University of Pennsylvania Libraries for OPenn
Transcription being verified
The manuscript is catalogued in full on OPenn (Lewis O 92); no published transcription of the deed has been traced.
What the model is told to watch for: Expect Naskh for prose and religious works and Nastaʿliq for poetry, correspondence, or Persianate literary material, with Arabic and Persian passages embedded in Ottoman Turkish. In Nastaʿliq, follow the descending diagonal baseline, compact word groups, and numerous discretionary ligatures; in Naskh, use more regular positional forms but expect inconsistent dotting. Note izafet constructions, Arabic abbreviation formulae, vowel signs used selectively, catchwords, interlinear glosses, and marginal corrections.

Free Library of Philadelphia, Rare Book Department, Lewis O 157, c. 1700 · Source · Public domain · Digitised by the University of Pennsylvania Libraries for OPenn
Transcription being verified
Bâkî’s divan is edited in modern Turkish scholarship; the gazels on this particular folio have not been identified against a published text.
What the model is told to watch for: Expect Naskh for prose and religious works and Nastaʿliq for poetry, correspondence, or Persianate literary material, with Arabic and Persian passages embedded in Ottoman Turkish. In Nastaʿliq, follow the descending diagonal baseline, compact word groups, and numerous discretionary ligatures; in Naskh, use more regular positional forms but expect inconsistent dotting. Note izafet constructions, Arabic abbreviation formulae, vowel signs used selectively, catchwords, interlinear glosses, and marginal corrections.

Free Library of Philadelphia, Rare Book Department, Lewis O 93, 1722 · Source · Public domain · Digitised by the University of Pennsylvania Libraries for OPenn
Transcription being verified
Selânikî’s chronicle survives in several manuscripts; this copy has not been collated against a published text.
What the model is told to watch for: Expect Naskh for prose and religious works and Nastaʿliq for poetry, correspondence, or Persianate literary material, with Arabic and Persian passages embedded in Ottoman Turkish. In Nastaʿliq, follow the descending diagonal baseline, compact word groups, and numerous discretionary ligatures; in Naskh, use more regular positional forms but expect inconsistent dotting. Note izafet constructions, Arabic abbreviation formulae, vowel signs used selectively, catchwords, interlinear glosses, and marginal corrections.

Walters Art Museum, Baltimore, W.591, 1512 · Source · Public domain
Transcription being verified
The Walters catalogue describes the notes and identifies the first signature; the text itself is unedited.
What the model is told to watch for: Expect Naskh for prose and religious works and Nastaʿliq for poetry, correspondence, or Persianate literary material, with Arabic and Persian passages embedded in Ottoman Turkish. In Nastaʿliq, follow the descending diagonal baseline, compact word groups, and numerous discretionary ligatures; in Naskh, use more regular positional forms but expect inconsistent dotting. Note izafet constructions, Arabic abbreviation formulae, vowel signs used selectively, catchwords, interlinear glosses, and marginal corrections.
The two hands of the Ottoman book
From the sixteenth century Ottoman scribes divided their work between scripts. Prose — histories, legal opinions, endowment deeds, religious manuals — is normally copied in naskh, an upright, evenly spaced hand with most of the dots present and headings picked out in red. Poetry and Persianate literary material, and much private correspondence, go into nastaʿliq: the baseline of each word descends from right to left, letters are joined in long sweeping groups, and a word can be written above the tail of the one before it. In verse the two hemistichs of a couplet sit in ruled columns, and the rhyme word is often lifted into the margin when the line runs out of room.
Ottoman Turkish of these centuries is Turkic in grammar and largely Arabic and Persian in vocabulary. The Persian izafet chains words together with a vowel that is usually not written at all, so bir nüsha-i şerife looks like three separate words on the page. Arabic supplies fixed formulae — the basmala, salutations after the Prophet’s name, sallallahu aleyhi ve sellem contracted to a few letters, the ﷽ of an opening — and Arabic plurals are used for Turkish nouns. Vowel signs appear only where a word might be mistaken, and shadda, hamza and the Arabic case endings turn up in quoted passages and then vanish.
The apparatus around the text matters as much as the text. Catchwords in the lower left corner of the verso chain the quires. Glosses run between the lines and in the margins, sometimes at a slant, sometimes in a later and quite different hand. Corrections are marked with a small ṣaḥḥ, omitted words are written in the margin with a caret-like sign, and endowment or ownership statements and seal impressions are stamped across the first and last pages. A page can therefore carry four or five separate texts, and separating them is part of the transcription rather than a step after it.
Why ordinary OCR struggles here
A baseline that will not stay still
Nastaʿliq compresses a word group into a diagonal that ends far below where it began, then starts the next group high again. Layout analysis built for horizontal type slices such a line in the wrong place, splitting words between two output lines or fusing a rhyme word in the margin onto the wrong hemistich.
Ligatures that hide letters
Discretionary ligatures such as lam-mim-ha or the tight sin-ha of a scribal hand fuse three or four letters into one continuous shape whose individual teeth are no longer drawn. Reading them depends on recognising the whole word, which is why a system that thinks in characters fails on precisely the commonest words.
Selective vowelling and drifting dots
A copyist vowels a Qurʾanic quotation carefully and leaves the surrounding Turkish bare, or drops a dot under a ye at the end of a word because nothing else could go there. The transcription has to reproduce what is on the page rather than a regularised version of it.
Marginalia, seals and later hands
Endowment stamps, reader’s notes, price marks and eighteenth-century collation notes all sit on the same folios as the sixteenth-century text. Unless the pass distinguishes them, a manuscript’s search index quietly mixes a scribe’s words with a later owner’s.
Who works with this material
Literary scholars editing divans
Producing a text of an Ottoman poet means collating several manuscripts couplet by couplet. What helps is a transcription that keeps each hemistich in its column, records the rhyme word even when it has migrated into the margin, and does not silently repair a defective metre.
Historians reading chronicles and şeriyye records
Court registers, chronicles and biographical dictionaries are consulted for names, dates, prices and offices. A first pass that captures the red rubrics and the marginal additions is enough to build a searchable index long before a full critical text exists.
Islamic manuscript cataloguers
Cataloguers need incipit, explicit, colophon, copyist, date and the endowment or ownership notes on the flyleaves. These are short, formulaic and often in Arabic, and getting them out of a codex reliably is what turns a shelf list into a catalogue.
Conservation and provenance research
Seal impressions, waqf statements and price notes trace a book from a library in Istanbul to a European sale room. Transcribing them, with a note of where on the page each sits, gives provenance work something to search.
Getting the best transcription
Name the script and the layout
Tell the model whether the folio is naskh prose in a single ruled frame or nastaʿliq verse in two columns, and how many columns and lines to expect. Layout is the largest single source of error on Ottoman book pages, and stating it removes most of it.
Decide what to do with izafet and abbreviations
Say whether contractions such as the honorific after a name should be expanded, marked or left as written, and whether an izafet vowel should be supplied in transliteration. Fixing this once keeps a whole manuscript internally consistent.
Ask for the margins as a separate layer
Glosses, corrections and later notes should come back labelled by position — right margin, upper margin, interlinear — rather than folded into the main text. That also makes it obvious when a gloss belongs to a different century than the text it comments on.
Start with the colophon
The last page usually gives the copyist, the place and the date in a formulaic sentence. It is short, checkable and full of proper names, so it tells you quickly whether the settings are right before you commit to three hundred folios.
Further reading on this hand
The palaeography guides the in-app selector points to for this domain, if you want to check a transcription against the standard references.
- Scripts
This University of Pennsylvania guide introduces Naskh, Nastaʿliq, and other scripts used across Islamic manuscript cultures. It describes Naskh as a common book hand and Nastaʿliq as a strongly sloping, highly ligatured script associated particularly with Persianate production. The guide compares stroke angle, proportion, joining behavior, and typical textual functions. These distinctions are useful for Ottoman manuscripts in which Turkish, Arabic, and Persian content may be assigned different scripts or levels of formality.
Frequently asked questions
- Does nastaʿliq need different handling from naskh?
- Yes. Naskh sits on a line and can be segmented conventionally; nastaʿliq has to be read as descending word groups, and its line detection has to allow one word to overlap the tail of the next. Saying which script is on the page is the single most useful instruction you can give.
- Will couplets stay in the right order?
- In a two-column verse layout the reading order is across the line, not down each column, and that has to be stated. The output can keep one couplet per line so that a later collation against a printed edition lines up.
- How are Persian and Arabic passages handled inside a Turkish text?
- They are transcribed in place and can be labelled by language, which matters because a Persian couplet quoted in a Turkish chronicle would otherwise pollute a Turkish word list. Vowel signs present in the quotation are kept.
- Can catchwords be excluded from the running text?
- They can be captured separately or dropped. Keeping them is useful when you are checking quire structure or looking for a missing leaf, but they should never be run on to the last line of the page.
- What about a folio where the ink has bled through from the other side?
- Show-through is a common problem on thin Ottoman paper. Reversed ghost letters can be recognised as coming from the verso and left out, and a note recorded where the show-through makes a real letter uncertain.
Related scripts and pages
Transcribe your Ottoman codices
Upload a folio, choose Classical Ottoman Naskh and Nastaʿliq Manuscripts, and settle the layout on one page before running the volume.
