Ottoman Turkish · c. 1350–1500
Read Turkish as the first Ottoman scribes wrote it
Between the fourteenth century and the reign of Bayezid II, Turkish was written in an alphabet built for Arabic, with vowels only half indicated and dots often left off altogether. Evidano reads these pages letter by letter, keeps the old spellings rather than silently modernising them, and shows every line beside the image so you can check a reading before it goes into an edition.
Sample pages
Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Haus-, Hof- und Staatsarchiv, Vienna, Repertorium XIV B, Nr. 943, 1489 · Source · Public domain
Transcription being verified
Edited with the Ottoman text and a German translation as Urkunde Nr. 15 in F. von Kraelitz-Greifenhorst, Osmanische Urkunden in türkischer Sprache aus der zweiten Hälfte des 15. Jahrhunderts (Vienna, 1922); the printed Ottoman text is being keyed against the plate.
What the model is told to watch for: Expect Ottoman Turkish written with an Arabic-Persian alphabet in Naskh or early Taʿliq, frequently interspersed with Arabic quotations and Persian constructions. Distinguish letters sharing the same skeleton through their sometimes faint or displaced dots, and attend to Persian letters پ, چ, ژ, and گ as well as Ottoman use of ك and ڭ-like forms. Note vowel ambiguity, Arabic grammatical diacritics in quoted passages, ligatures, honorific abbreviations, catchwords, and changes of script or ink marking language shifts.

Haus-, Hof- und Staatsarchiv, Vienna, Repertorium XIV B, Nr. 10, 1496 · Source · Public domain
Transcription being verified
Printed as Urkunde Nr. 23, with the four marginal vizier attestations translated separately, in Kraelitz-Greifenhorst, Osmanische Urkunden in türkischer Sprache (Vienna, 1922), pp. 100–104.
What the model is told to watch for: Expect Ottoman Turkish written with an Arabic-Persian alphabet in Naskh or early Taʿliq, frequently interspersed with Arabic quotations and Persian constructions. Distinguish letters sharing the same skeleton through their sometimes faint or displaced dots, and attend to Persian letters پ, چ, ژ, and گ as well as Ottoman use of ك and ڭ-like forms. Note vowel ambiguity, Arabic grammatical diacritics in quoted passages, ligatures, honorific abbreviations, catchwords, and changes of script or ink marking language shifts.

Chester Beatty Library, Dublin, Is 1492, f. 4v, 1457–58 · Source · Public domain
Transcription being verified
The Arabic is the opening of Sūrat al-Shūrā (Q 42); no published transcription of the interlinear Persian gloss on this folio has been traced.
What the model is told to watch for: Expect Ottoman Turkish written with an Arabic-Persian alphabet in Naskh or early Taʿliq, frequently interspersed with Arabic quotations and Persian constructions. Distinguish letters sharing the same skeleton through their sometimes faint or displaced dots, and attend to Persian letters پ, چ, ژ, and گ as well as Ottoman use of ك and ڭ-like forms. Note vowel ambiguity, Arabic grammatical diacritics in quoted passages, ligatures, honorific abbreviations, catchwords, and changes of script or ink marking language shifts.

Transcription being verified
The folio is published on Wikimedia Commons without a shelfmark and the poem has not been identified; no edition covering these couplets has been traced.
What the model is told to watch for: Expect Ottoman Turkish written with an Arabic-Persian alphabet in Naskh or early Taʿliq, frequently interspersed with Arabic quotations and Persian constructions. Distinguish letters sharing the same skeleton through their sometimes faint or displaced dots, and attend to Persian letters پ, چ, ژ, and گ as well as Ottoman use of ك and ڭ-like forms. Note vowel ambiguity, Arabic grammatical diacritics in quoted passages, ligatures, honorific abbreviations, catchwords, and changes of script or ink marking language shifts.

Russian state archive, Moscow, 1456 · Source · Public domain
Transcription
Kraelitz’s printed Ottoman text of the whole document — the invocation, the legend of the tughra, the nine lines of the body and the place formula — keyed from the 1922 edition, pp. 44–45. Angle brackets are the editor’s, marking what he supplies where the photograph leaves him unsure; his superscript footnote numbers are dropped, and the medda over long ā and the vowel points on the addressee’s name are kept as he prints them. The line breaks are the edition’s own and do not coincide with the sheet: his first line ends بغدان, the document’s ends المتميز.
⟨هو⟩ محمد بن مراد خان مظفر دا⟨ئما⟩ نشآن همآيون حكمى اولدركى شمدكيحآلده مفخر الامراء المتميّز بغدان ايلى بكى پِتِرْ وُيودآ يله بآرشقليق ايدوب ارآدن دشمنلغى كوتردوم وبيوردومكى آنوك ولآيتلرنده آق كرمانده اولآن بازركآنلر كميلريله كلآلر ادرنهده وبروسآده و استآنبولده خلقله معآمله ومبآيعه ايدوب بآزركآنلق ايدهلر كلمكده وكتمكده بنوم بكلرمدن وسو بآشيلرومدن وسپآهيلرومدن وقوللرومدن هيچ بريسى بونلروك جآننه وبآشنه ومآلنه ضرر وزيان كتورميه و الّا حكمومه مخآلفت ايدوب بر وجهله مضرت ايداجك اولرلآر اسهكى ايشتدم قول كوندرب عظيم بلآيه اوغرادرين بلمش اولآلر بتى تحقيق بلب اعتماد قلآلر تحريرا فى خامس رجب المرجب سنه ستين وثمانمائه بيورت شهر ردنيك
Transcription from F. von Kraelitz-Greifenhorst, Osmanische Urkunden in türkischer Sprache aus der zweiten Hälfte des 15. Jahrhunderts (Vienna, 1922), Urkunde Nr. 1, via archive.org (Public domain).
What an early Ottoman page looks like
The Turkish of the first Ottoman centuries — the language philologists call Old Anatolian Turkish — was copied in a rounded naskh for books and in a looser, faster chancery cursive for documents. Taʿlik, the sloping hand the Ottomans borrowed from Persian practice, appears in more literary and Persianate copies and in the headings of endowment deeds. The alphabet adds four letters to the Arabic set for sounds Arabic lacks: پ pe, چ çim, ژ je and گ gaf, and the last of these is frequently written as a plain ك kef, so that gel, kel and the suffix -ki can share a single graph. A nasal ñ, later written ڭ, is likewise often just a kef.
Vowels are the hardest part. Ottoman spelling uses the consonantal letters elif, vav and ye to carry a, o/ö/u/ü and ı/i, so a three-letter skeleton such as اولدی can be read oldı, öldi or uldı depending on the word. Early copies vowel Arabic and Qurʾanic quotations fully with fetha, kesra and damma while leaving the surrounding Turkish bare, and a page may switch from Turkish to Arabic and back inside a single sentence — a hadith, a legal formula, a doxology — sometimes with a change of pen or ink to mark the shift.
Layout is equally informative. Book pages carry catchwords in the lower margin to keep the quires in order, marginal glosses in a smaller hand, and rubricated headings. Documents of the period are narrow strips of paper written across the width, with an invocation such as هو at the top, an imperial tughra when the sultan issues them, and a date given by decade of the lunar month — evâil, evâsıt or evâhir — rather than by day. Attestations by witnesses or viziers run vertically up the right-hand edge, so the reading order of a single sheet is not simply top to bottom.
Why ordinary OCR struggles here
Dots that are missing, displaced or decorative
A scribe in a hurry writes the skeleton and trusts the reader. In an early Ottoman document ب, ت, ث, ن and ي can all appear as the same tooth, and a stray dot may belong to the letter above rather than the one it sits over. Generic OCR trained on printed Arabic assumes the dots are reliable and produces confident nonsense; a model that knows the hand has to weigh the shape, the word and the formula together.
One spelling, several readings
Because Turkish vowels ride on consonantal letters, the same graph sequence can yield more than one word, and early orthography had not settled which letters to use. Choosing between them is a matter of grammar and sense rather than shape, which is exactly what character-level recognition cannot do.
Three languages on one page
Turkic grammar carries a heavy Arabic and Persian vocabulary, and quoted Arabic is often pointed with vowel signs while the Turkish around it is not. A transcription has to hold the languages apart, keep the vocalisation where the scribe supplied it, and not normalise a Persian izafet into a Turkish suffix.
Documents that do not read straight down
Fifteenth-century fermans put the invocation above the tughra, the address in the first line, the substance in rising lines below, and witness attestations turned ninety degrees in the margin. Add a seal impression over the text and a dorsal registration note, and reading order becomes an editorial decision that has to be recorded, not guessed.
Who works with this material
Editors of Old Anatolian Turkish texts
Anyone preparing a critical edition of a fourteenth- or fifteenth-century mesnevi, a mevlid or an early prose translation needs a base reading that preserves archaic forms rather than replacing them with modern Turkish. That means keeping the old suffix spellings and flagging, not resolving, the places where the manuscript is ambiguous.
Diplomatists working on early Ottoman documents
Original Turkish-language documents from before 1500 are rare and scattered across the archives of the states the Ottomans dealt with — Dubrovnik, Venice, Vienna, Moscow. Scholars comparing formulae across such a corpus need each element labelled: invocation, tughra, address, narratio, command, date place and date.
Endowment and property historians
Early vakfiyes underpin the history of mosques, medreses and imarets, and they are read for boundary clauses, staff salaries and named witnesses. The useful output is a transcription that keeps the Arabic legal formulae distinct from the Turkish description of the property.
Catalogue and digitisation teams
Libraries adding early Ottoman codices to a discovery system need incipits, colophons, ownership notes and endowment stamps captured accurately enough to search, and they need to know which parts of the page the machine could not read.
Getting the best transcription
Say which orthography you want back
Ask explicitly for the manuscript spelling in Arabic script, for a scholarly transliteration with ā, ī, ū and the ayn and hamza marked, or for both in parallel. Without that instruction any system will drift towards modern Turkish forms, and the archaic spellings that date a copy will disappear.
Test on a dated page first
Choose a colophon, a witness list or a document with a decade date and check those lines before running the rest. Names and dates expose confusions between kef and gaf and between the tooth letters faster than continuous prose does.
Tell it how to treat missing dots
Decide in advance whether an undotted letter should be output bare, with the dots the sense demands, or with a marker. State the choice in the prompt and keep it constant across a manuscript so that a later collation compares like with like.
Handle marginal and rotated text separately
Run the main text block first, then ask for the margins, the vertical attestations and the dorsal notes as their own pass with their orientation stated. Merging them in one sweep is how lines end up interleaved in the wrong order.
Further reading on this hand
The palaeography guides the in-app selector points to for this domain, if you want to check a transcription against the standard references.
- Akis: Osmanlıca Transkripsiyon Aracı
This Sabancı University Digital Humanities Laboratory resource supports transcription of Ottoman Turkish written in the Arabic-Persian alphabet. It describes the language's use from the late fourteenth to the twentieth century and works with both manuscripts and post-1729 print. Its stated focus is Naskh, the script used for a large portion of Ottoman textual production. It is useful for early Ottoman material because it connects Arabic-script character recognition with Ottoman Turkish readings rather than modern Turkish orthography.
Frequently asked questions
- Can a model tell early Ottoman naskh from taʿlik on the same page?
- It can be told to record where the hand changes, which is what matters editorially. Headings and Persian passages in a sloping taʿlik beside a Turkish body in naskh are a common arrangement, and marking the switch is more useful than trying to name the style with certainty.
- What happens to Arabic quotations inside a Turkish sentence?
- They are transcribed as written, with any vowel signs the scribe supplied, and can be tagged as Arabic so that a later index does not treat a Qurʾanic phrase as Turkish vocabulary. Nothing is translated unless you ask for it.
- Will the fifteenth-century dating formulae come out correctly?
- Dates given as evâil, evâsıt or evâhir of a lunar month are kept in that form rather than converted to a single day, because the original only ever specified a ten-day span. A converted Gregorian range can be added alongside if you want one.
- Are tughras and seals transcribed?
- A tughra is described and its legend read where it is legible, and seal impressions are noted with whatever of the inscription can be made out. Both are recorded as separate elements so that the body text stays clean.
- How reliable is the reading of a nineteenth-century facsimile plate?
- A good collotype of a document reproduces the ink faithfully and is often easier to read than a modern colour photograph of stained paper, but line-block plates flatten faint strokes. Where a facsimile is the only witness, the transcription should say so.
Related scripts and pages
Transcribe your early Ottoman folios
Upload a scan, choose Early Ottoman Naskh and Taʿliq Manuscripts, and check the first page against the image before running a whole codex.
