Uzbek
Read Uzbek in whichever script and apostrophe convention it uses
A single wall in Samarkand can carry a mosque plaque with one sentence written out in Cyrillic, Latin and Arabic script, and a street sign next to a much older Soviet plaque only in Cyrillic. Evidano tells the scripts apart, keeps the Latin alphabet’s modifier-letter apostrophe distinct from an ordinary quotation mark, and reads Chagatai’s Nastaʿliq calligraphy on its own terms.
Sample pages
Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Cincinnati Art Museum, late 15th century · Source · CC0 1.0 · Photograph by Daderot
Transcription being verified
The museum’s object record does not give a line-by-line transcription of this folio’s verse, and no separately published edition matching this specific album page has been located.
What the model is told to watch for: Produce a diplomatic transcription in Uzbek.

Source · CC BY-SA 4.0 · Photograph by Akhemen
Transcription being verified
All three lines are legible in the photograph itself; a plaque of this kind has no separately published transcription to check against.
What the model is told to watch for: Produce a diplomatic transcription in Uzbek.

Transcription being verified
The sign and plaque text are both fully legible in the photograph; no separately published transcription of this specific plaque has been located.
What the model is told to watch for: Produce a diplomatic transcription in Uzbek.

Source · CC BY-SA 4.0 · Photograph by Umarxon III
Transcription being verified
The quotation is attributed to President Mirziyoyev on the box itself; no independently published transcription of this exact wording has been located to check it against.
What the model is told to watch for: Produce a diplomatic transcription in Uzbek.
Chagatai script, two Soviet alphabets, and a Latin standard still competing with Cyrillic
Before the twentieth century, the literary language of Timurid and post-Timurid Central Asia, now generally called Chagatai, was written in Perso-Arabic script in the same Nastaʿliq hand used for Persian court poetry. Mir Ali-Shir Nava'i, the fifteenth-century Herat statesman and poet whose Chagatai verse made the language a rival to Persian in prestige, is credited as calligrapher on surviving album folios: elegant, gold-flecked Nastaʿliq lines with no short vowels marked, exactly as literary and administrative Central Asian Turkic was written for the next four centuries.
Uzbek followed the same Soviet script sequence as its Turkic neighbours: a Latin alphabet from the late 1920s, then Cyrillic from 1940, with four letters — Ў, Қ, Ғ and Ҳ — added for sounds standard Russian Cyrillic does not distinguish. Independent Uzbekistan legislated a return to Latin script in 1993 and settled the alphabet still used today in a 1995 revision, built on plain Latin letters plus an apostrophe-modified oʻ and gʻ and the digraphs sh, ch and the velar ng.
That 1995 alphabet has never fully displaced Cyrillic in practice: older generations, much of the press, and plenty of everyday signage still use it, so a single building or a single official plaque can carry both scripts, sometimes with Arabic script added as a third for religious or ceremonial text. The apostrophe in oʻ and gʻ is itself a modifier letter, not a punctuation mark, and because most keyboards and fonts don’t distinguish it from a straight quote, a curly quotation mark or a grave accent, the same word turns up spelled several different ways across otherwise identical modern documents.
Why ordinary OCR struggles here
One letter, four different apostrophes
Oʻzbek can be typed with a proper modifier-letter apostrophe (ʻ), a plain typewriter apostrophe, a curly closing quote (’) or a grave accent, all visually similar and none of them wrong to the person typing. A recogniser has to normalise these to one form or a search for oʻzbek will silently miss oʼzbek and o’zbek.
Cyrillic letters a Russian-trained model doesn’t expect
Ў, Қ, Ғ and Ҳ were added to Uzbek Cyrillic for sounds Russian doesn’t have, and a model trained mainly on Russian text collapses them into У, К, Г and Х or Н, changing which word is on the page.
Scripts that switch inside a single object
A mosque plaque or an information board can carry the same sentence in Cyrillic, Latin and Arabic script one below the other, and different signs on the same street can each use a different one of the three. Reading it correctly means detecting the script region by region rather than assuming one script for the whole image.
Chagatai Nastaʿliq with no short vowels
Fifteenth-century Chagatai calligraphy joins letters into flowing, position-dependent shapes against a decorated ground, and like other Perso-Arabic hands it omits short vowels entirely, so reading it depends on knowing the Chagatai vocabulary a line is likely to contain, not just recognising letterforms.
Who works with this material
Civil records and family history researchers
Uzbek households can hold a Soviet Cyrillic birth certificate, a Latin-alphabet passport, and older documents in Perso-Arabic script within one family, and reconstructing a record means reading whichever script each document actually uses rather than only the current one.
Digital library and search projects
Making modern Uzbek Latin text properly searchable depends on normalising the apostrophe in oʻ and gʻ to one consistent character; without that, a catalogue silently splits into several spellings of the same word.
Heritage and mosque-documentation teams
Historic and religious buildings in Samarkand, Bukhara and Khiva carry inscriptions and plaques in some combination of Cyrillic, Latin and Arabic script, often on the same object, and a survey needs all three read and matched to each other rather than only the most recent one.
Scholars and translators of Chagatai literature
Editors working on Nava'i, the Baburnama or other Chagatai texts need a diplomatic first transcription of Nastaʿliq originals that keeps the manuscript’s own spelling, rather than one that silently modernises it into current Uzbek.
Getting the best transcription
Ask for the script to be identified, region by region if needed
A single plaque or building can carry Cyrillic, Latin and Arabic script side by side. Ask the model to label which script each block of text is in rather than assuming the whole image is one script.
Fix the apostrophe as one specific character
State that oʻ and gʻ should always use the modifier-letter apostrophe (ʻ), not a straight quote or a curly one, so that every page of a batch comes out searchable as the same spelling rather than several near-identical ones.
Keep Uzbek Cyrillic’s own letters distinct from Russian
Tell the model to preserve Ў, Қ, Ғ and Ҳ as written, or to convert them to a specific Latin equivalent, rather than letting them default to the nearest Russian letter.
For Nastaʿliq material, test the most calligraphic page first
Run a decorative album folio or a heavily stylised panel before a whole batch, and ask for supplied vowels to be flagged. If that comes back clean, plainer modern Latin or Cyrillic pages need far less checking.
Frequently asked questions
- Will Uzbek oʻ and gʻ come out with the correct apostrophe, not a plain quote mark?
- Yes, if you say so in the prompt. State that you want the modifier-letter apostrophe used consistently, and the output won’t drift between straight quotes, curly quotes and grave accents the way scanned or retyped text often does.
- Can it read a single sign that has the same text in Cyrillic, Latin and Arabic script?
- Yes. It can identify each script block separately and transcribe all three, which is useful for checking that an official plaque’s three versions actually say the same thing.
- How are the Cyrillic letters Ў, Қ, Ғ and Ҳ, which aren’t in Russian, handled?
- They’re recognised as their own letters rather than folded into the nearest Russian one. You can ask for them to be kept as written or converted to their Latin equivalents.
- Can it transcribe Chagatai manuscripts like Nava'i’s Nastaʿliq calligraphy?
- It can produce a reading, but the lack of short vowels and the highly stylised, decorative letterforms of literary Nastaʿliq make this the hardest of Uzbek’s scripts. Treat the output as a starting point for someone familiar with Chagatai to check.
- Since Uzbekistan’s official script is Latin, why would I need Cyrillic transcription at all?
- Because Cyrillic hasn’t gone away: older records, much of the press, and plenty of everyday signage are still in Cyrillic, sometimes on the same wall as a newer Latin sign. Both are handled, so an archive that mixes the two doesn’t need to be split into separate projects.
Related scripts and pages
Transcribe your Uzbek documents
Upload a Latin-script form, a Cyrillic newspaper or a Chagatai manuscript, and check which script and apostrophe convention the model has used before running the rest.
