Romanian
Romanian read across three alphabets
Romanian spent roughly three centuries written in a Cyrillic alphabet drawn from Church Slavonic, then passed through decades in which Cyrillic and Latin letters sat inside the same word, before settling on the Latin alphabet readers know today — except across the Dniester, where a Soviet-built Cyrillic form stayed official until 1989 and still turns up on signs. Modern spelling carries its own trap: ș and ț are properly written with a comma below the letter, but generations of substitute fonts and keyboards render them with a cedilla instead, a difference easy to miss on screen but damaging to search and matching. Evidano reads Romanian across all of these forms and keeps the source image next to every line.
Sample pages
Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Source · Public domain · Reproduced in Gheorghe T. Zaharia et al., Iași: A City of Great Destinies (Meridiane, Bucharest, 1986)
Transcription
The printed body text (transliterated into modern Latin-script Romanian on the Commons file page, up to where the page is cut off) followed by Ion Creangă’s marginal ownership note, quoted in full in the file’s description field.
PENTRU IZBODIREA GHEOGRAFIEI Gheografie (după cum și alte învățături) au avut asale tinerețe și adunări, după măsura înmulțirii oamenilor, care dintru întâi au început alăcui pământul. și multe numiri au avut, pentru că, când oamenii sau adunat întrun sângur loc îl numiră (topografie) (adecă, scrisoarea locului,) iară când să întinsără cu lăcuința mai pe mult loc, onumiră, Horografie) (adecă scrisoare de sate, și cetăți,) și iarăși când împliniră omare parte de apământului, onumiră (gheografie) (adecă scrisoarea pământului,) care poate săse facă istoricește, filosofește și mathimaticește. Gheografie istoricească, cuprinde întru sine arătările de trebuița țărilor, și lucrările care lau luminat. Gheografie filosofească cuprinde întru sine apropierile firești, alocurilor, (sau istoria cea firească alor.) Gheografie mathimaticească, este aceea care însem[…] Din cărţile subscrisului Ion Creangă 1878. Dăruită mie de dl. Mihail Eminescu, eminentul scriitor şi cel mai mare poet al Românilor. 1878. Nu lipseşte din ea nici o filă. I.Cr.
Transcription from Wikimedia Commons file page for Din cărţile subscrisului Ion Creangă 1878.jpg (Public domain).

Source · CC BY-SA 4.0 · Dan Rășcanu
Transcription being verified
A full run of Evenimentul (Iași, 1893–1940s) is held in Romanian research libraries; this issue has not yet been matched against a proofread edition.
What the model is told to watch for: Produce a diplomatic transcription in Romanian.

Source · Public domain · Rakoon
Transcription being verified
A personal document with no published transcription; the printed clauses on the left leaf follow the standard wording used on Romanian identity booklets of the early 1960s.
What the model is told to watch for: Produce a diplomatic transcription in Romanian.

Source · CC BY-SA 3.0 · Keizers
Transcription
A plain reading of the sign, transcribed as displayed in Cyrillic capitals; this follows the Commons file description, which gives the Latin-script Romanian equivalent, Bine ați venit!, and its English translation, “Welcome”, rather than a published scholarly transcription.
БИНЕ АЦЬ ВЕНИТ!
Transcription from Wikimedia Commons file page for SignInMoldovanCyrillic.JPG (CC BY-SA 3.0).
How Romanian moved through three alphabets
Romanian was written in a Cyrillic alphabet for roughly three centuries before the Latin alphabet familiar today became compulsory. Church Slavonic had been the liturgical and chancery language of Wallachia and Moldavia into the sixteenth century, and when Romanian itself began to appear in print and in charters it inherited Cyrillic letterforms and scribal habits wholesale: a titlo, a horizontal stroke over an abbreviated word, marks contractions exactly as it does in Church Slavonic manuscripts, and letters double as numerals under the same stroke rather than appearing as Arabic digits. Romanian Cyrillic added a handful of letters and ligatures that Slavonic did not need, for the schwa vowel ă and for the central vowel now written î or â, so a modern reader who knows Russian or Bulgarian Cyrillic still meets unfamiliar shapes on a Romanian page.
The move to Latin letters was neither sudden nor total. Reformers such as Ion Heliade Rădulescu promoted a transitional alphabet from the 1830s that spelled a single word with some letters in Cyrillic and some in Latin, on the theory that readers already literate in Cyrillic could be walked into the new alphabet one letter at a time; the effect on the page is a hybrid that neither an all-Cyrillic nor an all-Latin reading habit expects. Wallachia and Moldavia made Latin letters compulsory for official use in the early 1860s, but church printing lagged for a generation, so a Cyrillic-set book can carry a date as late as the 1870s or 1880s. Cyrillic never fully disappeared from Romanian either: Soviet Moldova wrote Romanian, officially called Moldovan, in a Cyrillic alphabet built on the Russian model from the 1920s until 1989, and the breakaway territory of Transnistria still uses it on signage today.
Modern printed Romanian is set in the Latin alphabet with five extra letters, ă, â, î, ș and ț, of which the last two are officially a comma set below the letter rather than a cedilla; decades of fonts, printers and keyboard layouts built for Turkish or for Central European cedilla letters instead put the wrong diacritic into circulation so widely that ș and ț with a cedilla appear throughout the twentieth-century press and in some digitised archives today. Newspapers before the Second World War still carry small, close-set type across several columns, and everyday Romanian handwriting on forms and identity papers favours a fast administrative cursive that abbreviates place names and standard phrases the way any bureaucratic hand does.
Why ordinary OCR struggles here
Cyrillic letters a modern model has never learned
Romanian Cyrillic reuses Church Slavonic letters that dropped out of Russian after its own eighteenth-century reform, plus letters invented for Romanian’s own vowels. A model trained only on modern Russian or Bulgarian civil script has no category for these shapes and either drops them or substitutes the nearest Russian letter, which is wrong far more often than it is right.
One word, two alphabets
The transitional spelling of the 1830s to 1850s puts Cyrillic and Latin letters inside the same word by design, so anything that assumes a document, or even a single line, uses one script will misread half of every hybrid word. Recognising the two alphabets together, rather than picking one and forcing the rest to fit, is the only way to get the word right.
A comma that looks like a cedilla
Ș and ț are properly written with a comma below the letter, but a great deal of printed and digitised Romanian instead carries a Turkish-style cedilla by font accident, and the two are close to indistinguishable on a scanned page at normal size. Getting the underlying character right matters for anyone who needs the output to match modern spelling standards or to be searchable against correctly encoded Romanian text.
Administrative cursive and inherited abbreviations
Identity booklets, land registers and army papers are filled in by hand in a fast administrative script that shortens common words such as str., nr., jud. and dl., and packs a birthplace, a district and a date into a few crowded lines. Reading it well means recognising the abbreviation as a known administrative term rather than treating it as a spelling the clerk happened to choose.
Who works with this material
Genealogists and family historians
Romanian family research usually runs through several reissues of the same identity booklet or parish register, each filled in by a different hand across decades, and further back it runs into Cyrillic-era church registers entirely. A transcription that reads both the cursive Latin script and the older Cyrillic consistently is what lets a family tree cross that boundary.
Editors of pre-1860s Cyrillic texts
Scholars preparing a critical edition of a Wallachian or Moldavian chronicle, chancery register or early printed book need a diplomatic transcription that keeps the Cyrillic letters, the titlo abbreviations and the letter-numerals as written, so the edition can be checked against the source rather than against someone’s earlier guess at it.
Press and library digitisation projects
Libraries turning nineteenth- and twentieth-century Romanian newspapers into searchable text need every ș and ț encoded the same way across an entire run, or full-text search silently fails on half the archive. Consistent diacritic handling is what makes a digitised newspaper collection actually searchable rather than merely browsable.
Researchers of Soviet Moldova and Transnistria
Historians and linguists working with Moldovan-language administrative or press material from the Soviet period need the Cyrillic Moldovan alphabet read on its own terms, not auto-corrected into Russian spelling, since the differences between the two are exactly what the research usually turns on.
Getting the best transcription
Say which alphabet you expect, or ask for both
Tell the model whether the page is Cyrillic, transitional or Latin, so it applies the right letter set from the first line instead of guessing partway through a hybrid word.
Fix the diacritic you want
Specify whether ș and ț should come back with the comma below or be normalised to whatever your downstream system expects, so the exported text matches the standard you are feeding it into.
Give the administrative vocabulary up front
List the abbreviations you expect to see, such as str., nr., jud. and n., so a cursive clerk’s shorthand is read as the known term rather than guessed letter by letter.
Test on the messiest field first
Run the crowded personal-details column of an identity booklet or the smallest newspaper type before a full batch, and check the reading against the image line by line before trusting the rest.
Frequently asked questions
- Can it read Romanian text written in Cyrillic letters?
- Yes. Romanian Cyrillic is read as its own alphabet, including the letters and ligatures added for sounds Church Slavonic did not have, and the output can be given as a diplomatic transcription or converted to modern Latin-script spelling.
- How does it handle the transitional alphabet that mixes Cyrillic and Latin?
- Both alphabets are recognised within the same word rather than one being forced to fit the other, which is what a mid-nineteenth-century transitional text actually requires.
- Will ș and ț come back with the comma or the cedilla?
- You choose. The default follows current Romanian orthography, a comma below the letter, but the output can be set to match whatever encoding your existing documents use.
- Can it read the Cyrillic alphabet used in Soviet Moldova and Transnistria?
- Yes, the Moldovan Cyrillic alphabet is treated as its own variant rather than mapped onto Russian, which matters because the two diverge in exactly the letters that carry the most historical evidence.
- What about handwritten Romanian on forms and identity papers?
- Administrative cursive is read with attention to the standard abbreviations clerks use, so a shortened place name or a hurried date is resolved as the term it stands for rather than left as an unclear scrawl.
Related scripts and pages
Turn your Romanian documents into text
Upload a Cyrillic manuscript, a pre-war newspaper or a handwritten identity booklet, and check the Romanian reading against the image line by line.
