Site Logo

Hungarian

Hungarian read letter by letter, name first

Hungarian print marks a long vowel with a double acute over ő and ű, a shape most OCR training data has never seen and flattens to ö or ü, quietly turning one word into another. Every Hungarian register, from a baptismal entry to a modern form, gives the family name before the given name, the reverse of the order a Western reading model expects, and the country’s oldest surviving script, rovásírás, is not in the Latin alphabet at all. Evidano reads Hungarian print, handwriting and runic signage in the order Hungarian actually writes it, and keeps the source image next to every line.

Sample pages

Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Parchment leaf with a red Latin rubric above rows of Old Hungarian runic letters and a Hebrew alphabet below
Sample 1. The Nikolsburg alphabet, a parchment leaf of the mid-fifteenth century reused as the endpaper of a Nuremberg book printed in 1483, now Országos Széchényi Könyvtár (National Széchényi Library), Budapest, MNy 70. A red rubric — Litterae Siculorum quas sculpunt vel cidunt in lignis, “letters of the Székelys, which they carve or cut into wood”, given here with the leaf’s own abbreviations expanded, so that its sculpūt reads sculpunt — introduces the earliest surviving list of the Old Hungarian runic alphabet, rovásírás, followed lower on the leaf by a comparative Hebrew alphabet and calendar.

Source · Public domain · Országos Széchényi Könyvtár, via nyelvemlekek.oszk.hu

Transcription being verified

A photograph of the leaf was published after p. 16 of Emil Jakubovich, “A székely rovásírás legrégibb ábécéi”, Magyar Nyelv 31 (1935); the letter names have not yet been matched line by line against this image.

What the model is told to watch for: Produce a diplomatic transcription in Hungarian.

Printed Latin baptismal certificate form with handwritten entries recording a Hungarian family’s 1847 birth
Sample 2. A copy, issued in 1886, of the 1847 baptismal register entry for the Hungarian politician Ábránfalvi Ugron Gábor, made out by the Roman Catholic parish office of Egrestő (Agrișteu), Transylvania, in the Latin-language printed form standard across Austro-Hungarian church registers, with the child’s name, parents, godparents and midwife filled in by hand.

Source · CC BY-SA 4.0 · Ugron András Gábor, family collection

Transcription being verified

The certificate cites its own source as Book II, page 62, entry 24 of the parish register of Egrestő (Agrișteu); the register itself has not been consulted to verify a full transcription.

What the model is told to watch for: Produce a diplomatic transcription in Hungarian.

Front page of the Hungarian newspaper Pesti Hirlap from 1867 with dense columns of print using o with double acute
Sample 3. The front page of the specimen issue of Pesti Hírlap (“Pest News”), a Hungarian political daily, Pest, 5 March 1867, published to attract subscribers as the paper began its first year. Double-acute ő and ű and the digraph spellings appear throughout the close-set columns of political commentary.

Source · Public domain · Lázár Kálmán, Frankenburg Adolf

Transcription being verified

Digitised runs of Pesti Hírlap are held in the Hungarian National Library’s Arcanum and ADT databases; this specimen issue has not yet been matched against a proofread text.

What the model is told to watch for: Produce a diplomatic transcription in Hungarian.

Town-limit road sign for Kalocsa written in Old Hungarian runic letters on a post painted with folk-art flowers
Sample 4. A town-limit sign for Kalocsa, Hungary, photographed in 2017, giving the town’s name in rovásírás, the Old Hungarian runic script, rather than the Latin alphabet used on the town’s standard road signs; the post below it is hand-painted with the folk-art floral motifs the region is known for.

Source · CC BY-SA 4.0 · Zemszo

Transcription

The town name shown on the sign, transliterated into Latin letters; this reading follows the Commons file title and description, not a scholarly transcription of the individual rovásírás characters, but the word is short enough to check directly against the photograph.

KALOCSA

Transcription from Wikimedia Commons file page for Kalocsa rovásírásos helységnévtábla.jpg (CC BY-SA 4.0).

What Hungarian writing looks like on the page

The Hungarian alphabet uses the Latin letters plus a run of digraphs and one trigraph that each stand for a single sound: cs, dz, dzs, gy, ly, ny, sz, ty and zs. In print and in handwriting these behave as one letter for the purposes of alphabetical order, so a Hungarian dictionary, index or register files Szabó after sz and before t, not among the plain s surnames, and any process that alphabetises letter by letter puts entries in the wrong place. Long vowels carry a separate mark of their own: á, é, í, ó and ú lengthen a vowel with a single acute, while ő and ű lengthen ö and ü with a double acute unique to Hungarian, a shape that OCR trained on other Latin-script languages regularly flattens to a single acute or drops altogether, turning a word such as tűz (fire) into tüz or the surname Kőrösi into Körösi.

Hungarian gives the family name before the given name in every register, ledger and modern form, the reverse of the convention a foreign reading model expects, so a baptismal or civil entry headed Nagy János names a man called János Nagy. Church registers kept under Habsburg administration compound the problem: the same Hungarian name sits inside a Latin-language printed form, a parish clerk might Latinise its ending, and Christian-name-first Latin word order appears directly beside a Hungarian family-name-first entry on the same page. Printed matter from the mid-nineteenth century onward, once censorship eased and the press expanded after 1867, already uses the double-acute vowels and the digraph spellings consistently, so a period newspaper column reads no differently from a modern one apart from its typeface.

Older than the Latin alphabet in Hungary is rovásírás, the Old Hungarian runic script incised rather than written with a pen, recorded as early as the fifteenth century in alphabet lists such as the Nikolsburg leaf and traditionally read right to left. It fell out of everyday administrative use once the Latin alphabet took hold, but it never disappeared: it is taught as a cultural curiosity in some schools today and turns up on town-limit signs, shopfronts and memorials, almost always alongside, not instead of, the Latin-alphabet form of the same name.

Why ordinary OCR struggles here

A double acute that gets flattened to a single one

ő and ű are letters in their own right, not ö and ü wearing an accent borrowed from another language, but generic OCR trained mostly on German or Turkish text keeps mapping the double acute onto the nearest diaeresis or single acute it recognises. The result reads as plausible Hungarian while quietly saying a different word.

Digraphs that a letter-by-letter reading breaks apart

Cs, gy, ly, ny, sz, ty, zs and dzs are single sounds spelled with two or three letters, and Hungarian alphabetical order treats each as one unit. A transcription that does not know this can split a digraph across a line break as if it were two unrelated consonants, or sort a list of surnames as though sz came before t only by coincidence of its first letter.

A name order every foreign template gets backwards

Parish registers, civil records and most Hungarian forms put the family name first, so reading Kovács Éva as a given name Kovács and a surname Éva reverses the person entirely. Getting genealogical or civil data right means recognising which register convention is in play, not assuming Western name order by default.

Runic letters unrelated to any Latin shape

Rovásírás shares no letterforms with the Latin alphabet it eventually replaced, runs right to left rather than left to right, and appears today mixed into otherwise Latin-alphabet signage. A model has to recognise it as its own script before it can even attempt a transliteration.

Who works with this material

Genealogists tracing Hungarian ancestry

Family history research routinely crosses from Hungarian-language civil records into Latin-language parish registers kept under Habsburg administration, with the same person’s name in family-name-first Hungarian order in one book and Christian-name-first Latin order in the next. Reading both conventions correctly is what keeps a family tree from silently swapping generations.

Newspaper and archive digitisation projects

Libraries turning nineteenth- and twentieth-century Hungarian newspapers into searchable text need every ő, ű and digraph handled the same way across an entire run, since a search index that treats sz as s plus z, or ő as ö, will miss or misfile a large share of its own material.

Museums and heritage projects cataloguing rovásírás

Collections of runic inscriptions, calendar sticks and modern revival signage need a transliteration alongside the original characters, with the reading direction and letterforms treated as their own system rather than approximated from the Latin alphabet.

Translators and notaries handling official Hungarian documents

Birth certificates, diplomas and other documents submitted for translation or legalisation have to reproduce a name exactly, double-acute vowels included, since a Kőrösi rendered as Körösi on an official translation can be rejected as a different name entirely.

Getting the best transcription

    Step 1

    Say which name order the record uses

    Tell the model whether the source is a Hungarian family-name-first register or a Latin-language form with Christian-name-first entries, and ask for given name and family name as separate fields so nothing downstream reverses them.

    Step 2

    Keep ő and ű distinct from ö and ü

    Ask explicitly for the double-acute vowels to be preserved rather than normalised, particularly for names and place names where the difference changes the word.

    Step 3

    Flag mixed-language pages

    Where Hungarian names sit inside a Latin-language printed form, ask for each stretch labelled by language so a Latinised ending is not carried over into the Hungarian reading of the name.

    Step 4

    Ask for a Latin transliteration of rovásírás alongside the original

    For runic signage or manuscripts, request the transliteration as a second field rather than a replacement, and confirm the reading direction the model has assumed before trusting a longer inscription.

Frequently asked questions

Does it keep ő and ű distinct from ö and ü?
Yes, the double-acute vowels are read as their own letters rather than normalised to the single-acute or umlaut forms, which matters most in names and place names where the two spellings mean different things.
Will it reorder a family-name-first record into given-name, family-name for a database?
It can, if you ask for the two parts as separate fields; by default it transcribes the name in the order the document actually presents it, which is safer when you have not yet confirmed which convention a particular register follows.
Can it read Old Hungarian runic script, rovásírás?
Yes, it is recognised as its own script rather than approximated from the Latin alphabet, and the output can include a Latin-letter transliteration alongside the runic transcription.
How are digraphs such as sz, cs and zs handled?
They are kept as the single sounds they represent, so alphabetised output and line-break decisions treat szabó as starting with sz rather than splitting it into s and z.
What about Hungarian names inside Latin-language church registers?
The Latin surrounding text and the Hungarian name are read as what they are, and a Latinised name ending can be flagged separately from the vernacular form on request, rather than the two being merged into one uncertain reading.

Related scripts and pages

Turn your Hungarian documents into text

Upload a parish register, a period newspaper or a photograph of a runic sign, and check the Hungarian reading, name order included, against the image line by line.