Hebrew · c. 900–1300
Fragments in three languages and one alphabet
The store-room of the Ben Ezra synagogue in Fustat held a thousand years of paper that nobody meant to keep: business letters, court records, drafts, receipts, school exercises and torn book leaves. Most of it is written in Hebrew characters, but the language underneath may be Hebrew, Aramaic or Arabic, and it can change halfway through a sentence. Evidano transcribes what is on the fragment, marks what is lost at the edges, and keeps the writing that runs sideways up the margin apart from the main text.
Sample pages
Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Cambridge University Library, Taylor-Schechter Collection, T-S 8J18.5, 12th century · Source · Public domain
Transcription being verified
A transcription by S. D. Goitein appears in “Autographs of Yehuda Hallevi”, Tarbiz 25 (1956), and the Princeton Geniza Project records the document; both are still in copyright, so a freely licensed transcription of this leaf is being prepared.
What the model is told to watch for: Expect Hebrew square, semi-cursive, or cursive scripts used for Hebrew, Aramaic, and Arabic-language texts, sometimes with language changes inside a single line. Distinguish similar pairs ב/כ, ד/ר, ה/ח, ו/ז, and final letters, while recognizing Arabic words represented through Hebrew consonants and variable use of diacritics. Note document reuse, writing across recto and verso, marginal additions, letter-fold addresses, legal formulas, signatures, and fragments whose present edges do not preserve the original layout.

Cambridge University Library, Taylor-Schechter Collection, T-S 8J5.5, 1104 · Source · Public domain · Cambridge University Library
Transcription being verified
The four sides of this bifolio are catalogued and transcribed in the Princeton Geniza Project from S. D. Goitein’s unpublished editions, which are not freely licensed; a transcription that can be published here is being made.
What the model is told to watch for: Expect Hebrew square, semi-cursive, or cursive scripts used for Hebrew, Aramaic, and Arabic-language texts, sometimes with language changes inside a single line. Distinguish similar pairs ב/כ, ד/ר, ה/ח, ו/ז, and final letters, while recognizing Arabic words represented through Hebrew consonants and variable use of diacritics. Note document reuse, writing across recto and verso, marginal additions, letter-fold addresses, legal formulas, signatures, and fragments whose present edges do not preserve the original layout.

University of Pennsylvania Libraries, Halper 462, f. 1r, 12th century · Source · Public domain
Transcription being verified
The genealogical lists of this type are edited in Adolf Neubauer, Mediaeval Jewish Chronicles (Anecdota Oxoniensia, 1887–95); the Judeo-Arabic section of this leaf is being matched against the Penn catalogue record before it is published here.
What the model is told to watch for: Expect Hebrew square, semi-cursive, or cursive scripts used for Hebrew, Aramaic, and Arabic-language texts, sometimes with language changes inside a single line. Distinguish similar pairs ב/כ, ד/ר, ה/ח, ו/ז, and final letters, while recognizing Arabic words represented through Hebrew consonants and variable use of diacritics. Note document reuse, writing across recto and verso, marginal additions, letter-fold addresses, legal formulas, signatures, and fragments whose present edges do not preserve the original layout.

Bodleian Library, Oxford, MS Heb. d. 26 (2), facsimile 1913 · Source · Public domain · Paul Kahle, Masoreten des Ostens (Leipzig, 1913), plate 3
Transcription
Deuteronomy 14:9–13, the Hebrew verses in the left-hand column, which alternate with the Aramaic Targum verses not reproduced here; the plate is captioned Dt 14,4–19. Consonantal text as in the standard Masoretic text; the pointing printed here is Tiberian, whereas the fragment vocalises the same consonants in the Babylonian supralinear system. One line per verse.
אֶת־זֶה֙ תֹּֽאכְל֔וּ מִכֹּ֖ל אֲשֶׁ֣ר בַּמָּ֑יִם כֹּ֧ל אֲשֶׁר־ל֛וֹ סְנַפִּ֥יר וְקַשְׂקֶ֖שֶׂת תֹּאכֵֽלוּ׃ וְכֹ֨ל אֲשֶׁ֧ר אֵֽין־ל֛וֹ סְנַפִּ֥יר וְקַשְׂקֶ֖שֶׂת לֹ֣א תֹאכֵ֑לוּ טָמֵ֥א ה֖וּא לָכֶֽם׃ כׇּל־צִפּ֥וֹר טְהֹרָ֖ה תֹּאכֵֽלוּ׃ וְזֶ֕ה אֲשֶׁ֥ר לֹֽא־תֹאכְל֖וּ מֵהֶ֑ם הַנֶּ֥שֶׁר וְהַפֶּ֖רֶס וְהָֽעׇזְנִיָּֽה׃ וְהָרָאָה֙ וְאֶת־הָ֣אַיָּ֔ה וְהַדַּיָּ֖ה לְמִינָֽהּ׃
Transcription from מקרא על פי המסורה (Miqra ‘al pi ha-Mesorah), Deuteronomy 14, Hebrew Wikisource (CC BY-SA 4.0).

Younes and Soraya Nazarian Library, University of Haifa · Source · CC BY 4.0 · Younes and Soraya Nazarian Library, University of Haifa, digital collections
Transcription
Esther 1:1–3, the ten ruled lines below the invocation; the leaf breaks off in verse 3 after מִשְׁתֶּה לְכׇל, and verse 3 is printed whole here. Vocalised and accented text from Miqra ‘al pi ha-Mesorah on Hebrew Wikisource, one line per verse rather than per manuscript line.
וַיְהִ֖י בִּימֵ֣י אֲחַשְׁוֵר֑וֹשׁ ה֣וּא אֲחַשְׁוֵר֗וֹשׁ הַמֹּלֵךְ֙ מֵהֹ֣דּוּ וְעַד־כּ֔וּשׁ שֶׁ֛בַע וְעֶשְׂרִ֥ים וּמֵאָ֖ה מְדִינָֽה׃ בַּיָּמִ֖ים הָהֵ֑ם כְּשֶׁ֣בֶת ׀ הַמֶּ֣לֶךְ אֲחַשְׁוֵר֗וֹשׁ עַ֚ל כִּסֵּ֣א מַלְכוּת֔וֹ אֲשֶׁ֖ר בְּשׁוּשַׁ֥ן הַבִּירָֽה׃ בִּשְׁנַ֤ת שָׁלוֹשׁ֙ לְמׇלְכ֔וֹ עָשָׂ֣ה מִשְׁתֶּ֔ה לְכׇל־שָׂרָ֖יו וַעֲבָדָ֑יו חֵ֣יל ׀ פָּרַ֣ס וּמָדַ֗י הַֽפַּרְתְּמִ֛ים וְשָׂרֵ֥י הַמְּדִינ֖וֹת לְפָנָֽיו׃
Transcription from מקרא על פי המסורה (Miqra ‘al pi ha-Mesorah), Esther 1, Hebrew Wikisource (CC BY-SA 4.0).
What Geniza material looks like
The documentary material clusters between the tenth and the thirteenth centuries, when Fustat was a hub of Mediterranean trade and its Jewish community wrote everything down. Parchment gives way to paper during the tenth century, and paper invites speed: the formal square letters of a book are reserved for Bibles and prayer books, while letters, accounts and drafts are written in a semi-cursive or a fully joined cursive that varies from one clerk to the next. Court documents keep an old Aramaic formulary, letters are usually in Judeo-Arabic, poems and commentaries in Hebrew, and a single leaf may carry all three because the back of a used sheet was too valuable to waste.
Judeo-Arabic is Arabic written in Hebrew characters, and the mapping is not one to one. Arabic has more consonants than Hebrew has letters, so scribes add a dot or a stroke to make ג׳ for jīm, ד׳ for dhāl, ז׳ for ẓāʾ, ט׳ or ץ׳ for ḍād, ח׳ for khāʾ and ת׳ for thāʾ — and they add them inconsistently, or not at all. Spelling often follows Arabic orthography rather than pronunciation, with the definite article written אל and joined to its noun, and hamza and tāʾ marbūṭa represented by whatever the writer preferred. Names are followed by strings of abbreviated blessings, dates are given in the Seleucid era with letters for numerals, and sums of money are written out in a shorthand of the trade.
The physical state of the material is part of the reading problem. A fragment may be a strip torn from the middle of a leaf, so that every line has lost its beginning or its end. Letters were folded, addressed on the outside and often continued up the right margin at a right angle to the main text, or upside down along the top. Thin paper lets the ink of the other side show through as a mirror image, and a leaf may have been scraped and rewritten, or cut up to stiffen a binding, before it reached the store-room.
Why ordinary OCR struggles here
Arabic hiding inside Hebrew letters
A line of Judeo-Arabic looks like Hebrew to any system trained on Hebrew and like nothing at all to a system trained on Arabic. Words are divided in Arabic fashion, the diacritic dots that separate ג from ג׳ are optional, and the same consonant string can be read as two different Arabic words. Deciding which language a stretch belongs to is the first step, not an afterthought.
A cursive that joins and abbreviates
In a quick documentary hand ד and ר differ by a hairline, ב and כ by whether the corner is square, and ה and ח by a gap that a fast pen often closes. Letters run into each other, the אל of the Arabic article turns into a single stroke, and the five final forms are the only reliable clue to where one word stops.
Text that leaves the page and comes back
A letter frequently continues in the margin, written along the edge, and then finishes in a few words squeezed above the opening line. Reading order is not top to bottom. Any transcription has to say which block is which, note the rotation, and keep the address panel on the other side separate from the body of the letter.
Edges that are simply gone
Because these are discards, most items are incomplete. Half a word at the tear, a hole where an insect ate through several lines, and ink from the verso showing through all look like letters to a system that must output something. What is needed instead is an honest lacuna: a mark saying this much is missing here.
Who works with this material
Social and economic historians
The India trade, the marriage market, charity lists and communal quarrels are all reconstructed from these scraps. A researcher working through a box of unedited fragments needs a rough reading fast enough to decide which ones repay a full edition, and accurate enough to search for a merchant’s name across thousands of images.
Corpus and cataloguing projects
Teams building searchable Geniza corpora need consistent handling of Judeo-Arabic orthography, of lacunae and of the join between fragments held in different libraries. Machine-readable text from the images is what makes two halves of the same letter, one in Cambridge and one in New York, findable at all.
Palaeographers identifying scribes
Attributing a document to a known court clerk depends on the shape of individual letters and on the formulae a writer habitually uses. A transcription that records what is written rather than what is expected gives the evidence for such an attribution instead of blurring it.
Genealogists and family historians
Marriage contracts, divorce deeds and legal releases name parties, fathers, witnesses and places. Extracting those names, with their honorifics and their Seleucid dates, turns a photograph of a torn deed into a record that can be indexed and searched.
Getting the best transcription
Name the languages you expect
Say that the document may contain Hebrew, Aramaic and Judeo-Arabic, all in Hebrew characters, and ask for a language label on each block or line. That one instruction stops Arabic words being silently corrected into Hebrew ones and makes the output usable for a bilingual index.
Decide what to do with the Arabic diacritics
Choose whether the pointed forms ג׳, ד׳, ז׳, ט׳, ח׳ and ת׳ should be reproduced only where the scribe wrote them, or supplied throughout for searching. Recording them as written is usually right for an edition, since their presence or absence is evidence about the writer.
Ask for the margins as separate blocks
Request the main text first, then each marginal or upside-down block in its own section with a note of where it sits and which way it runs. Trying to interleave marginal lines into the body from the image alone produces a text nobody can check against the fragment.
Test on your worst fragment
Start with a strip that has lost both edges and shows the verso through the paper. If the losses come back marked, the show-through is ignored and the surviving half-words are given as half-words, the rest of the box will be straightforward.
Further reading on this hand
The palaeography guides the in-app selector points to for this domain, if you want to check a transcription against the standard references.
- The Scribes of the Cairo Geniza
This University of Pennsylvania project applies digital methods to the scribes and scripts represented in Cairo Geniza fragments. It describes a multilingual corpus containing Hebrew, Aramaic, Judeo-Arabic, and other documentary languages written primarily in Hebrew characters. The project emphasizes script classification, writer identification, fragmentary layout, and transcription of difficult manuscript images. It is directly useful for Geniza documents because linguistic identification and palaeographical hand analysis must be performed together.
Frequently asked questions
- Can it read Judeo-Arabic, not just Hebrew?
- Yes. Judeo-Arabic is treated as Arabic written in Hebrew characters, with its own word division and its own spelling habits, and can be transcribed in Hebrew letters or transliterated into Arabic script if that is more useful for your index.
- How are missing edges and holes represented?
- As lacunae. A break in the middle of a word is marked at the point of loss, and nothing is invented to fill it. If you want a conjectural restoration you can ask for it separately, in brackets, so that it is never confused with what survives on the fragment.
- What happens to writing in the margins or upside down?
- It is kept as its own block, with a note of its position and orientation, and placed after the main text rather than woven into it. That keeps the reading order recoverable and lets you decide where a marginal sentence belongs.
- Does the ink showing through from the other side confuse it?
- Show-through reads as faint mirror-image letters and can be ignored on request, so that only the writing on the side you photographed is transcribed. If both sides matter, photograph and run them separately and the two texts stay apart.
- Can it date a document from the formulae it uses?
- It will transcribe a date where one is written, including Seleucid year numbers spelled with Hebrew letters, and can flag the standard opening and closing formulae of court documents. Deciding what those formulae imply about place and period remains a job for a historian.
Related scripts and pages
Transcribe your Geniza fragments
Upload a fragment, choose Cairo Geniza Documents as the domain, and see how the margins and the torn edges come back before you run a whole box.
