Site Logo

Hebrew · c. 900–1300

Every layer of a masoretic page, kept apart

A biblical codex of this period carries three texts at once: the consonants, a system of vowel points and cantillation signs written under, inside and above them, and a running commentary on the spelling of that very text, squeezed into the margins in letters a third of the size. Evidano reads the columns right to left, keeps the layers separate, and sets the result beside the folio so you can check a hireq against a sere before you trust it.

Sample pages

Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Two columns of the Aleppo Codex in formal square Hebrew, with tiny masoretic notes between the columns and two lines of Masorah magna above
Sample 1. The Aleppo Codex, written in Tiberias c. 930 and vocalised by Aaron ben Asher, photographed by William Wickes in 1887. Two of the folio’s three columns, with Masorah magna in the upper margin, Masorah parva in the space between the columns, and a closed section break before Genesis 27:1.

Ben-Zvi Institute, Jerusalem, c. 930 (photographed 1887) · Source · Public domain

Transcription

Genesis 26:34–27:10, the right-hand of the two columns shown, from the vocalised and accented text of Miqra ‘al pi ha-Mesorah on Hebrew Wikisource, which follows the Aleppo Codex where it survives. The column opens in the middle of verse 34, at the first הַחִתִּי, and breaks off in verse 10 after אֲשֶׁר; both verses are printed whole here. At 27:3 the codex writes the ketiv צידה and the edition prints the qere צָיִד. One line per verse, not per manuscript line.

וַיְהִ֤י עֵשָׂו֙ בֶּן־אַרְבָּעִ֣ים שָׁנָ֔ה וַיִּקַּ֤ח אִשָּׁה֙ אֶת־יְהוּדִ֔ית בַּת־בְּאֵרִ֖י הַֽחִתִּ֑י וְאֶת־בָּ֣שְׂמַ֔ת בַּת־אֵילֹ֖ן הַֽחִתִּֽי׃
וַתִּהְיֶ֖יןָ מֹ֣רַת ר֑וּחַ לְיִצְחָ֖ק וּלְרִבְקָֽה׃
וַֽיְהִי֙ כִּֽי־זָקֵ֣ן יִצְחָ֔ק וַתִּכְהֶ֥יןָ עֵינָ֖יו מֵרְאֹ֑ת וַיִּקְרָ֞א אֶת־עֵשָׂ֣ו ׀ בְּנ֣וֹ הַגָּדֹ֗ל וַיֹּ֤אמֶר אֵלָיו֙ בְּנִ֔י וַיֹּ֥אמֶר אֵלָ֖יו הִנֵּֽנִי׃
וַיֹּ֕אמֶר הִנֵּה־נָ֖א זָקַ֑נְתִּי לֹ֥א יָדַ֖עְתִּי י֥וֹם מוֹתִֽי׃
וְעַתָּה֙ שָׂא־נָ֣א כֵלֶ֔יךָ תֶּלְיְךָ֖ וְקַשְׁתֶּ֑ךָ וְצֵא֙ הַשָּׂדֶ֔ה וְצ֥וּדָה לִּ֖י צָֽיִד׃
וַעֲשֵׂה־לִ֨י מַטְעַמִּ֜ים כַּאֲשֶׁ֥ר אָהַ֛בְתִּי וְהָבִ֥יאָה לִּ֖י וְאֹכֵ֑לָה בַּעֲב֛וּר תְּבָרֶכְךָ֥ נַפְשִׁ֖י בְּטֶ֥רֶם אָמֽוּת׃
וְרִבְקָ֣ה שֹׁמַ֔עַת בְּדַבֵּ֣ר יִצְחָ֔ק אֶל־עֵשָׂ֖ו בְּנ֑וֹ וַיֵּ֤לֶךְ עֵשָׂו֙ הַשָּׂדֶ֔ה לָצ֥וּד צַ֖יִד לְהָבִֽיא׃
וְרִבְקָה֙ אָֽמְרָ֔ה אֶל־יַעֲקֹ֥ב בְּנָ֖הּ לֵאמֹ֑ר הִנֵּ֤ה שָׁמַ֙עְתִּי֙ אֶת־אָבִ֔יךָ מְדַבֵּ֛ר אֶל־עֵשָׂ֥ו אָחִ֖יךָ לֵאמֹֽר׃
הָבִ֨יאָה לִּ֥י צַ֛יִד וַעֲשֵׂה־לִ֥י מַטְעַמִּ֖ים וְאֹכֵ֑לָה וַאֲבָרֶכְכָ֛ה לִפְנֵ֥י יְהֹוָ֖ה לִפְנֵ֥י מוֹתִֽי׃
וְעַתָּ֥ה בְנִ֖י שְׁמַ֣ע בְּקֹלִ֑י לַאֲשֶׁ֥ר אֲנִ֖י מְצַוָּ֥ה אֹתָֽךְ׃
לֶךְ־נָא֙ אֶל־הַצֹּ֔אן וְקַֽח־לִ֣י מִשָּׁ֗ם שְׁנֵ֛י גְּדָיֵ֥י עִזִּ֖ים טֹבִ֑ים וְאֶֽעֱשֶׂ֨ה אֹתָ֧ם מַטְעַמִּ֛ים לְאָבִ֖יךָ כַּאֲשֶׁ֥ר אָהֵֽב׃
וְהֵבֵאתָ֥ לְאָבִ֖יךָ וְאָכָ֑ל בַּעֲבֻ֛ר אֲשֶׁ֥ר יְבָרֶכְךָ֖ לִפְנֵ֥י מוֹתֽוֹ׃

Transcription from מקרא על פי המסורה (Miqra ‘al pi ha-Mesorah), Genesis 26–27, Hebrew Wikisource (CC BY-SA 4.0).

Leningrad Codex folio with the end of the Song of the Sea laid out in a brick pattern of short blocks, prose beneath, and lines of tiny Masorah magna above
Sample 2. Leningrad Codex, St Petersburg, National Library of Russia, Firkovich B 19 A, f. 40v, copied in Cairo in 1008 by Samuel ben Jacob. The close of the Song of the Sea in the stichographic “half-brick over whole-brick” layout, with Masorah magna and a micrographic crown above and prose resuming below.

National Library of Russia, St Petersburg, Firkovich B 19 A, f. 40v, 1008 · Source · Public domain

Transcription

Exodus 15:14–21: the stichographic block and the first prose lines below it. The line breaks are the edition’s stichographic divisions, which follow the manuscript layout; ׀ marks a paseq. From Miqra ‘al pi ha-Mesorah on Hebrew Wikisource, with vowel points and cantillation signs.

שָֽׁמְע֥וּ עַמִּ֖ים יִרְגָּז֑וּן
חִ֣יל אָחַ֔ז יֹשְׁבֵ֖י פְּלָֽשֶׁת׃
אָ֤ז נִבְהֲלוּ֙ אַלּוּפֵ֣י אֱד֔וֹם
אֵילֵ֣י מוֹאָ֔ב יֹֽאחֲזֵ֖מוֹ רָ֑עַד
נָמֹ֕גוּ כֹּ֖ל יֹשְׁבֵ֥י כְנָֽעַן׃
תִּפֹּ֨ל עֲלֵיהֶ֤ם אֵימָ֙תָה֙ וָפַ֔חַד
בִּגְדֹ֥ל זְרוֹעֲךָ֖ יִדְּמ֣וּ כָּאָ֑בֶן
עַד־יַעֲבֹ֤ר עַמְּךָ֙ יְהֹוָ֔ה
עַֽד־יַעֲבֹ֖ר עַם־ז֥וּ קָנִֽיתָ׃
תְּבִאֵ֗מוֹ וְתִטָּעֵ֙מוֹ֙ בְּהַ֣ר נַחֲלָֽתְךָ֔
מָכ֧וֹן לְשִׁבְתְּךָ֛ פָּעַ֖לְתָּ יְהֹוָ֑ה
מִקְּדָ֕שׁ אֲדֹנָ֖י כּוֹנְנ֥וּ יָדֶֽיךָ׃
יְהֹוָ֥ה ׀ יִמְלֹ֖ךְ לְעֹלָ֥ם וָעֶֽד׃
כִּ֣י בָא֩ ס֨וּס פַּרְעֹ֜ה בְּרִכְבּ֤וֹ וּבְפָרָשָׁיו֙ בַּיָּ֔ם
וַיָּ֧שֶׁב יְהֹוָ֛ה עֲלֵהֶ֖ם אֶת־מֵ֣י הַיָּ֑ם
וּבְנֵ֧י יִשְׂרָאֵ֛ל הָלְכ֥וּ בַיַּבָּשָׁ֖ה בְּת֥וֹךְ הַיָּֽם׃
וַתִּקַּח֩ מִרְיָ֨ם הַנְּבִיאָ֜ה אֲח֧וֹת אַהֲרֹ֛ן אֶת־הַתֹּ֖ף בְּיָדָ֑הּ וַתֵּצֶ֤אןָ כׇֽל־הַנָּשִׁים֙ אַחֲרֶ֔יהָ בְּתֻפִּ֖ים וּבִמְחֹלֹֽת׃
וַתַּ֥עַן לָהֶ֖ם מִרְיָ֑ם שִׁ֤ירוּ לַֽיהֹוָה֙ כִּֽי־גָאֹ֣ה גָּאָ֔ה ס֥וּס וְרֹכְב֖וֹ רָמָ֥ה בַיָּֽם׃

Transcription from מקרא על פי המסורה (Miqra ‘al pi ha-Mesorah), Exodus 15, Hebrew Wikisource (CC BY-SA 4.0).

Two stained parchment columns of large Oriental square Hebrew with vowel points, and faint marginal notes at the outer edge
Sample 3. The Damascus Pentateuch, Jerusalem, National Library of Israel, Ms. Heb. 24°5702, leaf 5, an Oriental codex of the tenth century. Two of the three columns of the genealogies of Genesis 11 in a large, widely spaced square hand with Tiberian pointing; damp staining has thinned the ink across the middle of the page.

National Library of Israel, Jerusalem, Ms. Heb. 24°5702, 10th century · Source · Public domain · Ktiv project, National Library of Israel

Transcription

Genesis 11:17–22, the right-hand of the two columns shown; the column starts in the middle of verse 17, at שְׁלֹשִׁים שָׁנָה, and ends with verse 22 before the next וַיְחִי. Vocalised and accented text from Miqra ‘al pi ha-Mesorah on Hebrew Wikisource, one line per verse.

וַֽיְחִי־עֵ֗בֶר אַחֲרֵי֙ הוֹלִיד֣וֹ אֶת־פֶּ֔לֶג שְׁלֹשִׁ֣ים שָׁנָ֔ה וְאַרְבַּ֥ע מֵא֖וֹת שָׁנָ֑ה וַיּ֥וֹלֶד בָּנִ֖ים וּבָנֽוֹת׃
וַֽיְחִי־פֶ֖לֶג שְׁלֹשִׁ֣ים שָׁנָ֑ה וַיּ֖וֹלֶד אֶת־רְעֽוּ׃
וַֽיְחִי־פֶ֗לֶג אַחֲרֵי֙ הוֹלִיד֣וֹ אֶת־רְע֔וּ תֵּ֥שַׁע שָׁנִ֖ים וּמָאתַ֣יִם שָׁנָ֑ה וַיּ֥וֹלֶד בָּנִ֖ים וּבָנֽוֹת׃
וַיְחִ֣י רְע֔וּ שְׁתַּ֥יִם וּשְׁלֹשִׁ֖ים שָׁנָ֑ה וַיּ֖וֹלֶד אֶת־שְׂרֽוּג׃
וַיְחִ֣י רְע֗וּ אַחֲרֵי֙ הוֹלִיד֣וֹ אֶת־שְׂר֔וּג שֶׁ֥בַע שָׁנִ֖ים וּמָאתַ֣יִם שָׁנָ֑ה וַיּ֥וֹלֶד בָּנִ֖ים וּבָנֽוֹת׃
וַיְחִ֥י שְׂר֖וּג שְׁלֹשִׁ֣ים שָׁנָ֑ה וַיּ֖וֹלֶד אֶת־נָחֽוֹר׃

Transcription from מקרא על פי המסורה (Miqra ‘al pi ha-Mesorah), Genesis 11, Hebrew Wikisource (CC BY-SA 4.0).

Small parchment page of pointed Hebrew Psalms with a red and blue decorated word panel and Latin annotations crowded into every margin
Sample 4. A pocket Hebrew Psalter used as a schoolbook in England in the first half of the thirteenth century: Oxford, Bodleian Library, MS. Bodl. Or. 621, f. 2v. Vocalised Hebrew in a small square hand, the opening of Psalm 9 in a red and blue panel, and Latin and French notes on grammar and vocabulary added by at least three Christian readers.

Bodleian Library, Oxford, MS. Bodl. Or. 621, f. 2v, early 13th century · Source · Public domain

Transcription

Psalms 8:4–10, the six Hebrew lines above the decorated panel; the page then continues into Psalm 9. Vocalised text without cantillation signs, matching this copy, from Miqra ‘al pi ha-Mesorah on Hebrew Wikisource; that edition ends each verse with a full stop where the manuscript writes sof pasuq, and the manuscript abbreviates the divine name as a double yod where the edition prints יְהֹוָה. The column begins in mid-verse 4, at מַעֲשֵׂה אֶצְבְּעֹתֶיךָ.

כִּי אֶרְאֶה שָׁמֶיךָ מַעֲשֵׂה אֶצְבְּעֹתֶיךָ יָרֵחַ וְכוֹכָבִים אֲשֶׁר כּוֹנָנְתָּה.
מָה אֱנוֹשׁ כִּי תִזְכְּרֶנּוּ וּבֶן אָדָם כִּי תִפְקְדֶנּוּ.
וַתְּחַסְּרֵהוּ מְּעַט מֵאֱלֹהִים וְכָבוֹד וְהָדָר תְּעַטְּרֵהוּ.
תַּמְשִׁילֵהוּ בְּמַעֲשֵׂי יָדֶיךָ כֹּל שַׁתָּה תַחַת רַגְלָיו.
צֹנֶה וַאֲלָפִים כֻּלָּם וְגַם בַּהֲמוֹת שָׂדָי.
צִפּוֹר שָׁמַיִם וּדְגֵי הַיָּם עֹבֵר אׇרְחוֹת יַמִּים.
יְהֹוָה אֲדֹנֵינוּ מָה אַדִּיר שִׁמְךָ בְּכׇל הָאָרֶץ.

Transcription from מקרא על פי המסורה (Miqra ‘al pi ha-Mesorah), Psalms 8, Hebrew Wikisource (CC BY-SA 4.0).

What a masoretic folio looks like

The great pointed Bibles were made between roughly the late ninth and the thirteenth centuries, first in the eastern Mediterranean and later in Spain, Ashkenaz, Italy and Yemen. The model copies are Oriental: the Aleppo Codex, written in Tiberias around 930 and vocalised by Aaron ben Asher, and the Leningrad Codex, finished in Cairo in 1008 and still the base text of most printed critical Bibles. The page is parchment, ruled with a hard point, and laid out in two or three narrow columns of formal square script; the Pentateuch usually takes three columns, the poetic books two. Word division is generous, letters are stretched or compressed so that a line ends flush, and the scribe never breaks a word across lines.

The consonantal skeleton is only the beginning. Under and inside each letter sits the Tiberian pointing: seven vowel qualities plus shewa and the reduced hatef vowels, a dot for dagesh or mappiq, and a dot to the left or right of shin. Above and below run the te‘amim, the cantillation signs, which mark both melody and syntax, so that a single word can carry a consonant, a vowel, a dagesh and an accent stacked in one column of glyphs. Layout is meaningful as well: open and closed section breaks (petuhah and setumah) are shown by blank space, the Song of the Sea and the Song of Deborah are set out in a brick pattern of alternating written and blank blocks, and enlarged, diminished, suspended or inverted letters are copied exactly because the Masorah counts them.

Around the text block runs the Masorah itself. The Masorah parva sits in the side margins in tiny script: numeral letters counting how often a form occurs, the abbreviation ל for a word found once in the Bible, and qere notes recording that a word is read otherwise than it is written. The Masorah magna takes two or three lines above and below the columns, and in many later codices it is drawn out as micrography, the words themselves forming knotwork or arcades. It is not decoration to be skipped but the apparatus that kept the text fixed.

Why ordinary OCR struggles here

Three signals stacked on one letter

A vowel point, a dagesh and an accent can all belong to one consonant, sitting below it, inside it and above it. Ordinary page OCR drops the small marks as noise or attaches them to the neighbouring letter, which turns a pausal form into a different word. Each cluster has to be read as a unit and written out in a fixed order.

Letters distinguished by one short stroke

Square Hebrew separates ב from כ by a squared corner, ד from ר by a small tick at the top right, ה from ח by a gap in the left leg, and ו from ז by the length of the head. Five letters take a different shape at the end of a word (ך ם ן ף ץ), and final kaf and non-final nun look alike. On worn parchment a single fading stroke changes both the word and its grammar.

Marginal Masorah is a second document

The marginal notes use their own abbreviations and numeral letters, and they refer to words in the column beside them rather than to the line they sit on. Treating them as running text produces nonsense; leaving them out throws away the reason the manuscript was made. They belong in a separate stream, linked back to the word they annotate.

White space that carries meaning

A gap of nine letters in the middle of a line is a closed section; a line left blank is an open one. Poetry is set in columns of half-verses with blank blocks between them. Any pipeline that normalises whitespace or reflows columns destroys information that editors depend on, and column order is right to left, which most layout analysis gets backwards.

Who works with this material

Editors collating biblical manuscripts

Anyone comparing a codex against the Leningrad base text needs the consonants, the vowels and the accents as three separate, searchable layers, plus a note of every plene or defective spelling. A first pass that keeps them apart turns collation into checking rather than copying.

Masorah and grammar projects

Research on the Masorah depends on the marginal notes, not the biblical text: the counts, the ל sigla, the qere notes and the lists in the upper and lower margins. Pulling those into their own file, with numeral letters resolved and each note anchored to the word it comments on, is the slow part of the work.

Libraries describing Hebrew collections

Cataloguers need to say which verses a leaf carries before it can be identified, matched to sister leaves or aligned in a viewer. A transcription of even a few lines is enough to place a stray folio in a known codex or in the Genizah scatter.

Getting the best transcription

    Step 1

    Say which layers you want

    Ask for consonants only, consonants with vowels, or the fully accented text, and state the order in which the marks should be written out. Being explicit avoids a mixed result where some words carry accents and others do not, and it makes two runs of the same manuscript comparable.

    Step 2

    Give the column order and the margins a rule

    Tell the model that columns are read right to left, that the Masorah parva in the side margins is a separate stream, and whether the Masorah magna above and below the text block should be transcribed, skipped or marked as present. Otherwise the small script drifts into the biblical text.

    Step 3

    Test on a page with a section break and a qere

    Pick a folio that has an open or closed section, a word with a qere note, and a stretch of stichographic poetry. If the spacing, the marginal note and the brick layout all survive that page, ordinary prose folios will be straightforward.

    Step 4

    Ask for doubt to be marked, not smoothed

    On stained or rubbed parchment the difference between ד and ר, or between two dots and three, is often genuinely unclear. Request that uncertain marks be flagged rather than supplied from the standard text, so that the transcription records the manuscript and not the printed Bible.

Further reading on this hand

The palaeography guides the in-app selector points to for this domain, if you want to check a transcription against the standard references.

  • Digital Hebrew Palaeography: Script Types and Modes

    This scholarly article presents a systematic framework for classifying medieval Hebrew scripts in digital palaeography. It distinguishes square, semi-square, and cursive modes and regional traditions including Oriental, Sephardi, Ashkenazi, Italian, Byzantine, and Yemenite. The study explains how script classification depends on recurring letter construction rather than isolated decorative features. Its treatment of formal square modes is useful for locating biblical codices within regional and chronological Hebrew writing traditions.

Frequently asked questions

Can it transcribe the vowel points and cantillation, or only the consonants?
Both, and you choose. The default keeps everything the scribe wrote, including shewa, hatef vowels, dagesh, mappiq and the accents above and below the line. If you want a consonantal text for searching, ask for the pointing to be stripped and it is removed consistently rather than at random.
How is the Masorah in the margins handled?
It is treated as a second text. The notes in the side margins can be transcribed separately and tied to the line they face, and the lists above and below the columns can be captured or simply flagged as present, depending on what you ask for. Numeral letters can be left as letters or resolved into figures.
What happens with qere and ketiv?
The written form in the column and the note in the margin are recorded separately, so nothing is silently corrected. If you prefer a reading text you can ask for the qere to be shown in brackets after the ketiv, which keeps both and makes the choice visible.
Does it keep the section breaks and the poetic layout?
Yes, if you ask for layout to be preserved. Open and closed sections can be marked in the output, and the stichographic blocks of the Song of the Sea, the Song of Deborah or the poetic books can be kept as separate lines instead of being run together into prose.
Can it tell the divine name from its abbreviations?
The four-letter name, the double or triple yod used as an abbreviation in many medieval copies, and the substitute אֲדֹנָי are distinguished as written. Nothing is expanded or replaced unless you ask for a normalised form, and the choice is applied to every folio in the run.

Related scripts and pages

Transcribe your masoretic folios

Upload a scan, choose Masoretic Hebrew Biblical Codices as the domain, and check one pointed column before running the whole manuscript.