Japanese · c. 700–1100
Read Nara and Heian columns as the scribe set them down
One syllable may be written with any of a dozen Chinese characters, and a single sheet can run cursive kana down one column and Chinese word order with reading marks down the next. Evidano works through these pages column by column, keeps the graph the scribe actually chose rather than flattening it to modern kana, and puts the reading beside the image so every character can be checked.
Sample pages
Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

Shōsōin Repository, Nara, 752 · Source · Public domain
Transcription being verified
The Shōsōin administrative papers are printed in full in Dai Nihon komonjo, hennen monjo (Tokyo Imperial University, from 1901); this sheet is being matched to the entry for the fourth year of Tenpyō-shōhō.
What the model is told to watch for: Expect Chinese characters used phonographically as man'yōgana alongside early hiragana, katakana, and kanbun annotation systems. Distinguish the source kanji or character form underlying each cursive kana, since several historically different graphs can represent the same modern syllable. Note kundoku marks, phonetic glosses, repetition signs, voiced-sound marks used inconsistently, vertical column order, and text distributed around illustrations or main Chinese passages.

Tokyo National Museum, Heian period, 9th century · Source · CC0 1.0
Transcription
The whole indigo panel, from 爾時毗沙門天王 in the first column to 㝹醯 where the segment breaks off, given in the standard text of Kumārajīva’s translation as printed on Chinese Wikisource. That text uses the modern variants 咒 and 毗 where the ninth-century copy writes 呪 and 毘, and prints 拘那覆 where the copy has 拘那履; two rare syllables it leaves as a box have been restored to 㝹 from the image. The small numerals beside the dhāraṇī phrases belong to the manuscript and are not reproduced below.
爾時毗沙門天王護世者白佛言:「世尊,我亦為愍念眾生、擁護此法師故,說是陀羅尼。」即說咒曰: 阿梨 那梨 㝹那梨 阿那盧 那履 拘那覆 「世尊,以是神咒、擁護法師,我亦自當擁護持是經者,令百由旬內、無諸衰患。」 爾時持國天王、在此會中,與千萬億那由他乾闥婆眾,恭敬圍繞,前詣佛所,合掌白佛言:「世尊,我亦以陀羅尼神咒、擁護持法華經者。」即說咒曰: 阿伽禰 伽禰 瞿利 乾陀利 旃陀利 摩蹬耆 常求利 浮樓莎柅 頞底 「世尊,是陀羅尼神咒,四十二億諸佛所說,若有侵毀此法師者,則為侵毀是諸佛已。」 爾時有羅剎女等,一名藍婆,二名毗藍婆,三名曲齒,四名華齒,五名黑齒,六名多發,七名無厭足,八名持瓔珞,九名睾帝,十名奪一切眾生精氣,是十羅剎女,與鬼子母、並其子、及眷屬,俱詣佛所,同聲白佛言:「世尊,我等亦欲擁護讀誦受持法華經者,除其衰患,若有伺求法師短者,令不得便。」即於佛前,而說咒曰: 伊提履 伊提泯 伊提履 阿提履 伊提履 泥履 泥履 泥履 泥履 泥履 樓醯 樓醯 樓醯 樓醯 多醯 多醯 多醯 兜醯 㝹醯
Transcription from 妙法蓮華經 (Lotus Sutra), trans. Kumārajīva, chapter 26 (Dhāraṇī), via Chinese Wikisource (Public domain).

Tokyo National Museum, 867 (Jōgan 9) · Source · Public domain · Tokyo National Museum
Transcription being verified
The Enchin surname-change documents are printed among the ninth-century material in Takeuchi Rizō (ed.), Heian ibun, komonjo-hen; a column-by-column reading of this section is being checked against that text.
What the model is told to watch for: Expect Chinese characters used phonographically as man'yōgana alongside early hiragana, katakana, and kanbun annotation systems. Distinguish the source kanji or character form underlying each cursive kana, since several historically different graphs can represent the same modern syllable. Note kundoku marks, phonetic glosses, repetition signs, voiced-sound marks used inconsistently, vertical column order, and text distributed around illustrations or main Chinese passages.

Gotoh Museum, Tokyo, c. 1050 · Source · Public domain
Transcription
The two poems on the sheet, Kokin wakashū 1 and 2, in the kana text on Japanese Wikisource: the headnote and the poet’s name, then the poem in the mixed kanji-kana form and again as an all-kana reading divided into its five metrical units by dashes. The kana line is the one that corresponds to what the calligrapher actually wrote. The sheet also carries the scroll title 古今倭歌集巻第一 and the section heading 春歌上, which stand outside the edited poem text.
[詞書]ふるとしに春たちける日よめる 在原元方 としのうちに春はきにけりひととせをこそとやいはむことしとやいはむ としのうちに-はるはきにけり-ひととせを-こそとやいはむ-ことしとやいはむ [詞書]はるたちける日よめる 紀貫之 袖ひちてむすひし水のこほれるを春立つけふの風やとくらむ そてひちて-むすひしみつの-こほれるを-はるたつけふの-かせやとくらむ
Transcription from 古今和歌集 巻一 (Kokin wakashū, book 1), via Japanese Wikisource (Public domain).

Gotoh Museum, Tokyo, 11th century · Source · Public domain
Transcription
The two poems on the fragment, Man’yōshū 4:640 and 4:641, from the Japanese Wikisource text: the Chinese-character headnote, the man’yōgana original as it stands in the manuscript tradition, the modern reading, and the all-kana reading that corresponds to the cursive gloss on the sheet. Angle brackets mark graphs the editors emended; this sheet writes 焼太刀 where the edited text has 焼大刀.
[題詞]湯原王亦贈歌一首 [原文]波之家也思 不遠里乎 雲<居>尓也 戀管将居 月毛不經國 [訓読]はしけやし間近き里を雲居にや恋ひつつ居らむ月も経なくに [仮名]はしけやし まちかきさとを くもゐにや こひつつをらむ つきもへなくに [題詞]娘子復報贈<歌>一首 [原文]絶常云者 和備染責跡 焼大刀乃 隔付經事者 幸也吾君 [訓読]絶ゆと言はばわびしみせむと焼大刀のへつかふことは幸くや我が君 [仮名]たゆといはば わびしみせむと やきたちの へつかふことは さきくやあがきみ
Transcription from 万葉集 第四巻 (Man’yōshū, book 4), poems 640–641, via Japanese Wikisource (Public domain).
What an eighth- to eleventh-century Japanese page looks like
Japanese began as a language written entirely in Chinese characters. A clerk drafting a requisition in the Tōdai-ji construction office and a compiler assembling the poems of the Man’yōshū both used kanji in two ways on the same sheet: for their meaning, arranged in Chinese word order, and for their sound alone, which is what modern scholarship calls man’yōgana. The choice of sound-character was loose, so 波, 者 and 半 could all stand for ha, and a personal name might be spelled three ways in one register. Columns are written top to bottom and ordered right to left, so a document opens at the top right corner and closes at the bottom left; nothing marks the boundary between one word and the next.
Across the ninth and tenth centuries the sound-characters were written faster and smaller until the underlying kanji dissolved: 安 flattened into あ, 以 into い, 知 into ち. Because scribes had started from different characters, several shapes survived for the same syllable — the variants later called hentaigana — and a single poem sheet can spell ka from 加, 可 and 我 within a few lines. Katakana grew the opposite way, from fragments broken off a character (伊 giving イ, 多 giving タ) and jotted between the columns of Chinese texts as an aid to reading them aloud in Japanese.
Those aids form a system of their own. Kaeriten tell the reader to jump back up the column to follow Japanese word order; small katakana at the lower right of a character supply the inflectional ending; and wokototen, dots placed at fixed positions around a character, stand in for particles without writing a single kana. Physical formats differ as sharply as the scripts. Sutra copies are ruled in columns of seventeen characters and may be written in gold or silver on indigo-dyed paper; state documents carry square vermilion office seals stamped straight across the writing and set explanatory clauses in two half-width columns inside one ruled space; poem sheets use dyed, mica-printed or collaged papers with painted underdrawing that crosses the brush strokes.
Why ordinary OCR struggles here
One sound, many characters
Recognising a shape is only half the problem: the reader also has to decide whether a character is being used for its meaning or only for its sound. Generic OCR trained on modern Japanese returns the kanji it recognises and stops there, so a line of man’yōgana comes back as a string of unrelated nouns instead of a poem.
No spaces and almost no punctuation
Word and clause boundaries are not written. A column of forty characters has to be segmented from the language itself, and the segmentation changes the reading: a run that can be split as a place name or as a verb plus particle will be read one way by a model that has seen Nara-period documents and another way by one that has not.
Two writing systems interleaved
A Buddhist or administrative page often carries a Chinese main text in full-size characters with Japanese reading marks added in the gutters at a fraction of the size. Detecting where the small marks begin, keeping them out of the main column, and recording which character each one attaches to are three separate jobs that ordinary text recognition treats as one stream.
Seals, decoration and dyed grounds
Vermilion office seals overlap the characters they authenticate, gold ink on indigo inverts the usual contrast, and a poem sheet may be painted with grasses and birds before a word is written on it. Binarisation designed for black type on white paper loses strokes on all three.
Who works with this material
Waka and Man’yōshū textual scholars
Collating the Katsura-bon, Genryaku kōhon and later copies means recording which sound-character each witness used, not just which word it spells. A transcription that preserves the graph as written lets variants be counted mechanically instead of being read off the plates by eye.
Historians of the ritsuryō state
The Shōsōin papers and provincial petitions are the census returns, tax accounts and personnel files of eighth- and ninth-century Japan. Researchers building name and place indexes from them need the small half-width annotations and the dates captured as reliably as the main clauses.
Buddhist studies and manuscript cataloguers
Dated colophons, donor names and vow texts at the end of a sutra scroll are what make a copy datable. Getting those columns out accurately, and flagging where a passage diverges from the received canonical text, is more useful than a clean run of the scripture itself.
Historical linguists and kana researchers
Anyone tracking how a syllable’s repertoire of graphs narrowed between the eighth and the eleventh century needs counts of which character was used where. That means transcriptions that record the source character of every kana rather than normalising it away.
Getting the best transcription
Decide what the output should be
A page of this date can be rendered three ways: the characters as written, the kana reading, or a modern normalised text. Ask for one as the main output and, if you need it, a second as a parallel line, so that later comparison work is not guessing which convention a file follows.
Say how reading marks should be treated
State whether kaeriten, okurigana and interlinear dots should be dropped, listed separately, or recorded with the character they attach to. If you want the Chinese text reordered into Japanese word order, ask for that as an extra field rather than in place of the column-order reading.
Number the columns
Ask for one output line per written column, numbered from the right. Poem sheets in kana start columns at deliberately uneven heights, and a numbered column list is the only way to check quickly that nothing has been read out of order.
Test on a decorated or sealed sheet first
Run a poem sheet with painted underdrawing or a document with seals stamped across the writing before you run the plain ones. If strokes survive the background and the seal text is kept apart from the document text, the undecorated pages will give no trouble.
Further reading on this hand
The palaeography guides the in-app selector points to for this domain, if you want to check a transcription against the standard references.
- Kuzushiji (Japanese Cursive Characters)
This National Diet Library guide introduces resources for reading kuzushiji and historical Japanese character forms. It explains hentaigana by connecting cursive kana to their source characters and provides charts linking historical forms with modern hiragana. The guide also points readers to dictionaries, fonts, databases, and instructional materials for manuscript and early printed texts. Its treatment of source characters is useful for early kana because one phonetic value may be represented by several graphically distinct man'yōgana-derived forms.
Frequently asked questions
- Can it tell man’yōgana from characters used for their meaning?
- That judgement is made from context, and it is the part worth reviewing. Ask for the phonographic runs to be marked in the output; a poem line will then show as a sequence of sound-characters with its reading beside it, while a headnote or a date stays as ordinary Chinese-character text.
- Does the transcription keep the original kana graph or give modern hiragana?
- Either, and both if you want them. The default for eleventh-century material is the modern kana equivalent with the source character noted, because that is what most editions print; ask instead for the written variant and it will be recorded consistently across the run.
- How are kaeriten and other reading marks handled?
- They are treated as a separate layer from the main column. You can have them omitted, collected as a list keyed to the character each one sits beside, or used to produce a second reading of the passage in Japanese word order alongside the column-order text.
- What happens where a vermilion seal covers the writing?
- The seal is read as its own object and its legend transcribed if it is legible. Characters underneath it are given where enough of the strokes show and marked as uncertain where they do not, rather than being silently completed from what the formula usually says.
- Are the numerals beside a dhāraṇī or in a tax register recognised?
- Yes. Small counting numerals written beside transliterated syllables, and the Chinese numerals used for household counts and dates, are transcribed in place; because they often sit in a smaller size beside the main column, it is worth asking for them on their own output line so they can be checked as a series.
Related scripts and pages
Transcribe your Nara and Heian sheets
Upload a scan, choose Nara and Heian Man’yōgana and Early Kana as the domain, and check one column list before running the scroll.
