Gujarati
Gujarati letters, read without the line that holds them together
Gujarati split off from the same Nagari root as Devanagari but dropped the continuous head line, so a line of print or handwriting is a row of separate floating letters rather than words strung along a bar. Evidano reads Gujarati print, business-ledger cursive and everyday handwriting, keeps its own numerals as written, and sets the transcription beside the image so every floating cluster can be checked against the page.
Sample pages
Real pages from public collections, shown beside their transcriptions. Pages whose transcription is still being checked are marked.

British Library, Or. 2116C, c. 1650 · Source · Public domain
Transcription being verified
No line-by-line published transcription of this folio’s interlinear Gujarati commentary has been located; the British Library’s catalogue describes the manuscript but does not transcribe it.
What the model is told to watch for: Produce a diplomatic transcription in Gujarati.

Transcription being verified
The interview was reprinted in Yahya Hashim Bawany’s Rare Speeches & Documents of Quaid-e-Azam (Karachi, 1987), a source still in copyright, so it cannot be used here; no public-domain transcription of this exact page has been found.
What the model is told to watch for: Produce a diplomatic transcription in Gujarati.

Transcription being verified
A self-published clipping from an unidentified local newspaper; no separate published transcription exists to check against.
What the model is told to watch for: Produce a diplomatic transcription in Gujarati.

Source · CC BY-SA 3.0 · Maulik Joshi
Transcription being verified
A one-off community noticeboard; no published transcription exists for this specific sign.
What the model is told to watch for: Produce a diplomatic transcription in Gujarati.

Source · Public domain · City Library copy scanned for the Gujarati Wikisource proofreading project
Transcription
The whole of printed page 116, proofread and then validated on Gujarati Wikisource (quality level 4 of 4) against this same scan. The running head and wiki markup are dropped and the centred italic file heading is kept as its own two lines; elsewhere the breaks are those the Wikisource text records, which follow the printed lines closely but not everywhere.
એને સરત ન રહી. સુજાનગઢના બંગલામાં પેસતાં જ એણે બુઢ્ઢા ચાઊસને પૂછી જોયું : “માલુજીને કેમ છે ?” “બુખાર હૈ.” “કેટલોક ?” “થોરા ! બિલકુલ કમતી, હાં સા’બ !” પોતે માલુજીના ખાટલા પાસે ગયો. માલુજી ઘેનમાં હતા : શિવરાજે ઊંઘ સમજી લીધી. “આરામ છે ? ઊંઘે છે ? તો તો ઊંઘવા દો, ચાઊસ !” “હાં સા’બ.” સૂવાના ઓરડા તરફ જતાં જતાં બાજુના ઓરડામાં, પિતાના લખવાના મેજ પર, પેલી નેતરની છાબડી પડી હતી એમાંનાં કાગળિયાંની થપ્પીના લાલ પટા પર શિવરાજની નજર ગઈ. ગાડીવાળાએ ત્યાં સાચવીને મૂકી હતી. સૂતાં પહેલાં વાંચી તો જોઉં — એમ વિચારીને શિવરાજ મેજ પર બેઠો. દોરી ખોલી અને પહેલી જ ફાઈલ હાથમાં લીધી. તેના ઉપરના શબ્દો જોઈ ચમક્યો : સરકાર વિ. બાઈ અજવાળી વાઘા. બાળહત્યા બાબત. કમિટ કરવાનો કેસ. અક્ષરો પર શિવરાજની નજર સ્થિર ન રહી શકી. અક્ષરો અક્ષર મટી ગયા; એમાંથી આંખો, મોં, હાથ, પગ વગેરે માનવાકૃતિનાં અંગો રચાવા લાગ્યાં. ‘બાળહત્યા’ શબ્દ ભાષામાં બહુ સહેલાઈથી પેસી ગયો છે. પણ બધી જ આંખો એને એટલી સહેલાઈથી વાંચી નથી શકતી. અજવાળી વાઘા : એ જ એ જ : એણે બાળહત્યા કરી ! શિવરાજના પ્રાણનું તળિયું સળવળ્યું. આજ દસ મહિનાથી એ ભયની ભૂતાવળ પોતે મનની ઊંડી બખોલોમાં જોયા કરી હતી. એને પલે પલે અસ્વસ્થ બનાવનાર એ એક જ ફફડાટ હતો. ગઈ સાંજની જાહેર સભામાં એના શબ્દો કંઠમાં આવીને પાછા વળી ગયા હતા તે આ ભયના જ ભણકારને આભારી હતું. અજવાળી આવી પણ ગઈ ? એનો મુકદ્દમો શું મારે ચલાવવાનો છે ? એણે પિતાના મેજ પર માતાની તસવીર દેખી. એ ઊઠી ગયો; સૂવાના ઓરડા તરફ ચાલ્યો. એની આંખે તમ્મર આવ્યાં. ઉંબરમાં જ એના શરીરે પડતું મેલ્યું. ધબકારો બહાર સંભળાયો. ચાઊસ અને રસોઈયો દોડતા આવ્યા. અચેતન શિવરાજને ઉપાડી પથારીમાં સુવરાવ્યો. દાક્તર આવ્યા, દવા કરી. “કશું નથી; માનસિક થાક છે.” કહીને એણે અર્ધજાગ્રત શિવરાજને હિંમત આપી. “સૂઈ જવું છે, દાક્તર !” શિવરાજે એટલું કહીને પડખું ફેરવ્યું. અને દાક્તરે એના કપાળ પર હાથ ફેરવીને કહ્યું : પંદર દિવસ સુધી પૂરો આરામ લેવો પડશે.” કાંપમાં ખબર પહોંચ્યા. પંડિતસાહેબ અને સરસ્વતી હાજર થયાં. “આપણે બધાંએ જઈને એની હવા નથી બગાડવી, બેટા ! તું એકલી જ જઈ આવ.” એમ કહી પંડિતસાહેબ બહારના રૂમમાં જ બેઠા. ઓરડામાં પહેલું પ્રવેશ્યું સરસ્વતીનું હાસ્ય ને પાછળ સરસ્વતીના શરીરે પગ મૂક્યો, એવું હરકોઈને લાગે. સૂતેલા શિવરાજનું થાકેલું સ્મિત એના કદમોમાં વેરાયું. સરસ્વતી
Transcription from પૃષ્ઠ:Aparadhi - Gujarati Novel (1938).pdf/૧૨૦, Gujarati Wikisource (CC BY-SA 4.0).
What Gujarati text looks like on the page
Gujarati letters descend from the same medieval Nagari hand that gave North India Devanagari, and as late as the sixteenth and seventeenth centuries the two were close enough that a single manuscript could carry a Nagari main text with an Old Gujarati commentary squeezed between the lines. What became distinctly Gujarati grew out of the fast cursive that merchants and clerks across Gujarat used for correspondence and accounts: writing quickly and without lifting the pen, they let the continuous horizontal head line that ties Devanagari letters into a word wear away, so that by the time Gujarati type was first cut for print in the 1790s the script had settled into the head-line-less form still used today. The result is the mirror image of Devanagari’s problem: instead of learning to find letter boundaries under one continuous bar, a reader — or a model — has to learn to group separate, unconnected letters into words using spacing and shape alone.
Below that missing line, Gujarati keeps most of the apparatus Devanagari has: vowel signs attached above, below, before or after a consonant, conjuncts formed by stacking or by half-forms of one consonant before another, and a visarga, anusvara and virama doing the same jobs they do further north. Gujarati also has its own set of digits, ૦ to ૯, whose shapes diverge from both the Devanagari numerals and the Hindu-Arabic ones, so a page of prices or dates needs to be read as Gujarati and not silently corrected into a more familiar numeral set. Everyday cursive — a personal letter, a shop’s account book, a note dashed off in the margin of a form — compresses vowel signs into small hooks and runs letters together far more than a school hand does, without ever gaining a line to hold them steady.
In print, Gujarati has carried newspapers, novels and government notices since Bombay’s first Gujarati paper appeared in 1822, and it remains one of the working languages of state administration in Gujarat, printed today alongside a great deal of English on forms, signs and notices. A modern signboard or notice is as likely to mix Gujarati and English in the same few lines as to appear in Gujarati alone, and a transcription has to keep the two apart rather than merging them into one run of text.
Why ordinary OCR struggles here
Letters that float free of each other
With no head line to string them together, a Gujarati word is a row of separate glyphs held apart only by spacing and shape. A model that expects Devanagari’s continuous bar has nothing to anchor on, and a cramped or uneven hand can make the gap inside a word look the same as the gap between two words.
A numeral set of its own
Gujarati’s digits, ૦–૯, are shaped differently from both the Devanagari numerals and Hindu-Arabic ones. A model trained mainly on one of those two number systems will often force a Gujarati digit into the nearest shape it already knows rather than reading the figure that is actually there.
The vahi ledger hand
Merchant account books, known as vahis, are written at speed in a rounded, looping cursive that shrinks vowel signs to small hooks and runs consecutive letters into one continuous stroke. It is legible to someone who reads it often and nearly opaque to anyone trained only on printed Gujarati.
English wedged into a Gujarati line
Signs, forms and even personal letters routinely drop an English word or abbreviation into an otherwise Gujarati sentence — a recreation named in brackets, a greeting chalked in Latin letters under a Gujarati notice. The transcription has to keep each script as its own run rather than transliterating one into the other.
Who works with this material
Family historians and trading-community archives
Gujarati and Kutchi trading families, in India and across the diaspora, often hold vahis and correspondence going back generations. A transcription that can be checked against the ledger page turns a locked family archive into something a genealogist can actually search.
Newspaper and periodical digitisation
Gujarati has been a newspaper language since 1822, and archives working through decades of runs need a transcription that keeps headlines, columns and photo captions distinct so a scanned page becomes a searchable article rather than one undivided block of text.
Literary and political archives
Letters and manuscripts by Gujarati-writing public figures — a category that includes political leaders far beyond Gujarat itself — need a diplomatic transcription that keeps embedded English words and phrases visibly separate from the surrounding Gujarati.
Institutions and trusts with old paper records
Temple trusts, schools and small businesses in Gujarat sit on decades of registers, notices and account books. Transcribing them page by page makes it possible to search a register for a name or a date without rereading the whole volume by eye.
Getting the best transcription
Say there is no head line to rely on
If a prompt is written with Devanagari in mind, say explicitly that Gujarati letters are not joined by a line, so the model groups them by spacing and shape rather than expecting a bar to segment words for it.
Set a policy for Gujarati numerals
Decide whether ૦–૯ should be kept as written or converted to standard digits, and say so in the prompt. Dates and prices are the most common place this matters, and the default should keep the figures exactly as printed unless you need them normalised.
Test a vahi or cursive page before a whole ledger
Business-ledger cursive compresses far more than a school or print hand does. Run the fastest, least formal page in a batch first, and if the model separates its rows and columns correctly there, the clearer pages that follow need only a light check.
Flag embedded English on signs and forms
Tell the model to keep English words, abbreviations or English-script headings in their own script rather than transliterating them into Gujarati letters or dropping them, since forms and signs frequently switch scripts mid-line.
Frequently asked questions
- How does the transcription cope with Gujarati having no head line?
- It reads each letter and vowel sign as a separate unit and groups them into words by spacing and shape rather than looking for a connecting bar, which is the opposite of how a Devanagari transcription has to work.
- Are Gujarati numerals kept as written or converted?
- Kept as written by default, since ૦–૯ are a distinct set of digits from both Devanagari and Hindu-Arabic numerals. Ask for them converted to standard digits if that suits your project better, and the conversion is applied consistently.
- Can it read a vahi or other business-ledger cursive?
- Yes, though a very compressed ledger hand benefits from being tested on one representative page first. Tell the model this is fast cursive rather than a formal print or book hand so it does not expect the clearer spacing of typeset Gujarati.
- What happens with English words inside a Gujarati sentence or sign?
- They are transcribed in their own script and kept in place rather than transliterated into Gujarati letters, so a bracketed English word or a Latin-script greeting on a sign comes back exactly as it appears.
- Can it tell Gujarati apart from Devanagari automatically?
- Yes — the absence of a head line and the distinct letterforms and numerals are enough to identify Gujarati reliably, even on a page that also carries some Devanagari or Sanskrit text, as in an older manuscript with an interlinear commentary.
Related scripts and pages
Transcribe your Gujarati pages
Upload a ledger, a letter, a newspaper clipping or a sign and check the Gujarati transcription against the image before you export it.
