Indic Transliteration: IAST, ISO 15919 and Regional Scripts

Short Answer

A single Sanskrit word may be written in several South Asian scripts while retaining essentially the same linguistic form. For example, the Sanskrit word dharma may appear as: Script Form Devanagari धर्म Bengali ধর্ম Gujarati ધર્મ Gurmukhi ਧਰਮ Kannada ಧರ್ಮ Malayalam ധർമ്മ Odia ଧର୍ମ Tamil தர்ம Telugu ధర్మ Sinhala ධර්ම Latin transliteration dharma This apparent […]

A single Sanskrit word may be written in several South Asian scripts while retaining essentially the same linguistic form.

For example, the Sanskrit word dharma may appear as:

Script Form
Devanagari धर्म
Bengali ধর্ম
Gujarati ધર્મ
Gurmukhi ਧਰਮ
Kannada ಧರ್ಮ
Malayalam ധർമ്മ
Odia ଧର୍ମ
Tamil தர்ம
Telugu ధర్మ
Sinhala ධර්ම
Latin transliteration dharma

This apparent correspondence is possible because most major scripts of South Asia descend from Brahmi and share a broadly similar structural model. Yet converting them into Latin letters is not as simple as matching symbols in parallel columns.

An Indic transliteration system must account for:

  • Inherent vowels
  • Dependent vowel signs
  • Consonant clusters
  • Virāma or halant signs
  • Aspiration
  • Retroflex consonants
  • Vowel length
  • Nasalization
  • Script-specific letters
  • Language-specific pronunciation
  • Historical and regional spelling conventions

The two most important scholarly frameworks are IAST, the International Alphabet of Sanskrit Transliteration, and ISO 15919, the international standard for transliterating Devanagari and related Indic scripts.

They overlap substantially, but they are not identical.

IAST was developed primarily for Sanskrit and related classical languages. ISO 15919 was created to support a wider range of modern and classical languages across multiple Indic scripts, including distinctions that Sanskrit-focused IAST does not always need.

Quick Answer

IAST is the established scholarly transliteration system for Sanskrit and closely related classical languages. ISO 15919 is a broader international standard designed to transliterate multiple Indic scripts and languages consistently into Latin characters.

The principal difference is scope.

IAST works especially well for the Sanskrit sound system:

  • a, ā
  • i, ī
  • u, ū
  • ṛ, ṝ
  • e, ai
  • o, au
  • ṅ, ñ, ṇ, n, m
  • ś, ṣ, s
  • ṭ, ṭh, ḍ, ḍh

ISO 15919 extends this model to cover regional distinctions such as:

  • Short and long e/ē
  • Short and long o/ō
  • Additional Dravidian consonants
  • Script-specific vowel qualities
  • Expanded vocalic liquids
  • Letters used in Assamese, Tamil, Malayalam, Sinhala and other regional systems

ISO 15919:2001 covers Devanagari, Bengali—including characters used for Assamese—Gujarati, Gurmukhi, Kannada, Malayalam, Odia, Sinhala, Tamil and Telugu. The 2001 edition was reviewed and confirmed in 2022 and remained the published edition in July 2026, while a replacement draft was under development.

IAST vs ISO 15919

Feature IAST ISO 15919
Full name International Alphabet of Sanskrit Transliteration Transliteration of Indic scripts into Latin characters
Main focus Sanskrit and related classical languages Multiple classical and modern Indic languages
Script scope Commonly Devanagari and other scripts used for Sanskrit Devanagari plus major regional Indic scripts
Language scope Primarily Sanskrit, Pali and related philological work Indo-Aryan, Dravidian and other regional languages
Retroflex letters ṭ, ṭh, ḍ, ḍh, ṇ Same basic forms
Sanskrit vocalic r ṛ, ṝ Commonly represented structurally as r̥, r̥̄
Sanskrit vocalic l ḷ or related conventional form l̥, l̥̄
Sanskrit e and o e, o Usually ē, ō when length must be explicit
Short regional e and o Not central to traditional Sanskrit IAST e, o
Long regional e and o Often not distinguished from Sanskrit conventions ē, ō
Regional letters Limited Expanded coverage
Main use Sanskrit editions, dictionaries and scholarship Standards, multilingual data and cross-script conversion

The two systems may produce almost identical results for ordinary Sanskrit words, but differences emerge with vocalic liquids and with languages that distinguish short and long e and o.

What Are Indic Scripts?

The term Indic scripts usually refers to the Brahmi-derived writing systems of South Asia.

Major examples include:

  • Devanagari
  • Bengali or Bangla
  • Assamese
  • Gujarati
  • Gurmukhi
  • Odia
  • Tamil
  • Telugu
  • Kannada
  • Malayalam
  • Sinhala

Related Brahmi-derived traditions extend much farther, including historical scripts of India and many scripts of Tibet and Southeast Asia.

Unicode describes the principal Indic systems as abugidas. Their effective written unit is an orthographic syllable, usually built around a consonant and vowel core. A consonant normally carries an inherent vowel unless another vowel sign replaces it or a virāma suppresses it.

For example, the Devanagari letter क is not merely k. By itself, it normally represents ka.

Devanagari Analysis Transliteration
k + inherent a ka
का k + ā vowel sign
कि k + i vowel sign ki
कु k + u vowel sign ku
के k + e vowel sign ke
क् k + virāma k

This structural principle occurs throughout the major Indic script family, although the exact inherent vowel and its treatment differ by language.

Unicode notes that the major scripts of India share a parallel encoding arrangement derived from ISCII. This correspondence facilitates script mapping, but Unicode warns that implementations must not assume that every regional script behaves exactly like Devanagari.

What Is IAST?

IAST stands for the International Alphabet of Sanskrit Transliteration.

It is the conventional scholarly method for representing Sanskrit in Latin letters while preserving important distinctions in the original text.

Harvard’s South Asian library guidance identifies IAST as one of the principal systems readers encounter and notes that it is mostly—but not completely—identical to related library schemes.

IAST uses Latin letters with diacritics to distinguish sounds and source characters that would otherwise collapse into the same basic spelling.

For example:

Devanagari IAST
ta
ṭa
da
ḍa
na
ṇa
sa
śa
ṣa

Without the diacritics, three different sibilants would all become s, while dental and retroflex consonants would become indistinguishable.

IAST Vowels

The standard Sanskrit vowel inventory is commonly represented as:

Devanagari IAST
a
ā
i
ī
u
ū
e
ai
o
au

The macron distinguishes long vowels:

  • a vs ā
  • i vs ī
  • u vs ū

The underdot marks Sanskrit vocalic r and l:

In traditional Sanskrit phonology, e and o are treated as inherently long, so ordinary IAST does not normally write them as ē and ō.

That convention becomes insufficient when the system is extended to Dravidian languages, where short and long e and o can contrast.

IAST Consonants

IAST follows the traditional Indic phonetic ordering.

Velars

Devanagari IAST
ka
kha
ga
gha
ṅa

Palatals

Devanagari IAST
ca
cha
ja
jha
ña

Retroflexes

Devanagari IAST
ṭa
ṭha
ḍa
ḍha
ṇa

Dentals

Devanagari IAST
ta
tha
da
dha
na

Labials

Devanagari IAST
pa
pha
ba
bha
ma

Semivowels and liquids

Devanagari IAST
ya
ra
la
va

Sibilants and aspirate

Devanagari IAST
śa
ṣa
sa
ha

The current ALA-LC Sanskrit and Prakrit table uses this familiar Sanskrit-oriented inventory, including ā, ī, ū, ṅ, ñ, ṭ, ḍ, ṇ, ś, ṣ, ṃ and . It also directs catalogers to use the equivalent letters of another script when Sanskrit is written outside Devanagari.

What Is ISO 15919?

ISO 15919 is an international standard for transliterating Indic scripts into Latin characters.

Its purpose is broader than IAST. It must accommodate not only Sanskrit but also multiple Indo-Aryan and Dravidian languages written in structurally related scripts.

The current published title is:

ISO 15919:2001 — Information and documentation — Transliteration of Devanagari and related Indic scripts into Latin characters

Its coverage includes:

  • Devanagari
  • Bengali, including Assamese characters
  • Gujarati
  • Gurmukhi
  • Kannada
  • Malayalam
  • Odia
  • Sinhala
  • Tamil
  • Telugu

Why ISO 15919 Extends IAST

IAST was optimized around Sanskrit’s phonological and orthographic requirements.

Regional languages introduce contrasts that require additional forms.

For example, Kannada, Malayalam, Tamil and Telugu distinguish:

  • Short e
  • Long ē
  • Short o
  • Long ō

The current ALA-LC regional tables reflect these distinctions directly:

Script Short e Long e Short o Long o
Kannada ಎ e ಏ ē ಒ o ಓ ō
Malayalam എ e ഏ ē ഒ o ഓ ō
Tamil எ e ஏ ē ஒ o ஓ ō
Telugu ఎ e ఏ ē ఒ o ఓ ō

Traditional Sanskrit IAST writes Sanskrit ए and ओ as e and o, because Sanskrit does not use corresponding short vowel phonemes in the same way.

ISO 15919 avoids this cross-language conflict by allowing:

  • e for short e
  • ē for long e
  • o for short o
  • ō for long o

Consequently, Sanskrit text under strict ISO 15919 may use ē and ō, while conventional IAST normally uses e and o.

Vocalic R and L

Another familiar difference concerns vocalic liquids.

Traditional IAST commonly uses:

ISO 15919 uses combinations that make the vowel component explicit:

  • r̥̄
  • l̥̄

These forms can be represented as a base consonant plus a combining ring below and, where necessary, a macron.

For general Sanskrit readers, is usually more familiar than . For a cross-script standard, however, presents the symbol as a vocalic version of r and supports a more regular system.

Extended Regional Consonants

ISO 15919 also accommodates consonants not central to classical Sanskrit, including:

  • Retroflex lateral sounds
  • Alveolar or trill-like regional consonants
  • Dravidian approximants
  • Extended letters used in modern Indo-Aryan languages
  • Script-specific nasal and vowel signs

Examples include:

  • Modified letters used for Persian and Arabic loanwords

The exact form depends on the script, language and standard table being followed.

IAST and ISO 15919 Examples

For many Sanskrit words, ordinary IAST and ISO 15919 appear nearly identical.

Devanagari Meaning IAST Typical ISO 15919 distinction
धर्म dharma dharma dharma
योग yoga yoga yōga
संस्कृत Sanskrit saṃskṛta saṃskr̥ta
कृष्ण Krishna kṛṣṇa kr̥ṣṇa
राम Rama rāma rāma
गुरु teacher guru guru
शान्ति peace śānti śānti
वेद Veda veda vēda

The ISO forms in this table illustrate the systematic marking of Sanskrit long e/o and vocalic r. Actual editorial implementations sometimes retain familiar IAST-style forms, so a platform should identify whether it is generating strict ISO 15919 or a compatible scholarly variant.

Transliteration Is Not Pronunciation

A major source of confusion is the assumption that a Latin transliteration automatically shows how a modern language is spoken.

It may not.

Strict transliteration represents written structure. Modern pronunciation may delete or modify sounds that remain visible in the script.

Inherent Vowel and Schwa Deletion

A Devanagari consonant normally carries an inherent vowel unless it has another vowel sign or a virāma. Unicode defines this as a basic feature of the Indic abugida model.

The Hindi ALA-LC table consequently supplies a after consonants unless another vowel or a halant blocks it.

Consider भारत:

  • Orthographic transliteration: Bhārata
  • Common Hindi pronunciation-oriented form: Bhārat

The final written consonant त has no visible halant, so a strict script transliteration supplies its inherent a. Modern Hindi pronunciation generally drops that final schwa.

Similar differences occur in names and ordinary vocabulary.

A tool must distinguish:

  • Orthographic transliteration
  • Standard-language transcription
  • Established personal or geographic spelling

Different Languages, Same Script

Devanagari is used for Sanskrit, Hindi, Marathi, Nepali and numerous other languages. Unicode lists many languages using the script, and the Library of Congress maintains separate Sanskrit, Hindi and Marathi tables.

The sequence should not automatically be interpreted according to Sanskrit merely because it is in Devanagari.

For example:

  • Sanskrit may preserve every written inherent vowel.
  • Hindi applies systematic schwa deletion in pronunciation.
  • Marathi has its own pronunciation and lexical conventions.
  • Nepali may treat certain consonants and vowels differently.

The Hindi table explicitly says that when another language with its own romanization table appears in Devanagari, the table for that language should be used.

Core Diacritics in Indic Transliteration

Diacritics are not decorative. They preserve contrasts between different letters and sounds.

Macron

The macron marks vowel length:

  • ā
  • ī
  • ū
  • ē
  • ō

Examples:

Unmarked Marked Distinction
a ā short vs long a
i ī short vs long i
u ū short vs long u
e ē short vs long e in relevant languages
o ō short vs long o in relevant languages

Dot Below

The dot below commonly marks retroflex or related consonants:

These are separate from:

  • t
  • d
  • n
  • ś or s
  • l

Removing the dot can merge distinct words and source characters.

Dot Above

A dot above is used in forms such as:

  • ṅ for the velar nasal
  • ṁ in some conventions for nasalization or anusvāra

IAST and related systems more commonly use —m with a dot below—for anusvāra.

Tilde

The palatal nasal is:

  • ñ

It is distinct from:

  • n

Acute Accent

An acute may appear in specialized transliteration for:

  • Vedic pitch
  • Script-specific vowel distinctions
  • Particular regional or library conventions

Vedic transliteration may require several additional accents beyond ordinary IAST.

Ring Below

ISO-style vocalic liquids may use:

A following macron may indicate the long form:

  • r̥̄
  • l̥̄

Combining Marks and Unicode

A visibly identical transliterated letter may be encoded either as:

  • One precomposed Unicode character
  • A base letter followed by one or more combining marks

For example, can be encoded as a precomposed character, while r̥̄ generally requires combining marks.

Applications should normalize Unicode consistently. Search, sorting and duplicate detection can fail when two visually identical strings use different underlying character sequences.

Regional Script Differences

Devanagari

Devanagari is used for Sanskrit, Hindi, Marathi, Nepali and other languages.

Important features include:

  • An inherent vowel
  • Dependent vowel signs
  • Conjunct consonants
  • Virāma or halant
  • Nukta-modified letters
  • Anusvāra
  • Candrabindu
  • Visarga
  • Avagraha
  • Vedic signs

Unicode explains that Devanagari consonant clusters may have contextual and ligature forms, while the stored characters remain in logical phonetic order.

Modern Hindi uses nukta letters for sounds common in Persian and Arabic loanwords:

Devanagari Romanization
क़ q
ख़ kh or x in some systems
ग़ gh or ġ
ज़ z
फ़ f
ड़
ढ़ ṛh

The current ALA-LC Hindi table includes these extended letters and was issued as a 2025 version.

Bengali or Bangla

The Bengali script is used primarily for Bangla and also forms the basis of the Assamese script tradition.

Its transliteration includes familiar forms such as:

  • ā, ī, ū
  • ṅ, ñ, ṇ
  • ṭ, ḍ
  • ś
  • ṃ, ḥ

However, some letters have language-specific behavior. The Bengali letter ব functions as a labial and may also have a semivowel role in clusters. The ALA-LC Bengali table romanizes it as va when it appears as a later consonant in certain clusters.

Modern Bangla pronunciation often differs substantially from Sanskrit-oriented spelling. A strict transliteration may preserve etymological distinctions that ordinary pronunciation has merged.

For instance, written শ, ষ and স may have similar or overlapping pronunciations in modern Bangla, but scholarly transliteration can preserve them as:

  • ś
  • s

Assamese

Assamese shares most of its script with Bengali but contains distinctive letters and phonological values.

The ALA-LC Assamese table includes separate semivowel and liquid forms corresponding to Assamese ৰ and ৱ, in addition to the broader Bengali-derived inventory.

A converter should not label Assamese text as Bengali merely because the Unicode blocks and visible letter shapes substantially overlap.

Gujarati

Gujarati shares the broad northern Brahmic structure but has a visually distinct script without the continuous headline associated with Devanagari.

Its ALA-LC table includes:

  • a, ā, i, ī, u, ū
  • e, ê, ai
  • o, ô, au
  • ṭ, ḍ, ṇ
  • ś
  • ṃ and ḥ

Gujarati language pronunciation, especially inherent-vowel behavior, should not be inferred solely from Sanskrit.

Gurmukhi

Gurmukhi is used primarily for Panjabi in India.

Its structure differs from classical Sanskrit-oriented systems in several practical respects:

  • Vowel signs are organized through bearer letters.
  • Tonal contrasts are important in modern Panjabi.
  • The script includes extended dotted letters used for Persian-Arabic loan sounds.
  • The sign called adhak doubles the following consonant.

The ALA-LC Panjabi table states that determining whether the inherent a should be supplied can require knowledge of the language or suitable reference sources. It also specifies that the adhak doubles the following consonant.

This shows why a purely visual character converter may be insufficient.

Odia

The Odia script is used primarily for the Odia language and related regional languages.

The current ALA-LC table is the 2024 version and includes:

  • r̥ and r̥̄
  • l̥ and l̥̄
  • ṛ and ṛh
  • ṃ, ḥ and nasalization signs

The table also notes that the official English names Odia and Odisha replaced Oriya and Orissa through constitutional amendments in 2011.

Historical catalogs may therefore retain older terminology even when current references use Odia.

Tamil

Tamil presents one of the greatest challenges for a Sanskrit-derived one-to-one transliteration model.

The native Tamil consonant inventory is smaller than the Sanskrit inventory. A single Tamil letter may correspond to several phonetic realizations depending on position and linguistic context.

For example, க may be pronounced with values approximating:

  • k
  • g
  • a fricative sound in certain environments

A strict letter transliteration usually represents the source letter consistently rather than attempting to show each contextual pronunciation.

The ALA-LC Tamil table distinguishes:

  • e and ē
  • o and ō
  • Tamil ழ and ள
  • Additional letters used for Sanskrit sounds, including ஜ, ஶ, ஷ, ஸ and ஹ

It also identifies the puḷḷi as the sign suppressing the inherent vowel.

A pronunciation-oriented Tamil transcription therefore requires rules beyond IAST character mapping.

Telugu

Telugu retains a broad consonant inventory that maps readily to Sanskrit-derived categories while also distinguishing short and long e and o.

The current ALA-LC table includes:

  • e and ē
  • o and ō
  • ṭ, ḍ and ṇ
  • ś, ṣ and s
  • Script-specific historical letters
  • Nasal signs and visarga

Telugu vowel signs can display in positions that do not correspond visually to their logical character order, so software should transliterate the encoded sequence rather than reading glyphs from left to right.

Kannada

Kannada also distinguishes:

  • e vs ē
  • o vs ō
  • l vs ḷ
  • Additional historical lateral and rhotic forms

The ALA-LC Kannada table uses and a separate marked lateral form for an additional letter. It also treats the inherent vowel and vowel-suppression sign explicitly.

Kannada and Telugu are closely related historically and structurally, but their script characters must still be interpreted through the appropriate table.

Malayalam

Malayalam contains several script features requiring specialized handling:

  • Short and long e/o vowels
  • Retroflex lateral and approximant letters
  • Distinct alveolar consonants
  • Chillu consonants
  • Complex consonant clusters
  • Traditional and reformed orthographic styles

The ALA-LC Malayalam table treats chillu forms as consonants without an inherent vowel. It separately lists several modified consonantal forms that prevent the automatic addition of a.

A converter that automatically appends a after every Malayalam consonant will therefore produce incorrect results for chillus.

Sinhala

Sinhala is structurally related to the Indic script family but has a significantly expanded vowel and consonant inventory.

The ALA-LC Sinhalese table includes:

  • a and ā
  • ă and â
  • e and ē
  • o and ō
  • Sanskrit-derived consonant series
  • Sinhala-specific nasal constructions
  • Saññaka-marked forms

ISO 15919 includes Sinhala because a pan-Indic standard must represent these regional distinctions rather than assuming a Sanskrit-only inventory.

Sanskrit Can Be Written in Many Scripts

Sanskrit is strongly associated with Devanagari today, but it has historically been written in numerous regional scripts.

These include:

  • Bengali
  • Grantha
  • Gujarati
  • Kannada
  • Malayalam
  • Odia
  • Sharada
  • Sinhala
  • Tamil
  • Telugu
  • Newar scripts
  • Nandinagari
  • Siddham

The ALA-LC Sanskrit and Prakrit table says that when Sanskrit is written in another script, the corresponding letters in that script should be transliterated according to the Sanskrit table.

This means the transliteration should normally reflect the Sanskrit language value, not an unrelated modern regional pronunciation.

For example, a Sanskrit manuscript written in Malayalam script should be transliterated as Sanskrit rather than as ordinary Malayalam speech.

The workflow should therefore be:

  1. Identify the script.
  2. Identify the language.
  3. Determine the intended reading.
  4. Apply the appropriate language-level transliteration system.
  5. Preserve the original script.

Inherent Vowels and Virāma

The inherent vowel is one of the most important concepts in Indic transliteration.

A consonant symbol normally includes an implied vowel:

  • क = ka
  • க = ka
  • ಕ = ka
  • ക = ka
  • క = ka

The exact phonetic realization of that vowel may vary by language.

A virāma, halant, puḷḷi, hasanta or equivalent sign suppresses the vowel:

  • क् = k
  • க் = k
  • ಕ್ = k
  • ക് = k
  • క్ = k

Unicode describes the virāma as a combining sign used to suppress the inherent vowel. The name differs across languages: Hindi commonly uses halant, while Tamil uses puḷḷi.

Consonant Clusters

When a consonant with a suppressed vowel is followed by another consonant, the script may form:

  • A ligature
  • A stacked form
  • A half-form
  • A reduced sign
  • A visible virāma sequence

For example:

  • क् + ष = क्ष
  • क् + र = क्र
  • त् + र = त्र
  • श् + र = श्र

The Latin transliteration should reflect the underlying consonant sequence:

  • kṣa
  • kra
  • tra
  • śra

It should not attempt to transliterate the visual ligature as an indivisible symbol.

Nasal Signs

Indic scripts use several nasal signs whose interpretation depends on context.

Anusvāra

Anusvāra is commonly transliterated as:

However, some systems replace it with the homorganic nasal before a following consonant.

For example:

Context Nasal
Before velars
Before palatals ñ
Before retroflexes
Before dentals n
Before labials m

The ALA-LC Sanskrit, Hindi, Bengali, Gujarati, Kannada, Malayalam, Odia, Telugu and Sinhala tables include contextual rules of this type.

A tool should state whether it preserves the written anusvāra as or expands it according to pronunciation or cataloging rules.

Candrabindu and Nasalized Vowels

Candrabindu commonly marks vowel nasalization.

Possible transliterations include:

  • A vowel with a tilde in certain transcription systems

The precise output depends on the standard and phonetic context.

Visarga

Visarga is generally transliterated:

Its pronunciation may vary by phonetic context, but strict transliteration preserves the written sign.

Retroflex and Dental Consonants

Indic languages frequently distinguish dental and retroflex consonants.

Dental Retroflex
t
th ṭh
d
dh ḍh
n

These distinctions are essential.

For example:

  • त and ट are different letters.
  • द and ड are different letters.
  • न and ण are different letters.

Removing the dots may make several source spellings indistinguishable.

The terms cerebral and retroflex both appear in romanization references. Modern linguistics generally favors retroflex, while many older library tables retain the traditional heading cerebral.

Aspirated Consonants

Indic transliteration uses h to distinguish aspirated consonants:

  • k vs kh
  • g vs gh
  • c vs ch
  • j vs jh
  • ṭ vs ṭh
  • ḍ vs ḍh
  • t vs th
  • d vs dh
  • p vs ph
  • b vs bh

These are not ordinary English consonant-plus-h clusters.

For example:

  • th represents an aspirated dental stop, not the English sounds in thin or this.
  • ph represents an aspirated p, not necessarily an English f sound.

Modern languages may pronounce some historical aspirated consonants differently, but a scholarly transliteration generally follows the script distinction.

Transliteration vs Modern-Language Romanization

IAST and ISO 15919 are designed for relatively precise transliteration. Everyday romanized writing in South Asia is often much less systematic.

A name might appear in:

  • A full scholarly form
  • A national passport form
  • A simplified English spelling
  • A regional conventional spelling
  • An informal mobile-keyboard spelling

Examples include:

Native form Scholarly form Common form
कृष्ण Kṛṣṇa Krishna
शिव Śiva Shiva
रामायण Rāmāyaṇa Ramayana
महाभारत Mahābhārata Mahabharata
तिरुवनन्तपुरम् Tiruvanantapuram or a technical regional form Thiruvananthapuram
ਪੰਜਾਬ Pañjāba or Panjāb by system Punjab
தமிழ் Tamiḻ Tamil

The simplified forms are not necessarily incorrect. They serve a different purpose.

A personal or official name should normally follow the form used by the individual or authority, while the scholarly transliteration can be shown separately.

ASCII Transliteration Systems

Before Unicode diacritics became widely supported, several ASCII systems were developed for Sanskrit and Indic texts.

These include:

  • Harvard-Kyoto
  • ITRANS
  • Velthuis
  • SLP1
  • WX notation

Typical substitutions might use:

IAST Possible ASCII representation
ā A or aa
ī I or ii
ū U or uu
R, R^i or .r
T or .t
D or .d
N or .n
ś z, sh or “s
S or .s
M or .m

These systems remain useful for:

  • Legacy text collections
  • Keyboard entry
  • Computational linguistics
  • Source code
  • Search aliases
  • Systems with restricted character support

However, an unlabeled ASCII form can be ambiguous. S, z, R, or N may mean different things in different schemes.

A platform should always name the ASCII system rather than calling it simply “Roman Sanskrit.”

Reversibility

A transliteration is reversible when the original script sequence can be reconstructed from the Latin text.

IAST can be highly reversible for properly normalized Sanskrit, provided that:

  • All diacritics are retained
  • The language is known
  • Sandhi does not obscure source segmentation
  • The source uses a compatible character inventory
  • Editorial normalization has not changed the spelling

ISO 15919 is better suited to reversible conversion across multiple regional scripts because it expands the inventory.

Even so, the Latin text may not reveal the original script.

For example, dharma could represent equivalent text originally written in:

  • Devanagari
  • Bengali
  • Gujarati
  • Kannada
  • Malayalam
  • Odia
  • Tamil
  • Telugu
  • Sinhala

A language-level transliteration may preserve the linguistic sequence while losing the identity of the source script.

For true round-trip conversion, a database should retain:

  • Original script
  • Script code
  • Language
  • Transliteration standard
  • Standard version
  • Normalization form
  • Any editorial changes

Choosing Between IAST and ISO 15919

Use IAST When:

  • Publishing Sanskrit terms
  • Preparing Sanskrit editions
  • Citing classical Indian texts
  • Teaching Sanskrit
  • Building Sanskrit dictionaries
  • Working with Pali in an IAST-compatible tradition
  • Writing for readers familiar with Indology

Examples:

  • Bhagavadgītā
  • Mahābhārata
  • Upaniṣad
  • Śaṅkara
  • Kṛṣṇa
  • saṃskṛta

Use ISO 15919 When:

  • Converting multiple Indic scripts
  • Supporting both Indo-Aryan and Dravidian languages
  • Preserving short and long e/o distinctions
  • Building multilingual databases
  • Creating reversible script-conversion tools
  • Normalizing regional-script data
  • Designing international information systems

Use a Regional or Institutional Table When:

  • Cataloging library materials
  • Following a university or publisher style
  • Processing one modern language
  • Handling regional letters not fully described by a simplified IAST chart
  • Matching established bibliographic records

The Library of Congress currently maintains separate tables for Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Pali, Panjabi, Sanskrit and Prakrit, Sinhala, Tamil and Telugu. The table list was updated on July 24, 2026.

Use an Official or Conventional Form When:

  • Recording personal names
  • Recording geographic names
  • Displaying institutions
  • Processing passports
  • Publishing contemporary news
  • Preserving brand identities

A mechanically generated scholarly spelling should not overwrite an established identity form.

Best Practices for an Indic Transliteration Tool

Preserve the Original Script

The original text should always remain available.

Identify Both Language and Script

Do not assume:

  • All Devanagari is Sanskrit
  • All Bengali-script text is Bangla
  • All Gurmukhi text follows Sanskrit pronunciation
  • All Malayalam-script text is Malayalam rather than Sanskrit
  • All Tamil-script Sanskrit should be interpreted as ordinary Tamil

Name the Exact System

Use labels such as:

  • IAST
  • Strict ISO 15919
  • ALA-LC Sanskrit and Prakrit 2012
  • ALA-LC Hindi 2025
  • ALA-LC Bengali 2017
  • ALA-LC Odia 2024
  • Simplified ASCII
  • Harvard-Kyoto
  • ITRANS

Show the Inherent Vowel Policy

State whether the output:

  • Preserves every written inherent vowel
  • Applies modern-language schwa deletion
  • Uses dictionary pronunciation
  • Follows an official name
  • Represents Sanskrit morphology

Separate Transliteration and Pronunciation

A useful result might include:

Field Example
Original भारत
Script Devanagari
Language Hindi
Orthographic transliteration Bhārata
Standard Hindi-oriented form Bhārat
Common English form Bharat
IPA or pronunciation Separate field

Preserve Diacritics by Default

Show:

  • Kṛṣṇa, not only Krishna
  • Śiva, not only Shiva
  • Tamiḻ, not only Tamil
  • Pañjāb, not only Punjab

A simplified version can be offered separately.

Support Unicode Normalization

Normalize combining-mark sequences consistently and test:

  • Search
  • Sorting
  • Copy and paste
  • Database uniqueness
  • URL generation
  • Font rendering

Warn When a Result Is Ambiguous

Warnings may include:

  • Language not identified
  • Schwa deletion uncertain
  • Multiple lexical readings
  • Historical spelling
  • Script-specific distinction lost
  • Diacritics removed
  • Informal name overrides systematic form

Common Misconceptions

IAST and ISO 15919 Are Identical

They overlap heavily for Sanskrit but differ in scope and in some letter conventions, especially vocalic liquids and short-versus-long e/o.

IAST Is Only for Devanagari

IAST represents Sanskrit and related linguistic forms, not one visual script. Sanskrit written in Bengali, Malayalam, Telugu or another compatible script can also be transliterated into IAST.

Every Consonant Is Followed by a Pronounced A

The script contains an inherent vowel, but modern languages may suppress it in pronunciation even when no virāma appears.

Transliteration Shows Exact Pronunciation

Strict transliteration follows writing. Regional pronunciation, schwa deletion and sound changes may require a separate transcription.

Diacritics Are Optional Decoration

Diacritics distinguish different vowels and consonants. Removing them can destroy reversibility.

Tamil Can Be Converted With a Sanskrit Table Alone

Tamil has a different phonological and orthographic structure, including short and long e/o, script-specific consonants and context-dependent pronunciations.

All Brahmi-Derived Scripts Map Perfectly One to One

They share structural ancestry, but each script has additions, omissions and language-specific behavior. Unicode explicitly warns against assuming that South Indian scripts operate exactly like Devanagari.

Romanized Personal Names Should Be Corrected to IAST

Official and preferred identity spellings take priority. Scholarly transliteration can be offered as additional information.

Frequently Asked Questions

What does IAST stand for?

IAST stands for International Alphabet of Sanskrit Transliteration.

What is ISO 15919?

ISO 15919 is an international standard for transliterating Devanagari and related Indic scripts into Latin characters. The currently published edition is ISO 15919:2001.

Is ISO 15919 still current?

The 2001 edition was reviewed and confirmed in 2022 and remained the published edition in July 2026. ISO also had a replacement draft under development.

What is the main difference between IAST and ISO 15919?

IAST is optimized for Sanskrit. ISO 15919 covers a wider range of Indic scripts and languages, including regional vowel and consonant distinctions.

Why does IAST use ṛ while ISO may use r̥?

IAST treats vocalic r as a conventional precomposed scholarly letter. ISO’s form displays it more analytically as r with a ring below.

Why does ISO use ē and ō for Sanskrit?

ISO must distinguish long Sanskrit e/o from the short e/o found in several regional languages. Traditional IAST does not need that distinction within Sanskrit.

What is the difference between ṭ and t?

is retroflex. t represents the dental consonant in Sanskrit-oriented transliteration.

What does ṃ represent?

generally represents anusvāra, a nasal sign whose phonetic value may depend on the following consonant.

What does ḥ represent?

represents visarga.

Why is Krishna written Kṛṣṇa?

Kṛṣṇa preserves the vocalic ṛ, retroflex ṣ and retroflex ṇ of कृष्ण. Krishna is a simplified conventional English spelling.

Is Bharat or Bhārata correct?

They serve different purposes. Bhārata is an orthographically explicit Sanskrit-style transliteration. Bhārat reflects the common Hindi form. Bharat is a simplified official and public spelling.

Can IAST represent Tamil?

It can approximate many Tamil letters, but traditional Sanskrit IAST lacks some distinctions required for complete Tamil transliteration. ISO 15919 or a dedicated Tamil table is more appropriate.

Can one tool transliterate every Indic script?

Yes, provided it identifies the script and language, supports the complete character inventory, handles inherent vowels correctly and declares the chosen standard.

Should long vowels always keep macrons?

Keep them in the complete scholarly or standard form. A simplified unmarked form may be generated separately for search or technical use.

Is transliteration reversible after diacritics are removed?

Usually not reliably. Distinct letters such as t/ṭ, d/ḍ, n/ṇ, s/ś/ṣ and a/ā may collapse.

Conclusion

Indic transliteration is built on a shared structural inheritance, but it must respect the diversity of South Asia’s languages and scripts.

IAST provides the familiar scholarly alphabet used for Sanskrit and related classical traditions. It preserves distinctions such as:

  • a and ā
  • i and ī
  • u and ū
  • t and ṭ
  • d and ḍ
  • n, ṅ, ñ and ṇ
  • s, ś and ṣ
  • r and ṛ

ISO 15919 extends this framework across the major Indic scripts. It accommodates regional distinctions such as short and long e/o, expanded Dravidian consonants, script-specific signs and language inventories beyond classical Sanskrit.

Regional scripts introduce additional complexities:

  • Devanagari requires language-sensitive handling of inherent vowels.
  • Bengali and Assamese share a script tradition but differ linguistically.
  • Gurmukhi may require language knowledge to determine implicit vowels.
  • Tamil has a smaller native consonant inventory and context-sensitive pronunciation.
  • Malayalam contains chillu consonants.
  • Kannada and Telugu distinguish short and long e/o.
  • Sinhala contains additional vowels and nasal structures.
  • Odia, Gujarati and other scripts have their own extended characters.

A dependable transliteration platform should therefore preserve the original script and identify:

  1. The source script
  2. The source language
  3. The transliteration system
  4. The treatment of inherent vowels
  5. Whether pronunciation rules were applied
  6. Which diacritics are required
  7. Whether the result is reversible
  8. Whether an official or conventional name overrides the generated form

The goal is not to flatten many writing systems into simplified English spelling. It is to create a transparent bridge that preserves the distinctions carried by each language and script.

FAQ

What is the main difference between IAST and ISO 15919?

IAST is a scholarly system focused on Sanskrit and classical languages, while ISO 15919 is a broader international standard covering multiple Indic scripts and modern languages, including additional diacritics for regional distinctions.

Which transliteration system should I use for Sanskrit texts?

For pure Sanskrit scholarship, IAST is the most widely accepted. If you need to transliterate texts from other Indic scripts or modern languages, ISO 15919 offers broader coverage.

How do I handle inherent vowels and virāma when transliterating?

Both IAST and ISO 15919 assume the inherent vowel ‘a’ for a bare consonant. A virāma (halant) suppresses the vowel, so the consonant is rendered without a vowel, e.g., क् → ‘k’.

Are diacritics mandatory in these transliteration systems?

Yes. Diacritics are essential to preserve phonemic distinctions (e.g., ṭ vs t, ā vs a). Omitting them can lead to ambiguity and loss of information.

Where can I find the official ISO 15919 specification?

The ISO 15919:2001 standard can be purchased from the International Organization for Standardization (ISO) website or accessed through many university libraries that hold ISO standards collections.

Leave a Reply

Your email address will not be published. Required fields are marked *