Short Answer
A book title written in Arabic, Chinese, Japanese, Korean, Russian, Hindi, or another non-Latin script may need a Roman-letter form before it can be searched, sorted, cited, shelved, or exchanged consistently across library systems.
That is the problem the ALA-LC Romanization Tables are designed to solve.
The tables provide standardized rules for converting text from non-Latin writing systems into the Latin alphabet. They are used primarily in library cataloging, bibliographic databases, authority records, archival descriptions, and scholarly citations.
ALA-LC romanization is not intended to translate a title or provide a simple pronunciation guide. Its purpose is to create a consistent Latin-script representation of names, words, and titles so that records created by different catalogers and institutions can work together.
The current Library of Congress directory describes the tables as transliteration schemes approved by the Library of Congress and the American Library Association. The broader ALA-LC publication covers more than 145 languages and dialects written in non-Roman scripts.
Quick Answer
The ALA-LC Romanization Tables are language- and script-specific rules used to transliterate non-Latin text into Roman letters for library and bibliographic purposes.
They help standardize:
- Author names
- Book and journal titles
- Publisher names
- Series titles
- Organizations
- Geographic names in bibliographic records
- Subject and authority access points
- Searchable catalog data
A typical table contains:
- Source-script characters
- Their approved Latin equivalents
- Diacritic rules
- Context-dependent exceptions
- Word-division instructions
- Capitalization and punctuation guidance
- Examples showing how the rules are applied
The ALA-LC framework favors systematic transliteration over informal pronunciation spelling. Its current procedural goals include machine-assisted conversion, reversibility where practical, consistency with international or nationally approved standards, and careful consideration of the effect that revisions have on existing catalog records.
What Does ALA-LC Mean?
ALA-LC combines the names of two organizations:
- ALA: American Library Association
- LC: Library of Congress
The tables are developed and approved through cooperation between the Library of Congress and ALA cataloging bodies. Current procedures use a Romanization Tables Review Board and language-specific review subcommittees that can include librarians, linguists, local speakers, regional experts, and representatives of specialist library communities.
ALA-LC is therefore not one universal alphabet.
It is a collection of separate tables designed for particular languages, scripts, or groups of related languages.
Examples include:
- Arabic
- Armenian
- Bengali
- Bulgarian
- Chinese
- Greek
- Hebrew and Yiddish
- Hindi
- Japanese
- Korean
- Persian
- Russian
- Sanskrit and Prakrit
- Tamil
- Thai
- Tibetan
- Ukrainian
- Urdu
The current directory also includes tables for less widely supported writing systems, including ADLaM, Cham, Deseret, Meitei, N’Ko, Ol Chiki, and Tifinagh-based Moroccan Tamazight.
Why Libraries Use Romanization
Modern catalogs can increasingly store and display original scripts, but romanized data continues to serve several important functions.
Library of Congress documentation identifies uses including:
- Circulation
- Acquisitions
- Serials management
- Shelflisting
- Shelving
- Reference services
- Indexing
- Sorting
- Systems that do not fully support non-Latin scripts
Romanization also helps staff and users who cannot read the original writing system identify and retrieve relevant materials.
Consider a book whose title is written only as:
История русской литературы
A Russian reader can recognize it directly. A user who does not read Cyrillic may need a standardized Latin form, such as:
Istorii͡a russkoĭ literatury
The romanized form creates an access path. It does not replace the original title or translate it into English.
A translation might be:
History of Russian Literature
These three forms serve different purposes:
| Form | Function |
|---|---|
| История русской литературы | Original-script title |
| Istorii͡a russkoĭ literatury | ALA-LC romanization |
| History of Russian Literature | English translation |
Romanization Is Transliteration, Not Translation
The Program for Cooperative Cataloging defines transliteration as the systematic conversion of text from one script into another and romanization as transliteration specifically into Latin script.
Translation changes the language and communicates meaning.
Romanization generally preserves the original language but changes the writing system.
For example:
| Process | Result |
|---|---|
| Original Arabic | كتاب |
| ALA-LC romanization | kitāb |
| English translation | book |
The Arabic table represents كتاب as kitāb, preserving its Arabic linguistic form with a macron over the long vowel. The word book is a translation, not a romanization.
ALA-LC Is Usually Not a Pronunciation Guide
ALA-LC tables are designed primarily for bibliographic consistency.
The current procedural guidelines state that future tables should be transliteration schemes rather than pronunciation guides because pronunciation varies by region, period, dialect, and context. They also encourage mappings that can be reversed systematically whenever possible.
This distinction matters.
A Roman-letter form may contain symbols whose values are not obvious to an English-speaking reader:
- ĭ
- ḥ
- ḍ
- ʻ
- ś
- ṣ
- ṭ
- ū
- zh
- shch
The Russian table, for example, represents:
| Cyrillic | ALA-LC |
|---|---|
| Ё | Ë |
| Ж | Zh |
| Й | Ĭ |
| Х | Kh |
| Ч | Ch |
| Ш | Sh |
| Щ | Shch |
| Ъ | ʺ |
| Ь | ʹ |
These forms distinguish source letters consistently, but users must learn the conventions before treating the output as pronunciation guidance.
Why There Is No Single ALA-LC Table for Every Script
A script may be used by several languages.
Cyrillic, for example, is used for Russian, Ukrainian, Belarusian, Bulgarian, Serbian, Macedonian, Mongolian, and numerous non-Slavic languages. Arabic-derived scripts are used for Arabic, Persian, Urdu, Pashto, Kurdish, Sindhi, Jawi, and others.
Applying one language’s table to another can produce incorrect results.
The Library of Congress therefore provides separate tables such as:
- Russian
- Ukrainian
- Belarusian
- Bulgarian
- Serbian
- Macedonian
- Non-Slavic Languages in Cyrillic Script
It likewise provides separate Arabic-script tables for Arabic, Persian, Urdu, Pashto, Kurdish, Ottoman Turkish, Jawi-Pegon, Sindhi, and other languages.
The correct sequence is:
Identify the language and script first, then select the relevant table.
Do not simply choose “Cyrillic” or “Arabic” based on visual appearance.
What an ALA-LC Table Contains
ALA-LC documents range from short character charts to extensive manuals with dozens of pages of contextual rules.
Character Correspondence Table
The basic chart lists source characters and approved Latin equivalents.
A Russian example is straightforward:
| Russian | Romanization |
|---|---|
| А | A |
| Б | B |
| В | V |
| Г | G |
| Ж | Zh |
| Х | Kh |
| Ч | Ch |
A character chart may be enough for simple alphabetic text, but many languages require additional rules.
Context-Dependent Mappings
A character may be romanized differently depending on its function or position.
In Arabic, و may represent:
- w as a consonant
- ū as a long vowel
- Part of the diphthong aw
Similarly, ي may represent:
- y
- ī
- Part of ay
The Arabic table therefore contains application rules and vocabulary examples rather than only a one-line character mapping.
Readings
Logographic and mixed writing systems require the cataloger to determine how a character sequence is read.
The Chinese table follows Hanyu Pinyin principles, uses Standard Chinese pronunciation, and omits tone marks. It may require authoritative dictionary consultation when a character has multiple possible readings.
Japanese is even more context-sensitive because kanji may have several readings. The 2022 Japanese table says that romanization should mirror the original script and reading as closely as practical, but acknowledges that the process is not exact. It instructs catalogers to consider context and consult standard dictionaries when the reading is uncertain.
Word Division
Romanization is not complete when the characters have been converted. The cataloger must also decide where words begin and end.
Word division may affect:
- Search results
- Alphabetical sorting
- Name headings
- Title matching
- Series records
- Data conversion
- Duplicate detection
The Japanese table describes conversion and word division as the two principal components of library romanization.
The Chinese rules use a specialized library approach: individual syllables are generally separated, but personal names, geographic names, and certain proper nouns are joined. This differs from ordinary Pinyin word-division practice and was designed partly to support consistent bibliographic conversion.
Diacritics and Modifier Characters
ALA-LC uses diacritics and modifier letters to preserve distinctions that basic A–Z spelling cannot represent.
Examples include:
| Symbol | Possible function |
|---|---|
| ā | Long vowel |
| ḥ | Distinct Arabic consonant |
| ʻ | Arabic ʿayn |
| ʹ | Cyrillic soft sign |
| ʺ | Cyrillic hard sign |
| ĭ | Russian Й |
| ṣ | Distinct sibilant |
| ṭ | Retroflex or emphatic consonant |
| ū | Long u |
The review guidelines prefer mappings that minimize ambiguity and support reverse conversion. They allow extended Latin characters but recommend using them carefully so that text remains displayable and technically manageable.
Notes and Exceptions
Tables may contain rules for:
- Historical characters
- Loanwords
- Grammatical endings
- Prefixes and articles
- Abbreviations
- Numerals
- Capitalization
- Punctuation
- Personal names
- Geographic names
- Foreign words written in the source script
These notes are often more important than the initial alphabet chart.
How to Use an ALA-LC Romanization Table
Step 1: Identify the Language
Begin by determining the language of the material.
Do not rely solely on script detection.
For example:
- Cyrillic text could be Russian, Ukrainian, Bulgarian, or another language.
- Arabic-script text could be Arabic, Persian, Urdu, or Pashto.
- Devanagari text could be Sanskrit, Hindi, Marathi, or Nepali.
- Han characters could represent Chinese or occur within Japanese or Korean text.
The ALA-LC Index of Languages identifies which table should be used when a language is covered under a table with a different name.
Step 2: Select the Current Table
Check the official directory before beginning.
The Library of Congress page records the current version of each table and links earlier editions for reference. It explicitly warns that earlier versions should not be used for current cataloging.
As of the May 11, 2026 directory update, examples include:
| Table | Current listed version |
|---|---|
| Arabic | 2012 |
| Balinese | 2025 |
| Chinese | 2012 |
| Japanese | 2022 |
| Korean | 2009, with a minor revision in 2025 |
| Odia | 2024 |
| Russian | 2012 |
| Sindhi | 2024 |
| Tibetan | 2015 |
| Uighur | 2015 |
Romanization metadata should record the table version because revisions can alter characters, spacing, terminology, or application rules.
Step 3: Preserve the Original Script
Record or retain the original-script form before generating the Latin version.
The original text remains the strongest evidence for:
- Spelling
- Character identity
- Language
- Historical form
- Future reconversion
- Verification of disputed readings
In MARC practice described by PCC guidelines, original-script fields can be linked to parallel romanized fields, often through MARC 880 structures.
Step 4: Convert the Characters
Apply the table character by character or unit by unit.
For a simple alphabetic script, this may be direct.
For example, Russian:
- М → M
- о → o
- с → s
- к → k
- в → v
- а → a
Result:
Москва → Moskva
For a contextual script such as Arabic, Japanese, or an Indic script, do not stop at mechanical character substitution. Determine the letter’s grammatical, vocalic, or lexical function.
Step 5: Apply Contextual Rules
Check the rules following the main chart.
Questions may include:
- Is the Arabic letter a consonant or long vowel?
- Does a Japanese kanji have more than one reading?
- Is a Chinese character polyphonic?
- Is the Indic vowel inherent or suppressed?
- Is a consonant affected by its position?
- Is the text historical or modern?
- Does the table distinguish a name from an ordinary word?
The application notes determine whether a technically possible output is actually correct.
Step 6: Apply Word Division
Decide where spaces and hyphens belong.
Do not copy spacing assumptions from English.
Some source scripts:
- Do not normally use spaces between all words
- Combine particles with neighboring words
- Display syllable blocks rather than linear letters
- Use joining conventions different from Latin-script catalog records
Follow the selected table’s instructions rather than inventing segmentation.
Step 7: Apply Capitalization
The source script may not distinguish uppercase and lowercase.
Capitalization is therefore added according to cataloging and romanization rules.
A table may specify capitalization for:
- Personal names
- Corporate names
- Geographic names
- Titles
- Initial words
- Religious terms
- Administrative units
Do not assume that every Romanized word should follow ordinary English title capitalization.
Step 8: Insert Diacritics Precisely
Use the exact characters required by the table.
Do not replace:
- ḥ with h
- ʻ with an ordinary apostrophe
- ĭ with i
- ʹ with a quotation mark
- ṭ with t
- ā with a
Visually similar Unicode characters may have different meanings and behave differently in searching, sorting, and reverse conversion.
The official site recommends using its source documents rather than copying table data from PDFs, partly because specialized characters and fonts can be handled more reliably in the source files.
Step 9: Review the Complete Form
Check the result against:
- The source item
- The table
- Authoritative dictionaries
- Existing authority records
- Established cataloging practice
- Known personal or institutional forms
A correct letter mapping can still produce an incorrect record if the reading, word division, language, or name identity is wrong.
Practical Examples
Arabic
Source:
كتاب
ALA-LC:
kitāb
The macron marks the long vowel represented by alif. Removing it gives kitab, which may be useful as a simplified search form but is not the complete ALA-LC output.
Chinese
ALA-LC Chinese romanization is based on Hanyu Pinyin but omits tone marks and applies specialized bibliographic word-division rules.
For example, 北京 is represented as Beijing, while ordinary non-name syllables are generally treated according to the table’s separation rules. The system replaced the earlier Wade–Giles-based practice in Library of Congress cataloging.
Japanese
Japanese romanization requires both conversion and word division.
The 2022 table explains that a catalog form is intended to reflect the original Japanese text and reading but is not itself a pronunciation guide. When the source supports several readings, the cataloger must use context and authoritative reference works and may create variant access points.
Russian
Selected mappings include:
| Source | ALA-LC |
|---|---|
| Ж | Zh |
| Й | Ĭ |
| Х | Kh |
| Ц | T͡s |
| Ч | Ch |
| Ш | Sh |
| Щ | Shch |
| Ъ | ʺ |
| Ь | ʹ |
Historical Russian texts may contain letters removed in the 1918 spelling reform. The Russian table includes several obsolete characters and directs users to the Church Slavic table for others.
ALA-LC vs Other Romanization Standards
ALA-LC is one system family among many.
ALA-LC vs ISO
ISO standards often prioritize international data exchange and systematic character mapping.
ALA-LC may resemble an ISO system but can differ because library cataloging has its own requirements involving:
- Legacy records
- Authority control
- User searching
- Word division
- Existing bibliographic conventions
- Retrospective database maintenance
ALA-LC review procedures instruct developers to examine national and international standards and align with them when practical, but not at the expense of important bibliographic requirements.
ALA-LC vs BGN/PCGN
BGN/PCGN systems are designed mainly for geographic names used in maps, government databases, and gazetteers.
ALA-LC is designed for bibliographic records.
A place name may therefore have:
- ALA-LC form
- BGN/PCGN form
- National official form
- Conventional English name
The correct form depends on the task.
ALA-LC vs National Systems
A country may have an official romanization used for:
- Passports
- Road signs
- Public administration
- Geographic names
- Education
ALA-LC may adopt or adapt elements of that system, but library rules can add different word division, diacritics, or treatment of foreign names.
Chinese ALA-LC romanization, for example, is Pinyin-based but omits tones and applies library-specific syllable-division conventions.
Original Script and Romanization Should Work Together
Romanization was once essential because many library systems could not store or display large non-Latin character sets.
That technical environment has changed.
The Library of Congress has continued expanding original-script input in bibliographic and authority records, including extended Cyrillic, full CJK Unicode coverage, additional Indigenous North American languages, Armenian, Thai, and Mongolian.
This does not make romanization unnecessary.
Original script and romanization serve complementary purposes.
| Original script provides | Romanization provides |
|---|---|
| Authentic spelling | Latin-script access |
| Source character identity | Sorting and indexing |
| Cultural and linguistic fidelity | Access for users who cannot read the script |
| Evidence for verification | Compatibility with legacy data |
| Better native-language searching | Cross-system record exchange |
The strongest catalog record often includes both.
Can ALA-LC Romanization Be Automated?
Many mappings can be automated, but full reliability depends on the language.
Automation works relatively well when:
- Characters have stable one-to-one mappings
- Word boundaries are explicit
- Pronunciation does not determine spelling
- The language is identified correctly
- The input is normalized
Automation becomes harder when:
- Vowels are omitted
- Characters have multiple readings
- Word division requires linguistic analysis
- Names have established exceptions
- Historical spellings occur
- Sound changes affect output
- Several languages share the script
PCC guidance warns that macros and automatic transliteration tools cannot be trusted to produce correct results in every case. Their output should be reviewed by someone with specialist knowledge of the language and script.
A good automated workflow is:
- Detect the script
- Identify or request the language
- Apply the named table
- Flag ambiguous readings
- Apply word-division rules
- Compare with authority data
- Require human review for uncertain cases
How ALA-LC Affects Catalog Searching
Users searching a library catalog may encounter several versions of the same title or name:
- Original-script form
- ALA-LC romanization
- Simplified spelling without diacritics
- Conventional Latin name
- Historical romanization
- Translation
Search interfaces should ideally connect these forms.
For example, a user might search for:
- Чайковский
- Chaĭkovskiĭ
- Tchaikovsky
- Chaikovsky
These forms may refer to the same person but arise from different naming or romanization traditions.
Search aliases can improve discovery, but the authorized or descriptive form should not be silently replaced by the easiest spelling.
Common Mistakes
Choosing a Table by Script Alone
Arabic script does not necessarily mean Arabic. Cyrillic does not necessarily mean Russian.
Identify the language first.
Using an Earlier Table
Older table versions remain online for historical reference but are not intended for current cataloging.
Treating Romanization as Pronunciation
ALA-LC output may preserve spelling distinctions rather than approximate spoken language for English readers.
Ignoring Application Rules
The alphabet chart is only the beginning. Contextual notes, word division, and exceptions can change the final form.
Removing Diacritics
Diacritics may distinguish separate letters or vowel lengths.
Store a simplified search form separately rather than deleting the complete form.
Confusing Translation With Romanization
A translated title communicates meaning. A romanized title represents the original-language wording in Latin script.
Applying National Signage Rules to Library Records
Road signs, passports, maps, and library catalogs can legitimately use different standards.
Trusting Software Without Review
Automatic conversion can mishandle polyphonic characters, unmarked vowels, names, word boundaries, and historical text.
Replacing Established Personal Names
An author may publish under a conventional Latin spelling that differs from a mechanically generated ALA-LC form. Authority-control and descriptive rules must determine which form is used in each field.
Copying Visually Similar Punctuation
An ordinary apostrophe, modifier letter, soft-sign mark, hamza sign, and ʿayn sign may look similar but are not interchangeable.
Best Practices for Romanization.org
A practical ALA-LC reference tool should include the following features.
Table Finder
Allow users to search by:
- Language
- Script
- Region
- Alternative language name
- Table title
The tool should map languages covered under broader tables, following the official Index of Languages.
Version Label
Display:
- Table name
- Current version
- Approval or revision year
- Earlier versions
- Date last verified
Original and Romanized Text
Show both forms side by side.
Rule Explanation
For each conversion, explain:
- Which character mapping was used
- Whether a contextual rule applied
- Why a diacritic appears
- How word division was determined
- Whether the result is reversible
Alternative Outputs
Provide separate fields for:
- Complete ALA-LC form
- Diacritic-free search form
- ASCII form
- National romanization
- Common conventional spelling
- Translation, when supplied separately
Uncertainty Warnings
Warn when:
- The language is unknown
- A character has multiple readings
- Vowels are not shown
- Name authority data may override the generated form
- Word division is uncertain
- The input contains mixed scripts
- The source uses historical spelling
Correct Unicode Characters
Use the precise Unicode forms required by the tables and normalize text consistently.
Human Review Status
Label the output as:
- Automatically generated
- Dictionary verified
- Authority-record matched
- Manually reviewed
Frequently Asked Questions
What are the ALA-LC Romanization Tables?
They are transliteration schemes approved by the American Library Association and the Library of Congress for converting non-Latin scripts into Latin letters, primarily for library and bibliographic use.
How many languages do they cover?
The Library of Congress describes the wider collection as covering more than 145 languages and dialects written in non-Roman scripts.
Are ALA-LC tables the same as ISO standards?
No. They may align with ISO or national standards where practical, but they are specifically designed for bibliographic and library requirements.
Are the tables pronunciation guides?
Generally, no. Current procedural goals favor transliteration rather than pronunciation spelling.
Is ALA-LC romanization reversible?
Some tables are substantially reversible, while others lose information because of contextual readings, omitted vowels, word division, or merged characters. The procedural guidelines encourage reversibility where practical.
Should tone marks be used in Chinese ALA-LC romanization?
No. The ALA-LC Chinese system is based on Pinyin but omits tone marks.
Which Japanese system does ALA-LC use?
The current Japanese table applies its own detailed bibliographic romanization rules derived from modified Hepburn-related practice. It includes extensive guidance on readings, word division, long vowels, names, and historical spellings.
Why does ALA-LC use so many diacritics?
They preserve distinctions between source characters and support more precise searching, identification, and possible reverse conversion.
Can I omit the diacritics for a website?
A simplified form can be provided for search, URLs, or user convenience, but the complete ALA-LC form should be retained separately.
Which table should I use for Cyrillic?
Use the table for the actual language—such as Russian, Ukrainian, Bulgarian, Serbian, or the Non-Slavic Languages table—not a generic Russian mapping.
Which table should I use for Sanskrit written in a regional script?
Use the Sanskrit and Prakrit table according to its instructions rather than automatically applying the modern regional language’s pronunciation.
Can software apply ALA-LC automatically?
Software can assist, but PCC guidance requires specialist review because automatic conversion does not produce correct results in every case.
Where should the latest table be checked?
Use the official Library of Congress ALA-LC Romanization Tables directory. The directory provides current and historical versions and identifies which editions are valid for current use.
Should I copy characters from the PDF?
The Library of Congress recommends using the table’s source documents rather than the PDF when copying data, because the source files are the documents from which the published tables are produced.
Practical Checklist
Before creating an ALA-LC romanization, confirm:
- The source language is known.
- The source script is known.
- The correct ALA-LC table has been selected.
- The latest table version is being used.
- The original-script text has been preserved.
- Every contextual rule has been reviewed.
- Word division follows the table.
- Capitalization follows cataloging practice.
- Diacritics and modifier letters are exact.
- Historical characters have been identified.
- Proper names have been checked against authority records.
- Automated output has received human review.
- Simplified and ASCII forms are stored separately.
- The table name and version have been recorded.
- Romanization has not been confused with translation.
Conclusion
The ALA-LC Romanization Tables provide a shared bibliographic framework for representing non-Latin writing in Roman letters.
Their value lies in consistency.
When two catalogers encounter the same Arabic title, Russian surname, Chinese place name, or Sanskrit term, a standard table helps them produce compatible records instead of relying on personal spelling preferences.
The tables are not universal pronunciation guides, and they do not eliminate the need for original-script data. They are specialized tools for:
- Catalog access
- Authority control
- Sorting
- Shelving
- Citation
- Record exchange
- Cross-script discovery
Using them correctly requires more than consulting an alphabet chart. The cataloger or tool must identify the language, select the current table, understand contextual rules, divide words correctly, preserve diacritics, and review ambiguous readings.
For Romanization.org, the most useful presentation is therefore not simply:
Character X becomes Latin letter Y.
A dependable ALA-LC reference should explain:
- Which language and table apply
- Which version is current
- How the source characters are mapped
- Which contextual rules change the result
- How word division and capitalization work
- Which diacritics are required
- Whether the conversion is reversible
- How the original script and Roman form should be stored together
Used this way, ALA-LC romanization becomes more than a cataloging convention. It becomes a transparent bridge between writing systems, library records, and the people trying to find the world’s published knowledge.
Leave a Reply