Short Answer
A title written in Arabic, Chinese, Japanese, Korean, Cyrillic, or an Indic script may need several parallel representations before it can be cataloged, cited, searched, or preserved effectively.
Consider a Russian book title:
- Original script: История русской литературы
- Romanized form: Istorii͡a russkoĭ literatury
- English translation: History of Russian Literature
These three versions are related, but they perform different functions.
The original script preserves the work’s authentic spelling. The romanized form makes the title accessible through Latin-script systems and standardized catalogs. The translation communicates its meaning to readers of another language.
Libraries, academic publishers, and digital archives must keep these functions separate.
A library may use an ALA-LC romanization to create a predictable catalog record. A researcher may need to follow APA, Chicago, or a discipline-specific citation style. A digital repository must preserve the original script, record how a romanized form was produced, normalize Unicode text, and retain enough metadata for future systems to understand the relationship between each representation.
The strongest approach is therefore not to replace non-Latin text with Roman letters. It is to preserve the original and connect it to one or more clearly labeled access forms.
Quick Answer
Romanization helps libraries, scholars, and digital archives represent non-Latin text in Latin letters for discovery, citation, sorting, indexing, and interoperability.
Best practice is to store at least:
- The original-script text
- A standard romanization
- The name and version of the romanization system
- A translation when useful
- Alternative or historical spellings
- Language and script metadata
- A persistent identifier for the person or resource
- Provenance showing how and when the romanized form was created
The three environments have different priorities:
| Environment | Primary objective |
|---|---|
| Library catalog | Standardized discovery and authority control |
| Academic citation | Accurate identification and reader comprehension |
| Digital archive | Long-term preservation, provenance and interoperability |
A romanization that works well in one environment may not be suitable in another.
Romanization, Transliteration and Translation
The terms must be distinguished before discussing metadata or citation practice.
Romanization
Romanization is the representation of text from a non-Latin writing system using Latin letters.
Examples include:
- Москва → Moskva
- 東京 → Tōkyō
- 한국 → Hanguk
- كتاب → kitāb
- संस्कृत → saṃskṛta
Transliteration
Transliteration is the systematic conversion of text from one script into another.
The Program for Cooperative Cataloging defines romanization as transliteration specifically into Latin script. It defines non-standard romanization as a Latin-script conversion that does not follow the applicable ALA-LC table.
Translation
Translation changes the language and communicates meaning.
For example:
| Type | Form |
|---|---|
| Original Arabic | كتاب |
| Romanization | kitāb |
| English translation | book |
The Roman form remains Arabic linguistically. Only the writing system has changed.
Transcription
Transcription represents pronunciation rather than strictly preserving written characters.
This distinction matters because libraries typically prefer systematic transliteration, while language-learning materials may favor pronunciation-oriented transcription.
Why Romanization Still Matters in a Unicode World
Modern information systems can store and display original scripts through Unicode. That does not eliminate the need for romanization.
Romanized data continues to support:
- Users who cannot read the original script
- Latin-script searching
- Alphabetical arrangement
- Legacy library systems
- Citation styles that require Latin characters
- Authority control
- Cross-database matching
- URLs and technical identifiers
- International record exchange
- Discovery across multiple romanization traditions
The PCC recommends including enough non-Latin data to help users identify and locate materials while also maintaining required Latin-script access points. Its guidelines require ALA-LC romanization for PCC bibliographic records and warn that automatic conversion must be reviewed by specialists familiar with the language and script.
The objective is not to choose between original script and romanization. High-quality records use both.
Romanization in Library Catalogs
Libraries need consistent access points for millions of names, titles, subjects, publishers, series, and organizations.
Without shared rules, one cataloger might write an Arabic name with full diacritics, another might use a simplified media spelling, and a third might follow pronunciation. Records for the same author or work could become difficult to connect.
ALA-LC Romanization Tables
The ALA-LC Romanization Tables are transliteration schemes approved by the American Library Association and the Library of Congress.
The official directory contains separate tables for languages and scripts including:
- Arabic
- Armenian
- Bengali
- Chinese
- Greek
- Hebrew and Yiddish
- Hindi
- Japanese
- Korean
- Persian
- Russian
- Sanskrit and Prakrit
- Tamil
- Thai
- Tibetan
- Ukrainian
- Urdu
The directory was last updated on May 11, 2026, and records the approval or revision date of individual tables. Earlier versions are retained for historical reference but should not automatically be used for current cataloging.
Why ALA-LC Is Language-Specific
A writing system may be shared by several languages.
Cyrillic is used for Russian, Ukrainian, Bulgarian, Serbian, Mongolian, and other languages. Arabic-derived scripts are used for Arabic, Persian, Urdu, Pashto, Kurdish, and Sindhi. Devanagari is used for Sanskrit, Hindi, Marathi, Nepali, and other languages.
The same character can have different functions or Latin equivalents in different languages.
A reliable workflow must therefore identify:
- The script
- The language
- The applicable table
- The table version
Selecting a generic “Cyrillic” or “Arabic-script” converter is often insufficient.
Romanized and Original-Script Fields
PCC practice supports parallel Latin and non-Latin data.
In MARC 21 bibliographic records, field 880 contains a fully content-designated representation of another field in a different script. It is linked to the corresponding regular field through subfield $6. The 880 data may contain more than one script.
A simplified conceptual example is:
245 10 $a Istorii͡a russkoĭ literatury
880 10 $6 245-01 $a История русской литературы
The exact MARC syntax and linkage must follow current cataloging rules, but the principle is straightforward:
- One field supplies the standardized Latin representation.
- The linked field preserves the original script.
PCC guidelines state that when non-Latin descriptive data are supplied, parallel Latin and non-Latin forms are required for many important fields, including titles, edition statements, publication statements, and series statements.
Library Romanization Is Not Necessarily Public Spelling
A library form may differ from the spelling familiar to the public.
For the Russian composer Чайковский, possible forms include:
- ALA-LC: Chaĭkovskiĭ
- ISO-style: Čajkovskij
- Practical romanization: Chaykovskiy
- Established English name: Tchaikovsky
The library form is designed to create predictable records. It should not automatically replace a famous established name in general publishing.
Word Division Is Part of Romanization
Romanization involves more than converting characters.
Libraries must also determine:
- Word boundaries
- Hyphenation
- Capitalization
- Grammatical particles
- Personal-name structure
- Administrative terms
- Historical spellings
Chinese ALA-LC romanization is based on Pinyin but follows specialized bibliographic word-division rules. Japanese cataloging requires the cataloger to determine kanji readings and divide words according to detailed library guidance. Arabic romanization requires rules for articles, vowels, hamza, ʿayn, and grammatical endings.
Two records can use the same character table but still differ because one applies incorrect word division.
Authority Control
Authority control connects different forms of the same identity or work.
A name authority record may connect:
- Romanized name
- Original-script form
- Alternative romanization
- Former name
- Pseudonym
- Professional spelling
- Historical spelling
For example:
| Variant type | Example |
|---|---|
| Authorized Latin form | Xi, Jinping |
| Simplified Chinese | 习近平 |
| Traditional Chinese | 習近平 |
| Alternative spacing | Xi Jinping |
Current PCC work continues to address how non-Latin-script variants should be evaluated and represented in authority records, reflecting the importance of script variants for identification and retrieval.
Found Romanization vs Systematic Romanization
A resource may contain its own Latin spelling.
For example, an author may print an English-friendly spelling on the title page even though the ALA-LC rules would generate another form.
Library records should distinguish:
- Systematic romanization: generated according to the approved table
- Found romanization: copied from the item
- Conventional form: an established Latin name
- Original-script form: the name as written in the non-Latin script
Library of Congress guidance for authority-record source citations explicitly distinguishes a systematic romanization from a romanized form found on the item.
This distinction is also valuable outside library catalogs.
Romanization in Academic Citations
Academic citation has a different purpose from cataloging.
A citation must allow the reader to:
- Identify the source
- Find the version used
- Understand enough of the title to assess relevance
- Distinguish the author from similarly named people
- Connect the citation to a DOI, catalog, archive or database record
The appropriate format depends on the required citation style.
Follow the Required Style First
There is no universal academic rule requiring every field to use ALA-LC.
A journal or discipline may specify:
- ALA-LC
- ISO 9
- ISO 233
- IAST
- Hanyu Pinyin
- Modified Hepburn
- A house transliteration system
- Simplified established spellings
Before preparing a bibliography, check:
- Journal instructions
- Publisher style guide
- Departmental or university requirements
- Discipline-specific conventions
- The citation style being used
Consistency throughout the publication is more important than switching systems based on which spelling looks most familiar.
APA Practice for Non-Roman Titles
APA Style instructs writers to transliterate a work’s title when it is written in a non-Roman alphabet for inclusion in the reference list. An English translation can then be supplied to help readers understand the title.
A general pattern is:
Author. (Year). Transliterated title [English translation]. Publisher.
Example structure:
Ivanov, I. I. (2020). Istorii͡a russkoĭ literatury
[History of Russian literature]. Publisher.
The exact capitalization, punctuation, author formatting, and source information should follow the required edition of the style guide.
Original Script in Citations
Some styles or disciplines permit or encourage the original script alongside the romanization.
Possible formats include:
Romanized title (Original-script title) [English translation]
or:
Original-script title [Romanized title; English translation]
The best sequence depends on:
- The journal’s language
- Reader expectations
- The discipline
- Whether the bibliography must sort by Roman letters
- Whether the original script is supported in the publishing system
Where space permits, preserving both forms improves verification.
Cite the Version Actually Used
Do not cite an original-language edition when the research actually relied on a translation.
If the researcher consulted an English translation, the citation should identify that translation, including its translator, edition, publisher, and publication date where required.
The original title may be included as supporting information, but it should not obscure which manifestation was used.
Author Names
For contemporary authors, use the name under which the person publishes.
Do not mechanically replace an author’s chosen credit name with a newly generated transliteration.
A scholar may publish under:
- An official passport spelling
- A simplified spelling
- A name containing diacritics
- A historical family romanization
- Different name forms at different career stages
ORCID allows researchers to specify a published name and add multiple “also known as” variants, including names in other character sets.
An ORCID iD is name-independent and helps connect a researcher to their works even when their name changes or appears in different forms.
Use Persistent Identifiers
A name or title can vary, but a persistent identifier should remain stable.
Important identifiers include:
- DOI for publications and datasets
- ORCID iD for researchers
- ISBN for books
- ISSN for serials
- Archival collection identifiers
- Library authority identifiers
- Repository handles or ARKs
The DOI Foundation describes DOIs as persistent identifiers designed for reliable identification and access by both humans and machines.
A citation containing a DOI is more resilient to spelling variation than a citation based only on a romanized author and title.
Example Citation Layers
An academic database might preserve:
| Field | Value |
|---|---|
| Original author | Иван Иванов |
| Published author form | Ivan Ivanov |
| ALA-LC form | Ivanov, Ivan |
| Original title | История русской литературы |
| Romanized title | Istorii͡a russkoĭ literatury |
| Translated title | History of Russian Literature |
| DOI | Persistent identifier |
| Citation style | APA 7, Chicago, MLA or house style |
The formatted citation should be generated from these structured fields rather than stored as the only metadata record.
Romanization in Digital Archives
A digital archive has responsibilities extending beyond discovery.
It must preserve:
- The original digital object
- Its descriptive metadata
- Its technical characteristics
- Its provenance
- Changes made during processing
- Relationships between original and derivative files
- The standards used to generate access forms
A romanized title may help users discover a resource, but the archive must also be able to explain where that title came from.
Preserve the Original Script
The original-script form is the strongest evidence of the source text.
Romanization may lose:
- Tone
- Vowel length
- Source-character distinctions
- Historical spelling
- Script identity
- Word boundaries
- Grammatical information
For archival purposes, the romanization should be treated as a derivative metadata value, not as a replacement for the source.
Store Romanization as a Separate Metadata Value
A good metadata record may contain:
- Primary title in the source script
- Romanized title
- English translated title
- Historical title
- Abbreviated title
- Alternative romanization
- Language
- Script
- Romanization system
- Standard version
Dublin Core defines Alternative Title as an alternative name for a resource and recognizes that the distinction between the main and alternative title is application-specific. It also defines Bibliographic Citation as a reference containing enough detail to identify a resource unambiguously.
These properties can be used to keep original, romanized, and translated titles distinct.
Record Provenance
An archive should record:
- Who created the romanization
- Whether it was generated or transcribed manually
- Which standard was used
- Which version of the standard applied
- Whether diacritics were removed
- Whether a human reviewed the result
- When the conversion occurred
- Whether the value was later corrected
PREMIS is the international preservation-metadata standard intended to support the preservation and long-term usability of digital objects. It includes structures for documenting preservation events, agents, objects and relationships.
A romanization update can be documented as a metadata-creation or metadata-modification event, especially when it changes search access or replaces an earlier standard.
Preserve Previous Romanizations
Standards change.
An archive may have records created under:
- Wade–Giles and later converted to Pinyin
- McCune–Reischauer and later converted to Revised Romanization
- An older ALA-LC table
- A previous national geographic-name standard
- A local undocumented system
Do not silently overwrite the earlier value.
Instead, record:
| Field | Example |
|---|---|
| Current romanization | Beijing |
| Previous romanization | Pei-ching |
| Historical conventional form | Peking |
| Original script | 北京 |
| Status | Current, superseded or historical |
| Validity date | Date range |
| Source | Standard or authority |
Earlier spellings remain valuable for searching historical catalogs, filenames, citations and imported metadata.
Unicode and Text Normalization
Romanization frequently uses Latin characters with diacritics.
The same visible character may have more than one valid Unicode encoding. For example, ā may be stored as a precomposed character or as a followed by a combining macron.
Unicode normalization provides four standard forms:
- NFC
- NFD
- NFKC
- NFKD
Unicode explains that normalization allows canonically equivalent strings to receive a consistent binary representation. NFC performs canonical decomposition followed by composition, while NFD retains the decomposed form.
Recommended Storage Practice
For most descriptive and display metadata:
- Store original and romanized text in Unicode.
- Normalize to NFC at defined system boundaries.
- Preserve complete diacritics.
- Generate diacritic-free forms separately for searching.
- Do not use compatibility normalization indiscriminately on authoritative text.
Why Diacritics Must Be Preserved
Diacritics may distinguish:
- ā from a
- ḥ from h
- ṣ from s
- ṭ from t
- š from s
- ŏ from o
- ṛ from r
Removing them may destroy reversibility or merge different source characters.
A search field may contain a simplified form such as krsna, but the authoritative transliteration should remain Kṛṣṇa.
Metadata Model for Digital Collections
A single field called romanized_title is rarely sufficient.
A stronger structure is:
| Field | Purpose |
|---|---|
original_title | Title in the source script |
original_title_language | Language code |
original_title_script | Script code |
romanized_title | Standard Latin representation |
romanization_standard | ALA-LC, ISO, Pinyin or another system |
romanization_version | Table edition or rule version |
translated_title | Meaning in the interface language |
alternative_titles | Historical or variant names |
creator_original | Creator name in original script |
creator_published_form | Preferred credited name |
creator_romanized | Systematic romanization |
identifier | DOI, handle, ARK or local ID |
normalization_form | NFC or another documented policy |
conversion_agent | Person or software |
conversion_date | Date generated |
review_status | Automated, reviewed or authority-verified |
This model prevents original, romanized, translated and simplified forms from being confused.
Search and Discovery
Search users rarely know which romanization system a catalog uses.
A researcher may search for:
- Original script
- A standard romanization
- A simplified spelling
- An older system
- A conventional English form
- A spelling without diacritics
- A name copied from a bibliography
A strong search index should connect these forms.
Example: Chinese Place Name
| Type | Form |
|---|---|
| Original | 北京 |
| Pinyin | Beijing |
| Wade–Giles | Pei-ching |
| Historical conventional form | Peking |
Example: Russian Author
| Type | Form |
|---|---|
| Original | Чайковский |
| ALA-LC | Chaĭkovskiĭ |
| ISO-style | Čajkovskij |
| Simplified form | Chaikovskii |
| Established English form | Tchaikovsky |
Search Aliases vs Display Values
Search aliases improve retrieval but should not determine the primary label.
Use separate layers:
- Display value: authoritative or preferred form
- Search alias: alternative indexed spelling
- Sort value: normalized ordering form
- Identifier: stable machine-readable key
A catalog can be tolerant in search while precise in display.
Academic and Archival Name Disambiguation
Romanization can cause different original names to collapse into identical Latin forms.
Chinese Pinyin is a prominent example because many characters share the same pronunciation. Arabic short vowels may be omitted, producing several possible Latin readings. Simplified Indic and Cyrillic romanization can remove distinctions preserved by diacritics.
Do not treat a romanized name as a unique identifier.
Use additional evidence such as:
- ORCID iD
- Authority-record identifier
- Date of birth
- Institutional affiliation
- Coauthors
- Subject area
- Original-script name
- DOI-linked publication history
ORCID specifically allows a researcher to maintain a published name and multiple alternative names, including names in different character sets. This helps systems connect outputs credited under different spellings.
Choosing a Romanization Standard
Use the standard required by the environment.
| Use case | Recommended starting point |
|---|---|
| North American library cataloging | Applicable ALA-LC table |
| International script-conversion project | Applicable ISO standard |
| Academic journal | Journal or discipline style |
| Geographic collection | National, UNGEGN or BGN/PCGN standard |
| Personal-name authority | Preferred or established name plus authority data |
| Language-learning archive | Pronunciation-oriented system plus source script |
| Digital archive search | Standard form plus multiple aliases |
| Historical bibliography | Preserve the source-era spelling |
| URL or filename | Stable ASCII derivative |
| Machine-reversible conversion | Explicit reversible transliteration |
Do not apply one system globally merely because it is familiar.
Automated Romanization
Software can accelerate romanization, especially when:
- Character mappings are stable
- The language is known
- Word boundaries are explicit
- The text is normalized
- The standard has machine-readable rules
Automation is less reliable when:
- Vowels are omitted
- Characters have several readings
- Personal names are irregular
- Word segmentation is ambiguous
- Historical spelling is present
- Multiple languages share a script
- The output must follow established identity usage
PCC guidelines state that macros and automatic transliteration tools cannot be depended upon for correct results in every case and should be reviewed by catalogers with specialist language and script knowledge.
Unicode CLDR likewise distinguishes generic script transforms from variants associated with standards such as UNGEGN, BGN, ISO 9 and the U.S. Library of Congress. Generic and standard-specific transforms should not be treated as interchangeable.
Recommended Automation Workflow
- Preserve the exact input.
- Detect the script.
- Identify or request the language.
- Normalize Unicode text.
- Select a named standard and version.
- Generate the romanization.
- Apply word-division and capitalization rules.
- Normalize the result.
- Check authority data.
- Flag uncertainty.
- Require human review where necessary.
- Record provenance.
Common Problems
Using Romanization Instead of Original Script
Romanized data cannot preserve every source distinction. Keep the original whenever technically possible.
Treating Translation as Romanization
A translated title communicates meaning. A romanized title represents the original-language wording in Latin letters.
Using a Script Table Without Identifying the Language
Russian rules should not be applied automatically to Ukrainian. Arabic rules should not be applied to Persian. Sanskrit rules should not automatically govern modern Hindi.
Mixing Standards
A title should not combine one system’s consonants, another system’s vowel marks, and a third system’s spacing unless the hybrid is declared as an editorial house style.
Omitting the Standard Version
ALA-LC tables and national systems are revised. Store the exact edition or date.
Removing Diacritics From the Only Copy
Generate an unmarked search field instead.
Overwriting Historical Data
Older romanizations may remain essential for citation matching and archive discovery.
Using a Generated Name Over an Author’s Published Name
Use the author’s credited or verified form for attribution. Store systematic romanizations as variants.
Depending on Names Alone
Use ORCID, DOI and authority identifiers to reduce ambiguity.
Ignoring Unicode Normalization
Visually identical strings can be encoded differently and fail exact matching.
Storing Only a Formatted Citation
Store structured metadata so citations can be regenerated in different styles.
Practical Examples
Example 1: Arabic Book
| Metadata field | Value |
|---|---|
| Original title | كتاب الأدب |
| ALA-LC title | Kitāb al-adab |
| English translation | The Book of Literature |
| Language | Arabic |
| Script | Arabic |
| Romanization standard | ALA-LC Arabic |
| Simplified search form | kitab al adab |
| Identifier | ISBN or catalog ID |
The original title, romanization and translation should remain separate.
Example 2: Japanese Article
| Metadata field | Value |
|---|---|
| Original title | 日本語の歴史 |
| Reading | にほんごのれきし |
| Romanized title | Nihongo no rekishi |
| English translation | History of the Japanese Language |
| Romanization system | Declared Japanese library or publisher system |
| Alternative reading | Added only when supported |
| DOI | Persistent identifier |
The kanji reading should be verified rather than generated solely from individual character values.
Example 3: Russian Archival Item
| Metadata field | Value |
|---|---|
| Original title | История русской литературы |
| ALA-LC title | Istorii͡a russkoĭ literatury |
| Simplified title | Istoriia russkoi literatury |
| English translation | History of Russian Literature |
| Historical romanization | Retained if present in an older catalog |
| Normalization | NFC |
| Provenance | Converted and reviewed by named agent |
Example 4: Researcher Identity
| Metadata field | Value |
|---|---|
| Original-script name | 李明 |
| Published name | Ming Li |
| Systematic Pinyin | Li Ming |
| Other names | Li, Ming; 李明 |
| ORCID | Persistent researcher identifier |
| Affiliation | Used as supporting disambiguation |
| Source | Researcher-confirmed |
The published form should control credit, while the systematic and original-script variants improve discovery.
Recommended Policy for Romanization.org
Romanization.org should model library, citation and archive workflows as connected but distinct applications.
Library Mode
Display:
- Original script
- Applicable ALA-LC table
- Table version
- Full romanization
- Word-division notes
- Diacritic explanation
- MARC-friendly output
- Authority-record warnings
Citation Mode
Allow users to select:
- APA
- Chicago
- MLA
- Discipline-specific style
- Journal house style
The output should separate:
- Author’s published name
- Romanized title
- Original title
- English translation
- Edition used
- DOI or other identifier
Archive Mode
Store or export:
- Original script
- Language and script codes
- Romanization standard and version
- Alternative titles
- Conversion agent
- Conversion date
- Review status
- Unicode normalization
- Persistent identifier
- Preservation-event metadata
Comparison Mode
Show side-by-side outputs from:
- ALA-LC
- ISO
- National standard
- Historical standard
- Simplified form
- Common conventional spelling
Each form should be labeled by purpose.
Frequently Asked Questions
What is the best romanization system for libraries?
Use the system required by the cataloging institution. In PCC and many North American library workflows, that is the applicable ALA-LC Romanization Table.
Should libraries still store original scripts?
Yes. Original-script data improves identification, verification and native-language searching. MARC 21 supports alternate script representations through linked field 880.
Is library romanization a pronunciation guide?
Usually not. It is designed primarily for consistent bibliographic representation and retrieval.
Should a citation contain the original script?
That depends on the required style, discipline and publisher. Where supported, including the original script alongside a romanization can improve verification.
Should non-Roman titles be translated in a bibliography?
Many styles provide a translated title for reader comprehension. APA requires transliteration of titles written in non-Roman alphabets and supports adding an English translation in the reference.
Should an author’s name be systematically romanized?
Use the name under which the author publishes whenever possible. Record systematic and original-script forms as additional identifiers or variants.
Why use ORCID when the author’s name is already present?
Names can change, be abbreviated or appear in several scripts. ORCID provides a name-independent identifier that connects researchers with their work.
Why include a DOI?
A DOI provides a persistent identifier for a publication or other object and remains useful even when a title, URL or romanized spelling varies.
Can romanization be automated?
Partly. Automatic output should be reviewed where readings, vowels, names, segmentation or historical forms are uncertain.
Should digital archives remove diacritics?
No. Preserve the complete form and create a simplified derivative for search or ASCII-only systems.
What Unicode normalization should an archive use?
NFC is a common choice for stored and displayed text. The archive should document its policy and test canonical equivalence consistently.
What is the difference between an alternative title and a translated title?
A translated title communicates meaning in another language. An alternative title is a broader metadata category that may include a translation, abbreviation, historical title or different romanization.
Should older romanizations be deleted after a standard changes?
No. Mark them as historical, superseded or alternative forms and retain them for search and provenance.
Is a diacritic-free spelling still the same standard?
It is normally a simplified derivative of the full form. It should be labeled as diacritic-free or ASCII rather than presented as identical to the complete standard.
Final Checklist
Before publishing, cataloging or archiving romanized metadata, confirm:
- The original-script text is preserved.
- The source language is identified.
- The source script is identified.
- The correct romanization system is selected.
- The system version is recorded.
- Romanization is separated from translation.
- Word division follows the chosen standard.
- Diacritics are preserved in the authoritative form.
- The author’s published or preferred name has been checked.
- Original and Roman fields are linked appropriately.
- Alternative and historical spellings are retained.
- A DOI, ORCID or other persistent identifier is included where available.
- Unicode normalization is applied consistently.
- Search keys are separate from display values.
- Automated results have been reviewed.
- Conversion provenance has been recorded.
- Formatted citations can be regenerated from structured metadata.
- Changes to romanization do not erase earlier catalog or archive evidence.
Conclusion
Romanization serves different purposes in libraries, academic citations and digital archives.
Libraries use standardized romanization to support cataloging, authority control, sorting, shelving and cross-script discovery.
Academic writers use romanization to identify non-Latin sources within a citation system their readers can navigate. They may also include translations, original-script titles and persistent identifiers.
Digital archives must go further. They preserve the original script, connect it to all derivative representations, document the standard and version used, normalize Unicode consistently, retain previous spellings and record the provenance of every metadata transformation.
The strongest record therefore contains several layers:
- Original-script title and creator name
- Standard romanization
- Translation where useful
- Preferred or published personal-name form
- Alternative and historical spellings
- Language and script metadata
- Standard and version information
- Persistent identifiers
- Unicode and preservation metadata
- Human-review and provenance status
Romanization should not erase the source. It should create an additional path to it.
For Romanization.org, the goal is not merely to convert characters. It is to help librarians, researchers and archivists create transparent records that remain discoverable, citable and intelligible across scripts, systems and generations.
FAQ
What is the difference between romanization, transliteration, and transcription?
Romanization converts non‑Latin script into Latin letters, transliteration is a systematic script‑to‑script conversion (often romanization when the target script is Latin), and transcription represents pronunciation rather than exact spelling.
Why do libraries still need romanized forms when Unicode can store original scripts?
Romanized forms support users who cannot read the original script, enable Latin‑script searching, alphabetical sorting, legacy system compatibility, citation styles that require Latin characters, and facilitate authority control and cross‑database matching.
Which romanization system should I use for a given language?
Use the system approved for library cataloging by the ALA‑LC tables for that language, or the national standard (e.g., ISO 15919 for Indic scripts) when appropriate. Identify the script, language, and version of the table before applying.
How should original script and romanized titles be stored in MARC records?
Store the romanized title in the standard field (e.g., 245) and the original‑script title in field 880, linked via subfield $6. Both fields should include language and script metadata.
What metadata should accompany a romanized form in a digital archive?
Include the original‑script text, the romanized string, the name and version of the romanization system used, language/script codes, provenance (who/when created it), and a persistent identifier for the resource.
Leave a Reply