Romanization for Libraries, Academic Citations and Digital Archives

Short Answer

A title written in Arabic, Chinese, Japanese, Korean, Cyrillic, or an Indic script may need several parallel representations before it can be cataloged, cited, searched, or preserved effectively. Consider a Russian book title: Original script: История русской литературы Romanized form: Istorii͡a russkoĭ literatury English translation: History of Russian Literature These three versions are related, but […]

A title written in Arabic, Chinese, Japanese, Korean, Cyrillic, or an Indic script may need several parallel representations before it can be cataloged, cited, searched, or preserved effectively.

Consider a Russian book title:

  • Original script: История русской литературы
  • Romanized form: Istorii͡a russkoĭ literatury
  • English translation: History of Russian Literature

These three versions are related, but they perform different functions.

The original script preserves the work’s authentic spelling. The romanized form makes the title accessible through Latin-script systems and standardized catalogs. The translation communicates its meaning to readers of another language.

Libraries, academic publishers, and digital archives must keep these functions separate.

A library may use an ALA-LC romanization to create a predictable catalog record. A researcher may need to follow APA, Chicago, or a discipline-specific citation style. A digital repository must preserve the original script, record how a romanized form was produced, normalize Unicode text, and retain enough metadata for future systems to understand the relationship between each representation.

The strongest approach is therefore not to replace non-Latin text with Roman letters. It is to preserve the original and connect it to one or more clearly labeled access forms.

Quick Answer

Romanization helps libraries, scholars, and digital archives represent non-Latin text in Latin letters for discovery, citation, sorting, indexing, and interoperability.

Best practice is to store at least:

  1. The original-script text
  2. A standard romanization
  3. The name and version of the romanization system
  4. A translation when useful
  5. Alternative or historical spellings
  6. Language and script metadata
  7. A persistent identifier for the person or resource
  8. Provenance showing how and when the romanized form was created

The three environments have different priorities:

EnvironmentPrimary objective
Library catalogStandardized discovery and authority control
Academic citationAccurate identification and reader comprehension
Digital archiveLong-term preservation, provenance and interoperability

A romanization that works well in one environment may not be suitable in another.

Romanization, Transliteration and Translation

The terms must be distinguished before discussing metadata or citation practice.

Romanization

Romanization is the representation of text from a non-Latin writing system using Latin letters.

Examples include:

  • Москва → Moskva
  • 東京 → Tōkyō
  • 한국 → Hanguk
  • كتاب → kitāb
  • संस्कृत → saṃskṛta

Transliteration

Transliteration is the systematic conversion of text from one script into another.

The Program for Cooperative Cataloging defines romanization as transliteration specifically into Latin script. It defines non-standard romanization as a Latin-script conversion that does not follow the applicable ALA-LC table.

Translation

Translation changes the language and communicates meaning.

For example:

TypeForm
Original Arabicكتاب
Romanizationkitāb
English translationbook

The Roman form remains Arabic linguistically. Only the writing system has changed.

Transcription

Transcription represents pronunciation rather than strictly preserving written characters.

This distinction matters because libraries typically prefer systematic transliteration, while language-learning materials may favor pronunciation-oriented transcription.

Why Romanization Still Matters in a Unicode World

Modern information systems can store and display original scripts through Unicode. That does not eliminate the need for romanization.

Romanized data continues to support:

  • Users who cannot read the original script
  • Latin-script searching
  • Alphabetical arrangement
  • Legacy library systems
  • Citation styles that require Latin characters
  • Authority control
  • Cross-database matching
  • URLs and technical identifiers
  • International record exchange
  • Discovery across multiple romanization traditions

The PCC recommends including enough non-Latin data to help users identify and locate materials while also maintaining required Latin-script access points. Its guidelines require ALA-LC romanization for PCC bibliographic records and warn that automatic conversion must be reviewed by specialists familiar with the language and script.

The objective is not to choose between original script and romanization. High-quality records use both.

Romanization in Library Catalogs

Libraries need consistent access points for millions of names, titles, subjects, publishers, series, and organizations.

Without shared rules, one cataloger might write an Arabic name with full diacritics, another might use a simplified media spelling, and a third might follow pronunciation. Records for the same author or work could become difficult to connect.

ALA-LC Romanization Tables

The ALA-LC Romanization Tables are transliteration schemes approved by the American Library Association and the Library of Congress.

The official directory contains separate tables for languages and scripts including:

  • Arabic
  • Armenian
  • Bengali
  • Chinese
  • Greek
  • Hebrew and Yiddish
  • Hindi
  • Japanese
  • Korean
  • Persian
  • Russian
  • Sanskrit and Prakrit
  • Tamil
  • Thai
  • Tibetan
  • Ukrainian
  • Urdu

The directory was last updated on May 11, 2026, and records the approval or revision date of individual tables. Earlier versions are retained for historical reference but should not automatically be used for current cataloging.

Why ALA-LC Is Language-Specific

A writing system may be shared by several languages.

Cyrillic is used for Russian, Ukrainian, Bulgarian, Serbian, Mongolian, and other languages. Arabic-derived scripts are used for Arabic, Persian, Urdu, Pashto, Kurdish, and Sindhi. Devanagari is used for Sanskrit, Hindi, Marathi, Nepali, and other languages.

The same character can have different functions or Latin equivalents in different languages.

A reliable workflow must therefore identify:

  1. The script
  2. The language
  3. The applicable table
  4. The table version

Selecting a generic “Cyrillic” or “Arabic-script” converter is often insufficient.

Romanized and Original-Script Fields

PCC practice supports parallel Latin and non-Latin data.

In MARC 21 bibliographic records, field 880 contains a fully content-designated representation of another field in a different script. It is linked to the corresponding regular field through subfield $6. The 880 data may contain more than one script.

A simplified conceptual example is:

245 10  $a Istorii͡a russkoĭ literatury
880 10  $6 245-01 $a История русской литературы

The exact MARC syntax and linkage must follow current cataloging rules, but the principle is straightforward:

  • One field supplies the standardized Latin representation.
  • The linked field preserves the original script.

PCC guidelines state that when non-Latin descriptive data are supplied, parallel Latin and non-Latin forms are required for many important fields, including titles, edition statements, publication statements, and series statements.

Library Romanization Is Not Necessarily Public Spelling

A library form may differ from the spelling familiar to the public.

For the Russian composer Чайковский, possible forms include:

  • ALA-LC: Chaĭkovskiĭ
  • ISO-style: Čajkovskij
  • Practical romanization: Chaykovskiy
  • Established English name: Tchaikovsky

The library form is designed to create predictable records. It should not automatically replace a famous established name in general publishing.

Word Division Is Part of Romanization

Romanization involves more than converting characters.

Libraries must also determine:

  • Word boundaries
  • Hyphenation
  • Capitalization
  • Grammatical particles
  • Personal-name structure
  • Administrative terms
  • Historical spellings

Chinese ALA-LC romanization is based on Pinyin but follows specialized bibliographic word-division rules. Japanese cataloging requires the cataloger to determine kanji readings and divide words according to detailed library guidance. Arabic romanization requires rules for articles, vowels, hamza, ʿayn, and grammatical endings.

Two records can use the same character table but still differ because one applies incorrect word division.

Authority Control

Authority control connects different forms of the same identity or work.

A name authority record may connect:

  • Romanized name
  • Original-script form
  • Alternative romanization
  • Former name
  • Pseudonym
  • Professional spelling
  • Historical spelling

For example:

Variant typeExample
Authorized Latin formXi, Jinping
Simplified Chinese习近平
Traditional Chinese習近平
Alternative spacingXi Jinping

Current PCC work continues to address how non-Latin-script variants should be evaluated and represented in authority records, reflecting the importance of script variants for identification and retrieval.

Found Romanization vs Systematic Romanization

A resource may contain its own Latin spelling.

For example, an author may print an English-friendly spelling on the title page even though the ALA-LC rules would generate another form.

Library records should distinguish:

  • Systematic romanization: generated according to the approved table
  • Found romanization: copied from the item
  • Conventional form: an established Latin name
  • Original-script form: the name as written in the non-Latin script

Library of Congress guidance for authority-record source citations explicitly distinguishes a systematic romanization from a romanized form found on the item.

This distinction is also valuable outside library catalogs.

Romanization in Academic Citations

Academic citation has a different purpose from cataloging.

A citation must allow the reader to:

  • Identify the source
  • Find the version used
  • Understand enough of the title to assess relevance
  • Distinguish the author from similarly named people
  • Connect the citation to a DOI, catalog, archive or database record

The appropriate format depends on the required citation style.

Follow the Required Style First

There is no universal academic rule requiring every field to use ALA-LC.

A journal or discipline may specify:

  • ALA-LC
  • ISO 9
  • ISO 233
  • IAST
  • Hanyu Pinyin
  • Modified Hepburn
  • A house transliteration system
  • Simplified established spellings

Before preparing a bibliography, check:

  1. Journal instructions
  2. Publisher style guide
  3. Departmental or university requirements
  4. Discipline-specific conventions
  5. The citation style being used

Consistency throughout the publication is more important than switching systems based on which spelling looks most familiar.

APA Practice for Non-Roman Titles

APA Style instructs writers to transliterate a work’s title when it is written in a non-Roman alphabet for inclusion in the reference list. An English translation can then be supplied to help readers understand the title.

A general pattern is:

Author. (Year). Transliterated title [English translation]. Publisher.

Example structure:

Ivanov, I. I. (2020). Istorii͡a russkoĭ literatury
[History of Russian literature]. Publisher.

The exact capitalization, punctuation, author formatting, and source information should follow the required edition of the style guide.

Original Script in Citations

Some styles or disciplines permit or encourage the original script alongside the romanization.

Possible formats include:

Romanized title (Original-script title) [English translation]

or:

Original-script title [Romanized title; English translation]

The best sequence depends on:

  • The journal’s language
  • Reader expectations
  • The discipline
  • Whether the bibliography must sort by Roman letters
  • Whether the original script is supported in the publishing system

Where space permits, preserving both forms improves verification.

Cite the Version Actually Used

Do not cite an original-language edition when the research actually relied on a translation.

If the researcher consulted an English translation, the citation should identify that translation, including its translator, edition, publisher, and publication date where required.

The original title may be included as supporting information, but it should not obscure which manifestation was used.

Author Names

For contemporary authors, use the name under which the person publishes.

Do not mechanically replace an author’s chosen credit name with a newly generated transliteration.

A scholar may publish under:

  • An official passport spelling
  • A simplified spelling
  • A name containing diacritics
  • A historical family romanization
  • Different name forms at different career stages

ORCID allows researchers to specify a published name and add multiple “also known as” variants, including names in other character sets.

An ORCID iD is name-independent and helps connect a researcher to their works even when their name changes or appears in different forms.

Use Persistent Identifiers

A name or title can vary, but a persistent identifier should remain stable.

Important identifiers include:

  • DOI for publications and datasets
  • ORCID iD for researchers
  • ISBN for books
  • ISSN for serials
  • Archival collection identifiers
  • Library authority identifiers
  • Repository handles or ARKs

The DOI Foundation describes DOIs as persistent identifiers designed for reliable identification and access by both humans and machines.

A citation containing a DOI is more resilient to spelling variation than a citation based only on a romanized author and title.

Example Citation Layers

An academic database might preserve:

FieldValue
Original authorИван Иванов
Published author formIvan Ivanov
ALA-LC formIvanov, Ivan
Original titleИстория русской литературы
Romanized titleIstorii͡a russkoĭ literatury
Translated titleHistory of Russian Literature
DOIPersistent identifier
Citation styleAPA 7, Chicago, MLA or house style

The formatted citation should be generated from these structured fields rather than stored as the only metadata record.

Romanization in Digital Archives

A digital archive has responsibilities extending beyond discovery.

It must preserve:

  • The original digital object
  • Its descriptive metadata
  • Its technical characteristics
  • Its provenance
  • Changes made during processing
  • Relationships between original and derivative files
  • The standards used to generate access forms

A romanized title may help users discover a resource, but the archive must also be able to explain where that title came from.

Preserve the Original Script

The original-script form is the strongest evidence of the source text.

Romanization may lose:

  • Tone
  • Vowel length
  • Source-character distinctions
  • Historical spelling
  • Script identity
  • Word boundaries
  • Grammatical information

For archival purposes, the romanization should be treated as a derivative metadata value, not as a replacement for the source.

Store Romanization as a Separate Metadata Value

A good metadata record may contain:

  • Primary title in the source script
  • Romanized title
  • English translated title
  • Historical title
  • Abbreviated title
  • Alternative romanization
  • Language
  • Script
  • Romanization system
  • Standard version

Dublin Core defines Alternative Title as an alternative name for a resource and recognizes that the distinction between the main and alternative title is application-specific. It also defines Bibliographic Citation as a reference containing enough detail to identify a resource unambiguously.

These properties can be used to keep original, romanized, and translated titles distinct.

Record Provenance

An archive should record:

  • Who created the romanization
  • Whether it was generated or transcribed manually
  • Which standard was used
  • Which version of the standard applied
  • Whether diacritics were removed
  • Whether a human reviewed the result
  • When the conversion occurred
  • Whether the value was later corrected

PREMIS is the international preservation-metadata standard intended to support the preservation and long-term usability of digital objects. It includes structures for documenting preservation events, agents, objects and relationships.

A romanization update can be documented as a metadata-creation or metadata-modification event, especially when it changes search access or replaces an earlier standard.

Preserve Previous Romanizations

Standards change.

An archive may have records created under:

  • Wade–Giles and later converted to Pinyin
  • McCune–Reischauer and later converted to Revised Romanization
  • An older ALA-LC table
  • A previous national geographic-name standard
  • A local undocumented system

Do not silently overwrite the earlier value.

Instead, record:

FieldExample
Current romanizationBeijing
Previous romanizationPei-ching
Historical conventional formPeking
Original script北京
StatusCurrent, superseded or historical
Validity dateDate range
SourceStandard or authority

Earlier spellings remain valuable for searching historical catalogs, filenames, citations and imported metadata.

Unicode and Text Normalization

Romanization frequently uses Latin characters with diacritics.

The same visible character may have more than one valid Unicode encoding. For example, ā may be stored as a precomposed character or as a followed by a combining macron.

Unicode normalization provides four standard forms:

  • NFC
  • NFD
  • NFKC
  • NFKD

Unicode explains that normalization allows canonically equivalent strings to receive a consistent binary representation. NFC performs canonical decomposition followed by composition, while NFD retains the decomposed form.

For most descriptive and display metadata:

  • Store original and romanized text in Unicode.
  • Normalize to NFC at defined system boundaries.
  • Preserve complete diacritics.
  • Generate diacritic-free forms separately for searching.
  • Do not use compatibility normalization indiscriminately on authoritative text.

Why Diacritics Must Be Preserved

Diacritics may distinguish:

  • ā from a
  • ḥ from h
  • ṣ from s
  • ṭ from t
  • š from s
  • ŏ from o
  • ṛ from r

Removing them may destroy reversibility or merge different source characters.

A search field may contain a simplified form such as krsna, but the authoritative transliteration should remain Kṛṣṇa.

Metadata Model for Digital Collections

A single field called romanized_title is rarely sufficient.

A stronger structure is:

FieldPurpose
original_titleTitle in the source script
original_title_languageLanguage code
original_title_scriptScript code
romanized_titleStandard Latin representation
romanization_standardALA-LC, ISO, Pinyin or another system
romanization_versionTable edition or rule version
translated_titleMeaning in the interface language
alternative_titlesHistorical or variant names
creator_originalCreator name in original script
creator_published_formPreferred credited name
creator_romanizedSystematic romanization
identifierDOI, handle, ARK or local ID
normalization_formNFC or another documented policy
conversion_agentPerson or software
conversion_dateDate generated
review_statusAutomated, reviewed or authority-verified

This model prevents original, romanized, translated and simplified forms from being confused.

Search and Discovery

Search users rarely know which romanization system a catalog uses.

A researcher may search for:

  • Original script
  • A standard romanization
  • A simplified spelling
  • An older system
  • A conventional English form
  • A spelling without diacritics
  • A name copied from a bibliography

A strong search index should connect these forms.

Example: Chinese Place Name

TypeForm
Original北京
PinyinBeijing
Wade–GilesPei-ching
Historical conventional formPeking

Example: Russian Author

TypeForm
OriginalЧайковский
ALA-LCChaĭkovskiĭ
ISO-styleČajkovskij
Simplified formChaikovskii
Established English formTchaikovsky

Search Aliases vs Display Values

Search aliases improve retrieval but should not determine the primary label.

Use separate layers:

  • Display value: authoritative or preferred form
  • Search alias: alternative indexed spelling
  • Sort value: normalized ordering form
  • Identifier: stable machine-readable key

A catalog can be tolerant in search while precise in display.

Academic and Archival Name Disambiguation

Romanization can cause different original names to collapse into identical Latin forms.

Chinese Pinyin is a prominent example because many characters share the same pronunciation. Arabic short vowels may be omitted, producing several possible Latin readings. Simplified Indic and Cyrillic romanization can remove distinctions preserved by diacritics.

Do not treat a romanized name as a unique identifier.

Use additional evidence such as:

  • ORCID iD
  • Authority-record identifier
  • Date of birth
  • Institutional affiliation
  • Coauthors
  • Subject area
  • Original-script name
  • DOI-linked publication history

ORCID specifically allows a researcher to maintain a published name and multiple alternative names, including names in different character sets. This helps systems connect outputs credited under different spellings.

Choosing a Romanization Standard

Use the standard required by the environment.

Use caseRecommended starting point
North American library catalogingApplicable ALA-LC table
International script-conversion projectApplicable ISO standard
Academic journalJournal or discipline style
Geographic collectionNational, UNGEGN or BGN/PCGN standard
Personal-name authorityPreferred or established name plus authority data
Language-learning archivePronunciation-oriented system plus source script
Digital archive searchStandard form plus multiple aliases
Historical bibliographyPreserve the source-era spelling
URL or filenameStable ASCII derivative
Machine-reversible conversionExplicit reversible transliteration

Do not apply one system globally merely because it is familiar.

Automated Romanization

Software can accelerate romanization, especially when:

  • Character mappings are stable
  • The language is known
  • Word boundaries are explicit
  • The text is normalized
  • The standard has machine-readable rules

Automation is less reliable when:

  • Vowels are omitted
  • Characters have several readings
  • Personal names are irregular
  • Word segmentation is ambiguous
  • Historical spelling is present
  • Multiple languages share a script
  • The output must follow established identity usage

PCC guidelines state that macros and automatic transliteration tools cannot be depended upon for correct results in every case and should be reviewed by catalogers with specialist language and script knowledge.

Unicode CLDR likewise distinguishes generic script transforms from variants associated with standards such as UNGEGN, BGN, ISO 9 and the U.S. Library of Congress. Generic and standard-specific transforms should not be treated as interchangeable.

  1. Preserve the exact input.
  2. Detect the script.
  3. Identify or request the language.
  4. Normalize Unicode text.
  5. Select a named standard and version.
  6. Generate the romanization.
  7. Apply word-division and capitalization rules.
  8. Normalize the result.
  9. Check authority data.
  10. Flag uncertainty.
  11. Require human review where necessary.
  12. Record provenance.

Common Problems

Using Romanization Instead of Original Script

Romanized data cannot preserve every source distinction. Keep the original whenever technically possible.

Treating Translation as Romanization

A translated title communicates meaning. A romanized title represents the original-language wording in Latin letters.

Using a Script Table Without Identifying the Language

Russian rules should not be applied automatically to Ukrainian. Arabic rules should not be applied to Persian. Sanskrit rules should not automatically govern modern Hindi.

Mixing Standards

A title should not combine one system’s consonants, another system’s vowel marks, and a third system’s spacing unless the hybrid is declared as an editorial house style.

Omitting the Standard Version

ALA-LC tables and national systems are revised. Store the exact edition or date.

Removing Diacritics From the Only Copy

Generate an unmarked search field instead.

Overwriting Historical Data

Older romanizations may remain essential for citation matching and archive discovery.

Using a Generated Name Over an Author’s Published Name

Use the author’s credited or verified form for attribution. Store systematic romanizations as variants.

Depending on Names Alone

Use ORCID, DOI and authority identifiers to reduce ambiguity.

Ignoring Unicode Normalization

Visually identical strings can be encoded differently and fail exact matching.

Storing Only a Formatted Citation

Store structured metadata so citations can be regenerated in different styles.

Practical Examples

Example 1: Arabic Book

Metadata fieldValue
Original titleكتاب الأدب
ALA-LC titleKitāb al-adab
English translationThe Book of Literature
LanguageArabic
ScriptArabic
Romanization standardALA-LC Arabic
Simplified search formkitab al adab
IdentifierISBN or catalog ID

The original title, romanization and translation should remain separate.

Example 2: Japanese Article

Metadata fieldValue
Original title日本語の歴史
Readingにほんごのれきし
Romanized titleNihongo no rekishi
English translationHistory of the Japanese Language
Romanization systemDeclared Japanese library or publisher system
Alternative readingAdded only when supported
DOIPersistent identifier

The kanji reading should be verified rather than generated solely from individual character values.

Example 3: Russian Archival Item

Metadata fieldValue
Original titleИстория русской литературы
ALA-LC titleIstorii͡a russkoĭ literatury
Simplified titleIstoriia russkoi literatury
English translationHistory of Russian Literature
Historical romanizationRetained if present in an older catalog
NormalizationNFC
ProvenanceConverted and reviewed by named agent

Example 4: Researcher Identity

Metadata fieldValue
Original-script name李明
Published nameMing Li
Systematic PinyinLi Ming
Other namesLi, Ming; 李明
ORCIDPersistent researcher identifier
AffiliationUsed as supporting disambiguation
SourceResearcher-confirmed

The published form should control credit, while the systematic and original-script variants improve discovery.

Recommended Policy for Romanization.org

Romanization.org should model library, citation and archive workflows as connected but distinct applications.

Library Mode

Display:

  • Original script
  • Applicable ALA-LC table
  • Table version
  • Full romanization
  • Word-division notes
  • Diacritic explanation
  • MARC-friendly output
  • Authority-record warnings

Citation Mode

Allow users to select:

  • APA
  • Chicago
  • MLA
  • Discipline-specific style
  • Journal house style

The output should separate:

  • Author’s published name
  • Romanized title
  • Original title
  • English translation
  • Edition used
  • DOI or other identifier

Archive Mode

Store or export:

  • Original script
  • Language and script codes
  • Romanization standard and version
  • Alternative titles
  • Conversion agent
  • Conversion date
  • Review status
  • Unicode normalization
  • Persistent identifier
  • Preservation-event metadata

Comparison Mode

Show side-by-side outputs from:

  • ALA-LC
  • ISO
  • National standard
  • Historical standard
  • Simplified form
  • Common conventional spelling

Each form should be labeled by purpose.

Frequently Asked Questions

What is the best romanization system for libraries?

Use the system required by the cataloging institution. In PCC and many North American library workflows, that is the applicable ALA-LC Romanization Table.

Should libraries still store original scripts?

Yes. Original-script data improves identification, verification and native-language searching. MARC 21 supports alternate script representations through linked field 880.

Is library romanization a pronunciation guide?

Usually not. It is designed primarily for consistent bibliographic representation and retrieval.

Should a citation contain the original script?

That depends on the required style, discipline and publisher. Where supported, including the original script alongside a romanization can improve verification.

Should non-Roman titles be translated in a bibliography?

Many styles provide a translated title for reader comprehension. APA requires transliteration of titles written in non-Roman alphabets and supports adding an English translation in the reference.

Should an author’s name be systematically romanized?

Use the name under which the author publishes whenever possible. Record systematic and original-script forms as additional identifiers or variants.

Why use ORCID when the author’s name is already present?

Names can change, be abbreviated or appear in several scripts. ORCID provides a name-independent identifier that connects researchers with their work.

Why include a DOI?

A DOI provides a persistent identifier for a publication or other object and remains useful even when a title, URL or romanized spelling varies.

Can romanization be automated?

Partly. Automatic output should be reviewed where readings, vowels, names, segmentation or historical forms are uncertain.

Should digital archives remove diacritics?

No. Preserve the complete form and create a simplified derivative for search or ASCII-only systems.

What Unicode normalization should an archive use?

NFC is a common choice for stored and displayed text. The archive should document its policy and test canonical equivalence consistently.

What is the difference between an alternative title and a translated title?

A translated title communicates meaning in another language. An alternative title is a broader metadata category that may include a translation, abbreviation, historical title or different romanization.

Should older romanizations be deleted after a standard changes?

No. Mark them as historical, superseded or alternative forms and retain them for search and provenance.

Is a diacritic-free spelling still the same standard?

It is normally a simplified derivative of the full form. It should be labeled as diacritic-free or ASCII rather than presented as identical to the complete standard.

Final Checklist

Before publishing, cataloging or archiving romanized metadata, confirm:

  • The original-script text is preserved.
  • The source language is identified.
  • The source script is identified.
  • The correct romanization system is selected.
  • The system version is recorded.
  • Romanization is separated from translation.
  • Word division follows the chosen standard.
  • Diacritics are preserved in the authoritative form.
  • The author’s published or preferred name has been checked.
  • Original and Roman fields are linked appropriately.
  • Alternative and historical spellings are retained.
  • A DOI, ORCID or other persistent identifier is included where available.
  • Unicode normalization is applied consistently.
  • Search keys are separate from display values.
  • Automated results have been reviewed.
  • Conversion provenance has been recorded.
  • Formatted citations can be regenerated from structured metadata.
  • Changes to romanization do not erase earlier catalog or archive evidence.

Conclusion

Romanization serves different purposes in libraries, academic citations and digital archives.

Libraries use standardized romanization to support cataloging, authority control, sorting, shelving and cross-script discovery.

Academic writers use romanization to identify non-Latin sources within a citation system their readers can navigate. They may also include translations, original-script titles and persistent identifiers.

Digital archives must go further. They preserve the original script, connect it to all derivative representations, document the standard and version used, normalize Unicode consistently, retain previous spellings and record the provenance of every metadata transformation.

The strongest record therefore contains several layers:

  1. Original-script title and creator name
  2. Standard romanization
  3. Translation where useful
  4. Preferred or published personal-name form
  5. Alternative and historical spellings
  6. Language and script metadata
  7. Standard and version information
  8. Persistent identifiers
  9. Unicode and preservation metadata
  10. Human-review and provenance status

Romanization should not erase the source. It should create an additional path to it.

For Romanization.org, the goal is not merely to convert characters. It is to help librarians, researchers and archivists create transparent records that remain discoverable, citable and intelligible across scripts, systems and generations.

FAQ

What is the difference between romanization, transliteration, and transcription?

Romanization converts non‑Latin script into Latin letters, transliteration is a systematic script‑to‑script conversion (often romanization when the target script is Latin), and transcription represents pronunciation rather than exact spelling.

Why do libraries still need romanized forms when Unicode can store original scripts?

Romanized forms support users who cannot read the original script, enable Latin‑script searching, alphabetical sorting, legacy system compatibility, citation styles that require Latin characters, and facilitate authority control and cross‑database matching.

Which romanization system should I use for a given language?

Use the system approved for library cataloging by the ALA‑LC tables for that language, or the national standard (e.g., ISO 15919 for Indic scripts) when appropriate. Identify the script, language, and version of the table before applying.

How should original script and romanized titles be stored in MARC records?

Store the romanized title in the standard field (e.g., 245) and the original‑script title in field 880, linked via subfield $6. Both fields should include language and script metadata.

What metadata should accompany a romanized form in a digital archive?

Include the original‑script text, the romanized string, the name and version of the romanization system used, language/script codes, provenance (who/when created it), and a persistent identifier for the resource.

Leave a Reply

Your email address will not be published. Required fields are marked *