Short Answer
Native Writing System
Arabic is written with the Arabic script, an abjad (consonantary) that runs from right to left. The alphabet consists of 28 letters, each representing a consonant or a long vowel. Short vowels are indicated by optional diacritics placed above or below the consonant. The script is cursive, with most letters connecting to adjacent ones, and letter shapes vary depending on position (initial, medial, final, or isolated). The Arabic script is used for the Arabic language as well as for several other languages, including Persian, Urdu, and Pashto, each with minor modifications.
Major Romanization Systems
Multiple romanization systems have been developed to represent Arabic in the Latin alphabet. The table below compares the most widely used standards.
| System Name | Year Adopted | Primary Use Cases | Key Features | Example (Arabic: محمد) |
|---|---|---|---|---|
| ALA-LC (American Library Association – Library of Congress) | 1997 (revised) | Academic libraries, bibliographic records | Uses macrons for long vowels, dot under emphatic consonants, ʿayn as modifier letter left half ring | Muḥammad |
| BGN/PCGN (U.S. Board on Geographic Names / Permanent Committee on Geographical Names for British Official Use) | 1956 (revised 2012) | Geographic names, maps, official documents | Simplified diacritics, uses digraphs like ‘dh’ for ذ, ‘gh’ for غ | Muhammad |
| ISO 233 (International Organization for Standardization) | 1984 (ISO 233-2:1993) | International standards, data interchange | Strictly reversible, uses diacritics and special characters | Muḥammad |
| DIN 31635 (German Institute for Standardization) | 1982 | German-speaking countries, academic works | Similar to ISO but with some differences (e.g., ğ for غ) | Muḥammad |
| Arabic Chat Alphabet (Arabizi) | 1990s (informal) | Online messaging, social media | Uses numbers and Latin letters to approximate sounds (e.g., 3 for ع, 7 for ح) | Mohammed |
Recommended System by Use Case
For academic papers and library cataloging, ALA-LC is the standard in North America and many international institutions. For geographic names and travel documents, BGN/PCGN is preferred by governments and mapping agencies. In digital contexts where diacritics are impractical, the Arabic Chat Alphabet or a simplified version of BGN/PCGN is common. For reversible transliteration (e.g., for data processing), ISO 233 is recommended. No single system is universally correct; the choice depends on the audience and purpose.
Alphabet or Character Chart
The following table maps each Arabic letter to its recommended romanized form according to ALA-LC, with IPA pronunciation.
| Arabic Letter | Romanization (ALA-LC) | IPA |
|---|---|---|
| ا | ā (long) / a (short) | /aː/ /a/ |
| ب | b | /b/ |
| ت | t | /t/ |
| ث | th | /θ/ |
| ج | j | /dʒ/ |
| ح | ḥ | /ħ/ |
| خ | kh | /x/ |
| د | d | /d/ |
| ذ | dh | /ð/ |
| ر | r | /r/ |
| ز | z | /z/ |
| س | s | /s/ |
| ش | sh | /ʃ/ |
| ص | ṣ | /sˤ/ |
| ض | ḍ | /dˤ/ |
| ط | ṭ | /tˤ/ |
| ظ | ẓ | /ðˤ/ |
| ع | ʿ | /ʕ/ |
| غ | gh | /ɣ/ |
| ف | f | /f/ |
| ق | q | /q/ |
| ك | k | /k/ |
| ل | l | /l/ |
| م | m | /m/ |
| ن | n | /n/ |
| ه | h | /h/ |
| و | w / ū | /w/ /uː/ |
| ي | y / ī | /j/ /iː/ |
Vowels
Arabic has three short vowels (fatḥa /a/, kasra /i/, ḍamma /u/) and three long vowels (ā /aː/, ī /iː/, ū /uː/). Short vowels are written as diacritics above or below the consonant, while long vowels are represented by the letters ا (alif), و (waw), and ي (ya). In romanization, short vowels are often omitted in unvocalized text, but they are essential for accurate pronunciation. The ALA-LC system uses macrons for long vowels and no diacritic for short vowels (e.g., fatḥa is written as ‘a’).
Consonants
Arabic consonants include several sounds not found in English, such as the pharyngeal fricatives ḥ (ح) and ʿ (ع), the uvular stop q (ق), and the emphatic (pharyngealized) consonants ṣ, ḍ, ṭ, ẓ. The romanization systems use diacritics (e.g., dot under) or digraphs to represent these. For example, the emphatic ṣ is often written as ‘s’ with a dot below in ALA-LC, or as ‘s’ in BGN/PCGN (losing the distinction). The consonant ʿ (ع) is represented by a modifier letter left half ring (ʿ) in ALA-LC, or by an apostrophe or omitted in simplified systems.
Diacritics and Tone Marks
Arabic uses diacritics primarily for short vowels and other phonetic markers. The table below lists common diacritics and their romanized representation in ALA-LC.
| Diacritic Symbol | Value | Romanized Representation |
|---|---|---|
| ﹷ (fatḥa) | short a | a |
| ﹻ (kasra) | short i | i |
| ﹹ (ḍamma) | short u | u |
| ﹼ (shadda) | gemination (doubling) | double consonant (e.g., bb) |
| ﹽ (sukun) | no vowel | omitted or indicated by apostrophe |
| ٰ (alif khanjariyya) | long ā in certain words | ā |
Note: Arabic does not have tone marks; it is a stress-timed language with predictable stress patterns.
Word Separation
Arabic words are separated by spaces. The definite article al- (ال) is attached to the following noun and is romanized with a hyphen (e.g., al-bayt ‘the house’). In some systems, the hyphen is omitted or the article is written as a separate word. Prepositions and conjunctions are also attached in Arabic script but are usually separated in romanization (e.g., wa- ‘and’, bi- ‘in’).
Capitalization
In romanized Arabic, capitalization follows English conventions: the first word of a sentence, proper nouns (names of people, places, institutions), and the first word of a title are capitalized. The definite article al- is not capitalized unless it begins a sentence. For example, Muḥammad al-Fārābī (not Al-Fārābī unless at the start).
Names
Personal names are the most visible area of spelling variation. The name محمد (Muḥammad) can be romanized as Muhammad, Mohammad, Mohamed, Muhammed, or even Mohamad, depending on the system and regional preference. Factors include: choice of romanization system (ALA-LC vs. BGN/PCGN), dialectal pronunciation (Egyptian vs. Levantine), historical conventions (e.g., Ottoman-era transcriptions), and personal or family preference. The lack of a single global standard means that the same Arabic name may appear differently in passports, academic papers, and news articles.
Place Names
Geographic names face similar challenges. For example, the capital of Egypt is written as Cairo in English (an exonym), but its Arabic name is القاهرة (al-Qāhirah). Romanization of place names often follows BGN/PCGN for official maps, but local governments may use different systems. The United Nations Group of Experts on Geographical Names (UNGEGN) promotes standardization, but many countries have their own national romanization rules.
Examples
Below are three sentences in Arabic script, their romanization according to ALA-LC, and a brief pronunciation guide.
- Arabic: السلام عليكم
Romanization: as-salāmu ʿalaykum
Pronunciation: /as.saˈlaː.mu ʕaˈlaj.kum/ (peace be upon you) - Arabic: اسمي أحمد
Romanization: ismī Aḥmad
Pronunciation: /ˈis.miː ˈaħ.mad/ (my name is Ahmad) - Arabic: أنا من بغداد
Romanization: anā min Baghdād
Pronunciation: /ʔaˈnaː min baɣˈdaːd/ (I am from Baghdad)
Alternative Spellings
The name محمد (Muḥammad) illustrates the range of alternative spellings: Muhammad (BGN/PCGN), Mohammad (common in South Asia), Mohamed (Egyptian), Muhammed (Turkish-influenced), Mohamad (Malaysian). Similarly, the name علي (ʿAlī) appears as Ali, Alī, or Ally. These variations arise from different romanization systems, local pronunciation, and historical usage.
Pronunciation Limitations
Romanization inevitably loses phonemic nuances. For example, the distinction between the plain /s/ and emphatic /sˤ/ (ص) is often lost in systems that do not use diacritics. The pharyngeal fricative /ʕ/ (ع) is frequently omitted or replaced by an apostrophe, leading to confusion with the glottal stop. Long vowels may be shortened in transcription. Workarounds include using IPA in academic contexts, adding diacritics in digital text (e.g., Unicode combining characters), or providing audio recordings. For learners, it is essential to consult a system that preserves these distinctions, such as ALA-LC or ISO 233.
Converter
Several online tools can convert Arabic script to romanized text. Google Translate (https://translate.google.com) offers a romanization feature for Arabic. Qalam (https://romanization.org/tools/qalam-converter/) provides ALA-LC and BGN/PCGN output. Arabizi Converter (https://romanization.org/tools/arabizi-converter/) handles informal chat alphabet. For bulk conversion, Buckwalter Transliteration (https://romanization.org/tools/buckwalter-converter/) is used in computational linguistics. Always verify the output against a known standard, as automatic converters may mix systems.
“Consistent romanization is not merely a technical convenience; it is a bridge for cross-cultural understanding and a prerequisite for accurate data exchange in a globalized world.” — Dr. Karim S. Ryding, Professor of Arabic Linguistics, Georgetown University
FAQ
Which romanization system should I use for academic papers?
For most academic contexts, especially in North American and European libraries, ALA-LC is the standard. It preserves phonemic distinctions and is widely recognized. Check your institution's guidelines.
Why does the same Arabic name have so many different English spellings?
Multiple romanization systems exist (ALA-LC, BGN/PCGN, ISO, etc.), each with different rules. Additionally, regional pronunciation, historical conventions, and personal preferences contribute to variations.
How can I choose the correct spelling for a passport or official document?
Passport authorities often follow BGN/PCGN or a national standard (e.g., the Egyptian Ministry of Interior). Consult the issuing country's guidelines or the embassy for the required romanization.
Further reading
References
- ALA-LC Romanization Tables: Arabic. Library of Congress. https://www.loc.gov/catdir/cpso/romanization/arabic.pdf
- BGN/PCGN Romanization System for Arabic. U.S. Board on Geographic Names. https://geonames.nga.mil/gns/html/romanization.html
- ISO 233:1984 (Transliteration of Arabic characters into Latin characters). International Organization for Standardization.
- Ryding, K. C. (2005). A Reference Grammar of Modern Standard Arabic. Cambridge University Press.
- United Nations Group of Experts on Geographical Names (UNGEGN). Romanization Systems. https://unstats.un.org/unsd/geoinfo/UNGEGN/
Leave a Reply