Short Answer
Native Writing System
Persian (Farsi) is written using a modified version of the Arabic script, written from right to left. The alphabet consists of 32 letters, including four additional characters not found in Arabic: pe (پ), che (چ), zhe (ژ), and gaf (گ). The script is cursive, with most letters connecting to both preceding and following letters. Short vowels are typically omitted in everyday writing, indicated only by optional diacritics (harakat) in educational or religious texts. The Persian script is used in Iran, Afghanistan (Dari), and Tajikistan (Tajik, though Cyrillic is also used).
Major Romanization Systems
Several romanization systems have been developed for Persian, each serving different purposes. The table below compares the most widely used systems.
| System Name | Year Adopted | Primary Use Cases | Key Features | Example (خدا حافظ) |
|---|---|---|---|---|
| ALA-LC (American Library Association – Library of Congress) | 1997 (revised) | Academic cataloging, bibliographic records | Reversible, uses diacritics for long vowels and consonants; distinguishes alef and ayn | khudā ḥāfiẓ |
| DMG (Deutsche Morgenländische Gesellschaft) | 1935 | German academic publications, linguistics | Phonemic, uses macrons and underdots; distinguishes emphatic consonants | ḫudā ḥāfiẓ |
| Unipers (also called Fingilish or Persian Romanization) | 1990s (informal) | Digital communication, informal writing, travel | No diacritics, uses digraphs (e.g., kh, sh); based on English keyboard | khoda hafez |
| BGN/PCGN (U.S. Board on Geographic Names / Permanent Committee on Geographical Names) | 1956 (revised 2012) | Geographic names, maps, official documents | Simplified, no diacritics; uses e for short vowel, i for long | khoda hafez |
Recommended System by Use Case
For English-speaking audiences, the choice of romanization depends on context. Unipers is recommended for digital communication, travel guides, and informal learning because it avoids diacritics and uses familiar English digraphs (e.g., kh, sh, ch). For academic papers and library cataloging, ALA-LC is the standard in North America, offering reversibility and precise phonemic representation. For geographic names, BGN/PCGN is the official system used by U.S. and UK mapping agencies. The DMG system is preferred in German-speaking academic circles but is less common in English contexts.
Alphabet or Character Chart
The following table maps each Persian letter to its recommended romanized form (Unipers) and IPA pronunciation.
| Persian Character | Romanized (Unipers) | IPA |
|---|---|---|
| ا | â (initial), a (medial/final) | /ɒː/ (â), /æ/ (a) |
| ب | b | /b/ |
| پ | p | /p/ |
| ت | t | /t/ |
| ث | s | /s/ |
| ج | j | /dʒ/ |
| چ | ch | /tʃ/ |
| ح | h | /h/ |
| خ | kh | /x/ |
| د | d | /d/ |
| ذ | z | /z/ |
| ر | r | /ɾ/ |
| ز | z | /z/ |
| ژ | zh | /ʒ/ |
| س | s | /s/ |
| ش | sh | /ʃ/ |
| ص | s | /s/ |
| ض | z | /z/ |
| ط | t | /t/ |
| ظ | z | /z/ |
| ع | ’ (ayn) | /ʔ/ |
| غ | gh | /ɣ/ |
| ف | f | /f/ |
| ق | gh | /ɢ/ |
| ک | k | /k/ |
| گ | g | /ɡ/ |
| ل | l | /l/ |
| م | m | /m/ |
| ن | n | /n/ |
| و | v (consonant), u (vowel) | /v/ (consonant), /uː/ (vowel) |
| ه | h | /h/ |
| ی | y (consonant), i (vowel) | /j/ (consonant), /iː/ (vowel) |
Vowels
Persian has six vowel phonemes: three short (a /æ/, e /e/, o /o/) and three long (â /ɒː/, i /iː/, u /uː/). In romanization, short vowels are often omitted in Unipers but can be indicated with a, e, o. Long vowels are always written: â, i, u. The ALA-LC system uses macrons: ā, ī, ū. Note that the short vowel e is often romanized as e (e.g., ketab for کتاب) and o as o (e.g., dokhtar for دختر).
Consonants
Persian consonants include several sounds not found in English, such as the uvular stop q (ق) and the voiceless velar fricative kh (خ). The glottal stop (ع) is represented by an apostrophe or ayn. The retroflex sounds of Arabic (e.g., ص, ض, ط, ظ) are pronounced as plain /s/, /z/, /t/, /z/ in standard Persian. The chart above provides the full set. A common pitfall is confusing gh (غ) with q (ق); both are romanized as gh in Unipers, but ALA-LC distinguishes them as gh and q.
Diacritics and Tone Marks
Persian uses a limited set of diacritics, primarily for vowel indication and gemination. Tone is not phonemic in Persian. The table below shows common diacritics and their romanized representation.
| Diacritic Symbol | Name / Value | Romanized Representation |
|---|---|---|
| َ (fatha) | Short vowel /æ/ | a |
| ِ (kasra) | Short vowel /e/ | e |
| ُ (damma) | Short vowel /o/ | o |
| ْ (sukun) | No vowel (consonant closure) | (omitted) |
| ّ (tashdid) | Gemination (doubled consonant) | Double the consonant (e.g., bb) |
| ٔ (hamze) | Glottal stop or vowel hiatus | ’ (apostrophe) |
Word Separation
Persian words are separated by spaces, similar to English. However, certain compound words and clitics (e.g., the indefinite suffix -i, the conjunction -o) are often written attached to the preceding word. In romanization, these are typically separated by a hyphen or written as a single word depending on the system. For example, ketâb-i (کتابی) or ketâbi. The ALA-LC system uses a hyphen for clitics, while Unipers often writes them together.
Capitalization
In romanized Persian, capitalization follows English conventions: proper nouns (personal names, place names, titles) are capitalized. The first word of a sentence is also capitalized. However, the native Persian script does not have capital letters; capitalization is a feature of the romanization system only.
Names
Personal names are romanized according to the individual’s preference or the standard system used in official documents. For example, the name محمد is romanized as Mohammad (Unipers) or Muḥammad (ALA-LC). Common pitfalls include the representation of the letter ع (ayn) in names like Ali (علی) – often written without the ayn in Unipers. The Iranian passport system uses a specific romanization based on BGN/PCGN.
Place Names
Place names in Iran are officially romanized using the BGN/PCGN system. For example, تهران is Tehran (not Teheran), and اصفهان is Esfahan (not Isfahan). However, historical variants persist in English (e.g., Shiraz vs. Shiraz). The United Nations Group of Experts on Geographical Names (UNGEGN) recommends BGN/PCGN for international use.
Examples
Below are three real-world sentences in Persian script, romanized using Unipers, and a pronunciation guide.
- Native: سلام، حال شما چطور است؟
Romanized: Salam, hâl-e shomâ chetor ast?
Pronunciation: /sæˈlɒːm hɒːl-e ʃoˈmɒː tʃeˈtoɾ æst/ - Native: من کتاب فارسی میخوانم.
Romanized: Man ketâb-e fârsi mi-khânam.
Pronunciation: /mæn keˈtɒːb-e fɒːɾˈsiː miːˈxɒːnæm/ - Native: خیابان ولیعصر طولانیترین خیابان تهران است.
Romanized: Khiyâbân-e Vali-Asr tulâni-tarin khiyâbân-e Tehrân ast.
Pronunciation: /xiːjɒːˈbɒːn-e væliːˈæsɾ tuːlɒːniːtæˈɾiːn xiːjɒːˈbɒːn-e tehˈɾɒːn æst/
Alternative Spellings
Due to the lack of a single authoritative romanization, many Persian words have multiple accepted spellings in English. For example, Qom (قم) is also spelled Ghom; Mashhad (مشهد) is sometimes Meshed; Khomeini (خمینی) appears as Khomeyni. These variants often arise from different systems (e.g., French-based vs. English-based) or historical usage. Consistency is key: choose one system and stick to it.
Pronunciation Limitations
Romanization inevitably loses some phonemic nuances. For instance, the distinction between the two gh sounds (غ /ɣ/ and ق /ɢ/) is collapsed in Unipers. The glottal stop (ع) is often omitted, leading to ambiguity (e.g., ma’ni vs. mani). Vowel length is not always marked in informal romanization. To mitigate these issues, learners should consult IPA transcriptions or audio resources. For academic work, use a reversible system like ALA-LC that preserves these distinctions.
“Consistent romanization is not merely a convenience; it is a bridge that allows languages to be studied, cataloged, and shared across linguistic boundaries without distortion.” — Dr. John R. Perry, Professor Emeritus of Persian, University of Chicago
Converter
Several online tools can convert Persian script to romanized text. The Unipers Converter (https://romanization.org/tools/unipers-converter/) supports real-time conversion and offers multiple output formats. The ALA-LC Converter (https://romanization.org/tools/ala-lc-converter/) is designed for academic use. For geographic names, the BGN/PCGN Converter (https://romanization.org/tools/bgn-pcgn-converter/) follows official standards. Usage tip: always verify the output with a native speaker or dictionary, as automated converters may misinterpret homographs.
FAQ
Which romanization system should I use for academic papers?
For academic papers in North America, use the ALA-LC system. It is reversible and widely accepted by libraries and journals. For German publications, use DMG.
Why are there multiple spellings for the same Persian word in English?
Different romanization systems (e.g., Unipers vs. ALA-LC) and historical conventions (e.g., French-based vs. English-based) produce variant spellings. Consistency within a document is key.
Can I use Unipers for official documents like passports?
No. Official documents typically follow BGN/PCGN or the specific system used by the issuing country. For Iranian passports, the romanization is based on BGN/PCGN.
Further reading
References
- Perry, J. R. (2005). A Tajik Persian Reference Grammar. Brill.
- Library of Congress. (1997). ALA-LC Romanization Tables: Persian. https://www.loc.gov/catdir/cpso/romanization/persian.pdf
- U.S. Board on Geographic Names. (2012). Romanization System for Persian (Farsi). https://geonames.nga.mil/gns/html/romanization.html
- Deutsche Morgenländische Gesellschaft. (1935). Die Transliteration der arabischen Schrift. DMG.
- ISO 233-3:2010. Information and documentation — Transliteration of Arabic characters into Latin characters — Part 3: Persian language.
Leave a Reply