Persian spelling did not change overnight when the Academy issued its standards. Older manuscripts, lithographs, typewritten pages, early websites, and even current publishers display competing practices: affixes may be solid or spaced, ezafe after final ه may appear as a superscript sign, and Arabic ك/ي may hide inside otherwise Persian digital text. Academy conventions supply a shared modern default; they do not make every earlier document defective.
The editor's first decision is therefore not “Which form is newer?” but “What kind of text am I producing?” A modern reading edition normally regularizes spelling. A diplomatic transcription records what the witness actually shows. A quotation must not be silently modernized until it says something the source never said.
Standardization is a convention, not a time machine
Modern (written-standard) Iranian Persian uses Persian ک and ی, standard ZWNJ boundaries, and this guide's خانهی form for ezafe after silent ه. Before these conventions became widely supported in fonts and keyboards, printers solved the same problems with the equipment available to them. Variation also continued after official recommendations because house styles, habits, and software changed at different speeds.
| Editorial feature | Legacy or competing form | Guide's modern form | Nature of the change |
|---|---|---|---|
| imperfective prefix | می رود / میرود | میرود | space or fusion becomes ZWNJ |
| plural suffix | کتابها / کتاب ها | کتابها | morpheme boundary becomes ZWNJ |
| ezafe after silent ه | خانهٔ من / خانۀ من | خانهی من | superscript sign or precomposed glyph becomes ه + ZWNJ + ی |
| kaf and yeh in digital text | ك / ي / ى | ک / ی | Arabic or legacy code point becomes Persian code point |
| compound boundary | جستجو | جستوجو | lexical parts made visible with ZWNJ |
The right column is the production norm for this guide. The middle column is evidence you must recognize, not a menu of forms to mix casually.
The modern boundary system is introduced step by step on Spaces, half-spaces, and the ZWNJ.
این کتاب در سال ۱۳۴۸ چاپ شده است.
in ketâb dar sâl-e 1348 châp shode ast.
This book was printed in the year 1348 of the Solar Hijri calendar.
فاصلهگذاری این چاپ با معیار امروز فرق دارد.
fâsele-gozâri-ye in châp bâ me'yâr-e emruz farq dârad.
The spacing in this edition differs from today's standard.
Historical spacing and the rise of ZWNJ
Paper manuscripts do not contain Unicode characters. A scribe could leave a small gap or interrupt joining, but no hidden U+200C sat between the letters. Metal type, typewriters, and early computer systems likewise offered uneven ways to represent intermediate boundaries. That history helps explain می رود، میرود، and میرود without making them equivalent in a modern data field.
Current standardization makes morphological structure visible: میخواند is one verb form with a protected prefix boundary, not the phrase می خواند and not an undivided میخواند. The same logic applies to نمیدانم، کتابها، بزرگتر، and many compounds.
در متن امروزی میخواند را با نیمفاصله مینویسیم.
dar matn-e emruzi mi-khânad râ bâ nim-fâsele mi-nevisim.
In modern text, we write mi-khânad with a half-space.
کتابهای قدیمی را جداگانه فهرست کردهاند.
ketâb-hâ-ye qadimi râ jodâgâne fehrest karde-and.
They have catalogued the old books separately.
Ezafe after final he
After consonants, ezafe is normally heard but not written. After the silent final ه of خانه, however, the linker needs a visible carrier. Older and competing typography commonly uses خانهٔ or the precomposed-looking خانۀ. This guide uses خانهی, exactly ه + ZWNJ + Persian ی, because its letters and boundary remain explicit and searchable.
خانهی پدریاش هنوز در همان کوچه است.
khâne-ye pedari-ash hanuz dar hamân kuche ast.
His/Her family home is still on the same alley.
نسخهی اصلی در کتابخانه نگهداری میشود.
noskhe-ye asli dar ketâbkhâne negahdâri mi-shavad.
The original manuscript is kept in the library.
This is an orthographic choice, not a change in syntax: all three historical display strategies represent the same ezafe relation and the same standard pronunciation -ye. When transcribing a source diplomatically, record its form. When modernizing, select one declared convention consistently.
Legacy kaf and yeh are an encoding layer
Arabic ك (U+0643), ي (U+064A), and ى (U+0649) occur frequently in older Persian digital files. Modern Iranian Persian requires ک (U+06A9) and ی (U+06CC). Unlike a visible spelling reform, this replacement often corrects an encoding mismatch created by keyboards or software.
Even here, context matters. Preserve Arabic code points inside an Arabic quotation, and do not alter an archival byte-for-byte field. In a normalized Persian reading field, map them deliberately. Ordinary Unicode NFC does not perform this Persian-specific conversion.
کلیدواژههای متن قدیمی را برای جستوجو یکسانسازی کردیم.
kelidvâzhe-hâ-ye matn-e qadimi râ barâ-ye jost-o-ju yeksân-sâzi kardim.
We normalized the old text's keywords for searching.
Diplomatic, normalized, and critical editions
A diplomatic transcription reproduces a witness's spelling, abbreviations, and uncertain readings as closely as the project's method allows. A normalized reading edition gives readers consistent contemporary spelling. A critical edition may establish a text from several witnesses and report variants in an apparatus. These are different scholarly products.
Use separate fields whenever possible:
- source image or archival string;
- diplomatic transcription;
- normalized reading text;
- searchable key;
- editorial notes recording non-mechanical intervention.
صورت اصلی نام در متن و صورت معیار در نمایه آمده است.
surat-e asli-ye nâm dar matn va surat-e me'yâr dar namâye âmade ast.
The original form of the name appears in the text and the standard form in the index.
در نقل مستقیم، املای منبع را بیتوضیح تغییر ندهید.
dar naql-e mostaqim, emlâ-ye manba' râ bi-towzih taghyir nadahid.
In a direct quotation, do not change the source spelling without explanation.
Register and period
Academy-based spelling is expected in contemporary (formal), (academic), educational, and edited (written-standard) prose. Informal messages vary more, especially in their omission of ZWNJ, but this is not a separate Tehran grammar. Superscript ezafe forms remain visible in respected publishing traditions, while Arabic kaf and yeh in Persian usually signal legacy input rather than elevated register. Historical forms belong to their period; call them (obsolete) only when modern Iranian writers no longer select them except in reproduction.
Common Mistakes
1. Calling every older spelling an error
❌ نویسندهی قدیمی نیمفاصله را فراموش کرده است.
Misleading — a manuscript scribe could not omit a Unicode character that did not yet exist.
✅ کاتب میان دو جزء فاصلهی کوچکی گذاشته است.
kâteb miyân-e do joz' fâsele-ye kucheki gozâshte ast.
The scribe left a small gap between the two parts.
2. Mixing old and modern boundary conventions accidentally
❌ می رود اما نمیماند.
Inconsistent in a normalized edition — one verbal prefix is spaced and the other uses ZWNJ.
✅ میرود اما نمیماند.
mi-ravad ammâ nemi-mânad.
He/She goes but does not stay.
3. Using Arabic code points as though they marked an old register
❌ اين كتاب قديمی است.
Incorrect normalized Persian encoding — this string contains Arabic ي and ك instead of the Persian letters.
✅ این کتاب قدیمی است.
in ketâb qadimi ast.
This book is old.
4. Modernizing a quotation silently
❌ املای نقل را عوض کردم و همان را اصل نامیدم.
Incorrect editorial practice — the altered wording is being misrepresented as the source.
✅ املای نقل را حفظ کردم و تغییرها را توضیح دادم.
emlâ-ye naql râ hefz kardam va taghyir-hâ râ towzih dâdam.
I preserved the quotation's spelling and explained the changes.
5. Writing a short vowel as a new letter during modernization
❌ کیتابهای قدیمی
Incorrect — modernization does not write the short e of ketâb with ی.
✅ کتابهای قدیمی
ketâb-hâ-ye qadimi
old books
Key Takeaways
Modern standards regularize boundaries, ezafe spelling, and Persian character encoding, but historical evidence remains evidence. Identify the edition you are making, keep source and normalized layers distinct, apply the modern convention consistently within the normalized layer, and document every change that goes beyond safe encoding cleanup.
Now practice Farsi
Reading grammar gets you part of the way. The exercises are where it sticks — free, no signup needed.
Start learning Farsi →Related Topics
- Spaces, Half-Spaces, and the Zero-Width Non-JoinerA2 — Learn how Persian distinguishes word spaces, invisible half-spaces, and natural cursive breaks inside a word.
- Final He and the Written Ezâfe Form هیB1 — Distinguish final he pronounced as -e from final /h/, and write ezâfe after -e with the house form هی.
- Word Division in Compounds and DerivativesC1 — Choose solid spelling, ZWNJ, or a full space in Persian compounds and derivatives, and search safely across real-world variants.
- Unicode Forensics for Persian EditorsC2 — Diagnose invisible controls, presentation forms, confusable Persian characters, and clean text without changing its linguistic skeleton.
- Editing Mixed-Period PersianC2 — Separate transcription, normalized spelling, supplied reading, punctuation, annotation, and modern paraphrase without falsifying a historical source.
- Modernizing Classical Text into a Different MeaningC2 — Modernize classical Persian responsibly by separating transcription, normalization, interpretation, and paraphrase while preserving morphology, syntax, meter, and variants.