Converting old Esperanto text is not just a global search-and-replace operation. X-system spelling is usually easy to decode, but H-system spelling can be ambiguous, names and web addresses may contain literal x or h, and damaged character encoding can disguise the original letters entirely. Safe conversion separates detection, automated replacement, linguistic review, and final Unicode validation.
The goal is normal modern spelling with ĉ, ĝ, ĥ, ĵ, ŝ, and ŭ. Those are letters, not optional accents. Restoring them recovers Esperanto's productive one-letter–one-sound system and may also recover contrasts in meaning: celo “aim” is not ĉelo “cell,” and juro “law” is not ĵuro “oath.”
First identify what kind of source you have
Do not convert until you know the source convention. The dedicated guides to the H-system and the X-system explain their mappings in full. Three patterns are common:
| Source pattern | Likely system | Example |
|---|---|---|
| cx, gx, hx, jx, sx, ux | X-system (technical/informal input) | Cxu sxi venos? |
| ch, gh, hh, jh, sh; plain u for ŭ | H-system (technical/legacy) | Chu shi venos? |
| Ã, Ä, or replacement symbols near expected letters | Encoding damage, not a spelling system | Requires recovery from the original bytes |
Ĉu ŝi venos aŭ ne?
Will she come or not?
The same sentence may appear as Cxu sxi venos aux ne? in X-system text or Chu shi venos au ne? in H-system text. Detecting the convention before replacement prevents one system's rules from being applied to the other.
La ĝusta dosiero troviĝas en la malnova arkivo.
The correct file is in the old archive.
An X-system source would contain gxusta and trovigxas; an H-system source would contain ghusta and trovighas. A few repeated diagnostic sequences are usually more reliable than a single word.
Converting X-system text
Within a passage known to be Esperanto, the normal mapping is deterministic: cx → ĉ, gx → ĝ, hx → ĥ, jx → ĵ, sx → ŝ, ux → ŭ, with corresponding uppercase forms. The convenience comes from the fact that x is not a letter of the Esperanto alphabet, so it normally signals conversion.
Mi loĝas apud la stacidomo kaj ofte iras tien piede.
I live near the station and often go there on foot.
The X-source Mi logxas apud la stacidomo converts cleanly to loĝas. Apply the mapping inside Esperanto words, not blindly to every character sequence in the file.
Ŝi mendis teon kun citrono, sed sen sukero.
She ordered tea with lemon, but without sugar.
The X-source Sxi mendis... gives Ŝi mendis.... Initial Sx becomes uppercase Ŝ, not ŝ and not the two letters S plus x.
Protect material that is not ordinary Esperanto prose: URLs, email addresses, code, mathematical variables, product identifiers, and unassimilated names. In example.com/x, the x is literal. In the name Max, it is not an Esperanto conversion marker unless the editor has explicit evidence otherwise.
Converting H-system text
The consonant mappings ch → ĉ, gh → ĝ, hh → ĥ, jh → ĵ, sh → ŝ look straightforward, but h is itself an Esperanto letter. A sequence may cross a morpheme boundary or belong to a foreign name. Worse, H-system writes ŭ as plain u, so the source does not visibly distinguish u from ŭ.
La flughaveno estas proksima al la urbo.
The airport is close to the city.
The word flughaveno contains genuine g + h at the boundary between flug- and haveno. A carefully disambiguated H-system source may write flug'haveno or flug-haveno, but many legacy sources leave the boundary unmarked. In such a source, a blind gh → ĝ replacement would create the nonexistent fluĝaveno. Token-level linguistic review is therefore mandatory.
Ni aŭdis laŭtan bruon apud la aŭto.
We heard a loud noise near the car.
An H-system source may read Ni audis lautan bruon apud la auto. Nothing marks which plain u sequences should become ŭ. A dictionary, grammar, and sentence meaning must guide the restoration: aŭdis, laŭtan, aŭto.
A safe five-step workflow
1. Segment the document
Separate Esperanto prose from metadata, quotations in other languages, bibliographic records, code, and links. Record whether the edition should preserve historical capitalization or adopt a modern house style. Do not let conversion cross token boundaries: ĉi tie remains two words, while ĉiuj remains one.
2. Run constrained replacements
Use the detected system only in the Esperanto text fields. Preserve case: Cx or Ch → Ĉ; SX or SH in all-capitals material → Ŝ. Never run both H- and X-rules successively over an unreviewed mixed document.
ĈIUJ RAJTAS ENIRI.
EVERYONE MAY ENTER.
An all-caps X-source CXIUJ RAJTAS ENIRI or H-source CHIUJ RAJTAS ENIRI should yield ĈIUJ, with the word's capitalization preserved.
3. Review ambiguous tokens in context
Flag every H-system candidate, every plain au/eu sequence, unfamiliar proper name, and any replacement that produces a word absent from your reference lexicon. Automatic dictionary checking helps, but compounds and new formations are productive in Esperanto, so a missing dictionary entry is a review signal rather than proof of error.
La celo estas purigi la ĉelon sen damaĝi la aparaton.
The aim is to clean the cell without damaging the device.
Context confirms both celo and ĉelon. A converter that simply “adds missing accents” cannot decide this contrast without analyzing the words.
4. Normalize Unicode
Store the reviewed output in a consistent Unicode normalization form, normally NFC. A displayed ĉ can be encoded either as one precomposed code point or as c plus a combining circumflex; they can look identical while comparing differently in software. Normalization belongs after spelling conversion, not in place of it.
5. Compare, proofread, and sample aloud
Generate a diff against the source, inspect every changed token, then read a sample aloud. Because Esperanto spelling closely tracks pronunciation, forms that cannot be pronounced as intended often reveal a bad conversion. Finally search for leftover cx, sx, suspicious gh, or damaged characters.
— Ĉu la konvertado jam finiĝis? — Preskaŭ; mi ankoraŭ kontrolas la proprajn nomojn.
— Has the conversion finished yet? — Almost; I am still checking the proper names.
This connected exchange captures good editorial practice: conversion is not complete when the script stops; it is complete when the risky cases have been reviewed.
Mixed and damaged sources
Some archives combine conventions because several contributors used different keyboards. Convert each segment according to evidence, then regularize the final reading edition. In a scholarly quotation (academic), you may preserve the original spelling and identify it as H- or X-system. In ordinary publication (neutral/formal), use full Unicode.
Encoding damage requires a different remedy. If UTF-8 bytes were misread as another encoding, replacing visible gibberish may destroy recoverable information. Return to the original file or byte stream, identify the incorrect decoding, and reopen it correctly. Only after the actual characters are restored should you apply Esperanto spelling conversion.
La originala mesaĝo enhavis la frazon: ‘Ĝis revido!’
The original message contained the sentence: ‘See you again!’
If Ĝ appears as several strange symbols, do not guess each symbol manually. Repair the encoding first, then confirm that the recovered text says Ĝis revido!
Common Mistakes
1. Replacing every gh in an H-system file
❌ Ni renkontiĝos ĉe la fluĝaveno.
Incorrect: a blind replacement corrupted flughaveno at a compound boundary.
✅ Ni renkontiĝos ĉe la flughaveno.
We will meet at the airport.
2. Leaving plain u where H-system hid ŭ
❌ Ŝi audis la auton, sed ne la buson.
Incorrect conversion: the plain u sequences in aŭdis and aŭton have not been restored.
✅ Ŝi aŭdis la aŭton, sed ne la buson.
She heard the car, but not the bus.
3. Converting inside a protected identifier
❌ La retejo estas example.ĉ/x.
Incorrect: blind H-system conversion has altered the literal .ch domain.
✅ La retejo estas example.ch/x.
The website is example.ch/x.
The foreign address is quoted data and remains unchanged.
4. Mixing legacy and Unicode spelling in the final edition
❌ Ĉu shi ankau venos?
Incorrect edited prose: Unicode, H-system, and unconverted ŭ are mixed.
✅ Ĉu ŝi ankaŭ venos?
Will she come too?
Key Takeaways
Identify the source convention, protect non-prose data, convert only within known Esperanto text, and review every ambiguous H-system form. X-system is comparatively mechanical; H-system demands lexical and contextual judgment, especially for ŭ and genuine consonant-plus-h sequences. Preserve the original, normalize the reviewed Unicode output, and treat a human-checked diff—not a successful script run—as the end of conversion.
Now practice Esperanto
Reading grammar gets you part of the way. The exercises are where it sticks — free, no signup needed.
Start learning Esperanto →Related Topics
- The H-SystemA2 — How the Fundamento's fallback spelling represents Esperanto's six diacritic letters, and when modern writers should use it.
- The X-SystemA1 — How to encode Esperanto’s six diacritic letters with x, convert the result safely, and know when the fallback is appropriate.
- Typing Unicode EsperantoA1 — Practical ways to enter, verify, and publish the six Esperanto letters with their correct Unicode characters.
- Unicode Normalization and SearchA2 — Why visually identical Esperanto letters can behave differently in software, and how to store and search them reliably.
- The Six Diacritic LettersA1 — Master ĉ, ĝ, ĥ, ĵ, ŝ, and ŭ as six independent Esperanto letters with distinct sounds and spelling roles.