Esperanto's normal modern spelling uses Unicode ĉ, ĝ, ĥ, ĵ, ŝ, and ŭ. On screen, one of these letters may look perfectly correct yet fail an exact search, sort unexpectedly, or compare unequal to an apparently identical copy. The usual cause is Unicode normalization: software can represent the same visible letter with different sequences of code points.
This is a computing issue, not free orthographic variation. A reader sees the same Esperanto letter and pronounces the same sound, but software may see different data. Reliable systems therefore preserve the proper letters, normalize their encoding, and decide explicitly how broad each kind of search should be.
One visible letter, two possible encodings
Unicode can represent ĉ in at least two canonically equivalent ways:
| Form | Underlying sequence | Typical label |
|---|---|---|
| ĉ | one precomposed LATIN SMALL LETTER C WITH CIRCUMFLEX | NFC-style |
| c + ◌̂ | LATIN SMALL LETTER C plus COMBINING CIRCUMFLEX ACCENT | NFD-style |
The same issue applies to ĝ, ĥ, ĵ, ŝ, and to ŭ, whose second representation uses a combining breve. Canonical normalization converts equivalent sequences to a consistent form. NFC generally composes them; NFD generally decomposes them. Because all six Esperanto letters have precomposed Unicode characters, NFC is a practical storage and publication default (technical).
Mi serĉas la vorton ‘ĉirkaŭ’.
I am searching for the word ‘around.’
The visible word ĉirkaŭ may contain precomposed letters, decomposed sequences, or one of each. Normalize both the stored text and the query before exact comparison.
La dosiero aspektas ĝuste, sed la serĉilo ne trovas ĝin.
The file looks correct, but the search engine does not find it.
That is the classic symptom: display rendering combines the marks for the eye, while an unnormalized search compares raw sequences.
Orthographic contrast must survive search
English-oriented search tools often remove diacritics to make café match cafe. Applying that policy blindly to Esperanto collapses letters that the alphabet deliberately distinguishes. celo means “aim”; ĉelo means “cell.” juro means “law”; ĵuro means “oath.”
La celo estas renovigi la ĉelon antaŭ vendredo.
The aim is to renovate the cell before Friday.
An accent-insensitive index that reduces both words to the same sequence loses meaningful information. It may be useful as a secondary “fuzzy” layer, but it must not replace exact orthographic search.
La atestanto faris ĵuron antaŭ ol la advokato klarigis la juron.
The witness took an oath before the lawyer explained the law.
The contrast between ĵuron and juron is lexical. A search interface may offer “include spelling variants,” but results should still display the original spelling and preferably rank exact matches first.
Mi ne serĉas ‘sato’; mi serĉas ‘ŝato’.
I am not searching for ‘satiety’; I am searching for ‘liking.’
This negative context shows why mark-stripping is not a harmless convenience. sato and ŝato are different word families.
A layered search model
A robust Esperanto search can proceed in widening layers:
- Exact normalized search: normalize query and text to NFC, and compare the actual Esperanto letters.
- Case-insensitive search: apply Unicode-aware case folding so Ĉ matches ĉ, without changing the letter to c.
- Legacy-input expansion: when the user explicitly requests it, interpret X-system forms such as cxirkaux or H-system forms such as chirkau.
- Broad fuzzy search: optionally ignore distinctions to rescue uncertain queries, while labeling and ranking these matches as approximate.
Ĉu la serĉilo trovas kaj ‘Ĝis’ kaj ‘ĝis’?
Does the search engine find both ‘Ĝis’ and ‘ĝis’?
A case-insensitive search should; an exact case-sensitive search may not. This is separate from normalization, and both are separate from legacy-spelling conversion.
Mi tajpis ‘gxardeno’, sed la rezultoj montris ankaŭ ‘ĝardeno’.
I typed ‘gxardeno,’ but the results also showed ‘ĝardeno.’
That behavior is useful in a learner dictionary (pedagogical/technical) when X-system query expansion is intentional. It would be dangerous as an unconditional replacement rule across source code, names, or URLs.
Token boundaries and whole-word search
Normalization never inserts or removes spaces. Esperanto compounds, affixes, grammatical endings, and punctuation still determine what counts as a token. A whole-word search for ĵurnal should not automatically equal a search for ĵurnalo, ĵurnalon, and ĵurnalistoj; that is a separate stemming or morphological-analysis decision.
La ĵurnalisto verkis artikolon por la ĵurnalo.
The journalist wrote an article for the newspaper.
Both words contain the root ĵurnal-, but their endings and structure differ. Unicode normalization makes their ĵ consistent; it does not analyze their morphology.
Hyphenated forms, apostrophes, quoted phrases, and filenames also need explicit policy. Normalize text without moving punctuation or crossing boundaries. In mixed-language records, apply Esperanto legacy-input expansion only to fields or queries identified as Esperanto.
— Kial la titolo aperas dufoje? — Unu kopio ne estis normaligita.
— Why does the title appear twice? — One copy was not normalized.
Duplicate detection often fails for exactly this reason. Normalize before generating search keys, slugs, uniqueness checks, or dictionary index entries.
Storage, display, and collation
For new content, accept proper Unicode input, normalize on ingestion, preserve the original human-readable text where editorial history matters, and store a normalized search key. The practical input options are covered in Typing Unicode Esperanto. Do not store only an X-system transliteration: it is an input convention (technical/informal), not the ordinary published form.
Sorting is another operation. Raw code-point order is not necessarily Esperanto alphabetical order. A locale-aware collator should treat the alphabet's distinct letters in the intended sequence, while a simple binary database index may not. Normalization prevents equivalent forms from splitting apart, but it does not by itself choose the right alphabetic collation.
En la indekso, ‘ĉambro’ kaj ‘celo’ devas resti apartaj kapvortoj.
In the index, ‘room’ and ‘aim’ must remain separate headwords.
Similarly, rendering fonts should position the circumflex and breve correctly. A font problem does not justify replacing the letters with plain ASCII; choose a font with suitable Latin coverage.
Common Mistakes
1. Stripping all diacritics before search
❌ La sistemo traktas ‘celo’ kaj ‘ĉelo’ kiel la saman vorton.
Incorrect search behavior: two distinct Esperanto words have been collapsed.
✅ La sistemo distingas ‘celo’ de ‘ĉelo’, sed povas proponi proksimumajn rezultojn aparte.
The system distinguishes ‘aim’ from ‘cell’ but may offer approximate results separately.
2. Assuming visual identity means byte identity
❌ La du aperoj de ‘ŝi’ aspektas same, do ili certe estas same koditaj.
Incorrect: identical rendering does not prove identical code-point sequences.
✅ Ni normaligu ambaŭ tekstojn antaŭ la komparo.
Let us normalize both texts before the comparison.
3. Using ASCII fallback as the database's canonical spelling
❌ La kapvorto estas ‘cxirkaux’.
Incorrect as the canonical published headword: this is X-system input.
✅ La kapvorto estas ‘ĉirkaŭ’.
The headword is ‘around.’
4. Letting conversion cross into foreign identifiers
❌ La serĉilo ŝanĝas ĉiun literon x en retadresoj.
Incorrect: legacy conversion is being applied outside Esperanto words.
✅ La serĉilo vastigas X-sisteman demandon nur en la Esperanta serĉkampo.
The search engine expands an X-system query only in the Esperanto search field.
Key Takeaways
Use full Unicode Esperanto orthography and normalize equivalent code-point sequences, usually to NFC. Preserve the distinction between c/ĉ, g/ĝ, h/ĥ, j/ĵ, s/ŝ, and u/ŭ. Build exact, case-insensitive, legacy-aware, and fuzzy search as explicit layers; normalization supports all of them but substitutes for none of them.
Now practice Esperanto
Reading grammar gets you part of the way. The exercises are where it sticks — free, no signup needed.
Start learning Esperanto →Related Topics
- Typing Unicode EsperantoA1 — Practical ways to enter, verify, and publish the six Esperanto letters with their correct Unicode characters.
- Converting Legacy Text SafelyA2 — A reliable workflow for converting H-system, X-system, and damaged Esperanto text into reviewed Unicode spelling.
- The Six Diacritic LettersA1 — Master ĉ, ĝ, ĥ, ĵ, ŝ, and ŭ as six independent Esperanto letters with distinct sounds and spelling roles.
- The X-SystemA1 — How to encode Esperanto’s six diacritic letters with x, convert the result safely, and know when the fallback is appropriate.
- One Letter, One SoundA1 — Understand Esperanto’s regular sound–spelling relationship without confusing a stable phoneme with mechanically identical pronunciation.