Recent Posts
Archives

Posts Tagged ‘Normalization’

PostHeaderIcon [PyConUS2025] Why `len(‘😶‍🌫️’) == 4` and Other Unexpected Behaviors of Python Strings

Lecturer

Marie Roald is a researcher, data scientist, and educator affiliated with the Norwegian Language Bank at the National Library of Norway. She has more than eight years of experience teaching Python to secondary-school students, teachers, and professionals, and is a co-founder and organizer of PyLadies Oslo. Yngve Mardal Moe is an experienced Python educator, developer, and data-science consultant who previously led the redesign of an introductory Python course at the Norwegian University of Life Sciences; he currently serves as tech lead on automation projects for the Norwegian power grid. Together they bring complementary perspectives from language technology and software engineering to the practical difficulties of Unicode handling.

Abstract

Python strings appear simple until everyday operations—length measurement, equality testing, case conversion, and slicing—produce counter-intuitive results. This presentation traces those surprises to the underlying Unicode encoding model, the distinction between code points and grapheme clusters, the existence of multiple normalized forms, and the incomplete implementation of locale-sensitive operations in the language. The speakers supply concrete recommendations for robust comparison, normalization, and length measurement that avoid the most common pitfalls.

Encoding, Code Points, and the Limits of Naïve Operations

A computer stores only bits; any text representation is therefore a mapping from abstract characters onto sequences of numbers called code points. Early seven-bit ASCII proved insufficient for the world’s writing systems, prompting a proliferation of national encodings and, eventually, the Unicode standard. Unicode presently defines more than a million code points and is transmitted most commonly as the variable-length UTF-8 encoding. Python itself stores strings in one of three internal widths—1, 2, or 4 bytes per code point—chosen according to the highest code point present in the string.

Because a single visible character (a grapheme) may be composed of several code points, the built-in len function counts code points rather than user-perceived characters. The rainbow-flag emoji, for example, comprises a white flag, a variation selector, a zero-width joiner, and a rainbow emoji—four code points that render as one glyph. The same phenomenon appears in ordinary text: the Norwegian letter “å” may be stored either as the precomposed code point U+00E5 or as the sequence “a” plus a combining ring (U+030A). Equality tests and slicing that operate on the raw sequence therefore diverge from human expectations.

Case conversion is similarly subtle. The German sharp S (“ß”) upper-cases to “SS”, so a naïve round-trip through str.upper and str.lower fails to restore the original spelling. The Unicode-recommended solution is case folding (str.casefold), which maps characters to a canonical caseless form and correctly handles the 297 known special cases. Even case folding, however, is incomplete for languages such as Turkish, whose dotted and dotless “I” require locale-aware rules that Python does not implement.

Normalization, Security, and Practical Recommendations

Unicode defines four normalization forms. NFC and NFD perform canonical composition and decomposition; NFKC and NFKD additionally map compatibility characters (superscripts, stylistic variants, fractions) onto their plain counterparts. Normalization is essential before comparison or hashing: two strings that look identical to a user may otherwise compare unequal. Python’s identifier parser already applies NFKC, which is why “fancy” mathematical letters are silently rewritten to ordinary ASCII identifiers—an amusing demonstration of the same machinery.

Homoglyphs (characters that look alike but occupy distinct code points) introduce security considerations. The Unicode Consortium publishes confusable lists that applications may consult when validating user names or domain names. Because there is no fixed upper bound on the number of code points that may form a single grapheme cluster, length limits expressed solely in graphemes remain vulnerable to pathological input; a practical defense is to impose both a grapheme limit and a modest code-point ceiling.

For everyday work the speakers recommend a short checklist: always exchange text as UTF-8; compare caselessly with casefold; normalize to a chosen form (usually NFC) before equality tests or storage; treat len and slicing as code-point operations and, when visual length matters, employ a library such as regex or PyICU that understands extended grapheme clusters; and remain aware that Unicode is still evolving and that Python’s support, while extensive, is not exhaustive. Written language is inherently complex; the apparent oddities of Python strings are simply the language’s honest reflection of that complexity.

Links: