pyunormalize 18.0.0


pip install pyunormalize

  Latest version

Released: Sep 26, 2026


Meta
Author: Marc Lodewijck
Requires Python: >=3.9

Classifiers

Development Status
  • 5 - Production/Stable

Intended Audience
  • Developers

Operating System
  • OS Independent

Programming Language
  • Python :: 3
  • Python :: 3 :: Only
  • Python :: 3.9
  • Python :: 3.10
  • Python :: 3.11
  • Python :: 3.12
  • Python :: 3.13
  • Python :: 3.14

Topic
  • Software Development :: Internationalization
  • Software Development :: Localization
  • Software Development :: Libraries :: Python Modules
  • Text Processing
  • Text Processing :: Linguistic
  • Utilities

pyunormalize

The pyunormalize package is a pure-Python, zero-dependency implementation of the Unicode normalization algorithm.

It supports the four standard Unicode normalization forms:

  • NFC
  • NFD
  • NFKC
  • NFKD

The package uses generated lookup tables derived from the Unicode Character Database (UCD) for Unicode 18.0.0, released in September 2026. It is therefore independent of the unicodedata module and the specific Unicode database version bundled with the Python runtime on which it is installed.

All four normalization forms are tested against the official Unicode NormalizationTest.txt test file.

Requirements

Python 3.9 or newer.

Installation

Install the package with:

pip install pyunormalize

Upgrade to the latest version with:

pip install --upgrade pyunormalize

Public API

The public API exposes the four normalization functions and a generic dispatcher:

from pyunormalize import NFC, NFD, NFKC, NFKD, normalize

It also exposes constants for the Unicode version in use:

from pyunormalize import UCD_VERSION, UNICODE_VERSION

Note that UCD_VERSION and UNICODE_VERSION are aliases referring to the same version string ('18.0.0').

Usage examples

Below are practical examples illustrating how to normalize Unicode strings using either dedicated functions or the generic dispatcher.

Dedicated functions

Use the convenience functions NFC, NFD, NFKC, or NFKD for direct normalization:

from pyunormalize import NFC, NFD, NFKC, NFKD

text = "désaffiliât"
assert text == NFC(text)

def hex_repr(string):
    return " ".join([f"{ord(c):04X}" for c in string])

print(f"NFD  : {hex_repr(NFD(text))}")
print(f"NFKD : {hex_repr(NFKD(text))}")

print(f"NFC  : {hex_repr(NFC(text))}")
print(f"NFKC : {hex_repr(NFKC(text))}")

Output:

NFD  : 0064 0065 0301 0073 0061 FB03 006C 0069 0061 0302 0074
NFKD : 0064 0065 0301 0073 0061 0066 0066 0069 006C 0069 0061 0302 0074
NFC  : 0064 00E9 0073 0061 FB03 006C 0069 00E2 0074
NFKC : 0064 00E9 0073 0061 0066 0066 0069 006C 0069 00E2 0074

Generic normalize function

When the normalization form is specified dynamically at runtime, use normalize(form, text):

from pyunormalize import normalize

text = "fluffiness"

# The `form` parameter accepts "NFC", "NFD", "NFKC", or "NFKD" (case-sensitive)
nfkd_text = normalize("NFKD", text)

print(nfkd_text)

Output:

fluffiness

Related resources

This implementation is based on the following resources:

Changelog

See the CHANGELOG for the latest updates and changes.

Licenses

The code is available under the terms of the MIT License.

The use of Unicode data files is governed by the UNICODE TERMS OF USE. Further specifications of rights and restrictions pertaining to the use of the Unicode data files and software can be found in the Unicode License v3, a copy of which is included as UNICODE-LICENSE.

Wheel compatibility matrix

Platform Python 3
any

Files in release

Extras:
Dependencies: