[Unicode] Unicode 18.0.0 Tech Site | Site Map | Search
 

Unicode® 18.0.0

2026 September 16 (Announcement)

STATUS: This is a preliminary draft page for an upcoming release. Some details may be missing or incorrect, and some links may be wrong or broken. During the beta review period, feedback about errors on this page will be helpful and appreciated.

This page summarizes the important changes for the Unicode Standard, Version 18.0.0. This version supersedes all previous versions of the Unicode Standard.

A. Summary

Unicode 18.0 adds 13,007 characters, for a total of 172,808 characters. The new additions include three new scripts:

  • Proto-Cuneiform (numerals)
  • Jurchen
  • Seal (= “Small Seal”)

New Data Files for Unicode 18.0

  • JurchenSources.txt
  • SealSources.txt

Synchronization

Several other important Unicode specifications have been updated for Version 18.0. The following five Unicode Technical Standards are versioned in synchrony with the Unicode Standard, because their data files cover the same repertoire. All have been updated to Version 18.0:

Specification Scope Data Files
UTS #10, Unicode Collation Algorithm Sorting Unicode text UCA data
UTS #39, Unicode Security Mechanisms Reducing Unicode spoofing Security data
UTS #46, Unicode IDNA Compatibility Processing Compatible processing of non-ASCII URLs IDNA data
IDNA 2008 derived data
UTS #51, Unicode Emoji Emoji and their behavior Emoji data
UTS #58, Unicode Link Detection and Formatting: URLs and Email Addresses Linkification of URLs and email addresses Linkification data

Some of the changes in Version 18.0 and associated Unicode Technical Standards may require modifications to implementations. For more information, see the migration and modification sections of those specifications.

See Sections D through H below for additional details regarding the changes in this version of the Unicode Standard, its associated annexes, and the other synchronized Unicode specifications.

See the following resource links for general information about Unicode versions and other information about the Unicode Standard and other publications of the Unicode Consortium.

B. Technical Overview

Version 18.0 of the Unicode Standard consists of:

  • The core specification
  • The code charts for this version
  • The Unicode Standard Annexes
  • The Unicode Character Database (UCD)

The core specification gives the general principles, requirements for conformance, and guidelines for implementers. The code charts show representative glyphs for all the Unicode characters. The Unicode Standard Annexes supply detailed normative information about particular aspects of the standard. The Unicode Character Database supplies normative and informative data for implementers to allow them to implement the Unicode Standard.

Core Specification

Core Specification (HTML) Authoritative version. Chapter-by-chapter web pages.
Core Specification (PDF, 13 MB) Archival copy. Consolidated single file.

The core specification for Version 18.0 is available for browsing online as per-chapter web pages. Because the full table of contents for the core specification is provided, with interactive links, no separate bookmarks page is provided, nor are separate chapter links provided directly in this summary page for the Unicode Standard. Anchors for chapters, sections, tables, and figures in the core specification are shown with the convention of a "#" in the left margin of the heading or caption. Those anchors can be clicked on to provide custom bookmarks to any particular portion of the text, down to the level of subsections.

The HTML version of the core specification is authoritative. However, for convenience of reference, an archival version of the core specification is also available, formatted as a single pdf.

Code Charts

Several sets of code charts are available. They serve different purposes:

Chart Type Description
Code Charts Block-by-block code charts for Version 18.0.0. The charts are organized by scripts and blocks for easy reference.
Delta Code Charts These charts show the new blocks and any blocks in which characters were added specifically for Unicode 18.0.0. The new characters and any major updates to the representative glyphs are visually highlighted in these charts.
Consolidated Code Charts (179 MB) These charts are distributed as a single large pdf file containing the entire set of characters, names and representative glyphs at the time of publication of Unicode 18.0.0.
Auxiliary Code Charts The auxiliary charts display information about collation and casing for the repertoire of the scripts for this release.

The block-by-block, delta, and consolidated code charts are a stable part of this release of the Unicode Standard. They will never be updated. The auxiliary code charts are provided for information, and have no stability guarantees.

Index Type Description
Character Name Index An interactive character name index with incremental matching, which also enables lookup by names list annotations, chart headings, and code points.

Han Radical-Stroke Indices

There are a number of radical-stroke indices available to assist in the lookup of Han ideographs in the code charts.

Index Type Description
Interactive An interactive CJK character lookup page that supports lookup either by code point or by radical and stroke values.
IICore (4.6 MB) A static radical-stroke index PDF file limited to only the IICore repertoire. (This RS index is seldom updated.)
Unihan Core 2020 (8.6 MB) A static radical-stroke index PDF file limited to only the Unihan Core 2020 repertoire. (This RS index is seldom updated.)
Complete (35 MB) A static radical-stroke index PDF file that covers the entire CJK ideograph repertoire for Unicode 18.0.
Complete A static data file that corresponds to the complete radical-stroke index for Unicode 18.0.

The complete radical-stroke index is a stable part of this release of the Unicode Standard. It will never be updated.

Unicode Standard Annexes

STATUS: During the alpha review and beta review periods, links to individual UAXes (or UTSes) point to the proposed update for that document, if any. If no proposed update has been posted for the document, links point to the last published version of the document, for reference.

Links to the individual Unicode Standard Annexes for this version are available in Section I, List of Components below. The summary list of significant changes in the content of each Unicode Standard Annex for Version 18.0 can be found in Section G, Changes in the Unicode Standard Annexes below.

Unicode Character Database

STATUS: During the beta review period, the draft of UCD data includes data for the complete, planned character repertoire of Unicode 18.0, including all data changes approved by UTC for version 18.0.

Data files for Version 18.0 of the Unicode Character Database are available. Detailed documentation about the data files can be found in UAX #44, Unicode Character Database.

Version References

Version 18.0.0 of the Unicode Standard should be referenced as:

The Unicode Consortium. The Unicode Standard, Version 18.0.0, (South San Francisco: The Unicode Consortium, 2026. ISBN 978-1-936213-36-8)
https://www.unicode.org/versions/Unicode18.0.0/

The terms “Version 18.0” or “Unicode 18.0” are abbreviations for the full version reference, Version 18.0.0.

The citation and permalink for the latest published version of the Unicode Standard is:

The Unicode Consortium. The Unicode Standard.
https://www.unicode.org/versions/latest/

A complete specification of the contributory files for Unicode 18.0 is found below in Section I, List of Components. For examples of how to cite particular portions of the Unicode Standard, see also the Reference Examples.

Errata

Errata incorporated into Unicode 18.0 are listed by date in a separate table. For corrigenda and errata after the release of Unicode 18.0, see the list of current Updates and Errata.

C. Stability Policy Update

A new property stability policy has been added for the ID_Compat_Math_Start and ID_Compat_Math_Continue properties. See the Character Encoding Stability Policies.

D. Textual Changes and Character Additions

Changes in the Unicode Standard Annexes are listed in Section G.

Character Assignment Overview

13,007 characters have been added. Most character additions are in new blocks, but there are also character additions to a number of existing blocks. For details, see the delta code charts.

New Blocks

The following blocks are newly defined in Version 18.0:

Range Block Name
11DF0..11DFF Bengali Supplement
12550..1268F Archaic Cuneiform Numerals
18E00..1919F Jurchen
191A0..191DF Jurchen Radicals
1D250..1D28F Musical Symbols Supplement
1DB00..1DBFF Miscellaneous Symbols and Arrows Extended
3D000..3FC3F Seal

Note that the block name for the Seal script is “Seal”. The Script property value is also “Seal”. However, the code chart headers and core specification documentation use the extended name “Small Seal”. This is intentional and is not a discrepancy to be reported.

E. Conformance Changes

The conformance section of the standard has been updated with new definitions and requirements regarding the use of variation selectors and variation sequences.

F. Changes in the Unicode Character Database

Some of the important impacts of data changes on implementations migrating from earlier versions of the standard are highlighted in Section M.

G. Changes in the Unicode Standard Annexes

In Version 18.0, some of the Unicode Standard Annexes have had significant revisions. The most important of these changes are listed below. For the full details of all changes, see the Modifications section of each UAX, linked directly from the following list of UAXes.

Unicode Standard Annex Changes
UAX #9
Unicode Bidirectional Algorithm
No significant changes in this version.
UAX #11
East Asian Width
The unassigned code point ranges in Section 6.1 were adjusted.
UAX #14
Unicode Line Breaking Algorithm
Rule LB12a was changed to disallow a break between BA and GL. The Line_Break assignment of FIGURE DASH and EN DASH was changed from HH to BA and SOFT HYPHEN from BA to HH for better linebreaking behavior for those characters.
UAX #15
Unicode Normalization Forms
No significant changes in this version.
UAX #24
Unicode Script Property
Added explanation of mixed-script script codes.
UAX #29
Unicode Text Segmentation
Rule GB9c (“Do not break within certain combinations with Indic_Conjunct_Break (InCB)=Linker.”) has been revised.
UAX #31
Unicode Identifiers and Syntax
The new scripts in Unicode 18.0 were added to Table 4, Excluded Scripts. A note was added indicating that the Property and Algorithms Working Group (PAG) is the primary point of contact for script reclassification, and pointing to the guidelines approved by the UTC.
UAX #34
Unicode Named Character Sequences
No significant changes in this version.
UAX #38
Unicode Han Database (Unihan)
The provisional kJapaneseNewVariant and kJapaneseOldVariant properties were added. The provisional properties kIRGDaeJaweon and kIRGKangXi were removed. The delimiter of 23 properties was updated. The syntax of the kGB5 property was updated. The descriptions of the kIRG_KSource, kIRG_UKSource, kJinmeiyoKanji, kOtherNumeric, and kTang properties were updated. The syntax and description of the kIRG_GSource property were updated. U+2B81E was added to the first table in Section 4.4.
UAX #41
Common References for Unicode Standard Annexes
All references were updated for Unicode 18.0.
UAX #42
Unicode Character Database in XML
New code point attributes, values, and patterns were added for Unicode 18.0.
UAX #44
Unicode Character Database
A new section was added to document UAX #60 and the data files for the large East Asian scripts it covers (Seal, Jurchen, Nushu, Tangut). Additions were made to Table 5 for the new data files JurchenSources.txt and SealSources.txt. The section regarding the directory structure for UCD and non-UCD files distributed under /Public/draft/ and versioned directories was rewritten.
UAX #45
U-Source Ideographs
No significant changes in this version.
UAX #50
Unicode Vertical Text Layout
Table 3 was updated to add a character that is now assigned the property value vo=Tu (U+1B168).
UAX #53
Unicode Arabic Mark Rendering
U+10EF4, U+10EF6, and U+10EF9 were added to MCM.
UAX #57
Unicode Egyptian Hieroglyph Database (Unikemet)
Many small corrections for details about the data and regex values.
UAX #60
Data for East Asian Scripts
New in this release.

H. Changes in Synchronized Unicode Technical Standards

There are also significant revisions in the Unicode Technical Standards whose versions are synchronized with the Unicode Standard. The most important of these changes are listed below. For the full details of all changes, see the Modifications section of each UTS, linked directly from the following list of UTSes.

Unicode Technical Standard Changes
UTS #10
Unicode Collation Algorithm
Jurchen and Small Seal were added to the table for computing implicit weights. Missing Tibetan contractions were added to DUCET. A discussion of the mapping and behavior for U+FFFE and U+FFFF was added. The Shift-Trimmed option was removed.
UTS #39
Unicode Security Mechanisms
Support was added for the Hntl script code. Rule A1 was tightened by removing the allowance for Transparent characters between ZWNJ and the following joining characters. The term “nonspacing mark” has been clarified, and some outdated text has been removed.
UTS #46
Unicode IDNA Compatibility Processing
No significant changes in this version.
UTS #51
Unicode Emoji
The guidance regarding emoji/text presentation selectors and default presentation style was updated.
UTS #58
Unicode Link Detection and Formatting: URLs and Email Addresses
No significant changes in this version.

I. List of Components

This section lists the components of Version 18.0.0 of the Unicode Standard. The version numbering and the role of each component are explained in Versions of The Unicode Standard.

Core Specification
Authoritative HTML
Archival PDF: UnicodeStandard-18.0.pdf (size: 13 MB)
Code Charts and Radical-Stroke Index
Code Charts (size: 179 MB)
Radical-Stroke Index (size: 35 MB)
Radical-Stroke Index data
Unicode Standard Annexes
UAX #9: Unicode Bidirectional Algorithm
UAX #11: East Asian Width
UAX #14: Unicode Line Breaking Algorithm
UAX #15: Unicode Normalization Forms
UAX #24: Unicode Script Property
UAX #29: Unicode Text Segmentation
UAX #31: Unicode Identifiers and Syntax
UAX #34: Unicode Named Character Sequences
UAX #38: Unicode Han Database (Unihan)
UAX #41: Common References for Unicode Standard Annexes
UAX #42: Unicode Character Database in XML
UAX #44: Unicode Character Database
UAX #45: U-Source Ideographs
UAX #50: Unicode Vertical Text Layout
UAX #53: Unicode Arabic Mark Rendering
UAX #57: Unicode Egyptian Hieroglyph Database (Unikemet)
UAX #60: Data for East Asian Scripts
Unicode Character Database
https://www.unicode.org/Public/18.0.0/
(Note that not all subdirectories of this versioned data directory are formally part of the UCD. See Table 2a in UAX #44 for details.)
Documentation
Index.txt
NamesList.html
ReadMe.txt
Core Data
ArabicShaping.txt
BidiBrackets.txt
BidiMirroring.txt
Blocks.txt
CJKRadicals.txt
CompositionExclusions.txt
DoNotEmit.txt
EastAsianWidth.txt
EmojiSources.txt
EquivalentUnifiedIdeograph.txt
HangulSyllableType.txt
IndicPositionalCategory.txt
IndicSyllabicCategory.txt
Jamo.txt
LineBreak.txt
NameAliases.txt
NamedSequences.txt
NamedSequencesProv.txt
NamesList.txt
NormalizationCorrections.txt
PropertyAliases.txt
PropertyValueAliases.txt
PropList.txt
Scripts.txt
ScriptExtensions.txt
SpecialCasing.txt
StandardizedVariants.txt
UnicodeData.txt
VerticalOrientation.txt
Data for UAX #38: Unihan Database (Unihan.zip)
Unihan_DictionaryIndices.txt
Unihan_DictionaryLikeData.txt
Unihan_IRGSources.txt
Unihan_NumericValues.txt
Unihan_OtherMappings.txt
Unihan_RadicalStrokeCounts.txt
Unihan_Readings.txt
Unihan_Variants.txt
Data for UAX #45
USourceData.txt
USourceGlyphs.pdf
USourceRSChart.pdf
Data for UAX #57
Unikemet.txt
Data for UAX #60
JurchenSources.txt
NushuSources.txt
SealSources.txt
TangutSources.txt
Derived Data
CaseFolding.txt
DerivedAge.txt
DerivedCoreProperties.txt
DerivedNormalizationProps.txt
Extracted Data
DerivedBidiClass.txt
DerivedBinaryProperties.txt
DerivedCombiningClass.txt
DerivedDecompositionType.txt
DerivedEastAsianWidth.txt
DerivedGeneralCategory.txt
DerivedJoiningGroup.txt
DerivedJoiningType.txt
DerivedLineBreak.txt
DerivedName.txt
DerivedNumericType.txt
DerivedNumericValues.txt
Conformance Test Data
BidiCharacterTest.txt
BidiTest.txt
NormalizationTest.txt
Auxiliary Data for UAX #14 and UAX #29
GraphemeBreakProperty.txt
GraphemeBreakTest.txt
LineBreakTest.txt
SentenceBreakProperty.txt
SentenceBreakTest.txt
WordBreakProperty.txt
WordBreakTest.txt
Documentation for Auxiliary Data
GraphemeBreakTest.html
LineBreakTest.html
SentenceBreakTest.html
WordBreakTest.html
Emoji Data
emoji-data.txt
emoji-variation-sequences.txt

M. Implications for Migration

STATUS: During the beta review period, the following section is incomplete. For issues which may impact migration, see the detailed notes presented under Notable Issues for Beta Reviewers on the 18.0 beta review page.

There are a significant number of changes in Unicode 18.0 which may impact implementations upgrading to Version 18.0 from earlier versions of the standard. The most important of these are listed and explained here, to help focus on the issues most likely to cause unexpected trouble during upgrades.

Core Specification Changes

New content has been added for Unicode 18.0, and many other improvements have been made to the text.

Wording in the core specification of earlier versions was not completely clear regarding variation sequences and conformance. To provide greater clarity, the text describing variation sequences and related conformance requirements has been revised. See Section 3.6.2 in the core specification for details. There are some related changes to Section 5.20 and Section 23.4, as well.

Script-related Changes

  • There are three new scripts encoded in Unicode 18.0. Two of these scripts, Jurchen and Seal (= “Small Seal”), are large ideographic scripts.
  • For Proto-Cuneiform, only the archaic numerals are added in Unicode 18.0; encoding of additional signs for non-numerics is anticipated for a future version.

General Character Property Issues

  • Two new UCD data files, JurchenSources.txt and SealSources.txt, are associated with two new scripts, Jurchen and Seal (= “Small Seal”). The latter data file includes kSEAL_THXSrc, kSEAL_CCZSrc, kSEAL_DYCSrc, and kSEAL_QJZSrc as new normative properties.

Reorganization of Some Data Files

  • The /Public/18.0.0/ directory now contains a linkification subdirectory, which contains the data files associated with UTS #58, Unicode Link Detection and Formatting: URLs and Email Addresses.
  • The /Public/18.0.0/charts/ directory has been substantially expanded to include all of the versioned block charts and the auxiliary charts, as well as the navigation pages for code charts.

Other UCD Issues

  • The UCD file Index.txt has not been updated for Unicode Version 18.0; the file published in Version 18.0 is identical to the file for Version 17.0. This file is obsolete and will be removed altogether in Unicode Version 19.0.
  • The permuted index generated from the manually-curated Index.txt has been replaced with a search tool based on the names list and other UCD data. This search tool is available where the static index used to be, at https://www.unicode.org/charts/charindex.html.

Security and Identifier-related Issues (See UAX #31 and UTS #39.)

  • Many lines of unused data in confusables.txt have been removed. These lines corresponded to characters that never appear in NFD form and thus were never exercised by the Confusable Detection algorithm in UTS #39.
  • Several corrections and additions to confusables data have been made, incorporating parts of a large backlog of public contributions to confusables data. Emphasis has been on confusable pairs with at least one side having Identifier_Status=Allowed.
  • The entries in the file IdentifierType.txt are no longer grouped by sets of values, but are instead ordered by code point; this is similar to the change made to ScriptExtensions.txt in Unicode Version 16.0.
  • Support was added for the Hntl script code.
  • Rule A1 was tightened by removing the allowance for Transparent characters between ZWNJ and the following joining characters. This can affect the processing result for an implementation using Rule A1.

Segmentation (See UAX #14 and UAX #29.)

  • There has been a change to line breaking rule LB12a and to the Line_Break property assignments of some dashes and hyphens, including SOFT HYPHEN. This fixes a regression in the behaviour of an EN DASH set aside from preceding text by a NO-BREAK SPACE that had been introduced in Unicode Version 5.1.
  • Grapheme cluster breaking rule GB9c, which binds Indic conjuncts, has been changed to eliminate the requirement for context before the linker. This improves grapheme cluster breaking for Balinese.
  • The derivation of the Indic_Conjunct_Break property has been changed, correcting a regression in the behaviour of Zanabazar Square grapheme cluster breaking that had been introduced in Unicode Version 17.0.
  • An additional change to the derivation of the Indic_Conjunct_Break has been made: the derivation uses Script_Extensions instead of Script. This improves grapheme cluster breaking for Bengali.

Numeric Property Issues

  • The new Archaic Cuneiform Numerals block contains a very large set of numeric characters. Specialist implementations dealing with cuneiform text should be aware of these characters, which also pose challenges for formatting and for font design.

CJK/Unihan Changes (See UAX #38.)

  • There have been many changes to various Unihan properties. See Unicode Character Database, and UAX #38, Unicode Han Database (Unihan) for further details on these changes.
  • One more unified ideograph was added to the CJK Unified Ideographs Extension D block, extending the assigned range to U+2B81E, instead of U+2B81D. Implementations which use hard-coded ranges for CJK unified ideographs need to be updated.
  • The RSIndex.txt data file now uses a semicolon-delimited syntax. See the documentation in the file header of RSIndex.txt for details.

Standardized Variation Sequences

  • 17 variation sequences have been added for CJK strokes.
  • 10 variation sequences have been added for various math operators and a script small z.
  • Many changes have been made to standardized variation sequences for Mongolian. These sequences are now synchronized with the Chinese standards for Mongolian. See UTN #57, "Encoding and Shaping of the Mongolian Script" for more details.

Changes to Code Charts

  • The chart fonts of Armenian and Khitan Small Script have been changed to Noto Serif Armenian and Khitan Small Linear.
  • The glyph of U+06C4 ARABIC LETTER WAW WITH RING has been corrected per 187-C13.
  • There are a number of other Han glyph updates.
  • Other glyph updates are listed explicitly in the delta charts index page.
  • The two code charts for Egyptian hieroglyphs contain extensive functional and phonetic information derived from the data file Unikemet.txt, and have notable further updates for Version 18.0.
  • The version-specific code charts are now distributed under the same /Public/18.0.0/charts/ directory as the consolidated charts, and an explicitly versioned code charts index page is available to help with access to them.
  • The auxiliary charts have been completely reconstituted, and are now also distributed under the /Public/18.0.0/charts/ directory. See Collation and Casing Charts.
  • To explain these changes, the Code Charts Help page has been substantially reorganized and is also now explicitly versioned along with the charts.

Collation-related Changes

The Default Unicode Collation Element Table (DUCET) was updated to the Unicode 18.0.0 repertoire for UCA 18.0.0.

The two new large siniform ideographic scripts, Jurchen and Seal (= “Small Seal”), are given collation weights using implicit weighting. This requires a small change to the implicit weighting algorithm, to add new base weights. Implementations of UCA should be aware of this change.

Several special mappings have been added to the DUCET. These had already been added in the CLDR root collation tailoring.

The Shift-Trimmed option was not previously recommended. It has now been removed completely from the UCA specification.

Linkification-related Issues

The recently added UTS #58, Unicode Link Detection and Formatting: URLs and Email Addresses is published in synchronization with the Unicode Standard starting with Unicode 18.0.0.

Emoji Changes

For details about emoji changes, see the Unicode 18.0 emoji charts and Emoji Recently Added, v18.0.

 


Access to Copyright and terms of use