Skip to main content
Resources

String Similarity Evaluation

The String Similarity Evaluation (SSE) evaluates top-level domains to prevent user confusion and loss of confidence in the Domain Name System (DNS) that would result from delegation of visually similar strings in the root zone.

The SSE is required as a part of the New gTLD Program: 2026 Round application evaluation as per the details found in the New gTLD Program: 2026 Round Applicant Guidebook (AGB), Section 7.10. This is based on the Generic Names Supporting Organization (GNSO) Final Report on the New gTLD Subsequent Procedures Policy Development Process (Section 24) and the Phase 1 Final Report on the Internationalized Domain Names Expedited Policy Development Process.

The SSE is conducted by analyzing comparisons between strings along with their variant strings. This is accomplished through a manual review process based on the pre-screening report generated by the SSE tool using the SSE data, and following the SSE guidelines.

SSE Guidelines

SSE guidelines: String Similarity Evaluation Guidelines for the New gTLD Program: 2026 Round version 1.0

The SSE guidelines will provide direction for the SSE panel on how to manually conduct the string similarity evaluation and how to use the pre-screening report generated by the SSE tool during this analysis. The SSE guidelines have been finalized after public comment.

SSE Data

The SSE data identifies the pairs of code points which are similar, as determined by script experts, and defines the level of similarity between them. The SSE data covers the analysis of the full repertoire of the Root Zone Label Generation Rules (RZ-LGR). The SSE data is tabulated in a machine-readable XML version, intended to be used by the SSE tool, and in a human-readable HTML version.

For details, see the overview document String Similarity Evaluation Data. The SSE data have been finalized after public comment. The SSE data files can be collectively downloaded with this package.

Script Similarity Data
Common HTML XML
Arabic HTML XML
Armenian HTML XML
Bangla (Bengali) HTML XML
Chinese HTML XML
Cyrillic HTML XML
Devanagari HTML XML
Ethiopic HTML XML
Georgian HTML XML
Greek HTML XML
Gujarati HTML XML
Gurmukhi HTML XML
Hebrew HTML XML
Japanese (Han + Hiragana + Katakana) HTML XML
Kannada HTML XML
Khmer HTML XML
Korean (Han + Hangul) HTML XML
Lao HTML XML
Latin HTML XML
Malayalam HTML XML
Myanmar HTML XML
Oriya HTML XML
Sinhala HTML XML
Tamil HTML XML
Telugu HTML XML
Thaana HTML XML
Thai HTML XML

SSE Tool

The SSE tool uses the SSE data to determine if any of the input strings are similar to other strings in the scope of comparison. Such cases are identified by the SSE tool as potential contention sets.

The detailed workflow of the SSE tool is included in the SSE guidelines, Appendix A: Workflow of the SSE Tool.

Domain Name System
Internationalized Domain Name ,IDN,"IDNs are domain names that include characters used in the local representation of languages that are not written with the twenty-six letters of the basic Latin alphabet ""a-z"". An IDN can contain Latin letters with diacritical marks, as required by many European languages, or may consist of characters from non-Latin scripts such as Arabic or Chinese. Many languages also use other types of digits than the European ""0-9"". The basic Latin alphabet together with the European-Arabic digits are, for the purpose of domain names, termed ""ASCII characters"" (ASCII = American Standard Code for Information Interchange). These are also included in the broader range of ""Unicode characters"" that provides the basis for IDNs. The ""hostname rule"" requires that all domain names of the type under consideration here are stored in the DNS using only the ASCII characters listed above, with the one further addition of the hyphen ""-"". The Unicode form of an IDN therefore requires special encoding before it is entered into the DNS. The following terminology is used when distinguishing between these forms: A domain name consists of a series of ""labels"" (separated by ""dots""). The ASCII form of an IDN label is termed an ""A-label"". All operations defined in the DNS protocol use A-labels exclusively. The Unicode form, which a user expects to be displayed, is termed a ""U-label"". The difference may be illustrated with the Hindi word for ""test"" — परीका — appearing here as a U-label would (in the Devanagari script). A special form of ""ASCII compatible encoding"" (abbreviated ACE) is applied to this to produce the corresponding A-label: xn--11b5bs1di. A domain name that only includes ASCII letters, digits, and hyphens is termed an ""LDH label"". Although the definitions of A-labels and LDH-labels overlap, a name consisting exclusively of LDH labels, such as""icann.org"" is not an IDN."