1
0
mirror of https://github.com/alexandrev/xslt-lab.git synced 2026-09-18 19:43:16 +00:00
Files
xslt-lab/site/content/xslt/functions/xpath-string-to-codepoints.md
T
alexandrev-tibco 53e90ef86e feat(blog): complete XSLT/XPath reference — 229 function pages
Full coverage of XSLT 1.0, 2.0 and 3.0 elements and XPath functions:
- 59 XSLT elements (xsl:stylesheet → xsl:use-accumulators)
- 170 XPath functions (1.0 node/string/numeric/boolean, 2.0 sequence/
  date/QName/string, 3.0 HOF/map/array/JSON/streaming)
Each page: description, parameters table, return value, 2 runnable
Saxon examples with input XML + stylesheet + output, notes, cross-links.
xsltCompletions.js: all 229 entries now have blogSlug for hover links.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-19 09:28:05 +02:00

3.4 KiB

title, description, date, version, versionLabel, category, syntax, tags
title description date version versionLabel category syntax tags
string-to-codepoints() Returns a sequence of integers representing the Unicode codepoints of each character in a string. 2026-04-18T00:00:00Z 2.0 XSLT 2.0 string function string-to-codepoints(string)
xslt
reference
xslt2
xpath

Description

string-to-codepoints() decomposes a string into its individual Unicode characters and returns their integer codepoints as a sequence of xs:integer values. The sequence length equals the number of Unicode characters (codepoints) in the string, which may differ from the byte length in UTF-8 or UTF-16 encodings.

It is the inverse of codepoints-to-string() and enables character-level manipulation — inspecting, filtering, or transforming individual characters by their numeric values.

If the argument is an empty sequence or an empty string, the function returns an empty sequence.

Parameters

Parameter Type Required Description
string xs:string? Yes The string to decompose into codepoints.

Return value

xs:integer* — a sequence of Unicode codepoint integers, one per character.

Examples

Inspecting character codepoints

Stylesheet:

<?xml version="1.0" encoding="UTF-8"?>
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
  <xsl:output method="xml" indent="yes"/>

  <xsl:template match="/">
    <codepoints>
      <xsl:for-each select="string-to-codepoints('Hello!')">
        <cp value="{.}"/>
      </xsl:for-each>
    </codepoints>
  </xsl:template>
</xsl:stylesheet>

Output:

<codepoints>
  <cp value="72"/>
  <cp value="101"/>
  <cp value="108"/>
  <cp value="108"/>
  <cp value="111"/>
  <cp value="33"/>
</codepoints>

Filtering non-ASCII characters

Input XML:

<?xml version="1.0" encoding="UTF-8"?>
<texts>
  <text>Héllo Wörld</text>
  <text>Plain ASCII only</text>
</texts>

Stylesheet:

<?xml version="1.0" encoding="UTF-8"?>
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
  <xsl:output method="xml" indent="yes"/>

  <xsl:template match="/texts">
    <analysis>
      <xsl:for-each select="text">
        <xsl:variable name="cps" select="string-to-codepoints(.)"/>
        <text ascii-only="{if (every $cp in $cps satisfies $cp le 127) then 'yes' else 'no'}">
          <xsl:value-of select="."/>
        </text>
      </xsl:for-each>
    </analysis>
  </xsl:template>
</xsl:stylesheet>

Output:

<analysis>
  <text ascii-only="no">Héllo Wörld</text>
  <text ascii-only="yes">Plain ASCII only</text>
</analysis>

Notes

  • Each item in the returned sequence is the codepoint of one Unicode character, not one byte. For multi-byte UTF-8 characters (e.g., © is 2 bytes) the function still returns one integer.
  • Surrogate pairs as used in UTF-16 are presented as their actual codepoint (e.g., U+1F600 emoji returns 128512, not two surrogate integers).
  • Combining with codepoints-to-string() allows lossless character-by-character transformations.
  • count(string-to-codepoints($s)) gives the number of Unicode characters, equivalent to string-length($s).

See also