文本、Bidi 与 IMEText, Bidi & IME
从字节、字素簇、双向段落到输入法合成:zenit 把文本当作系统能力,而不是一串按码点排开的 glyph。From bytes to grapheme clusters, bidirectional paragraphs and IME composition: zenit treats text as a system capability, not a row of glyphs laid out per code point.
一条文本,四层职责One line, four layers
存储、编辑边界、方向解析和平台塑形各有唯一的 authority,最后统一成可命中、可选择的视觉行。任何一层都不替另一层做决定。Storage, edit boundaries, direction resolution and platform shaping each have exactly one authority, and together produce visual lines you can hit-test and select. No layer makes decisions on another’s behalf.
src/text_coresrc/text_core/grapheme.zigsrc/i18n/bidi.zigsrc/render/text_renderer.zig字素不是码点Graphemes, not code points
一个家庭 emoji 是 7 个 Unicode scalar(4 个人加 3 个 ZWJ),é 可以是 e 加一个组合重音,一面旗帜是两个区域指示符。用户眼里它们各是一个字符。zenit 用扩展字素簇作为光标移动、删除、选区和容量截断的原子。A family emoji is seven Unicode scalars (four people and three ZWJs), é can be e plus a combining acute, and a flag is two regional indicators. To the user each is one character. zenit uses extended grapheme clusters as the atom for caret movement, deletion, selection and capacity truncation.
双向文本Bidirectional text
文本在内存里按逻辑顺序存放,屏幕上按视觉顺序显示。UAX #9 先为每个字符解析嵌入层级,再按规则 L2 从最高层级开始逐级反转。下例是 LTR 段落中的 abc אבג 123:希伯来文得到层级 1,紧随其后的数字得到层级 2。Text is stored in logical order and shown in visual order. UAX #9 first resolves an embedding level for every character, then rule L2 reverses runs from the highest level down. Below, abc אבג 123 sits in an LTR paragraph: the Hebrew resolves to level 1 and the digits that follow it to level 2.
UAX #9 解析器是 zenit 自研的可移植实现,位于 src/i18n/bidi.zig,属性表由固定的 Unicode 17.0.0 数据生成。macOS 上最终的 glyph 塑形与 RTL 连字仍整段交给 CoreText,文本渲染器不会把 RTL 片段逐码点拆开。The UAX #9 resolver is zenit’s own portable implementation in src/i18n/bidi.zig, with property tables generated from pinned Unicode 17.0.0 data. On macOS, final glyph shaping and RTL joining are still handed to CoreText a whole run at a time — the text renderer never splits an RTL run per code point.
Unicode 17 一致性门禁Unicode 17 conformance
运行时属性表由仓库固定的 Unicode 17.0.0 数据生成;官方测试文件原样放在 vendor/unicode/17.0.0,并在 SHA256SUMS 中用 SHA-256 固定。这里的「支持」不是挑几个 emoji 写单元测试,而是跑完整的官方用例。Runtime property tables are generated from Unicode 17.0.0 data pinned in the repo; the official test files are vendored unmodified under vendor/unicode/17.0.0 and pinned by SHA-256 in SHA256SUMS. “Supported” here doesn’t mean a handful of emoji unit tests — it means running the complete official suites.
python3 tools/generate_grapheme_data.py --check
python3 tools/generate_bidi_data.py --check
(cd vendor/unicode/17.0.0 && shasum -a 256 -c SHA256SUMS)
zig build test-text-core
zig build test-bidi-conformance--check 确认生成的 Zig 表与数据一致;test-text-core 跑完 GraphemeBreakTest 全部用例;test-bidi-conformance 跑两个完整的官方 bidi 文件(到 UAX #9 规则 L2),它也包含在 zig build test-headless 中。--check confirms the generated Zig tables match the data; test-text-core runs every GraphemeBreakTest case; test-bidi-conformance runs both complete official bidi files (through UAX #9 rule L2) and is also part of zig build test-headless.
IME 是一等公民IME as a first-class citizen
IME 不是把最终汉字伪装成一次键盘输入。平台事件保留阶段语义:ime_preedit 携带合成文本和合成区内的光标偏移,ime_commit 携带最终文本。合成期间 Input 与 Textarea 渲染带下划线的 marked text,但 canonical buffer 只在提交时改变。没有单独的「取消」事件:空的 preedit 就表示取消合成。IME isn’t the final characters disguised as a keystroke. Platform events keep the phases: ime_preedit carries the composing text plus a caret offset inside it, and ime_commit carries the final text. While composing, Input and Textarea render underlined marked text, but the canonical buffer changes only on commit. There is no separate “discard” event — an empty preedit cancels the composition.
ime_preedit / ime_commit,不丢阶段信息;空 preedit 取消合成ime_preedit / ime_commit keep the phase; an empty preedit cancelsime_phase、ime_preedit_len、buffer、cursor_pos、anchorThe harness injects preedit / commit separately and reads back ime_phase, ime_preedit_len, buffer, cursor_pos, anchorscripts/verify_ime.sh 用系统拼音 / 日文输入源走完整的 NSTextInputClient 路径scripts/verify_ime.sh drives the full NSTextInputClient path with the system Pinyin / Japanese input sources测试输入法Testing IME
E2E 客户端把两个阶段作为独立的 RPC 暴露出来,inputState 按 test id 回读输入框的真实状态。The E2E client exposes both phases as separate RPCs, and inputState reads an input’s real state back by test id.
import { imeCommit, imePreedit, inputState } from "./client";
// The input must already have focus (e.g. click it first).
await imePreedit("nihongo");
let s = await inputState("story.input.name");
// s.ime_phase === "composing", s.ime_preedit_len === 7, s.buffer === ""
await imeCommit("日本語");
s = await inputState("story.input.name");
// s.buffer === "日本語", s.ime_preedit_len === 0
await imePreedit(""); // empty preedit cancels an active compositionHarness 的完整用法见 E2E 自动化。See E2E harness for the full harness workflow.


