StreamTN

A Low-Latency Streaming Chinese Text Normalization Model for Streaming TTS in Dialogue Systems

Abstract Text-to-Speech (TTS) is an essential module that provides spoken responses in an LLM-centered spoken dialogue system (SDS). To ensure accurate TTS synthesis, responses generated by an LLM must be converted into TTS-readable formats via a Text Normalization (TN) module, imposing strict low-latency requirements in real-time SDS scenarios. Existing TN solutions are largely rule-based, rely on manual engineering, and generalize poorly to unseen patterns. Although an LLM itself can perform TN through prompt engineering, it faces key limitations: high first-token latency due to non-streaming processing, hallucination risks, and degraded intelligence or reasoning when the core LLM module is fine-tuned solely for TN. To address these challenges, we propose StreamTN, a lightweight LLM-based Chinese streaming TN model. Built on Qwen3-0.6B, StreamTN employs a dual-track streaming framework in which input tokens and output tokens are processed on two parallel tracks, enabling low-latency real-time inference without complex prompting. Moreover, task-specific fine-tuning yields superior TN performance and fewer hallucinations than rule-based systems and general-purpose LLMs. We also introduce a TN benchmark that spans diverse text scenarios, providing a comprehensive evaluation standard for speech generation in spoken dialogue systems. Experiments demonstrate the effectiveness of StreamTN in accuracy and inference latency.

This page is for research demonstration purposes only.

System Overview

Figure 1. The architecture of our proposed StreamTN model.

Case Studies

Integers and Decimals

Input Text Reference Normalization StreamTN Output BiLSTM Output WeTextProcessing Output Raw Audio StreamTN Audio
实验室温度计显示-273.15℃ 实验室温度计显示负二百七十三点一五摄氏度 实验室温度计显示负二百七十三点一五摄氏度 实验室温度计显示负两百七十三点一五摄氏度 实验室温度计显示负两百七十三点一五摄氏度
电脑的售价是4899元,比原价便宜200元。 电脑的售价是四千八百九十九元,比原价便宜二百元。 电脑的售价是四千八百九十九元,比原价便宜二百元。 电脑的售价是四千八百九十九元,比原价便宜两百元。 电脑的售价是四千八百九十九元,比原价便宜两百元。

Ordinal Numbers

Input Text Reference Normalization StreamTN Output BiLSTM Output WeTextProcessing Output Raw Audio StreamTN Audio
本次马拉松比赛的第2名选手因伤退赛,冠军由第3名递补获得。 本次马拉松比赛的第二名选手因伤退赛,冠军由第三名递补获得。 本次马拉松比赛的第二名选手因伤退赛,冠军由第三名递补获得。 本次马拉松比赛的第两名选手因伤退赛,冠军由第三名递补获得。 本次马拉松比赛的第两名选手因伤退赛,冠军由第三名递补获得。
在年度排名中,第2名的得分与第4名相差百分,差距显著。 在年度排名中,第二名的得分与第四名相差百分,差距显著。 在年度排名中,第二名的得分与第四名相差百分,差距显著。 在年度排名中,第两名的得分与第四名相差百分,差距显著。 在年度排名中,第两名的得分与第四名相差百分,差距显著。

Phone Numbers

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
配送电话13109826543,请注意接听 配送电话一三一零九八二六五四三,请注意接听 配送电话一三一零九八二六五四三,请注意接听 配送电话幺三幺零九八二六五四三,请注意接听
如需了解活动详情,请拨打14785236901进行咨询。 如需了解活动详情,请拨打一四七八五二三六九零一进行咨询。 如需了解活动详情,请拨打一四七八五二三六九零一进行咨询。 如需了解活动详情,请拨打幺四七八五二三六九零幺进行咨询。

Time

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
日落将在19:23分准时出现 日落将在十九点二十三分准时出现 日落将在十九点二十三分准时出现 日落将在十九比二三分准时出现
演出将在8:45开始,请提前入场并关闭手机以免影响他人。 演出将在八点四十五分开始,请提前入场并关闭手机以免影响他人。 演出将在八点四十五分开始,请提前入场并关闭手机以免影响他人。 演出将在八点四十五分开始,请提前入场并关闭手机以免影响他人。

Military Passwords

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
敌军在405高地部署了重武器,准备发起攻击。 敌军在四洞五高地部署了重武器,准备发起攻击。 敌军在四洞五高地部署了重武器,准备发起攻击。 敌军在四零五高地部署了重武器,准备发起攻击。
任务编号为3017,开始行动 任务编号为三洞幺拐,开始行动 任务编号为三洞幺拐,开始行动 任务编号为三千零一十七,开始行动

Capital Amounts

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
发票金额显示为贰仟叁佰元整,需重新核对。 发票金额显示为二千三百元整,需重新核对。 发票金额显示为二千三百元整,需重新核对。 发票金额显示为贰仟叁佰元整,需重新核对。
客户账户余额显示为叁万捌仟元整,请及时核对 客户账户余额显示为三万八千元整,请及时核对 客户账户余额显示为三万八千元整,请及时核对 客户账户余额显示为叁万捌仟元整,请及时核对

Physical Units

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
手机电池的容量是3000mAh,续航能力不错。 手机电池的容量是三千毫安时,续航能力不错。 手机电池的容量是三千毫安时,续航能力不错。 手机电池的容量是三千米Ah,续航能力不错。
旅行包的容量是30L,足够装下一周的衣物和日常用品,适合短途旅行,尤其是周末出游。 旅行包的容量是三十升,足够装下一周的衣物和日常用品,适合短途旅行,尤其是周末出游。 旅行包的容量是三十升,足够装下一周的衣物和日常用品,适合短途旅行,尤其是周末出游。 旅行包的容量是三十L,足够装下一周的衣物和日常用品,适合短途旅行,尤其是周末出游。

Chemical Formulas and Equations

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
在298K时,AgBr的溶度积Ksp=5.0×10⁻¹³,其溶解平衡为AgBr(s) ⇌ Ag⁺ + Br⁻,该平衡常用于控制AgBr的沉淀与溶解 在二百九十八开尔文时,溴化银的溶度积Ksp等于五点零乘以十的负十三次方,其溶解平衡为溴化银固体可逆生成银离子和溴离子,该平衡常用于控制溴化银的沉淀与溶解 在二百九十八开尔文时,溴化银的溶度积Ksp等于五点零乘以十的负十三次方,其溶解平衡为溴化银固体可逆生成银离子和溴离子,该平衡常用于控制溴化银的沉淀与溶解 在二九八K时,AgBr的溶度积Ksp等于五点零乘十⁻¹³,其溶解平衡为AgBr(s) ⇌ Ag⁺ 加 Br⁻,该平衡常用于控制AgBr的沉淀与溶解
在晶体结构中,NaCl晶胞为面心立方结构,每个晶胞包含4个Na⁺和4个Cl⁻,晶胞参数为a = 564 pm,晶格能为787 kJ/mol,表明其具有较强的离子键。 在晶体结构中,氯化钠晶胞为面心立方结构,每个晶胞包含四个钠离子和四个氯离子,晶胞参数为a等于五百六十四皮米,晶格能为七百八十七千焦每摩尔,表明其具有较强的离子键。 在晶体结构中,氯化钠晶胞为面心立方结构,每个晶胞包含四个钠离子和四个氯离子,晶胞参数为a等于五百六十四皮米,晶格能为七百八十七千焦每摩尔,表明其具有较强的离子键。 在晶体结构中,NaCl晶胞为面心立方结构,每个晶胞包含四个Na⁺和四个Cl⁻,晶胞参数为a 等于 五六四 pm,晶格能为七八七 kJ/mol,表明其具有较强的离子键。

Daily Chemicals

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
烘焙饼干时,加入的NaHCO₃受热分解,生成CO₂气体使饼干松软。 烘焙饼干时,加入的碳酸氢钠受热分解,生成二氧化碳气体使饼干松软。 烘焙饼干时,加入的碳酸氢钠受热分解,生成二氧化碳气体使饼干松软。 烘焙饼干时,加入的NaHCO₃受热分解,生成CO₂气体使饼干松软。
洗衣粉中的Na₂CO₃能中和水中的Ca²⁺和Mg²⁺,提高洗涤效果。 洗衣粉中的碳酸钠能中和水中的钙离子和镁离子,提高洗涤效果。 洗衣粉中的碳酸钠能中和水中的钙离子和镁离子,提高洗涤效果。 洗衣粉中的Na₂CO₃能中和水中的Ca²⁺和Mg²⁺,提高洗涤效果。

Simple Mathematical Equations

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
若x=3,则x⁴=81,x²=9,x³=27 若x等于三,则x的四次方等于八十一,x的平方等于九,x的三次方等于二十七 若x等于三,则x的四次方等于八十一,x的平方等于九,x的三次方等于二十七 若x等于三,则x⁴等于八十一,x²等于九,x³等于二十七
计算体积V=½×10×5×4=100cm³ 计算体积V等于二分之一乘十乘五乘四等于一百立方厘米 计算体积V等于二分之一乘十乘五乘四等于一百立方厘米 计算体积V等于½乘十乘五乘四等于十零立方厘米

Fractions and Percentages

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
调查显示约2/3的受访者支持环保政策 调查显示约三分之二的受访者支持环保政策 调查显示约三分之二的受访者支持环保政策 调查显示三分之约二的受访者支持环保政策
报告显示,75%的用户对产品表示满意。 报告显示,百分之七十五的用户对产品表示满意。 报告显示,百分之七十五的用户对产品表示满意。 报告显示,百分之七十五的用户对产品表示满意。

Years and Dates

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
2025年国际人工智能大会将在上海召开 二零二五年国际人工智能大会将在上海召开 二零二五年国际人工智能大会将在上海召开 二零二五年国际人工智能大会将在上海召开
2021年夏季奥运会申办城市名单确定 二零二一年夏季奥运会申办城市名单确定 二零二一年夏季奥运会申办城市名单确定 二零二一年夏季奥运会申办城市名单确定

Amount Readings

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
她买了一件衣服,价格是£150.75,还购买了配饰。 她买了一件衣服,价格是一百五十点七五英镑,还购买了配饰。 她买了一件衣服,价格是一百五十点七五英镑,还购买了配饰。 她买了一件衣服,价格是一百五十点七五英镑,还购买了配饰。
这本书的价格是€23.5,比原价便宜了,因为促销活动。 这本书的价格是二十三点五欧元,比原价便宜了,因为促销活动。 这本书的价格是二十三点五欧元,比原价便宜了,因为促销活动。 这本书的价格是二十三点五欧元,比原价便宜了,因为促销活动。

Complex Mathematical Equations

Input Text Reference Normalization StreamTN Output WeTextProcessing Output Raw Audio StreamTN Audio
设 f(x) = ∫₀ˣ e⁻ᵗ² dt / log(√x + 1),求x→0时的极限 设函数 f(x) 是分子从零到x的定积分e的负t平方次方dt,分母是自然对数x开平方加一,求x趋近于零时的极限 设函数 f(x) 是分子从零到x的定积分e的负t平方次方dt,分母是自然对数x开平方加一,求x趋近于零时的极限 设函数 f(x) 是分子从零到x的定积分e的负t平方次方dt,分母是自然对数x开平方加一,求x趋近于零时的极限
2×2矩阵行列式det([[a,b],[c,d]])=ad−bc 2乘2矩阵的行列式等于a乘d减b乘c 2乘2矩阵的行列式等于a乘d减b乘c 2乘2矩阵的行列式等于a乘d减b乘c

Delay Frame Analysis

Latency measurement setup. The upstream LLM (Qwen3-32B) is served with vLLM on 2×A6000 GPUs (tensor_parallel_size=2, greedy decoding, gpu_memory_utilization=0.9, dtype=auto), producing roughly 20 tokens/s. StreamTN and all local TN baselines run on a single NVIDIA A6000 GPU. Encoding all test inputs with the Qwen3 tokenizer gives a mean length of 27.58 tokens. FPD (ms) denotes First Packet Delay, decomposed as upstream waiting (the time to receive the first d tokens) plus StreamTN computation. The configuration marked with * is adopted in our main experiments.

Strategy Delay Frame Micro-F1 Micro-P FPD (ms)
Non-Streaming - 0.9639 ± 0.0011 0.9620 ± 0.0012 -
Streaming 1 frame 0.7030 ± 0.0010 0.7135 ± 0.0018 75
2 frames 0.7952 ± 0.0015 0.8062 ± 0.0026 121
4 frames * 0.8937 ± 0.0009 0.8931 ± 0.0022 213
8 frames 0.9072 ± 0.0011 0.9093 ± 0.0020 394
16 frames 0.9239 ± 0.0008 0.9303 ± 0.0015 756

Appendix: Supplementary Experiments

Detailed latency decomposition, full evaluation metrics, and ablations referenced in the rebuttal.

A. Latency Comparison

FPD is the sum of the upstream waiting time and the TN computation time. Streaming StreamTN starts after receiving the first d tokens from Qwen3-32B (~20 tokens/s on 2×A6000). Non-streaming methods must wait for the complete input (a mean of 27.58 tokens → roughly 1.4 s at 20 tokens/s) before producing any output, which is the architectural latency advantage of the dual-track design.

Method Micro-F1 Upstream Wait (ms) TN Compute (ms) FPD (ms)
StreamTN (streaming, d=1) 0.7030 70 5 75
StreamTN (streaming, d=2) 0.7952 115 6 121
StreamTN (streaming, d=4) * 0.8937 206 7 213
StreamTN (streaming, d=8) 0.9072 385 9 394
StreamTN (streaming, d=16) 0.9239 744 12 756
WeTextProcessing 0.7853 ≈1,400 10.5 ≈1,400
Qwen3-0.6B (prompted) 0.4872 ≈1,400 50.3 ≈1,400
FlatTN 0.7569 ≈1,400 24.9 ≈1,400
BiLSTM 0.8941 ≈1,400 2.8 ≈1,400
StreamTN (non-streaming) 0.9639 ≈1,400 37.0 ≈1,400

* configuration adopted in our main experiments. For streaming rows, “Upstream Wait” is the time to receive the first d tokens and “TN Compute” is StreamTN’s incremental first-token compute. For non-streaming methods, “Upstream Wait” is the time to receive the complete input, a simulated estimate (≈1.4 s) obtained from the mean input length (27.58 tokens) divided by the upstream token rate (~20 tokens/s) rather than a directly measured value. FPD is therefore dominated by this wait, and the method’s own compute (mean, shown in the “TN Compute” column, ≤50 ms) adds only a negligible amount. TN compute for BiLSTM and non-streaming StreamTN is the time to the first output token; for the other non-streaming methods it is the full-output processing time.

B. Full Evaluation Metrics

Character-level Micro-F1 is the primary metric because it accommodates pronunciation-equivalent variants (e.g., “2900¥” → “两千九百元” or “二千九百元” are both correct for TTS). Sentence-level exact match and CER are reported as complementary, stricter metrics.

Method Micro-P Micro-R Micro-F1 CER Sentence Acc.
WeTextProcessing 0.7587 0.8139 0.7853 0.2780 0.4350
Qwen3-0.6B (prompted) 0.5282 0.4521 0.4872 0.4774 0.2655
BiLSTM 0.8677 0.9222 0.8941 0.1698 0.6038
FlatTN 0.7390 0.7758 0.7569 0.3115 0.3328
StreamTN (non-streaming) 0.9620 0.9658 0.9639 0.0490 0.6838
StreamTN (streaming, d=4) 0.8931 0.8943 0.8937 0.1399 0.5523

C. Prompt Used for the General-Purpose LLM Baseline

The exact prompt used to evaluate the non-fine-tuned Qwen3-0.6B baseline:

messages = [
    {"role": "system", "content": "你是一个专业的文本标准化助手,请为我做一个文本标准化任务,将原始文本中的数字时间符号等全部标准化为汉字,用于后续TTS系统合成。"},
    {"role": "user", "content": raw_text.strip()}
]

More detailed rule-based prompts were also tested but substantially increased per-item latency while still causing frequent insertion, deletion, and out-of-input hallucination errors.

D. Zero-Vector Ablation: Shared Zero vs. Learned Delay Embedding

In the dual-track design, the input and output embeddings are added element-wise. During the delay phase the output track contributes a zero vector, so the summed embedding carries only the raw input; once the input ends, the input track contributes a zero vector, so the summed embedding carries only the output history. The phase is therefore recoverable from whichever track is nonzero, which is why sharing a single zero vector for delay and pad does not confuse the model. Replacing it with a learned delay embedding yields no consistent gain.

Delay Frame Micro-F1 (shared zero) Micro-F1 (learned delay)
1 frame 0.7030 ± 0.0010 0.7024 ± 0.0013
2 frames 0.7952 ± 0.0015 0.7950 ± 0.0008
4 frames * 0.8937 ± 0.0009 0.8954 ± 0.0011
8 frames 0.9072 ± 0.0011 0.9082 ± 0.0010
16 frames 0.9239 ± 0.0008 0.9210 ± 0.0009

E. LoRA Rank / Alpha Ablation

With learning rate 1e-5 and the q/k/v/o projections fixed, a rank sweep (α=4r) raises Micro-F1 from 0.5228 (r=2) to 0.7292 (r=32), still below full fine-tuning (0.8937 at d=4). The gap is therefore configuration-dependent rather than an inherent LoRA limitation.

Rank (α=4r) Micro-P Micro-R Micro-F1 CER Sentence Acc.
r=2 (α=8) 0.4959 0.5527 0.5228 0.6480 0.0135
r=4 (α=16) 0.5649 0.6043 0.5839 0.5483 0.0349
r=8 (α=32) 0.6251 0.6423 0.6336 0.4727 0.0658
r=16 (α=64) 0.6761 0.6859 0.6810 0.4057 0.1173
r=32 (α=128) 0.7319 0.7264 0.7292 0.3441 0.1783

F. Error Analysis

On the 1,262-sample test set, 565 samples (44.8%) contain at least one character-level error. Character-level Levenshtein alignment (tie priority S>D>I) yields 33,684 matches, 2,736 substitutions (51.9%), 1,264 deletions (24.0%), and 1,271 insertions (24.1%), for 5,271 total edits.

Category Samples M S D I Errors (%)
Integers and Decimals10018277013221.99
Ordinal Numbers100231813836594.42
Phone Numbers100259316768465.33
Years and Dates9726205225602.60
Time Expressions10025032637812.73
Fractions and Percentages9920957051312.88
Military Code1001268991472.28
Uppercase Amounts701326264231.01
Amount Readings7814178731222.66
Physical Units681751278130.91
Chemical Formulas and Equations100557199542856137.64
Simple Chemicals1002823244139688.56
Mathematical Expressions10026959876313.89
Advanced Mathematical Expressions50287763733424723.11

Complex chemistry and advanced mathematics account for 11.9% of samples but 60.7% of total edits, indicating that whole-expression rewrites and omissions (hallucination / context loss) are the dominant failure mode for these categories, rather than tokenizer fragmentation alone.

High-Frequency Edits

Substitutions (ref → hyp)
CountEdit
41两 → 二
32一 → 二
25加 → 和
23勾 → 九
15一 → 1
13二 → 2
11三 → 3
11千 → 百
10十 → 0
10加 → 与
10甲 → 乙
9四 → 4
9幺 → 一
9百 → 千
8。 → ,
8洞 → 零
7五 → 5
7三 → 二
7拐 → 七
6六 → 6
6千 → 万
6四 → 三
6十 → 百
5九 → 9
5二 → 两
5三 → 四
5生 → 等
5成 → 于
5价 → 亚
4万 → 百
Deletions
CountCharacter
38的
36二
35,
28分
25一
24十
23百
16在
16(space)
15于
15数
14加
13以
13三
13酸
12个
12。
12应
11时
11之
11为
11生
11乘
11x
10五
10号
10锰
10离
10子
9九
Insertions
CountCharacter
47零
43分
42,
40一
29反
28的
27应
25为
24十
22子
21个
21该
18二
17百
16离
15千
15锰
15化
141
14氧
140
11四
11中
11乙
11于
10和
10价
9是
9六
9x