Metadata-Version: 2.4
Name: g2p-mix
Version: 0.7.1
Summary: Mixed Mandarin/Cantonese and English grapheme-to-phoneme conversion
Author-email: Zhendong Peng <pzd17@tsinghua.org.cn>
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/pengzhendong/g2p-mix
Project-URL: Documentation, https://github.com/pengzhendong/g2p-mix#readme
Project-URL: BugTracker, https://github.com/pengzhendong/g2p-mix/issues
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Operating System :: OS Independent
Classifier: Topic :: Scientific/Engineering
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: click
Requires-Dist: g2p-en
Requires-Dist: jieba
Requires-Dist: nltk
Requires-Dist: pycantonese
Requires-Dist: pyopenhc
Requires-Dist: pypinyin
Requires-Dist: ToJyutping>=3.2
Requires-Dist: wetext<0.2,>=0.1.4
Requires-Dist: wordsegment
Provides-Extra: g2pw
Requires-Dist: torch; extra == "g2pw"
Requires-Dist: modelscope; extra == "g2pw"
Requires-Dist: g2pw>=0.1.1; extra == "g2pw"
Provides-Extra: similarity
Requires-Dist: panphon<0.23,>=0.22.2; extra == "similarity"
Provides-Extra: test
Requires-Dist: build<1.4,>=1.2.1; extra == "test"
Requires-Dist: pip<26,>=23.1; extra == "test"
Requires-Dist: pytest>=8; extra == "test"
Requires-Dist: setuptools<81,>=77; extra == "test"
Requires-Dist: wheel<0.47,>=0.45; extra == "test"
Provides-Extra: dev
Requires-Dist: build<1.4,>=1.2.1; extra == "dev"
Requires-Dist: pip<26,>=23.1; extra == "dev"
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: ruff==0.16.0; extra == "dev"
Requires-Dist: setuptools<81,>=77; extra == "dev"
Requires-Dist: wheel<0.47,>=0.45; extra == "dev"
Dynamic: license-file

# g2p-mix

[![PyPI](https://img.shields.io/pypi/v/g2p-mix)](https://pypi.org/project/g2p-mix/)
[![License](https://img.shields.io/github/license/pengzhendong/g2p-mix)](LICENSE)

Mixed Chinese–English grapheme-to-phoneme conversion with source alignment.

The package intentionally supports two modes:

- Mandarin + English
- Cantonese + English

Python 3.10 or newer is required.

## Installation

```bash
pip install g2p-mix
```

## Quick start

```python
from g2p_mix import G2P

g2p = G2P()
result = g2p("你这个 idea，不太 make sense。")

print(result.phones)
```

```text
('n', 'i3', 'zh', 'e4', 'g', 'e5', 'AY0', 'D', 'IY1', 'AH0', 'b', 'u2', 't', 'ai4', 'M', 'EY1', 'K', 'S', 'EH1', 'N', 'S')
```

Mandarin is the default. Use Cantonese by changing only the mode:

```python
g2p = G2P("cantonese")
print(g2p("你好 idea").phones)
```

```text
('n', 'ei5', 'h', 'ou2', 'AY0', 'D', 'IY1', 'AH0')
```

`result.phones` is always the final, directly usable output. Numeric tones are
attached in native mode; IPA tone letters and English stress marks are attached
in IPA mode.

Arabic numbers and other written forms are normalized automatically by WeText:

```python
result = G2P()("版本1.0发布于2026年")
print(result.normalized_text)
```

```text
版本一点零发布于二零二六年
```

Unicode compatibility letters and numbers are normalized before TN, so
full-width input such as `ＡＢＣ １２３` is supported without changing Chinese
punctuation. Decomposable English diacritics are folded only for pronunciation
lookup, so `café` retains its spelling and source spans while using the
CMUdict entry for `cafe`. Latin text that cannot be folded safely still raises
`G2PError` instead of silently losing phones.

When a sentence contains a covered POS-dependent homograph, the English backend
tags the complete English projection once and shares that context across its
tokens. For example, `record` receives different pronunciations in `I record
music` and `This is a record`. Chinese islands remain visible to the tagger as
one `<ZH>` placeholder and punctuation is retained. Unambiguous sentences skip
POS tagging, while words outside CMUdict continue through segmentation and the
`g2p-en` OOV predictor. Project-reviewed corrections to upstream homograph data
live in an external resource rather than being embedded in backend code.

Unknown Chinese characters are strict by default. Use `preserve` when a
partially pronounced result is preferable to rejecting the whole sentence:

```python
result = G2P(unknown="preserve")("你𲎯好")
print(result.phones)
print(result.warnings)
```

The unknown character remains as a source-aligned unit with
`is_unknown=True`, empty `phones`, and a warning; no pronunciation is
invented. A compatible secondary backend can instead be selected explicitly:

```python
g2p = G2P(
    backend="g2pw",
    fallback_backend="pypinyin",
)
```

## IPA

```python
from g2p_mix import G2P

g2p = G2P(output="ipa", tone_sandhi=False)
result = g2p("中国 idea")

print(result.phones)
```

```text
('ʈ͡ʂ', 'ʊ', 'ŋ˥˥', 'k', 'w', 'o˧˥', 'a', 'ɪ', 'd', 'ˈi', 'ə')
```

Base phones without tone or stress remain available separately:

```python
print(result.base_phones)
```

```text
('ʈ͡ʂ', 'ʊ', 'ŋ', 'k', 'w', 'o', 'a', 'ɪ', 'd', 'i', 'ə')
```

## Backends

Built-in Chinese backends are selected by name:

| Mode | Default | Alternatives |
| --- | --- | --- |
| `mandarin` | `pypinyin` | `g2pw` |
| `cantonese` | `tojyutping` | `pycantonese` |

```python
g2p = G2P("mandarin", backend="g2pw")
```

G2PW is optional:

```bash
pip install "g2p-mix[g2pw]"
```

The G2PW backend keeps the upstream ONNX session by default; pass
`G2PWBackend(onnx_threads=N)` to rebuild it with a different intra-op thread
count (upstream hardcodes 2). The knob is latency-only and never changes
predictions.

Two opt-in switches trade a sliver of accuracy for speed. Both need the
extra graph files produced by `scripts/prepare_g2pw_split.py` (which also
verifies that the split graphs are bit-exact on the upstream per-position
inputs):

- `G2PWBackend(quantize=True)` loads a dynamic-INT8 copy of the model
  (606 MB -> 159 MB) with the same windowed inference semantics;
  clean-CPP target accuracy measured 96.0269% vs 96.0716% fp32.
- `G2PWBackend(use_split_model=True)` runs the BERT encoder once per
  sentence instead of once per polyphonic position; combine with
  `quantize=True` for a quantized encoder. Sharing the encoder sees the
  whole sentence instead of the training-time ±16 character window, so
  predictions drift slightly on longer sentences (clean-CPP 95.7359%).

Both numbers and the default fp32 path's latency are listed in
[`benchmarks/baseline.md`](benchmarks/baseline.md). Downloads are scoped to
what a configuration actually runs: the default path skips the split/INT8
graphs and the 393 MB tokenizer checkpoint, the split path skips the
monolithic graph, and INT8 paths skip their fp32 counterparts (so
`use_split_model=True, quantize=True` pulls about 155 MB).

A custom backend object can be passed through the same argument:

```python
g2p = G2P("mandarin", backend=MyMandarinBackend())
```

## Mandarin phrase lexicon

`g2p_mix/dict/phrases.txt` pins curated readings for polyphonic phrases, one
entry per line (`word syllable ...`, with trailing tone digits and `5` for the
neutral tone). The lexicon is re-matched against the analyzed tokens after each
Mandarin backend — left to right, longest match first — so an entry keeps
working for both `pypinyin` and `g2pw` even when the segmenter splits the
phrase apart. It only rewrites listed phrases and never changes anything else.

## Detailed results

Most applications only need `result.phones`. `result.base_phones` always
removes Mandarin and Cantonese tones as well as English stress. Source-aligned
units are available when more detail is required:

```python
for unit in result.units:
    print(unit.text, unit.phones, unit.tone, unit.source_spans)
```

IPA units also retain their original alphabet and phones:

```python
for unit in result.units:
    print(unit.source_alphabet, unit.source_phones)
```

The input remains losslessly reconstructable:

```python
assert result.reconstruct_original() == "中国 idea"
```

## Phonetic similarity

Install the optional PanPhon backend:

```bash
pip install "g2p-mix[similarity]"
```

Then compare text directly through the same `G2P` object:

```python
g2p = G2P("mandarin", tone_sandhi=False)

near = g2p.compare("西", "she")
far = g2p.compare("西", "key")

assert near.score > far.score
print(near.score, near.alignment)
```

Similarity currently uses base phones. Tone and stress remain available in the
structured result but are intentionally excluded from the score. PanPhon is
loaded lazily and is not required for G2P or IPA output.

## CLI

```bash
g2p_mix "你这个 idea。"
g2p_mix "你这个 idea。" --mode cantonese
g2p_mix "你这个 idea。" --output ipa
g2p_mix "银行 ATM" --backend g2pw --format json
```

## Advanced usage

The root package exposes only the simple API:

```python
from g2p_mix import G2P, G2PError, G2PResult
```

Backend protocols, structured models, transcription, projections, and the
internal pipeline live in their respective submodules. See
[Architecture and extension points](https://github.com/pengzhendong/g2p-mix/blob/master/docs/architecture.md).

## Development

```bash
python -m pip install -U pip
python -m pip install -e ".[dev]"
ruff check .
ruff format --check .
python -m pytest
```

The normal suite uses an injected converter for G2PW tests and does not
download a model. Run the real-model smoke test explicitly:

```bash
python -m pip install -e ".[g2pw,test]"
G2P_MIX_TEST_G2PW=1 python -m pytest -m g2pw
```

Quality evaluation is separate from unit tests and the published wheel:

```bash
python -m benchmarks
python -m benchmarks --json
python -m benchmarks --fail-under 1.0
python -m benchmarks --corpus cpp --max-cases 100 --seed 42
python -m benchmarks --corpus hkcancor --cantonese-backend tojyutping
python -m benchmarks benchmarks/data/mandarin_normalization_sandhi.json
```

See [benchmarks/README.md](benchmarks/README.md) for the dataset schema, backend
comparison options, reproducible CPP/HKCanCor adapters, metrics, and measured
baselines.
