Metadata-Version: 2.4
Name: tokmor
Version: 1.2.10.post20260204
Summary: Dependency-free, fast deterministic tokenizer + morphology splitter for 375 languages (~4.6MB)
Author-email: tokmor <tokmor@tokmor.com>
License: MIT License
        
        Copyright (c) 2026 TokMor Contributors
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
        
Project-URL: Homepage, https://github.com/tokmorlab/tokmor
Project-URL: Documentation, https://github.com/tokmorlab/tokmor#readme
Project-URL: Source, https://github.com/tokmorlab/tokmor
Project-URL: Issues, https://github.com/tokmorlab/tokmor/issues
Keywords: nlp,preprocessing,multilingual,tokenizer,tokenization,segmentation,morphology,lemmatization,ner,rag,information-extraction,offline
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Text Processing :: Linguistic
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Dynamic: license-file

# TokMor (Core)

**375 languages • dependency-free • ~4.6MB • fast • deterministic • MIT**

[![PyPI](https://img.shields.io/pypi/v/tokmor.svg)](https://pypi.org/project/tokmor/)
[![Python](https://img.shields.io/pypi/pyversions/tokmor.svg)](https://pypi.org/project/tokmor/)
[![License: MIT](https://img.shields.io/pypi/l/tokmor.svg)](https://pypi.org/project/tokmor/)

TokMor is a **dependency-free** multilingual preprocessing library for **375 languages**.
It provides **fast, deterministic tokenization + offsets**, **best-effort rule-based morphology**, and **NER-friendly preprocessing hints**.

TokMor is **not** a BERT/Transformer model.
TokMor is **not** a linguistic POS tagger and does **not** run ML models at inference time.

- ✅ **Zero runtime dependencies** (easy installs)
- ✅ **~4.6MB wheel** (`tokmor-1.2.10-py3-none-any.whl`)
- ✅ **Fast** (~35k tokens/sec in the benchmark below)
- ✅ **375 languages** (broad coverage)
- ✅ **Deterministic** (reproducible outputs)
- ✅ **MIT License** (commercial use OK)

Demo (text):

**CODE**

```python
import tokmor

out = tokmor.unified_tokenize("Apple announced products in Seoul!!!", lang="en", sns=False)
print(out["tokens"][:8])
```

**OUTPUT (preview: first 8 tokens)**

```json
[
  {"text": "Apple", "start": 0, "end": 5},
  {"text": "announced", "start": 6, "end": 15},
  {"text": "products", "start": 16, "end": 24},
  {"text": "in", "start": 25, "end": 27},
  {"text": "Seoul", "start": 28, "end": 33},
  {"text": "!", "start": 33, "end": 34},
  {"text": "!", "start": 34, "end": 35},
  {"text": "!", "start": 35, "end": 36}
]
```

## Multilingual demos

- Repo doc: [`docs/DEMO_LANGS.md`](https://github.com/tokmorlab/tokmor/blob/master/docs/DEMO_LANGS.md) (**multilingual** sentence demo + **100-language** one-page tokenize demo)

Examples (one-line):

- **ja** — IN: `私たちは2025-01-10にソウルを訪れました。`  
  OUT: `私たち | は | 2025-01-10 | に | ソウル | を | 訪れました | 。`
- **zh** — IN: `我们在2025-01-10访问了首尔。`  
  OUT: `我们 | 在 | 2025-01-10 | 访问 | 了 | 首尔 | 。`
- **ko** — IN: `우리는 2025-01-10에 서울을 방문했다.`  
  OUT: `우리 | 는 | 2025-01-10 | 에 | 서울 | 을 | 방문했다 | .`
- **th** — IN: `เราไปโซลเมื่อ 2025-01-10`  
  OUT: `เรา | ไป | โซล | เมื่อ | 2025-01-10`
- **lo** — IN: `ພວກເຮົາໄປຢ້ຽມຢາມໂຊລ 2025-01-10`  
  OUT: `ພວກເຮົາ | ໄປ | ຢ້ຽມຢາມ | ໂຊລ | 2025-01-10`
- **my** — IN: `ကျွန်ုပ်တို့သည် 2025-01-10 တွင် ဆိုးလ်ကို သွားခဲ့သည်။`  
  OUT: `ကျွန်ုပ်တို့ | သည် | 2025-01-10 | တွင် | ဆိုး | လ် | ကို | သွားခဲ့ | သည် | ။`
- **km** — IN: `យើងបានទៅទស្សនា សេអ៊ូល នៅ 2025-01-10។`  
  OUT: `យើងបាន | ទៅ | ទស្សនា | សេអ៊ូល | នៅ | 2025-01-10 | ។`
- **bo** — IN: `ང་ཚོས 2025-01-10 ཉིན་སོལ་ལ་བསྐྱོད་པ་ཡིན།`  
  OUT: `ང | ་ | ཚོས | 2025-01-10 | ཉིན | ་ | སོལ | ་ | ལ | ་ | བསྐྱོད | ་ | པ | ་ | ཡིན | །`
- **ar** — IN: `زرنا سيول في 2025-01-10.`  
  OUT: `زرنا | سيول | في | 2025-01-10 | .`
- **he** — IN: `ביקרנו בסיאול ב-2025-01-10.`  
  OUT: `ביקרנו | בסיאול | ב | - | 2025-01-10 | .`
- **fa** — IN: `ما در 2025-01-10 از سئول بازدید کردیم.`  
  OUT: `ما | در | 2025-01-10 | از | سئول | بازدید | کردیم | .`
- **ur** — IN: `ہم نے 2025-01-10 کو سیول کا دورہ کیا۔`  
  OUT: `ہم | نے | 2025-01-10 | کو | سیول | کا | دورہ | کیا۔`
- **dv** — IN: `އަހަރެން 2025-01-10 ދުވަސް ސިއޫލަށް ދިޔައީ.`  
  OUT: `އަހަރެން | 2025-01-10 | ދުވަސް | ސިއޫލަށް | ދިޔައީ | .`
- **hi** — IN: `हमने 2025-01-10 को सियोल का दौरा किया।`  
  OUT: `हमने | 2025-01-10 | को | सियोल | का | दौरा | किया।`
- **bn** — IN: `আমরা 2025-01-10 তারিখে সিউল ভ্রমণ করেছি।`  
  OUT: `আমরা | 2025-01-10 | তারিখে | সিউল | ভ্রমণ | করেছি | ।`
- **ta** — IN: `நாங்கள் 2025-01-10 அன்று சியோலைப் பார்த்தோம்.`  
  OUT: `நாங்கள் | 2025-01-10 | அன்று | சியோலைப் | பார்த்தோம் | .`
- **te** — IN: `మేము 2025-01-10 న సియోల్‌ను సందర్శించాము.`  
  OUT: `మేము | 2025-01-10 | న | సియోల్‌ను | సందర్శించాము | .`
- **mr** — IN: `आम्ही 2025-01-10 रोजी सिओলला भेट दिली.`  
  OUT: `आम्ही | 2025-01-10 | रोजी | सिओलला | भेट | दिली | .`
- **gu** — IN: `અમે 2025-01-10 ના રોજ સિયોલની મુલાકાત લીધી.`  
  OUT: `અમે | 2025-01-10 | ના | રોજ | સિયોલની | મુલાકાત | લીધી | .`
- **pa** — IN: `ਅਸੀਂ 2025-01-10 ਨੂੰ ਸਿਓਲ ਗਏ।`  
  OUT: `ਅਸੀਂ | 2025-01-10 | ਨੂੰ | ਸਿਓਲ | ਗਏ | ।`
- **ne** — IN: `हामीले 2025-01-10 मा सियोल भ्रमण गर्‍यौं।`  
  OUT: `हामीले | 2025-01-10 | मा | सियोल | भ्रमण | गर्‍यौं।`
- **si** — IN: `අපි 2025-01-10 දින සෝල් වෙත ගියෙමු.`  
  OUT: `අපි | 2025-01-10 | දින | සෝල් | වෙත | ගියෙමු | .`
- **el** — IN: `Επισκεφθήκαμε τη Σεούλ στις 2025-01-10.`  
  OUT: `Επισκεφθήκαμε | τη | Σεούλ | στις | 2025-01-10 | .`
- **ru** — IN: `Мы посетили Сеул 2025-01-10.`  
  OUT: `Мы | посетили | Сеул | 2025-01-10 | .`
- **uk** — IN: `Ми відвідали Сеул 2025-01-10.`  
  OUT: `Ми | відвідали | Сеул | 2025-01-10 | .`
- **bg** — IN: `Посетихме Сеул на 2025-01-10.`  
  OUT: `Посетихме | Сеул | на | 2025-01-10 | .`
- **ka** — IN: `ჩვენ 2025-01-10-ს სეულს ვესტუმრეთ.`  
  OUT: `ჩვენ | 2025-01-10 | - | ს | სეულს | ვესტუმრეთ | .`
- **hy** — IN: `Մենք այցելեցինք Սեուլ 2025-01-10-ին։`  
  OUT: `Մենք | այցելեցինք | Սեուլ | 2025-01-10 | - | ին | ։`
- **am** — IN: `እኛ በ2025-01-10 ሴኡልን ጎበኘን።`  
  OUT: `እኛ | በ | 2025-01-10 | ሴኡልን | ጎበኘን | ።`
- **ti** — IN: `ንሕና ብ2025-01-10 ሴኡል ጎበኘና።`  
  OUT: `ንሕና | ብ | 2025-01-10 | ሴኡል | ጎበኘና | ።`

## What it is

- Deterministic tokenization/segmentation with offsets
- Best-effort, rule-based morphology (language-specific analyzers + fallbacks)

## What it is NOT

- A POS tagger
- An NER system
- A machine-learning / model-based tokenizer
- An LLM token-usage monitoring / observability tool

## Install

```bash
pip install tokmor
```

## Quick start (Python)

```python
import tokmor

out = tokmor.unified_tokenize("We visited Seoul on 2025-01-10.", lang="en", sns=False)
print(out["tokens"][:5])
```

## Speed (benchmark)

Measured on this repo’s CI server:
- CPU: Intel(R) Xeon(R) w9-3595X
- Python: 3.13.3
- Workload: 12,000 short mixed-language lines, calling `tokmor.unified_tokenize(..., lang="auto", sns=True)`
- Result: **~35k tokens/sec** (best-of-5)

Reproduce:

```bash
python -m pip install -U tokmor
python - <<'PY'
import time, sys
import tokmor

samples = [
  "We visited Seoul on 2025-01-10, LOL!!!",
  "米拉·万·托斯抵达维伦德尔港。",
  "こんにちは世界。アップルが新製品を発表した。",
  "مرحبا بالعالم! زرنا سيول.",
  "हैलो वर्ल्ड! हम सियोल गए।",
]
texts = samples * 2000

for i in range(200):
  tokmor.unified_tokenize(texts[i], lang="auto", sns=True)

t0 = time.perf_counter()
tokc = 0
for s in texts:
  out = tokmor.unified_tokenize(s, lang="auto", sns=True)
  tokc += len(out.get("tokens", []))
dt = time.perf_counter() - t0
print("python:", sys.version.split()[0])
print("texts:", len(texts))
print("tokens:", tokc)
print("sec:", round(dt, 4))
print("tokens_per_sec:", int(tokc/dt))
PY
```

## NER preprocessing helper

```python
import tokmor

prep = tokmor.ner_preprocess("LOL!!! Apple announced new products in Seoul...", lang="en")
print(prep["tokens"][:8])
```

## Docs / repo

See the repo docs:
- `docs/DATA_SOURCES_AND_LICENSES.md`
- `docs/DOMAIN_LEXICONS.md`
- `docs/FAQ.md`

## License

MIT (see `LICENSE`)
