PySET
A high-performance, zero-dependency library for sentence boundary detection. Built for speed, accuracy, and simplicity.
Features
Zero
Dependencies
2-3x
Faster than PySBD
100%
52 Language Accuracy
85
Rules
Quick Example
from pyset import TokenBoundaryDetector
detector = TokenBoundaryDetector()
sentences = detector.split("Hello world. How are you? I'm doing great!")
print(sentences)
# ['Hello world.', 'How are you?', "I'm doing great!"]
Why PySET?
| Feature | PySET | PySBD | Others |
|---|---|---|---|
| Dependencies | Zero | 1 | 1-5 |
| Speed | Fastest | 2-3x slower | Varies |
| Accuracy | 100% | Lower | Varies |
| Languages | 52+ | 20+ | Varies |
Performance


Supported Languages
PySET supports 52+ languages including:
English, German, French, Spanish, Italian, Portuguese, Dutch
Russian, Chinese, Japanese, Korean, Arabic, Hebrew
Hindi, Bengali, Tamil, Telugu, Malayalam
And 35+ more...
Documentation
| Guide | Description |
|---|---|
| Quick Start | Get up and running in 5 minutes |
| Installation | pip install and setup |
| API Reference | Full API documentation |
| Configuration | Customize behavior |
| Rules | How sentence detection works |
| Benchmarks | Performance data |
| About | License and credits |