PySET Logo

PySET

A high-performance, zero-dependency library for sentence boundary detection. Built for speed, accuracy, and simplicity.

Features

Zero
Dependencies
2-3x
Faster than PySBD
100%
52 Language Accuracy
85
Rules

Quick Example

from pyset import TokenBoundaryDetector

detector = TokenBoundaryDetector()
sentences = detector.split("Hello world. How are you? I'm doing great!")

print(sentences)
# ['Hello world.', 'How are you?', "I'm doing great!"]

Why PySET?

Feature PySET PySBD Others
Dependencies Zero 1 1-5
Speed Fastest 2-3x slower Varies
Accuracy 100% Lower Varies
Languages 52+ 20+ Varies

Performance

Execution Time

Words per Second

Supported Languages

PySET supports 52+ languages including:

English, German, French, Spanish, Italian, Portuguese, Dutch
Russian, Chinese, Japanese, Korean, Arabic, Hebrew
Hindi, Bengali, Tamil, Telugu, Malayalam
And 35+ more...

Documentation

Guide Description
Quick Start Get up and running in 5 minutes
Installation pip install and setup
API Reference Full API documentation
Configuration Customize behavior
Rules How sentence detection works
Benchmarks Performance data
About License and credits