Metadata-Version: 2.4
Name: textmeasures
Version: 0.1.0
Summary: Quantitative text measurement tools for frequency distributions and corpus statistics.
Author-email: Tsy Yih <yihtsy@outlook.com>
Project-URL: Homepage, https://github.com/Yihtsy/textmeasures
Project-URL: Repository, https://github.com/Yihtsy/textmeasures
Project-URL: Issues, https://github.com/Yihtsy/textmeasures/issues
Keywords: text analysis,quantitative linguistics,corpus linguistics,frequency distribution,Zipf,CoNLL-U
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: numpy>=1.23
Requires-Dist: scipy>=1.9

# textmeasures

`textmeasures` is a Python package for quantitative text measurement. The current public preview focuses on frequency distributions, CoNLL-U based symbol extraction, rank-frequency curves, vocabulary richness, concentration/evenness metrics, and Zipf-family model fitting.

This is an early public release (`0.1.0`). The API is usable, but it may still evolve before a stable `1.0.0` release.

## Installation

```bash
pip install textmeasures
```

For local development from source:

```bash
git clone https://github.com/Yihtsy/textmeasures.git
cd textmeasures
pip install -e .
```

## Quick Start

```python
from textmeasures import FreqDist, entropy, repeat_rate, gini, normalized_entropy

freqs = FreqDist([10, 5, 3, 1, 1])

print(freqs.to_list())
print(entropy(freqs))
print(repeat_rate(freqs))
print(gini(freqs))
print(normalized_entropy(freqs))
```

## Distribution Objects

```python
from textmeasures import FreqDist

fd = FreqDist([10, 5, 3, 1, 1])

relative = fd.to_rel_freqdist()
cumulative = fd.to_cum_freqdist()
cumulative_relative = fd.to_cum_rel_freqdist()
spectrum = fd.to_freq_spectrum()

print(relative.to_list())
print(cumulative.to_list())
print(cumulative_relative.to_list())
print(spectrum.to_list())
```

## CoNLL-U Input

`textmeasures` can read CoNLL-U files and convert selected linguistic units into symbols or frequency distributions.

```python
from textmeasures import conllu_to_freqdist, conllu_to_symbols

symbols = conllu_to_symbols(
    "sample.conllu",
    linguistic_unit="lemma",
    language="en",
    exclude_upos=("PUNCT", "SYM", "X"),
)

freqs = conllu_to_freqdist(
    "sample.conllu",
    linguistic_unit="word",
    language="en",
)

print(symbols[:10])
print(freqs.to_list())
```

Supported languages are `en` and `zh`. Supported linguistic units include `word`, `lemma`, `letter`, `character`, `upos`, `deprel`, and `n_gram`; availability depends on the selected language.

## Zipf-Family Fitting

```python
from textmeasures import (
    zipf_fitted_parameters,
    zipf_mandelbrot_fitted_parameters,
    zipf_alekseev_fitted_parameters,
)

freqs = [100, 53, 31, 19, 12, 8, 5, 3, 2, 1]

print(zipf_fitted_parameters(freqs))
print(zipf_mandelbrot_fitted_parameters(freqs))
print(zipf_alekseev_fitted_parameters(freqs))
```

Fitting methods support `"log_ols"`, `"linear_ols"`, and `"chi_square"`.

## API Overview

| Category | Main APIs |
| --- | --- |
| Distribution objects | `FreqDist`, `RelFreqDist`, `CumFreqDist`, `CumRelFreqDist`, `FreqSpectrum` |
| CoNLL-U utilities | `conllu_to_symbols`, `conllu_to_freqdist`, `conllu_token_dict`, `conllu_token_is_valid` |
| Entropy and repetition | `entropy`, `repeat_rate`, `inverse_repeat_rate`, `normalized_entropy`, `simpson`, `normalized_simpson` |
| Curves | `pareto_curve`, `lorenz_curve`, `yih_curve` |
| Points and richness | `h_point`, `k_point`, `n_point`, `m_point`, `r1`, `r2`, `r4`, `indicator_b` |
| Concentration and evenness | `gini`, `camargo_evenness`, `sheldon_equitability`, `inverse_simpson_evenness`, `robin_hood` |
| QUITA and related indicators | `curve_length`, `lambda_indicator`, `b1`, `b2`, `b3`, `b4`, `b5`, `b6`, `b8`, `b10` |
| Thematic concentration | `thematic_concentration`, `secondary_thematic_concentration`, `proportional_thematic_concentration` |
| Moments and Ord criteria | `origin_moment`, `central_moment`, `moments`, `ords_i`, `ords_s`, `ords_criterion` |
| Zipf-family fitting | `zipf_fitted_parameters`, `zipf_mandelbrot_fitted_parameters`, `zipf_alekseev_fitted_parameters`, `rtmza_fitted_parameters` |

## Accepted Inputs

Most functions accept descending integer frequency sequences or `FreqDist`-like objects. Integer sequences are treated as raw counts and sorted in descending order when needed.

Floating-point sequences are treated as probability or relative-frequency distributions. They must be finite, non-negative, one-dimensional, and sum to approximately 1.

## Development Notes

This public release intentionally includes the frequency-distribution functionality only. Additional modules for length and syntactic-complexity measures are being held back until their external dependencies and documentation are ready for public installation.

## License

License information will be added here.
