Metadata-Version: 2.4
Name: pythonfaster
Version: 1.8.0
Summary: Transparent Python acceleration — compile .py to native code on import, zero source changes
Author: pythonfaster contributors
License: MIT
Project-URL: Homepage, https://github.com/15525730080/pythonfaster
Project-URL: Documentation, https://github.com/15525730080/pythonfaster#readme
Project-URL: Repository, https://github.com/15525730080/pythonfaster
Project-URL: Issues, https://github.com/15525730080/pythonfaster/issues
Keywords: python,acceleration,cython,aot,compilation,performance,speedup,native,import-hook
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Cython
Classifier: Topic :: Software Development :: Compilers
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: Cython>=3.0
Requires-Dist: setuptools>=68.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: build; extra == "dev"
Requires-Dist: twine>=4.0; extra == "dev"
Dynamic: license-file

# pythonfaster

> Transparent Python acceleration. Import your project once — it runs compiled.

`pythonfaster` hooks Python's import machinery and compiles your `.py` modules to native machine code via Cython, automatically. No decorators, no type annotations, no source changes. Type inference, class-layout conversion, and method-call lowering happen behind the import hook. If any step is unsafe or fails, it silently falls back to normal Python — your program never breaks.

[中文文档](README_zh-CN.md)

## What is this project?

`pythonfaster` is a project that makes Python fast. The goal in one sentence: without changing a single line of code, lift Python's runtime performance to the same tier as Node.js and Java. Python traditionally sits in the bottom tier of language performance — pythonfaster closes that gap through compilation.

## How much faster?

Measured on all 10 [Computer Language Benchmarks Game](https://salsa.debian.org/benchmarksgame-team/benchmarksgame/) (CLBG) programs at official reference workload sizes, then calibrated to the official CLBG language ranking (C = 1.00×; the multiplier means "how many times slower than C"):

![Cross-language ranking](benchmarks/results/current/language_comparison.png)

**CPython 3.12 runs the CLBG suite 55.1× slower than C. The same code under pythonfaster runs 7.04× slower than C — a 7.82× geometric-mean speedup that moves Python from the bottom tier into the same tier as Node.js (6.63×), within 2× of Java (3.69×), and 2.5× faster than PyPy.**

In the official 31-language CLBG ranking, that is a jump from 29th place to 19th — right between Node.js (18th) and Racket (20th). Key languages are excerpted below (full 31-language table in the appendix at the end):

| Rank | Language | vs C |
|---:|---|---:|
| 1 | C clang | 1.00× |
| 13 | Java | 3.69× |
| 18 | Node.js | 6.63× |
| **19** | **pythonfaster** | **7.04×** |
| 22 | PyPy | 17.41× |
| 29 | CPython 3.12 | 55.10× |

### Per-benchmark results (official workloads)

All 10 benchmarks pass strict output verification. Apple Silicon M3, Python 3.12.0, strict-mode pre-check enabled (default), fastest of multiple fresh-subprocess runs:

| Benchmark | Workload | CPython 3.12 | pythonfaster | Speedup |
|---|---:|---:|---:|---:|
| binary_trees | 21 | 43.97s | 0.37s | **117.6×** |
| spectral_norm | 5,500 | 133.08s | 1.62s | **82.3×** |
| mandelbrot | 16,000 | 300.58s | 10.83s | **27.8×** |
| fannkuch_redux | 12 | 373.77s | 21.32s | **17.5×** |
| nbody | 50,000,000 | 122.78s | 9.61s | **12.8×** |
| pidigits | 10,000 | 2.33s | 0.40s | **5.8×** |
| revcomp | 100M stdin | 14.05s | 10.02s | 1.40× |
| regex_redux | 5M stdin | 12.26s | 9.08s | 1.35× |
| fasta | 25,000,000 | 21.51s | 17.67s | 1.22× |
| knucleotide | 25M stdin | 142.02s | 133.01s | 1.07× |
| | | | **Geometric mean** | **7.82×** |

Compute-heavy code gets 5.8×–117.6×. I/O-heavy code (fasta, knucleotide, revcomp) sees modest gains because its hot path already runs inside C extensions of the standard library.

![Per-benchmark comparison](benchmarks/results/current/benchmark_comparison.png)

### Object-intensive workloads

Seven class-layout benchmarks (CPython 3.12 vs pythonfaster vs Java on the same programs):

| Benchmark | CPython | pythonfaster | Java | Speedup |
|---|---:|---:|---:|---:|
| linked_tree | 0.280s | 0.050s | 0.004s | **5.63×** |
| particle_sim | 0.140s | 0.039s | 0.006s | **3.62×** |
| typed_attrs | 0.762s | 0.286s | 0.012s | **2.67×** |
| interp_objects | 0.481s | 0.190s | 0.007s | **2.53×** |
| banking | 0.335s | 0.142s | 0.007s | **2.36×** |
| rat_arith | 0.225s | 0.215s | 0.002s | 1.05× |
| shapes_polymorphism | 0.148s | 0.221s | 0.009s | 0.67× |

Geometric mean 2.17×. The `shapes_polymorphism` regression is a known limit: virtual dispatch through non-cdef subclasses defeats method-call lowering, and that code is skipped from optimization — the cost is call-site indirection, not correctness.

## Quick start

```bash
pip install pythonfaster
pythonfaster enable
```

After that, Python processes auto-activate pythonfaster only for modules below a detected project root (a `pyproject.toml`, `pythonfaster.toml`, or `.git` marker in the CWD tree). Disable globally with `pythonfaster disable`, or per-process with `PYTHONFASTER_DISABLE=1`.

For explicit per-process activation:

```python
from pythonfaster import Config, install

install(Config(root="."))
```

## Requirements

- Python 3.10 or newer
- Cython 3.0 or newer (installed automatically as a package dependency)
- A working native C/C++ build toolchain compatible with the active Python interpreter (Xcode Command Line Tools on macOS, or GCC/Clang and Python development headers on Linux)
- A project-root marker: `pyproject.toml`, `pythonfaster.toml`, or `.git`

The package is designed for CPython. The cache key includes Python ABI, Cython version, compiler flags, CPU target, source bytes, and strategy-set version.

## Pipeline and safety model

```mermaid
flowchart TD
    A[import module.py] --> B{module under project root?}
    B -->|no| Z[standard Python import]
    B -->|yes| C[compute cache key]
    C --> D{cached extension?}
    D -->|yes| E[load extension]
    D -->|no| F{compatibility pre-check}
    F -->|unsafe pattern| Z
    F -->|safe| G[conservative AST transforms]
    G --> H{Cython build}
    H -->|transformed build fails| I[retry verbatim source]
    I -->|fails| Z
    H -->|success| E
    I -->|success| E
    E --> J{extension initialization succeeds?}
    J -->|yes| K[accelerated module]
    J -->|no| L[execute original .py and record skip]
    L --> K
```

The degradation path is intentional:

1. **Compatibility pre-check** skips known CPython/Cython semantic mismatches before compilation.
2. **Transform fallback** uses an unchanged `.py` copy if analysis declines or a transform fails.
3. **Build fallback** returns control to Python if Cython compilation fails.
4. **Runtime fallback** executes the original source, removes the failing artifact, and persistently skips that exact cache key after extension initialization fails.

The persistent cache and skip index live under `~/.cache/pythonfaster/<implementation>-<major><minor>/` by default. Set `PYTHONFASTER_CACHE_DIR` to relocate it.

## Optimization scope

The transform engine selectively applies optimizations only when its preconditions can be established:

- local integer/float/range type inference;
- cdef-class conversion, typed attributes, and direct method-call lowering;
- selected numeric-container unboxing;
- expression-helper inlining and `sum(generator)` loopification;
- recursive cdef lowering and integer-only loop-invariant `sum` folding.

Global Cython directives remain conservative. Per-function directives are injected only by transforms that establish their safety conditions.

## Reproducing the numbers

```bash
# Development-sized profile; suitable for a quick local check.
python benchmarks/full_compare.py --runs 3

# CLBG reference profile; expensive and may take a long time.
python benchmarks/full_compare.py --reference-profile --runs 3 --no-resume

# Render HTML only from the current JSON snapshots.
python benchmarks/render_report.py

# Object-intensive comparison (the authoritative object benchmark entry point).
python benchmarks/obj_compare.py
```

For publication, record macOS/Linux version, CPU, RAM, Python/Cython/compiler versions, compiler flags, exact command, run count, and summary statistic. Do not combine results from separate runs in one ranking.

## Testing

```bash
# Project regression tests: runtime extension failure must fall back to source.
python -m pytest -q

# CPython-standard-library compatibility probe (requires a CPython test-suite installation).
python tests/run_cpython_tests.py
```

`tests/run_cpython_tests.py` and `tests/speed_compare.py` clear the cache for the active interpreter dynamically; they do not assume a specific CPython minor version. The standard-library probe is an integration/compatibility measurement, not a replacement for focused unit tests of individual AST passes.

## Configuration

Configuration is loaded in this order (later sources override earlier ones):

1. `[pythonfaster]` in `pythonfaster.toml`, or `[tool.pythonfaster]` in `pyproject.toml`;
2. environment variables;
3. defaults.

| Option | Default | Meaning |
|---|---:|---|
| `enabled` | `true` | Master switch |
| `exclude` | built-in patterns | Additional `fnmatch` paths not to compile |
| `aggressive` | `false` | Higher-risk transforms (off by default; currently no-op) |

Useful environment variables: `PYTHONFASTER_DISABLE=1`, `PYTHONFASTER_STRICT=0|1`, `PYTHONFASTER_VERBOSE=1`, `PYTHONFASTER_CACHE_DIR`, `PYTHONFASTER_MARCH_NATIVE=1`, and `PYTHONFASTER_KEEP_INTERMEDIATE=1`.

`PYTHONFASTER_STRICT=1` is the default. It enables all known compatibility guards. Setting it to `0` may improve coverage but accepts known Cython semantic risks. `PYTHONFASTER_MARCH_NATIVE=1` can improve local performance but makes cached artifacts CPU-specific; use it only for a cache that is not shared between different machines.

## Limitations

- Compilation adds cold-start cost; it is most appropriate for repeatedly imported or long-running numeric workloads.
- Python/Cython semantic compatibility is not universal. Unsupported or risky modules deliberately fall back to Python.
- String- and I/O-bound code often sees limited benefit because its expensive work already occurs in C extensions or the standard library.
- Native compilation requires a compatible local toolchain; failure to build is non-fatal and falls back to Python.
- Benchmarks are sensitive to hardware, thermal state, compiler, input size, and cache warmness. Treat checked-in JSON as evidence for its recorded environment only.

## Project layout

```text
pythonfaster/
├── pythonfaster/              # import hook, config, cache, compiler, transforms
├── benchmarks/
│   ├── full_compare.py        # multi-runtime CLBG comparison
│   ├── obj_compare.py         # object-intensive comparison
│   ├── render_report.py       # render report from current JSON data
│   └── results/current/       # latest raw benchmark snapshots and generated report
├── tests/
│   ├── test_runtime_fallback.py
│   ├── run_cpython_tests.py
│   └── speed_compare.py
└── pyproject.toml
```

## Publishing to PyPI

### One-time setup

```bash
pip install build twine
# or: pip install -e ".[dev]"
```

Configure PyPI credentials (create an API token at [pypi.org/manage/account](https://pypi.org/manage/account)):

```bash
# ~/.pypirc
[pypi]
username = __token__
password = pypi-XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
```

### Build and upload

```bash
# 1. Clean previous builds
rm -rf dist/ build/ *.egg-info

# 2. Build source distribution + wheel
python -m build

# 3. Verify the package
python -m twine check dist/*

# 4. Upload to Test PyPI first (recommended)
python -m twine upload --repository testpypi dist/*

# 5. Upload to PyPI
python -m twine upload dist/*
```

### Verify the release

```bash
pip install --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple pythonfaster
# or for the production release:
pip install pythonfaster
```

### Automated release (GitHub Actions)

A sample `.github/workflows/publish.yml` for automatic release on tag push:

```yaml
name: Publish to PyPI
on:
  push:
    tags:
      - 'v*'
jobs:
  publish:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: '3.12'
      - run: pip install build twine
      - run: python -m build
      - run: python -m twine upload dist/*
        env:
          TWINE_USERNAME: __token__
          TWINE_PASSWORD: ${{ secrets.PYPI_API_TOKEN }}
```

## License

MIT

## Appendix: Full 31-language CLBG ranking

<details>
<summary>Full 31-language ranking with pythonfaster inserted</summary>

Calibrated to the official CLBG language ranking (C = 1.00×; the multiplier means "how many times slower than C"). Inserted, pythonfaster jumps from CPython 3.12's 29th place to 19th:

| Rank | Language | vs C |
|---:|---|---:|
| 1 | C clang | 1.00× |
| 2 | C++ g++ | 1.12× |
| 3 | Rust | 1.17× |
| 4 | Julia | 2.32× |
| 5 | C# .NET | 2.39× |
| 6 | Chapel | 2.52× |
| 7 | Fortran | 2.55× |
| 8 | Ada 2012 GNAT | 2.89× |
| 9 | F# .NET | 3.23× |
| 10 | Go | 3.35× |
| 11 | Haskell GHC | 3.45× |
| 12 | Free Pascal | 3.63× |
| 13 | Java | 3.69× |
| 14 | OCaml | 3.94× |
| 15 | Swift | 5.50× |
| 16 | Lisp SBCL | 5.55× |
| 17 | Dart | 6.37× |
| 18 | Node.js | 6.63× |
| **19** | **pythonfaster** | **7.04×** |
| 20 | Racket | 7.56× |
| 21 | PHP | 12.74× |
| 22 | PyPy | 17.41× |
| 23 | Pyston | 28.60× |
| 24 | Erlang | 33.66× |
| 25 | Ruby | 39.01× |
| 26 | VW Smalltalk | 41.11× |
| 27 | CinderX JIT 3.14 | 47.59× |
| 28 | CPython 3.14 | 53.27× |
| 29 | CPython 3.12 | 55.10× |
| 30 | Lua | 61.34× |
| 31 | Perl | 67.97× |

</details>
