Metadata-Version: 2.4
Name: nexilis
Version: 1.0.2
Summary: Automatic SageMaker Pipeline Generation from DAG Specifications
Author-email: Tianpei Xie <unidoctor@gmail.com>
Maintainer-email: Tianpei Xie <unidoctor@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/TianpeiLuke/nexilis
Project-URL: Documentation, https://nexilis.readthedocs.io/
Project-URL: Repository, https://github.com/TianpeiLuke/nexilis
Project-URL: Issues, https://github.com/TianpeiLuke/nexilis/issues
Project-URL: Changelog, https://github.com/TianpeiLuke/nexilis/blob/main/CHANGELOG.md
Keywords: sagemaker,pipeline,dag,machine-learning,aws,automation,mlops,data-science,workflow,orchestration
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Information Technology
Classifier: Operating System :: OS Independent
Classifier: Environment :: Console
Classifier: Natural Language :: English
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Software Development :: Code Generators
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: System :: Distributed Computing
Classifier: Topic :: System :: Monitoring
Classifier: Typing :: Typed
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: boto3>=1.39.0
Requires-Dist: botocore>=1.39.0
Requires-Dist: sagemaker<3,>=2.248.0
Requires-Dist: pydantic<3,>=2.11.0
Requires-Dist: PyYAML>=6.0.0
Requires-Dist: networkx>=2.8.0
Requires-Dist: click>=8.1.0
Requires-Dist: pytz
Provides-Extra: processing
Requires-Dist: pandas>=2.1.0; extra == "processing"
Requires-Dist: numpy>=1.26.0; extra == "processing"
Requires-Dist: scikit-learn>=1.3.0; extra == "processing"
Requires-Dist: pyarrow>=14.0.0; extra == "processing"
Provides-Extra: pytorch
Requires-Dist: torch>=2.0.0; extra == "pytorch"
Requires-Dist: pytorch-lightning>=2.0.0; extra == "pytorch"
Requires-Dist: lightning>=2.0.0; extra == "pytorch"
Provides-Extra: gbm
Requires-Dist: xgboost>=2.0.0; extra == "gbm"
Requires-Dist: lightgbm>=4.0.0; extra == "gbm"
Requires-Dist: pandas>=2.1.0; extra == "gbm"
Requires-Dist: numpy>=1.26.0; extra == "gbm"
Requires-Dist: scikit-learn>=1.3.0; extra == "gbm"
Provides-Extra: nlp
Requires-Dist: transformers>=4.30.0; extra == "nlp"
Requires-Dist: tokenizers>=0.15.0; extra == "nlp"
Provides-Extra: jupyter
Requires-Dist: jupyter>=1.0.0; extra == "jupyter"
Requires-Dist: ipywidgets>=8.0.0; extra == "jupyter"
Requires-Dist: plotly>=5.0.0; extra == "jupyter"
Requires-Dist: nbformat>=5.0.0; extra == "jupyter"
Requires-Dist: jinja2>=3.0.0; extra == "jupyter"
Requires-Dist: seaborn>=0.12.0; extra == "jupyter"
Requires-Dist: jupyterlab>=4.0.0; extra == "jupyter"
Requires-Dist: ipython>=8.0.0; extra == "jupyter"
Provides-Extra: viz
Requires-Dist: matplotlib>=3.8.0; extra == "viz"
Requires-Dist: seaborn>=0.12.0; extra == "viz"
Requires-Dist: plotly>=5.0.0; extra == "viz"
Provides-Extra: dev
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0.0; extra == "dev"
Requires-Dist: pytest-mock>=3.10.0; extra == "dev"
Requires-Dist: black>=23.0.0; extra == "dev"
Requires-Dist: isort>=5.12.0; extra == "dev"
Requires-Dist: flake8>=6.0.0; extra == "dev"
Requires-Dist: mypy>=1.0.0; extra == "dev"
Requires-Dist: pre-commit>=3.0.0; extra == "dev"
Provides-Extra: docs
Requires-Dist: sphinx<9,>=7.2; extra == "docs"
Requires-Dist: furo==2025.12.19; extra == "docs"
Requires-Dist: myst-parser>=2.0.0; extra == "docs"
Requires-Dist: linkify-it-py>=2.0; extra == "docs"
Requires-Dist: sphinx-click>=6.0; extra == "docs"
Requires-Dist: jinja2>=3.1; extra == "docs"
Requires-Dist: pyyaml>=6.0; extra == "docs"
Provides-Extra: mcp
Requires-Dist: mcp<2,>=1.2.0; extra == "mcp"
Requires-Dist: anyio>=4.0.0; extra == "mcp"
Provides-Extra: all
Requires-Dist: nexilis[gbm,jupyter,mcp,nlp,processing,pytorch,viz]; extra == "all"
Dynamic: license-file

# Nexilis: Automatic SageMaker Pipeline Generation

[![Python 3.9+](https://img.shields.io/badge/python-3.9+-blue.svg)](https://www.python.org/downloads/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

**Turn a pipeline graph plus a JSON config into a complete, production-ready SageMaker pipeline — automatically.**

Nexilis is a specification-driven pipeline generation system for Amazon SageMaker. You describe your ML workflow as a **DAG of step names**; Nexilis resolves the dependencies between steps, wires their inputs and outputs, looks up each step's declarative interface, and assembles a runnable `sagemaker.workflow.pipeline.Pipeline`. You say *what* the pipeline is — Nexilis figures out *how* to build it.

> **📖 Full documentation:** the [`docs/`](https://github.com/TianpeiLuke/nexilis/tree/main/docs) tree — getting-started, tutorials, concepts & architecture, and the complete API / CLI / MCP / step-catalog reference. It builds as a Sphinx site with `make -C docs html`.

---

## Installation

Install from a clone — the first PyPI release is still ahead:

```bash
git clone https://github.com/TianpeiLuke/nexilis.git
cd nexilis
pip install -e .
```

Requires **Python 3.9+** and targets the **SageMaker Python SDK 2.x** (`sagemaker>=2.248.0,<3`).

Optional extras keep heavy ML/data libraries out of the core install — pull in only what your steps run:

```bash
pip install -e ".[processing]"   # pandas / numpy data-processing utilities
pip install -e ".[pytorch]"      # PyTorch / Lightning
pip install -e ".[gbm]"          # XGBoost / LightGBM
pip install -e ".[nlp]"          # tokenizers / transformers
pip install -e ".[mcp]"          # MCP server for LLM agents (mcp + anyio)
pip install -e ".[all]"          # everything
```

Verify:

```bash
nexilis --version
python -c "import nexilis; print(nexilis.__version__)"
```

---

## Quick Start

There are three ways in, highest-level first. They all compile through the same engine, so the same DAG + config produces the same pipeline.

Every path needs two ingredients: a **DAG** (step names + edges) and a **config JSON** whose `metadata.config_types` maps each node to a configuration class. Compilation is offline; you only need AWS credentials to *deploy* (`--upsert`) or *run* (`--start`).

### 1. Start from the pre-built pipeline catalog

Nexilis ships **60 validated DAGs** across 8 frameworks. Let the router recommend one and build it:

```python
from nexilis.pipeline_catalog import recommend_dag, load_shared_dag
from nexilis import PipelineDAGCompiler

# recommend_dag returns a ranked list of matches (dicts with 'id', 'score', ...)
recommendations = recommend_dag(framework="xgboost", task_type="end_to_end")
dag = load_shared_dag(recommendations[0]["id"])

pipeline, report = PipelineDAGCompiler(config_path="config.json").compile_with_report(dag)
print(pipeline.name, "-", len(pipeline.steps), "steps")
```

### 2. Compile from the command line

Reproducible, no-glue path — point it at a DAG JSON and a config JSON:

```bash
# compile only (writes the SageMaker pipeline definition to a file)
nexilis compile -d my_dag.json -c my_config.json -o pipeline.json

# validate DAG <-> config alignment without compiling
nexilis compile -d my_dag.json -c my_config.json --validate-only

# compile, deploy to SageMaker, and start an execution
nexilis compile -d my_dag.json -c my_config.json \
    --upsert --start --role arn:aws:iam::123456789012:role/MySageMakerRole
```

### 3. Build a DAG in Python

```python
from nexilis.api import PipelineDAG
from nexilis.core import compile_dag_to_pipeline

# Nodes are step names; edges are data dependencies
dag = PipelineDAG()
for node in ["DummyDataLoading", "TabularPreprocessing", "XGBoostTraining"]:
    dag.add_node(node)
dag.add_edge("DummyDataLoading", "TabularPreprocessing")
dag.add_edge("TabularPreprocessing", "XGBoostTraining")

# config.json maps each node -> a config class (metadata.config_types)
pipeline = compile_dag_to_pipeline(dag=dag, config_path="config.json")

# Deploy / run when ready
pipeline.upsert(role_arn="arn:aws:iam::123456789012:role/MySageMakerRole")
pipeline.start()
```

See the [Quickstart guide](https://github.com/TianpeiLuke/nexilis/blob/main/docs/getting_started/quickstart.md) for the full walkthrough.

---

## Key Features

- **🎯 Graph-to-pipeline automation** — a DAG of step names compiles straight to a SageMaker pipeline; the SageMaker step objects, wiring, and naming are generated for you.
- **🧠 Intelligent dependency resolution** — Nexilis infers step connections and data flow by matching each step's declared outputs to the next step's declared inputs (semantic scoring), instead of hand-wiring `properties` paths.
- **📄 Declarative, data-driven steps** — every step is a single `<step>.step.yaml` interface unifying the script contract (I/O, env vars, job arguments) and the dependency spec; step builders are synthesized at runtime, with no hand-written builder classes to maintain.
- **📦 A pre-built pipeline catalog** — 60 ready-to-use DAGs across XGBoost, PyTorch, LightGBM, Bedrock and more, discoverable by framework, task type, and complexity.
- **🧩 Extensible via step packs** — define your own steps in a folder *outside* the installed package and Nexilis discovers them as native, strictly additively (built-in steps are never removed).
- **🛡️ Built-in validation** — DAG↔config alignment, interface conformance, and dependency resolution are checked before you deploy (`nexilis validate`, `nexilis alignment`, `--validate-only`).
- **🤖 Agent-ready (MCP)** — a framework-neutral, self-documenting tool surface of **55 JSON-in/JSON-out tools** across 11 namespaces mirrors the CLI/API for LLM agents. `pip install -e ".[mcp]"`, then wire `nexilis-mcp` into any MCP host (Claude Desktop, Cursor, Kiro). **Read-only by default** — state-changing tools are opt-in via env vars — with per-tool safety annotations. See the [MCP server guide](https://github.com/TianpeiLuke/nexilis/blob/main/src/nexilis/mcp/README.md); `nexilis mcp help` to inspect the tools.

---

## How It Works

A DAG + config flows through layered subsystems to a SageMaker pipeline:

| Subsystem | Package | Responsibility |
|---|---|---|
| **DAG model** | `nexilis.api.dag` | `PipelineDAG` — nodes (step names) + edges (dependencies) |
| **Compiler** | `nexilis.core.compiler` | `PipelineDAGCompiler` / `compile_dag_to_pipeline` — validate → resolve → build → assemble |
| **Assembler** | `nexilis.core.assembler` | Turns resolved steps into a `sagemaker` `Pipeline` |
| **Dependency resolver** | `nexilis.core.deps` | Matches producer outputs to consumer inputs (semantic scoring) |
| **Step interfaces** | `nexilis.core.base` + `nexilis.steps.interfaces` | Declarative `<step>.step.yaml`; builders synthesized at runtime |
| **Registry & discovery** | `nexilis.registry` + `nexilis.step_catalog` | Canonical step table, derived interface-first; step-pack discovery |
| **Config system** | `nexilis.core.config_fields` + `nexilis.api.factory` | Pydantic config classes; `metadata.config_types` node→class map |
| **Pipeline catalog** | `nexilis.pipeline_catalog` | Pre-built shared DAGs + router (`recommend_dag` / `load_shared_dag`) |
| **Validation** | `nexilis.validation` | Alignment / interface / dependency checks |
| **Agent surface** | `nexilis.mcp` | The MCP tool surface |
| **CLI** | `nexilis.cli` | 13 command groups |

Read the [Concepts, Architecture & Design](https://github.com/TianpeiLuke/nexilis/blob/main/docs/concepts/index.md) docs for the full picture.

---

## What's Included

| | Count | |
|---|---|---|
| **Step types** | 53 registered (50 declarative `.step.yaml` interfaces) | data loading, preprocessing, training, eval, calibration, packaging, … |
| **Pre-built DAGs** | 60 across 8 frameworks | XGBoost, XGBoost-MT, PyTorch, LightGBM, LightGBM-MT, Bedrock, Dummy, Generic |
| **CLI command groups** | 13 | `compile`, `dag`, `config`, `catalog`, `steps`, `strategies`, `pipeline-catalog`, `validate`, `alignment`, `registry`, `projects`, `exec-doc`, `mcp` |
| **MCP agent tools** | 55 across 11 namespaces | discover, construct, validate, compile, author — for LLM agents |

---

## Command-Line Interface

```bash
nexilis compile -d dag.json -c config.json -o pipeline.json   # DAG + config -> pipeline
nexilis compile -d dag.json -c config.json --validate-only    # dry-run alignment report
nexilis pipeline-catalog recommend --framework xgboost        # discover a pre-built DAG
nexilis pipeline-catalog get-dag xgboost_complete_e2e         # inspect a catalog DAG
nexilis catalog list                                          # browse available step types
nexilis steps io XGBoostTraining                              # a step's declared I/O
nexilis dag resolve TabularPreprocessing XGBoostTraining      # score dependency edges
nexilis validate step-interface --all                         # validate every interface
nexilis projects list                                         # discover pipeline projects
nexilis mcp help                                              # explore the agent tool surface
```

Every group, subcommand, and flag is in the [CLI reference](https://github.com/TianpeiLuke/nexilis/blob/main/docs/cli.rst) — rendered from the live Click app when the docs build — or a `nexilis --help` away.

---

## Installation Options

| Extra | Installs | For |
|---|---|---|
| *(core)* | DAG model, compiler, catalog, registry, CLI | everything except heavy ML/data libs |
| `processing` | pandas, numpy | data-processing utilities & scripts |
| `pytorch` | torch, lightning | PyTorch training/eval steps |
| `gbm` | xgboost, lightgbm | gradient-boosting steps |
| `nlp` | tokenizers, transformers | text steps |
| `mcp` | mcp, anyio | the `nexilis-mcp` agent server |
| `jupyter` | notebook tooling | interactive development |
| `viz` | plotting libraries | reports/visualizations |
| `docs` | sphinx, furo, sphinx-click, … | building this documentation |
| `dev` | test/lint toolchain | contributing |
| `all` | pytorch + gbm + nlp + processing + jupyter + viz + mcp | full runtime install |

```bash
pip install -e ".[all]"          # full ML runtime
pip install -e ".[dev]"          # contributor toolchain
```

---

## 📖 Documentation

### 📚 [The `docs/` tree](https://github.com/TianpeiLuke/nexilis/tree/main/docs)
**The full documentation, versioned with the code.** It is a Sphinx site — `pip install -e ".[docs]"`, then `make -C docs html` and open `docs/_build/html/index.html`:
[Getting Started](https://github.com/TianpeiLuke/nexilis/blob/main/docs/getting_started/index.md) ·
[Tutorials](https://github.com/TianpeiLuke/nexilis/blob/main/docs/tutorials/index.md) ·
[Concepts & Architecture](https://github.com/TianpeiLuke/nexilis/blob/main/docs/concepts/index.md) ·
[How-to Guides](https://github.com/TianpeiLuke/nexilis/blob/main/docs/guides/index.md) ·
[Step Catalog](https://github.com/TianpeiLuke/nexilis/blob/main/docs/steps-catalog/index.md) ·
[Migration](https://github.com/TianpeiLuke/nexilis/blob/main/docs/migration/index.md)

The Python API, CLI, and MCP-tool references are generated from the installed package at build time ([`docs/api/`](https://github.com/TianpeiLuke/nexilis/blob/main/docs/api/index.rst), [`docs/cli.rst`](https://github.com/TianpeiLuke/nexilis/blob/main/docs/cli.rst), [`docs/reference/`](https://github.com/TianpeiLuke/nexilis/blob/main/docs/reference/index.md)), so they never drift from the version you have. Build the site to read them, or ask the terminal: `nexilis --help`, `nexilis mcp help`.

### Design & developer notes (in-repo)
- **[Author a Custom Step](https://github.com/TianpeiLuke/nexilis/blob/main/docs/tutorials/author_a_step.md)** — the three files a new step needs, end to end
- **[Define a Step Pack](https://github.com/TianpeiLuke/nexilis/blob/main/docs/guides/define_a_step_pack.md)** — extending Nexilis with your own steps from outside the package
- **[Concepts, Architecture & Design](https://github.com/TianpeiLuke/nexilis/blob/main/docs/concepts/index.md)** — subsystem mechanics, plus the [design principles](https://github.com/TianpeiLuke/nexilis/blob/main/docs/concepts/design_principles.md) behind them
- **[Pipeline Catalog](https://github.com/TianpeiLuke/nexilis/blob/main/src/nexilis/pipeline_catalog/README.md)** — the prebuilt-DAG collection
- **[MCP Server](https://github.com/TianpeiLuke/nexilis/blob/main/src/nexilis/mcp/README.md)** — the agent tool surface and its safety model
- **[Changelog](https://github.com/TianpeiLuke/nexilis/blob/main/CHANGELOG.md)** — release history

---

## Who Should Use Nexilis?

- **Data scientists & ML practitioners** — go from a workflow sketch to a running SageMaker pipeline without writing SageMaker step glue; start from a catalog template and customize.
- **Platform & ML engineers** — standardize pipeline construction, enforce DAG↔config alignment in CI, and extend the step library with your own step packs.
- **Organizations** — a consistent, validated, reproducible path from graph to production pipeline, with less bespoke SageMaker code to maintain.

---

## 🤝 Contributing

Contributions are welcome! Start with `pip install -e ".[dev]"`, then:

- **[Installation](https://github.com/TianpeiLuke/nexilis/blob/main/docs/getting_started/installation.md)** — install options and how to verify one
- **[Author a Custom Step](https://github.com/TianpeiLuke/nexilis/blob/main/docs/tutorials/author_a_step.md)** — the interface, config, and script a new step needs
- **[Define a Step Pack](https://github.com/TianpeiLuke/nexilis/blob/main/docs/guides/define_a_step_pack.md)** — adding steps without touching the package
- **[Validate a Pipeline](https://github.com/TianpeiLuke/nexilis/blob/main/docs/guides/validate_a_pipeline.md)** — the checks to run before you push, and how to read them
- **[Design Principles](https://github.com/TianpeiLuke/nexilis/blob/main/docs/concepts/design_principles.md)** — the load-bearing decisions the codebase follows

Then let the test suite and the validators have a look: `pytest`, `nexilis validate step-interface --all`, `nexilis alignment validate-all`.

Or author a step the guided way with the [`nexilis mcp`](https://github.com/TianpeiLuke/nexilis/blob/main/docs/tutorials/agentic.md) agent tools.

## 📄 License

Licensed under the MIT License — see [LICENSE](https://github.com/TianpeiLuke/nexilis/blob/main/LICENSE).

## 🔗 Links

- **Documentation**: [`docs/`](https://github.com/TianpeiLuke/nexilis/tree/main/docs) — Sphinx source; `make -C docs html` to build the site
- **GitHub**: https://github.com/TianpeiLuke/nexilis
- **Issues**: https://github.com/TianpeiLuke/nexilis/issues
- **Changelog**: [CHANGELOG.md](https://github.com/TianpeiLuke/nexilis/blob/main/CHANGELOG.md)
