Metadata-Version: 2.4
Name: trellis-flow
Version: 0.1.0
Summary: Declarative DAG orchestration for agent and skill workflows.
Author: karolsabier06-cmyk
License: MIT
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: PyYAML>=6.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Dynamic: license-file

# Trellis

Declarative DAG orchestration for agent and skill workflows.

Trellis sits between "chain some CLI calls together by hand" and a full
batch-job platform like Airflow: you describe a workflow as a YAML file —
steps, dependencies, retries, timeouts — and Trellis runs it, persisting
every step's status and output so you can inspect what happened after the
fact, even if the run already finished.

It's built for the shape of problem that comes up when composing
agents/skills rather than batch data jobs: steps that call out to shell
commands (CLI-driven agents, Claude Code skills), HTTP endpoints (MCP
servers, other services), or in-process Python functions (custom agents
already living in your codebase) — with real-time scheduling, not a nightly
cron tick.

## Why not just a shell script?

A shell script with `&&` gives you sequencing but not: parallelism between
independent steps, retries with backoff, per-step timeouts, a record of what
ran and what it returned, or a way to answer "which step failed, and what
did the step before it output?" six hours after the process exited. Trellis
gives you all of that from one YAML file with no server to run.

## Install

```bash
pip install -e .
```

This registers the `trellis` command (see `pyproject.toml` — it's a normal
`pip`-installable package, `trellis-flow` on PyPI naming, importable as
`trellis`). Requires Python 3.9+. The only runtime dependency is PyYAML.

## Quickstart

```bash
trellis validate examples/claude_code_skill_chain.yaml
trellis run examples/claude_code_skill_chain.yaml
trellis runs
trellis show <run-id>
```

## Writing a workflow

```yaml
name: my-workflow

steps:
  - id: fetch
    type: shell
    command: "curl -s https://api.example.com/data"
    retry:
      max_attempts: 3
      backoff_seconds: 5
    timeout_seconds: 30

  - id: process
    type: python
    depends_on: [fetch]
    module: my_agents
    function: process_data
    params:
      raw: "{{ steps.fetch.output.stdout }}"

  - id: notify
    type: http
    depends_on: [process]
    method: POST
    url: "http://localhost:8080/notify"
    body:
      message: "{{ steps.process.output.summary }}"
```

Steps run as soon as their `depends_on` are satisfied — independent steps
run concurrently (bounded by `trellis run --workers N`, default 4).

### Step types

| Type     | Use for                                                | Config fields |
|----------|---------------------------------------------------------|---------------|
| `shell`  | CLI-driven agents, Claude Code skills, any command       | `command`, `cwd`, `env` |
| `http`   | MCP servers, REST APIs, any HTTP-reachable agent          | `method`, `url`, `headers`, `body`, `allow_error_status` |
| `python` | Custom Python agents already in your codebase, in-process | `module`, `function`, `params` |

All fields except `id`, `type`, `depends_on`, `retry`, `timeout_seconds`,
`allow_failure` are passed through to the step's own config and support
`{{ steps.<id>.output... }}` references, resolved just before the step runs.

### Retries, timeouts, failure handling

- `retry: {max_attempts: N, backoff_seconds: S}` — retries with linear
  backoff (`S * attempt_number` between attempts). Default: no retries.
- `timeout_seconds` — enforced for `shell` (kills the subprocess) and `http`
  (socket timeout). For `python` steps this is best-effort: the call isn't
  forcibly killed if it hangs (a CPython thread limitation) — use `shell` if
  you need hard timeout enforcement for long-running code.
- `allow_failure: true` — a step that exhausts its retries is recorded as
  `failed_allowed` and the workflow continues instead of failing.
- A step that fails (without `allow_failure`) marks every step that
  transitively depends on it as `skipped`; independent branches still run.

### Referencing earlier output

`{{ steps.<id>.output }}` or `{{ steps.<id>.output.<path> }}` (dot-path into
the step's output). Output shape per step type:

- `shell` → `{"returncode": int, "stdout": str, "stderr": str}`
- `http` → `{"status_code": int, "body": <parsed JSON or raw text>}`
- `python` → whatever the function returns; non-dict returns are wrapped as `{"result": ...}`

## Observability

Every run gets an id. State (run + per-step status, attempts, timing,
output, errors) is persisted to SQLite at `.trellis/state.db`. A structured
JSONL trace is written to `.trellis/runs/<run-id>/trace.jsonl` as the run
happens (and mirrored to stdout live).

```bash
trellis runs                # recent runs
trellis show <run-id>       # per-step status, duration, output, errors
trellis logs <run-id>       # raw trace events
trellis logs <run-id> --step process   # filter to one step
```

Override the state location with `trellis --state-dir /path/to/dir <cmd>`.

## Examples

Three runnable examples in `examples/`, one per step type / integration
target named in the brief:

- `claude_code_skill_chain.yaml` — chains shell-invoked steps the way you'd
  chain Claude Code skills or any CLI-driven agent, with retries and
  output-passing between steps.
- `mcp_server_call.yaml` + `mock_mcp_server.py` — calls an MCP-style tool
  server over HTTP. Start the mock server first (`python
  examples/mock_mcp_server.py`), then run the workflow in another terminal.
  Point `url` at a real MCP gateway route to use it for real.
- `custom_python_agent.yaml` + `python_agents.py` — calls custom Python
  agent functions in-process, passing one step's output into the next
  step's params.

Run them from the project root so relative module/paths resolve:

```bash
trellis run examples/claude_code_skill_chain.yaml
trellis run examples/custom_python_agent.yaml
python examples/mock_mcp_server.py &        # separate terminal in practice
trellis run examples/mcp_server_call.yaml
```

## Development

```bash
pip install -e ".[dev]"
pytest
```

24 tests cover DAG validation (cycles, unknown refs, duplicate ids),
dependency-aware scheduling, retry/backoff behavior, failure propagation to
downstream steps, `allow_failure`, output-reference resolution, and state
persistence (including across separate `StateStore` instances, i.e. after a
process restart).

## Known limitations (honest, not hidden)

- Single-machine only — no distributed workers. Fine for orchestrating calls
  *out* to other services/agents; not a replacement for a distributed task
  queue.
- `python` step timeouts are best-effort (see above).
- No built-in scheduling/cron — Trellis runs a workflow once, on demand;
  wrap `trellis run` in cron/systemd if you need periodic execution.
- No workflow-level UI yet — `show`/`logs` are CLI-only.
