Metadata-Version: 2.5
Name: controlloop
Version: 0.1.0
Summary: Local-first security scanner and deterministic CI gate for AI agents
License-Expression: Apache-2.0
License-File: LICENSE
Requires-Python: <3.13,>=3.12
Requires-Dist: pathspec<2,>=1.1.1
Requires-Dist: pydantic<3,>=2.12
Requires-Dist: python-hcl2>=8.1.2
Requires-Dist: pyyaml>=6.0.3
Requires-Dist: typer<1,>=0.21
Description-Content-Type: text/markdown

# ControlLoop CLI

**A local-first scanner and deterministic CI gate for what your AI agents can
actually do.**

An agent's real authority is spread across an agent card, an MCP config, an
OpenAPI spec, an IAM policy -- and, for most agents, the framework code
itself: a CrewAI crew, a LangGraph graph, an OpenAI Agents SDK handoff
chain. No single file tells you that adding one
skill just gave it the ability to move money. ControlLoop reads those files,
builds one graph of the agent's capabilities, compares it to the state you
approved, and fails your build when the agent's authority grows past the
ceiling you declared.

It runs entirely on your machine. It never executes your code, never starts
your MCP servers, never contacts a declared endpoint, and never sends your
source anywhere.

> **Development preview.** `init`, `bom`, `diff`, `scan`, `gate`, `rules`, and
> `example` all work
> and are covered by tests. There is no published release yet — no PyPI
> package, no Homebrew tap, no container image, no tagged version. Install
> from source as below. See [What is not built yet](#what-is-not-built-yet).

---

## Why use it

- **Catch capability growth in review, not after.** A pull request that adds
  a bulk-refund tool is blocked with the reason, the path, and a retained
  evidence artifact — before it merges.
- **Declare a ceiling and have it hold.** Pattern rules only catch what
  someone anticipated. A capability ceiling blocks authority you never
  imagined, through routes no rule matches.
- **Deterministic.** Same inputs, same pinned versions, same digest, same
  verdict. No model decides whether your build passes.
- **Nothing protected leaves.** Source, prompts, credentials, environment
  values, and customer data never appear in output, logs, errors, identity
  material, or artifacts.
- **Free and local.** Core scanning works with the network disabled.

## Install and first run

Requires **Python 3.12**, [**uv**](https://docs.astral.sh/uv/), and Git.

```bash
cd cli
uv sync --locked --all-groups
uv run controlloop --version
```

```
ControlLoop 0.1.0
Supported OASF versions: none yet
Rule pack: controlloop-community 0.2.0 (10 rules)
```

Point it at a repository containing an agent:

```bash
uv run controlloop scan /path/to/your/agent-repo
```

Nothing to try it on? The public
[reference agent](https://github.com/danny-tijerina2/controlloop-reference-agent)
is a realistic, deliberately imperfect agent built for exactly this:

```bash
git clone https://github.com/danny-tijerina2/controlloop-reference-agent
uv run controlloop scan controlloop-reference-agent/scenarios/order-support-agent
```

## The commands

New here? Run `controlloop example` — it prints a five-step walkthrough and
points at the [Example walkthrough](../docs/guides/example-walkthrough.md)
guide.

| Command | Use it to |
| --- | --- |
| `example` | Print a guided walkthrough of how ControlLoop works |
| `rules` | List the rule pack in use and what each rule looks for |
| [`init`](#init) | Draft a `controlloop.yaml` from what's actually in the repo |
| [`bom`](#bom) | Generate or validate the capability graph (ACBOM) |
| [`diff`](#diff) | See what changed against a baseline |
| [`scan`](#scan) | Full report for a human: graph, policy, findings |
| [`gate`](#gate) | One verdict and one exit code, for CI |

`scan` is what you run at your desk. `gate` is what you run in CI. Both do
the same analysis; `gate` never prompts, always writes an evidence artifact,
and returns a verdict designed to be acted on by a machine.

### `init`

```bash
uv run controlloop init [REPOSITORY] [--dry-run]
```

Discovers candidate agents and writes a reviewable `controlloop.yaml`. Every
field is present even where ControlLoop could not determine a value — those
are marked `TODO` with a comment saying what to supply. It evaluates no
policy and enforces nothing.

`--dry-run` prints the draft instead of writing it. Real output:

```
# controlloop.yaml
# ControlLoop security manifest -- generated by `controlloop init`.
...
# Manifest schema version 1.1.0.

schema_version: "1.1.0"

# TODO: owner
# ControlLoop cannot determine who is accountable for this
# repository's agents. Replace the placeholder below with an
# identifier for the responsible team or individual (for
# example: "platform-security" or "jane.doe").
owner: "TODO-owner"

trust_boundaries:
  # Discovered candidate agents (from repository discovery):
  #   - invoice-agent  (tools observed: Send an invoice)
  - name: "TODO-trust-boundary-name"
```

### `bom`

```bash
uv run controlloop bom [REPOSITORY] [-o PATH] [--validate] [--format acbom|cyclonedx]
```

Builds the **Agent Capability Bill of Materials** — the canonical graph of
agents, tools, resources, identities, and permissions, with evidence for
every node and edge.

```bash
uv run controlloop bom ./agent-repo -o acbom.json
uv run controlloop bom acbom.json --validate
```

```
✓ VALID  acbom.json is a valid ACBOM
```

`--validate` checks an existing file against this reader's own schema
version. Use `--format cyclonedx` for a CycloneDX projection.

### `diff`

```bash
uv run controlloop diff [FIRST] [SECOND] [--baseline REF] [--format text|json]
```

- `diff <repo>` — compare the repository against its resolved baseline
  (a committed `.controlloop/baseline.json`, else the merge-base of your
  default branch).
- `diff <baseline.json> <current.json>` — compare two ACBOM files directly.
- `--baseline` — an explicit baseline file *or* git revision.

```
• BASELINE   committed baseline: .controlloop/baseline.json
• SCOPE      effective authority only -- node and edge changes were not evaluated

• ADDED    capability  clp:agent:sha256:2321fe51…  invokes  clp:tool:sha256:ccc68b72…

1 change: 1 added
```

### `scan`

```bash
uv run controlloop scan [REPOSITORY] [--format text|json|sarif]
```

Discovery, graph, baseline comparison, and all ten policy rules, rendered as
a full report. `--format json` gives you the complete report;
`--format sarif` gives a SARIF 2.1.0 log for code-scanning tools.

### `gate`

```bash
uv run controlloop gate [REPOSITORY] \
  [--evidence-path PATH] [--sarif-path PATH] \
  [--default-branch NAME] [--format text|json|sarif]
```

The CI form. Returns a four-state verdict (see
[Reading the output](#reading-the-output)) and **always** writes an evidence
artifact — including on a failing run, which is exactly when you need it.

- `--evidence-path` — where the artifact goes (default `evidence.json`).
  The parent directory must already exist.
- `--sarif-path` — additionally write a SARIF log, carrying the blast-radius
  statement as `runs[0].properties.controlloop.blastRadius`.
- `--default-branch` — override which branch the merge-base baseline is taken
  from, for CI checkouts where `origin/HEAD` is unset.

`gate` needs real git history for an automatic baseline, so check out with
`fetch-depth: 0`.

## Reading the output

Here is a real passing run against the reference agent:

```
! WORST REACHABLE  financial-action: widgetworks-order-support-agent can invokes
  Issue a refund

✓ PASSED

• material change: not evaluated

! MEDIUM  policy.undeclared-tools: Tool 'Look up an order' is reachable but not
  declared in controlloop.yaml.
    at agent-card.json

Suppressions: none declared
```

Exit code `0`.

**1. The blast-radius line comes first.** `WORST REACHABLE` is the single
worst thing this agent can do *right now* — not what changed. It is the line
you forward to your manager. Here: this agent can take a `financial-action`
by invoking `Issue a refund`. If nothing is classified, it reads
`• no material capability is reachable`.

**2. The verdict.** `gate` returns one of four:

| Verdict | Exit | Meaning |
| --- | --- | --- |
| `✓ PASSED` | 0 | Nothing blocking, nothing warned |
| `! PASSED WITH WARNINGS` | 0 | Findings below the blocking threshold |
| `! REVIEW REQUIRED` | 3 | Analysis was incomplete and enforcement is set to fail |
| `✕ DEPLOYMENT BLOCKED` | 1 | A blocking finding |

A report can additionally show `• NOT EVALUATED`, meaning no policy was
evaluated at all — which is not the same as evaluating and finding nothing.

Note the glyphs: `PASSED WITH WARNINGS` uses `!`, not `✓`. Every state is
distinguishable by symbol and by word, never by colour alone.

**3. Material change** — what changed against the baseline, by class.
`not evaluated` means no comparison was possible, which is different from
"compared and found nothing" (`• no material capability change`). ControlLoop
never reports an unmade comparison as a clean one.

**4. Findings.** Each carries a severity glyph and label (`✕ HIGH`,
`! MEDIUM`), the rule ID, a plain-language message, and the evidence path.
Symbols and words both, never colour alone.

**5. Suppressions** are always listed, never hidden.

### A blocked run

```
! WORST REACHABLE  destructive-action: widgetworks-order-support-agent can
  invokes Bulk refund orders

✕ DEPLOYMENT BLOCKED

✕ HIGH    policy.destructive-or-financial-action: Agent gained new direct
  invokes access to 'Bulk refund orders', classified as destructive-action and
  financial-action.
    at agent-card.json

Completeness enforcement: warn (default)
Evidence: evidence.json
```

Exit code `1`. Two things to notice:

- The blast radius **got worse** — `financial-action` became
  `destructive-action`. The agent's worst case degraded, and that is the
  first thing the report says.
- The finding says **"gained new"**. It is not objecting to refunds existing;
  it is objecting to authority this change *introduces* against the approved
  baseline. In `--format json` that appears as `"change_state": "added"`.

### In JSON

```json
{
  "rule_id": "policy.undeclared-tools",
  "severity": "medium",
  "confidence": "high",
  "message": "Tool 'Look up an order' is reachable but not declared in controlloop.yaml.",
  "change_state": "not-compared",
  "remediation": "Declare this tool in controlloop.yaml ... or remove the agent's access to it if it was not intentional.",
  "documentation_url": "https://docs.controlloop.ai/reference/policy-rules#undeclared-tools"
}
```

Every finding carries its own `remediation` and a stable `documentation_url`.

### Exit codes

| Code | Name | Meaning |
| --- | --- | --- |
| 0 | `PASS` | No blocking finding |
| 1 | `POLICY_FAIL` | A blocking finding |
| 2 | `INVALID_INPUT` | Bad manifest, bad path, unreadable ACBOM |
| 3 | `INCOMPLETE_ANALYSIS` | Analysis incomplete and enforcement set to fail (`REVIEW REQUIRED`) |
| 4 | `ENTITLEMENT` | Reserved |
| 5 | `INTERNAL_ERROR` | A bug — please report it |

## Telling ControlLoop what your tools do

**This is the step people miss.** Discovery reads your agent card and your
OpenAPI spec, but nothing in a repository states that issuing a refund moves
money. Without that, every capability is unclassified and every scan reports
`no material capability is reachable`.

Declare it in `controlloop.yaml`:

```yaml
schema_version: "1.1.0"
owner: support-eng

tool_capabilities:
  - tool: Issue a refund        # the DISCOVERED tool name, not its id
    financial_action: true
    read_only: false
    sensitivity: confidential
  - tool: Purge ticket history
    destructive: true
```

Match the **name discovery found**, which for an A2A skill is its card
`name`, not its `id` — `Issue a refund`, not `issue-refund`. Run
`controlloop bom` and read the node names if you are unsure; a declaration
naming a tool that does not exist is reported, not silently ignored.

Every boolean also accepts `"unknown"`, which the taxonomy ranks **worst of
all**. An unknown risk is never treated as safer than a known one.

## Declaring a ceiling

The ceiling is the part that catches what nobody anticipated:

```yaml
ceilings:
  - agent: widgetworks-order-support-agent
    prohibited_resources:
      - customer-payment-methods
    prohibited_delegation_targets:
      - unapproved-agent
```

Any reachable capability that breaches this blocks the gate — even with every
pattern rule suppressed. A capability-ceiling breach **cannot** be suppressed.

## In CI

```yaml
- uses: actions/checkout@v7.0.1
  with:
    fetch-depth: 0          # gate needs real history for a baseline
- uses: danny-tijerina2/controlloop@<ref>
  with:
    path: ${{ github.workspace }}
```

The action runs `gate`, fails the check on a blocking verdict, uploads SARIF,
and writes a check summary. See
[the GitHub Action reference](../docs/reference/github-action.md). A
[pre-commit hook](../docs/reference/pre-commit-hook.md) is also available.

## What is not built yet

Stated plainly, so you can plan:

- **No published release.** No PyPI package, Homebrew tap, container image, or
  tagged version. Source install only.
- **No Team tier.** Private-repository gates, shared policy bundles, history,
  attestations, dashboards, and billing do not exist.
- **No runtime protection.** ControlLoop analyzes repositories. It does not
  intercept, sandbox, or kill a running agent.
- **Windows is unsupported.**
- **Committed baselines carry no graph**, so `scan`/`gate` report
  `material change: not evaluated` against one. Only the merge-base path
  computes a material change summary.

## Learn more

- [CLI reference](../docs/reference/cli.md) — every command and flag
- [Terminal report format](../docs/reference/terminal-report-format.md)
- [Manifest schema](../docs/reference/manifest-schema.md)
- [Capability ceiling](../docs/reference/capability-ceiling.md)
- [Policy rules](../docs/reference/policy-rules.md) — all ten
- [Material capability taxonomy](../docs/reference/material-capability-taxonomy.md)
- [Privacy, security, and limitations](../docs/reference/privacy-security.md)

---

## For maintainers

The lockfile is authoritative. Do not install unpinned project dependencies
outside `uv`.

```bash
uv sync --locked --all-groups
uv run ruff format --check .
uv run ruff check .
uv run mypy
uv run pytest
uv build
```

`uv run pytest` takes roughly two to three minutes locally, including real
PyInstaller builds and a wall-clock performance benchmark that is sensitive
to machine load.
