Metadata-Version: 2.4
Name: fingate
Version: 0.2.0
Summary: AI Agent Payment Trust Layer - Deterministic authorization for agent payments
Home-page: https://github.com/arnav7897/GateKeeper---AI-Agent-Trust-Layer
Author: GateKeeper Team
License: MIT
Project-URL: Homepage, https://github.com/arnav7897/GateKeeper---AI-Agent-Trust-Layer
Project-URL: Documentation, https://github.com/arnav7897/GateKeeper---AI-Agent-Trust-Layer/blob/main/README.md
Project-URL: Repository, https://github.com/arnav7897/GateKeeper---AI-Agent-Trust-Layer
Project-URL: Bug Tracker, https://github.com/arnav7897/GateKeeper---AI-Agent-Trust-Layer/issues
Keywords: ai,agents,payments,security,authorization,trust-layer
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Topic :: Security
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic>=2.0.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: click>=8.0.0
Provides-Extra: razorpay
Requires-Dist: razorpay>=2.0.0; extra == "razorpay"
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.5.0; extra == "anthropic"
Provides-Extra: gemini
Requires-Dist: google-generativeai>=0.3.0; extra == "gemini"
Provides-Extra: dev
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0.0; extra == "dev"
Requires-Dist: black>=23.0.0; extra == "dev"
Requires-Dist: mypy>=1.0.0; extra == "dev"
Dynamic: home-page
Dynamic: license-file
Dynamic: requires-python

# GateKeeper — AI Agent Trust Layer

Deterministic, explainable authorization layer that evaluates AI-agent payment
requests **before** money moves. Policy enforcement + explainable risk scoring
+ optional LLM intent analysis → `APPROVE` / `REVIEW` / `BLOCK`, every decision
fully audited.

**Core guarantee:** the payment gateway is contacted only after GateKeeper
returns APPROVE. The LLM can make GateKeeper smarter, but deterministic rules
make it trustworthy — the LLM never authorizes anything.

---

## The problem

Businesses increasingly let AI agents spend money — procurement bots, finance
agents, autonomous subscription renewers. Payment infrastructure still assumes
a human is making the purchase, which breaks in four specific ways:

1. **No authority verification.** Nothing checks *whether the agent is allowed*
   to make this purchase. A bot (or an attacker who has hijacked one) can buy
   anything with the attached card.
2. **No semantic validation.** Nothing compares *what the agent said it was
   buying* with *what is actually being charged*. A prompt injection can dress
   up a gift-card purchase as a software subscription.
3. **No anomaly detection.** A compromised agent can fire rapid, high-value
   payments and no system notices the pattern.
4. **No explainability.** When money moves, nobody can answer *why* — which
   rule allowed it, what risk signals were present, what authorized it.

## The solution

GateKeeper inserts one deterministic decision point between every agent
payment request and the payment gateway:

```text
AI agent asks to spend money
        ↓
GateKeeper verifies authority + intent + risk
        ↓
Decision is explainable (rule-by-rule, feature-by-feature)
        ↓
Only an APPROVED transaction reaches Razorpay/Stripe
        ↓
The payment outcome is audited
```

**How each gap is closed:**

| Gap | Mechanism |
|---|---|
| Authority | Per-agent policy as code: amount caps, daily limits, category allow/block, merchant allow/block, country restrictions, human-approval thresholds — pure deterministic checks |
| Semantics | Optional LLM intent analyzer (Gemini/Anthropic) scores mismatch between declared intent and actual merchant description; flags prompt-injection-style language. **Advisory only** — its output is one of five risk features; it can never approve, block, or override |
| Anomalies | Explainable risk score 0–100 from five weighted features (policy violations 0.4, amount anomaly 0.2, merchant risk 0.15, velocity 0.15, intent deviation 0.1) with per-feature contributions |
| Explainability | Every decision stores decision, risk score, per-check PASS/FAIL, violations, reason codes, and timestamp in a queryable JSONL audit log |

**Fail-closed everywhere:** hard policy violation → BLOCK regardless of risk
score; evaluation error or missing safety data → BLOCK, never a silent APPROVE.

## Features

- **`protect()` wrapper** — wrap any payment function with one line; it becomes the only path to money movement
- **`evaluate_payment_intent()`** — dry-run decisions with zero side effects
- **MCP tool** — native `evaluate_payment` Model Context Protocol tool (`GateKeeperMCPTool`) for Claude-style agents and any MCP runtime; fail-closed handler, Anthropic tool_use-spec schema
- **Policy as code** — `gatekeeper.policy.yaml`; `gatekeeper init` scaffolds it interactively and generates plain-English `POLICY.md`
- **Gateway adapters** — Mock (offline dev), Razorpay (test/production payment links), Stripe; implement `PaymentGatewayAdapter` for anything else
- **Pluggable intent analysis** — Gemini and Anthropic presets behind one `IntentAnalyzer` protocol; or none at all
- **Audit trail** — JSONL backend, one queryable record per decision/payment/webhook; same format in library and cloud mode
- **Framework-agnostic** — works with LangChain, raw MCP, cron bots, or plain Python (see `examples/`)
- **Cloud mode** — FastAPI + React dashboard, SSE live logs, Razorpay test-mode payment links + idempotent signature-verified webhooks, API-key auth, HTTPS
- **100-case eval suite** — precision/recall/FPR/amount-weighted-FN/latency p50/p95, honest labels

## Architecture

```text
AI agent request (any framework)
    ↓
validate PaymentIntent
    ↓
Policy Engine     (deterministic rules — no LLM)
    ↓
Risk Engine       (explainable weighted score, 0–100)
    ↓
Decision Engine   (APPROVE / REVIEW / BLOCK, fail-closed)
    ↓ APPROVE only
Payment Gateway   (Razorpay Test Mode / Stripe / Mock)
    ↓
Webhook → audited final status
```

```mermaid
graph TD
    A[AI Agent] -->|PaymentIntent| B[GateKeeper protect / MCP tool / API]
    B --> C[Policy Engine]
    C --> D[Risk Engine]
    D --> F[Decision Engine]
    F -->|APPROVE| G[Payment Gateway]
    F -->|REVIEW| H[Human approval]
    F -->|BLOCK| I[Blocked - nothing executes]
    F -.->|audit| J[(Audit log)]
    G -.->|webhook| J
```

## Two ways to use it

### 1. `gatekeeper` package (library, zero infrastructure)

```bash
pip install fingate
gatekeeper init    # scaffolds gatekeeper.policy.yaml + POLICY.md
```

**Wrap any payment function:**

```python
from gatekeeper import protect
from gatekeeper.adapters import MockAdapter

def execute_payment(payment_intent, adapter, **kw):
    return adapter.create_payment_link(payment_intent.amount, payment_intent.currency)

safe_payment = protect(execute_payment, policy_config=my_policy, gateway_adapter=MockAdapter())
result = safe_payment(intent)   # result.decision.is_approve() → link created; else nothing ran
```

**Dry-run evaluation:**

```python
from gatekeeper import evaluate_payment_intent
decision = evaluate_payment_intent(intent, my_policy)   # Decision, no execution
```

**Native MCP tool** (Claude-style agents, any MCP runtime):

```python
from gatekeeper import GateKeeperMCPTool, build_claude_tool_spec

tool = GateKeeperMCPTool.from_policy_file("./gatekeeper.policy.yaml")
result = tool.handler({                       # plain dict in, dict out, fail-closed
    "amount": 8000, "category": "developer_tools",
    "declared_intent": "GitHub Team seat",
    "actual_description": "GitHub Team Plan - annual",
    "merchant_name": "GitHub",
})
# result["decision"] == "APPROVE" | "REVIEW" | "BLOCK"
# Anthropic Messages API: tools=[build_claude_tool_spec()]
```

### 2. GateKeeper Cloud (FastAPI + React dashboard)

Multi-agent management UI, PostgreSQL audit trail, Razorpay webhooks. Same
engines underneath — see [Running Cloud](#running-cloud).

---

## Components

| Component | File | Role |
|---|---|---|
| Policy Engine | `gatekeeper/policy_engine.py` | Deterministic checks: amount limit, daily limit, category allow/block, merchant allow/block, country, approval threshold, velocity |
| Risk Engine | `gatekeeper/risk_engine.py` | Explainable 0–100 score, 5 weighted features, per-feature contribution |
| Intent Analyzer | `gatekeeper/intent/` | Optional LLM semantic analysis (Gemini / Anthropic presets). Advisory only — never decides |
| Decision Engine | `gatekeeper/decision_engine.py` | Final authority. Hard policy violation → BLOCK always. Fail-closed on any error |
| Protect wrapper | `gatekeeper/protect.py` | Primary API: wraps any payment function |
| MCP tool | `gatekeeper/mcp_tool.py` | `evaluate_payment` tool, Anthropic tool_use spec compliant |
| Adapters | `gatekeeper/adapters/` | `PaymentGatewayAdapter`: Mock, Razorpay, Stripe |
| Audit | `gatekeeper/audit/jsonl.py` | JSONL audit backend (pluggable) |
| CLI | `gatekeeper/cli.py` | `init` (scaffold policy + POLICY.md), `validate` |

**Decision model:**

```text
hard policy violation        → BLOCK  (always, overrides everything)
critical missing safety data → REVIEW / BLOCK
risk score 0–29              → APPROVE
risk score 30–69             → REVIEW
risk score 70–100            → BLOCK
evaluation error             → BLOCK (fail-closed)
```

**Risk features (weights):** policy violation 0.4, amount anomaly 0.2,
merchant risk 0.15, velocity 0.15, intent deviation 0.1.

## Integration examples (`examples/`)

| Example | What it proves | Run |
|---|---|---|
| `langchain_agent/` | `protect()` composes with LangChain `@tool`; LLM requests, GateKeeper authorizes | `python3 agent.py` |
| `mcp_raw_agent/` | `GateKeeperMCPTool` drives a raw Anthropic tool_use loop | `python3 agent.py --dry` (no API key) or real with `ANTHROPIC_API_KEY` |
| `cron_recurring_bot/` | Autonomous scheduled payments stay within daily limit across restarts | `python3 bot.py` (simulated day) or `bot.py --tick` under cron |

All examples run offline with `MockAdapter`; swap in Razorpay/Stripe adapter for
real payments.

## Quickstart (package)

```bash
pip install -e .
gatekeeper init                       # creates gatekeeper.policy.yaml
gatekeeper validate                   # checks policy config
python3 examples/langchain_agent/agent.py
python3 -m pytest tests/ -q           # 76 tests
```

## Policy configuration (`gatekeeper.policy.yaml`)

```yaml
agent:
  name: procurement-bot
  owner: engineering-team
limits:
  max_transaction_amount: 25000      # paise
  daily_limit: 100000
  require_human_approval_above: 25000
categories:
  allowed: [developer_tools, saas_subscriptions, cloud_services]
  blocked: [gift_cards, cryptocurrency, gambling]
countries:
  allowed: [IN, US]
merchants:
  allowlist_enabled: false
  blocked_merchants: []
```

## Environment variables

### Package (`gatekeeper`) — all optional

| Variable | Needed for | Default |
|---|---|---|
| *(none)* | Core engines, `protect()`, MCP tool, MockAdapter, CLI | — |
| `RAZORPAY_KEY_ID` | RazorpayAdapter (`create_razorpay_adapter()` with no config) | — |
| `RAZORPAY_KEY_SECRET` | RazorpayAdapter | — |
| `RAZORPAY_WEBHOOK_SECRET` | Webhook signature verification | — |
| `RAZORPAY_TEST_MODE` | Defaults to `true` | `true` |
| `STRIPE_API_KEY` | StripeAdapter | — |
| `STRIPE_WEBHOOK_SECRET` | Stripe webhooks | — |
| `GEMINI_API_KEY` | Gemini intent analyzer preset | — |
| `ANTHROPIC_API_KEY` | Anthropic intent preset / MCP raw agent example (live mode) | — |

**Zero keys required** for: policy/risk/decision engines, `protect()`,
`evaluate_payment_intent`, MCP tool, MockAdapter, all three examples
(`--dry` for the MCP agent), full 76-test suite.

### GateKeeper Cloud (backend) — in `.env`

| Variable | Needed for |
|---|---|
| `DATABASE_URL` | PostgreSQL audit trail (`docker-compose` default prewired) |
| `RAZORPAY_KEY_ID` / `RAZORPAY_KEY_SECRET` | Test Mode payment link creation |
| `RAZORPAY_WEBHOOK_SECRET` | Webhook signature verification |
| `GEMINI_API_KEY` | Cloud intent analysis (optional — conservative fallback) |
| `SECRET_KEY` | Backend session signing |
| `GATEKEEPER_API_KEYS` | API-key auth on cloud endpoints (comma-separated; empty = auth off, dev mode) |
| `AUDIT_LOG_PATH` | Package JSONL audit log location (default `./gatekeeper_audit.jsonl`) |
| `REDIS_URL` | Optional caching/rate limiting |
| `FRONTEND_URL`, `CORS_ORIGINS` | Dashboard connectivity |

Never commit real secrets. `.env` is gitignored; copy from `.env.example`.

## Running Cloud

```bash
cp .env.example .env                 # add Razorpay Test Mode keys
./scripts/generate_dev_certs.sh      # self-signed TLS cert for local HTTPS
docker-compose up -d                 # nginx :443, API :8000, frontend :3000, Postgres
docker-compose exec backend python scripts/seed_demo.py
# API via HTTPS: https://localhost/api/health (HTTP 301-redirects to HTTPS)
# API docs: http://localhost:8000/docs
```

**Auth:** set `GATEKEEPER_API_KEYS=key1,key2` to require `X-API-Key: key1` on
all endpoints except `/api/health` and Razorpay webhooks (signature-verified).
SSE clients can pass `?api_key=`. Empty = auth disabled.

Key endpoint:

```bash
curl -X POST http://localhost:8000/api/transactions/evaluate \
  -H "Content-Type: application/json" \
  -d '{"agent_id":"<uuid>","merchant_id":"<uuid>","amount":8000,
       "category":"developer_tools","declared_intent":"GitHub subscription",
       "actual_description":"GitHub Enterprise Plan"}'
```

## Live demo runbook (judges)

End-to-end payment demo in under 5 minutes:

1. **Start**: `docker compose up -d`, then `docker-compose exec backend python scripts/seed_demo.py`.
   Add Razorpay Test Mode keys + `RAZORPAY_WEBHOOK_SECRET` to `.env`.
2. **Webhook tunnel** so Razorpay can reach your machine:
   `cloudflared tunnel --url http://localhost:8000` (or ngrok http 8000).
   In the Razorpay dashboard, set the webhook URL to
   `https://<tunnel-host>/api/webhooks/razorpay` with the same secret.
3. **Open the dashboard** at `http://localhost:3000` → **Try It Live**.
4. **Attack demo**: click *Gift-card attack* → instant `BLOCK` with the failing
   policy check highlighted. Click *Intent mismatch* → semantic deviation raises
   the risk score.
5. **Approve → pay → PAID, live**: run *Normal purchase* → `APPROVE` with a
   Razorpay test payment link → click **Pay now** and complete the test payment
   → the Live Logs feed shows the decision and then the `payment.paid` webhook,
   and the transaction flips to `paid` in real time.
6. **Human-in-the-loop**: run *Limit overshoot* → `REVIEW` → open the
   **Approvals** page → Reject one, Approve another. Approval re-runs the full
   deterministic pipeline (fail-closed: if the re-check no longer approves, a
   409 is returned and no payment link is created).
7. **Library talking point**: same engines run in-process — see the
   `pip install gatekeeper` showcase card at the bottom of **Try It Live**, or
   run `examples/mcp_raw_agent/agent.py --dry` in a terminal.

## Example decisions

Approved:

```json
{"decision": "APPROVE", "risk_score": 13,
 "reason": "Low risk transaction: matching established patterns",
 "risk_breakdown": {"policy_violation": {"contribution": 0.0},
                    "amount_anomaly": {"contribution": 4.9},
                    "merchant_risk": {"contribution": 7.5},
                    "velocity": {"contribution": 0.75},
                    "intent_deviation": {"contribution": 0.0}}}
```

Blocked:

```json
{"decision": "BLOCK", "policy_violations": ["category_blocked"],
 "safety_note": "Only proceed to payment execution if decision == 'APPROVE'."}
```

## Testing

```bash
python3 -m pytest tests/ -q                      # package: 76 tests (policy/risk/decision/MCP)
python3 -m pytest backend/tests/ -q              # backend engines: 30 tests
python3 evals/run_tests.py                       # 100 eval cases + precision/recall/FPR/latency
python3 evals/generate_dataset.py                # regenerate expanded dataset
```

**Eval metrics tracked:** TP/TN/FP/FN, precision, recall, false-positive rate,
amount-weighted false negatives, latency p50/p95. Labels are
deterministic-only (no LLM); attaching a Gemini/Anthropic intent analyzer
raises sensitivity to semantic attacks beyond these numbers.

## Project structure

```text
gatekeeper/            # installable package (zero infra)
├── policy_engine.py  risk_engine.py  decision_engine.py
├── protect.py         # primary API
├── mcp_tool.py        # MCP evaluate_payment tool
├── cli.py             # init / validate
├── models/  adapters/  intent/  audit/  templates/
backend/               # GateKeeper Cloud (FastAPI)
frontend/              # React + TS dashboard
examples/              # LangChain, raw MCP, cron bot
evals/                 # datasets + runner
tests/                 # package test suite (76)
scripts/               # seed / demo / reset
```

## Roadmap

- [x] Package: engines, protect(), adapters, audit, CLI
- [x] MCP tool wrapper (`GateKeeperMCPTool`)
- [x] Integration examples (LangChain, raw MCP, cron bot)
- [x] Eval expansion: 100 cases with honest metrics
- [x] Cloud backend: package audit format, API-key auth, HTTPS in Docker
- [ ] Deploy Cloud, record demo

## Limitations

- Razorpay Test Mode only; no real money movement
- Daily limit is a rolling 24h window via transaction history — no persistent store in package mode
- Dashboard stats/transactions pages fetch the most recent 50 rows; API auth is static keys (no user accounts/roles)
- Local HTTPS uses self-signed certs (swap real certs into `nginx/certs/` for production)

## License

MIT — see LICENSE.
