Metadata-Version: 2.4
Name: polymath-agent
Version: 0.4.1
Summary: Multi-model AI orchestrator built on the nexus brain. Multi-model routing, ReAct pipeline, plan mode, custom slash commands, subagents.
Author: Ayushi Gupta
License: MIT
Project-URL: Homepage, https://github.com/ayushigupta-29/polymath
Project-URL: Repository, https://github.com/ayushigupta-29/polymath
Project-URL: Brain, https://github.com/ayushigupta-29/nexus
Keywords: ai,llm,orchestrator,multi-model,claude,gemini,openai,ollama,agent
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: nexus-md>=0.3.0
Requires-Dist: httpx>=0.27.0
Provides-Extra: orchestrator
Requires-Dist: anthropic>=0.40.0; extra == "orchestrator"
Requires-Dist: google-genai>=1.0.0; extra == "orchestrator"
Requires-Dist: openai>=1.50.0; extra == "orchestrator"
Requires-Dist: rich>=13.0.0; extra == "orchestrator"
Requires-Dist: prompt_toolkit>=3.0.0; extra == "orchestrator"
Requires-Dist: tiktoken>=0.7.0; extra == "orchestrator"
Requires-Dist: sqlite-vec>=0.1.6; extra == "orchestrator"
Provides-Extra: embed
Requires-Dist: google-genai>=1.0.0; extra == "embed"
Requires-Dist: openai>=1.50.0; extra == "embed"
Provides-Extra: dev
Requires-Dist: polymath-agent[orchestrator]; extra == "dev"
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.21; extra == "dev"
Dynamic: license-file

# polymath

A multi-model AI orchestrator built on top of [**nexus**](https://github.com/ayushigupta-29/nexus), the long-term project brain.

Polymath routes tasks across every model available on your device — Claude, Gemini, GPT, Ollama local models, LM Studio — through a single CLI with unified context, an autonomous agentic loop, cross-model review, custom slash commands, subagents, and shared team project folders. The team's knowledge (rules, decisions, glossary, code patterns) lives in nexus and is shared with every AI tool, not just polymath.

## Two-layer architecture

```
┌─────────────────────────────────────────────────────────────┐
│ polymath (this repo) — multi-model orchestrator             │
│   • Routing across Claude / Gemini / GPT / Ollama           │
│   • Agentic ReAct pipeline + plan mode                      │
│   • Slash commands, subagents, cross-model review           │
│   • Ensemble + fanout for concurrent multi-model execution  │
└─────────────────────────────────────────────────────────────┘
                              ↓ depends on
┌─────────────────────────────────────────────────────────────┐
│ nexus — long-term team brain                                │
│   • 9-typed buckets, per-owner sub-files, date stamps       │
│   • Auto-generated CLAUDE.md / AGENTS.md / .cursor / Copilot│
│   • Append-only, schema-versioned, forward-compatible       │
│   • Lives at github.com/ayushigupta-29/nexus                │
└─────────────────────────────────────────────────────────────┘
```

The brain is its own package because **every** AI tool benefits from a shared team brain — not just polymath. You can `pip install nexus-md` and use the `nexus` CLI with Cursor / Claude Code / Codex without ever installing polymath.

---

## Install

**Brain only (just the team memory, works with any AI tool):**
```bash
pip install nexus-md
cd your-repo
nexus init
```

**Full orchestrator (brain + multi-model routing):**
```bash
pip install "polymath-agent[orchestrator]"
# the command is still `polymath`; only the PyPI name differs
# or from source:
git clone https://github.com/ayushigupta-29/polymath
cd polymath
bash install.sh
```

Requires Python 3.10+.

On startup polymath reports any missing dependency and tells you what to
install. It does not install anything itself. Set `POLYMATH_AUTO_INSTALL=1`
to get the old self-repairing behaviour, which is meant for local
development in a virtualenv you own.

---

## Getting API Keys (three ways — pick one)

### Option 1 — CLI auto-detection (zero config)

If you already use any of these CLIs, polymath picks up their credentials automatically at startup. No setup needed.

| CLI | Install | What polymath reads |
|-----|---------|-------------------|
| [Claude Code](https://claude.ai/code) | `npm i -g @anthropic-ai/claude-code` | macOS Keychain / `~/.claude/config.json` (Linux) / `%APPDATA%\Claude\` (Windows) |
| [Gemini CLI](https://github.com/google-gemini/gemini-cli) | `npm i -g @google/gemini-cli` | `~/.gemini/oauth_creds.json` — auto-refreshed via `refresh_token` |
| [Codex CLI](https://github.com/openai/codex) | `npm i -g @openai/codex` | `~/.codex/auth.json` — ChatGPT OAuth mode |

### Option 2 — Environment variables

```bash
export ANTHROPIC_API_KEY=sk-ant-...
export GOOGLE_API_KEY=AIza...
export OPENAI_API_KEY=sk-...
```

### Option 3 — Setup wizard

```bash
polymath --setup
```

Walks you through entering API keys. Saved to `~/.polymath/config.json`.

---

## Platform support

| OS | Claude | Gemini | OpenAI |
|----|--------|--------|--------|
| macOS | Keychain auto-detect ✓ | `~/.gemini/` ✓ | `~/.codex/` ✓ |
| Linux | `~/.claude/config.json` ✓ | `~/.gemini/` ✓ | `~/.codex/` ✓ |
| Windows | `%APPDATA%\Claude\` ✓ | `~/.gemini/` ✓ | `~/.codex/` ✓ |

All providers also fall back to environment variables on every platform.

---

## Start

```bash
polymath                          # interactive REPL
polymath --project <name>         # start with a detached local project
polymath --models                 # list all detected models + availability
polymath --sessions               # list past sessions
polymath --session <id>           # resume a session
polymath --setup                  # re-run setup wizard
polymath ask "question"           # one-shot, no REPL
```

---

## How polymath behaves

polymath is designed to feel like one terminal assistant, not a manual model switcher.

- simple runtime-state questions are answered from local polymath state
- simple chat prompts use a direct answer path
- multi-step work uses the full planning/execution/review pipeline
- if a model fails because of auth, quota, or context pressure, polymath walks to the next viable model
- some internal work can run concurrently, but the user still sees one shared conversation

This means the user-facing experience stays unified even when multiple providers and model roles are involved behind the scenes.

---

## Task Pipeline

Multi-step tasks flow through a structured pipeline. Complexity is classified automatically.

```
SIMPLE   plan → [plan approval] → clarify → execute → review → output

MEDIUM   clarify → plan → [plan approval] → clarify → execute → review → output

COMPLEX  understand → clarify → plan → [plan approval] → clarify
         → execute loop → review → corrections → re-review → output
```

**Plan approval** pauses before execution on every task. The model's plan is shown; type:
- `y` / Enter — proceed
- `refine: <feedback>` — revise the plan (up to 3 rounds)
- `cancel` — abort

The **execute loop** is a full ReAct (Reason / Act / Observe) agentic loop — up to 10 tool-use iterations. Models can read files, list directories, and run shell commands until the task is complete.

On COMPLEX tasks, up to 5 key decisions are automatically extracted from the output and appended to `decisions.md` with a UTC timestamp.

Simple chat and runtime-state questions bypass this full pipeline when that would produce worse UX.

---

## REPL — what you see

```
polymath > write a fastapi server for user auth

⋯ code · complex                          ← processing indicator, immediate
                                           ← prompt stays at bottom while streaming

polymath ⋯ [1↑] > explain last output    ← next command already queued
                                           ← ⋯ = processing, [1↑] = 1 queued

claude-sonnet-3-5: 4.2k↑ 1.1k↓  ·  gemini-2.0-flash: 3.1k↑ 0.4k↓
                                           ← per-model token usage, right-aligned
steps: plan → clarify → execute → review  | primary: Claude Sonnet 3.5
```

**Command queue**: type and submit a new command any time — it queues and runs after the current one finishes.
- Up arrow on an empty prompt pulls the most recent queued command back into the draft for editing
- `/queue` shows queued commands
- `/queue drop <n>` removes one
- `/queue move <from> <to>` reorders them

**Inline interaction modes**: clarification, plan approval, and permission prompts happen in the same REPL surface with `clarify >`, `plan >`, and `permission >` prompts instead of nested modal sessions.

**Persistent state**: the footer shows repo/project provenance, queue preview, and git state while the REPL is running.

**Single-agent behavior**: polymath should feel like one assistant even when multiple provider/model workers are involved internally. Runtime-state questions are answered locally, normal chat uses a direct-answer path, and provider failures are reduced to short orchestration messages instead of raw tracebacks.

---

## Routing

### Force a provider

```
@claude <task>      use Claude
@gemini <task>      use Gemini
@openai <task>      use OpenAI/GPT
@ollama <task>      use local Ollama
```

### Flags

```
--cheap <task>      cost-first routing
--fast  <task>      speed-first routing
--ask <question>    skip pipeline — direct Q&A
--verify            run verify subagent after pipeline
--simplify          run code-simplifier subagent after pipeline
```

### Priority profiles

Set during `/setup`.

| Profile | Order |
|---------|-------|
| `quality-first` | best quality → cheaper fallbacks |
| `cost-first` | free → cheap → premium |
| `speed-first` | fastest available |
| `balanced` | quality + speed averaged |

---

## Supported Models

| Provider | Models |
|----------|--------|
| Anthropic | Claude Sonnet 3.5, Claude Haiku 3.5, Claude Opus 3 |
| Google | Gemini 2.0 Flash, Gemini 1.5 Pro, Gemini 1.5 Flash |
| OpenAI | GPT-4o, GPT-4o Mini, o1 Mini |
| Ollama | Any locally pulled model (auto-detected at runtime) |
| LM Studio | Any loaded model (auto-detected at runtime) |

---

## Built-in Tools

Models use these autonomously during the execute loop:

| Tool | What it does | Requires approval |
|------|-------------|-------------------|
| `list_directory` | List files and folders | No |
| `read_file` | Read any file | No |
| `write_file` | Create or overwrite a file | **Yes** (unless pre-allowed) |
| `run_shell_command` | Run any shell command | **Yes** (unless pre-allowed) |

Pre-allow tools and shell patterns via `.polymath/settings.json` (see [Project Folders](#shared-project-folders)).

---

## Cross-Model Review

After every execution, a **different model** automatically reviews the output — verdict, issues, and suggestions — then corrections are applied and re-reviewed. Claude writes → Gemini reviews → Claude corrects → Gemini re-reviews.

For some phases, polymath can also run internal workers concurrently:

- post-run subagents can execute in parallel
- backup/fallback model attempts are tracked so the orchestrator does not loop
- the visible output still stays on one shared terminal thread

---

## Typed Project Context

9 context files per project. Only the relevant ones are injected per task type — no token bloat.

| File | Purpose | Injected for |
|------|---------|-------------|
| `rules.md` | Golden rules, hard constraints | **Always** |
| `goals.md` | Objectives, success criteria | **Always** |
| `code.md` | Patterns, conventions, architecture | `code` tasks |
| `stack.md` | Tech stack, dependencies, versions | `code` tasks |
| `logic.md` | Business / domain logic | `analysis`, `research` |
| `data.md` | Schemas, field definitions, samples | `analysis` tasks |
| `glossary.md` | Domain-specific terms | `analysis`, `research` |
| `decisions.md` | Past decisions + rationale (ADR-style) | `planning` |
| `personas.md` | Tone, behavior, communication style | `creative` tasks |

### `@context` mentions

```
polymath > @rules what are the constraints on this project?
polymath > @stack which framework handles auth?
```

---

## Chunked Memory

Polymath maintains a **chunked, embedded, team-shareable memory** for each project — semantic retrieval over your typed context plus durable facts auto-extracted from sessions.

### How it differs from typed context alone

The 9 typed context files are still the **source of truth** (and what humans edit). Memory is the **derived, retrieval-friendly layer** built on top:

- Each `.md` file is split into heading-aware chunks
- Each chunk is embedded (Gemini `text-embedding-004` by default, OpenAI `text-embedding-3-small` fallback — Claude has no embeddings API)
- At injection time, only the **top-K chunks relevant to your prompt** are sent to the model — not the whole file
- After every pipeline run, a **writeback** step extracts up to 8 durable facts from the conversation as new candidate chunks
- `confirmed` chunks rank higher than `candidate` chunks; teammates promote candidates via `/memory confirm <id>`

### Storage layout (committed to git)

```
.polymath/memory/
├── chunks/
│   ├── rules.jsonl          ← one chunk per line, per type
│   ├── code.jsonl
│   ├── decisions.jsonl
│   ├── extracted.jsonl      ← writeback output lands here
│   └── ... (one file per typed context concern)
├── manifest.json            ← embed_provider/model/dim, last writeback
└── .gitignore               ← ignores cache.db
```

Each chunk: `id`, `type`, `source_file`, `heading_path`, `content`, base64-encoded float32 `embedding`, `embed_provider`/`model`/`dim`, `created_by_model`/`user`, `created_at`, `last_validated_at`, `scope` (session/project/global), `status` (candidate/confirmed).

A local SQLite + sqlite-vec cache lives at `~/.polymath/cache/<project_id>/chunks.db` and is **derived** from the JSONL files. New devs auto-build it on first run; embeddings are committed so re-embed cost is zero.

### `/memory` commands

| Command | What it does |
|---------|-------------|
| `/memory show` | Counts by type + status, embedder info, last writeback |
| `/memory rebuild` | Re-chunk + re-embed all context files for the active project |
| `/memory search <query>` | Debug what the retriever returns for a query |
| `/memory confirm <id>` | Promote a `candidate` chunk → `confirmed` (boosts ranking) |
| `/memory diff` | Preview what writeback would extract from the current session |
| `/memory test` | Verify your embedder config end-to-end |

### Setup

Set one embedding key (Claude has no embeddings API, so this is required):

```bash
# Option A — Gemini (free tier, recommended)
export GOOGLE_API_KEY=AIza...   # https://aistudio.google.com/apikey

# Option B — OpenAI ($0.02 / 1M tokens, paid)
export OPENAI_API_KEY=sk-...
```

Codex CLI tokens are chat-only and **cannot embed** — set a real `OPENAI_API_KEY` if going that route.

On first launch in a project with existing context, polymath asks once:
```
Found 9 context files (~340 chunks) not yet indexed for memory.
Migrate to chunked memory now? Free Gemini embedding tier. (y / n)
```

### Curation workflow (team)

1. Daily use — writeback auto-creates `candidate` chunks in `.polymath/memory/chunks/extracted.jsonl`
2. Reviewer (you, on PR) sees the diff — line-by-line in git
3. `/memory confirm <id>` promotes good candidates to `confirmed`; bad ones get pruned automatically after 30 days

### Settings (`~/.polymath/config.json` or `.polymath/settings.json`)

```json
{
  "memory": {
    "embedding_provider": "gemini",
    "embedding_model": "text-embedding-004",
    "retrieval_k": 12,
    "dedup_threshold": 0.92,
    "writeback_min_chars": 200,
    "stale_candidate_days": 30
  }
}
```

---

## Shared Project Folders

`.polymath/` is a directory you commit to git — shared across your whole team.

```
.polymath/
├── context/          ← 9 typed context .md files
├── commands/         ← custom slash command .md files
├── subagents/        ← custom subagent configs (.md or .json)
├── handoffs/         ← session briefings from /handoff
└── settings.json     ← allowed tools, allowed shell patterns, model preferences
```

polymath walks up from `cwd` to find `.polymath/`, like git.

### Repo-first project model

polymath is repo-first:

- If the current git repo contains `.polymath/`, that repo is treated as the primary shared project
- Shared context, commands, subagents, and settings live in that repo and are meant to be committed
- Sessions, history, caches, and exports remain local under `~/.polymath/`
- Detached local projects under `~/.polymath/projects/` still work, but they are secondary to repo-backed projects

Preferred workflow:

```bash
cd your-app-repo
polymath /project init
git add .polymath
git commit -m "Add polymath project metadata"
```

### Working as a team

For concurrent human collaboration, keep the repo as the shared source of truth and use git normally:

- commit `.polymath/` so context, commands, subagents, and handoffs are shared
- keep session transcripts and caches local under `~/.polymath/`
- use separate branches or worktrees for parallel implementation work
- use `/handoff` to drop a brief into `.polymath/handoffs/` when another person needs to continue the thread
- avoid editing the same code paths from multiple terminals unless you intend to resolve the merge at git level

Recommended team model:

- `.polymath/` is committed and shared
- `~/.polymath/` remains private and local
- coordination happens through git, not a shared backend
- handoffs are explicit, reviewable files in the repo

### `settings.json`

```json
{
  "allowed_tools": ["list_directory", "read_file", "write_file"],
  "allowed_shell_commands": ["git *", "pytest *", "npm *"],
  "model_preferences": {
    "primary": "claude",
    "reviewer": "gemini"
  }
}
```

---

## Slash Commands

| Command | What it does |
|---------|-------------|
| `/ship` | AI-writes commit message → `git add -A && git commit` → `git push` → `gh pr create` |
| `/learn [text]` | Save a lesson to `decisions.md`. No text = AI extracts key takeaway from last output |
| `/handoff` | AI writes a structured briefing (Context / Current state / Next steps / Watch out for), saved to `.polymath/handoffs/` |
| `/simplify` | Simplify and tighten last output |
| `/explain` | Explain last output in plain language |

### Custom slash commands

Drop a `.md` file in `.polymath/commands/` (project) or `~/.polymath/commands/` (global):

```markdown
---
name: deploy
description: Deploy to staging
shell_before: npm run build
shell_after: echo "{{ai_output}}" | pbcopy
---
Deploy the following changes. Session context: {{session_output}}
```

---

## Subagents

Post-pipeline focused passes on the output.

| Subagent | Trigger | What it does |
|----------|---------|-------------|
| `code-simplifier` | `--simplify` | Removes complexity, improves naming |
| `verify` | `--verify` | Runs tests, fixes failures — up to 3 iterations |
| `security-scan` | security tasks | Flags injections, hardcoded creds |
| `test-writer` | code tasks | Generates test cases |
| `docs-writer` | code tasks | Generates inline documentation |
| `brainstorm-critic` | planning tasks | Challenges assumptions |
| `perf-reviewer` | code tasks | Spots performance hotspots |
| `diff-explainer` | code tasks | Plain-language diff explanation |

Custom subagents: `.md` files in `.polymath/subagents/` or `~/.polymath/subagents/`.

---

## Auth & Accounts

```
/accounts                          show all authenticated accounts per provider
/accounts switch gemini <email>    switch active Gemini account
/accounts switch claude            instructions to switch Claude account
/accounts switch openai            instructions to switch OpenAI/Codex account
```

When a token expires mid-session, polymath catches the 401, prints the exact re-login command, and re-probes automatically. For Gemini, a silent token refresh is attempted first via `refresh_token` — no interruption for routine expiry.

---

## Session Management

```
/sessions             list all past sessions
/session <id>         resume a session
/session delete <id>  delete a session
/new                  start a new session
```

All sessions stored locally in `~/.polymath/context.db` (SQLite). Exported to Markdown in `~/.polymath/sessions/`. Nothing sent anywhere beyond the model API calls.

Use `/state` in the REPL to inspect:
- active session
- active project source (`repo` vs `local`)
- git repo / branch / dirty state
- loaded context files
- command and subagent provenance
- allowed tools from repo settings

---

## Parallel Sessions

Run polymath in multiple terminal tabs — one per workstream. OS notifications fire on pipeline completion so you can switch tabs without watching the screen.

```
/parallel       show how many other polymath sessions are running
```

---

## All REPL Commands

```
/models               list detected models
/state                show active session, project, repo, and permissions state
/queue                list queued commands
/queue drop <n>       remove queued command
/queue move <a> <b>   reorder queued commands
/sessions             list past sessions
/session <id>         resume a session
/session delete <id>  delete a session
/new                  new session
/project new <name>   create a new project
/project use <name>   switch to a project
/project list         list all projects
/project init         create .polymath/ in current directory
/context show <type>  display a context file
/context add <t> <e>  append entry to context
/context edit <type>  open in $EDITOR
/ship                 commit + push + PR
/learn [text]         save a lesson to decisions.md
/handoff              write a teammate handoff brief
/simplify             simplify last output
/explain              explain last output
/<command>            run any custom slash command
/commands             list all slash commands
/subagents            list all subagents
/accounts             list authenticated accounts
/accounts switch ...  switch account (see above)
/parallel             show other active polymath sessions
/memory               show chunk counts, embedder, last writeback
/memory rebuild       re-chunk + re-embed all context
/memory search <q>    debug retrieval
/memory confirm <id>  promote candidate chunk → confirmed
/memory diff          preview writeback for current session
/memory test          verify embedder config
/setup                re-run setup wizard
/exit                 exit
```

---

## Global Config (`~/.polymath/`)

```
~/.polymath/
├── config.json               # settings, API keys, priority profile
├── history.txt               # REPL command history
├── context.db                # SQLite: all sessions + messages
├── sessions/                 # Markdown session exports
├── commands/                 # global custom slash commands
├── subagents/                # global custom subagents
├── cache/                       # local-only, derived from .polymath/memory/
│   └── <project_id>/
│       └── chunks.db            # sqlite + sqlite-vec retrieval index
└── projects/
    └── <project-name>/
        ├── context/
        │   ├── rules.md
        │   ├── logic.md
        │   ├── code.md
        │   ├── stack.md
        │   ├── data.md
        │   ├── goals.md
        │   ├── decisions.md
        │   ├── glossary.md
        │   └── personas.md
        └── memory/              # detached fallback when no .polymath/ exists
            ├── chunks/*.jsonl
            └── manifest.json
```

---

## Architecture

```
polymath/
├── main.py                startup wiring + REPL bootstrap
├── orchestrator/
│   ├── run_controller.py  request orchestration, mode selection, shared agent flow
│   ├── attempt_ledger.py  model-attempt tracking to prevent fallback loops
│   ├── output_policy.py   concise user-facing orchestration messages
│   ├── state_responder.py answer runtime/state questions from local polymath state
│   ├── race.py            with_failover / race_first_success / gather_all helpers
│   └── worker_pool.py     manages concurrent model/subagent workers
├── memory/
│   ├── store.py           per-type JSONL chunks + sqlite-vec cache
│   ├── embedder.py        Gemini primary, OpenAI fallback
│   ├── chunker.py         heading-aware markdown splitter
│   ├── retriever.py       top-K hybrid retrieval, drop-in for build_context_injection
│   ├── writer.py          end-of-turn extraction → candidate chunks
│   ├── migrate.py         one-shot bootstrap from existing .md context
│   └── sync.py            re-index a single source after edit
├── command_service.py     built-in + custom REPL command handling
├── execution_service.py   ask/pipeline execution on top of orchestrator policy
├── command_registry.py    REPL completion + command registry helpers
├── ui_state.py            footer, queue, and status rendering helpers
├── domain.py              Project, SessionState, ModelSelection, RunContext
├── pipeline.py            task pipeline + ReAct agentic loop
├── router.py              task classification heuristics
├── model_policy.py        model selection, adapter creation, fallback policy
├── permissions.py         permission parsing + repo allow-list checks
├── detector.py            model auto-detection, token refresh, account management
├── project_runtime.py     repo-first project resolution + git provenance
├── config.py              MODEL_REGISTRY, ModelInfo, TaskType, CostTier
├── tools.py               tool registry: list_dir, read_file, write_file, run_shell
├── context_store.py       SQLite session storage, message log, cost tracking
├── context_manager.py     typed project context: read/write/inject
├── project_config.py      .polymath/ discovery, settings, tool allow-lists
├── slash_commands.py      slash command registry + /ship, /learn, /handoff
├── subagents.py           subagent registry, verify loop, auto-selection
├── compressor.py          context window management + summarization
├── workspace.py           project structure scan
├── setup_wizard.py        first-run setup + model configuration
├── bootstrap.py           report missing dependencies on startup
└── adapters/
    ├── base.py            BaseAdapter, Message, ToolCall, AuthExpiredError
    ├── claude.py          Anthropic async adapter
    ├── gemini.py          Google Gemini async adapter (OAuth + API key)
    ├── openai_adapter.py  OpenAI-compatible async adapter (OpenAI + LM Studio)
    └── ollama.py          Ollama local async adapter
```

### Orchestrator responsibilities

- `run_controller.py` decides whether a request is local state, direct chat, or pipeline work
- `attempt_ledger.py` records which models already failed for a phase so fallback does not loop
- `output_policy.py` turns provider failures into short user-facing orchestration messages
- `worker_pool.py` manages concurrent internal workers without changing the single-threaded terminal UX
