Metadata-Version: 2.5
Name: citar
Version: 0.1.4
Summary: Civ Inspired Tool for AI Research - a Civilization V-style 4X game for benchmarking language models
Project-URL: Homepage, https://github.com/jprodgers/CITAR
Project-URL: Documentation, https://github.com/jprodgers/CITAR/wiki
Project-URL: Source, https://github.com/jprodgers/CITAR
Project-URL: Issues, https://github.com/jprodgers/CITAR/issues
Project-URL: Changelog, https://github.com/jprodgers/CITAR/blob/main/CHANGELOG.md
Author: Jimmie Rodgers
License-Expression: MPL-2.0
License-File: LICENSE
License-File: NOTICE.md
Keywords: 4x,agents,ai-research,benchmark,civilization,evaluation,game,llm,mcp,strategy
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Web Environment
Classifier: Framework :: FastAPI
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Mozilla Public License 2.0 (MPL 2.0)
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Games/Entertainment :: Turn Based Strategy
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.11
Requires-Dist: alembic>=1.13
Requires-Dist: argon2-cffi>=23.1
Requires-Dist: email-validator>=2.1
Requires-Dist: fastapi>=0.115
Requires-Dist: httpx>=0.27
Requires-Dist: itsdangerous>=2.2
Requires-Dist: python-multipart>=0.0.9
Requires-Dist: sqlalchemy>=2.0
Requires-Dist: tzlocal>=5.2
Requires-Dist: uvicorn[standard]>=0.30
Provides-Extra: all
Requires-Dist: anthropic>=1.0; extra == 'all'
Requires-Dist: authlib>=1.3; extra == 'all'
Requires-Dist: cryptography>=42.0; extra == 'all'
Requires-Dist: keyring>=25.0; extra == 'all'
Requires-Dist: mcp>=2.0; extra == 'all'
Requires-Dist: openai>=1.40; extra == 'all'
Requires-Dist: websockets>=12.0; extra == 'all'
Provides-Extra: anthropic
Requires-Dist: anthropic>=1.0; extra == 'anthropic'
Provides-Extra: dev
Requires-Dist: anthropic>=1.0; extra == 'dev'
Requires-Dist: authlib>=1.3; extra == 'dev'
Requires-Dist: build>=1.2; extra == 'dev'
Requires-Dist: cryptography>=42.0; extra == 'dev'
Requires-Dist: keyring>=25.0; extra == 'dev'
Requires-Dist: mcp>=2.0; extra == 'dev'
Requires-Dist: mkdocs-material>=9.5; extra == 'dev'
Requires-Dist: mkdocstrings[python]>=0.25; extra == 'dev'
Requires-Dist: openai>=1.40; extra == 'dev'
Requires-Dist: pip-audit>=2.7; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Requires-Dist: twine>=5.1; extra == 'dev'
Requires-Dist: websockets>=12.0; extra == 'dev'
Provides-Extra: keyring
Requires-Dist: cryptography>=42.0; extra == 'keyring'
Requires-Dist: keyring>=25.0; extra == 'keyring'
Provides-Extra: mcp
Requires-Dist: mcp>=2.0; extra == 'mcp'
Provides-Extra: oauth
Requires-Dist: authlib>=1.3; extra == 'oauth'
Provides-Extra: openai
Requires-Dist: openai>=1.40; extra == 'openai'
Provides-Extra: postgres
Requires-Dist: psycopg[binary]>=3.1; extra == 'postgres'
Provides-Extra: server
Requires-Dist: anthropic>=1.0; extra == 'server'
Requires-Dist: authlib>=1.3; extra == 'server'
Requires-Dist: cryptography>=42.0; extra == 'server'
Requires-Dist: keyring>=25.0; extra == 'server'
Requires-Dist: websockets>=12.0; extra == 'server'
Provides-Extra: worker
Requires-Dist: websockets>=12.0; extra == 'worker'
Description-Content-Type: text/markdown

<div align="center">

<img src="docs/assets/icon.png" alt="" width="96">

# CITAR

**Civ Inspired Tool for AI Research**

A Civilization V-style 4X game whose players can be language models.
Play against them, watch them play each other, and measure how well they do it.

[![tests](https://github.com/jprodgers/CITAR/actions/workflows/test.yml/badge.svg)](https://github.com/jprodgers/CITAR/actions/workflows/test.yml)
[![PyPI](https://img.shields.io/pypi/v/citar.svg)](https://pypi.org/project/citar/)
[![Python](https://img.shields.io/pypi/pyversions/citar.svg)](https://pypi.org/project/citar/)
[![Licence: MPL-2.0](https://img.shields.io/badge/licence-MPL--2.0-blue.svg)](LICENSE)

[Install](#install) · [Quick start](docs/QUICKSTART.md) · [Documentation](docs/) · [Run a server](docs/server/DEPLOY.md)

</div>

---

## What it is

A full Civilization V ruleset — nine eras, 35 civilizations, 40 city-states, religion, social
policies, great people, espionage, the United Nations, nuclear weapons — with an interface built so
that a language model can sit in any seat.

Humans play in the browser. Models join through **MCP** (Claude Code, Claude Desktop, any MCP
client), through the **Anthropic API**, or through any **OpenAI-compatible endpoint** — LM Studio,
Ollama, llama.cpp, vLLM. A scripted bot is included to play against and to measure models against,
so the game is fully playable with no model at all.

The rules and numbers come from [UnCiv](https://github.com/yairm210/Unciv)'s "Civ V – Gods & Kings"
ruleset, which is why they are faithful enough to be worth measuring against.

## Why it exists

Most model evaluations are short. A question, an answer, a score. A game of Civilization is the
opposite: hundreds of turns, imperfect information, an opponent who reacts, and consequences that
arrive forty turns after the decision that caused them. That is a different thing to be good at,
and it is hard to fake.

So CITAR is built as an instrument, not just a game:

- **Benchmarks** run full games across several models on identical maps and score them against the
  scripted bot, with confidence intervals and per-turn timing.
- **Scenario probes** put a model in a prepared position — an offer to accept, a war to decide on —
  and record what it does, one decision at a time, repeatably.
- **Metrics** record every tool call, every rejected order, every loop, and how each turn ended.
- **Reports** price it all: hardware time, electricity, tokens, and what a given result cost to
  produce.

None of it phones home. CITAR has no telemetry and no accounts unless you run a server on purpose.

## Install

### Just want to play

**Windows** — download [`CITAR-setup.exe`](https://github.com/jprodgers/CITAR/releases/latest) and
run it. Nothing else needed; there is no Python to install.

**macOS and Linux**

```bash
curl -fsSL https://raw.githubusercontent.com/jprodgers/CITAR/main/install.sh | bash
```

**Anywhere with Python 3.11+**

```bash
pipx install "citar[all]"     # or: pip install "citar[all]"
citar setup                   # finds your models and configures CITAR
citar                         # starts the game and opens your browser
```

Also on [Homebrew, Scoop and winget](docs/INSTALL.md).

### Want to run a server

```bash
curl -fsSL https://raw.githubusercontent.com/jprodgers/CITAR/main/install.sh | bash -s -- --server
```

or with Docker:

```bash
curl -O https://raw.githubusercontent.com/jprodgers/CITAR/main/docker-compose.yml
curl -o .env https://raw.githubusercontent.com/jprodgers/CITAR/main/.env.docker.example
$EDITOR .env                  # domain, secret key
docker compose up -d
```

Either way you get accounts, invitations, single sign-on, per-user budgets, TLS, and a way for
people to lend their own GPUs to the server without opening a port at home. See
[docs/server/DEPLOY.md](docs/server/DEPLOY.md).

## First game in five minutes

```bash
citar                         # opens http://127.0.0.1:8765
```

Create a game, give one seat to yourself and the rest to **scripted bots**, and press Create. That
works with nothing installed.

To give a seat to a model, pick **LLM** and choose a server and model — `citar setup` will have
found LM Studio or Ollama if either is running. To hand a seat to Claude Code instead, choose **MCP
client**, then copy the `claude mcp add …` command from **Join / Seats**.

[The full quick start](docs/QUICKSTART.md) walks through all three.

## What you can do with it

| | |
|---|---|
| **Play** | Full Civ V rules in the browser: tech tree, policies, religion, espionage, city-states, diplomacy with a deal builder, five victory conditions. |
| **Watch** | Games with no human seat can be watched with full vision, paused, slowed down, and viewed as any civilization — including each AI's recorded reasoning. |
| **Benchmark** | Suites of models against the same seeded maps, run sequentially or in parallel, resumable across restarts and reboots. |
| **Probe** | Prepared scenarios replayed case by case: offers, messages, or a whole turn, with expected outcomes and pass rates. |
| **Measure** | Per-seat turn times, tool mix, error and loop rates, token counts, and how turns ended. Exportable as CSV. |
| **Cost** | A usage ledger priced at report time, so correcting an electricity rate corrects every report ever made. |
| **Edit** | A map editor and a scenario editor, so you can build the position you want to test. |
| **Tune** | A parallel bot simulator and a resumable experiment lab, because the bot is the yardstick. |

## Documentation

| | |
|---|---|
| [Quick start](docs/QUICKSTART.md) | First game, first model, first benchmark |
| [Installing](docs/INSTALL.md) | Every install route, and how to remove it |
| [Playing](docs/PLAYING.md) | The browser client, keyboard shortcuts, every screen |
| [AI players](docs/AI_PLAYERS.md) | LLM seats, MCP, providers, prompts, guard rails |
| [Benchmarks](docs/BENCHMARKS.md) | Suites, scheduling, scoring, what the numbers mean |
| [Scenarios and probes](docs/SCENARIOS.md) | The map editor, scenarios, and repeatable decision tests |
| [Servers, costs and reports](docs/REPORTS.md) | The machine registry, the usage ledger, costed reports |
| [Scripted bots](docs/BOTS.md) | How the bot plays, balancing it, the experiment lab |
| [Configuration](docs/CONFIGURATION.md) | Every environment variable, and where files live |
| [Running a server](docs/server/DEPLOY.md) | Domain, TLS, accounts, workers, backups, the runbook |
| [Architecture](docs/ARCHITECTURE.md) | How the pieces fit, for contributors |
| [Modding](docs/MODDING.md) | Adding rules, units and mechanics |
| [HTTP and tool API](docs/API.md) | Driving CITAR from your own code |
| [Troubleshooting](docs/TROUBLESHOOTING.md) | When something does not work |
| [FAQ](docs/FAQ.md) | Including the honest answers about model strength |

The same pages are on the [wiki](https://github.com/jprodgers/CITAR/wiki) and at
[jprodgers.github.io/CITAR](https://jprodgers.github.io/CITAR).

## How good are the models, actually?

Not very, yet — and that is the interesting part.

A small local model can found cities, research sensibly and hold a conversation about a trade, and
will still lose to a scripted bot that has no idea what it is doing beyond a few hundred lines of
heuristics. Larger models play better and cost more per turn. CITAR exists to put numbers on that
rather than anecdotes, which means the measurement has to be honest about its own limits: see
[FAQ](docs/FAQ.md#how-strong-are-the-models) and
[KNOWN_ISSUES.md](KNOWN_ISSUES.md).

## Contributing

Bug reports, ideas and pull requests are all welcome — see [CONTRIBUTING.md](CONTRIBUTING.md). The
short version:

```bash
git clone https://github.com/jprodgers/CITAR && cd CITAR
pip install -e ".[dev]"
python -m unittest discover -s tests     # 272 tests, about 90 seconds
citar serve --debug
```

Most content is data: the ruleset is UnCiv-style JSON, and a new unit or building is a JSON entry,
not code. [docs/MODDING.md](docs/MODDING.md) covers that;
[docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) covers the rest.

## Licence and credits

CITAR is licensed under the [Mozilla Public License 2.0](LICENSE).

Its rules, numbers and much of its game logic are derived from
**[UnCiv](https://github.com/yairm210/Unciv)** by Yair Morgenstern and contributors, also MPL-2.0.
No UnCiv graphics, sounds or flavour text are included. Full attribution is in
[NOTICE.md](NOTICE.md).

*Sid Meier's Civilization* is a trademark of Take-Two Interactive. CITAR is an independent project,
not affiliated with or endorsed by Take-Two, 2K, Firaxis Games or the UnCiv project, and contains
no assets from any commercial Civilization title.
