Metadata-Version: 2.4
Name: slpilot
Version: 0.1.0
Summary: A dependency-free Slurm parameter sweep and artifact engine
License: MIT
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# slpilot

Temporary AI-generated README below. 
It uses the old working name s-pilot.

Will replace with 100% human-written soon :)
---

`s-pilot` is a Python 3.9+ Slurm parameter-sweep runner and artifact collector.
It has no runtime dependencies: planning, task execution, JSON manifests, CSV
merging, gzip output, and artifact hashing all use the Python standard library.

Install the future published package with `pip install s-pilot`, then submit a
declarative configuration with:

```bash
s-pilot submit path/to/sweep.py
```

For local development in this checkout:

```bash
PYTHONPATH=src python -m s_pilot submit examples/sweep.py
```

The example defaults to dry-run mode, so it does not call Slurm.

## Training-script contract

A compatible training script must expose an `argparse`-style command-line
interface. Each `s-pilot` invocation passes one value for every configured sweep
flag, one seed, and one output directory. If data staging is configured, it also
passes the staged data directory. The script must execute one condition/seed
combination, return a nonzero exit code on failure, and write declared artifacts
below the supplied output directory.

`s-pilot` does not import PyTorch, pandas, safetensors, or any other model
library. Checkpoints such as `.safetensors` are declared as `opaque` artifacts:
they are copied into the published run tree and recorded with a SHA-256 hash,
but their contents are not decoded.

## Config shape

Every config is a Python module containing uppercase TOML-like constants. It is
declarative: `s-pilot` does not accept callbacks, shell snippets, or command
templates. Relative filesystem paths are resolved relative to the config file.

```python
EXPERIMENT_NAME = "example"
ENTRYPOINT = "train.py"
SEEDS = range(4)
PARTITION = "normal"
TIME = "00:10:00"
MEMORY = "1G"
WORKERS_PER_TASK = 2
CPUS_PER_WORKER = 1
MAX_ACTIVE_TASKS = 2
RESULTS_ROOT = "../results"
WIDTH = (8, 16)  # sequences create sweep axes; scalars are fixed arguments
CSV_ARTIFACTS = ("metrics.csv",)
OPAQUE_ARTIFACTS = ("checkpoint.bin",)
```

Each sweep value is rendered as `FLAG VALUE`, so the training script owns type
parsing. `seed_count` generates seeds `0` through `seed_count - 1`. Configured
sweep order determines Cartesian-product order; seeds are the inner axis.

## Run layout and guarantees

The submitter holds the compute array, writes `submission.json` and
`task-plan.json`, submits an `afterany` collector, and only then releases the
array. Every task has a generic `task-######` identity and writes to an isolated
staging directory.

The collector writes `manifest.json` for complete, partial, failed, and aborted
runs. It checks one status per planned task and required artifact presence. CSV
artifacts are merged per non-seed condition with `s_pilot_task_id`,
`s_pilot_condition_id`, `s_pilot_seed`, and `s_pilot_<axis>` columns added. Opaque
artifacts are published unchanged per condition/task. Final published files and
log archives are atomically replaced; staging is removed only after a complete
run.

## Deliberately reserved for later

The task-plan and manifest schemas reserve source hashes/execution entrypoints,
storage mode, and modular generated wrappers. This lets future releases add
code snapshotting, scratch-mode runs, cluster profiles from `s-pilot init`, GPU
binding with `SLURM_LOCALID`/`CUDA_VISIBLE_DEVICES` or `--gpu-bind`, and
manifest query commands without changing experiment identity. Those features
are not implemented in this release.
