Metadata-Version: 2.4
Name: ladon-ml
Version: 0.1.0a2
Summary: Guardian of Distributed Compute: Zombie utilization detection for ML infra.
Author-email: Saptarshi Bhattacharjee <bsaptarshi52@gmail.com>
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/bsaptarshi/ladon-ml
Keywords: ml,gpu,observability,infrastructure
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: System :: Monitoring
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: click>=8.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: nvidia-ml-py>=12.535.133
Requires-Dist: prometheus_client
Requires-Dist: psutil>=5.9
Requires-Dist: kubernetes>=27.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Dynamic: license-file

# Ladon-ML 🐉

[![CI](https://github.com/bsaptarshi/ladon-ml/actions/workflows/ci.yml/badge.svg)](https://github.com/bsaptarshi/ladon-ml/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/ladon-ml.svg)](https://pypi.org/project/ladon-ml/)
[![Python versions](https://img.shields.io/pypi/pyversions/ladon-ml.svg)](https://pypi.org/project/ladon-ml/)
[![License: Apache 2.0](https://img.shields.io/badge/license-Apache%202.0-blue.svg)](LICENSE)

**The Hundred-Headed Guardian of Distributed Compute.**

`ladon-ml` is a lightweight, infrastructure-level Python library designed to identify and reclaim **Zombie Utilization** in high-performance ML clusters. It is purpose-built for teams managing large-scale accelerator fleets (GPUs or specialized silicon like Trainium) where stalled kernels and silent hangs can lead to massive resource waste.

> **Status: Pre-Alpha.** The full Collector → Brain → Reaper pipeline described below is implemented and tested (NVIDIA/Neuron/mock collectors, the sliding-window Brain, local and Kubernetes reapers, the CLI, Prometheus metrics), but it has not yet run against a real production training cluster. `LocalReaper`/`KubernetesReaper` default to dry-run/log-only behavior - see [Configuration](#configuration) before enabling live process/pod termination.

---

## The Problem: Zombie Utilization

In distributed training, a "Zombie" is a node that reports **100% Compute Utilization** and high power draw but provides **0% Functional Throughput**. This usually happens due to:

* **Stalled Kernels:** An infinite loop in a custom op pinning the chip.
* **Collective Deadlocks:** A node waiting for a synchronization signal (Gloo/NCCL) that never arrives.
* **Silent HBM Failures:** Memory errors that stop the training loop but keep the hardware active.



---

## Core Features

* **Zero-Trust Telemetry:** Correlates hardware-level metrics (Power, Thermal, Compute) with software-level progress (Network IO, Memory Delta, Gradient Updates).
* **Inverse Correlation Detection:** Identifies the signature of a zombie: high power/utilization + near-zero network/collective-communication traffic.
* **Modular Reapers:** Pluggable actions to "drain" a node, kill a specific process group, or signal orchestrators (Kubernetes/Slurm).
* **Low Overhead:** Written to run as a sidecar or a lightweight daemon with minimal impact on training latency.

---

## Architecture

Ladon follows a **Collector → Brain → Reaper** pattern:

1.  **Collectors:** Interface with hardware drivers (e.g., `NVML`, `Neuron SDK`) to pull raw telemetry.
2.  **The Brain:** Applies statistical analysis to detect anomalies and "Zombie" patterns over a sliding time window.
3.  **The Reaper:** Executes the reclamation policy (Log, Alert, or Kill).

---

## Quick Start

### Installation

```bash
pip install ladon-ml
```

### CLI

```bash
# One-shot telemetry snapshot (use --mock to try it without real GPU/Neuron hardware)
ladon status --mock healthy

# Run the continuous Collector -> Brain -> Reaper daemon + Prometheus exporter (:9100)
ladon run --mock zombie
```

`ladon status` prints current telemetry per detected device - it's a single snapshot, not a Zombie verdict (the Brain needs a sliding window of samples over time to compute a Z_s score). `ladon run` is the real continuous detection loop.

### Library usage

```python
from ladon.brain import Brain, DeviceStatus
from ladon.collector import detect_collectors
from ladon.reaper.local import LocalReaper

collectors = detect_collectors()  # auto-detects NVIDIA/Neuron hardware present on this host
brain = Brain(window_seconds=30.0, threshold=50.0)
reaper = LocalReaper(dry_run=True)  # stays log-only until you explicitly opt into live kills

telemetry = [chip for collector in collectors for chip in collector.collect()]
for event in brain.ingest(telemetry):
    if event.status == DeviceStatus.ZOMBIE_DETECTED:
        print(f"Zombie detected on device {event.device_id} (Z_s={event.z_score:.2f})")
        reaper.reap(event.active_pids)
```

See [`examples/simulate_cluster.py`](examples/simulate_cluster.py) for a full interactive demo of this detect → reap → recover cycle, using simulated telemetry so it runs without any real hardware.

### Configuration

Ladon can be configured via `ladon_conf.yaml` ([shipped example](ladon_conf.yaml)), with environment variable overrides on top (e.g. `LADON_BRAIN_THRESHOLD=75`):

```yaml
brain:
  window_seconds: 30.0   # sliding window duration, in seconds
  threshold: 50.0        # Z_s score above which a device is flagged as a zombie
  epsilon: 0.001         # smoothing constant to avoid divide-by-zero on idle interconnects
  power_max_w: 400.0     # device TDP used to normalize the power ratio

reaper:
  dry_run: true               # log-only by default; set to false to allow live process termination
  grace_period_seconds: 5.0   # SIGTERM grace period before escalating to SIGKILL
```

## Why "Ladon"?
In Greek mythology, Ladon was the dragon with one hundred heads that guarded the golden apples. He never slept and each head watched a different part of the garden. Ladon-ML serves as that unblinking guardian for your compute garden, ensuring no single chip goes rogue and wastes your "golden" training time.

License
Apache 2.0
