Metadata-Version: 2.5
Name: modal-cuda
Version: 0.2.0
Summary: CLI that compiles and runs CUDA C programs on Modal GPUs
Author: ExpressGradient
License: MIT License
        
        Copyright (c) 2025 ExpressGradient
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Requires-Python: >=3.12
Requires-Dist: modal>=1.2.2
Requires-Dist: pathspec>=0.12.1
Description-Content-Type: text/markdown

# mcc

Compile and run CUDA C/C++ programs on a [Modal](https://modal.com) GPU.
Requires Python 3.12+ and a Modal account with GPU access.

```bash
uv tool install modal-cuda
uvx modal token new
mcc sample.cu --gpu T4
```

For a one-off run: `uvx --from modal-cuda mcc sample.cu`.

## Usage

```bash
mcc input.cu [options] [-- program arguments]

mcc sample.cu --gpu A100 --nvcc-arg=-O3 -- --size 4096
```

Everything after `--` is passed unchanged to the executable. Commands run without
a shell. Standard input is closed; interactive programs are not supported.

| Option | Meaning |
| --- | --- |
| `--gpu GPU` | GPU type; defaults to `T4`. See `mcc --help` for choices. |
| `--image IMAGE` | CUDA development image; defaults to `nvidia/cuda:12.4.1-devel-ubuntu22.04`. |
| `--app NAME` | Modal app name; defaults to the input filename. |
| `--timeout SECONDS` | Remote job limit, including compilation and downloads; defaults to 600. Range: 10–86400. Image builds and GPU scheduling are outside this limit. |
| `--nvcc-arg=FLAG` | Extra compiler argument; repeat for multiple arguments. Use `=` for flags starting with `-`. |
| `--project DIR` | Upload a project and use its root as the remote working directory. |
| `--source FILE` | Additional C/C++ or CUDA source to compile; repeat as needed. Requires `--project`. |
| `--download PATH` | Download an exact file path relative to the remote working directory; repeat as needed. |
| `--output-dir DIR` | Local download directory; defaults to `mcc-output`. |
| `--version` | Print the installed version. |

Choose an image and `--nvcc-arg` architecture flags compatible with your GPU.
The CLI does not select a toolkit or architecture automatically.

## Projects and results

```bash
mcc src/main.cu --project . --source src/helper.cu \
  --nvcc-arg=-Iinclude --nvcc-arg=-O3 \
  --download results/output.bin --output-dir runs/first \
  -- --input data/input.bin
```

Local source paths are relative to your shell's working directory. All sources
must be inside `--project`. The project keeps its directory layout, so headers
and data files are available remotely. Without `--project`, only the input file
is uploaded and the executable runs beside that file.

Uploads respect the project root's `.gitignore`, then `.mccignore`, using Git-style
patterns. Nested ignore files are not read. `.git`, `.venv`, Python/test/lint
caches, symlinks, and the selected output directory are skipped. An explicitly
selected source excluded by these rules is an error. Add private files and large
unused data to `.mccignore` before uploading a project.

Downloads accept regular files only, with no globs, directories, symlinks, absolute
paths, or `..`. Parent directories are created locally. Existing destination files
are never overwritten. Transfers are staged; incomplete transfers are discarded.
Completed files are published individually once the remote stream finishes.
Files can still be downloaded after a nonzero program exit, which preserves that
exit code. Compile failures, timeouts, and cancellation skip downloads.

## Output and failures

Compiler and program output stream as bytes arrive, preserving stdout and stderr.
CLI status messages go to stderr, so redirection works:

```bash
mcc sample.cu >stdout.txt 2>stderr.txt
```

The program controls its own buffering. Call `fflush(stdout)` after progress
messages in C/C++ when you need them immediately. Output order is preserved within
each stream; ordering between stdout and stderr is not guaranteed.

Compiler and program exit codes are returned to your shell. Signals use `128 +
signal`; timeout returns `124`, Ctrl-C `130`, invalid CLI arguments `2`, and other
CLI/transfer failures `1`. The remote app is stopped when the command ends.

## Development

```bash
uv sync
uv run python -m mcc sample.cu
uv run pytest -q
uv run ruff check .
uv run ruff format --check .
uv build
```

The default tests use local files and real subprocesses, with no mocks. Run the
GPU tests separately; they make real Modal calls on a T4 and incur GPU usage:

```bash
MCC_TEST_MODAL=1 uv run pytest -q tests/test_modal.py
```

These cover a CUDA kernel, multiple source files, ignored files, arguments, binary
downloads, live output, compiler/program failures, timeout, and cancellation.

## License

MIT © ExpressGradient
