Metadata-Version: 2.4
Name: user-data-ingest-cli
Version: 0.1.140
Summary: CLI tool for ingesting user data via the API
Author: ASTRON SDC
Classifier: License :: OSI Approved :: Apache Software License
Requires-Python: >=3.13
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: click
Requires-Dist: requests
Requires-Dist: fsspec
Requires-Dist: PySide6
Dynamic: license-file

# User Data Ingest CLI (udicli)

Command-line client for uploading data into LOFAR 2.0's Long-Term Archive. Built for projects with non-standard pipelines or independent processing resources that need to submit data while adhering to LTA schemas.

See the [backend documentation](https://git.astron.nl/astron-sdc/user-ingest-backend) for more context. There's also a [web interface](https://sdchealth.fuse-astron.src.surf-hosted.nl/user-data-ingest/) if you prefer that.

## Getting Started

Python 3.13+ required ([download here](https://www.python.org)).

```bash
pip install user-data-ingest-cli
udicli session login
```

`udicli` is now on your PATH.

## Commands

```bash
udicli session login              Authenticate with an SRAM application token
udicli session logout             Remove cached credentials
udicli session whoami             Show your account details
udicli system status              Check server connectivity
udicli projects list              View projects you are a member of
udicli requests create            Create and submit an ingest request
udicli requests list              List ingest requests
udicli requests show <id>         Show an ingest request
udicli requests status <id>       Show an ingest request's current state
udicli requests watch <id>        Follow an ingest request's state
udicli requests validate <id>     Signal that an upload is ready for validation
udicli storage list               List supported storage backends
udicli storage connect [backend]  Select a storage backend
udicli config                     Show or manage persistent defaults
udicli completion <bash|zsh|fish> Generate shell completion
```

For help on any command:

```bash
udicli <topic> --help
udicli <topic> <action> --help
```

Use `-v` flag with any command to see raw requests and responses:

```bash
udicli -v system status
```

Commands that return structured data support `--output table` (the default) and `--output json`.
Use the global `--quiet` option to suppress non-essential progress output.

## Configuration

Persistent, non-secret defaults are stored in `~/.config/user-ingest-cli/config.json` (or under
`$XDG_CONFIG_HOME`). Credentials remain in their separate protected files.

```bash
udicli config show
udicli config set environment prod
udicli config set default-project APPPP_001
udicli config set default-site surf
udicli config set default-checksum-type MD5
udicli config set output json
udicli config unset default-project
udicli config path
```

Values are resolved in this order: command-line option, environment variable, saved setting,
built-in default.

## Inspecting Requests

```bash
udicli requests list
udicli requests list --state UPLOADING --project APPPP_001 --since 2026-08-01 --limit 20
udicli requests show <request-id>
udicli requests status <request-id> --output json
udicli requests watch <request-id>
udicli requests watch <request-id> --until VALIDATING --until FAILED --interval 5 --timeout 300
udicli requests validate <request-id>
```

## Shell Completion

Generate a completion script for Bash, Zsh, or Fish. For example, with Zsh:

```bash
eval "$(udicli completion zsh)"
```

## Authentication

Signing in to User Data Ingest with `session login` prompts you to log in with an account from an institution you are affiliated with.
If you do not have an account with an institution that is part of the EduGAIN federation, you can create an account with one of the eduID services (e.g. the Dutch one).

To request access to this service, you can use this link : https://sram.surf.nl/registration?collaboration=e80b9e70-62f8-4b05-890e-c40dd9d2ac68

After the membership is approved you can go to the [Application token page](https://sram.surf.nl/collaborations/7733/tokens) and click " Create application token". Make sure to copy the given token somewhere safe as this will not be visible
a second time. To then log in via the cli tool you can either put it in a `.ingestingrc` file in your home directory or give it as an argument to the login command, as explained below.
![Create an application token page](docs/images/create-token.png)

`session login` looks for your SRAM application token in this order:

1. `--token` (optional) flag: `udicli session login --token <your-token>`

2. `.ingestingrc` file in your home directory with `api_token=<your-token>`
   macOS/Linux: `~/.ingestingrc`
   Windows: `C:\Users\<username>\.ingestingrc`
   usage: `udicli session login`

3. Interactive prompt: provides a link to SRAM and asks you to paste the token

Once verified, the token is cached locally at `~/.config/user-ingest-cli/credentials-<env>`, so you won't need to authenticate again.

## Ingesting Data

An ingest request has this structure: project → data products → files.

Files and metadata are collected (flags, prompts, CSV file or `--inputfile`) and the request structure is verified. You then get a preview in the input file format, which you can save as `.json` and reuse with `--inputfile` (see [JSON Config File](#json-config-file)).
Next you are asked whether to proceed (skipped with `--noinput`).

Example preview:

```json
{
  "project_id": "APPPP_001",
  "site": "surf",
  "checksum_type": "MD5",
  "data_products": [
    {
      "data_product_type": "VisibilityDataProduct",
      "files_format": "MEASUREMENT_SET",
      "files": ["/path/to/file1.m5", "/another/path/to/file2.m5"]
    }
  ]
}
```

### Interactive Mode

Start with no arguments to be guided through the whole process:

```bash
udicli requests create
```

### Automated Mode

**Note:** file locations can be absolute or relative. Anywhere below that accepts files, glob
patterns like `"data/*.m5"` are supported (matching all `.m5` files in `data`); only `*` and `?`
are recognized, not `[..]` character classes or `**` recursion.

Provide every value up front so nothing needs to be typed interactively: via command-line flags,
a piped or CSV file list, or a JSON config file.

#### Command-Line Flags

Provide all details via flags .

```bash
udicli requests create --project-id APPPP_001 --site surf --checksum-type MD5 --files-format MEASUREMENT_SET --data-product-type VisibilityDataProduct --files test_files/file1.m5 --files test_files/file2.m5
udicli requests create --project-id APPPP_001 --site surf --checksum-type MD5 --files-format MEASUREMENT_SET --data-product-type VisibilityDataProduct --files test_files/file1.m5 --files /home/<user>/Downloads/file2.m5


```

Repeat `--files` for multiple files. Don't repeat other flags (they won't stack; only the last one counts).

#### `--noinput` and `--dry-run`

Two independent flags for running `requests create` non-interactively.

**`--noinput`** disables every prompt, including `Proceed?`. Anything not supplied as a flag,
`--inputfile` or saved config makes the command fail immediately.
Use it for scripts, cron and CI. Piping input (`--files-from -`) implies it.

**`--dry-run`** only runs verification of the request structure.

```bash
udicli requests create --project-id APPPP_001 --site surf --checksum-type MD5 \
  --data-product-type VisibilityDataProduct --files-format MEASUREMENT_SET \
  --files "test_files/*.m5" --dry-run
```

#### Piping File Lists

Use `--files-from -` to read from stdin:

```bash
find test_files -name "*.m5" -mtime -1 | udicli requests create --project-id APPPP_001 --site surf --checksum-type MD5 --data-product-type VisibilityDataProduct --files-format MEASUREMENT_SET --files-from -
find test_files -name "*.m5" | udicli requests create --project-id APPPP_001 --site surf --checksum-type MD5 --data-product-type VisibilityDataProduct --files-format MEASUREMENT_SET --files-from -
```

#### CSV File

Pass a file with one path per line:

```bash
udicli requests create --project-id APPPP_001 --site surf --checksum-type MD5 --data-product-type VisibilityDataProduct --files-format MEASUREMENT_SET --files-from test_config_files/config3.csv
```

Empty lines in the file are ignored.

#### JSON Config File

For multiple data products, filters, or repeat runs.
Tip: Save the preview printed by `requests create` as a `.json` file to reuse it.

```json
{
  "project_id": "APPPP_001",
  "site": "surf",
  "checksum_type": "MD5",
  "data_products": [
    {
      "data_product_type": "VisibilityDataProduct",
      "files_format": "MEASUREMENT_SET",
      "files": [
        "test_files/file*.m5"
      ]
    },
    {
      "data_product_type": "VisibilityDataProduct",
      "files_format": "MS",
      "files": [
        "test_files/*.MS"
      ],
      "filters": {
        "min_size": "1000000",
        "max_size": 100000000,
        "modified_after": "2024-01-01",
        "exclude": [
          "*test*.m5",
          "*backup*.m5"
        ]
      }
    }
  ]
}
```

Run with:

```bash
udicli requests create --inputfile test_config_files/inputFile1.json
udicli requests create --inputfile test_config_files/inputFile2.json
```

#### Filters (JSON Config Only)

Filters refine file selection after glob patterns expand. All are optional; a file must pass every filter to be included.

```json
"filters": {
  "min_size": 1000000,
  "max_size": 100000000,
  "modified_after": "2026-07-22",
  "include": ["file1.m5", "file2.m5"],
  "exclude": ["*test*.m5", "*backup*.m5"]
}
```

`min_size` and `max_size` accept numbers or strings (e.g., `1000000` or `"1000000"`).

`modified_after` takes ISO 8601 format: `"2026-07-22"` or `"2026-07-22T14:00:00"`.

`include` and `exclude` use glob patterns (`*.m5`, `file?.m5`). Character classes `[..]` and recursive globs `**` are not supported.

Exist status :
0 - success
1 - error
2 - invalid
command syntax or option values

## Starting Validation at the Long Term Archive
After all files for a request have been uploaded to its landing space, signal that the backend may start validation:
TBD : Tools for uploading files to lta locations.

```bash
udicli requests validate <request-id>
```

The request must belong to an accessible project and be in the `UPLOADING` state. This command starts
the backend validation workflow; it does not validate files locally.

## Testing and Feedback

We're actively developing this tool.

1. What do your processing resources typically include and how do you connect to them?
2. Note: File format verification isn't implemented yet. We're waiting on the LOFAR Data Working Group to approve the list of valid data products and formats.
3. Should a data product with zero files be rejected with an error or allowed with a warning?

## Development

Clone and set up:

```bash
git clone <this-repo>
cd user-data-ingest-cli
python -m venv .venv
source .venv/bin/activate
make bootstrap
make init
pip install -e .
```

For development start UDI backend (https://git.astron.nl/astron-sdc/user-ingest-backend) locally ( it defaults to port 8005) and connect to it using `udicli --env dev session login`
Switching environments with `--env dev` or `--env prod` keeps sessions separate.

Runtime dependencies are declared in `pyproject.toml`; the files in `requirements/` are pinned lockfiles compiled from it. After changing dependencies, recompile and reinstall:

```bash
make compile-requirements
pip install -r requirements/dev.txt
pip install -e .
```

Run linting, formatting, type-checking, and tests:

```bash
ruff check src tests
black src tests
mypy src
pytest
```

### Local State

The SRAM token obtained via `session login` is written to `~/.config/user-ingest-cli/credentials-{env}` (a separate file per `--env`) and stays there until you run `udicli session logout` or delete the file yourself; there's no client side expiry check, so a revoked or expired token will just fail on the next request to the API.
Project codes are never cached: `projects list` and the interactive project prompt in `requests create` both call the API every time, so they always reflect your current memberships.
Everything else, such as which environment you're targeting or which API base URL to call, lives in the in memory `Config` object, rebuilt from `constants.py` and CLI flags/environment variables on every invocation.

## Deployment

Every pipeline builds the package (`package_files`) with an auto-incrementing version (`0.1.<commit count>` — no manual version action needed). From there, two manual jobs are available in the pipeline's `publish` stage:

**Test deploy to GitLab** — click `publish_on_gitlab` to start the pipeline. Uploads to this project's own Package Registry, authenticated automatically via `CI_JOB_TOKEN`

```bash
pip install user-data-ingest-cli --index-url https://__token__:<your-personal-access-token>@git.astron.nl/api/v4/projects/1019/packages/pypi/simple
```

**Release to PyPI** — push a tag, then click `publish_on_pypi`:

```bash
git tag v0.2.0
git push origin v0.2.0
```

That job only appears on tag pipelines and uploads to the real [pypi.org](https://pypi.org/project/udicli/) using a `PYPI_TOKEN` CI/CD variable (set under Settings → CI/CD → Variables, masked + protected). Once it's up:

```bash
pip install user-data-ingest-cli
```
