Metadata-Version: 2.4
Name: data_check
Version: 0.20.0
Summary: simple data validation
Author-email: Andreas Rjasanow <andrjas@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://andrjas.github.io/data_check/
Project-URL: Repository, https://github.com/andrjas/data_check
Keywords: data,validation,testing,quality
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Other Audience
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Database
Classifier: Topic :: Software Development
Classifier: Topic :: Software Development :: Testing
Requires-Python: <3.14,~=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: SQLAlchemy<3,>=2.0.19
Requires-Dist: pandas<3,>=2.2.1
Requires-Dist: pyyaml~=6.0
Requires-Dist: click<9,>=8.1.5
Requires-Dist: colorama<0.5,>=0.4.6
Requires-Dist: Jinja2<4,>=3.1.2
Requires-Dist: openpyxl<4,>=3.1.2
Requires-Dist: click-default-group<2,>=1.2.2
Requires-Dist: Faker>=37.1.0
Requires-Dist: pydantic<3,>=2
Provides-Extra: postgres
Requires-Dist: psycopg2-binary<3,>=2.9.6; extra == "postgres"
Provides-Extra: oracle
Requires-Dist: cx_Oracle<9,>=8.3.0; extra == "oracle"
Provides-Extra: oracledb
Requires-Dist: oracledb<4,>=3.1.0; extra == "oracledb"
Provides-Extra: mysql
Requires-Dist: pymysql[rsa]<2,>=1.1.0; extra == "mysql"
Provides-Extra: mssql
Requires-Dist: pyodbc<6,>=5; extra == "mssql"
Provides-Extra: duckdb
Requires-Dist: duckdb-engine<0.18,>=0.17.0; extra == "duckdb"
Provides-Extra: databricks
Requires-Dist: databricks-sqlalchemy<3,>=2.0.5; extra == "databricks"
Dynamic: license-file

# data_check

data_check is a simple data validation tool. In its most basic form it will execute SQL queries and compare the results against CSV or Excel files. But there are more advanced features:

## Features

* [CSV checks](https://andrjas.github.io/data_check/csv_checks/): compare SQL queries against CSV files
* Excel support: Use Excel (xlsx) instead of CSV
* multiple environments (databases) in the configuration file
* [populate tables](https://andrjas.github.io/data_check/loading_data/) from CSV or Excel files
* [execute any SQL files on a database](https://andrjas.github.io/data_check/sql/)
* more complex [pipelines](https://andrjas.github.io/data_check/pipelines/)
* run any script/command (via pipelines)
* simplified checks for [empty datasets](https://andrjas.github.io/data_check/csv_checks/#empty-dataset-checks) and [full table comparison](https://andrjas.github.io/data_check/csv_checks/#full-table-checks)
* [lookups](https://andrjas.github.io/data_check/csv_checks/#lookups) to reuse the same data in multiple queries
* [test data generation](https://andrjas.github.io/data_check/test_data/)

## Database support

data_check is tested with these databases:

- PostgreSQL
- MySQL
- SQLite
- Oracle
- Microsoft SQL Server

Partially supported:

- DuckDB
- Databricks

Other databases supported by [SQLAlchemy](https://docs.sqlalchemy.org/en/20/dialects/) might also work.

## Quickstart

You need Python 3.9 or above to run data_check. The easiest way to install data_check is via [pipx](https://github.com/pipxproject/pipx):

`pipx install data-check`

The data_check Git repository is also a sample data_check project. Clone the repository, switch to the folder and run data_check:

```
git clone git@github.com:andrjas/data_check.git
cd data_check/example
data_check
```

This will run the tests in the _checks_ folder using the default connection as set in data_check.yml.

See the [documentation](https://andrjas.github.io/data_check) how to install data_check in different environments with additional database drivers and other usages of data_check.

## Project layout

data_check has a simple layout for projects: a single configuration file and a folder with the test files. You can also organize the test files in subfolders.

    data_check.yml    # The configuration file
    checks/           # Default folder for data tests
        some_test.sql # SQL file with the query to run against the database
        some_test.csv # CSV file with the expected result
        subfolder/    # Tests can be nested in subfolders

## CSV checks

This is the default mode when running data_check. data_check expects a SQL file and a CSV file. The SQL file will be executed against the database and the result is compared with the CSV file. If they match, the test is passed, otherwise it fails.

## Pipelines

If data_check finds a file named _data\_check\_pipeline.yml_ in a folder, it will treat this folder as a pipeline check. Instead of running [CSV checks](#csv-checks) it will execute the steps in the YAML file.

Example project with a pipeline:

    data_check.yml
    checks/
        some_test.sql                # this test will run in parallel to the pipeline test
        some_test.csv
        sample_pipeline/
            data_check_pipeline.yml  # configuration for the pipeline
            data/
                my_schema.some_table.csv       # data for a table
            data2/
                some_data.csv        # other data
            some_checks/             # folder with CSV checks
                check1.sql
                check1.csl
                ...
            run_this.sql             # a SQL file that will be executed
            cleanup.sql
        other_pipeline/              # you can have multiple pipelines that will run in parallel
            data_check_pipeline.yml
            ...

The file _sample\_pipeline/data\_check\_pipeline.yml_ can look like this:

```yaml
steps:
    # this will truncate the table my_schema.some_table and load it with the data from data/my_schema.some_table.csv
    - load: data
    # this will execute the SQL statement in run_this.sql
    - sql: run_this.sql
    # this will append the data from data2/some_data.csv to my_schema.other_table
    - load:
        file: data2/some_data.csv
        table: my_schema.other_table
        mode: append
    # this will run a python script and pass the connection name
    - cmd: "python3 /path/to/my_pipeline.py --connection {{CONNECTION}}"
    # this will run the CSV checks in the some_checks folder
    - check: some_checks
```

Pipeline checks and simple CSV checks can coexist in a project.

## Documentation

See the [documentation](https://andrjas.github.io/data_check) how to setup data_check, how to create a new project and more options.

## License

[MIT](LICENSE)
