Metadata-Version: 2.4
Name: imgcheck
Version: 0.1.0
Summary: Perceptual-hash based image deduplication and train/test leakage detection for image classification datasets
Author-email: Rama Narasimha Kanduri <kramanarasimha1@gmail.com>, CS Hari Krishna <csharikrishna1806@gmail.com>
License: MIT License
        
        Copyright (c) 2026 Rama Narasimha Kanduri & CS Hari Krishna
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
        -------------------------------------------------------------------------------
        
        ACKNOWLEDGEMENT REQUEST (not legally binding, but appreciated)
        
        If you use imgcheck in your research, project, or product, we would really
        appreciate it if you could — if possible — give a small credit in any of
        these ways, whichever feels right to you:
        
          Option 1 — Credit the tool:
              "Dataset auditing was performed using imgcheck
               (https://pypi.org/project/imgcheck)."
        
          Option 2 — Credit the authors:
              "Dataset auditing was performed using imgcheck
               (Rama Narasimha Kanduri & CS Hari Krishna, 2026)."
        
          Option 3 — Credit both:
              "Dataset auditing was performed using imgcheck by
               Rama Narasimha Kanduri & CS Hari Krishna
               (https://pypi.org/project/imgcheck)."
        
        Any acknowledgement — the tool name, the authors, or both — is warmly
        appreciated. Thank you for using imgcheck!
        
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: Pillow>=9.0
Requires-Dist: ImageHash>=4.3
Requires-Dist: pandas>=1.5
Requires-Dist: openpyxl>=3.0
Provides-Extra: pdf
Requires-Dist: fpdf2>=2.7; extra == "pdf"
Dynamic: license-file

# imgcheck

> Perceptual-hash based tools for auditing and cleaning image classification datasets —
> find duplicates, detect train/test leakage, and catch cross-label mislabelling, all from one CLI.

---

## What it does

Image classification datasets silently suffer from three data quality problems that hurt model accuracy:

| Problem | What it means | Detect | Remove |
|---|---|---|---|
| **Within-class duplicates** | Same image appears multiple times in one class, inflating dataset size | `imgcheck clean` | `imgcheck clean-inplace` |
| **Train/test leakage** | A test image is a near-duplicate of a train image — model has effectively "seen" it, inflating test accuracy | `imgcheck leakage` | `imgcheck leakage-fix` |
| **Cross-label duplicates** | Same image appears under two *different* class labels — contradictory training signal | `imgcheck cross-label` | `imgcheck cross-label-fix` |

Plus a fast no-hashing sanity check:

| Command | What it gives you |
|---|---|
| `imgcheck stats` | Per-class image counts, resolutions, class imbalance ratio, corrupt files |

---

## Real-world example (Brain Tumor MRI Dataset)

Suppose your dataset looks like:

```
dataset/
├── glioma/        (1000 images)
├── meningioma/    (800 images)
├── pituitary/     (900 images)
└── no_tumor/      (500 images)

train/
├── glioma/
├── meningioma/
├── pituitary/
└── no_tumor/

test/
├── glioma/
├── meningioma/
├── pituitary/
└── no_tumor/
```

### Step 1 — Quick sanity check (no hashing, instant)

```bash
imgcheck stats --dataset dataset/ --report reports/stats.xlsx
```

Output tells you: image counts per class, resolution ranges, class imbalance ratio (e.g. `2.0x`), and any corrupt files.

---

### Step 2 — Find and remove within-class duplicates

**Option A — Safe: copy survivors to a new clean folder (original untouched)**
```bash
imgcheck clean \
    --dataset dataset/ \
    --output  dataset_clean/ \
    --report  reports/duplicates.xlsx \
    --html    reports/duplicates.html
```

**Option B — In-place: delete duplicates directly from the dataset**
```bash
imgcheck clean-inplace \
    --dataset dataset/ \
    --report  reports/duplicates_removed.xlsx
```

Both generate a report showing exactly which images were removed and which were kept, with Cluster IDs grouping chains of similar images.

---

### Step 3 — Find and remove train/test leakage

**First, detect and review (read-only)**
```bash
imgcheck leakage \
    --train  train/ \
    --test   test/ \
    --report reports/leakage.xlsx \
    --html   reports/leakage.html
```

**Then, remove — Option A: delete leaked test images in-place**
```bash
imgcheck leakage-fix \
    --train  train/ \
    --test   test/ \
    --report reports/leakage_removed.xlsx \
    --mode   inplace
```

**Or — Option B: copy only clean test images to a new folder**
```bash
imgcheck leakage-fix \
    --train   train/ \
    --test    test/ \
    --report  reports/leakage_removed.xlsx \
    --mode    copy \
    --output  test_clean/
```

---

### Step 4 — Find and remove cross-label mislabelling

**First, detect and review (read-only)**
```bash
imgcheck cross-label \
    --dataset dataset/ \
    --report  reports/crosslabel.xlsx \
    --html    reports/crosslabel.html
```

**Then, remove — Option A: delete the mislabelled copy in-place**
```bash
imgcheck cross-label-fix \
    --dataset dataset/ \
    --report  reports/crosslabel_removed.xlsx \
    --mode    inplace
```

**Or — Option B: copy clean dataset (minus mislabelled images) to a new folder**
```bash
imgcheck cross-label-fix \
    --dataset dataset/ \
    --report  reports/crosslabel_removed.xlsx \
    --mode    copy \
    --output  dataset_clean/
```

> **Which copy gets removed?** For each cross-label pair, imgcheck keeps the image in the alphabetically-earlier class and removes the copy in the later class. This is deterministic regardless of file discovery order.

---

## How it works

All hashing commands use **perceptual hashing** — not cryptographic hashing. A perceptual hash encodes the *visual appearance* of an image as a short bitstring. Two images that look identical (even if they differ in JPEG quality, slight cropping, or minor brightness) produce hashes with a low **Hamming distance** (number of differing bits).

**Threshold guide:**

| `--threshold` | What gets matched |
|---|---|
| `0` | Exact pixel-perfect duplicates only |
| `1–3` | Near-duplicates (default: `3`) |
| `4–10` | Resized, re-compressed, or slightly edited variants |

**Clustering** uses **Union-Find**, so a transitive chain A ~ B ~ C is correctly collapsed into one group with one representative — not three confusing separate pairs.

---

## Dataset structure required

All commands expect a root folder of **class subfolders**:

```
dataset/
├── cats/
│   ├── img001.jpg
│   └── img002.png
├── dogs/
│   └── img003.jpg
└── birds/
    └── ...
```

Class names can be **anything** — the tool reads them dynamically from disk. Nothing is hardcoded.

For `leakage` and `leakage-fix`, train and test roots must have **matching class subfolder names**:

```
train/              test/
├── cats/           ├── cats/
├── dogs/           ├── dogs/
└── birds/          └── birds/
```

---

## Install

From the project root (where `pyproject.toml` lives):

```bash
pip install -e .
```

Installs the `imgcheck` command with all required dependencies:
`Pillow`, `ImageHash`, `pandas`, `openpyxl`

### Optional: PDF reports

```bash
pip install -e ".[pdf]"
```

Adds `fpdf2` for the `--pdf` flag. The `--html` flag needs no extra install.

---

## Full command reference

### `imgcheck stats`

```bash
imgcheck stats \
    --dataset path/to/dataset \
    --report  path/to/stats.xlsx        # optional
```

### `imgcheck clean` *(detection + copy)*

```bash
imgcheck clean \
    --dataset   path/to/dataset \
    --output    path/to/dataset_clean \
    --report    path/to/duplicates.xlsx \
    --threshold 3 \
    --algo      phash \
    --html      path/to/duplicates.html \
    --pdf       path/to/duplicates.pdf \
    --formats   csv,json
```

### `imgcheck clean-inplace` *(removal — destructive)*

```bash
imgcheck clean-inplace \
    --dataset   path/to/dataset \
    --report    path/to/removed.xlsx \
    --threshold 3 \
    --algo      phash \
    --formats   csv,json
```

### `imgcheck leakage` *(detection only)*

```bash
imgcheck leakage \
    --train     path/to/train \
    --test      path/to/test \
    --report    path/to/leakage.xlsx \
    --threshold 3 \
    --algo      phash \
    --html      path/to/leakage.html \
    --pdf       path/to/leakage.pdf \
    --formats   csv,json
```

### `imgcheck leakage-fix` *(removal)*

```bash
# inplace — deletes from test folder directly
imgcheck leakage-fix \
    --train     path/to/train \
    --test      path/to/test \
    --report    path/to/leakage_removed.xlsx \
    --mode      inplace \
    --threshold 3 \
    --algo      phash

# copy — writes clean test set to a new folder
imgcheck leakage-fix \
    --train     path/to/train \
    --test      path/to/test \
    --report    path/to/leakage_removed.xlsx \
    --mode      copy \
    --output    path/to/test_clean \
    --threshold 3
```

### `imgcheck cross-label` *(detection only)*

```bash
imgcheck cross-label \
    --dataset   path/to/dataset \
    --report    path/to/crosslabel.xlsx \
    --threshold 3 \
    --algo      phash \
    --html      path/to/crosslabel.html \
    --pdf       path/to/crosslabel.pdf \
    --formats   csv,json
```

### `imgcheck cross-label-fix` *(removal)*

```bash
# inplace — deletes the mislabelled copy from dataset directly
imgcheck cross-label-fix \
    --dataset   path/to/dataset \
    --report    path/to/crosslabel_removed.xlsx \
    --mode      inplace \
    --threshold 3

# copy — writes clean dataset (minus mislabelled images) to a new folder
imgcheck cross-label-fix \
    --dataset   path/to/dataset \
    --report    path/to/crosslabel_removed.xlsx \
    --mode      copy \
    --output    path/to/dataset_clean \
    --threshold 3
```

---

## Options reference

### Hashing options (`clean`, `clean-inplace`, `leakage`, `leakage-fix`, `cross-label`, `cross-label-fix`)

| Flag | Default | Description |
|---|---|---|
| `--threshold N` | `3` | Max Hamming distance to count as a match. `0` = exact only, `1–3` = near-duplicates, `4+` = resized/re-compressed variants |
| `--algo NAME` | `phash` | Perceptual hash algorithm |

| Algorithm | Best for |
|---|---|
| `phash` | General use, most robust (default) |
| `ahash` | Fastest, least accurate |
| `dhash` | Gradient/edge-heavy images |
| `whash` | Wavelet-based, handles compression artifacts well |

### Report output options (`clean`, `leakage`, `cross-label`)

| Flag | Description |
|---|---|
| `--html PATH` | Self-contained HTML gallery, no extra dependency |
| `--pdf PATH` | Visual PDF report, requires `pip install imgcheck[pdf]` |
| `--formats csv,json` | Also export report tables as CSV and/or JSON |

### Removal mode options (`leakage-fix`, `cross-label-fix`)

| Flag | Description |
|---|---|
| `--mode inplace` | Delete flagged images from source folder (default, destructive) |
| `--mode copy` | Copy surviving images to a new clean folder (safe) |
| `--output PATH` | Required when `--mode copy` |

---

## Using imgcheck as a Python library

```python
from imgcheck.stats import compute_dataset_stats
from imgcheck.clean import clean_duplicates, clean_duplicates_inplace
from imgcheck.leakage import check_train_test_leakage
from imgcheck.leakage_fix import fix_train_test_leakage
from imgcheck.cross_label import check_cross_label_duplicates
from imgcheck.cross_label_fix import fix_cross_label_duplicates

# Health check
compute_dataset_stats("dataset/", report_path="stats.xlsx")

# Find duplicates → copy survivors to new folder
clean_duplicates("dataset/", "dataset_clean/", "duplicates.xlsx", threshold=3)

# Remove duplicates in-place
clean_duplicates_inplace("dataset/", "duplicates_removed.xlsx", threshold=3)

# Detect leakage
check_train_test_leakage("train/", "test/", "leakage.xlsx", threshold=3)

# Remove leaked test images in-place
fix_train_test_leakage("train/", "test/", "leakage_removed.xlsx", mode="inplace", threshold=3)

# Remove leaked test images → copy clean set to new folder
fix_train_test_leakage("train/", "test/", "leakage_removed.xlsx", mode="copy", output_root="test_clean/")

# Detect cross-label mislabelling
check_cross_label_duplicates("dataset/", "crosslabel.xlsx", threshold=3)

# Remove cross-label duplicates → copy clean dataset to new folder
fix_cross_label_duplicates("dataset/", "crosslabel_removed.xlsx", mode="copy", output_root="dataset_clean/")
```

---

## Recommended audit workflow

```bash
# 1. Sanity check — instant, no hashing
imgcheck stats --dataset dataset/

# 2. Remove within-class duplicates (safe copy)
imgcheck clean --dataset dataset/ --output dataset_clean/ --report reports/duplicates.xlsx --html reports/duplicates.html

# 3. Remove cross-label mislabelling (safe copy)
imgcheck cross-label-fix --dataset dataset/ --mode copy --output dataset_fixed/ --report reports/crosslabel_removed.xlsx

# 4. Check and remove train/test leakage (safe copy)
imgcheck leakage-fix --train train/ --test test/ --mode copy --output test_clean/ --report reports/leakage_removed.xlsx
```

---

## Project layout

```
imgcheck/
├── pyproject.toml
├── README.md
└── src/imgcheck/
    ├── __init__.py          # package version and authors
    ├── hashing.py           # image discovery, perceptual hashing, Hamming distance
    ├── grouping.py          # Union-Find transitive duplicate clustering
    ├── clean.py             # within-class duplicate detection + copy or inplace removal
    ├── leakage.py           # train/test leakage detection (read-only)
    ├── leakage_fix.py       # train/test leakage removal (inplace or copy)
    ├── cross_label.py       # cross-label duplicate detection (read-only)
    ├── cross_label_fix.py   # cross-label duplicate removal (inplace or copy)
    ├── stats.py             # fast dataset health check (no hashing)
    ├── export.py            # optional CSV/JSON export alongside Excel
    ├── report.py            # optional PDF (fpdf2) + HTML visual reports
    └── cli.py               # imgcheck <command> ... entry point
```

---

## Authors

**Rama Narasimha Kanduri**
kramanarasimha1@gmail.com

**CS Hari Krishna**
csharikrishna1806@gmail.com
