Metadata-Version: 2.4
Name: sparpartner
Version: 3.6.0
Summary: A deterministic, benchmark-driven stratified sampler profile-matching engine for tabular data
Author: Henry
Author-email: Henry <osas2henry@gmail.com>
License: All Rights Reserved
Classifier: License :: Other/Proprietary License
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas>=1.3
Dynamic: author
Dynamic: license-file
Dynamic: requires-python

# sparpartner

A deterministic, benchmark-driven stratified sampler profile-matching
engine for tabular data. You define a benchmark
profile (the thing you actually care about matching, however you want
to define it), `sparpartner` ranks every row in your data by how
closely it resembles that profile, and `slicer` gives you a
principled, optional way to decide what "resembles" even means,
instead of you guessing weights by feel.

Train/test splitting is one thing this is good for. It's not the only
one. See [Use cases](#use-cases) below for others, like finding the
customers, matches, or candidates that most resemble an ideal profile
you define.

## Why this exists

Ranking rows against a benchmark, instead of scoring them in
isolation, turns out to answer a few different questions depending on
what you feed it as the benchmark:

- **"Does my model generalize, or did it just memorize a
  neighborhood?"** A random train/test split assumes the test set
  should look statistically like the train set. That's the wrong
  question if what you actually want to know is whether the model
  generalizes past one specific profile. Set the benchmark to the
  toughest, most representative case, and the rows most like it become
  your test set, the harder and more honest check.
- **"Which of my customers/rows most resemble my ideal profile?"**
  Same ranking machinery, different benchmark. Set the benchmark to
  an ideal customer profile instead of a "hardest test case" profile,
  and the exact same ranking now tells you who to prioritize, not who
  to hold out. See [Use cases](#use-cases) for a full walkthrough.

Either way, `sparpartner` answers it by ranking every row in your
data by how closely it resembles a benchmark ("bench_marks") you
define, then sorting closest-to-benchmark first. You then slice the
sorted frame yourself, or let `sparring_n` and `proximity_gates` do it
for you (see [Proximity gates](#proximity-gates-proximity_gates)
below):

- **Lookalikes as test, the rest as train**: the rows most like the
  benchmark are always at the top of the sorted result, so `head()`
  gives you the test set and `tail()` (or everything past your cut)
  gives you the train set. This is the harder, more honest check: it
  tells you whether the model actually learned something general, or
  only performs well near cases it's already seen a lot of.
- **Just the lookalikes, nothing else**: pass `sparring_n` and skip
  the manual slice entirely. `sample()` hands you back only the top
  `sparring_n` rows, already ranked, plus a report on exactly that
  set (see the sparring report section below).
- **Only the lookalikes that also clear a closeness bar**: pass
  `proximity_gates` alongside `sparring_n` (or on its own) when you
  want every row you get back to individually clear a minimum
  closeness score, either on specific signals or on its own final
  total match score. See
  [Proximity gates](#proximity-gates-proximity_gates).

`sparpartner` only produces the ranking (and, optionally, a
proximity-gated pool, and, optionally, the top-N slice on top of
that). What you do with that ranking, a train/test cut, a shortlist
of best-fit rows, or something else entirely, is on your side (see
[Use cases](#use-cases) and the usage examples below).

## Where the idea comes from

A fighter in camp doesn't spar with whoever's free in the gym, they
specifically look for a sparring partner who moves, reaches, and
hits like the opponent they're about to face. Training against a
random partner tells you nothing about how you'll actually do;
training against someone who resembles the real threat does.
`sparpartner` applies that same logic to a model: instead of a
random holdout, it finds the rows that resemble the toughest, most
relevant "opponent" profile and holds those back as the real test,
so what's left to train on is everything *unlike* that opponent,
and the test genuinely checks whether the model can handle the
match it's actually walking into.

## Use cases

`sample()` and `slicer()` don't know or care what your benchmark
represents, that's entirely up to what you put in `bench_marks`. A
few different framings, same two functions.

### 1. Train/test split (the hardest, most honest holdout)

Set the benchmark to the toughest, most representative case you can
define. The rows most like it become your test set, everything else
trains the model. This is the use case covered in depth throughout
the rest of this README, see [Sample usage](#sample-usage) below for
the full walkthrough.

### 2. Profile matching (finding your best-fit rows)

Same ranking, different intent. Instead of asking "which rows should
I hold out as a hard test," ask "which rows most resemble the profile
I actually want more of." Point `bench_marks` at an ideal customer,
an ideal candidate, an ideal match, whatever "ideal" means for your
data, and the ranking now tells you who to prioritize.

```python
customers = df   # your customer table

ideal_customer = {
    "country": "Nigeria",
    "age": 32,
    "income": 450000,
    "last_active_date": "2026-08-20",
    "channel": "referral",
}

weights = slicer(
    targets="country",
    temporal={"recency": "last_active_date"},
    causatives=["income", "age", "channel"],
    decay_causatives=True,
)
# {'country': 50.0, 'last_active_date': 25.0,
#  'income': 12.5, 'age': 8.33, 'channel': 4.17}

best_customers, report = sample(
    customers,
    bench_marks=ideal_customer,
    custom_weights=weights,
    best_first=True,
)
```

`best_customers` comes back sorted so the rows most like
`ideal_customer` are first, exactly the same mechanics as the
train/test case, just pointed at a different kind of benchmark. This
`weights` dict came from `slicer`, an optional helper for exactly
this situation: turning a hierarchy of "what matters most" into
concrete numbers instead of guessing them by feel. See
[Generating `custom_weights` with `slicer`](#generating-custom_weights-with-slicer)
further down for the full explanation of how it works.

`bystanders` fits naturally here too, a feature can be observed in
`sparring_report` without being allowed to influence who gets
selected:

```
TARGETS
country
   |
TEMPORAL
last_active_date
   |
CAUSATIVES
income
age
channel
   |
BYSTANDERS (reported, not weighted)
customer_id
region
```

```python
weights = slicer(
    targets="country",
    temporal={"recency": "last_active_date"},
    causatives=["income", "age", "channel"],
    decay_causatives=True,
    bystanders=["customer_id", "region"],
)
```

`region` and `customer_id` now show up in `sparring_report` as
`spar_region` / `spar_customer_id`, so you can see how close matches
tend to be on those dimensions too, without either one moving who
actually gets ranked as a best-fit customer.

### 3. Anything else that reduces to "rank rows by resemblance to X"

Candidate screening against an ideal-hire profile, lead scoring
against your best-converting customer, match-finding against a
target opponent profile, the underlying operation is always the same
"rank by resemblance to a benchmark," only the benchmark and what you
do with the ranking changes.

## How the scoring works

You give it:

- `df`: your data. Any column names are fine, including ones
  starting with `spar`; `sample()` never touches or overwrites
  your own columns (see [Validation](#validation)).
- `bench_marks`: a dict of `{column_name: benchmark_value}`, one
  entry per signal you care about. Every value must be a real,
  non-null value, a `None`, `NaN`, or infinite benchmark is rejected
  upfront (see [Validation](#validation)), since a broken benchmark
  poisons every row's distance for that signal, not just some of
  them.
- `custom_weights`: a **dict** of `{column_name: weight}`, saying
  which columns matter and how much. Insertion order is preserved
  and drives both the `show_progress` readout order and, combined
  with descending weight, the tie-break cascade order (see below).

For each weighted column, `sparpartner` auto-detects the column's
type and scores every row's distance to the benchmark on a 0-1
scale (1.0 = exact match, 0.0 = as far as possible):

| Detected type | How distance is measured |
|---|---|
| **numeric** | `abs(value - bench)`, capped by the column's own max observed distance from bench, or by a fixed cap you supply (see [Fixed caps](#fixed-caps-fixed_caps)) |
| **date** | both sides converted to "age in days" relative to the benchmark date, capped by the column's own max observed age, or by a fixed cap you supply |
| **string** | exact match = 1, anything else = 0. A missing value never counts as a match, see [String scoring, in detail](#string-scoring-in-detail) |

Date detection is automatic. A column is only treated as a date if
its values look date-shaped (contain a separator like `-`, `/`, `.`
or a recognizable month name) **and** parse successfully at least
98% of the time. Bare numeric-looking strings (e.g. `"12345"`) never
even reach the date-parsing attempt, and object-dtype columns of
digit strings fall through to the exact-match string path instead of
being mistaken for numbers. Only a real numeric dtype gets the
numeric path.

### String scoring, in detail

Strings never need manual 1/0 encoding before you hand them to
`sample()`, the string path does that for you automatically. Any
column that isn't numeric, isn't datetime, and doesn't parse as a
date gets scored by exact match against its benchmark:

```python
import pandas as pd
from sparpartner import sample

df = pd.DataFrame({
    "id": [1, 2, 3, 4, 5],
    "country": ["US", "US", "CA", "US", "MX"],
})

bench_marks = {"country": "US"}
custom_weights = {"country": 1}

result, report = sample(
    df,
    bench_marks=bench_marks,
    custom_weights=custom_weights,
    best_first=True,
)

print(report)
# {'spar_country': 60.0, 'spar_min_score': 0.0, 'spar_max_score': 100.0, 'spar_mean_score': 60.0}
```

Under the hood, each row's `country` value is compared to `"US"`
with a plain equality check:

| `country` value | matches bench `"US"`? | per-row score |
|---|---|---|
| `"US"` | yes | `1.0` |
| `"US"` | yes | `1.0` |
| `"CA"` | no | `0.0` |
| `"US"` | yes | `1.0` |
| `"MX"` | no | `0.0` |

That gives 3 exact matches out of 5 rows, an average of `0.6`, which
is exactly the `spar_country: 60.0` you see in the report (scores in
`sparring_report` are always shown on a 0-100 scale, not 0-1). Rows
with a matching `country` sort ahead of rows that don't, same as any
other signal.

This is exact-match only, not fuzzy or partial-similarity matching.
`"Manchester United"` vs `"Man United"` scores `0.0`, the same as
`"Manchester United"` vs `"Real Madrid"`, there's no partial credit
for near-matches the way numeric or date columns get graduated
distance-based scoring.

A missing or null value in a string signal is treated the same way
missing values are treated everywhere else in `sparpartner`, it is
never silently counted as a non-match that happens to score `0.0`
for the "wrong reason." It's flagged as a genuinely missing input,
the same as a missing numeric or date value, so `drop_nan=True`
catches it and drops that row, and `drop_nan=False` scores it `0.0`
for that signal specifically while leaving the row's other signals
untouched, see [Missing values](#missing-values-drop_nan).

Each column's 0-1 score is multiplied by its weight and summed into
one raw score per row. That raw sum is divided by the total weight
to get each row's overall match score, **but only if the total
weight is > 0**. In that normal case the match score always lands in
the 0-1 range, however many signals or weights you used. If the
weights sum to `<= 0`, normalization is skipped entirely and the raw,
unnormalized weighted sum is used instead (not guaranteed to fall in
0-1).

All of this (the per-signal 0-1 scores and the per-row overall
match score) is working state `sample()` uses internally to sort,
tie-break, and slice. **None of it is added as columns to the `df`
you get back.** The df you receive always contains only your
original columns, reordered (and filtered/sliced, if `sparring_n`
and/or `proximity_gates` are set). Score info comes back separately,
as aggregates in `sparring_report`, see
[The sparring report](#the-sparring-report).

The result is always sorted by match score descending, the row
closest to the benchmark is always first, this is also exactly what
`best_first=True` means, see
[Row order (`best_first`)](#row-order-best_first) below.

### Missing values (`drop_nan`)

A missing or unparseable input for a weighted signal is handled
differently depending on `drop_nan`:

- `drop_nan=True` (the default): any row with a missing input for
  even one signal is dropped completely, before that row's total
  match score, sorting, or the `spar_score` gate ever see it.
- `drop_nan=False`: the row is kept. Whichever signal(s) had a
  missing input score `0.0` for that signal specifically, not `NaN`.
  Every other signal on that same row still scores normally from its
  own actual value. The row's total match score is then the usual
  weighted sum of every signal, including the `0.0`s for the missing
  ones, so a row with one missing signal can still end up with a
  solid overall score if its other signals are close to the
  benchmark.

Either way, a missing input is never silently ignored or excluded
from the weight total, it either drops the whole row
(`drop_nan=True`) or counts as the worst possible score for that one
signal while the rest of the row still scores normally
(`drop_nan=False`). This applies the same way across numeric, date,
and string signals, none of the three types has a blind spot where a
missing input goes undetected.

### Tie-breaking

Rows that land on the exact same match score aren't left to random or
arbitrary order. Ties are broken by the per-signal score of the
**highest-weight** signal first (higher wins), then the next-highest,
cascading down the weight-sorted signal list until the tie resolves.
Signals that share the same weight are compared in the order they
appear in `custom_weights` (dict insertion order). Only if every
signal is exhausted and rows are still tied does it fall back to
pandas' stable sort (original row order).

## Row order (`best_first`)

By default (`best_first=True`), the returned `df` is sorted with the
closest match to the benchmark first. Pass `best_first=False` and, as
the very last step before returning, `sample()` flips that same set
of rows so the worst-of-selection is first and the best-of-selection
is last.

This only changes **presentation order**. It never changes which
rows get selected (e.g. via `sparring_n` or `proximity_gates`), never
touches scoring or tie-breaking, and `sparring_report` is identical
either way, since it's an order-independent aggregate.

It exists to replace a manual post-hoc flip like:

```python
result = result.sort_index(ascending=False).reset_index(drop=True)
```

which is fragile against how `sample()`'s own indexing/reset works.
Use `best_first=False` instead when you want the worst-of-selection
row first:

```python
result, report = sample(
    df,
    bench_marks=bench_marks,
    custom_weights=custom_weights,
    best_first=False,
)
```

## The sparring report

`sample()` doesn't just return the ranked frame, it returns a
`(df, sparring_report)` tuple. `sparring_report` is a flat dict
summarizing match quality, with one `spar_<name>` key per signal
plus `spar_min_score` / `spar_max_score` / `spar_mean_score`, e.g.:

```python
{
    "spar_age": 55.0, "spar_income": 57.5,   # one key per signal, avg score
    "spar_min_score": 0.0,
    "spar_max_score": 91.93,
    "spar_mean_score": 56.45,
}
```

Every value is on a 0-100 scale (100 = perfect match to benchmark)
and rounded to 2 decimal places. This is the *only* place score
information comes back to you: individual per-row scores are never
returned, only these aggregates.

A signal with `weight=0` is excluded from ranking/sorting entirely
(it can never move the match score or break a tie), but it still
gets its own `spar_<name>` entry in `sparring_report`, so you can
track how close rows are on a signal without letting that signal
influence which rows are considered "closest".

By default (`sparring_n=None`) the returned `df` and the report
cover every row that made it through scoring, `proximity_gates` (if
set), and `drop_nan`, see
[Proximity gates](#proximity-gates-proximity_gates) below. Pass an
int and `sample()` slices that same sorted, already-gated, already
cleaned result down to just the top `sparring_n` rows (the ones
closest to the benchmark). That slice is what you get back as `df`,
and it's also exactly what `sparring_report` and the score
distribution are computed on. There's no separate "full set" kept
around once `sparring_n` is set; if you need the rest of the rows
too (e.g. to build the train set), take them from your original `df`
yourself, or call `sample()` again with `sparring_n=None`.

### Progress readout (`show_progress`)

When `show_progress=True`, the sparring report section of the
printed readout marks each signal's average score, and the min /
max / mean of the score distribution, with a traffic-light emoji:

- red for an avg score in the bottom third (0-33.3)
- yellow for an avg score in the middle third (33.3-66.7)
- green for an avg score in the top third (66.7-100)

```
    age                  avg score= 64.33  (yellow)
    signup_date          avg score= 38.00  (yellow)
    country              avg score= 50.00  (yellow)

  SCORE DISTRIBUTION
  ------------------
    spar_min_score       =   6.67  (red)
    spar_max_score       =  82.00  (green)
    spar_mean_score      =  50.78  (yellow)
```

This is purely a print-time visual, it doesn't change anything about
`sparring_report`'s actual values.

If `proximity_gates` includes any per-signal gates,
`show_progress=True` prints a `PROXIMITY GATES` block right after the
per-signal scoring and before the `drop_nan` check: one line per gate
with its target and how many rows failed that gate specifically
(counted independently, so a row that fails two gates counts toward
both), plus a single summary line with rows before, dropped, and kept
for the combined filter.

If `proximity_gates` includes the `spar_score` key,
`show_progress=True` also prints a separate `SPAR_SCORE GATE` block
later in the readout, right after the sort step: target, rows before,
rows dropped, and rows kept.

If the gates (combined with `drop_nan`) leave zero rows at any point,
the rest of the readout (top rows, sparring report, best_first flip)
is skipped and a single warning line is printed instead, so you're
not shown a wall of empty output.

## Proximity gates (`proximity_gates`)

`sparring_n` cuts by a fixed count: give me the top 300, regardless of
how each one individually measures up. `proximity_gates` is a
different kind of cut: only keep rows that clear a closeness bar, row
by row, no matter where they'd otherwise land in the overall ranking.

Two independent kinds of gate can live in the same `proximity_gates`
dict:

- **Per-signal gates**, keyed by an existing signal name in
  `custom_weights`. These target that signal's own per-row closeness
  score.
- **The `spar_score` gate**, a single reserved key that targets each
  row's own final total match score instead of any one signal.

Both kinds can be set together in the same call.

### What per-signal gates do

Right after each signal's per-row score is computed, and before
`drop_nan`, before the weighted total is normalized, before sorting,
and before any `sparring_n` slicing, `sample()` checks every row
against every per-signal gate you've defined. A gate is a minimum
closeness bar on one signal's own per-row score. A row is kept only if
it clears every per-signal gate; if it falls short on even one, it's
dropped and never reaches `drop_nan`, the weighted total, sorting, the
`spar_score` gate, `sparring_n`, or the sparring report.

This is a plain, independent per-row filter, not a running or
group-level check. Each row is evaluated against its own scores only,
nothing is accumulated across rows, and the result doesn't depend on
row order or on the df already being sorted.

### What the `spar_score` gate does

The reserved key `"spar_score"` doesn't refer to any single signal, it
refers to the row's own final, normalized total match score, the same
number that drives the sort. Because that total only exists once
every signal has been scored, the weighted sum computed, and
normalization applied, the `spar_score` gate runs later in the
pipeline than the per-signal gates above: right after the sort step,
before `sparring_n` slicing.

A row is kept only if its final total match score is at least the
`spar_score` target. Since the result is always sorted highest score
first, this effectively keeps the top slice of rows that clear a
score bar, rather than a fixed count.

### `proximity_gates`

A dict of `{key: target}`. Most keys must already be a key in
`custom_weights`, targeting that signal's own per-row closeness score,
there's no separate raw-column lookup and no other aggregate alias.
One key is reserved, `"spar_score"`, it targets the row's final total
match score instead, and it does not need to be (and cannot be) a
signal in `custom_weights`. `target` is always a number from 0 to 100
(the same 0-100 scale `sparring_report` uses).

Any key that is neither an existing `custom_weights` signal nor the
reserved `"spar_score"` key raises a `ValueError` during validation,
before any scoring runs, the same way an unmatched `custom_weights`
key does.

Note: `"spar_score"` is always treated as the reserved total-score
gate, even if a signal happens to be literally named `spar_score` in
`custom_weights`. That signal can still be scored and weighted
normally, it just can't also be targeted by its own per-signal gate
under that same key inside `proximity_gates`.

### Example

```python
proximity_gates = {
    "age": 60,          # this row's own age-closeness score must be >= 60
    "income": 50,       # and its income-closeness score must be >= 50
    "spar_score": 70,   # and its final total match score must be >= 70
}

result, report = sample(
    df,
    bench_marks=bench_marks,
    custom_weights=custom_weights,
    best_first=True,
    proximity_gates=proximity_gates,
)
```

Every row in `result` individually scored at least 60 on
age-closeness, at least 50 on income-closeness, and at least 70 on its
own final total match score. `report` reflects only whatever survived
every gate (and, if `sparring_n` is also set, whatever survived the
gates and the slice), not the original ungated pool.

### Requirements and edge cases

- **Per-signal gates require `drop_nan=True`.** Passing a per-signal
  gate together with `drop_nan=False` raises a `ValueError`
  immediately, before scoring starts. The `spar_score` gate on its own
  has no such requirement, it can be used freely with `drop_nan=False`,
  see [Missing values](#missing-values-drop_nan).
- **No floor.** If every row fails at least one gate, the result is an
  empty df. This is intentional, there's no keep at least one row
  fallback.
- **Order matters.** Per-signal gates run first, before `drop_nan` and
  normalization. The `spar_score` gate runs later, after sorting.
  `sparring_n`, if also set, is taken from whatever survives both.
- **Vectorized, not sequential.** Every row is checked against its own
  scores independently, so this scales the same way no matter how
  large the pool is, there's no row-by-row walk that could slow down
  as the pool grows.

## Fixed caps (`fixed_caps`)

For numeric and date signals, the "cap" is the worst-case distance
from benchmark that still counts for something, past that distance a
row bottoms out at a `0.0` score for that signal. By default, that
cap is recalculated fresh on every call, as the largest distance from
benchmark actually observed in the current pool
(`distance.max()`). That's convenient, but it means the same row can
score differently on the same signal depending on what else happens
to be in the pool with it. Score a row alongside a thousand others
today and alongside three others tomorrow, and its `age` score for
the exact same `age` value can come out different both times, because
the "worst case" it's being measured against moved.

`fixed_caps` lets you pin that denominator yourself instead, so a
signal's score reflects a fixed, meaningful worst-case tolerance you
define, not whatever the current pool's own extremes happen to be.

```python
fixed_caps = {
    "age": 40,          # 40+ years from bench = 0 score, regardless of pool
    "income": 200000,   # $200k+ from bench = 0 score, regardless of pool
}

result, report = sample(
    df,
    bench_marks=bench_marks,
    custom_weights=custom_weights,
    best_first=True,
    fixed_caps=fixed_caps,
)
```

A few things worth knowing:

- **Only relevant to numeric and date signals.** String signals score
  by exact match, there's no distance/cap concept for them, so a
  `fixed_caps` entry for a string signal is not meaningful. It's
  still accepted through validation (any signal name valid in
  `custom_weights` is a valid `fixed_caps` key), but only ever affects
  the actual scoring math for a numeric or date signal.
- **Per-signal, opt-in.** Any signal named in `fixed_caps` uses your
  fixed value as its cap. Any signal you leave out falls back to the
  normal pool-relative `distance.max()` behavior, unchanged. You can
  fix some signals and leave others pool-relative in the same call.
- **Values must be finite and strictly greater than 0.** A cap of `0`
  or a negative number has no valid meaning here, distance from
  benchmark is always `>= 0` (via `.abs()`), so a cap `<= 0` either
  forces the degenerate "no variance" fallback (every non-missing row
  scores `1.0` for that signal) or, for a negative cap, flips the
  sign and always clips to a meaningless `1.0` regardless of actual
  distance. This is rejected outright at validation time rather than
  silently producing a flat, uninformative score.
- **Why use it:** mainly for comparability across separate calls,
  scoring one row at a time vs. scoring a thousand at once, or
  comparing this week's run to last week's, should give the same row
  the same score on the same signal, rather than a score that moves
  around depending on what else was in the pool that particular time.
  It also lets you encode a real-world tolerance directly ("more than
  40 years from the benchmark age just doesn't matter") instead of
  letting the pool's own extremes define what "doesn't matter" means.

### `fixed_caps`

An optional dict of `{signal_name: cap_value}`. `None` by default,
meaning every numeric/date signal uses the original pool-relative
cap. Every key must already be a signal in `custom_weights`. Every
value must be a finite real number greater than `0`. Keys are
case-sensitively matched but checked for case-insensitive duplicates
(two keys that only differ by case, e.g. `"Age"` and `"age"`, raise a
`ValueError` rather than silently letting one shadow the other).

### Raises

`TypeError` if:
- `fixed_caps` isn't a dict (when it isn't `None`)
- a `fixed_caps` key isn't a string

`ValueError` if:
- a `fixed_caps` key doesn't match an existing key in `custom_weights`
- a `fixed_caps` value isn't a finite real number
- a `fixed_caps` value is `NaN`, `inf`, or `-inf`
- a `fixed_caps` value is `<= 0`
- two `fixed_caps` keys collide once case is ignored

An empty `fixed_caps` dict (`{}`) is treated the same as `None`, no
signal gets a fixed cap, every signal falls back to the pool-relative
default.

## Generating `custom_weights` with `slicer`

Hand-picking numbers for `custom_weights` (`{"country": 2, "signup_date": 1,
"income": 1}`) works fine for a handful of signals, but it gets
arbitrary fast: why 2 and not 3? why does `income` get the same
weight as `signup_date`? `slicer` exists to replace that guesswork
with a principled cascade, so the weights you hand to `sample()` come
from a deliberate hierarchy of "how much does this signal matter"
instead of numbers picked by feel.

This whole section is optional. `sample()` only ever needs a plain
`custom_weights` dict, however you produce it; `slicer` is a
convenience for building that dict when you don't want to guess the
numbers yourself, not a required step.

### The philosophy (tiered weighting)

`slicer` treats your signals as belonging to tiers, not a flat list,
in a fixed priority order:

- **targets** (the anchor): the primary feature (or set of features)
  the whole weighting is built around, e.g. `country`. Covers the
  anchor broadly, who, what, or where the weighting originates from.
  Highest priority tier whenever it's active.
- **temporal** (the WHEN): recency and/or seasonality, e.g.
  `signup_date`. Optional, second priority.
- **causatives** (the WHY/WHAT): explanatory features that context
  the anchor further, e.g. `income`, `channel`. Optional, lowest
  priority of the three.

Unlike an older, decay-rate-driven version of this cascade, **there
is no `decay` parameter anymore, and no tier "eats into" a fixed
prior pool at a fixed rate.** Instead, the 100-point pool is split
directly across whichever tiers are actually active (`targets`,
`temporal`, `causatives`), using a Fibonacci-weighted proportional
split:

- 1 tier active -> that tier gets the whole pool (100%).
- 2 tiers active -> split evenly, 50% / 50%.
- 3 tiers active -> split `[50%, 25%, 25%]`, in priority order
  (`targets` > `temporal` > `causatives`), so `targets` always gets
  the largest share whenever more than one tier is active, and
  `temporal`/`causatives` split the remaining half evenly between
  them.

This split is fixed by *how many tiers you actually use*, not by any
number you tune. Skip `temporal` entirely and use only `targets` +
`causatives`, and those two split 50/50, exactly as if `temporal` had
never existed, there's no leftover "unused temporal pool" hanging
around.

Within a tier that has more than one item (multiple `targets`, or
`recency`/`season` inside `temporal`, or a `causatives` list), a
second, independent Fibonacci-based split (via the same fib-descending
helper) can further divide that tier's own pool. See each tier's
section below for exactly when that applies.

The result is a flat `{feature_name: weight}` dict, on the same
0-100-ish scale `sample()` expects for `custom_weights`, ready to
pass straight through.

### Usage

```python
from sparpartner import slicer

weights = slicer(
    targets="country",
    temporal={"recency": "last_active_date", "season": "signup_month"},
    causatives=["income", "channel"],
    decay_causatives=True,   # rank-based split: income > channel
)
# {'country': 50.0, 'last_active_date': 16.67, 'signup_month': 8.33,
#  'income': 16.67, 'channel': 8.33}

result, sparring_report = sample(
    df,
    bench_marks=bench_marks,
    custom_weights=weights,
    best_first=True,
)
```

A few things worth knowing about how the tiers behave:

- **`targets` accepts either a single feature name or a list of
  them.** With a single name (a plain `str`, or a one-item list),
  it simply behaves as a single-item tier, same as always. With a
  list of 2 or more names, `pool_targets` must be set explicitly,
  there's no default: `pool_targets=False` gives every name in the
  list the tier's full pool independently (no sharing at all, each
  target stands on its own), `pool_targets=True` treats the tier's
  pool as one shared pool, split evenly across every name in the
  list. There's no rank-based option for `targets`, order never
  matters here.
- **`temporal` has two mutually exclusive modes.** Either pass a
  single `temporal="<feature>"` (it takes the whole temporal pool),
  or pass a **dict**: `temporal={"recency": "<feature>", "season":
  "<feature>"}`. Unlike a plain-arguments approach, the dict doesn't
  require both keys, you can pass just `recency`, just `season`, or
  both. Whichever of `recency`/`season` are actually present get a
  fresh, count-sized Fibonacci-descending split of the temporal
  pool: 1 present -> that one gets 100% of the temporal pool; 2
  present -> `recency` gets ~66.67%, `season` gets ~33.33%
  (`recency` is always the larger share when both are given). Any
  dict key other than `recency`/`season` raises a `ValueError`.
- **`causatives` accepts a single feature name (`str`), a flat list of
  any length, or a nested list for grouped/step mode.** A plain `str`
  is treated the same as a one-item list. Passing a `dict` for
  `causatives` is not supported. There's no cap on how many you can
  pass, `causative_pool` is a fixed quota (this tier's share of the
  100-point pool, per the Fibonacci tier split above) no matter how
  many names split it, more causatives just means each one gets a
  thinner slice of that same quota. How multiple causatives split
  depends on `decay_causatives`, which has no default and must be set
  explicitly whenever `causatives` resolves to 2 or more items (flat
  list form only, see below for the nested form):
  - `False` splits evenly, order doesn't matter.
  - `True` splits by rank using a Fibonacci-descending weight per
    position (e.g. 3 items -> raw weights `[3, 2, 1]`, normalized
    against their own sum), first item in the list gets the most.
  - a positive **int `N`** is a hybrid: only the first `N` causatives
    (by list order) get the same Fibonacci rank-weighting `True`
    would use across the whole list, then whatever's left of
    `causative_pool` after that is split evenly across the rest.
    Handy when you want the top few causatives to dominate but don't
    want a long thinning tail past that. If `N` is greater than or
    equal to how many causatives you passed, this behaves identically
    to `True`, there's no tail left to flatten.
  - a **list of positive ints** (e.g. `[3, 2, 1]`) is step/group mode:
    each int is the size of one group, taken in order off the flat
    `causatives` list. Group height always follows a Fibonacci
    ranking by group *position* (not by group size), e.g. 3 groups
    always ranks `[3, 2, 1]` regardless of how many items sit in each
    group. Every member of a group gets that group's height. The list
    must sum to exactly the number of causatives passed, there's no
    repeating, tiling, or leftover flattened tail here.

  With exactly one causative, `decay_causatives` is ignored entirely,
  there's no rank to decay across a single item, it gets the whole
  causative pool either way.

- **`causatives` can also be passed pre-grouped, as a nested list**,
  which is a cleaner alternative to the `decay_causatives=[ints]` step
  mode above, no separate count list to keep in sync with how many
  causatives you actually passed:

  ```python
  causatives=[['a', 'b'], ['c', 'd', 'e'], 'f']
  # equivalent to causatives=['a','b','c','d','e','f'],
  # decay_causatives=[2, 3, 1]
  ```

  Each top-level item is either a plain feature name (a group of 1)
  or a list of feature names (a group of however many). Group sizes
  come straight off the nesting, so there's nothing to keep in sync
  with the flat count of causatives the way a manual
  `decay_causatives` count list could drift out of. `decay_causatives`
  must be `None` (or left out) whenever `causatives` is passed this
  way, the grouping already came from the nesting, so
  `decay_causatives` has nothing left to say, setting it to anything
  else raises. Group weighting works exactly like the
  `decay_causatives=[ints]` mode above: Fibonacci height by group
  position, not group size.
- **There is no `decay` parameter.** The tier-level split (targets vs
  temporal vs causatives) is fully determined by which tiers are
  active, see the Fibonacci tier split above. Within-tier splits
  (`recency`/`season`, and rank-based `causatives`) use the same
  fixed Fibonacci-descending weighting, not a tunable rate. If you
  want a different split ratio than the fixed ones described here,
  that isn't currently configurable, only which tiers/items you pass
  changes the outcome.
- **`bystanders` are along for the ride only.** Names passed here
  show up in the returned dict at weight `0`, taking no part in any
  split. This is the same shape `sample()` already accepts, a
  `custom_weights` entry with `weight=0`, which per `sample()`'s own
  contract still generates a `spar_<name>` entry in `sparring_report`,
  so you can pass a `slicer` bystander straight into `sample()` to
  track how close rows are on it, without letting it influence which
  rows are ranked closest.
- **`display`** (bool, default `False`) prints a colored tree
  breakdown of the finished weights, grouped by tier
  (TARGETS/TEMPORAL/CAUSATIVES/BYSTANDERS), after every weight has
  already been computed. It's a summary render, not a step-by-step
  progress readout, purely visual, it doesn't change anything about
  the returned dict.
- **Every feature name must be unique across all groups** (`targets`,
  `temporal` (including `recency`/`season` inside it), `causatives`,
  `bystanders`); reusing a name across two groups raises a
  `ValueError`, and comma-joined strings passed instead of a real
  list (`targets`/`causatives`/`bystanders`) are rejected for the
  same reason `sample()` rejects them.
- **Calling `slicer()` with nothing at all raises.** At least one of
  `targets`, `temporal`, `causatives`, or `bystanders` must be
  passed.

### Usage examples

`slicer` can be called a lot of different ways depending on how many
tiers you actually need. A single `targets` on its own is a valid
call, every other tier is optional and only shows up in the result
if you pass it.

**1. Simplest call, targets only**

```python
weights = slicer(targets="country")
# {'country': 100.0}
```

**2. `targets` + `temporal` as one combined feature**

```python
weights = slicer(targets="country", temporal="signup_recency_blend")
# {'country': 50.0, 'signup_recency_blend': 50.0}
# 2 active tiers -> even 50/50 split
```

**3. `targets` + `temporal` dict, `recency`/`season` split**

```python
weights = slicer(
    targets="country",
    temporal={"recency": "days_since_signup", "season": "signup_month"},
)
# {'country': 50.0,
#  'days_since_signup': 33.33,
#  'signup_month': 16.67}
# 2 active tiers (targets, temporal) -> 50/50
# within temporal: recency/season split ~66.67/33.33 of that 50
```

**4. Full cascade: `targets` -> `temporal` -> `causatives`, rank-decayed**

```python
weights = slicer(
    targets="country",
    temporal={"recency": "days_since_signup", "season": "signup_month"},
    causatives=["income", "channel", "referral_source"],
    decay_causatives=True,
)
# {'country': 50.0,
#  'days_since_signup': 16.67, 'signup_month': 8.33,
#  'income': 12.5, 'channel': 8.33, 'referral_source': 4.17}
# 3 active tiers -> [50%, 25%, 25%] (targets, temporal, causatives)
# temporal's 25 splits ~66.67/33.33 -> 16.67 / 8.33
# causatives' 25 splits by fib weights [3,2,1] -> 12.5 / 8.33 / 4.17
```

**5. Same cascade, `causatives` split evenly instead of by rank**

```python
weights = slicer(
    targets="country",
    temporal="signup_recency_blend",
    causatives=["income", "channel"],
    decay_causatives=False,
)
# {'country': 50.0, 'signup_recency_blend': 25.0,
#  'income': 12.5, 'channel': 12.5}
# 3 active tiers -> [50%, 25%, 25%]; causatives' 25 split evenly
```

**5b. Hybrid split, rank-decay the top 2 causatives, flatten the rest**

```python
weights = slicer(
    targets="country",
    temporal="recency_blend",
    causatives=["income", "channel", "referral_source", "device_type", "region"],
    decay_causatives=2,
)
# causative_pool here = 25 (this tier's share of the 100-point pool)
# {'country': 50.0, 'recency_blend': 25.0,
#  'income': 10.53, 'channel': 6.58,
#  'referral_source': 2.63, 'device_type': 2.63, 'region': 2.63}
# income/channel are fib-rank-decayed against each other (weights [8,5]
# out of the 5-item fib set [8,5,3,2,1]); whatever's left of the 25-pool
# after that (~7.89) is split evenly across the remaining 3 causatives
# instead of continuing to decay them into a long, thinning tail
```

**6. A single `causatives` feature, passed as a plain `str`**

```python
weights = slicer(targets="country", causatives="income")
# {'country': 50.0, 'income': 50.0}
# 2 active tiers -> 50/50; decay_causatives isn't required, only one
# item to rank
```

**7. Multi-target, `pool_targets=False`, each target keeps the tier's full pool**

```python
weights = slicer(targets=["home_team", "away_team"], pool_targets=False)
# {'home_team': 100.0, 'away_team': 100.0}
# 1 active tier (targets only) -> tier pool is 100; pool_targets=False
# means each name gets that full 100 independently
```

**8. Multi-target, `pool_targets=True`, the tier's pool is shared evenly**

```python
weights = slicer(targets=["home_team", "away_team"], pool_targets=True)
# {'home_team': 50.0, 'away_team': 50.0}
```

**9. Step decay: `decay_causatives` as a list of ints**

```python
weights = slicer(
    targets="country",
    causatives=["a", "b", "c", "d", "e", "f"],
    decay_causatives=[3, 2, 1],
)
# {'country': 50.0,
#  'a': 21.43, 'b': 21.43, 'c': 21.43,
#  'd': 14.29, 'e': 14.29,
#  'f': 7.14}
# causative_pool here = 50 (2 active tiers -> 50/50)
# 3 groups of size [3, 2, 1] -> Fibonacci height by position [3, 2, 1]
# group 1 (3 members) gets the tallest step, group 3 (1 member) the
# shortest, group size never dilutes the step's own height
```

**9b. The same step decay, expressed with nested `causatives` instead**

```python
weights = slicer(
    targets="country",
    causatives=[["a", "b", "c"], ["d", "e"], "f"],
    # decay_causatives left out entirely -- it must be None here
)
# identical result to example 9 above
```

**10. With `bystanders`, reported at weight 0, no effect on the cascade**

```python
weights = slicer(
    targets="country",
    causatives="income",
    bystanders=["signup_channel", "referral_code"],
)
# {'country': 50.0, 'income': 50.0,
#  'signup_channel': 0, 'referral_code': 0}
```

**11. `display=True`, printing the finished breakdown**

```python
weights = slicer(
    targets="country",
    temporal={"recency": "days_since_signup", "season": "signup_month"},
    causatives=["income", "channel"],
    decay_causatives=True,
    display=True,
)
# prints a colored TARGETS/TEMPORAL/CAUSATIVES tree of the weights
# above, purely a summary render after the fact, doesn't change the
# returned dict at all
```

**A note on `pool_targets`, when it's required and when it's ignored**

`pool_targets` only matters once `targets` is a list of 2 or more
names, that's the only situation where it must be set explicitly
(`True` or `False`, no default). If `targets` is a plain `str`, or a
list with just one name in it, `pool_targets` is ignored entirely
(with a warning printed if you set it anyway), that single name
always takes the tier's full pool regardless of what (or whether)
`pool_targets` is set:

```python
# pool_targets omitted, targets is a single str, no error
slicer(targets="country")
# {'country': 100.0}

# pool_targets omitted, targets is a one-item list, still no error
slicer(targets=["country"])
# {'country': 100.0}

# pool_targets omitted, targets has 2+ names, this raises
slicer(targets=["home_team", "away_team"])
# ValueError: pool_targets must be explicitly set to True or False when
# targets is a list of more than one feature name. there is no default

# pool_targets now provided, works fine
slicer(targets=["home_team", "away_team"], pool_targets=True)
# {'home_team': 50.0, 'away_team': 50.0}
```

### `slicer` parameters

| Name | Type | Default | What it does |
|---|---|---|---|
| `targets` | str, list, or `None` | `None` | The anchor tier. A single name takes the whole tier pool. A list of 2+ names requires `pool_targets` to be set explicitly |
| `pool_targets` | bool or `None` | `None` | Only relevant when `targets` resolves to 2+ names. `True` splits the tier's pool evenly across them; `False` gives each name the tier's full pool independently. Ignored (with a warning) if `targets` resolves to a single name |
| `temporal` | str, dict, or `None` | `None` | A single `str` takes the whole temporal tier pool. A dict `{"recency": ..., "season": ...}` splits that pool by a fixed Fibonacci-descending weighting between whichever of the two keys are present (1 present = 100%, 2 present = ~66.67% / ~33.33%, `recency` always the larger share). Any other dict key raises |
| `causatives` | str, list, nested list, or `None` | `None` | A single `str`, a flat list of feature names, or a nested list (e.g. `[['a','b'], 'c']`) for grouped/step mode. Dict form is not supported. Flat lists of 2+ items require `decay_causatives` to be set explicitly; nested lists require `decay_causatives` to be `None` |
| `decay_causatives` | bool, positive int, list of positive ints, or `None` | `None` | Required whenever `causatives` resolves to 2+ flat items. `True` = Fibonacci rank-decay, biggest share to the first item; `False` = even split; a positive int `N` = rank-decay the first `N` by the same weighting, then split whatever's left evenly across the rest; a list of positive ints (e.g. `[3,2,1]`) = step/group mode, each int a group size, group height by Fibonacci position. Must be `None` when `causatives` is passed in nested/grouped form |
| `bystanders` | list or `None` | `None` | Feature names included in the result at weight `0`, no part in any split |
| `display` | bool | `False` | Prints a colored tree summary of the finished weights, grouped by tier. Pure printing, doesn't affect the returned dict |

### Raises

`TypeError` if:
- `targets` is not a `str`, a list, or `None`
- any `targets` list item is not a `str`
- `temporal` is not a `str`, a dict, or `None`
- a dict `temporal` value (`recency`/`season`) is not a `str`
- `causatives` is not a `str`, a list, or `None`
- any `causatives` list item is not a `str`, or (for grouped/step
  mode) not a list of `str`
- any item inside a `causatives` group is not a `str`
- `bystanders` is not a list or `None`
- `decay_causatives` is not a bool, an int, a list of positive ints,
  or `None`
- `display` is not a bool or `None`
- `pool_targets` is not a bool or `None`

`ValueError` if:
- `targets` is an empty list
- a `targets` list item contains a comma, or is empty/whitespace-only
- `targets` is a list with more than one item but `pool_targets` is
  `None`
- an invalid key is used in the `temporal` dict (anything other than
  `recency`/`season`)
- a dict `temporal` value contains a comma, or is empty/whitespace-only
- a `causatives` or `bystanders` item is not a `str`, contains a
  comma, or is empty/whitespace-only
- a `causatives` group (nested list) is empty
- `causatives` resolves to 2 or more flat items but
  `decay_causatives` is `None`
- `decay_causatives` is an int that is `<= 0`
- `decay_causatives` is set while `causatives` was not passed at all
- `decay_causatives` is a list containing a non-int item
- `decay_causatives` is a list containing a non-positive int
- `decay_causatives` is an empty list
- `decay_causatives` is a list whose items don't sum to the number of
  causatives features
- `decay_causatives` is not `None` while `causatives` was passed in
  nested/grouped form
- the same feature name appears in more than one group
- none of `targets`, `temporal`, `causatives`, or `bystanders` were
  provided at all

### Returns

A flat dict: `{feature_name: weight, ...}`, ready to pass straight
into `sample()`'s `custom_weights`.

## Usage

### Sample usage

`df` is the only argument you can pass positionally. Every other
argument, including `bench_marks`, `custom_weights`, and
`best_first`, must be passed by keyword (see
[Keyword-only arguments](#keyword-only-arguments) below).

```python
import pandas as pd
from sparpartner import sample

df = pd.DataFrame({
    "id": [1, 2, 3, 4, 5],
    "age": [25, 30, 47, 52, 33],
    "signup_date": ["2023-01-15", "2023-03-02", "2022-11-20", "2023-01-10", "2023-06-01"],
    "country": ["US", "US", "CA", "US", "MX"],
})

bench_marks = {
    "age": 30,
    "signup_date": "2023-01-01",
    "country": "US",
}

custom_weights = {
    "age": 2,
    "signup_date": 1,
    "country": 1,
}

result, sparring_report = sample(
    df,
    bench_marks=bench_marks,
    custom_weights=custom_weights,
    best_first=True,      # required, no default, True = best match first; False = worst-of-selection first
    sparring_n=None,      # None = every row scored, sorted, and returned
    drop_nan=True,
    show_progress=True,   # prints the full scoring breakdown
)

print(result)
print(sparring_report)
```

`age=30`, `signup_date="2023-01-01"`, and `country="US"` closely
match row `id=1` (age 25, close date, US), so that row lands at or
near the top of the sorted output (or the bottom, if
`best_first=False`). The `country` column here is the string
exact-match path in action: rows `1`, `2`, and `4` score `1.0`
against bench `"US"`, rows `3` and `5` (`"CA"`, `"MX"`) score `0.0`,
see [String scoring, in detail](#string-scoring-in-detail) above for
the full walkthrough.

```python
# post-sample: turn the ranking into an actual train/test split.
# Use best_first=False: worst-of-selection first, best match (closest to
# benchmark) last. That puts the lookalike rows in one contiguous block
# at the tail, so the split is just a slice off the end, no need to
# track which end is which.
ranked, _ = sample(
    df,
    bench_marks=bench_marks,
    custom_weights=custom_weights,
    best_first=False,
)

# Slice however large you want the test set to be, e.g. the top 30%:
cut = int(len(ranked) * 0.3)
test = ranked.iloc[-cut:]    # lookalikes, the harder, honest test set
train = ranked.iloc[:-cut]   # everything unlike the benchmark
```

### Parameters

| Name | Type | Default | What it does |
|---|---|---|---|
| `df` | DataFrame | required, positional | Must contain a column for every name (key) in `custom_weights`. No restriction on your own column names, `spar`-prefixed columns are fine |
| `bench_marks` | dict | `None`, but required (raises if left `None`) | `{column_name: benchmark_value}`. Every value must be a real, non-null, finite value. Keyword-only |
| `custom_weights` | dict of `{name: weight}` | `None`, but required (raises if left `None`) | Which columns to score and how much each contributes. Keyword-only |
| `best_first` | bool | `None`, but required (raises if left `None`) | Applied last, after everything else (including `sparring_n` slicing and `proximity_gates`). `True` = best match first. `False` = flips that same set of rows to worst-of-selection first, best last. Never changes which rows are selected; see [Row order](#row-order-best_first). Keyword-only |
| `sparring_n` | int or `None` | `None` | `None` = every row scored, sorted, and returned (or, if `proximity_gates` is set, every row that survives the gates). An int slices that same sorted, already-gated result down to the top `sparring_n` rows. That slice is what's returned as `df`, and what the report/distribution are computed on. Runs after `proximity_gates`, not before. Keyword-only |
| `drop_nan` | bool | `True` | If `True`, drops any row with a missing or unparseable input for any weighted signal, after printing (if `show_progress`) a sanity check of what was dropped and why. If `False`, a missing input scores `0.0` for that signal only, every other signal on the row still scores normally, and the row's total match score is computed from whatever signals it did have, see [Missing values](#missing-values-drop_nan). Required to be `True` whenever `proximity_gates` includes a per-signal gate, not required for the `spar_score` gate alone, see [Proximity gates](#proximity-gates-proximity_gates). Keyword-only |
| `proximity_gates` | dict or `None` | `None` | `{key: target}`. Most keys must already be a key in `custom_weights` (that signal's own per-row closeness score, 0-100 scale). One key is reserved, `"spar_score"`, targeting the row's final total match score instead. `target` is the minimum score that must be cleared for the row to survive. Per-signal gates run right after per-signal scoring, before `drop_nan`, normalization, and sorting; the `spar_score` gate runs after sorting, before `sparring_n`. A row is dropped if it fails any gate. See [Proximity gates](#proximity-gates-proximity_gates). Keyword-only |
| `fixed_caps` | dict or `None` | `None` | `{signal_name: cap_value}`. Pins the worst-case distance-from-benchmark denominator for a numeric or date signal, instead of it being recalculated from the current pool's own max distance. Any signal not listed falls back to the pool-relative default. Every value must be a finite number `> 0`. See [Fixed caps](#fixed-caps-fixed_caps). Keyword-only |
| `show_progress` | bool | `False` | Prints a full readout, in run order: header (including a `signals used` count), per-column type/cap/sample scores, a `PROXIMITY GATES` block (if any per-signal gates are set), drop_nan check (if `drop_nan=True`), normalize check, sort-apply readout, a `SPAR_SCORE GATE` block (if the `spar_score` key is set), sparring_n slice readout (if `sparring_n` is set), top N ranked rows with weighted contributions (`N` in the header always matches the number of rows actually shown), the sparring report itself with red/yellow/green markers next to each score flagging any signal contributing zero separation, and finally a best_first flip readout (only printed if `best_first=False` and at least one row remains). If no rows survive `proximity_gates` and/or `drop_nan`, a single warning line is printed instead, and the top rows, sparring report, and best_first flip sections are skipped. Keyword-only |

Note: `sample()`'s progress flag is named `show_progress` because it
narrates a multi-stage pipeline as it runs; `slicer()`'s progress
flag is named `display` because it renders a one-shot summary of
already-finished weights rather than narrating stages. They're
different enough in behavior that they're intentionally named
differently, rather than forced to share one name across the two
functions.

### Returns

`sample()` returns a `(df, sparring_report)` tuple, not just a
DataFrame. The `df` always contains only your original columns
(reordered/filtered/sliced), no score columns are ever attached to
it. See [The sparring report](#the-sparring-report) for how score
information comes back to you instead.

### Validation

Input validation runs upfront, before any scoring starts, in four
passes: general parameter checks, `custom_weights` checks,
`proximity_gates` checks (only if `proximity_gates` is set), then
`fixed_caps` checks (only if `fixed_caps` is set).

Raises `TypeError` if:
- `df` isn't a pandas DataFrame
- `bench_marks` isn't a dict
- `custom_weights` isn't a dict
- `drop_nan`, `show_progress`, or `best_first` isn't a bool
- `sparring_n` isn't an int or `None` (bools are rejected too)
- `proximity_gates` isn't a dict (when it isn't `None`)
- `fixed_caps` isn't a dict (when it isn't `None`)
- a `fixed_caps` key isn't a string

Raises `ValueError` if:
- `bench_marks` is left as `None` (its default)
- `custom_weights` is left as `None` (its default)
- `best_first` is left as `None` (its default)
- `df` has no rows
- `custom_weights` is an empty dict
- `sparring_n` isn't a positive integer
- a `custom_weights` key isn't a string, or doesn't match a column
  in `df`
- a `custom_weights` key has no matching entry in `bench_marks`
- a `bench_marks` value is `None`, `NaN`, `inf`, or `-inf`
- a `custom_weights` key is literally named `"score"`. `sample()`
  keeps its own overall-total working score internally, and a
  signal named `"score"` would generate the exact same internal name,
  corrupting that total instead of just shadowing a per-signal value.
  Rename that column in `df` (and its entries in
  `bench_marks`/`custom_weights`) before calling `sample()`. Names
  that merely *contain* "score", like `test_score` or `score_pct`,
  are unaffected, only an exact match on `"score"` collides
- two or more `custom_weights` signal names would generate colliding
  internal columns (e.g. one signal literally named
  `nan_<other signal name>`)
- a weight isn't numeric (bools are rejected too, a `bool` is
  technically an `int` in Python but was never meant as a weight)
- a weight is `NaN`, `inf`, or `-inf`
- `proximity_gates` is an empty dict
- a per-signal `proximity_gates` gate is set while `drop_nan=False`
  (the reserved `spar_score` gate has no such requirement)
- a `proximity_gates` key isn't a string
- a `proximity_gates` key doesn't match an existing key in
  `custom_weights` and isn't the reserved `spar_score` key
- a `proximity_gates` target isn't a finite real number between 0
  and 100
- a `fixed_caps` key doesn't match an existing key in
  `custom_weights`
- a `fixed_caps` value isn't a finite real number
- a `fixed_caps` value is `NaN`, `inf`, or `-inf`
- a `fixed_caps` value is `<= 0`
- two `fixed_caps` keys collide once case is ignored

**Note on duplicate signal names:** since `custom_weights` is now a
dict, keys are inherently unique, so a repeated column name can no
longer be passed in the first place, Python itself resolves a
repeated key in a dict literal (keeping only the last value) before
`sample()` ever sees it. There's nothing left for validation to
catch here.

## A couple of things worth knowing

- **Your own column names are unrestricted**: `sample()` computes
  its working scores under internally-generated names that can't
  collide with anything you'd realistically name a column, and those
  working columns are always dropped before the df is returned. You
  can freely have your own columns named `spar_score`, `spar_age`,
  or anything else, `sample()` won't touch, rename, or overwrite
  them.
- **Object-dtype numeric strings**: a column of strings like
  `"100"`, `"200"` (object dtype, no separator) is scored as an
  exact-match string column, *not* auto-converted to numeric. Only
  genuine numeric dtypes (`int`, `float`) get the numeric distance
  path.
- **`fixed_caps` doesn't change what counts as a missing input**:
  fixing a signal's cap only changes the denominator used to turn
  distance into a score. Whether a given row's input is treated as
  missing (and therefore dropped, if `drop_nan=True`, or scored
  `0.0` for that signal, if `drop_nan=False`) is decided independently
  of `fixed_caps`, see [Missing values](#missing-values-drop_nan).
- **Keyword-only arguments**: `df` is the only argument `sample()`
  accepts positionally. Every other argument, `bench_marks`,
  `custom_weights`, `best_first`, `sparring_n`, `drop_nan`,
  `proximity_gates`, `fixed_caps`, and `show_progress`, must be
  passed by name. This is enforced by Python itself: a positional
  call like `sample(df, weights, benchmarks)` fails immediately with
  a `TypeError`, before any of `sample()`'s own code runs. It exists
  specifically to rule out accidentally swapping `bench_marks` and
  `custom_weights`, which are both dicts and can't be told apart by
  type alone. `bench_marks`, `custom_weights`, and `best_first`
  additionally default to `None` but are not actually optional,
  leaving any of them out (or passing `None` explicitly) raises a
  `ValueError` naming exactly which one is missing.
