Metadata-Version: 2.4
Name: surrogate-index
Version: 0.2.0
Summary: Efficient-Influence-Function (EIF) utilities for surrogate-index causal inference.
Project-URL: Homepage, https://github.com/kideokkwon/surrogate-index
Project-URL: Documentation, https://github.com/kideokkwon/surrogate-index#readme
Project-URL: Issues, https://github.com/kideokkwon/surrogate-index/issues
Project-URL: Source, https://github.com/kideokkwon/surrogate-index
Author-email: Kideok Kwon <kideokk16@gmail.com>
License: MIT
License-File: LICENSE.txt
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Topic :: Scientific/Engineering :: Mathematics
Requires-Python: >=3.10
Requires-Dist: numpy>=1.26
Requires-Dist: pandas>=2.2
Requires-Dist: scikit-learn>=1.4
Provides-Extra: dev
Requires-Dist: black>=24.3; extra == 'dev'
Requires-Dist: mypy>=1.9; extra == 'dev'
Requires-Dist: pandas-stubs>=2.1.1.230316; extra == 'dev'
Requires-Dist: pytest>=8.1.1; extra == 'dev'
Requires-Dist: ruff>=0.4; extra == 'dev'
Provides-Extra: ml
Requires-Dist: xgboost>=2.0; extra == 'ml'
Description-Content-Type: text/markdown

# surrogate-index

[![PyPI version](https://img.shields.io/pypi/v/surrogate-index.svg)](https://pypi.org/project/surrogate-index/)

## Introduction

This package provides an implementation of the **Surrogate Index Estimator** introduced by [Athey et al. (2016)](https://arxiv.org/pdf/1603.09326), a causal inference method for estimating long-term treatment effects using short-term randomized controlled trials (e.g., A/B tests).

The core idea is to **combine a randomized experimental dataset with an external observational dataset** to estimate the **Average Treatment Effect (ATE)** on a long-term outcome that is not directly observed in the experiment (e.g., annual revenue, long-term retention). This is particularly useful in settings where long-term metrics are delayed, costly, or infeasible to measure during the experiment window.

This package implements an estimator based on the **Efficient Influence Function (EIF)** derived by [Chen & Ritzwoller (2023)](https://arxiv.org/pdf/2107.14405), leveraging the **Double/Debiased Machine Learning (DML)** framework of [Chernozhukov et al. (2016)](https://arxiv.org/abs/1608.00060). EIF-based estimators enable valid inference while incorporating flexible machine learning models for nuisance components, such as short-term outcome regressions and propensity scores, without compromising asymptotic efficiency or introducing first-order bias.

## Brief Mathematical Background

Given the terms:
- $w\in\\{0,1\\}$: binary treatment indicator 
- $s$: a vector of an arbitrary number of short-term outcomes (typically used as the "metrics of interest" in an A/B Test)
- $x$: a vector of pre-treatment covariates.
- $y$: long-term outcome
- $g$: binary indicator for if the user is in the observational sample ($g=1$) or the experimental sample ($g=0$)

the corresponding influence function for the ATE $\tau_0$ is as follows: 

$$\xi_0(b,\tau_0,\varphi)=\frac{g}{1-\pi}\left[\frac{1-\gamma(s,x)}{\gamma(s,x)}\cdot\frac{(\varrho(s,x)-\varrho(x))(y-\nu(s,x))}{\varrho(x)(1-\varrho(x))}\right]+\frac{1-g}{1-\pi}\left[\frac{w(\nu(s,x)-\bar\nu_1(x))}{\varrho(x)}-\frac{(1-w)(\nu(s,x)-\bar\nu_0(x))}{1-\varrho(x)}+(\bar\nu_1(x)-\bar\nu_0(x))-\tau_0\right]$$

where:
- $\nu(s,x)=E[Y|S,X,G=1]$
- $\varrho(s,x)=P(W=1|S,X,G=0)$
- $\varrho(x)=P(W=1|X,G=0)$
- $\gamma(s,x)=P(G=1|S,X)$
- $\pi=P(G=1)$
- $\bar\nu_w(x)=E[\nu(S,X)|W=w, X,G=0]$
---
## Table of Contents
- [Installation](#installation)
- [Usage](#usage)
- [Planned Features](#planned-features)
- [License](#license)

---

## Installation

```bash
# simplest
pip install surrogate-index

# with ML extras (e.g. XGBoost)
pip install "surrogate-index[ml]"

# Conda users
conda install -c conda-forge xgboost scikit-learn pandas numpy
pip install surrogate-index
```
## Usage

```python
from surrogate_index import efficient_influence_function

df_exp = ...  # experimental sample
df_obs = ...  # observational sample

results_df = efficient_influence_function(
    df_exp=df_exp,
    df_obs=df_obs,
    y="six_month_revenue",
    w="treatment",
    s_cols=[...],   # list of surrogate metrics
    x_cols=[...],   # list of covariate names
    classifier=..., # e.g., GradientBoostingClassifier()
    regressor=...,  # e.g., XGBRegressor()
)
print(results_df)
```

## Planned Features
- Convert structure to an Object-based one (scikit-learn style)
- Add diagnostic checks
- Add alternative estimators provided in Athey et al. 2016
- etc.

## License

Distributed under the MIT License. See `LICENSE` for details.