Metadata-Version: 2.5
Name: browsergym-knows
Version: 1.2.0
Summary: KNOWS: a Google Workspace (Docs, Sheets, Slides) benchmark for web agents, packaged for BrowserGym
Project-URL: Homepage, https://alexgill321.github.io/KNOWS-benchmark
Project-URL: Repository, https://github.com/alexgill321/KNOWS-benchmark
Project-URL: Issues, https://github.com/alexgill321/KNOWS-benchmark/issues
Project-URL: Dataset, https://huggingface.co/datasets/utahnlp/knows-benchmark
Author: Alexander Gill, Md Farhan Ishmam
License: Apache-2.0
License-File: LICENSE
Keywords: benchmark,browsergym,google-workspace,llm-evaluation,web-agents
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Requires-Dist: arxiv
Requires-Dist: beautifulsoup4
Requires-Dist: browsergym-core
Requires-Dist: curl-cffi
Requires-Dist: google-api-python-client>=2.100.0
Requires-Dist: google-auth-httplib2
Requires-Dist: google-auth-oauthlib>=1.0.0
Requires-Dist: google-auth<3.0.0,>=2.23.0
Requires-Dist: google-cloud-secret-manager>=2.16.0
Requires-Dist: google-genai>=1.0.0
Requires-Dist: html2text
Requires-Dist: huggingface-hub
Requires-Dist: imagehash
Requires-Dist: numpy
Requires-Dist: opencv-python-headless
Requires-Dist: pandas
Requires-Dist: pillow
Requires-Dist: playwright
Requires-Dist: pymupdf
Requires-Dist: rapidfuzz
Requires-Dist: reportlab
Requires-Dist: requests
Requires-Dist: svglib
Requires-Dist: tabulate
Provides-Extra: local-models
Requires-Dist: accelerate; extra == 'local-models'
Requires-Dist: llama-cpp-python; extra == 'local-models'
Requires-Dist: qwen-vl-utils; extra == 'local-models'
Requires-Dist: torch; extra == 'local-models'
Requires-Dist: transformers; extra == 'local-models'
Description-Content-Type: text/markdown

# KNOWS Benchmark

KNOWS is a benchmark for evaluating web agents on realistic, open-ended **Google Workspace** tasks: writing documents, building spreadsheets, and composing slide decks that require web research, multi-step tool use, and faithful grounding in retrieved sources.

- **22 task templates × 5 instances = 110 tasks** across Google Docs (5 templates), Sheets (9), and Slides (8).
- **Hybrid, white-box evaluation**: each task ships a programmatic evaluator (`evaluator.py`) that scores the produced artifact step by step, combining deterministic checks, fuzzy/tolerance matching, geometric layout tests, document-structure tests, browsing-trace checks, and LLM/VLM judgements.
- **Step-level failure categories**: every evaluation step records the mechanism that decided its outcome (`StepCategory` in `src/browsergym/knows/eval/eval_utils/scoring.py`), enabling quantitative failure-mode analysis of any run via `Result.get_category_summary()`.

**Project page:** [alexgill321.github.io/KNOWS-benchmark](https://alexgill321.github.io/KNOWS-benchmark) · **Dataset:** [`utahnlp/knows-benchmark`](https://huggingface.co/datasets/utahnlp/knows-benchmark) on Hugging Face

The Hugging Face dataset carries the prompts and the structured evaluation rubric as data, for analysis or task selection. It does not score anything — the evaluators live here.

## Repository layout

```
src/browsergym/knows/
  __init__.py            # BrowserGym task registration (knows.<family>.<n>)
  task.py                # BrowserGym task classes (setup, prompting, grading glue)
  doc_setup.py           # Workspace provisioning & Drive-sharing helpers (CLI included)
  eval/eval_utils/       # Shared evaluation utilities (scoring, text/image/table/chart, LLM judge)
  eval/tasks/<family>/   # One directory per task template
    utils.py             #   template-shared evaluation helpers (where present)
    instance_N/          #   5 instances per template
      task.md            #     the agent prompt (verbatim)
      checkpoints.md     #     human-readable evaluation criteria
      evaluator.py       #     the automated evaluator (standalone CLI)
      data/              #     gold reference assets used by the evaluator
analysis/                # Scripts + data to reproduce the paper's dataset statistics
RUNNING_STANDALONE.md    # Run tasks on ANY agent harness (Comet, proprietary, ...)
ASSETS.md                # External Drive assets some tasks depend on, and how to host your own copies
```

## Quickstart

### Option A — Run with the reference harnesses

The maintained agent harnesses live in companion repos:

- **[BrowserGym-Knows](https://github.com/farhanishmam/BrowserGym-Knows)** — a BrowserGym fork with the `knows` backend, benchmark splits (`knows_docs_1`, `knows_sheets_7`, ...), runner scripts, and environment setup. **Start here**; this repository is consumed as its `browsergym/knows` submodule.
- **[AgentLab-Knows](https://github.com/farhanishmam/AgentLab-Knows)** — an AgentLab fork with agent configurations used in the paper. It requires the BrowserGym-Knows setup; see its README.

```bash
git clone https://github.com/farhanishmam/BrowserGym-Knows
cd BrowserGym-Knows
git submodule update --init --recursive   # pulls this repo into browsergym/knows/
# then follow BrowserGym-Knows' README for installation and runs
```

### Option B — Run on your own harness

Every evaluator is a standalone script: give your agent the `task.md` prompt plus a Google file it can edit, record the URLs it visits, then run the evaluator against the resulting file ID. **See [RUNNING_STANDALONE.md](RUNNING_STANDALONE.md)** for the complete protocol (provisioning, sharing, prompting, history capture, grading, and output parsing).

### Option C — Install as a package

```bash
pip install browsergym-knows          # Python >= 3.10
python -m playwright install chromium
```

This installs the evaluators and registers all 110 tasks as BrowserGym environments (`browsergym/knows.<family>.<instance>`); it is what the upstream BrowserGym `knows` backend installs. The ~210 MB of gold evaluation data is not in the wheel: it is downloaded once from this repository's GitHub release on first use into `~/.cache/browsergym-knows/` (override the location with `KNOWS_DATA_DIR`), or prefetch it explicitly:

```bash
python -c "import browsergym.knows; browsergym.knows.ensure_gold_data()"
```

Add the `[local-models]` extra to run local judge models instead of the Gemini API. The Setup section below applies to all three options.

## Setup (required for all options)

Evaluators read the agent's artifact through the Google Workspace APIs and use a Gemini model as the LLM/VLM judge.

1. **Google Cloud project** with the **Drive, Docs, Sheets, and Slides APIs** enabled.
2. **Evaluator credentials** — one of:
   - *Service account (recommended)*: create a service account, download its JSON key to `auth-data/service-account.json` (or set `SERVICE_ACCOUNT_PATH`). Every graded file must be **shared with the service account's email** (see RUNNING_STANDALONE.md — the harnesses do this automatically).
   - *OAuth*: place an OAuth desktop-client `credentials.json` in `auth-data/`; the first run opens a browser consent flow and caches `auth-data/token.json`.
3. **Judge model key**: create a [Google AI Studio](https://aistudio.google.com/) API key and export `GOOGLE_AI_API_KEY` (or configure Vertex AI application-default credentials with `GOOGLE_CLOUD_PROJECT`).
4. `cp .env.example .env` and fill in the values (some task families use additional optional keys).
5. Install dependencies:

```bash
pip install -r requirements.txt
python -m playwright install chromium   # only needed for harness runs / doc_setup provisioning
```

`auth-data/` is git-ignored — never commit credentials.

### 6. Provision write targets (required for two task families)

Most tasks only *read* from Drive. Two families ask the agent to **write** into it:

| Family | What the agent writes |
|---|---|
| `sheets_10_paper_sorting` | Uploads paper PDFs and Figure 1 screenshots |
| `slides_17_removeimagesaddplaceholders` | Saves images extracted from a deck, and edits a copy of it |

A write destination can't be shared between users, so their prompts ship with
`{{PLACEHOLDER}}` tokens instead of URLs. Run this once before each benchmark pass to create the
folders in **your** Drive and fill the tokens in:

```bash
python src/browsergym/knows/eval/tasks/provision_run_targets.py --all
```

This creates a fresh `run_NNNN` per instance, so one pass never sees another's uploads:

```
KNOWS-runs/                                    <- created in your My Drive
  sheets_10_paper_sorting/
    instance_1/run_0001/
      pdfs/                                    <- {{OUTPUT_FOLDER_URL}} points here
      figures/                                    (the prompt names both subfolders)
  slides_17_removeimagesaddplaceholders/
    instance_1/run_0001/
      images/                                  <- {{IMAGES_FOLDER_URL}}
      KNOWS ... run_0001                       <- {{WORKING_COPY_URL}} (a deck copy)
    instance_2/run_0001/
      images/                                  <- {{IMAGES_FOLDER_URL}}
      copies/                                  <- {{OUTPUT_FOLDER_URL}}
```

Useful flags: `--family <name>` / `--instance <N>` to do a subset, `--parent_folder_id <id>` to
build somewhere other than a new `KNOWS-runs` folder, and `--auth oauth|service` to choose
credentials.

> **Use OAuth for `slides_17`.** Instance 1 needs a *copy* of a presentation, and service accounts
> have no Drive storage quota, so they cannot own files. Folder-only provisioning
> (`sheets_10`, `slides_17` instances 2–5) works with either credential.

The resulting IDs are written to `run_targets.json` in the repository root. The harness substitutes
them into the prompt at episode start and the evaluators read the same file when grading, so there
is nothing to edit by hand. It is git-ignored — it holds locations only you can write to.

If a task runs without provisioning, it fails immediately with the exact command to fix it. Set
`KNOWS_SKIP_PROVISION=1` to reuse the existing folders (e.g. when re-grading a finished run) and
`KNOWS_RUN_TARGETS=/path/to/file.json` to keep the config elsewhere.

## External task assets

Six task families reference source documents/folders on Google Drive from their prompts. Those
sources are hosted view-only and work as-is; the two families above additionally need the write
targets described in step 6. Full details — including how to rehost the sources from the released
assets bundle if a link ever breaks — are in [ASSETS.md](ASSETS.md).

## Reproducing paper analyses

`analysis/` contains the dataset-statistics script (`analyze_task_instances.py`) and the step-type taxonomy labels (`analysis/taxonomy/`). Step-level failure categories are produced natively by the evaluators (`Result.get_category_summary()`; see `eval/eval_utils/scoring.py`).

## Maintenance, Versioning, and Issue Reporting

KNOWS tasks are curated to be time-agnostic and independent of any single website, but some tasks reference live web pages and Drive-hosted assets that can change over time. Our maintenance policy:

- **Periodic freshness checks.** We maintain an internal list of the external URLs and website-extracted gold data our evaluators rely on, and manually verify them at periodic intervals after release (automated checks will replace manual ones once validated against them).
- **Updates and deprecation.** If a website change or outage significantly alters a task's difficulty or solvability, we will update the affected task or replace it with a new task of similar complexity and scope.
- **Versioning.** Every change to tasks or evaluators is released as a new tagged version on GitHub. Prior versions remain permanently available via their release tags, so results reported against any version stay interpretable and reproducible. Always report the benchmark version (release tag) alongside your results.

### Reporting outdated tasks or evaluator issues

Found a task whose referenced website changed, a dead link, or an evaluator that scores incorrectly? Please [open a GitHub issue](../../issues/new/choose) using the provided templates. Include:

1. the task and instance (e.g. `sheets_7_running_analysis/instance_3`) and the benchmark version (release tag);
2. what you observed vs. what you expected (for evaluator issues: the step name and the evaluator's printed output — scores, step details, and category summary);
3. for outdated-content reports: the affected URL and what changed;
4. if relevant and shareable: a link to the graded document (shared as view-only).

We triage reports against the policy above; fixes ship as new tagged releases.

## Citation

If you use KNOWS, please cite:

```bibtex
@inproceedings{gill2026knows,
  title     = {The Hard Part Comes After Search: Benchmarking Web Agents on
               Synthesizing, Organizing, and Displaying Knowledge},
  author    = {Gill, Alexander and Ishmam, Md Farhan and Nguyen, Xuyen and
               Bhat, Neha and DeYoung, Parker Henry and
               Hashemi Chaleshtori, Fateme and Stringham, Nathan and
               Marino, Kenneth and Marasovi\'{c}, Ana},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026}
}
```

Please also state which benchmark version you ran (see [Maintenance, Versioning, and Issue Reporting](#maintenance-versioning-and-issue-reporting)).

## License

Apache License 2.0 — see [LICENSE](LICENSE).
