Metadata-Version: 2.5
Name: fossick
Version: 0.1.16
Summary: Web search, fetch, crawl, and browser automation for humans and agents
Project-URL: Repository, https://github.com/vedicreader/fossick
Project-URL: Documentation, https://vedicreader.github.io/fossick/
Author-email: Karthik <karthik.rajgopal@hotmail.com>
License: Apache-2.0
License-File: LICENSE
Keywords: nbdev
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Requires-Python: >=3.12
Requires-Dist: certifi>=2026.7.22
Requires-Dist: curl-cffi>=0.15.0
Requires-Dist: ddgs>=9.14.4
Requires-Dist: diskcache>=5.6.3
Requires-Dist: fastcdp>=0.0.5
Requires-Dist: fastcore>=1.12.31
Requires-Dist: html2text>=2025.4.15
Requires-Dist: liteparse>=2.1.1
Requires-Dist: mcp<2,>=1.2.0
Requires-Dist: pdf-oxide>=0.3.67
Requires-Dist: readability-lxml>=0.8.4.1
Requires-Dist: scrapling[fetchers]>=0.4.8
Requires-Dist: yt-dlp>=2026.3.17
Description-Content-Type: text/markdown

# fossick


<!-- WARNING: THIS FILE WAS AUTOGENERATED! DO NOT EDIT! -->

[`search`](https://vedicreader.github.io/fossick/cli.html#search) finds sources. [`read`](https://vedicreader.github.io/fossick/core.html#read) turns a supported source into markdown with one result shape.
[`fetch`](https://vedicreader.github.io/fossick/cli.html#fetch) and [`to_md`](https://vedicreader.github.io/fossick/core.html#to_md) handle web pages, from plain HTTP through browser-backed requests.
[`cdp_connect`](https://vedicreader.github.io/fossick/cdp.html#cdp_connect) drives Chrome for authenticated or multi-step work. Dedicated readers handle
YouTube, arXiv, GitHub, and PDFs.

## Install

``` sh
uv add fossick
```

Text/image/news search need no Docker. JS rendering, stealth fetch, and [`google()`](https://vedicreader.github.io/fossick/search.html#google) use a bundled headless browser.

## Quick start

``` python
results = search('koshas github fts semantic code graph', method='flashrank', n=5)
for r in results: print(r['score'], r['engines'], r['title'], r['href'])
```

    0.030536 ['brave', 'yahoo'] github.com › codegraph-ai › CodeGraphGitHub - codegraph-ai/CodeGraph: CodeGraph builds a semantic... https://github.com/codegraph-ai/CodeGraph
    0.016129 ['yahoo'] vedicreader.github.io › kosha › graphgraph – koshas - vedicreader.github.io https://vedicreader.github.io/kosha/graph.html
    0.015873 ['brave'] fts · GitHub Topics · GitHub https://github.com/topics/fts?o=desc&s=updated
    0.016393 ['yahoo'] pypi.org › project › koshaskoshas · PyPI https://pypi.org/project/koshas/
    0.015625 ['brave'] GitHub - colbymchenry/codegraph: Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, and Hermes Agent — fewer tokens, fewer tool calls, 100% local https://github.com/colbymchenry/codegraph

``` python
res = research('sqlite WAL mode vs journal mode', n=5)
print(res['digest'][:500])
print([s['href'] for s in res['sources']])
```

    ## Write-Ahead Logging
    https://www.sqlite.org/wal.html

    Write-Ahead Logging

    Table Of Contents

    # 1\. Overview

    The default method by which SQLite implements atomic commit and rollback is a rollback journal. Beginning with version 3.7.0 (2010-07-21), a new "Write-Ahead Log" option (hereafter referred to as "WAL") is available.

    There are advantages and disadvantages to using WAL instead of a rollback journal. Advantages include:

    1. WAL is significantly faster in most scenarios. 
      2. WAL provid
    ['https://www.sqlite.org/wal.html', 'https://blog.sqlite.ai/journal-modes-in-sqlite', 'https://mohit-bhalla.medium.com/understanding-wal-mode-in-sqlite-boosting-performance-in-sql-crud-operations-for-ios-5a8bd8be93d2', 'https://til.simonwillison.net/sqlite/enabling-wal-mode', 'https://fly.io/blog/wal-mode-in-litefs/']

``` python
page = read('https://en.wikipedia.org/wiki/Web_scraping')
print(page.text[:400])
```

    The legality of web scraping varies across the world. In general, web scraping may be against the terms of service of some websites, but the enforceability of these terms is unclear.[11]

    In the United States, website owners can use three major legal claims to prevent undesired web scraping: (1) copyright infringement (compilation), (2) violation of the Computer Fraud and Abuse Act ("CFAA"), and (

## One door

[`what_is`](https://vedicreader.github.io/fossick/core.html#what_is) classifies a target as `dir`, `file`, `arxiv`, `youtube`, `github`, `ghfile`, `pdf`,
or `web`. [`read`](https://vedicreader.github.io/fossick/core.html#read) selects that reader and returns `ok`, `kind`, `title`, `source`, `text`,
`skipped`, and `meta`.

``` python
target = str(repo_root()/'README.md')
assert what_is(target) == 'file'
result = read(target)
assert sorted(result) == ['kind', 'meta', 'ok', 'skipped', 'source', 'text', 'title']
assert (result.ok, result.kind, result.title) == (True, 'file', 'fossick')
assert '# fossick' in result.text and result.skipped is None
result.kind, result.title, result.ok
```

Directories and GitHub repositories are trees. Their result has a local path in `meta` and no
`text`. Other successful readers return markdown in `text`. For PDFs, `pages=True` returns
`(page_number, text)` pairs for page-level citations.

A recognized target can still fail to yield content. Bot walls, missing YouTube transcripts, and
invalid PDF responses set `ok=False` and put the reason in `skipped`. Unsupported targets raise
`ValueError` during classification.

## Modules

| notebook | for |
|----|----|
| `00_core` | [`fetch`](https://vedicreader.github.io/fossick/cli.html#fetch), [`to_md`](https://vedicreader.github.io/fossick/core.html#to_md), [`crawl`](https://vedicreader.github.io/fossick/cli.html#crawl), readers (yt/arxiv/gh), hidden APIs |
| `01_cdp` | Chrome over DevTools: `snapshot`, `fill_form`, `act` |
| `02_search` | metasearch, RRF/BM25/flashrank, [`google`](https://vedicreader.github.io/fossick/search.html#google), [`research`](https://vedicreader.github.io/fossick/cli.html#research) |
| `03_cli` | `fossick` CLI |
| `04_mcp` | MCP server for agents |
| `05_shop` | cart/checkout automation on a live page |
| `06_quality` | [`plan`](https://vedicreader.github.io/fossick/quality.html#plan), [`curate`](https://vedicreader.github.io/fossick/quality.html#curate), authority, diversity |

## One-liners

``` python
search('query', n=10, region='auto', timelimit='y')
research('query', n=5)                    # search, read, and build a cited digest
fetch(url, auto=True)                     # escalate past bot walls
crawl(url, max_pages=5, same_domain=True)
google('query', n=10)                     # real Google via stealth browser
cdp_connect(); pg = await cdp.new_page(url); print(await pg.snapshot())
```

CLI: `fossick search "q"`, `fossick fetch URL --auto`, `fossick research "q"`.
MCP: `uvx fossick-mcp` (or `fossick-mcp --http`).

## Debug Chrome

Use a persistent Chrome profile for logged-in fetches (`session=True`) and CDP:

``` sh
fossick cdp-install   # launch agent and keep-alive
```

Cookies survive restarts. Log in once by hand.
