Metadata-Version: 2.4
Name: remove-paywall-mcp
Version: 1.3.0
Summary: MCP server that removes article paywalls via internet archives
Author: RemovePaywall
License-Expression: MIT
Project-URL: Repository, https://github.com/jasval/remove-paywall-mcp
Keywords: mcp,paywall,archive,wayback-machine,claude,opencode
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Internet :: WWW/HTTP
Classifier: Topic :: Utilities
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: mcp[cli]>=1.0.0
Requires-Dist: httpx>=0.27.0
Requires-Dist: beautifulsoup4>=4.12.0
Requires-Dist: readability-lxml>=0.8.0
Requires-Dist: aiosqlite>=0.20.0
Dynamic: license-file

# remove-paywall-mcp

MCP server that removes article paywalls by searching internet archives. Give it a URL, get back the article text.

## How it works

1. You give it a paywalled article URL
2. Tracking params are stripped and the URL is normalized
3. It searches internet archives in parallel (Wayback Machine CDX API, archive.is mirrors, Wayback Availability API) for an archived copy — archives don't have paywalls because they were crawled without login gates
4. It extracts the article body with [readability-lxml](https://github.com/buriy/python-readability), stripping navigation, ads, and sidebar cruft
5. Post-extraction check: if the result still contains paywall text (e.g., the snapshot captured the paywall itself), it retries with the next archive
6. It returns clean text with the title and snapshot URL

It also learns from every attempt — success rates per domain per archive source are tracked in a local SQLite database, and archive search order is re-ranked automatically using Laplace smoothing.

## Install

```bash
# zero-install (recommended — works everywhere uvx is available)
uvx remove-paywall-mcp

# from PyPI
pip install remove-paywall-mcp

# from source
pip install git+https://github.com/jasval/remove-paywall-mcp.git

# Docker
docker run -i --rm remove-paywall-mcp
docker compose up -d  # HTTP mode on port 8000
```

## Platform configs

Once installed, add this to your MCP client config:

### OpenCode

```json
{
  "mcp": {
    "remove-paywall": {
      "type": "local",
      "command": ["uvx", "remove-paywall-mcp"],
      "enabled": true
    }
  }
}
```

### Claude Desktop

```json
{
  "mcpServers": {
    "remove-paywall": {
      "command": "uvx",
      "args": ["remove-paywall-mcp"]
    }
  }
}
```

### LiteLLM

```yaml
mcp_tools:
  remove_paywall:
    type: "stdio"
    command: "uvx"
    args: ["remove-paywall-mcp"]
```

### Docker (any client)

```json
{"command": "docker", "args": ["run", "-i", "--rm", "remove-paywall-mcp"]}
```

## Tools

### `remove_paywall`

Main tool. Removes a paywall from an article URL and returns clean article text.

| Parameter | Type | Description |
|-----------|------|-------------|
| `url` | string | The paywalled article URL |

### `search_archives`

Search all archive sources for snapshots without extracting content. Useful to see what's available.

| Parameter | Type | Description |
|-----------|------|-------------|
| `url` | string | The article URL to search for |

### `get_from_archive`

Fetch from a specific archive source.

| Parameter | Type | Description |
|-----------|------|-------------|
| `url` | string | The article URL |
| `source` | string | `wayback`, `archive_is`, or `wayback_available` |

### `domain_info`

Look up a domain in the knowledge base — paywall status, notes, and per-archive success rates.

| Parameter | Type | Description |
|-----------|------|-------------|
| `domain` | string | Domain name (e.g. `nytimes.com`) |

### `add_domain`

Register a domain in the knowledge base. Mark paywalled domains so archives are searched first, or non-paywalled domains so the live page is fetched directly.

| Parameter | Type | Description |
|-----------|------|-------------|
| `domain` | string | Domain name |
| `has_paywall` | boolean | `true` if the site has a paywall |
| `notes` | string? | Optional description |

## Prompts

The server provides 3 prompt templates for LLMs to use the tools effectively.

### `remove_paywall_prompt`

Full instruct for bypassing a specific URL. Tells the assistant to use `remove_paywall`, fall back to `search_archives`, and check `domain_info`.

| Parameter | Type | Description |
|-----------|------|-------------|
| `url` | string | The paywalled article URL |

### `bypass_paywall`

Short alias — just tells the assistant to call `remove_paywall` on the URL.

| Parameter | Type | Description |
|-----------|------|-------------|
| `url` | string | The paywalled article URL |

### `handle_paywalls`

System prompt fragment. No arguments — returns instructions for the assistant to automatically call `remove_paywall` whenever it encounters a paywall, login wall, or metered content. Paste this into your system prompt or load it as a prompt at session start.

## Domain knowledge base

Seeded with 30 well-known paywalled domains (NYT, WSJ, Bloomberg, Medium, etc.), stored in SQLite at `~/.remove-paywall-mcp/domains.db`. Tracks every archive success/failure per domain and re-ranks archive search order automatically — domains where archive.is consistently fails won't waste time on it.

### Env vars

| Variable | Default | Description |
|----------|---------|-------------|
| `MCP_TRANSPORT` | `stdio` | `stdio` or `streamable-http` |
| `MCP_HOST` | `0.0.0.0` | Bind address (HTTP mode) |
| `MCP_PORT` | `8000` | Port (HTTP mode) |
| `MCP_DB_DIR` | `~/.remove-paywall-mcp` | Database directory |

## Archive sources

| Source | Priority | Notes |
|--------|----------|-------|
| Wayback Machine | 1 | CDX API, newest-first (`limit=-5`), dedup via `collapse=digest`, HTML-only |
| archive.is mirrors | 2 | Tries newest/oldest across archive.is, archive.today, archive.ph, archive.md |
| Wayback Availability | 3 | `archive.org/wayback/available` — single closest snapshot as fast fallback |

Priority is dynamically re-ranked per domain based on historical success rates recorded in the knowledge base.

## Architecture

```
MCP client (Claude/OpenCode/LiteLLM)
       │  stdio or HTTP
       ▼
┌─────────────────┐
│    server.py     │  MCPServer with 5 tools
│  +3 prompts       │
└────────┬────────┘
         │
    ┌────┴────┐
    ▼         ▼
┌────────┐ ┌──────────┐
│archives│ │ extractors│
│  .py   │ │   .py     │
│        │ │           │
│ wayback│ │readability│
│ archive│ │Beautiful  │
│ .is    │ │Soup       │
│ wayback│ │           │
│ avail  │ │           │
└───┬────┘ └──────────┘
    │
    ▼
┌──────────────┐
│domain_store  │
│   .py        │
│              │
│ SQLite knows │
│ which domains│
│ have paywalls│
│ and which    │
│ archives work│
│ best (Laplace│
│ smoothed)    │
└──────────────┘
```

## License

MIT
