Metadata-Version: 2.4
Name: urlps
Version: 0.7.0
Summary: Lightweight URL parsing and building helpers (RFC 3986-like).
Author: urlps maintainers
License-Expression: MIT
Project-URL: Homepage, https://github.com/i3iorn/urlps
Project-URL: Repository, https://github.com/i3iorn/urlps
Project-URL: Documentation, https://github.com/i3iorn/urlps#readme
Project-URL: Changelog, https://github.com/i3iorn/urlps/blob/master/changelogs/0.7.0.md
Project-URL: Issues, https://github.com/i3iorn/urlps/issues
Project-URL: Bug Tracker, https://github.com/i3iorn/urlps/issues
Keywords: url,parser,builder,rfc3986,validation,idna,ipv6,punycode,ipv4
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Natural Language :: English
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Topic :: Internet :: WWW/HTTP
Classifier: Topic :: Software Development :: Libraries
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: idna
Requires-Dist: idna>=3.11; extra == "idna"
Provides-Extra: dev
Requires-Dist: pytest>=7.4; extra == "dev"
Requires-Dist: pytest-cov>=4.1; extra == "dev"
Requires-Dist: hypothesis>=6.98; extra == "dev"
Requires-Dist: mypy>=1.19; extra == "dev"
Requires-Dist: bandit>=1.9; extra == "dev"
Requires-Dist: ruff>=0.7; extra == "dev"
Requires-Dist: idna>=3.11; extra == "dev"
Dynamic: license-file

# urlps

Lightweight, secure URL parsing and building library with RFC 3986 compliance. Features comprehensive security protections including SSRF prevention, DNS rebinding detection, path traversal protection, and homograph attack detection.

## Installation

```bash
pip install urlps
```

Development setup:
```bash
python -m venv .venv
. .venv/Scripts/activate  # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
```

Search cleanup: repository searches can use `.rgignore` to skip local/IDE/build artifacts.

## Quick Start

```python
from urlps import parse_url, build

# Secure by default - blocks SSRF, private IPs, localhost
url = parse_url("https://api.example.com/data?token=abc#section")
print(url.host)  # api.example.com
print(url.query_params)  # [("token", "abc")]

# Build URLs
url_str = build("https", "example.com", port=8443, path="/api", query="x=1")
# https://example.com:8443/api?x=1

# Immutable with functional updates
url = parse_url("https://example.com/path")
new_url = url.with_host("other.com").with_port(8080)
print(new_url)  # https://other.com:8080/path

# Policy-based validation
strict_url = parse_url("https://example.com", policy="strict")
balanced_url = parse_url("HTTP://EXAMPLE.com", policy="balanced")
```

### Security

`parse_url()` blocks by default:
- Private IPs (192.168.x.x, 10.x.x.x, 172.16.x.x)
- Localhost and loopback addresses
- Link-local addresses (169.254.x.x)
- `.local` and `.internal` domains
- Path traversal patterns (`../`)
- Double-encoded characters
- Mixed Unicode scripts (homograph attacks)
- URL parser confusion attacks

`parse_url(..., policy="strict")` additionally blocks:
- Query parameter injection
- Dangerous ports (commonly exploited)
- Non-canonical URL forms (filter bypass prevention)
- Credentials in URL userinfo by default (`user:pass@host`)

Use `parse_url_unsafe()` for internal/development URLs:
```python
from urlps import SecurityPolicy, parse_url_unsafe

dev_url = parse_url_unsafe("http://localhost:3000/api")
internal = parse_url_unsafe("http://192.168.1.100/metrics")

# If policy is passed, parse_url_unsafe uses it exactly.
trusted_policy = SecurityPolicy.internal(check_dns=True)
internal_checked = parse_url_unsafe("http://intranet.local/service", policy=trusted_policy)
```

Need selective hardening? Use policy presets:
- `policy="strict"`: maximum protections, DNS connect checks fail-closed by default
- `policy="balanced"` (default): fewer false positives, DNS connect checks fail-open by default
- `policy="internal"`: trusted/internal traffic

DNS connect behavior can be customized per policy:

```python
from urlps import SecurityPolicy, parse_url

policy = SecurityPolicy.strict(check_dns=True, dns_fail_open_on_connect_error=True)
url = parse_url("https://api.example.com", policy=policy)
```

Recommended for multi-tenant or concurrent applications: inject a dedicated DNS limiter.

```python
from urlps import DNSRateLimiter, DNSRateLimiterConfig, parse_url

limiter = DNSRateLimiter(
    DNSRateLimiterConfig(max_lookups_per_second=20, max_lookups_per_host=50)
)

url = parse_url(
    "https://api.example.com",
    policy="strict",
    check_dns=True,
    dns_rate_limiter=limiter,
)
```

## Core Features

### Immutable URL Objects

```python
from urlps import parse_url

url = parse_url("https://user:pass@example.com:8080/path?token=abc")
print(url.netloc)         # user:pass@example.com:8080
print(url.effective_port) # 8080

# with_* methods return new URL objects
url2 = url.with_netloc("admin@example.com")
url3 = url.with_host("other.com").with_port(443).with_path("/api")
url4 = url.with_query_param("new", "value")
url5 = url.without_query_param("token")
```

### Query strings round-trip exactly

Parsing never rewrites the query. This matters if you verify signatures over a
raw query string, or proxy URLs onward:

```python
from urlps import parse_url

url = parse_url("https://api.example.com/search?sig=aGVsbG8%3D&q=a+%26+b")
print(url.query)         # sig=aGVsbG8%3D&q=a+%26+b   (byte-for-byte)
print(str(url) )         # ...unchanged...
print(url.query_params)  # [('sig', 'aGVsbG8='), ('q', 'a & b')]
```

`q=a+%26+b` is **one** parameter whose value contains `&`. Re-encoding is only
performed when you explicitly change the query (`with_query_param()`,
`canonicalize()`), and never turns one parameter into two.

### Reference resolution (RFC 3986 §5)

`join()` is the security-preserving equivalent of `urllib.parse.urljoin` — the
resolved target is validated, so resolution can't be used to slip past the
checks `parse_url()` applies:

```python
from urlps import join

join("https://example.com/a/b", "../c")     # https://example.com/c
join("https://example.com/a/b", "?q=1")     # https://example.com/a/b?q=1
join("https://example.com/a/b", "#frag")    # https://example.com/a/b#frag

# '..' can never escape the authority
join("https://example.com/a/b", "../../../../etc/passwd")
# https://example.com/etc/passwd

# A protocol-relative reference legitimately replaces the host, which is
# exactly why the *result* is re-validated rather than trusted:
join("https://example.com/a/", "//localhost/admin")   # raises InvalidURLError
```

### Security Checks

```python
from urlps import parse_url, InvalidURLError

# SSRF protection (enabled by default)
try:
    parse_url("http://localhost/admin")  # Blocked
except InvalidURLError as e:
    print(f"Rejected: {e}")

# DNS rebinding detection (optional - rate-limited to prevent DoS)
url_dns = parse_url("https://api.example.com/", check_dns=True)

# URL canonicalization
url_raw = parse_url("HTTP://EXAMPLE.COM:80/path?z=1&a=2")
canonical = url_raw.canonicalize()
print(canonical.scheme)  # "http"
print(canonical.host)    # "example.com"
print(canonical.port)    # None (default port removed)
print(canonical.query)   # "a=2&z=1" (sorted)

# Password masking
url = parse_url("https://admin:secret123@api.example.com/")
print(url.as_string(mask_password=True))  # https://admin:***@api.example.com/
```

### Audit Logging

Audit callbacks are supplied per call via `AuditConfig`, so different callers
can log differently without sharing global state:

```python
import logging
from urlps import AuditConfig, parse_url

def audit_url_parsing(logged_url, parsed_url, exception):
    if exception:
        logging.warning(f"Failed to parse URL: {exception}")
    else:
        logging.info(f"Parsed URL to host: {parsed_url.host}")

url = parse_url(
    "https://api.example.com/data",
    audit=AuditConfig(callback=audit_url_parsing),
)
```

Structured event callback:

```python
from urlps import AuditConfig, parse_url

def on_event(event):
    # event includes: timestamp, level, operation, raw_url, host,
    # error_type, error_code, correlation_id
    print(event)

url = parse_url(
    "https://api.example.com/data",
    correlation_id="request-42",
    audit=AuditConfig(event_callback=on_event),
)
```

URLs are redacted before being passed to callbacks (credentials and sensitive
query values are masked). Pass `AuditConfig(..., redact_urls=False)` to opt out.
A callback that raises is recorded as a failure and never breaks the parse.

The same `audit=` parameter is accepted by `parse_url_unsafe()`, `join()` and
`build_secure()`.

### Component Length Limits

Conservative limits to prevent DoS attacks:

| Component | Max Length |
|-----------|------------|
| URL (total) | 32 KB |
| Scheme | 16 chars |
| Host | 253 chars |
| Path | 4 KB |
| Query | 8 KB |
| Fragment | 1 KB |
| Userinfo | 128 chars |

## Environment Variables

Override length limits via environment variables:

```bash
# PowerShell
$env:URLPS_MAX_URL_LENGTH = "65536"
python -c "import urlps.constants as c; print(c.MAX_URL_LENGTH)"

# Bash
export URLPS_MAX_URL_LENGTH=65536
python -c 'import urlps.constants as c; print(c.MAX_URL_LENGTH)'
```

Supported variables:
- `URLPS_MAX_URL_LENGTH`
- `URLPS_MAX_SCHEME_LENGTH`
- `URLPS_MAX_HOST_LENGTH`
- `URLPS_MAX_PATH_LENGTH`
- `URLPS_MAX_QUERY_LENGTH`
- `URLPS_MAX_FRAGMENT_LENGTH`
- `URLPS_MAX_USERINFO_LENGTH`
- `URLPS_MAX_IPV6_STRING_LENGTH`

## API Reference

### Main Functions

| Function | Description |
| --- | --- |
| `parse_url(url, *, allow_custom_scheme=False, check_dns=False, check_phishing=False, dns_rate_limiter=None, policy=None, correlation_id=None, audit=None)` | Parse URL with policy-aware security checks (recommended) |
| `parse_url_unsafe(url, *, allow_custom_scheme=False, debug=False, check_dns=False, dns_rate_limiter=None, policy=None, correlation_id=None, audit=None)` | Parse URL for trusted/internal input with optional policy overrides |
| `join(base, reference, *, policy=None, strict_resolution=True, ...)` | Resolve a reference against a base URI (RFC 3986 §5), then validate |
| `build(*scheme_and_host, port=None, path="/", query=None, fragment=None, userinfo=None)` | Build URL string from components |
| `build_secure(*scheme_and_host, policy=None, check_dns=False, check_phishing=False, dns_rate_limiter=None, correlation_id=None, audit=None, ...)` | Build and then validate a URL under a selected security policy |
| `compose_url(components)` | Build URL from components dict |

Note: `get_dns_rate_limiter()` and `reset_dns_rate_limiter()` remain available for compatibility, but explicit `dns_rate_limiter=` injection is preferred.

### URL Methods

| Method | Description |
| --- | --- |
| `url.as_string(mask_password=False)` | Convert to string, optionally masking password |
| `url.canonicalize()` | Return canonicalized copy |
| `url.is_semantically_equal(other)` | Compare URLs by meaning after canonicalization |
| `url.same_origin(other)` | Check if URLs have same origin |
| `url.origin` | Return origin string (e.g., `https://example.com`) |
| `url.copy(**overrides)` | Create copy with optional component overrides |
| `url.with_*()` | Functional updates: `with_scheme`, `with_host`, `with_port`, `with_path`, `with_fragment`, `with_userinfo`, `with_netloc`, `with_query_param`, `without_query_param` |

### Cache Management

```python
from urlps import get_cache_info, clear_all_caches

# Get cache statistics
stats = get_cache_info()
print(stats['parser']['normalize_path']['hits'])

# Clear all caches (useful for long-running apps)
previous = clear_all_caches()
```

## Comparison with urllib.parse

| Feature | urllib.parse | urlps |
| --- | --- | --- |
| Basic URL parsing | ✓ | ✓ |
| RFC 3986 strict compliance | Partial | ✓ |
| SSRF protection | ✗ | ✓ |
| DNS rebinding detection | ✗ | ✓ (with rate limiting) |
| Path traversal detection | ✗ | ✓ |
| Homograph detection | ✗ | ✓ |
| URL parser confusion protection | ✗ | ✓ |
| Query parameter injection detection | ✗ | ✓ |
| Dangerous port validation | ✗ | ✓ |
| Canonical form validation | ✗ | ✓ |
| Immutable URL objects | ✗ | ✓ |
| URL canonicalization | ✗ | ✓ |
| Password masking | ✗ | ✓ |
| Audit logging | ✗ | ✓ |
| Component length limits | ✗ | ✓ |

**Use urllib.parse when:** You need zero dependencies and basic parsing is sufficient.

**Use urlps when:** Security matters, you need RFC 3986 strict compliance, or you want immutable URL objects with ergonomic manipulation methods.

## Exceptions

```python
from urlps import InvalidURLError, URLParseError, parse_url

user_input = "https://example.com"

try:
    url = parse_url(user_input)
except URLParseError:
    print("Malformed URL")
except InvalidURLError:
    print("Rejected by security policy")
```

Exception hierarchy:
- `InvalidURLError` — Base exception for all URL errors
- `URLParseError` — Parsing errors
- `URLBuildError` — Building errors
- `HostValidationError` / `PortValidationError` — Component validation errors
- `QueryParsingError`, `FragmentEncodingError`, `UserInfoParsingError`, `UnsupportedSchemeError` — Specific errors

## Running Tests

```bash
pytest
pytest -v -k "test_parse"     # Run specific tests
pytest -m ipv6                # Run IPv6 tests
pytest -m idna                # Run IDNA tests
```

## Changelog

See [CHANGELOG.md](CHANGELOG.md) for a summary of every release, and [changelogs/](changelogs/) for detailed per-release notes.

## License

MIT
