Metadata-Version: 2.5
Name: mwextract
Version: 0.1.0
Summary: Convert a MediaWiki XML dump (.7z / .xml / .xml.bz2) into structured Markdown, JSON, or XML.
Project-URL: Homepage, https://github.com/adityaparab/mwextract
Project-URL: Repository, https://github.com/adityaparab/mwextract
Project-URL: Issues, https://github.com/adityaparab/mwextract/issues
Author-email: Aditya Parab <mradityaparab@gmail.com>
License: MIT
License-File: LICENSE
Keywords: converter,dump,fandom,markdown,mediawiki,rag,wiki,wikia,wikitext
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Text Processing :: Markup
Classifier: Topic :: Utilities
Requires-Python: >=3.10
Requires-Dist: mwparserfromhell>=0.6
Requires-Dist: mwxml>=0.3.3
Requires-Dist: py7zr>=0.20
Provides-Extra: dev
Requires-Dist: pytest>=7; extra == 'dev'
Description-Content-Type: text/markdown

# mwextract

Convert a **MediaWiki XML dump** (as shipped by Fandom/Wikia and other MediaWiki
sites) into clean, structured **Markdown**, **JSON**, or **XML** — one command,
no pandoc, no other binaries.

It uses only the Wikimedia parsing stack:

- [`mwxml`](https://pypi.org/project/mwxml/) streams pages out of the dump
  (memory-safe, namespace-aware).
- [`mwparserfromhell`](https://pypi.org/project/mwparserfromhell/) parses
  wikitext: infobox/template fields become metadata, and the article body is
  rendered to clean Markdown.
- [`py7zr`](https://pypi.org/project/py7zr/) extracts the `.7z` archive the dump
  usually comes in.

## Install

```bash
uv add mwextract
# or
pip install mwextract
```

## Get a dump

`mwextract` does **not** download anything — you fetch the dump yourself. Fandom
wikis expose current-page dumps under a URL like:

```
https://s3.amazonaws.com/wikia_xml_dumps/<x>/<xx>/<wiki>_pages_current.xml.7z
```

For example, the Game of Thrones wiki:
`https://s3.amazonaws.com/wikia_xml_dumps/g/ga/gameofthrones_pages_current.xml.7z`

Download the `.7z` and point `mwextract` at it.

## Use it from Python

```python
import mwextract

result = mwextract.convert(
    "gameofthrones_pages_current.xml.7z",   # .7z, .xml, or .xml.bz2
    "out/",                                  # output directory
    base_url="https://gameofthrones.fandom.com/wiki/",  # optional
    fmt="md",                                # "md" | "json" | "xml"
    layout="combined",                       # "combined" | "per-file"
)

print(result.pages, "pages ->", result.outputs)
```

`convert()` extracts the archive to a temporary directory (removed afterwards),
streams every main-namespace, non-redirect article, and writes the chosen format.

## Use it from the command line

```bash
mwextract gameofthrones_pages_current.xml.7z out/ \
  --base-url https://gameofthrones.fandom.com/wiki/ \
  --format md \
  --layout combined
```

```
usage: mwextract [-h] [-b BASE_URL] [-f {md,json,xml}]
                 [-l {combined,per-file}] [-q] [-V]
                 dump output_dir
```

| Option | Default | Meaning |
| --- | --- | --- |
| `dump` | — | Path to the dump: `.7z`, `.xml`, or `.xml.bz2`. |
| `output_dir` | — | Directory to write into (created if missing). |
| `-b`, `--base-url` | *(none)* | Wiki URL prefix; when set, each page gets a `url`. |
| `-f`, `--format` | `md` | `md`, `json`, or `xml`. |
| `-l`, `--layout` | `combined` | `combined` (one file) or `per-file` (one per article). |
| `-q`, `--quiet` | off | Only warnings/errors. |

## Output

### Layouts

- **`combined`** (default) — a single file for the whole corpus:
  `corpus.md`, `corpus.json` (a JSON array), or `corpus.xml`
  (`<pages><page>…</page></pages>`).
- **`per-file`** — one file per article, named after the title:
  `Daenerys_Targaryen.md`, `.json`, or `.xml`.

Both layouts are written incrementally as pages stream in, so memory stays flat
no matter how large the dump is.

### Formats

**Markdown** — YAML front-matter (title, url, infobox fields) + an optional
summary line + the rendered body:

```markdown
---
title: "Jon Snow"
url: https://gameofthrones.fandom.com/wiki/Jon_Snow
house: "Stark"
allegiance: "Night's Watch"
---

# Jon Snow

> **House:** Stark · **Allegiance:** Night's Watch

Jon Snow is the illegitimate son of ...
```

**JSON** — the same content, structured:

```json
{
  "title": "Jon Snow",
  "url": "https://gameofthrones.fandom.com/wiki/Jon_Snow",
  "metadata": { "house": "Stark", "allegiance": "Night's Watch" },
  "body": "Jon Snow is the illegitimate son of ..."
}
```

**XML** — the same content as elements:

```xml
<page>
  <title>Jon Snow</title>
  <url>https://gameofthrones.fandom.com/wiki/Jon_Snow</url>
  <metadata>
    <field name="house">Stark</field>
    <field name="allegiance">Night's Watch</field>
  </metadata>
  <body>Jon Snow is the illegitimate son of ...</body>
</page>
```

`body` is the rendered Markdown body only — the structured `metadata`/`url`
carry everything the Markdown front-matter would, so nothing is duplicated.

## What gets kept / dropped

- **Kept:** main-namespace articles, section hierarchy, bold/italic, list items,
  wikilink and external-link visible text, infobox fields (as metadata).
- **Dropped:** redirects and non-article namespaces, file/image/category/media
  links, citation footnotes (`<ref>`), templates (after their infobox fields are
  extracted), and comments.

## License

MIT © Aditya Parab
