Metadata-Version: 2.5
Name: invenio-archive-it
Version: 0.1.4
Summary: Links to archived versions of InvenioRDM records in Archive-It.
Project-URL: Homepage, https://codeberg.org/front-matter/invenio-archive-it
Project-URL: Repository, https://codeberg.org/front-matter/invenio-archive-it
Project-URL: Bug Tracker, https://codeberg.org/front-matter/invenio-archive-it/issues
Author-email: Martin Fenner <martin@front-matter.de>
License-Expression: MIT
License-File: LICENSE
Keywords: archive-it,internet archive,invenio,inveniordm,wayback
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Web Environment
Classifier: Framework :: Flask
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Internet :: WWW/HTTP :: Dynamic Content
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.11
Requires-Dist: invenio-app-rdm>=14.0.0
Requires-Dist: invenio-jobs>=10.0.0
Requires-Dist: requests>=2.32
Requires-Dist: structlog>=24.1
Provides-Extra: dev
Requires-Dist: black>=24.0; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Provides-Extra: tests
Requires-Dist: invenio-search[opensearch2]>=3.0.0; extra == 'tests'
Requires-Dist: pytest>=8.0; extra == 'tests'
Requires-Dist: responses>=0.25; extra == 'tests'
Description-Content-Type: text/markdown

# invenio-archive-it

Links to archived versions of InvenioRDM records in [Archive-It](https://archive-it.org).

Archiving is set up per community. A community names the Archive-It collection
that archives its website, for example
[collection 22103](https://archive-it.org/collections/22103) for
<https://jabberwocky.weecology.org/>. A job run from the admin UI looks up that
website in the collection and stores on each of the community's records which
of its URLs were captured, and when. The record's landing page sidebar lists
them under **External resources → Archived in**, the same way Zenodo shows
software archived in Software Heritage.

Requires InvenioRDM v14 or later.

## How it works

1. A community admin sets the community's **Archive-It collection**
   (`ia:collection`). The website to look up is the community's own website
   (`metadata.website`). Example: collection ID `22103` for
   <https://jabberwocky.weecology.org/>.
2. An admin runs the **Update Archive-It captures** job for one community, or
   for every community with a collection. Each community gets one Celery task,
   rate-limited to **6 per minute** (as Archive-It's Wayback service is shared
   across all users).
3. The task sends one prefix query to
   `https://wayback.archive-it.org/{collection}/timemap/cdx` for the website,
   which returns every capture under it. It then matches the captures to the
   URLs of the records whose default community this is. Those are identifiers
   and related identifiers with scheme `url`; Archive-It playback URLs are
   first reduced to the page they show.
4. Each record's matched captures (first and last capture per URL) are stored
   in the record custom field `ia:captures`. Only records whose captures changed
   are written.
5. The sidebar shows one **Archive-It** entry per captured URL, with the URL and
   capture dates beneath it. The entry links to the Wayback calendar, e.g.
   `https://wayback.archive-it.org/22103/*/https://jabberwocky.weecology.org/…`.

If the lookup fails or Archive-It refuses it, every record keeps its previous
captures.

## Installation

```sh
uv add git+https://codeberg.org/front-matter/invenio-archive-it
```

Add the custom fields and the sidebar link to `invenio.cfg`. Merge them with
any namespaces, custom fields and external links you already have:

```python
from invenio_archive_it.custom_fields import (
    ARCHIVE_IT_COMMUNITY_CUSTOM_FIELDS,
    ARCHIVE_IT_COMMUNITY_CUSTOM_FIELDS_UI,
    ARCHIVE_IT_CUSTOM_FIELDS,
    ARCHIVE_IT_NAMESPACE,
)
from invenio_archive_it.links import archived_versions_render

RDM_NAMESPACES = {**ARCHIVE_IT_NAMESPACE}
RDM_CUSTOM_FIELDS = [*ARCHIVE_IT_CUSTOM_FIELDS]
COMMUNITIES_CUSTOM_FIELDS = [*ARCHIVE_IT_COMMUNITY_CUSTOM_FIELDS]
COMMUNITIES_CUSTOM_FIELDS_UI = [ARCHIVE_IT_COMMUNITY_CUSTOM_FIELDS_UI]
APP_RDM_RECORD_LANDING_PAGE_EXTERNAL_LINKS = [
    {"id": "archive_it", "render": archived_versions_render},
]
```

`COMMUNITIES_NAMESPACES` defaults to `RDM_NAMESPACES`, so `ia` covers both.
To add the collection field to an existing community settings section instead
of its own "Web archiving" section, add its entry from
`ARCHIVE_IT_COMMUNITY_CUSTOM_FIELDS_UI["fields"]` to that section.

Create the search mappings:

```sh
invenio rdm-records custom-fields init -f ia:captures
invenio communities custom-fields init -f ia:collection
```

Copy the sidebar icon into the instance's static folder with `invenio collect`
(`invenio-cli assets build` runs it too).

To run lookups, open **Update Archive-It captures** under Administration → Jobs
and start a run. Pick a community, or leave the field empty for every community
with a collection. The module adds no Celery beat schedule.

## Configuration

| Setting | Default | Meaning |
| --- | --- | --- |
| `ARCHIVE_IT_WAYBACK_URL` | `https://wayback.archive-it.org` | Wayback service for lookups and links; validated during startup |
| `ARCHIVE_IT_ICON` | `images/invenio_archive_it/archive-it.svg` | Sidebar icon (Archive-It logo), as a path in the static folder |
| `ARCHIVE_IT_TIMEOUT` | `120` | Seconds to wait for a CDX response |
| `ARCHIVE_IT_USER_AGENT` | `invenio-archive-it (+…)` | User-Agent sent with lookups |
| `ARCHIVE_IT_CDX_LIMIT` | `50000` | Captures asked for per lookup. Archive-It answers at most 50,000, and says it held some back only when a limit is asked for |

**Important**: Misconfiguring `ARCHIVE_IT_WAYBACK_URL` will cause sidebar links to fail
silently. The module validates the URL on startup; check the logs if configured endpoints
are unreachable.

## Troubleshooting

- **Lookups fail silently**: Archives at Archive-It's Wayback service are behind bot
  protection. Check logs for `error=blocked`; verify your server can reach the CDX API.
  Failed lookups preserve existing captures; no data is lost.
- **Lookup errors in logs**: Monitor warnings with `error=http_error`, `error=network`,
  or `error=invalid_response` to catch misconfigurations or API changes.
- **Answer cut short**: a lookup asks only for pages and their revisit records, but a site
  can still have more than Archive-It's 50,000-capture maximum. The warning
  `Archive-It answered only part of the collection` then names the last URL reached;
  records past it keep the captures they had rather than losing them.
- **Sidebar links appear broken**: Ensure `ARCHIVE_IT_WAYBACK_URL` is correct. The module
  validates it on startup; misconfiguration will be logged at warning level.
- **Specific community not found**: Only records whose default community matches are
  updated. Verify the community is set correctly on records.
- Captures are written to the published record directly. This doesn't create
  a new version or a draft. An open draft of the record gets the same captures,
  so publishing it doesn't bring back old ones.
- Anything else that updates records must keep `ia:captures`. An update through
  the records service replaces all custom fields with what it sends.
- Anyone who can edit a community's settings can set its collection. To make
  it admin-only, restrict the field in the instance.

## Development

```sh
uv sync --all-extras
uv run pytest
uv run ruff check .
uv run black --check .
```

## License

MIT
