MCP server exposing 3 tools for arquivo-pt.
This URL is a JSON-RPC 2.0 endpoint over HTTP. Issue POST requests with a JSON-RPC body. Browsers and search crawlers land here on GET.
POST https://gateway.pipeworx.io/arquivo-pt/mcp
Content-Type: application/json
{"jsonrpc":"2.0","id":1,"method":"tools/list"}
arquivo_search_pages — Full-text search of archived web pages in Arquivo.pt, the Portuguese web archive (captures since 1996), when the URL is unknown. Finds pages whose TEXT contains the given terms or a quoted phrase and returns, per hit, the page title, original URL, capture date, a text snippet with the matching terms, the archived (replay) URL, and links to the extracted text and a screenshot. Coverage: centred on the Portuguese web (.pt sites and Portuguese-language pages) but international English-language sites that Portuguese pages link to are captured too — probes for "climate change", "quantum computing" and "Federal Reserve interest rates" return roughly 50M, 1M and 2M estimated hits, with ipcc.ch, microsoft.com and federalreserve.gov among the top results. The full-text index lags the crawl by years (probed 2026-10: most hits end in 2020, a few reach 2023) while arquivo_url_history carries current captures, so use that for anything recent at a known URL. Supports a date window (YYYY or YYYYMMDD), restricting to one site, restricting to a file type (pdf, html, doc), exclusion with a leading minus (Albert -Einstein), and paging by offset. Do not pass a URL as the query; use arquivo_url_history for that.arquivo_url_history — List every preserved capture of a specific URL in Arquivo.pt, the Portuguese web archive, newest first — the version history of a page or domain from 1996 to the current crawl. Each capture carries the capture timestamp (UTC), the archived replay URL, the crawl HTTP status, MIME type, content length and content digest, plus links to the extracted text, the raw original file and a screenshot. Accepts a bare domain (publico.pt), a host (www.publico.pt) or a full URL with path and query. Use this to see how a Portuguese page changed over time or to find the capture timestamp to pass to arquivo_page_text; use arquivo_search_pages when you only know words from the page, not its address.arquivo_page_text — Read the extracted plain text of one archived page capture held by Arquivo.pt, the Portuguese web archive — the page content with HTML stripped, as the archive indexed it. Takes the original URL and the capture timestamp (YYYYMMDDhhmmss) that arquivo_search_pages or arquivo_url_history returned, and returns the text (optionally truncated to max_chars) with the archived replay URL. Use it to quote or summarise what a Portuguese page said on a given date without fetching the archived HTML.Code samples (curl / TypeScript / one-click client install), schemas, and the live playground are on the pack page:
https://pipeworx.io/packs/arquivo-pt/
Pipeworx is an open MCP gateway connecting AI agents to live data. pipeworx.io