ArchiveBoxWeb Archive

Open-source self-hosted web archiving

Preserve the web.On infrastructure you control.

ArchiveBox saves websites, bookmarks, RSS feeds, social posts, media, source code, and research material in durable files like HTML, PDF, PNG, TXT, JSON, WARC, MP4, and SQLite.

View screenshots ↗   ·   View demo ↗




Open source MIT badge Active development badge GitHub stars badge Docker pulls badge PyPI installs badge Chrome store users badge
# Docker Compose is the recommended setup
$ mkdir -p ~/archivebox/data && cd ~/archivebox
$ curl -fsSL 'https://docker-compose.archivebox.io' > docker-compose.yml
$ docker compose up -d --wait
-> initialized ./data/index.sqlite3
ok listening on http://127.0.0.1:5797
$ 
ArchiveBox Web UI showing saved pages with thumbnails and capture formats ArchiveBox on iPhone, with its collection, tags, and server connection inside an iPhone frame

Browse and save from any device.
Free, self-hosted, 0 analytics.

Snapshots grid — desktopSnapshot View (header collapsed) — desktopAI agent — desktopArchive results — desktopAdd URLs — desktopSnapshots table — desktopFirst-time setup wizard — desktopLogin — desktopPublic snapshot list — desktopSnapshot View (capture in progress) — desktopAdmin dashboard — desktopSnapshot admin detail — desktopSnapshot files — desktopArchive result detail — desktopTags — desktopTag detail — desktopUsers — desktopUser detail — desktopCrawls — desktopCrawl detail — desktopCrawl schedules — desktopCrawl schedule detail — desktopPersonas — desktopPersona detail — desktopMachines — desktopMachine detail — desktopNetwork interfaces — desktopNetwork interface detail — desktopBinaries — desktopBinary detail — desktopProcesses — desktopProcess detail — desktopAPI tokens — desktopAPI token detail — desktopWebhooks — desktopWebhook detail — desktopEnvironment — desktopConfiguration — desktopConfiguration detail — desktopDependencies — desktopDependency detail — desktopPlugins — desktopWorkers — desktopWorker detail — desktopLogs — desktopLog detail — desktopSnapshot View (singlefile) — desktopSnapshot View (screenshot) — desktopSnapshot View (wget) — desktopSnapshot View (dom) — desktopSnapshot View (pdf) — desktopSnapshot View (readability) — desktopSnapshot View (archivewebpage) — desktopSnapshot View (responses) — desktopSnapshot View (ytdlp) — desktopSnapshot View (chrome_mhtml) — desktopSnapshot View (defuddle) — desktopSnapshot View (mercury) — desktopSnapshot View (tlsnotary) — desktopSnapshot View (opentimestamps) — desktopSnapshot View (chrome) — desktopSnapshot View (consolelog) — desktopSnapshot View (dns) — desktopSnapshot View (sslcerts) — desktopSnapshot View (redirects) — desktopSnapshot View (headers) — desktopSnapshot View (seo) — desktopSnapshot View (accessibility) — desktopSnapshot View (htmltotext) — desktopSnapshot View (trafilatura) — desktopSnapshot View (parse_html_urls) — desktopSnapshot View (parse_txt_urls) — desktopSnapshot View (parse_dom_outlinks) — desktopSnapshot View (hashes) — desktop
CLI Web UI REST API Webhooks Browser extension Filesystem access

Why ArchiveBox

Designed to make your archived data useful today.

Saves tabs in seconds, stores them for decades

Collect bookmarks to read later, or preserve evidence for legal cases with verifiable captures from our TLSNotary signing plugin.

Tag & search millions of captures instantly

Organize your collection with tags and find what you need with full-text search powered by Sonic and ripgrep.

Auto-extract embedded media to standard formats

Automatically extract audio, video, git archives, forum posts, SEO metadata, and more into ordinary files you can open and reuse.

Connect with other systems easily

Tell your AI to inspect a page with abx-dl, pull an RSS feed every 24 hours, or orchestrate complex crawls through MCP and the REST API.

Who it is for

Powerful features for individuals, professionals, and institutions.

Personal archivists and self-hosters

Save bookmarks, browser history, RSS feeds, social media, form content, videos, podcasts, music, photos, and personal knowledge collections.

  • Own your data and keep it on local or remote storage you control.
  • Use the browser extension, CLI, Web UI, and scheduled imports together.
  • Export static HTML or browse the filesystem directly.
iPhone share sheet confirming a URL was saved with tags Android share sheet showing a successful save, suggested tags, and confirmation that tags were saved

Lawyers and journalists

Keep copies of articles, source material, and public records, even after the original pages change or disappear.

  • Store screenshots, PDFs, headers, WARC files, and text extraction.
  • Tag and review sources through the self-hosted web interface.
  • Use ZK proofs with TLSNotary to prove content authenticity.
TLSNotary output verifying an archived response on the snapshot details page

Researchers and institutions

Support OSINT, social media research, AI-powered research agents, libraries, governments, and collection-building teams.

  • Automate imports through the CLI, REST API, webhooks, and schedules.
  • Keep machine-readable metadata in JSON and SQLite.
  • Extend extraction pipelines through the ArchiveBox ecosystem.
Archive depth, domain and subpath scope, URL allowlist, and URL denylist options on the Add page

Get ArchiveBox

Quickstart

Choose how you want to run your archive. Docker Compose is recommended for most installs.

ArchiveBox — Quickstart
ArchiveBox on iPhone showing its collection and server tools
ArchiveBox for iPhone & iPad
1

Install ArchiveBox for iOS

Get ArchiveBox.app for iOS

Open the TestFlight link on your iPhone or iPad to install ArchiveBox.

2

Connect your server

The first-run guide helps you choose a server, or select I already have a server. Scan the connection QR code from ArchiveBox Server.app on your Mac, or enter your server URL and API key.

3

Save your first link

In Safari or another app, choose Share → ArchiveBox. Your server saves the page so you can browse it from your phone.

iOS setup guide ↗

ArchiveBox on Android showing its home screen, connected server, and collection tools
ArchiveBox for Android
1

Install ArchiveBox for Android

Get ArchiveBox for Android

Android beta APKs will appear here when published. Open the APK and allow installation from your browser if Android asks.

2

Connect your server

Use the first-run guide to choose a server, or select I already have a server. In Connection Settings, enter your server URL and API key, then verify the connection.

3

Save your first link

Choose Share → ArchiveBox from your browser or another app, add tags, and save. Keep your server reachable to save and browse pages.

Android setup guide ↗

ArchiveBox.app
Applications
ArchiveBox.app connection settings with a verified local server
macOS 26+ · Apple Silicon
1

Download and open the app

Get ArchiveBox.app for macOS

Move it to Applications, and open it.

2

Follow the setup guide

Choose Set up on this Mac in the setup guide, then download and open ArchiveBox Server.app. Already have a server? Connect with its URL and API key.

3

Connect your apps

Use the Web UI, REST API, or CLI. Connect the mobile app, browser extension, and other clients using your server URL and API key.

Windows
ArchiveBox
ArchiveBox Desktop for Windows
1

Install Docker Desktop

Download Docker Desktop and start it on your computer. Keep Docker running while you use ArchiveBox.

2

Download ArchiveBox for Windows

Open the installer, launch ArchiveBox, and create your administrator account. The first launch prepares your local collection.

3

Start saving

Add a URL in the app to archive your first page.

Desktop app guide ↗

Docker Desktop
ArchiveBox Desktop on Linux showing its saved-page collection, captured from the packaged app in CI
ArchiveBox Desktop · Real Linux app
1

Install Docker Desktop

Download Docker Desktop and start it on your computer. Keep Docker running while you use ArchiveBox.

2

Download ArchiveBox for Linux

Open the installer, launch ArchiveBox, and create your administrator account. The first launch prepares your local collection.

3

Start saving

Add a URL in the app to archive your first page.

Desktop app guide ↗

Docker Compose
ArchiveBox
ArchiveBox web interfacedocker-compose.yml
Recommended
1

Docker Compose Recommended

Compose keeps your settings in one file and makes updates easier. Install Docker, then create a directory and get the config.

mkdir -p ~/archivebox/data && cd ~/archivebox
curl -fsSL 'https://docker-compose.archivebox.io' > docker-compose.yml
2

Start ArchiveBox

docker compose pull
docker compose up -d --wait

A new collection initializes automatically.

3

Finish setup and archive your first page

Open the admin UI to create your first admin account and finish the setup wizard. On a remote server, open /admin/ on its hostname or IP instead.

docker compose exec archivebox archivebox add 'https://example.com'
Use raw Docker instead docker run

Have a Compose file? Convert it to docker run commands with Decomposerize.

1

Create a collection and start the server

Install Docker first.

mkdir -p ~/archivebox/data && cd ~/archivebox/data
docker run -d --name archivebox -v "$PWD:/data" -p 5797:5797 archivebox/archivebox:dev
2

Finish setup

Open the admin UI to create your first admin account and finish the setup wizard. On a remote server, open /admin/ on its hostname or IP instead.

3

Archive your first page

docker exec archivebox archivebox add 'https://example.com'
pip / uv
ArchiveBox
HTMLPDFPNGWARC
Python package
1

Install the Python package

Install uv first.

uv tool install --python 3.13 --prerelease explicit --upgrade 'archivebox>=0.9.0rc0,<0.10'
archivebox version
2

Create a collection and start the server

mkdir -p ~/archivebox/data && cd ~/archivebox/data
archivebox init
archivebox install
archivebox server 0.0.0.0:5797
3

Finish setup

Open the admin UI to create your first admin account and finish the setup wizard. On a remote server, open /admin/ on its hostname or IP instead.

Or use pip instead of uv
1

Install with pip

With Python 3.13 installed, create a virtual environment and install ArchiveBox:

python3.13 -m venv ~/.venvs/archivebox
source ~/.venvs/archivebox/bin/activate
python -m pip install --upgrade 'archivebox>=0.9.0rc0,<0.10'
archivebox version

Then follow steps 2 and 3 above. In each new terminal, run source ~/.venvs/archivebox/bin/activate before using archivebox.

apt
ArchiveBox
HTMLPDFPNGWARC
Ubuntu / Debian
1

Add the repository and install

echo 'deb [trusted=yes] https://archivebox.github.io/debian-archivebox dev main' | sudo tee /etc/apt/sources.list.d/archivebox.list
sudo apt update
sudo apt install archivebox
(cd /tmp && archivebox version)
2

Create a collection and start the server

mkdir -p ~/archivebox/data && cd ~/archivebox/data
archivebox init
sudo archivebox install
archivebox server 0.0.0.0:5797
3

Finish setup

Open the admin UI to create your first admin account and finish the setup wizard. On a remote server, open /admin/ on its hostname or IP instead.

Homebrew
ArchiveBox
HTMLPDFPNGWARC
macOS / Linux
1

Install from the ArchiveBox tap

Install Homebrew first. Run these commands as your normal user, without sudo.

brew tap archivebox/archivebox
brew trust archivebox/archivebox
brew install archivebox
archivebox version
2

Create a collection and start the server

mkdir -p ~/archivebox/data && cd ~/archivebox/data
archivebox init
archivebox install
archivebox server 0.0.0.0:5797
3

Finish setup

Open the admin UI to create your first admin account and finish the setup wizard. On a remote server, open /admin/ on its hostname or IP instead.

Arch Linux
FreeBSD
Nix
ArchiveBox
Community packages
1

Choose your platform

Community packages may lag behind the Docker and Python releases.

Arch Linux (AUR)

yay -S archivebox

FreeBSD (uv shortcut)

curl -fsSL 'https://get.archivebox.io' | bash

Nix

nix-env --install archivebox

Guix

guix install archivebox
2

Set up your collection

Follow your package’s setup instructions, then use the CLI usage guide to initialize and manage your archive.

$curl get.archivebox.io
| sh
ArchiveBox
One command to get started
1

Run the setup script

curl -fsSL 'https://get.archivebox.io' | bash

You can read the script before running it.

2

Follow the prompts and finish setup

Follow the installer’s instructions to start using your collection.

Open the admin UI to create your first admin account and finish the setup wizard. On a remote server, open /admin/ on its hostname or IP instead.

ArchiveBox
NASHome server
On your own hardware
1

Choose your platform

Community integrations may lag behind the Docker and Python releases.

2

Install and open ArchiveBox

Follow the platform’s installation guide, configure persistent storage, then open the app to finish setup.

For TrueNAS, see the README’s platform notes before choosing an integration.

ArchiveBox
Your serverAnywhere
Cloud hosting / VPS
1

Choose a provider

Check the provider’s current plans, supported versions, and storage options.

2

Deploy and finish setup

Use the provider’s ArchiveBox deployment flow, then open your instance to create an admin account. For an unmanaged VPS, follow the Docker Compose quickstart on your server.

ArchiveBox browser extension saved URLs
Save from your browser
Chrome · Firefox · Edge · Safari
1

Install the browser extension

2

Connect your ArchiveBox server

Open the extension’s Configuration page, enter your server URL and API key, then test the connection. Use one of the server install tabs if you don’t have a server yet.

3

Save a page

Open a page, click the ArchiveBox extension, and add your tags. You can also import bookmarks and history from the extension’s options.

Extension guide ↗

Configuration

50+ plugins & flexible configuration options.

Browse plugins and extractors ↗

ArchiveBox can be configured via the admin UI, CLI, ArchiveBox.conf, or by importing config and cookies from Chrome / other browsers. The same configuration model works in Docker, Docker Compose, bare-metal installs, scheduled jobs, and one-off CLI runs.

Adjust how pages are captured

Use TIMEOUT for slow networks, CHECK_SSL_VALIDITY for sites with broken certificates, PUBLIC_INDEX, PUBLIC_SNAPSHOTS, and PUBLIC_ADD_VIEW for publishing policy, and browser user-agent settings for sites that block obvious bots.

archivebox config
$ archivebox config                         # view full config
$ archivebox config --get CHROME_BINARY
$ archivebox config --set TIMEOUT=240
$ archivebox config --set PUBLIC_INDEX=False
$ env CHROME_BINARY=chromium archivebox add 'https://example.com'

Inputs and outputs

Integrate with tools you already use.

Browse an archive in the Web UI

Access

Use the same collection wherever you work:

Disk layout, exporting, and security

Browse your captures as normal files without needing ArchiveBox.

Your captures are ordinary HTML, images, PDFs, and media files. Open them from Finder, search them with any filesystem search tools, let AI agents browse them directly, and back them up as regular files.

Designed to last 100yr+

Each snapshot has a folder for every capture method. HTML lives in folders such as dom/ and singlefile/; images, PDFs, and media have their own folders too.

Metadata is stored in SQLite and JSON, with snapshot metadata alongside the outputs. Your collection’s database and configuration live outside the snapshot folders, and your files remain readable without a hosted service.

macOS Finder with the date, news.ycombinator.com, and snapshot folders expanded down to dom/output.html, with the saved HTML file selected

Static exports and publishing

You can export the archive index as static HTML, JSON, or CSV with archivebox list so collections can be reviewed without running the web server. Keep generated exports next to the archive/ folder so relative snapshot paths continue to work.

$ archivebox list --html --with-headers > index.html
$ archivebox list --json --with-headers > index.json
$ archivebox list --csv=timestamp,url,title > index.csv
Archived page rendered from saved HTML

Archived JavaScript

ArchiveBox uses a Chrome browser during archiving to run JavaScript and capture the rendered page.

ArchiveWeb.page records pages for replay, including their scripts and network responses. Replay captured pages with JavaScript using ReplayWeb.page, Webrecorder’s archive viewer.

Snapshot details and saved archive size

Storage requirements

ArchiveBox can use roughly 1 GB to 50 GB per 1,000 snapshots depending on media, video, audio, and extractor settings. Tune YTDLP_ENABLED and YTDLP_MAX_SIZE for media-heavy collections.

Keep index.sqlite3 on local storage or SSD when possible. Large archive/ folders can live on NFS, SMB, FUSE, S3-backed, or HDD storage, but Docker and fileshare setups may need ownership and root-squash adjustments. Avoid older filesystems like EXT3 or FAT for very large archives.

Plan for large archives

Background and motivation

Keep a copy of the web you care about.

ArchiveBox exists because link rot, platform churn, censorship, and disappearing media routinely erase useful knowledge. The project aims to make the web content you care about viewable with common software in 50 to 100 years, without requiring ArchiveBox or a hosted replay service to understand your files.

Self-hosted ArchiveBox dashboard

Own all your data

Centralized public archives are essential, but not every page belongs in a global public service. ArchiveBox lets individuals and organizations save public or private material they can access, keep it locally or within their institution, and decide what to publish case by case.

Run it as a Docker web app, use one-off CLI commands, or automate with APIs while keeping private and public data neatly separated. Use ordinary files for backups and exports.

Ethical and legal context

Responsible archiving depends on your jurisdiction, your use case, and your publishing policy. ArchiveBox is a tool; operators remain responsible for handling private data, copyright, DMCA or GDPR requests, and local legal requirements.

Public instances should publish contact and removal information, avoid monetizing copied content, and review the security and publishing documentation before exposing snapshots to anonymous viewers.

Related web archiving tools and communities

Looking for specialized crawling, replay, or bookmark management? Explore these other web archiving projects.