Wallabag, a self-hosted read-later and article archive app

Half the articles you saved last year are already dead. Sites redesign, paywalls close, and companies quietly delete their old archives. Bookmarks pointing at gone pages are worse than useless because they hide the fact that you cannot re-read what you thought you saved. The fix is to keep the article, not the link.

We use eight desktop apps to build and search a personal archive of what we read. Some are self-hosted, some are cloud services with a strong export path, all of them capture the full page (text, images, and metadata) instead of just a URL. Between them, one library will hold the last five years of reading with search across every word.

What to look for in a personal web archive app

The eight below meet at least four of the five. The rank is by how much of the pipeline they cover on their own.

Quick comparison

App Best for Free plan Starting price Self-hosted
Wallabag Self-hosted read-later with clean parsing Community server 11 euros/yr hosted Yes
ArchiveBox Snapshot-style deep archive Self-hosted only Free Yes
Karakeep Bookmark manager with offline copies Self-hosted only Free Yes
Omnivore Cross-device read-later with highlights Yes Free Yes (fork)
SingleFile One-click page-to-HTML in a browser Free Free Standalone
Readwise Reader Curated reading and highlights Trial $9.99/mo No
Zotero Academic capture with citation Yes $20/yr storage Partial
Pocket Cross-platform save queue Yes $4.99/mo No

1. Wallabag, best for a self-hosted read-later with clean parsing

Wallabag is the self-hosted read-later app that keeps parsing clean pages long after the source rots. It pulls the article text, stores the images, and runs a clean reading view across desktop and mobile. The hosted “Wallabag.it” plan is 11 euros a year for anyone who does not want to run their own server.

Where it falls short: parsing occasionally misses on javascript-heavy pages. The community filters catch most of them within a few days.

Pricing:

Platforms: Web app, browser extensions, plus native clients for Windows, macOS, and Linux via Docker or the standalone builds.

Download: wallabag.org

Bottom line: The default recommendation for a self-hosted archive that respects the article, not the tracking script.

2. ArchiveBox, best for a snapshot-style deep archive

ArchiveBox takes a URL and stores the full page (HTML, screenshot, WARC, PDF, and a media snapshot) in a local folder you own. Search runs across every stored copy, and the metadata layer keeps titles, tags, and timestamps.

Where it falls short: the archive is deep, and disks fill up fast. A year of daily saves can push into terabytes if video is included.

Pricing:

Platforms: Windows, macOS, Linux (Docker or Python install).

Download: archivebox.io

Bottom line: The right archive for anyone who wants a permanent copy, not a read-later queue.

3. Karakeep, best for a bookmark manager with offline copies

Karakeep takes the bookmark manager pattern and adds a page snapshot per entry. Save a link, get a searchable copy that survives a dead source. Tags and lists sit on top, and the mobile companion apps keep the offline copies in sync.

Where it falls short: the storage layer is per-user; sharing an archive with a partner or a small team requires a small ops step.

Pricing:

Platforms: Windows, macOS, Linux (Docker), plus mobile companions.

Download: karakeep.app

Bottom line: The choice for anyone who wants their bookmarks to survive when the bookmarked page does not.

4. Omnivore, best for cross-device read-later with highlights

Omnivore is the open-source read-later that sits in the same neighborhood as Instapaper and Pocket. Save from a browser, iOS, or Android; read offline; highlight; export to Markdown or Obsidian. After the hosted service wound down in early 2025, community forks kept the app running.

Where it falls short: the hosted flagship shut down. The forks work well but require a small self-host lift.

Pricing:

Platforms: Web app, browser extensions, native Windows/macOS/Linux via a wrapper.

Download: omnivore.app

Bottom line: The pick for a read-later with highlights and a clean export.

5. SingleFile, best for one-click page-to-HTML

SingleFile is a browser extension that turns any page into a single HTML file, images and all. Save it locally, drop it in a folder, and grep it later. It is the smallest tool on this list, and it works without a server or an account.

Where it falls short: it is a saver, not a library. No cross-device sync, no tagging, no search. Pair it with something like Recoll or a local desktop search.

Pricing:

Platforms: Firefox, Chrome, and Chromium-derived browsers on Windows, macOS, Linux.

Download: github.com/gildas-lormeau/SingleFile

Bottom line: The right one-click for anyone who wants a single portable file per page and nothing extra.

6. Readwise Reader, best for curated reading and highlights

Readwise Reader is the paid, hosted read-later that comes with Readwise’s highlight system. Save from anywhere, read in a distraction-free mode, and every highlight flows to the daily review or an Obsidian sync. Tagging is fast and search is quick.

Where it falls short: it is hosted-only, and the export is not one-click. A subscription lapse is a lock-out until the export runs.

Pricing:

Platforms: Web app, browser extensions, Windows/macOS/Linux via a desktop wrapper.

Download: readwise.io/read

Bottom line: The right pick for anyone who highlights heavily and lives in the Readwise flow.

7. Zotero, best for academic capture with citation

Zotero was built for research citation, and it does the archive job on the side. Save a page and it grabs the metadata, attaches a PDF or HTML snapshot, and puts it in a searchable library. Group libraries let a team share the same corpus.

Where it falls short: the UI feels like a bibliography manager because that is what it is. The archive is a byproduct.

Pricing:

Platforms: Windows, macOS, Linux.

Download: zotero.org

Bottom line: The overpowered but underrated choice when the reading is research and citations matter.

8. Pocket, best for a cross-platform save queue

Pocket is the hosted save queue Mozilla still runs. Save from anywhere, read on any device, and the export API returns a JSON of every saved URL. It stopped being interesting years ago, but the export path is clean and the archive holds.

Where it falls short: article parsing is aging. Long-form pages sometimes lose figure captions and structural HTML.

Pricing:

Platforms: Web app, browser extensions, native Windows/macOS via a wrapper.

Download: getpocket.com

Bottom line: A safe fallback for anyone already saving to Pocket, with a clean export path when the day comes.

How to pick the right one

FAQ

What is the difference between a read-later app and a web archive?

A read-later app queues a URL for later reading; a web archive keeps the page even if the URL dies. Wallabag, Karakeep, and ArchiveBox all keep the page. Pocket and Readwise Reader lean toward the queue.

Which one keeps offline copies without a server?

SingleFile is the simplest: install the extension, click save, and get a full-page HTML file. For a library with search, Karakeep and Wallabag both run locally in Docker.

Can I self-host on a Raspberry Pi or a NAS?

Wallabag, ArchiveBox, Karakeep, and community Omnivore forks all run in Docker on a Pi 4 or later, or on any NAS with Docker support.

Does any of these run without a browser extension?

ArchiveBox and Zotero both take URLs from the command line or a paste box. Wallabag and Karakeep have API endpoints so a scripted save works fine.

What happens if the app disappears?

Wallabag, ArchiveBox, Karakeep, and Omnivore are open-source with local storage, so an app shutdown does not lose data. Readwise Reader, Pocket, and hosted Wallabag ship export APIs. Only the closed cloud stack (Pocket without export) is a real risk.

Which of these handles paywalled pages?

None of them bypass a paywall. If a page is paywalled at capture time, only the fragment you can see gets archived. Save from your logged-in browser session for the paid content.