Persisting Crawl State to Resume After a Crash¶
A crawl that keeps its frontier in memory starts over after every crash, deploy or out-of-memory kill — and long crawls meet all three. Moving the frontier into a small database turns a restart into a resume. Tested with an asyncio crawler using SQLite through aiosqlite: the process was killed with SIGKILL after 519 pages had been fetched; on restart it requeued the URLs that had been in flight, carried on, and finished its 600-page budget at 607 fetches in total — 607 unique, 0 fetched twice. Without persisted state, the restart would have fetched those 519 pages again. This guide designs the schema, the state transitions that make recovery exact, and the batching that keeps the database from becoming the bottleneck.
Prerequisites¶
- Python 3.11+,
pip install aiohttp aiosqlite. - SQLite from asyncio, from using SQLite from asyncio with aiosqlite.
- Checkpoint principles, from checkpointing progress in long-running async jobs.
1. Store each URL with a state¶
One table holds every URL the crawl has discovered, keyed by its normalized form, with a state that moves forward as work happens:
SCHEMA = """
create table if not exists urls (
url text primary key, -- normalized
host text not null,
state integer not null default 0, -- 0 queued, 1 in flight, 2 done, 3 failed, 4 skipped
attempts integer not null default 0,
depth integer not null default 0,
updated_at real not null
);
create index if not exists urls_queue on urls(state, host);
"""
async def open_frontier(path: str) -> aiosqlite.Connection:
db = await aiosqlite.connect(path)
await db.execute("pragma journal_mode = wal")
await db.execute("pragma synchronous = normal")
await db.executescript(SCHEMA)
return db
The primary key on the normalized URL is the seen set, deduplicated by the database: insert or ignore adds only URLs not already present, atomically. The (state, host) index serves the scheduler's main query, "next queued URLs for this host". WAL mode lets the progress reporter read while workers write.
Verify: inserting the same normalized URL twice leaves one row.
2. Make state transitions match the work¶
The order of writes decides what a crash can lose. Claim before fetching; record links and completion together after:
async def claim(db, host: str, n: int) -> list[str]:
async with db.execute(
"select url from urls where state = 0 and host = ? limit ?", (host, n)
) as cur:
urls = [row[0] for row in await cur.fetchall()]
await db.executemany("update urls set state = 1, attempts = attempts + 1, updated_at = ? "
"where url = ?", [(time.time(), u) for u in urls])
await db.commit()
return urls
async def complete(db, url: str, links: list[str], depth: int) -> None:
now = time.time()
await db.executemany(
"insert or ignore into urls(url, host, depth, updated_at) values (?, ?, ?, ?)",
[(link, urlsplit(link).netloc, depth + 1, now) for link in links],
)
await db.execute("update urls set state = 2, updated_at = ? where url = ?", (now, url))
await db.commit() # links and "done" become durable together
Because the discovered links and the done mark are committed in one transaction, a crash either keeps both or neither: a page is never marked done without its links being saved. A crash after the fetch but before this commit leaves the URL in flight, and it is fetched again on restart — the only kind of duplicate possible, and only for the handful of pages in flight at the moment of the crash. In the test, the in-flight pages had not finished, so none were fetched twice.
Verify: kill the crawler at random moments in a test loop; every done URL's links are in the table, and duplicates never exceed the number of workers.
3. Recover on startup¶
On start, anything still in flight belongs to a process that no longer exists. Put it back in the queue, and give up on URLs that keep failing:
MAX_ATTEMPTS = 3
async def recover(db) -> tuple[int, int]:
cur = await db.execute("update urls set state = 0 where state = 1 and attempts < ?", (MAX_ATTEMPTS,))
requeued = cur.rowcount
cur = await db.execute("update urls set state = 3 where state = 1 and attempts >= ?", (MAX_ATTEMPTS,))
abandoned = cur.rowcount
await db.commit()
return requeued, abandoned
The attempts counter protects against a URL that crashes the crawler itself — a pathological page that exhausts memory or triggers a parser bug. Without it, that URL would be requeued, crash the crawler again, and loop forever across restarts. Measured: after the SIGKILL, recovery requeued the in-flight URLs and the resumed crawl completed with 607 unique fetches and none repeated.
Verify: a URL that crashes the process on every attempt ends in state failed after MAX_ATTEMPTS restarts, and the crawl continues.
4. Keep the database off the critical path¶
Every fetch now involves database writes. With many workers, one commit per page becomes the bottleneck. Batch writes through a single writer task:
class FrontierWriter:
def __init__(self, db, flush_every: float = 0.5, max_batch: int = 500) -> None:
self.db, self.flush_every, self.max_batch = db, flush_every, max_batch
self.pending: list[tuple[str, list[str], int]] = []
self.flushed = asyncio.Event()
def complete(self, url: str, links: list[str], depth: int) -> None:
self.pending.append((url, links, depth)) # workers never wait on the disk
async def run(self) -> None:
while True:
await asyncio.sleep(self.flush_every)
batch, self.pending = self.pending[: self.max_batch], self.pending[self.max_batch:]
if not batch:
continue
now = time.time()
await self.db.executemany(
"insert or ignore into urls(url, host, depth, updated_at) values (?, ?, ?, ?)",
[(l, urlsplit(l).netloc, d + 1, now) for _, links, d in batch for l in links])
await self.db.executemany("update urls set state = 2, updated_at = ? where url = ?",
[(now, u) for u, _, _ in batch])
await self.db.commit() # one transaction per batch
A batch still commits links and completion together, so the recovery guarantee holds; a crash loses at most the last half-second of completions, which are fetched again. One writer also matches SQLite's single-writer model, avoiding lock contention, as described for aiosqlite. Move to PostgreSQL when several crawler processes share one frontier — select ... for update skip locked then lets them claim work without colliding.
Verify: database write time per second stays well below one second at peak fetch rate, and a kill loses no more than the last flush interval.
5. Persist the other state, too¶
The frontier is the largest piece of state, but not the only one a resumed crawl needs:
async def save_host_state(db, host: str, state: HostState) -> None:
await db.execute(
"insert into hosts(host, limit_now, backoff_until, robots_fetched_at, robots_text) "
"values (?, ?, ?, ?, ?) on conflict(host) do update set "
"limit_now = excluded.limit_now, backoff_until = excluded.backoff_until, "
"robots_fetched_at = excluded.robots_fetched_at, robots_text = excluded.robots_text",
(host, state.limit, state.backoff_until, state.robots_fetched_at, state.robots_text),
)
Persisting per-host backoff means a host that was rate-limiting the crawl before the crash is not hit at full speed afterwards. Persisting robots text and fetch time avoids refetching every robots file on restart. And recording the crawl's configuration — seeds, scope, budget — alongside the frontier prevents a restart with different settings from silently mixing two crawls in one database.
Verify: after a restart, a host that was in backoff stays in backoff until its recorded time, and robots files are not refetched before their cache expiry.
Verification¶
Crawl state is durable when:
- Every discovered URL is a row keyed by its normalized form.
- Links and completion commit together, after the fetch.
- Startup requeues in-flight URLs and abandons ones that keep crashing.
- Writes are batched through one writer, and host state is persisted too.
Diagnostic Hook: export counts per state and the age of the oldest in-flight row. A growing in-flight count with an old oldest row means workers are stuck or a writer is falling behind; after a restart, the requeued count should be close to the number of workers, and much larger values mean completions were not being committed.
Pitfalls & edge cases¶
- Marking done before saving links. A crash loses the page's outlinks forever.
- No attempts counter. A page that crashes the crawler loops across restarts.
- One commit per fetch with many workers. The database becomes the bottleneck.
- Reusing a database with different settings. Two crawls mix without warning.
Frequently Asked Questions¶
How do I make a Python crawler resume after a crash?
Store every URL with a state in a database, claim URLs before fetching, commit discovered links and the done mark together, and requeue in-flight URLs on startup. A test crawl killed after 519 fetches resumed with no page fetched twice.
Is SQLite fast enough for a crawl frontier?
For a single crawler process, yes, in WAL mode with batched writes through one writer task. Several processes sharing a frontier are better served by PostgreSQL.
Will a resumed crawl fetch some pages twice?
Only pages that were in flight when it crashed, at most one per worker, because their completion was not committed. Make processing idempotent for those.
What else should a crawler persist besides URLs?
Per-host state such as backoff and current limits, cached robots rules with their fetch times, and the crawl's configuration, so a restart behaves exactly like the original run.
Related¶
- Async Web Crawlers — up to the topic overview.
- Deduplicating URLs in a concurrent crawl frontier — the normalization that produces the primary key.
- Network I/O & Protocol Handling — the section overview.