Short answer
To resume a large crawl safely, persist the crawl frontier and output state outside the worker process. Give every segment, page, URL, and record a stable identity; use explicit task states and expiring worker leases; write raw responses and parsed records idempotently; and reconcile all discovered work before delivery. A process restart should reclaim unfinished tasks—not rebuild the queue from memory or append duplicate rows.
Resumability is an architecture property. It cannot be added reliably by saving one page number at shutdown.
Why one checkpoint is not enough
A large crawl contains several progress streams:
- Segments or search partitions
- Listing pages or cursors
- Discovered detail URLs or record IDs
- Network attempts and retries
- Raw responses
- Parsed source observations
- Normalized records
- Export or delivery state
Saving “last page = 250” does not explain which earlier pages failed, which details remain, whether page 250 committed its output, or whether the source changed ordering before the restart.
Persist work at the same granularity that can fail and be retried independently.
The durable crawl frontier
A crawl frontier is the system of record for discovered work.
Each task should include:
- Stable
task_key - Source
- Task type
- Segment or parent task
- Normalized URL or request identity
- Pagination state
- Priority
- Status
- Attempt count
- Eligibility time for the next retry
- Lease owner and lease expiry
- First discovered and last updated times
- Terminal reason
A useful state model is:
| State | Meaning | Can be claimed? |
|---|---|---|
pending | Ready for work | Yes |
in_progress | Owned by a worker lease | No, until lease expires |
retry_wait | Temporary failure; waiting until eligible | Later |
completed | Output committed successfully | No |
failed_terminal | Retry budget exhausted or permanent failure | No |
excluded | Deliberately not collected, with reason | No |
Do not represent all terminal states as done. Delivery needs to distinguish completed work from missing or intentionally excluded work.
Create stable task keys
The same logical request may be discovered through several paths. A stable key prevents it from entering the queue repeatedly.
Examples:
source-a|segment|city=chicago&category=cardiology
source-a|listing|segment-123|cursor=abc
source-a|detail|provider-id=987
source-b|detail|normalized-url=https://example.com/item/42Build the key from normalized logical identity, not temporary values such as worker ID, attempt number, or timestamp.
URL normalization should be conservative. Removing a query parameter can merge two genuinely different records. Define allowed tracking parameters and identity rules per source.
Use idempotent task insertion
When a listing page discovers a detail URL, insert it with a unique constraint on the task key. If another page discovers the same URL, the insertion should leave one task rather than append a duplicate.
The same principle applies to source observations and normalized records. Use a source-aware unique key and an upsert or conditional insert.
INSERT INTO crawl_tasks (task_key, source, task_type, status)
VALUES (:task_key, :source, :task_type, 'pending')
ON CONFLICT (task_key) DO NOTHING;Idempotent does not always mean “ignore the duplicate.” For a recurring dataset, an existing entity may need a new source observation for the current run. Separate stable entity identity from run-specific observation identity.
Make output commits atomic
A task should be marked completed only after its required outputs have committed.
A safe transaction may:
- Store or reference the raw response.
- Write parsed source observations.
- Create newly discovered child tasks.
- Record validation results.
- Mark the parent task completed.
If the process stops before commit, the transaction rolls back and the task lease eventually expires. If it stops after commit, the completed state prevents duplicate processing.
When raw response storage is external, write it under a deterministic key and keep its checksum or version in the database before completion.
Use leases instead of permanent locks
A worker can crash while holding a task. A permanent in_progress flag then leaves that task stuck forever.
Use a lease with:
- Worker identity
- Claimed time
- Expiry time
- Optional heartbeat for long tasks
If the lease expires without a completion transaction, another worker can reclaim the task. Set the lease longer than normal task duration and extend it deliberately for known long operations.
Lease expiry should create a metric and log event. Frequent reclamation may indicate worker instability or underestimated task time.
Store pagination state explicitly
The checkpoint format depends on the source.
| Pagination model | State to persist |
|---|---|
| Page number | Segment, page number, page fingerprint, next-page state |
| Offset | Segment, offset, limit, sort, observed total |
| Cursor | Query identity, cursor used, next cursor received, hasNextPage |
| Infinite scroll | Underlying request state or browser checkpoint plus unique records |
| Sitemap | Sitemap URL, child index, URL task keys |
Never invent the next cursor after a restart. Store the cursor chain and the response that produced it where permitted.
Read Web Scraping Pagination Patterns for termination and coverage rules.
Separate discovery from extraction
Discovery and detail extraction should be different task types.
This allows the system to:
- Resume failed details without repeating successful listing pages
- Re-run discovery when source coverage changes
- Compare discovered IDs with parsed records
- Prioritize newly found work
- Reparse stored raw responses without network requests
- Track missing details separately from missing listings
In Orzaen's published Kramp crawler and GraphQL result, HTML discovery, URL validation, and GraphQL extraction were measured separately. That is the correct mental model for resumable large crawls.
Preserve raw responses where appropriate
Raw responses make recovery cheaper when parsing logic changes after collection.
Store or reference:
- Source and request identity
- Collection time
- Status and content type
- Response checksum
- Body or approved storage location
- Session-independent metadata
- Parser version later applied
Raw storage must still follow source, privacy, security, and retention requirements. Do not retain sensitive or unnecessary data merely because storage is available.
Make parsing repeatable
A parser should accept a raw observation and produce the same source record for the same parser version.
Record:
- Parser version
- Schema version
- Normalization rules version
- Validation result
- Exception details
If a website changes halfway through a crawl, stored responses and parser versions let you identify which records used each layout and reprocess the affected subset.
Handle retries without duplicate side effects
A retry can happen after the source responded but before the worker recorded success. The same task may therefore execute more than once.
Protect:
- Raw-response writes with deterministic keys
- Task insertion with unique constraints
- Source observations with run-specific identities
- Normalized records with explicit upsert rules
- Notifications and downstream actions with an idempotency ledger
- Exports with versioned artifact names and manifests
Do not rely on “the queue delivers each message once.” Design workers so repeated delivery is safe.
The 403, 429, and temporary-failure guide shows how to decide which tasks should re-enter the queue.
Resume browsers carefully
Browser state may include cookies, local storage, tokens, open pagination, and route affinity. Persist only what the source and security policy allow.
Prefer restarting from a durable logical task such as:
- Search segment plus cursor
- Detail URL
- Stable record ID
- Saved storage state tied to a controlled session
Trying to restore a browser from the exact scroll position is more fragile than recovering the underlying pagination request. Inspect the network model before making visual state the checkpoint.
Reconcile before delivery
A resumed run can finish without crashing and still lose work. Reconcile:
- Planned segments against terminal segments
- Discovered tasks against completed, failed, excluded, and deferred tasks
- Unique discovered record IDs against source observations
- Raw responses against parser outputs
- Accepted records against export rows
- Current snapshot against expected source totals or prior healthy ranges
Keep terminal failures visible. Do not delete them merely because most of the dataset completed.
A practical recovery sequence
When a worker or service restarts:
- Open a new run process against the existing
run_id. - Identify and expire stale leases after the defined timeout.
- Move eligible
retry_waittasks back topending. - Leave completed and terminal tasks untouched.
- Recreate only temporary worker and session state.
- Claim pending tasks transactionally.
- Continue producing idempotent outputs.
- Run full reconciliation before marking the run complete.
Do not clear the queue or reconstruct all URLs unless the run was deliberately abandoned and a new run is being created.
What to monitor during recovery
Track:
- Stale leases reclaimed
- Duplicate task insertions rejected
- Retry attempts
- Completed tasks reprocessed unexpectedly
- Raw-response checksum conflicts
- Upsert conflicts
- Queue depth by state
- Time since last completed task
- Reconciliation gaps
Connect these to Web Scraping Pipeline Monitoring so a resumed run is distinguishable from a clean first pass.
Common recovery mistakes
Keeping the queue only in memory
A process restart then loses discovered work and retry history. Persist frontier state externally.
Appending directly to CSV from many workers
Concurrent or repeated writes can corrupt order and create duplicates. Persist source observations first, then build a validated export.
Marking the task complete before writing output
A crash between those steps creates missing data that the queue considers finished. Commit output and completion atomically where possible.
Rebuilding the entire queue after every restart
This repeats discovery and can mix source changes into one run. Reclaim unfinished tasks from durable state.
Using URL alone as every record key
One URL may contain several records or change while the source identity remains stable. Define task, observation, and entity keys separately.
Frequently asked questions
Can Scrapy pause and resume a crawl?
Yes. Scrapy's job directory can persist scheduler state, duplicate filters, and crawl state for pausing and resuming. Custom pipelines still need durable output writes, compatible code, and careful handling of cookies or callbacks that may not serialize safely.
What is an idempotent scraper?
It is a scraper whose tasks and writes can be repeated without creating unintended duplicate work or records. Stable keys, unique constraints, deterministic raw storage, and explicit upsert rules make retries safe.
Should I save the last page number?
Save it as part of pagination evidence, but not as the only checkpoint. Persist every page or cursor task, discovered detail identity, retry state, and committed output so gaps can be recovered.
How do I avoid duplicate rows after a crash?
Write records to durable storage using source-aware unique keys and idempotent upserts. Build the final CSV or workbook from the validated database or observation store rather than appending blindly during collection.
Can I resume a crawl after the website changes?
Possibly, but first determine which responses and parsers were affected. Preserve raw responses and parser versions, pause incompatible tasks, update and test the parser, then reprocess the affected subset and reconcile the run.
Next step
Use the scalable web-scraping pipeline as the architecture overview, then connect durable state to pipeline monitoring and data validation before delivery. For a crawl that must survive long runs and recurring refreshes, review Orzaen's Web Scraping & Data Extraction service.
Sources
- Scrapy pausing and resuming crawls (opens in a new tab)
- Scrapy scheduler documentation (opens in a new tab)
- PostgreSQL INSERT documentation: ON CONFLICT (opens in a new tab)
- PostgreSQL transaction isolation (opens in a new tab)
- MongoDB unique indexes (opens in a new tab)
- MongoDB retryable writes (opens in a new tab)

