Healthcare Provider Data Collection — Custom physician and provider roster datasets for healthcare research, consulting, analytics, and health-tech projects.Learn more
OrzaenOrzaen

How to Resume a Large Web Crawl Without Duplicating Work

Engineering2026-07-1813 minAhmad Raza

A durable crawl-state design using stable task keys, checkpoints, idempotent writes, retry queues, leases, raw responses, and final reconciliation.

Short answer

To resume a large crawl safely, persist the crawl frontier and output state outside the worker process. Give every segment, page, URL, and record a stable identity; use explicit task states and expiring worker leases; write raw responses and parsed records idempotently; and reconcile all discovered work before delivery. A process restart should reclaim unfinished tasks—not rebuild the queue from memory or append duplicate rows.

Resumability is an architecture property. It cannot be added reliably by saving one page number at shutdown.

Why one checkpoint is not enough

A large crawl contains several progress streams:

  • Segments or search partitions
  • Listing pages or cursors
  • Discovered detail URLs or record IDs
  • Network attempts and retries
  • Raw responses
  • Parsed source observations
  • Normalized records
  • Export or delivery state

Saving “last page = 250” does not explain which earlier pages failed, which details remain, whether page 250 committed its output, or whether the source changed ordering before the restart.

Persist work at the same granularity that can fail and be retried independently.

The durable crawl frontier

A crawl frontier is the system of record for discovered work.

Each task should include:

  • Stable task_key
  • Source
  • Task type
  • Segment or parent task
  • Normalized URL or request identity
  • Pagination state
  • Priority
  • Status
  • Attempt count
  • Eligibility time for the next retry
  • Lease owner and lease expiry
  • First discovered and last updated times
  • Terminal reason

A useful state model is:

StateMeaningCan be claimed?
pendingReady for workYes
in_progressOwned by a worker leaseNo, until lease expires
retry_waitTemporary failure; waiting until eligibleLater
completedOutput committed successfullyNo
failed_terminalRetry budget exhausted or permanent failureNo
excludedDeliberately not collected, with reasonNo

Do not represent all terminal states as done. Delivery needs to distinguish completed work from missing or intentionally excluded work.

Create stable task keys

The same logical request may be discovered through several paths. A stable key prevents it from entering the queue repeatedly.

Examples:

text
source-a|segment|city=chicago&category=cardiology
source-a|listing|segment-123|cursor=abc
source-a|detail|provider-id=987
source-b|detail|normalized-url=https://example.com/item/42

Build the key from normalized logical identity, not temporary values such as worker ID, attempt number, or timestamp.

URL normalization should be conservative. Removing a query parameter can merge two genuinely different records. Define allowed tracking parameters and identity rules per source.

Use idempotent task insertion

When a listing page discovers a detail URL, insert it with a unique constraint on the task key. If another page discovers the same URL, the insertion should leave one task rather than append a duplicate.

The same principle applies to source observations and normalized records. Use a source-aware unique key and an upsert or conditional insert.

sql
INSERT INTO crawl_tasks (task_key, source, task_type, status)
VALUES (:task_key, :source, :task_type, 'pending')
ON CONFLICT (task_key) DO NOTHING;

Idempotent does not always mean “ignore the duplicate.” For a recurring dataset, an existing entity may need a new source observation for the current run. Separate stable entity identity from run-specific observation identity.

Make output commits atomic

A task should be marked completed only after its required outputs have committed.

A safe transaction may:

  1. Store or reference the raw response.
  2. Write parsed source observations.
  3. Create newly discovered child tasks.
  4. Record validation results.
  5. Mark the parent task completed.

If the process stops before commit, the transaction rolls back and the task lease eventually expires. If it stops after commit, the completed state prevents duplicate processing.

When raw response storage is external, write it under a deterministic key and keep its checksum or version in the database before completion.

Use leases instead of permanent locks

A worker can crash while holding a task. A permanent in_progress flag then leaves that task stuck forever.

Use a lease with:

  • Worker identity
  • Claimed time
  • Expiry time
  • Optional heartbeat for long tasks

If the lease expires without a completion transaction, another worker can reclaim the task. Set the lease longer than normal task duration and extend it deliberately for known long operations.

Lease expiry should create a metric and log event. Frequent reclamation may indicate worker instability or underestimated task time.

Store pagination state explicitly

The checkpoint format depends on the source.

Pagination modelState to persist
Page numberSegment, page number, page fingerprint, next-page state
OffsetSegment, offset, limit, sort, observed total
CursorQuery identity, cursor used, next cursor received, hasNextPage
Infinite scrollUnderlying request state or browser checkpoint plus unique records
SitemapSitemap URL, child index, URL task keys

Never invent the next cursor after a restart. Store the cursor chain and the response that produced it where permitted.

Read Web Scraping Pagination Patterns for termination and coverage rules.

Separate discovery from extraction

Discovery and detail extraction should be different task types.

This allows the system to:

  • Resume failed details without repeating successful listing pages
  • Re-run discovery when source coverage changes
  • Compare discovered IDs with parsed records
  • Prioritize newly found work
  • Reparse stored raw responses without network requests
  • Track missing details separately from missing listings

In Orzaen's published Kramp crawler and GraphQL result, HTML discovery, URL validation, and GraphQL extraction were measured separately. That is the correct mental model for resumable large crawls.

Preserve raw responses where appropriate

Raw responses make recovery cheaper when parsing logic changes after collection.

Store or reference:

  • Source and request identity
  • Collection time
  • Status and content type
  • Response checksum
  • Body or approved storage location
  • Session-independent metadata
  • Parser version later applied

Raw storage must still follow source, privacy, security, and retention requirements. Do not retain sensitive or unnecessary data merely because storage is available.

Make parsing repeatable

A parser should accept a raw observation and produce the same source record for the same parser version.

Record:

  • Parser version
  • Schema version
  • Normalization rules version
  • Validation result
  • Exception details

If a website changes halfway through a crawl, stored responses and parser versions let you identify which records used each layout and reprocess the affected subset.

Handle retries without duplicate side effects

A retry can happen after the source responded but before the worker recorded success. The same task may therefore execute more than once.

Protect:

  • Raw-response writes with deterministic keys
  • Task insertion with unique constraints
  • Source observations with run-specific identities
  • Normalized records with explicit upsert rules
  • Notifications and downstream actions with an idempotency ledger
  • Exports with versioned artifact names and manifests

Do not rely on “the queue delivers each message once.” Design workers so repeated delivery is safe.

The 403, 429, and temporary-failure guide shows how to decide which tasks should re-enter the queue.

Resume browsers carefully

Browser state may include cookies, local storage, tokens, open pagination, and route affinity. Persist only what the source and security policy allow.

Prefer restarting from a durable logical task such as:

  • Search segment plus cursor
  • Detail URL
  • Stable record ID
  • Saved storage state tied to a controlled session

Trying to restore a browser from the exact scroll position is more fragile than recovering the underlying pagination request. Inspect the network model before making visual state the checkpoint.

Reconcile before delivery

A resumed run can finish without crashing and still lose work. Reconcile:

  • Planned segments against terminal segments
  • Discovered tasks against completed, failed, excluded, and deferred tasks
  • Unique discovered record IDs against source observations
  • Raw responses against parser outputs
  • Accepted records against export rows
  • Current snapshot against expected source totals or prior healthy ranges

Keep terminal failures visible. Do not delete them merely because most of the dataset completed.

A practical recovery sequence

When a worker or service restarts:

  1. Open a new run process against the existing run_id.
  2. Identify and expire stale leases after the defined timeout.
  3. Move eligible retry_wait tasks back to pending.
  4. Leave completed and terminal tasks untouched.
  5. Recreate only temporary worker and session state.
  6. Claim pending tasks transactionally.
  7. Continue producing idempotent outputs.
  8. Run full reconciliation before marking the run complete.

Do not clear the queue or reconstruct all URLs unless the run was deliberately abandoned and a new run is being created.

What to monitor during recovery

Track:

  • Stale leases reclaimed
  • Duplicate task insertions rejected
  • Retry attempts
  • Completed tasks reprocessed unexpectedly
  • Raw-response checksum conflicts
  • Upsert conflicts
  • Queue depth by state
  • Time since last completed task
  • Reconciliation gaps

Connect these to Web Scraping Pipeline Monitoring so a resumed run is distinguishable from a clean first pass.

Common recovery mistakes

Keeping the queue only in memory

A process restart then loses discovered work and retry history. Persist frontier state externally.

Appending directly to CSV from many workers

Concurrent or repeated writes can corrupt order and create duplicates. Persist source observations first, then build a validated export.

Marking the task complete before writing output

A crash between those steps creates missing data that the queue considers finished. Commit output and completion atomically where possible.

Rebuilding the entire queue after every restart

This repeats discovery and can mix source changes into one run. Reclaim unfinished tasks from durable state.

Using URL alone as every record key

One URL may contain several records or change while the source identity remains stable. Define task, observation, and entity keys separately.

Frequently asked questions

Can Scrapy pause and resume a crawl?

Yes. Scrapy's job directory can persist scheduler state, duplicate filters, and crawl state for pausing and resuming. Custom pipelines still need durable output writes, compatible code, and careful handling of cookies or callbacks that may not serialize safely.

What is an idempotent scraper?

It is a scraper whose tasks and writes can be repeated without creating unintended duplicate work or records. Stable keys, unique constraints, deterministic raw storage, and explicit upsert rules make retries safe.

Should I save the last page number?

Save it as part of pagination evidence, but not as the only checkpoint. Persist every page or cursor task, discovered detail identity, retry state, and committed output so gaps can be recovered.

How do I avoid duplicate rows after a crash?

Write records to durable storage using source-aware unique keys and idempotent upserts. Build the final CSV or workbook from the validated database or observation store rather than appending blindly during collection.

Can I resume a crawl after the website changes?

Possibly, but first determine which responses and parsers were affected. Preserve raw responses and parser versions, pause incompatible tasks, update and test the parser, then reprocess the affected subset and reconcile the run.

Next step

Use the scalable web-scraping pipeline as the architecture overview, then connect durable state to pipeline monitoring and data validation before delivery. For a crawl that must survive long runs and recurring refreshes, review Orzaen's Web Scraping & Data Extraction service.

Sources

Tags

Web ScrapingCrawl RecoveryIdempotencyData Pipelines

Need this fixed?

Have the same system problem?

Share what is manual, messy, broken, or disconnected. We’ll review the cleanest next step.

Get System Review