Healthcare Provider Data Collection — Custom physician and provider roster datasets for healthcare research, consulting, analytics, and health-tech projects.Learn more
OrzaenOrzaen

How to Build a Scalable Web Scraping Pipeline

Engineering2026-07-1814 minAhmad Raza

A production architecture for URL discovery, extraction, raw evidence, validation, failure recovery, and repeatable delivery across large web crawls.

Short answer

A scalable web scraping pipeline separates source discovery, fetching, parsing, normalization, validation, storage, and delivery. It keeps durable crawl state, records why every URL succeeded or failed, preserves raw source observations where appropriate, and measures coverage against an expected set. Scale is not merely sending more concurrent requests. It is being able to stop, resume, inspect, and repeat a crawl without losing control of the dataset.

For a one-time list of 500 stable pages, a single script may be enough. For hundreds of thousands of pages, several source types, recurring updates, or a dataset that affects business decisions, the surrounding system matters more than the first selector.

What makes a scraping pipeline scalable?

Four properties matter.

  • Throughput: it processes the required volume within the available collection window.
  • Recoverability: a worker, proxy, session, or process can fail without forcing the crawl to restart.
  • Observability: the team can see coverage, failures, missing fields, response patterns, and changes.
  • Maintainability: source-specific logic can change without rewriting the queue, schema, validation, or delivery system.

A fast crawler that cannot explain 12,000 missing records is not scalable. Neither is a browser farm that finishes the crawl but costs more to run than the dataset is worth.

The production pipeline at a glance

LayerMain responsibilityDurable output
Scope and source reviewDefine approved sources, target entities, fields, freshness, and exclusionsSource register and data contract
DiscoveryFind canonical listing and detail URLsCrawl frontier
FetchingRetrieve HTML or structured responses with source-aware limitsResponse metadata and raw artifacts
ParsingTurn source-specific responses into source-shaped recordsParsed observations
NormalizationMap observations into a canonical schemaNormalized records plus raw values
ValidationTest structure, meaning, coverage, and reconciliationValid records and exception records
StoragePreserve state, lineage, and current or historical snapshotsDatabase, object storage, or files
DeliveryProduce the format and cadence the consumer needsCSV, Excel, JSON, API, database, or webhook
MonitoringTrack operational and data-quality signalsMetrics, alerts, and run summaries

These layers can run in one process at small scale. They should still remain separate concepts, because each fails differently.

1. Define the data contract before writing the crawler

Start with the dataset, not the technology.

For every requested field, define:

  • Meaning
  • Source location
  • Data type
  • Whether it is required or optional
  • Whether it may contain several values
  • How absence should be represented
  • Whether it can be normalized
  • Which raw value must be retained

For example, price is ambiguous until the contract answers whether it means list price, sale price, member price, price range, or the price for a particular variant and currency.

The same applies to coverage. "Collect all products" needs an expected universe: every URL in an approved sitemap, every leaf category reached from the catalog, every identifier supplied by the client, or another testable definition.

A useful source register includes:

FieldExample
source_idsupplier_de_014
base_urlApproved source domain
source_familyMagento catalog
discovery_methodSitemap plus category traversal
extraction_surfaceProduct JSON response
refresh_cadenceWeekly
rate_policySource-specific concurrency and delay
ownerMaintainer responsible for the adapter
last_verified_atLast manual source review

Before collection, review source terms, access instructions, robots directives, privacy requirements, permitted use, and rate expectations. The Robots Exclusion Protocol (opens in a new tab) defines how crawlers are requested to treat URI paths, but it is not an authorization system or a complete legal review.

2. Choose the cheapest reliable data surface

Do not begin by choosing Selenium, Playwright, or Requests. First determine where the record exists.

The useful data may be available in:

  • Server-rendered HTML
  • Embedded JSON in the document
  • JSON-LD or another structured-data block
  • An XHR or fetch response
  • A GraphQL response
  • A downloadable file
  • A page that genuinely requires browser state or interaction

The Chrome DevTools Network panel (opens in a new tab) can show document, XHR, fetch, JavaScript, and other requests made while a page loads. Playwright can also monitor or handle page network activity through its network APIs (opens in a new tab).

The usual preference is:

  1. An approved official API or export
  2. A stable structured response used by the public page
  3. Server-rendered HTML
  4. Browser automation for the interactions or state that cannot be reproduced more simply

This is not a rule that APIs are always best. An API may omit fields visible in HTML, apply restrictive pagination, or change independently. Source selection needs a field-by-field comparison.

Read API vs HTML vs Browser Automation for the diagnostic process, Requests vs Playwright vs Selenium for execution trade-offs, and How to Scrape JavaScript Websites Without Using a Browser for Every Request for the hybrid implementation pattern.

3. Separate URL discovery from detail extraction

Large crawls are easier to reason about when discovery produces a durable frontier rather than immediately extracting every record inside one nested loop.

A crawl frontier should identify at least:

  • Canonical URL or source key
  • Source and entity type
  • Discovery path
  • Priority
  • Current status
  • Attempt count
  • Last HTTP status or failure class
  • Next eligible attempt time
  • First and last seen timestamps
  • Parser version

Typical statuses are discovered, scheduled, in_progress, fetched, parsed, validated, failed_retryable, failed_terminal, and excluded.

This state prevents three common failures:

  • A crash discards every URL not yet processed.
  • Retrying the whole crawl creates duplicate records.
  • The final row count cannot be reconciled with the discovered universe.

Scrapy supports persisted scheduled requests, a persisted duplicate filter, and spider state through its `JOBDIR` persistence facilities (opens in a new tab). The exact implementation may instead use PostgreSQL, MongoDB, Redis, a queue service, or files. What matters is that pending work and completed work survive the process.

For task keys, worker leases, idempotent writes, and restart behavior, see How to Resume a Large Web Crawl Without Duplicating Work.

What Orzaen's Kramp crawl demonstrates

In Orzaen's published enterprise e-commerce crawl and GraphQL extraction result, URL discovery, HTML crawling, GraphQL response capture, deduplication, and validation were treated as separate responsibilities. The system processed more than 1.25 million HTML pages and 5 million GraphQL responses while validating more than 864,000 product and category URLs.

The lesson is not the headline volume. It is that discovery and structured extraction created different artifacts and needed different reconciliation rules.

4. Design fetching around each source

Concurrency should be a source policy, not one global number.

For every source or domain, configure:

  • Maximum concurrency
  • Request delay or adaptive throttle
  • Connection and read timeouts
  • Retryable status codes and exceptions
  • Maximum attempts
  • Session or cookie requirements
  • Cache behavior
  • Proxy requirements, if legitimately needed
  • Maximum response size
  • Allowed content types

HTTP 429 Too Many Requests means the client has exceeded a rate limit. The response may include Retry-After, which can specify either an HTTP date or a delay in seconds under HTTP semantics (opens in a new tab). A production client should record the response and respect source feedback rather than instantly retrying from every worker.

Retries need classification.

FailureTypical action
Connection reset or temporary DNS errorRetry with bounded backoff
429 with Retry-AfterPause according to the response and reduce pressure
Temporary 5xx responseRetry within a limit
404 for a previously valid itemRecord as a source observation; do not retry forever
Login expiredRefresh the approved session, then retry once or by policy
Parser produced no required fieldsSend to data-quality exceptions, not the network retry loop
Access is no longer permittedStop that source and escalate for review

Network success and extraction success are separate. A page can return 200 with an error template, consent screen, login page, or empty shell.

Use the 403, 429, and temporary-failure guide to define failure classes and retry budgets. If the source design genuinely requires distributed routes, add the proxy rotation architecture only after source policy, rate control, and session affinity are understood.

5. Use a hybrid HTTP-and-browser architecture when appropriate

Browser automation is valuable for login, interactive discovery, JavaScript state, and pages whose data cannot be obtained reliably another way. It is also resource-intensive compared with direct HTTP requests.

A hybrid pattern is often more efficient:

  1. Start a browser for the approved state or discovery step.
  2. Observe the relevant requests and required session context.
  3. Reuse the structured response or HTTP surface where appropriate.
  4. Return to the browser only when state must be refreshed or an interaction is required.

In Orzaen's YouTube hybrid proxy rotation result, browser-based discovery was separated from HTTP extraction. The published result reports an approximately fivefold speed improvement over the earlier browser-heavy approach.

The correct tool depends on the source. The optimization is not "remove every browser"; it is "pay the browser cost only where the browser provides necessary value."

Pagination is part of this source decision. Page numbers, offsets, cursors, infinite scroll, and capped searches need distinct termination and reconciliation rules; see Web Scraping Pagination Patterns.

6. Preserve raw observations before normalizing

The parser should produce a source observation before a canonical record.

An observation can include:

json
{
  "source_id": "catalog_014",
  "source_url": "https://example.com/item/123",
  "collected_at": "2026-07-22T08:42:11Z",
  "parser_version": "catalog_014@7",
  "source_key": "123",
  "title_raw": "18V Drill — Tool Only",
  "price_raw": "€129,95",
  "availability_raw": "Available in 2–4 days",
  "raw_artifact_key": "responses/2026-07-22/abc123.json.gz"
}

Normalization can then produce price_amount: 129.95, currency: EUR, and an agreed availability category without destroying the source text.

Raw HTML or JSON snapshots are useful when permitted and practical because they allow:

  • Parser changes without another network request
  • Investigation of a disputed field
  • Comparison before and after a source change
  • Evidence of what the source returned at collection time

Storage policies should account for source rules, sensitive data, size, retention, and security. Keeping everything forever is not automatically responsible architecture.

7. Make parsing deterministic and versioned

A parser should accept a known input and produce the same observation under the same version.

Keep source-specific extraction inside adapters:

text
fetch result
  -> identify response type
  -> source adapter
  -> source observation
  -> canonical mapping
  -> validation

Avoid spreading selectors across queue workers, delivery scripts, and normalization notebooks. When a source changes, the team should be able to identify which adapter and parser version created every affected record.

For heterogeneous sources, use a source-family template with explicit overrides. Do not force unrelated websites through one universal selector configuration if their semantics differ.

8. Validate the dataset at four levels

Validation is not a final dropna() call.

Structural validation

Check types, required fields, allowed values, formats, and nested shapes. JSON Schema (opens in a new tab) defines a vocabulary for asserting structural constraints on JSON instances.

Semantic validation

Check whether the value makes sense for this field and source: a price is non-negative, a latitude falls within range, a date parses under the documented format, or a product identifier follows the source's expected pattern.

Coverage validation

Compare discovered, attempted, fetched, parsed, valid, invalid, and excluded entities. A non-empty file can still omit an entire category, state, alphabet segment, or pagination branch.

Reconciliation validation

Ensure output records correspond to the requested input or frontier. This is particularly important for ordered input enrichment, where retries or asynchronous workers can otherwise shift results.

The Instagram bulk enrichment result describes preserving input order while processing more than 400,000 handles with batching and retries. That is a reconciliation requirement, not merely a formatting preference.

See How to Validate Scraped Data Before Delivery for a full QA model.

9. Monitor data signals as well as infrastructure

CPU, memory, and worker health are necessary, but they will not detect a parser that returns empty strings successfully.

Track metrics such as:

  • URLs discovered, attempted, and completed
  • Response count by status and content type
  • Retry and terminal-failure rates
  • Records parsed and validated
  • Required-field coverage by source
  • Duplicate keys
  • Records per listing page or category
  • Response-size and latency distributions
  • Parser version in use
  • New, changed, missing, and unchanged records on refresh

Scrapy's Stats Collector (opens in a new tab) provides counters and key/value statistics inside a crawl. Whatever framework is used, record both operational and dataset metrics.

Alerts should compare against a source baseline. Ten empty phone fields might be normal for one directory and a critical regression for another.

How to Detect Website Changes Before a Scraper Silently Loses Data covers sentinel pages, field-coverage drift, response signatures, and the recovery workflow.

Web Scraping Pipeline Monitoring expands this into run identity, crawl-frontier progress, retry backlogs, data-quality baselines, and delivery verification.

10. Design delivery for the consumer

CSV and Excel are appropriate when a person needs to review, filter, or import a bounded dataset. JSON or a database is better when relationships, nested values, history, or repeated updates matter. An API or webhook makes sense when another product needs programmatic access.

Every delivery should include a run summary:

  • Scope and collection window
  • Source list
  • Output schema
  • Record counts by source and status
  • Required-field coverage
  • Duplicate and exception counts
  • Known limitations
  • Refresh or change logic
  • Source URL and timestamp fields

Do not hide invalid records to make the final file look clean. Deliver a valid dataset plus a separate exception report when exceptions matter to coverage.

A practical reference architecture

text
source register
  -> discovery workers
  -> durable frontier
  -> source-aware fetch workers
  -> raw artifact storage
  -> versioned source parsers
  -> canonical normalization
  -> validation and reconciliation
  -> current snapshot + history
  -> file, database, API, or webhook delivery
  -> run summary, metrics, and alerts

At smaller scale, SQLite and object files may implement several boxes. At larger scale, PostgreSQL or MongoDB, a queue, object storage, and independent workers may be appropriate. Choose infrastructure from the recovery, throughput, and delivery requirements—not from a desire to make the diagram look distributed.

Common architecture mistakes

Scaling concurrency before measuring coverage

More workers can make an incomplete crawl finish faster. Establish the expected URL universe and reconciliation first.

Treating every failure as a retry

A changed selector, denied source, invalid input, and temporary network error require different actions. Blind retries increase cost and hide the cause.

Using browser automation for every request

This can be correct for a small interactive source, but expensive for a large crawl. Inspect the network and page source before deciding.

Overwriting raw values

If normalized values replace source text, later audits and mapping changes become difficult. Keep both.

Delivering only the successful rows

Without the expected count and exception set, the client cannot distinguish complete collection from silent loss.

Building one parser for thousands of unrelated sites

Shared infrastructure is valuable; false uniformity is not. Use source families, adapters, and explicit overrides.

For the full source-portfolio design, read How to Scrape Thousands of Websites With Different Structures.

Frequently asked questions

Do I need Scrapy for a scalable web scraper?

No. Scrapy provides a mature crawling framework, asynchronous scheduling, statistics, throttling, and job persistence, but a scalable pipeline can also be built from HTTP clients, browser workers, queues, and databases. The architecture and operating controls matter more than the framework name.

Should scraped responses be stored?

Store raw responses when the source permits it and the audit, reprocessing, or change-detection value justifies the cost. Define retention and security rules instead of keeping all content indefinitely by default.

How many concurrent requests should a scraper send?

There is no universal safe number. Set concurrency per source based on its published limits, response behavior, terms, server feedback, collection window, and approved use. Track latency and error changes, and reduce pressure when the source signals rate limiting or instability.

How can a large crawl resume after a crash?

Persist the crawl frontier, visited set, attempt count, and output state outside the worker process. Restart workers from pending or retryable records rather than regenerating the entire crawl blindly.

What is the most important web-scraping metric?

For a dataset project, it is usually coverage against an explicit expected set—not requests per second. Throughput matters only after the team can explain what was discovered, attempted, delivered, excluded, and missed.

Next step

If your project involves a large source, several websites, recurring refreshes, or a dataset that must be reconciled, start with Orzaen's web scraping and data extraction service. Use the pipeline monitoring guide and crawl-resumption guide as the operational continuation. For concrete architecture examples, review the Kramp crawl and GraphQL result, VRBO property extraction result, and YouTube hybrid extraction result.

Sources

Tags

Web ScrapingData PipelinePythonData Quality

Need this fixed?

Have the same system problem?

Share what is manual, messy, broken, or disconnected. We’ll review the cleanest next step.

Get System Review