Short answer
A scalable web scraping pipeline separates source discovery, fetching, parsing, normalization, validation, storage, and delivery. It keeps durable crawl state, records why every URL succeeded or failed, preserves raw source observations where appropriate, and measures coverage against an expected set. Scale is not merely sending more concurrent requests. It is being able to stop, resume, inspect, and repeat a crawl without losing control of the dataset.
For a one-time list of 500 stable pages, a single script may be enough. For hundreds of thousands of pages, several source types, recurring updates, or a dataset that affects business decisions, the surrounding system matters more than the first selector.
What makes a scraping pipeline scalable?
Four properties matter.
- Throughput: it processes the required volume within the available collection window.
- Recoverability: a worker, proxy, session, or process can fail without forcing the crawl to restart.
- Observability: the team can see coverage, failures, missing fields, response patterns, and changes.
- Maintainability: source-specific logic can change without rewriting the queue, schema, validation, or delivery system.
A fast crawler that cannot explain 12,000 missing records is not scalable. Neither is a browser farm that finishes the crawl but costs more to run than the dataset is worth.
The production pipeline at a glance
| Layer | Main responsibility | Durable output |
|---|---|---|
| Scope and source review | Define approved sources, target entities, fields, freshness, and exclusions | Source register and data contract |
| Discovery | Find canonical listing and detail URLs | Crawl frontier |
| Fetching | Retrieve HTML or structured responses with source-aware limits | Response metadata and raw artifacts |
| Parsing | Turn source-specific responses into source-shaped records | Parsed observations |
| Normalization | Map observations into a canonical schema | Normalized records plus raw values |
| Validation | Test structure, meaning, coverage, and reconciliation | Valid records and exception records |
| Storage | Preserve state, lineage, and current or historical snapshots | Database, object storage, or files |
| Delivery | Produce the format and cadence the consumer needs | CSV, Excel, JSON, API, database, or webhook |
| Monitoring | Track operational and data-quality signals | Metrics, alerts, and run summaries |
These layers can run in one process at small scale. They should still remain separate concepts, because each fails differently.
1. Define the data contract before writing the crawler
Start with the dataset, not the technology.
For every requested field, define:
- Meaning
- Source location
- Data type
- Whether it is required or optional
- Whether it may contain several values
- How absence should be represented
- Whether it can be normalized
- Which raw value must be retained
For example, price is ambiguous until the contract answers whether it means list price, sale price, member price, price range, or the price for a particular variant and currency.
The same applies to coverage. "Collect all products" needs an expected universe: every URL in an approved sitemap, every leaf category reached from the catalog, every identifier supplied by the client, or another testable definition.
A useful source register includes:
| Field | Example |
|---|---|
source_id | supplier_de_014 |
base_url | Approved source domain |
source_family | Magento catalog |
discovery_method | Sitemap plus category traversal |
extraction_surface | Product JSON response |
refresh_cadence | Weekly |
rate_policy | Source-specific concurrency and delay |
owner | Maintainer responsible for the adapter |
last_verified_at | Last manual source review |
Before collection, review source terms, access instructions, robots directives, privacy requirements, permitted use, and rate expectations. The Robots Exclusion Protocol (opens in a new tab) defines how crawlers are requested to treat URI paths, but it is not an authorization system or a complete legal review.
2. Choose the cheapest reliable data surface
Do not begin by choosing Selenium, Playwright, or Requests. First determine where the record exists.
The useful data may be available in:
- Server-rendered HTML
- Embedded JSON in the document
- JSON-LD or another structured-data block
- An XHR or
fetchresponse - A GraphQL response
- A downloadable file
- A page that genuinely requires browser state or interaction
The Chrome DevTools Network panel (opens in a new tab) can show document, XHR, fetch, JavaScript, and other requests made while a page loads. Playwright can also monitor or handle page network activity through its network APIs (opens in a new tab).
The usual preference is:
- An approved official API or export
- A stable structured response used by the public page
- Server-rendered HTML
- Browser automation for the interactions or state that cannot be reproduced more simply
This is not a rule that APIs are always best. An API may omit fields visible in HTML, apply restrictive pagination, or change independently. Source selection needs a field-by-field comparison.
Read API vs HTML vs Browser Automation for the diagnostic process, Requests vs Playwright vs Selenium for execution trade-offs, and How to Scrape JavaScript Websites Without Using a Browser for Every Request for the hybrid implementation pattern.
3. Separate URL discovery from detail extraction
Large crawls are easier to reason about when discovery produces a durable frontier rather than immediately extracting every record inside one nested loop.
A crawl frontier should identify at least:
- Canonical URL or source key
- Source and entity type
- Discovery path
- Priority
- Current status
- Attempt count
- Last HTTP status or failure class
- Next eligible attempt time
- First and last seen timestamps
- Parser version
Typical statuses are discovered, scheduled, in_progress, fetched, parsed, validated, failed_retryable, failed_terminal, and excluded.
This state prevents three common failures:
- A crash discards every URL not yet processed.
- Retrying the whole crawl creates duplicate records.
- The final row count cannot be reconciled with the discovered universe.
Scrapy supports persisted scheduled requests, a persisted duplicate filter, and spider state through its `JOBDIR` persistence facilities (opens in a new tab). The exact implementation may instead use PostgreSQL, MongoDB, Redis, a queue service, or files. What matters is that pending work and completed work survive the process.
For task keys, worker leases, idempotent writes, and restart behavior, see How to Resume a Large Web Crawl Without Duplicating Work.
What Orzaen's Kramp crawl demonstrates
In Orzaen's published enterprise e-commerce crawl and GraphQL extraction result, URL discovery, HTML crawling, GraphQL response capture, deduplication, and validation were treated as separate responsibilities. The system processed more than 1.25 million HTML pages and 5 million GraphQL responses while validating more than 864,000 product and category URLs.
The lesson is not the headline volume. It is that discovery and structured extraction created different artifacts and needed different reconciliation rules.
4. Design fetching around each source
Concurrency should be a source policy, not one global number.
For every source or domain, configure:
- Maximum concurrency
- Request delay or adaptive throttle
- Connection and read timeouts
- Retryable status codes and exceptions
- Maximum attempts
- Session or cookie requirements
- Cache behavior
- Proxy requirements, if legitimately needed
- Maximum response size
- Allowed content types
HTTP 429 Too Many Requests means the client has exceeded a rate limit. The response may include Retry-After, which can specify either an HTTP date or a delay in seconds under HTTP semantics (opens in a new tab). A production client should record the response and respect source feedback rather than instantly retrying from every worker.
Retries need classification.
| Failure | Typical action |
|---|---|
| Connection reset or temporary DNS error | Retry with bounded backoff |
429 with Retry-After | Pause according to the response and reduce pressure |
Temporary 5xx response | Retry within a limit |
404 for a previously valid item | Record as a source observation; do not retry forever |
| Login expired | Refresh the approved session, then retry once or by policy |
| Parser produced no required fields | Send to data-quality exceptions, not the network retry loop |
| Access is no longer permitted | Stop that source and escalate for review |
Network success and extraction success are separate. A page can return 200 with an error template, consent screen, login page, or empty shell.
Use the 403, 429, and temporary-failure guide to define failure classes and retry budgets. If the source design genuinely requires distributed routes, add the proxy rotation architecture only after source policy, rate control, and session affinity are understood.
5. Use a hybrid HTTP-and-browser architecture when appropriate
Browser automation is valuable for login, interactive discovery, JavaScript state, and pages whose data cannot be obtained reliably another way. It is also resource-intensive compared with direct HTTP requests.
A hybrid pattern is often more efficient:
- Start a browser for the approved state or discovery step.
- Observe the relevant requests and required session context.
- Reuse the structured response or HTTP surface where appropriate.
- Return to the browser only when state must be refreshed or an interaction is required.
In Orzaen's YouTube hybrid proxy rotation result, browser-based discovery was separated from HTTP extraction. The published result reports an approximately fivefold speed improvement over the earlier browser-heavy approach.
The correct tool depends on the source. The optimization is not "remove every browser"; it is "pay the browser cost only where the browser provides necessary value."
Pagination is part of this source decision. Page numbers, offsets, cursors, infinite scroll, and capped searches need distinct termination and reconciliation rules; see Web Scraping Pagination Patterns.
6. Preserve raw observations before normalizing
The parser should produce a source observation before a canonical record.
An observation can include:
{
"source_id": "catalog_014",
"source_url": "https://example.com/item/123",
"collected_at": "2026-07-22T08:42:11Z",
"parser_version": "catalog_014@7",
"source_key": "123",
"title_raw": "18V Drill — Tool Only",
"price_raw": "€129,95",
"availability_raw": "Available in 2–4 days",
"raw_artifact_key": "responses/2026-07-22/abc123.json.gz"
}Normalization can then produce price_amount: 129.95, currency: EUR, and an agreed availability category without destroying the source text.
Raw HTML or JSON snapshots are useful when permitted and practical because they allow:
- Parser changes without another network request
- Investigation of a disputed field
- Comparison before and after a source change
- Evidence of what the source returned at collection time
Storage policies should account for source rules, sensitive data, size, retention, and security. Keeping everything forever is not automatically responsible architecture.
7. Make parsing deterministic and versioned
A parser should accept a known input and produce the same observation under the same version.
Keep source-specific extraction inside adapters:
fetch result
-> identify response type
-> source adapter
-> source observation
-> canonical mapping
-> validationAvoid spreading selectors across queue workers, delivery scripts, and normalization notebooks. When a source changes, the team should be able to identify which adapter and parser version created every affected record.
For heterogeneous sources, use a source-family template with explicit overrides. Do not force unrelated websites through one universal selector configuration if their semantics differ.
8. Validate the dataset at four levels
Validation is not a final dropna() call.
Structural validation
Check types, required fields, allowed values, formats, and nested shapes. JSON Schema (opens in a new tab) defines a vocabulary for asserting structural constraints on JSON instances.
Semantic validation
Check whether the value makes sense for this field and source: a price is non-negative, a latitude falls within range, a date parses under the documented format, or a product identifier follows the source's expected pattern.
Coverage validation
Compare discovered, attempted, fetched, parsed, valid, invalid, and excluded entities. A non-empty file can still omit an entire category, state, alphabet segment, or pagination branch.
Reconciliation validation
Ensure output records correspond to the requested input or frontier. This is particularly important for ordered input enrichment, where retries or asynchronous workers can otherwise shift results.
The Instagram bulk enrichment result describes preserving input order while processing more than 400,000 handles with batching and retries. That is a reconciliation requirement, not merely a formatting preference.
See How to Validate Scraped Data Before Delivery for a full QA model.
9. Monitor data signals as well as infrastructure
CPU, memory, and worker health are necessary, but they will not detect a parser that returns empty strings successfully.
Track metrics such as:
- URLs discovered, attempted, and completed
- Response count by status and content type
- Retry and terminal-failure rates
- Records parsed and validated
- Required-field coverage by source
- Duplicate keys
- Records per listing page or category
- Response-size and latency distributions
- Parser version in use
- New, changed, missing, and unchanged records on refresh
Scrapy's Stats Collector (opens in a new tab) provides counters and key/value statistics inside a crawl. Whatever framework is used, record both operational and dataset metrics.
Alerts should compare against a source baseline. Ten empty phone fields might be normal for one directory and a critical regression for another.
How to Detect Website Changes Before a Scraper Silently Loses Data covers sentinel pages, field-coverage drift, response signatures, and the recovery workflow.
Web Scraping Pipeline Monitoring expands this into run identity, crawl-frontier progress, retry backlogs, data-quality baselines, and delivery verification.
10. Design delivery for the consumer
CSV and Excel are appropriate when a person needs to review, filter, or import a bounded dataset. JSON or a database is better when relationships, nested values, history, or repeated updates matter. An API or webhook makes sense when another product needs programmatic access.
Every delivery should include a run summary:
- Scope and collection window
- Source list
- Output schema
- Record counts by source and status
- Required-field coverage
- Duplicate and exception counts
- Known limitations
- Refresh or change logic
- Source URL and timestamp fields
Do not hide invalid records to make the final file look clean. Deliver a valid dataset plus a separate exception report when exceptions matter to coverage.
A practical reference architecture
source register
-> discovery workers
-> durable frontier
-> source-aware fetch workers
-> raw artifact storage
-> versioned source parsers
-> canonical normalization
-> validation and reconciliation
-> current snapshot + history
-> file, database, API, or webhook delivery
-> run summary, metrics, and alertsAt smaller scale, SQLite and object files may implement several boxes. At larger scale, PostgreSQL or MongoDB, a queue, object storage, and independent workers may be appropriate. Choose infrastructure from the recovery, throughput, and delivery requirements—not from a desire to make the diagram look distributed.
Common architecture mistakes
Scaling concurrency before measuring coverage
More workers can make an incomplete crawl finish faster. Establish the expected URL universe and reconciliation first.
Treating every failure as a retry
A changed selector, denied source, invalid input, and temporary network error require different actions. Blind retries increase cost and hide the cause.
Using browser automation for every request
This can be correct for a small interactive source, but expensive for a large crawl. Inspect the network and page source before deciding.
Overwriting raw values
If normalized values replace source text, later audits and mapping changes become difficult. Keep both.
Delivering only the successful rows
Without the expected count and exception set, the client cannot distinguish complete collection from silent loss.
Building one parser for thousands of unrelated sites
Shared infrastructure is valuable; false uniformity is not. Use source families, adapters, and explicit overrides.
For the full source-portfolio design, read How to Scrape Thousands of Websites With Different Structures.
Frequently asked questions
Do I need Scrapy for a scalable web scraper?
No. Scrapy provides a mature crawling framework, asynchronous scheduling, statistics, throttling, and job persistence, but a scalable pipeline can also be built from HTTP clients, browser workers, queues, and databases. The architecture and operating controls matter more than the framework name.
Should scraped responses be stored?
Store raw responses when the source permits it and the audit, reprocessing, or change-detection value justifies the cost. Define retention and security rules instead of keeping all content indefinitely by default.
How many concurrent requests should a scraper send?
There is no universal safe number. Set concurrency per source based on its published limits, response behavior, terms, server feedback, collection window, and approved use. Track latency and error changes, and reduce pressure when the source signals rate limiting or instability.
How can a large crawl resume after a crash?
Persist the crawl frontier, visited set, attempt count, and output state outside the worker process. Restart workers from pending or retryable records rather than regenerating the entire crawl blindly.
What is the most important web-scraping metric?
For a dataset project, it is usually coverage against an explicit expected set—not requests per second. Throughput matters only after the team can explain what was discovered, attempted, delivered, excluded, and missed.
Next step
If your project involves a large source, several websites, recurring refreshes, or a dataset that must be reconciled, start with Orzaen's web scraping and data extraction service. Use the pipeline monitoring guide and crawl-resumption guide as the operational continuation. For concrete architecture examples, review the Kramp crawl and GraphQL result, VRBO property extraction result, and YouTube hybrid extraction result.
Sources
- Requests advanced usage (opens in a new tab)
- Selenium WebDriver documentation (opens in a new tab)
- Playwright network documentation (opens in a new tab)
- Scrapy documentation (opens in a new tab)
- Scrapy pausing and resuming crawls (opens in a new tab)
- Scrapy statistics collection (opens in a new tab)
- Chrome DevTools Network panel (opens in a new tab)
- JSON Schema validation specification (opens in a new tab)
- RFC 9309: Robots Exclusion Protocol (opens in a new tab)
- RFC 9110: HTTP Semantics (opens in a new tab)

