Short answer
To scrape thousands of differently structured websites, use shared crawl infrastructure with source-specific adapters. Define one canonical schema, inventory and classify the sources, group similar sites into families, preserve raw source observations, and keep discovery, parsing, validation, and scheduling separate. Do not build one universal selector and do not create thousands of unrelated scripts with no shared operating layer.
The system should make adding source 1,001 routine: register the source, assign or create an adapter, map its fields, run fixtures, establish a baseline, and schedule it.
This article goes deep on heterogeneous sources. The scalable web scraping pipeline guide covers the surrounding discovery, queueing, storage, delivery, and monitoring architecture.
Why multi-website scraping is a different problem
One website can often be modeled as a page hierarchy. Thousands of websites are a changing source portfolio.
They differ in:
- Site platform and rendering
- Listing and detail-page structure
- Labels and field meanings
- Pagination
- Geographic or category navigation
- Identifier quality
- Update frequency
- Rate expectations
- Missing-value conventions
- Access and permission constraints
The challenge is not making every page look identical. It is producing comparable records without erasing what each source actually said.
Current community questions regularly ask how a small team could refresh data from tens or hundreds of thousands of independent websites. The answer is not a larger CSS-selector file. It is source inventory, adapters, durable state, and exception-driven operations.
The core architecture
| Shared component | Source-specific component |
|---|---|
| Canonical schema | Discovery rules |
| Queue and crawl state | Request and session policy |
| Raw artifact storage | Parser and field mapping |
| Retry framework | Pagination logic |
| Normalization library | Source-specific validation expectations |
| Metrics and alerts | Baseline and change thresholds |
| Delivery pipeline | Exceptions and known limitations |
Shared infrastructure provides consistency. Adapters preserve source truth.
1. Build a source inventory
Create a row for every website before writing hundreds of spiders.
Recommended fields include:
| Field | Purpose |
|---|---|
source_id | Stable internal identifier |
domain | Source authority |
source_type | Hospital, retailer, practice, supplier, directory, property site, or other class |
platform_family | WordPress, Shopify, custom React, shared vendor template, unknown |
entity_type | Provider, product, location, property, company, creator, or another entity |
discovery_method | Sitemap, category traversal, search, supplied URLs, API |
extraction_surface | HTML, embedded JSON, XHR, GraphQL, browser |
estimated_volume | Planning estimate, not a promised final count |
refresh_cadence | Daily, weekly, monthly, one-time |
access_review | Terms, robots, permission, and source notes |
adapter | Parser family or custom adapter |
status | Proposed, approved, active, blocked, paused, retired |
last_verified_at | Last manual verification |
Without this register, teams lose track of sources that never ran, were silently skipped, or require special handling.
2. Define one canonical schema—and keep raw fields
The canonical schema expresses the buyer's data model. Each adapter maps source fields into it.
For a business-directory project, the canonical schema might contain:
entity_id
name_raw
name_normalized
category_raw
category_normalized
address_raw
street
city
region
postal_code
country
phone_raw
phone_normalized
source_id
source_url
collected_at
parser_versionRaw and normalized fields should coexist. If one source says Cardiology and another says Cardiovascular Medicine, a reporting taxonomy may map both to one category while preserving the original labels.
Do not force every source into a flat row when the entity has repeated relationships. Products can have variants; providers can have several locations; properties can have many amenities; companies can expose several contacts. Normalize relationships in a database or preserve nested structures in JSON, then create a flat export only for a defined consumer.
3. Classify sources into families
Thousands of websites are rarely 100% unique. Common families include:
- The same vendor platform used by several organizations
- Shared CMS themes with local customization
- Common e-commerce platforms
- Static directory tables
- Search-driven React applications
- Sitemap-backed profile sites
- Organization sites with JSON-LD plus custom biography HTML
Use three implementation levels:
- Configuration only: the family parser works and only base URLs, labels, or paths differ.
- Family override: most behavior is shared, but one pagination or field rule differs.
- Custom adapter: the source has genuinely different discovery, state, or semantics.
This avoids two extremes: thousands of duplicated scripts and one brittle parser overloaded with conditionals.
4. Give every adapter a clear contract
An adapter should answer:
- How are entity URLs or identifiers discovered?
- What source surface is fetched?
- How is pagination completed?
- Which fields are parsed and where?
- What constitutes a valid record?
- Which missing fields are normal?
- Which response states mean removed, blocked, empty, or retryable?
- What fixtures prove the adapter still works?
A simplified interface might look like:
class SourceAdapter:
def discover(self, seed): ...
def build_requests(self, item): ...
def parse(self, response): ...
def normalize(self, observation): ...
def validate(self, record): ...The exact class design is less important than keeping source rules outside shared queue and delivery code.
5. Use a durable crawl frontier
For a multi-source crawl, every work item needs a durable identity.
An item can be keyed by:
(source_id, entity_type, source_key, collection_cycle)Store its discovery path, priority, attempt count, status, timestamps, and failure class. Workers should claim items atomically so two workers do not process the same entity unless the design intentionally allows it.
This makes it possible to report:
- 3,000 sources approved
- 2,950 sources started
- 2,870 sources completed
- 45 paused by source policy or access review
- 35 failed and require engineering review
The numbers are illustrative. The point is that source coverage and record coverage are both first-class outputs.
For framework support, Scrapy documents persisted scheduling and duplicate state through pausing and resuming crawls (opens in a new tab). A database-backed frontier can provide similar control across custom workers.
For stable task keys, leases, retry state, and idempotent output writes, see How to Resume a Large Web Crawl Without Duplicating Work.
6. Separate discovery from detail extraction
Discovery finds the expected entity universe. Detail extraction fills the record.
Discovery inputs may include:
- XML sitemaps
- Category and subcategory links
- Alphabetical indexes
- Geographic partitions
- Search results
- Supplied identifiers or URLs
- Public APIs or downloadable indexes
The Sitemaps protocol (opens in a new tab) supports URL lists and sitemap indexes. Sitemaps are useful seeds but should be reconciled with visible navigation or another expected set because source implementations vary.
In Orzaen's VRBO USA property extraction result, state, city, and subarea traversal created the discovery coverage needed to handle a search-result cap. Discovered URLs were queued before detail extraction. The final published result contained 201,717 US property records.
When discovery and extraction are fused, a failure halfway through a category makes it difficult to know which entity URLs were never seen.
7. Apply policy and concurrency per source
A large source portfolio needs fairness. One slow or rate-limited domain should not block all workers, and one large domain should not consume every available connection.
Configure by source or domain:
- Concurrent requests
- Delay or adaptive throttle
- Retry policy
- Session handling
- Browser allocation
- Collection window
- Maximum response size
- Source-specific exclusions
Scrapy's AutoThrottle (opens in a new tab) adjusts delay using observed latency while respecting configured concurrency and delay limits. Even with adaptive tooling, source instructions and explicit limits remain the governing constraints.
Use a domain-aware scheduler or separate queues so work is interleaved appropriately.
Apply the 403, 429, and temporary-failure model per source. If a permitted source requires multiple routes or regions, the proxy rotation guide explains session affinity and health scoring.
8. Do not render every site in a browser
For each source family, inspect:
- Server HTML
- Embedded structured data
- XHR, fetch, GraphQL, or file responses
- Browser-only state
Some sources will use Requests; others will use Playwright or Selenium; some will use a hybrid workflow. The method belongs in source configuration.
Requests vs Playwright vs Selenium compares the execution choices, while API vs HTML vs Browser Automation explains how to find the underlying data surface.
For JavaScript-heavy source families, use browser discovery with controlled HTTP extraction where the source and session allow it.
9. Preserve source observations and provenance
For each observation, keep:
source_id- Source URL or request identity
- Collection timestamp
- Parser version
- Raw source fields
- Normalized fields
- Raw artifact reference when retained
- Validation status
- Exception codes
Provenance becomes essential when two websites disagree, a normalized value is questioned, or a parser change affects only one source family.
It also prevents a dangerous pattern: filling one source's missing value from another source without disclosing that enrichment.
10. Validate by source and by portfolio
One global missingness threshold is not useful for heterogeneous sources.
Use source-level baselines:
- Expected listing-page count
- Typical entities per page
- Required-field coverage
- Optional-field coverage
- Response-type distribution
- Duplicate-key rate
- Detail-page success rate
Then validate the combined dataset:
- Canonical key uniqueness
- Cross-source duplicate candidates
- Normalized category distribution
- Geographic coverage
- Source representation
- Field lineage
For the full QA model, read How to Validate Scraped Data Before Delivery.
11. Build an onboarding process for new sources
A repeatable source-onboarding ticket should include:
- Confirm the source and intended use are approved.
- Sample listing, detail, missing-field, and edge-case pages.
- Identify discovery and extraction surfaces.
- Assign a source family or create a custom adapter.
- Map raw fields to the canonical schema.
- Add recorded fixtures where appropriate.
- Define expected-field and coverage baselines.
- Run a limited crawl.
- Review exceptions manually.
- Activate the production cadence.
Do not move a source from "we parsed one page" directly to "active weekly collection."
12. Operate by exceptions
At thousands of sources, nobody should manually inspect every successful page. The system should surface the smallest useful exception set.
Exception classes can include:
- Discovery returned zero entity URLs
- Response template changed
- Required field coverage dropped
- Pagination repeated a cursor
- Login or session expired
- Unexpected content type
- Duplicate source keys increased
- Record count moved outside the source baseline
- Source rules or availability changed
Each exception should have an owner, severity, first-seen time, last-seen time, affected source family, and reprocessing status.
How AI-assisted extraction fits
Language models or visual extraction can help classify unknown pages, propose mappings, or process long-tail layouts. They do not remove the need for:
- Approved sources
- A canonical schema
- Expected coverage
- Confidence and exception policies
- Raw evidence
- Deterministic reprocessing where required
- Cost measurement
- Human review for ambiguous fields
Use AI as one adapter strategy, not as a substitute for source control and dataset accountability.
Common design mistakes
One universal scraper
A universal extractor may be useful for discovery or low-stakes text capture. It is risky as the only layer when fields have source-specific meanings and completeness matters.
One script per website
This duplicates retry, storage, logging, normalization, and delivery logic. Keep source differences in adapters on top of shared infrastructure.
One global schedule
Sources change at different rates and impose different costs. Refresh cadence should follow business need and source volatility.
Treating zero records as a valid success
A job that exits without an exception can still have failed. Compare results with the expected source set and baseline.
Normalizing during parsing
This destroys source values and makes later mapping changes difficult. Parse first; normalize in a versioned downstream step.
Fixing only the first broken source
If several sources share a platform family, diagnose whether one family adapter changed before creating local patches.
Frequently asked questions
Do I need one scraper for every website?
You need source-specific behavior, but not necessarily a separate application. Group similar sites into adapter families, use configuration where possible, and keep custom code for genuinely different sources.
Can CSS selectors be generated automatically?
They can be proposed or learned, but production use still needs field definitions, confidence rules, fixtures, coverage checks, and exceptions. A selector that returns text is not proof that it returned the correct field.
How do I handle websites that have no sitemap?
Use another approved discovery method such as navigation traversal, search partitions, alphabetical indexes, supplied URLs, public APIs, or known identifiers. Record how completeness will be measured for that source.
Should all sources use the same refresh schedule?
No. Base cadence on business freshness requirements, source volatility, cost, and allowed access. Store the collection time for every observation.
What database works best for multi-source scraping?
Choose from the data and operating requirements. PostgreSQL fits relational entities, constraints, and reporting; MongoDB can fit varied source observations and queues; object storage fits raw artifacts. Many systems combine them. A spreadsheet is a delivery format, not a durable crawl-state system at this scale.
Next step
If your dataset spans many independent websites, Orzaen's web scraping and data extraction service can begin with source inventory, schema, coverage, and refresh design. Use the pipeline monitoring guide to operate the source portfolio by exceptions. For large discovery and extraction examples, see the Kramp crawl result, VRBO USA result, and recurring hospital directory result.
Sources
- Scrapy broad crawls (opens in a new tab)
- Scrapy pausing and resuming crawls (opens in a new tab)
- Scrapy AutoThrottle (opens in a new tab)
- Scrapy statistics collection (opens in a new tab)
- Sitemaps protocol (opens in a new tab)
- RFC 9309: Robots Exclusion Protocol (opens in a new tab)
- Chrome DevTools Network panel (opens in a new tab)

