Healthcare Provider Data Collection — Custom physician and provider roster datasets for healthcare research, consulting, analytics, and health-tech projects.Learn more
OrzaenOrzaen

How to Scrape Thousands of Websites With Different Structures

Engineering2026-07-1812 minHira Arif

A maintainable architecture for collecting one canonical dataset from thousands of websites with different layouts, fields, pagination, and failure modes.

Short answer

To scrape thousands of differently structured websites, use shared crawl infrastructure with source-specific adapters. Define one canonical schema, inventory and classify the sources, group similar sites into families, preserve raw source observations, and keep discovery, parsing, validation, and scheduling separate. Do not build one universal selector and do not create thousands of unrelated scripts with no shared operating layer.

The system should make adding source 1,001 routine: register the source, assign or create an adapter, map its fields, run fixtures, establish a baseline, and schedule it.

This article goes deep on heterogeneous sources. The scalable web scraping pipeline guide covers the surrounding discovery, queueing, storage, delivery, and monitoring architecture.

Why multi-website scraping is a different problem

One website can often be modeled as a page hierarchy. Thousands of websites are a changing source portfolio.

They differ in:

  • Site platform and rendering
  • Listing and detail-page structure
  • Labels and field meanings
  • Pagination
  • Geographic or category navigation
  • Identifier quality
  • Update frequency
  • Rate expectations
  • Missing-value conventions
  • Access and permission constraints

The challenge is not making every page look identical. It is producing comparable records without erasing what each source actually said.

Current community questions regularly ask how a small team could refresh data from tens or hundreds of thousands of independent websites. The answer is not a larger CSS-selector file. It is source inventory, adapters, durable state, and exception-driven operations.

The core architecture

Shared componentSource-specific component
Canonical schemaDiscovery rules
Queue and crawl stateRequest and session policy
Raw artifact storageParser and field mapping
Retry frameworkPagination logic
Normalization librarySource-specific validation expectations
Metrics and alertsBaseline and change thresholds
Delivery pipelineExceptions and known limitations

Shared infrastructure provides consistency. Adapters preserve source truth.

1. Build a source inventory

Create a row for every website before writing hundreds of spiders.

Recommended fields include:

FieldPurpose
source_idStable internal identifier
domainSource authority
source_typeHospital, retailer, practice, supplier, directory, property site, or other class
platform_familyWordPress, Shopify, custom React, shared vendor template, unknown
entity_typeProvider, product, location, property, company, creator, or another entity
discovery_methodSitemap, category traversal, search, supplied URLs, API
extraction_surfaceHTML, embedded JSON, XHR, GraphQL, browser
estimated_volumePlanning estimate, not a promised final count
refresh_cadenceDaily, weekly, monthly, one-time
access_reviewTerms, robots, permission, and source notes
adapterParser family or custom adapter
statusProposed, approved, active, blocked, paused, retired
last_verified_atLast manual verification

Without this register, teams lose track of sources that never ran, were silently skipped, or require special handling.

2. Define one canonical schema—and keep raw fields

The canonical schema expresses the buyer's data model. Each adapter maps source fields into it.

For a business-directory project, the canonical schema might contain:

text
entity_id
name_raw
name_normalized
category_raw
category_normalized
address_raw
street
city
region
postal_code
country
phone_raw
phone_normalized
source_id
source_url
collected_at
parser_version

Raw and normalized fields should coexist. If one source says Cardiology and another says Cardiovascular Medicine, a reporting taxonomy may map both to one category while preserving the original labels.

Do not force every source into a flat row when the entity has repeated relationships. Products can have variants; providers can have several locations; properties can have many amenities; companies can expose several contacts. Normalize relationships in a database or preserve nested structures in JSON, then create a flat export only for a defined consumer.

3. Classify sources into families

Thousands of websites are rarely 100% unique. Common families include:

  • The same vendor platform used by several organizations
  • Shared CMS themes with local customization
  • Common e-commerce platforms
  • Static directory tables
  • Search-driven React applications
  • Sitemap-backed profile sites
  • Organization sites with JSON-LD plus custom biography HTML

Use three implementation levels:

  1. Configuration only: the family parser works and only base URLs, labels, or paths differ.
  2. Family override: most behavior is shared, but one pagination or field rule differs.
  3. Custom adapter: the source has genuinely different discovery, state, or semantics.

This avoids two extremes: thousands of duplicated scripts and one brittle parser overloaded with conditionals.

4. Give every adapter a clear contract

An adapter should answer:

  • How are entity URLs or identifiers discovered?
  • What source surface is fetched?
  • How is pagination completed?
  • Which fields are parsed and where?
  • What constitutes a valid record?
  • Which missing fields are normal?
  • Which response states mean removed, blocked, empty, or retryable?
  • What fixtures prove the adapter still works?

A simplified interface might look like:

python
class SourceAdapter:
    def discover(self, seed): ...
    def build_requests(self, item): ...
    def parse(self, response): ...
    def normalize(self, observation): ...
    def validate(self, record): ...

The exact class design is less important than keeping source rules outside shared queue and delivery code.

5. Use a durable crawl frontier

For a multi-source crawl, every work item needs a durable identity.

An item can be keyed by:

text
(source_id, entity_type, source_key, collection_cycle)

Store its discovery path, priority, attempt count, status, timestamps, and failure class. Workers should claim items atomically so two workers do not process the same entity unless the design intentionally allows it.

This makes it possible to report:

  • 3,000 sources approved
  • 2,950 sources started
  • 2,870 sources completed
  • 45 paused by source policy or access review
  • 35 failed and require engineering review

The numbers are illustrative. The point is that source coverage and record coverage are both first-class outputs.

For framework support, Scrapy documents persisted scheduling and duplicate state through pausing and resuming crawls (opens in a new tab). A database-backed frontier can provide similar control across custom workers.

For stable task keys, leases, retry state, and idempotent output writes, see How to Resume a Large Web Crawl Without Duplicating Work.

6. Separate discovery from detail extraction

Discovery finds the expected entity universe. Detail extraction fills the record.

Discovery inputs may include:

  • XML sitemaps
  • Category and subcategory links
  • Alphabetical indexes
  • Geographic partitions
  • Search results
  • Supplied identifiers or URLs
  • Public APIs or downloadable indexes

The Sitemaps protocol (opens in a new tab) supports URL lists and sitemap indexes. Sitemaps are useful seeds but should be reconciled with visible navigation or another expected set because source implementations vary.

In Orzaen's VRBO USA property extraction result, state, city, and subarea traversal created the discovery coverage needed to handle a search-result cap. Discovered URLs were queued before detail extraction. The final published result contained 201,717 US property records.

When discovery and extraction are fused, a failure halfway through a category makes it difficult to know which entity URLs were never seen.

7. Apply policy and concurrency per source

A large source portfolio needs fairness. One slow or rate-limited domain should not block all workers, and one large domain should not consume every available connection.

Configure by source or domain:

  • Concurrent requests
  • Delay or adaptive throttle
  • Retry policy
  • Session handling
  • Browser allocation
  • Collection window
  • Maximum response size
  • Source-specific exclusions

Scrapy's AutoThrottle (opens in a new tab) adjusts delay using observed latency while respecting configured concurrency and delay limits. Even with adaptive tooling, source instructions and explicit limits remain the governing constraints.

Use a domain-aware scheduler or separate queues so work is interleaved appropriately.

Apply the 403, 429, and temporary-failure model per source. If a permitted source requires multiple routes or regions, the proxy rotation guide explains session affinity and health scoring.

8. Do not render every site in a browser

For each source family, inspect:

  1. Server HTML
  2. Embedded structured data
  3. XHR, fetch, GraphQL, or file responses
  4. Browser-only state

Some sources will use Requests; others will use Playwright or Selenium; some will use a hybrid workflow. The method belongs in source configuration.

Requests vs Playwright vs Selenium compares the execution choices, while API vs HTML vs Browser Automation explains how to find the underlying data surface.

For JavaScript-heavy source families, use browser discovery with controlled HTTP extraction where the source and session allow it.

9. Preserve source observations and provenance

For each observation, keep:

  • source_id
  • Source URL or request identity
  • Collection timestamp
  • Parser version
  • Raw source fields
  • Normalized fields
  • Raw artifact reference when retained
  • Validation status
  • Exception codes

Provenance becomes essential when two websites disagree, a normalized value is questioned, or a parser change affects only one source family.

It also prevents a dangerous pattern: filling one source's missing value from another source without disclosing that enrichment.

10. Validate by source and by portfolio

One global missingness threshold is not useful for heterogeneous sources.

Use source-level baselines:

  • Expected listing-page count
  • Typical entities per page
  • Required-field coverage
  • Optional-field coverage
  • Response-type distribution
  • Duplicate-key rate
  • Detail-page success rate

Then validate the combined dataset:

  • Canonical key uniqueness
  • Cross-source duplicate candidates
  • Normalized category distribution
  • Geographic coverage
  • Source representation
  • Field lineage

For the full QA model, read How to Validate Scraped Data Before Delivery.

11. Build an onboarding process for new sources

A repeatable source-onboarding ticket should include:

  1. Confirm the source and intended use are approved.
  2. Sample listing, detail, missing-field, and edge-case pages.
  3. Identify discovery and extraction surfaces.
  4. Assign a source family or create a custom adapter.
  5. Map raw fields to the canonical schema.
  6. Add recorded fixtures where appropriate.
  7. Define expected-field and coverage baselines.
  8. Run a limited crawl.
  9. Review exceptions manually.
  10. Activate the production cadence.

Do not move a source from "we parsed one page" directly to "active weekly collection."

12. Operate by exceptions

At thousands of sources, nobody should manually inspect every successful page. The system should surface the smallest useful exception set.

Exception classes can include:

  • Discovery returned zero entity URLs
  • Response template changed
  • Required field coverage dropped
  • Pagination repeated a cursor
  • Login or session expired
  • Unexpected content type
  • Duplicate source keys increased
  • Record count moved outside the source baseline
  • Source rules or availability changed

Each exception should have an owner, severity, first-seen time, last-seen time, affected source family, and reprocessing status.

How AI-assisted extraction fits

Language models or visual extraction can help classify unknown pages, propose mappings, or process long-tail layouts. They do not remove the need for:

  • Approved sources
  • A canonical schema
  • Expected coverage
  • Confidence and exception policies
  • Raw evidence
  • Deterministic reprocessing where required
  • Cost measurement
  • Human review for ambiguous fields

Use AI as one adapter strategy, not as a substitute for source control and dataset accountability.

Common design mistakes

One universal scraper

A universal extractor may be useful for discovery or low-stakes text capture. It is risky as the only layer when fields have source-specific meanings and completeness matters.

One script per website

This duplicates retry, storage, logging, normalization, and delivery logic. Keep source differences in adapters on top of shared infrastructure.

One global schedule

Sources change at different rates and impose different costs. Refresh cadence should follow business need and source volatility.

Treating zero records as a valid success

A job that exits without an exception can still have failed. Compare results with the expected source set and baseline.

Normalizing during parsing

This destroys source values and makes later mapping changes difficult. Parse first; normalize in a versioned downstream step.

Fixing only the first broken source

If several sources share a platform family, diagnose whether one family adapter changed before creating local patches.

Frequently asked questions

Do I need one scraper for every website?

You need source-specific behavior, but not necessarily a separate application. Group similar sites into adapter families, use configuration where possible, and keep custom code for genuinely different sources.

Can CSS selectors be generated automatically?

They can be proposed or learned, but production use still needs field definitions, confidence rules, fixtures, coverage checks, and exceptions. A selector that returns text is not proof that it returned the correct field.

How do I handle websites that have no sitemap?

Use another approved discovery method such as navigation traversal, search partitions, alphabetical indexes, supplied URLs, public APIs, or known identifiers. Record how completeness will be measured for that source.

Should all sources use the same refresh schedule?

No. Base cadence on business freshness requirements, source volatility, cost, and allowed access. Store the collection time for every observation.

What database works best for multi-source scraping?

Choose from the data and operating requirements. PostgreSQL fits relational entities, constraints, and reporting; MongoDB can fit varied source observations and queues; object storage fits raw artifacts. Many systems combine them. A spreadsheet is a delivery format, not a durable crawl-state system at this scale.

Next step

If your dataset spans many independent websites, Orzaen's web scraping and data extraction service can begin with source inventory, schema, coverage, and refresh design. Use the pipeline monitoring guide to operate the source portfolio by exceptions. For large discovery and extraction examples, see the Kramp crawl result, VRBO USA result, and recurring hospital directory result.

Sources

Tags

Web ScrapingMulti-Source DataData ArchitecturePython

Need this fixed?

Have the same system problem?

Share what is manual, messy, broken, or disconnected. We’ll review the cleanest next step.

Get System Review