Healthcare Provider Data Collection — Custom physician and provider roster datasets for healthcare research, consulting, analytics, and health-tech projects.Learn more
OrzaenOrzaen

How to Detect Website Changes Before a Scraper Silently Loses Data

Engineering2026-07-1812 minHira Arif

Monitor response signatures, field coverage, fixtures, and sentinel pages so website changes trigger alerts before a scraper delivers incomplete data.

Short answer

Detect scraper-breaking website changes by monitoring both responses and extracted data. Track response status, content type, size, template signatures, required-field coverage, records per page, duplicate rates, and expected partitions. Run representative sentinel pages before or during each production cycle, version parsers, preserve raw fixtures, and alert on sustained deviations from each source's baseline. A successful request or completed process is not proof that the data is still correct.

The most dangerous scraper failure is not a visible crash. It is a clean run that quietly returns empty, shifted, duplicated, or incomplete fields.

This article focuses on source drift after deployment. The scalable web scraping pipeline guide explains the full architecture that produces the crawl state, raw observations, and validation metrics used here.

What can change?

Change layerExampleLikely symptom
Access and responseLogin introduced, redirect, 403, 429, or error templateStatus or response-signature change
NavigationCategory path, sitemap, or pagination changesDiscovery count drops
MarkupClass, label, table, or page template changesField coverage falls or wrong text is parsed
Structured responseJSON path, GraphQL operation, nesting, or data type changesParser errors or null fields
Browser workflowButton, frame, timing, consent, or interaction changesStep timeout or wrong page state
Business semanticsField label stays but meaning changesPlausible yet incorrect values
Source inventoryEntity added, moved, merged, or removedSnapshot differences

Different changes need different signals. A selector test will not detect an entire category disappearing from discovery.

The monitoring layers

Use five layers together.

  1. Transport: request status, latency, redirects, content type, and retries.
  2. Response: template, size, structured paths, and recognizable page state.
  3. Parsing: required fields, extraction exceptions, and parser version.
  4. Dataset: counts, field coverage, distributions, duplicates, and partitions.
  5. Business freshness: added, changed, missing, and unchanged entities across snapshots.

Infrastructure monitoring belongs underneath these layers, but healthy CPU and memory cannot prove healthy data.

1. Establish a baseline for every source

A baseline describes normal behavior, not a universal target.

Record:

  • Typical response statuses
  • Accepted content types
  • Response-size range
  • Redirect behavior
  • Listing pages or partitions
  • Approximate entity-count range
  • Required-field coverage
  • Optional-field coverage
  • Records per listing page
  • Duplicate-key rate
  • Typical run duration and latency
  • Last verified parser version

Use a time window or known-good runs, not one page collected once. Sources can vary by weekday, inventory, geography, or content type.

Thresholds should express both absolute and relative change. A drop from 20 records to zero is critical even though the absolute difference is small; a drop from 200,000 to 199,900 may be normal business change.

2. Create representative sentinel pages

A sentinel is a known source page or request selected to expose likely failures early.

For each source family, include examples such as:

  • Listing page with pagination
  • Detail page with every common field
  • Detail page with intentionally missing optional fields
  • Multi-location or multi-variant record
  • Later pagination page
  • Geographic or category edge case
  • Authenticated or interactive workflow, when approved

Before a large crawl, run sentinels and check:

  • Expected response state
  • Required fields
  • Stable identifiers
  • Record count range
  • Pagination signal
  • Structured-response paths

Sentinels reduce the chance of discovering a broken parser after millions of requests. They do not replace full-run coverage monitoring because a source can change only one template or partition.

3. Detect response changes before parsing

Status and redirect checks

Record the final URL, redirect chain, status, and headers. A source can start redirecting detail pages to a login, consent, regional, or generic error page while still returning 200 at the end.

Content-type checks

An expected JSON response may become HTML. A file download may become a small error document. Reject or quarantine unexpected types before the parser treats them as data.

Size and shape checks

Track response-size distributions per endpoint or page type. A sharp collapse can indicate an empty shell or error template. Size is an alert signal, not proof; legitimate records vary.

For JSON or GraphQL, monitor required paths and value types. For HTML, monitor stable structural features such as page title pattern, canonical URL, entity identifier, primary heading, or known container.

Template signatures

Create a normalized signature from stable structural features rather than hashing the entire response.

An exact HTML hash changes because of timestamps, recommendations, ads, or session values. A structural signature might include:

  • Ordered heading levels
  • Stable element roles or labels
  • Presence of required JSON keys
  • Selected tag paths
  • Known page-state markers

Use the signature to route an unfamiliar template to review; do not automatically declare every difference a break.

4. Monitor extracted-field coverage

Field coverage is one of the strongest silent-failure signals.

For each source and parser version, calculate:

text
coverage(field) = non-null valid values / valid source records

Track required and optional fields separately. Alert on:

  • A required field falling below its hard threshold
  • An optional field moving far outside its baseline
  • Several related fields dropping together
  • One worker batch producing a different pattern
  • Coverage changing immediately after a parser deployment

Examples:

  • product_name falls from nearly universal to zero: likely parser or template failure.
  • sale_price falls during a promotion change: possibly legitimate; review source behavior.
  • phone falls only for one provider-directory family: likely a family template change.

Monitor both nulls and invalid values. A selector shift can extract navigation text into a required field, leaving coverage at 100% while correctness collapses.

5. Monitor distributions and invariants

Useful invariants include:

  • Price is non-negative and currency is expected for the source.
  • Coordinates remain in the expected region.
  • Pagination cursors do not repeat indefinitely.
  • Entity IDs remain unique within the source snapshot.
  • Input enrichment returns one status per input row.
  • Delivered rows equal valid plus explicitly accepted exception categories.
  • Every discovered partition reaches a terminal status.

Useful distributions include:

  • Records by source, geography, category, or status
  • String length by field
  • Numeric ranges and quantiles
  • Categorical value frequencies
  • Entities per listing page
  • Response and parse time

Compare distributions with prior runs and source expectations. Do not alert on any change; alert on unexplained change that crosses a material threshold.

6. Reconcile the crawl frontier

A monitor should be able to answer:

text
discovered = validated + invalid + retryable + terminal + excluded + in_progress

When the run closes, in_progress should be zero and every other bucket should have explicit identities.

Track attempts separately from unique entities. Otherwise a retry surge can make request volume look healthy while unique coverage falls.

Scrapy provides a key/value statistics collection facility (opens in a new tab), and its job persistence can retain scheduled and visited request state for pausing and resuming crawls (opens in a new tab). Custom systems should expose equivalent frontier and item counters.

This reconciliation belongs inside the broader Web Scraping Pipeline Monitoring system, which also tracks worker health, retries, data quality, and delivery. Change detection remains responsible for identifying source drift and the affected parser scope.

7. Preserve fixtures and raw evidence

For each adapter, keep approved representative fixtures or response snapshots where permitted. Tie them to:

  • Source ID
  • Page or response type
  • Collection date
  • Parser version
  • Expected parsed fields
  • Redacted or storage-safe form when necessary

Run fixture tests before deploying a parser change. Then run live sentinels because stored fixtures cannot show current source behavior.

Raw response references also shorten incident diagnosis. When coverage fell at 03:10, the team can inspect what the source actually returned instead of trying to reproduce the same state hours later.

8. Version parsers and schemas

Every observation should identify the parser and schema version that produced it.

Versioning allows the team to:

  • Compare old and new parser output on the same fixture
  • Limit reprocessing to affected sources
  • Identify which records need correction
  • Roll back a bad parser release
  • Explain why a field changed between snapshots

Do not edit a parser in place and overwrite all current records without retaining the relationship to the prior source observations.

9. Use a change-detection pipeline for recurring data

Website-change detection and business-data change detection are related but different.

After validating the new snapshot, match records using stable source keys and classify:

StateMeaning
addedKey appears in the new snapshot only
changedKey exists in both; one or more monitored fields changed
unchangedKey exists in both; monitored fields match
missingKey existed previously but was not observed now
unresolvedCurrent collection was insufficient to classify the entity

Do not immediately convert missing into deleted. The page may have moved, the partition may have failed, or access may have changed. Require source-specific confirmation rules.

Compare normalized fields for business changes, but retain raw values to investigate formatting-only differences.

10. Set actionable alerts

An alert should include:

  • Source and adapter family
  • Metric and current value
  • Baseline or threshold
  • First affected run and time
  • Example URLs or record keys
  • Parser version
  • Response status, type, or signature where relevant
  • Estimated affected record count
  • Link to raw artifacts or logs
  • Suggested owner and severity

Alert levels can be:

SeverityExample
CriticalDiscovery returns zero; required identifier coverage collapses; access becomes prohibited
HighOne large partition is missing; response template changes across most pages
MediumOptional field coverage drops materially; duplicate rate rises
LowNew categorical value or isolated unknown template appears

Avoid one notification per failed URL. Aggregate related failures into one incident while retaining the underlying exception list.

11. Respond to a detected change safely

Use this incident sequence:

  1. Pause or limit the affected source if continuing could create bad data or inappropriate traffic.
  2. Confirm whether the signal is a real source change, temporary incident, or bad deployment.
  3. Inspect representative raw responses and current live behavior.
  4. Determine whether discovery, fetching, parsing, normalization, or validation changed.
  5. Update the adapter and fixtures.
  6. Test old fixtures, new fixtures, and live sentinels.
  7. Reprocess stored responses when possible.
  8. Recrawl only the affected frontier when necessary and appropriate.
  9. Reconcile the corrected output.
  10. Record the incident, impact, and prevention rule.

Do not patch the selector and immediately rerun the entire source without calculating the affected range.

An example monitoring record

json
{
  "run_id": "source_014_2026-07-22",
  "source_id": "source_014",
  "parser_version": "source_014@7",
  "urls_discovered": 12840,
  "urls_fetched": 12802,
  "records_valid": 12761,
  "records_exception": 41,
  "required_field_coverage": {
    "source_key": 1.0,
    "name": 0.998,
    "category": 0.941
  },
  "response_statuses": {
    "200": 12770,
    "404": 12,
    "429": 20
  }
}

These values are illustrative. Production thresholds should come from the actual source contract and baseline.

What Orzaen's recurring work implies

Orzaen's public results span large crawls, bulk enrichment, property collection, social-platform workflows, and recurring provider directories. The operating lesson across them is that the deliverable is not just a parser.

The Kramp crawl result describes deduplication and URL validation across HTML and GraphQL artifacts. The Instagram bulk enrichment result describes batching, retries, and input-order preservation. The recurring hospital directory result makes freshness and repeatability part of the work.

Those controls should become measurable monitors rather than remain implementation details.

Common monitoring mistakes

Alerting only on crashes

The process may finish while the parser returns blanks. Track field coverage and reconciliation.

Hashing the entire page

Dynamic content creates constant noise. Use stable structural features and dataset signals.

Using one threshold for every source

Normal volume and field coverage differ. Baselines belong to the source and page type.

Ignoring discovery

Perfect detail-page parsing cannot recover entity URLs that were never discovered.

Calling every missing key deleted

Absence in one crawl is not proof of a real-world deletion. Distinguish missing from confirmed_removed.

Monitoring requests instead of unique work

Retries can increase request counts while successful unique entities decline.

Keeping no parser version

Without version lineage, the team cannot identify which records were produced before or after a change.

Frequently asked questions

How often should sentinel pages run?

Run them before each material production cycle and after parser, browser, or source-policy changes. High-frequency or high-impact pipelines may run sentinels continuously or before each partition batch.

Can visual screenshots detect scraper changes?

Screenshots can help diagnose browser workflows and major template changes, but they are difficult to validate at scale. Combine them with response signatures, structured-path checks, field coverage, and crawl reconciliation.

Should a scraper stop when one field is missing?

It depends on the field and source contract. A missing required identifier may justify quarantining the record or pausing the source. A normally optional field should be recorded and monitored without stopping all work.

How do I distinguish a website change from normal data change?

Website changes affect response structure, paths, templates, or field distributions across many records. Business changes affect entity values within a stable structure. Raw observations, parser versions, sentinels, and cross-record patterns help distinguish them.

What should happen to records collected by a broken parser?

Identify the affected source, parser version, run, and field set. Reparse stored raw responses when possible; otherwise recrawl only the affected scope when appropriate. Do not silently merge corrected and uncorrected records.

Next step

For a recurring scraper that needs run summaries, source baselines, exceptions, and refresh logic, review Orzaen's web scraping and data extraction service. Implement the complete pipeline monitoring model and preserve restart state with the crawl-resumption guide. Related proof includes the Kramp crawl result, Instagram bulk enrichment result, and recurring hospital-directory result.

Sources

Tags

Web ScrapingMonitoringChange DetectionData Quality

Need this fixed?

Have the same system problem?

Share what is manual, messy, broken, or disconnected. We’ll review the cleanest next step.

Get System Review