Healthcare Provider Data Collection — Custom physician and provider roster datasets for healthcare research, consulting, analytics, and health-tech projects.Learn more
OrzaenOrzaen

Provider Directory Refresh: Full Recrawls, Changes, and QA

Engineering2026-07-2111 minHira Arif

Learn how to refresh hospital and provider directories using snapshots, stable hashes, incremental inputs, missing-record rules, provenance, and QA gates.

Short answer

Refresh provider-directory data by creating a new source snapshot, not by editing the previous export in place. Compare stable record keys and canonical field hashes to classify added, changed, unchanged, and missing records. Use source-specific QA gates before publishing the new snapshot. Full recollection is often safer for hospital directories; trustworthy incremental files can be used when an authoritative source, such as NPPES, explicitly provides them.

Why refreshes fail

A refresh is not simply "run the scraper again."

Common failures include:

  • The source changed its page structure.
  • Pagination stopped early.
  • A search endpoint returned only a default subset.
  • The new parser produced empty fields.
  • A transient error made profiles appear deleted.
  • Normalization changed and generated false differences.
  • Records were updated in place, erasing the prior source state.
  • Two different providers were assigned the same match key.

A safe refresh treats collection, comparison, and publication as separate decisions.

Full snapshot versus incremental update

MethodHow it worksStrengthRisk
Full snapshotRecollect every approved record from the sourceDetects broad structural and roster changesMore requests and processing
Source-provided incrementalIngest an authoritative change file or feedEfficient and explicitDepends on correct baseline and update application
Sitemap-guided refreshFetch URLs whose reliable modification signal changedCan reduce worklastmod may be absent or unreliable
Profile-level monitoringRevisit a known URL setGood for defined rostersMisses newly added profiles without discovery
HybridRefresh discovery fully and fetch changed or selected profilesBalances coverage and costMore operational complexity

For many hospital and health-system directories, a full snapshot is the most defensible method because the source does not expose a complete public change log.

CMS is different. The NPPES downloadable-file page (opens in a new tab) currently supplies a monthly full Version 2 file, weekly incremental files, and a monthly deactivation update. A pipeline can use the full file as a baseline and apply the documented updates while preserving release metadata.

Choosing a refresh cadence

There is no universal "provider data must be scraped every X days" rule for this type of project.

Choose cadence based on:

  • Business decision frequency
  • Expected source volatility
  • Cost of stale data
  • Source size and access constraints
  • Whether the dataset supports a one-time study or a live product
  • Whether a reliable update feed exists
  • Client tolerance for missing and unconfirmed changes

Examples:

Use casePossible cadenceReasoning
One-time market studyOne fresh collectionHistorical monitoring may add no value
Recurring consulting marketFresh collection per engagement, sometimes after 6–12 monthsThe decision is tied to a new client snapshot
Provider-search productScheduled source-specific cadenceProduct freshness is an ongoing requirement
High-value monitored rosterWeekly or monthly, if permitted and justifiedChanges affect an active workflow
NPPES warehouseWeekly increments plus periodic full reconciliationCMS supplies both update modes

State the collection date prominently. "Current" without a date is not a useful freshness claim.

Snapshot architecture

Each run should have an immutable identity.

text
collection_run
├── source_id
├── started_at
├── completed_at
├── adapter_version
├── discovery_count
├── fetch_metrics
├── parse_metrics
├── qa_status
└── publication_status

Source records should be unique within a run:

text
(run_id, canonical_source_key)

The current published dataset is then a pointer to an approved run or derived release. If QA fails, the earlier release remains available.

Stable source keys

Choose the strongest key the source provides:

  1. Stable source-native profile ID
  2. Canonical profile URL
  3. Published NPI plus source-specific relationship key
  4. Conservative composite key

Do not use normalized provider name alone. Two people can share a name, and one person's displayed name can change.

The source key identifies the source record. It does not necessarily identify the canonical provider across all sources.

Canonical comparison payload

Exclude fields that change on every run, such as collected_at, run ID, retry count, and raw response storage path.

python
from __future__ import annotations

import hashlib
import json
from collections.abc import Mapping, Sequence
from typing import Any


def sorted_unique(values: Sequence[str]) -> list[str]:
    return sorted({value.strip() for value in values if value.strip()})


def comparison_payload(record: Mapping[str, Any]) -> dict[str, Any]:
    return {
        "provider_name_raw": record.get("provider_name_raw"),
        "credentials_raw": sorted_unique(record.get("credentials_raw", [])),
        "specialties_raw": sorted_unique(record.get("specialties_raw", [])),
        "organization_name_raw": record.get("organization_name_raw"),
        "locations": record.get("locations", []),
        "phone_raw": record.get("phone_raw"),
        "fax_raw": record.get("fax_raw"),
        "npi_raw": record.get("npi_raw"),
    }


def payload_hash(payload: Mapping[str, Any]) -> str:
    serialized = json.dumps(
        payload,
        sort_keys=True,
        separators=(",", ":"),
        ensure_ascii=False,
    ).encode("utf-8")
    return hashlib.sha256(serialized).hexdigest()

This code is illustrative, not a complete production comparator. Nested locations must also be canonicalized, and some arrays are order-sensitive while others are not.

Version the comparison logic. A change in normalization should not masquerade as a real source change.

Change classification

For every stable source key:

Previous snapshotCurrent snapshotClassification
AbsentPresentAdded
Present, same hashPresent, same hashUnchanged
Present, different hashPresent, different hashChanged
PresentAbsentMissing/unconfirmed
Explicit inactive stateExplicit inactive stateSource-confirmed inactive

Then create field-level events for changed records.

json
{
  "source_key": "provider-123",
  "change_type": "field_changed",
  "field": "phone_raw",
  "previous_value": "212-555-0100",
  "current_value": "212-555-0199",
  "previous_run_id": "run-a",
  "current_run_id": "run-b"
}

Example values are fictional.

Missing is not deleted

A source record can disappear because:

  • The provider profile was removed.
  • The URL changed.
  • The source search omitted it temporarily.
  • Pagination failed.
  • The directory added a filter.
  • The request was blocked or returned an error page.
  • The parser rejected the page.

Use confirmation rules such as:

  • Missing in two successful full snapshots
  • Old URL returns an explicit permanent not-found response
  • Source publishes an inactive or removed status
  • Redirect resolves to a successor profile
  • Manual review confirms removal for high-impact records

Until then, classify the record as missing_unconfirmed rather than deleted.

Discovery QA before record comparison

Compare the run with the last known-good source baseline.

Record-count checks

  • Total discovered profiles
  • Profiles per specialty, state, or organization partition
  • Percentage change from previous successful run
  • New and missing URL ratio

URL checks

  • Duplicate URLs
  • New path patterns
  • Unexpected query parameters
  • Redirect rate
  • Profile URL validation sample

Coverage checks

  • Pagination endpoints reached
  • Expected geographic or specialty partitions visited
  • Sitemap and listing overlap where relevant
  • Empty result pages explained

Do not calculate provider changes when discovery itself is incomplete.

Fetch and parser QA

Fetch metrics

  • Successful responses
  • Redirects
  • Not found
  • Rate limited
  • Forbidden or blocked
  • Timeout and retry exhaustion
  • Median and percentile response size

Parser metrics

  • Required name coverage
  • Specialty coverage
  • Location coverage
  • Phone and fax coverage
  • NPI coverage where the source publishes it
  • Page-template distribution
  • Unknown-template count

A provider source may legitimately have low NPI coverage. The alert should compare the source with its own baseline, not with a universal expectation.

Publication gates

A refresh should be promoted only when:

  • Discovery completed across the approved scope.
  • Fetch errors are within reviewed thresholds.
  • Required-field coverage is acceptable.
  • Structural changes have been investigated.
  • Added and missing counts are plausible.
  • Sample records passed source comparison.
  • Match and normalization versions are recorded.
  • The delivery includes collection dates and source provenance.

For a recurring product, use release states:

text
collected -> parsed -> compared -> qa_review -> approved -> published

Handling source redesigns

A redesign should create a new adapter version.

Recommended workflow:

  1. Pause publication for the affected source.
  2. Preserve the failed run and raw responses.
  3. Identify new page templates or endpoints.
  4. Update fixtures and parser tests.
  5. Run the adapter against historical and current samples.
  6. Compare field coverage with the previous adapter.
  7. Perform a controlled full recollection.
  8. Approve the new adapter version and release.

Do not patch selectors until one sample page works and then assume the directory is fixed. Large sources often contain several templates.

Sitemaps and lastmod

The Sitemaps protocol (opens in a new tab) allows a URL entry to include lastmod. It can be useful when the publisher maintains it accurately.

But verify its behavior before relying on it:

  • Does every provider URL appear?
  • Does lastmod change when meaningful profile data changes?
  • Is it the page modification time or sitemap generation time?
  • Are removed URLs represented anywhere?
  • Does the source use several sitemap indexes?

A trustworthy lastmod can optimize fetching. It should not eliminate periodic discovery validation.

NPPES reconciliation

For an NPPES warehouse:

  1. Load a known full monthly baseline.
  2. Store the release date and file identity.
  3. Apply weekly increments in order.
  4. Apply the documented deactivation update.
  5. Record every applied release.
  6. Reconcile periodically against a later full file.
  7. Keep NPPES history separate from hospital-directory observations.

If an incremental file is missed or applied twice, the warehouse can drift. Idempotent release tracking is mandatory.

Refresh delivery report

Provide more than a replacement CSV.

text
Source: Example Health System
Previous collection: 2025-09-10
Current collection: 2026-07-21
Previous profiles: 4,812
Current profiles: 4,905
Added: 132
Changed: 407
Missing/unconfirmed: 39
Unchanged: 4,327
Fetch exceptions: 3
Unknown templates: 0

Example values are fictional. The structure is what matters.

Frequently asked questions

Should hospital provider directories always be fully recollected?

Not always, but full snapshots are often the clearest approach when a source lacks a reliable public change feed. A hybrid method can reduce work if discovery remains complete and change signals have been tested.

How frequently should physician directory data be refreshed?

Base the cadence on the business use, source volatility, cost of staleness, and access constraints. A recurring consulting engagement may need a fresh six- or twelve-month snapshot, while an active product may require a source-specific scheduled cadence.

Can sitemap lastmod drive incremental updates?

Only after validating that the publisher maintains it consistently and that the sitemap contains the complete in-scope URL set. Continue monitoring full discovery because lastmod does not represent removed URLs reliably in every implementation.

When should a missing profile be marked inactive?

Use an explicit confirmation rule. Missing from one run is not enough. Confirm across successful snapshots, source status, redirects, permanent not-found responses, or manual review depending on the project's risk.

Should refreshes overwrite old records?

No. Keep immutable source observations or snapshot history and promote an approved current release. Historical records are required to explain when a value changed and to repair matching or normalization decisions.

Next step

Read How to Build a Multi-Source Healthcare Provider Data Pipeline for the full architecture. For recurring source collection, review Orzaen's Healthcare Provider Data Collection offer and three-year provider-directory delivery case study.

Sources

Tags

Provider DataChange DetectionWeb ScrapingData Quality

Need this fixed?

Have the same system problem?

Share what is manual, messy, broken, or disconnected. We’ll review the cleanest next step.

Get System Review