Short answer
Refresh provider-directory data by creating a new source snapshot, not by editing the previous export in place. Compare stable record keys and canonical field hashes to classify added, changed, unchanged, and missing records. Use source-specific QA gates before publishing the new snapshot. Full recollection is often safer for hospital directories; trustworthy incremental files can be used when an authoritative source, such as NPPES, explicitly provides them.
Why refreshes fail
A refresh is not simply "run the scraper again."
Common failures include:
- The source changed its page structure.
- Pagination stopped early.
- A search endpoint returned only a default subset.
- The new parser produced empty fields.
- A transient error made profiles appear deleted.
- Normalization changed and generated false differences.
- Records were updated in place, erasing the prior source state.
- Two different providers were assigned the same match key.
A safe refresh treats collection, comparison, and publication as separate decisions.
Full snapshot versus incremental update
| Method | How it works | Strength | Risk |
|---|---|---|---|
| Full snapshot | Recollect every approved record from the source | Detects broad structural and roster changes | More requests and processing |
| Source-provided incremental | Ingest an authoritative change file or feed | Efficient and explicit | Depends on correct baseline and update application |
| Sitemap-guided refresh | Fetch URLs whose reliable modification signal changed | Can reduce work | lastmod may be absent or unreliable |
| Profile-level monitoring | Revisit a known URL set | Good for defined rosters | Misses newly added profiles without discovery |
| Hybrid | Refresh discovery fully and fetch changed or selected profiles | Balances coverage and cost | More operational complexity |
For many hospital and health-system directories, a full snapshot is the most defensible method because the source does not expose a complete public change log.
CMS is different. The NPPES downloadable-file page (opens in a new tab) currently supplies a monthly full Version 2 file, weekly incremental files, and a monthly deactivation update. A pipeline can use the full file as a baseline and apply the documented updates while preserving release metadata.
Choosing a refresh cadence
There is no universal "provider data must be scraped every X days" rule for this type of project.
Choose cadence based on:
- Business decision frequency
- Expected source volatility
- Cost of stale data
- Source size and access constraints
- Whether the dataset supports a one-time study or a live product
- Whether a reliable update feed exists
- Client tolerance for missing and unconfirmed changes
Examples:
| Use case | Possible cadence | Reasoning |
|---|---|---|
| One-time market study | One fresh collection | Historical monitoring may add no value |
| Recurring consulting market | Fresh collection per engagement, sometimes after 6–12 months | The decision is tied to a new client snapshot |
| Provider-search product | Scheduled source-specific cadence | Product freshness is an ongoing requirement |
| High-value monitored roster | Weekly or monthly, if permitted and justified | Changes affect an active workflow |
| NPPES warehouse | Weekly increments plus periodic full reconciliation | CMS supplies both update modes |
State the collection date prominently. "Current" without a date is not a useful freshness claim.
Snapshot architecture
Each run should have an immutable identity.
collection_run
├── source_id
├── started_at
├── completed_at
├── adapter_version
├── discovery_count
├── fetch_metrics
├── parse_metrics
├── qa_status
└── publication_statusSource records should be unique within a run:
(run_id, canonical_source_key)The current published dataset is then a pointer to an approved run or derived release. If QA fails, the earlier release remains available.
Stable source keys
Choose the strongest key the source provides:
- Stable source-native profile ID
- Canonical profile URL
- Published NPI plus source-specific relationship key
- Conservative composite key
Do not use normalized provider name alone. Two people can share a name, and one person's displayed name can change.
The source key identifies the source record. It does not necessarily identify the canonical provider across all sources.
Canonical comparison payload
Exclude fields that change on every run, such as collected_at, run ID, retry count, and raw response storage path.
from __future__ import annotations
import hashlib
import json
from collections.abc import Mapping, Sequence
from typing import Any
def sorted_unique(values: Sequence[str]) -> list[str]:
return sorted({value.strip() for value in values if value.strip()})
def comparison_payload(record: Mapping[str, Any]) -> dict[str, Any]:
return {
"provider_name_raw": record.get("provider_name_raw"),
"credentials_raw": sorted_unique(record.get("credentials_raw", [])),
"specialties_raw": sorted_unique(record.get("specialties_raw", [])),
"organization_name_raw": record.get("organization_name_raw"),
"locations": record.get("locations", []),
"phone_raw": record.get("phone_raw"),
"fax_raw": record.get("fax_raw"),
"npi_raw": record.get("npi_raw"),
}
def payload_hash(payload: Mapping[str, Any]) -> str:
serialized = json.dumps(
payload,
sort_keys=True,
separators=(",", ":"),
ensure_ascii=False,
).encode("utf-8")
return hashlib.sha256(serialized).hexdigest()This code is illustrative, not a complete production comparator. Nested locations must also be canonicalized, and some arrays are order-sensitive while others are not.
Version the comparison logic. A change in normalization should not masquerade as a real source change.
Change classification
For every stable source key:
| Previous snapshot | Current snapshot | Classification |
|---|---|---|
| Absent | Present | Added |
| Present, same hash | Present, same hash | Unchanged |
| Present, different hash | Present, different hash | Changed |
| Present | Absent | Missing/unconfirmed |
| Explicit inactive state | Explicit inactive state | Source-confirmed inactive |
Then create field-level events for changed records.
{
"source_key": "provider-123",
"change_type": "field_changed",
"field": "phone_raw",
"previous_value": "212-555-0100",
"current_value": "212-555-0199",
"previous_run_id": "run-a",
"current_run_id": "run-b"
}Example values are fictional.
Missing is not deleted
A source record can disappear because:
- The provider profile was removed.
- The URL changed.
- The source search omitted it temporarily.
- Pagination failed.
- The directory added a filter.
- The request was blocked or returned an error page.
- The parser rejected the page.
Use confirmation rules such as:
- Missing in two successful full snapshots
- Old URL returns an explicit permanent not-found response
- Source publishes an inactive or removed status
- Redirect resolves to a successor profile
- Manual review confirms removal for high-impact records
Until then, classify the record as missing_unconfirmed rather than deleted.
Discovery QA before record comparison
Compare the run with the last known-good source baseline.
Record-count checks
- Total discovered profiles
- Profiles per specialty, state, or organization partition
- Percentage change from previous successful run
- New and missing URL ratio
URL checks
- Duplicate URLs
- New path patterns
- Unexpected query parameters
- Redirect rate
- Profile URL validation sample
Coverage checks
- Pagination endpoints reached
- Expected geographic or specialty partitions visited
- Sitemap and listing overlap where relevant
- Empty result pages explained
Do not calculate provider changes when discovery itself is incomplete.
Fetch and parser QA
Fetch metrics
- Successful responses
- Redirects
- Not found
- Rate limited
- Forbidden or blocked
- Timeout and retry exhaustion
- Median and percentile response size
Parser metrics
- Required name coverage
- Specialty coverage
- Location coverage
- Phone and fax coverage
- NPI coverage where the source publishes it
- Page-template distribution
- Unknown-template count
A provider source may legitimately have low NPI coverage. The alert should compare the source with its own baseline, not with a universal expectation.
Publication gates
A refresh should be promoted only when:
- Discovery completed across the approved scope.
- Fetch errors are within reviewed thresholds.
- Required-field coverage is acceptable.
- Structural changes have been investigated.
- Added and missing counts are plausible.
- Sample records passed source comparison.
- Match and normalization versions are recorded.
- The delivery includes collection dates and source provenance.
For a recurring product, use release states:
collected -> parsed -> compared -> qa_review -> approved -> publishedHandling source redesigns
A redesign should create a new adapter version.
Recommended workflow:
- Pause publication for the affected source.
- Preserve the failed run and raw responses.
- Identify new page templates or endpoints.
- Update fixtures and parser tests.
- Run the adapter against historical and current samples.
- Compare field coverage with the previous adapter.
- Perform a controlled full recollection.
- Approve the new adapter version and release.
Do not patch selectors until one sample page works and then assume the directory is fixed. Large sources often contain several templates.
Sitemaps and lastmod
The Sitemaps protocol (opens in a new tab) allows a URL entry to include lastmod. It can be useful when the publisher maintains it accurately.
But verify its behavior before relying on it:
- Does every provider URL appear?
- Does
lastmodchange when meaningful profile data changes? - Is it the page modification time or sitemap generation time?
- Are removed URLs represented anywhere?
- Does the source use several sitemap indexes?
A trustworthy lastmod can optimize fetching. It should not eliminate periodic discovery validation.
NPPES reconciliation
For an NPPES warehouse:
- Load a known full monthly baseline.
- Store the release date and file identity.
- Apply weekly increments in order.
- Apply the documented deactivation update.
- Record every applied release.
- Reconcile periodically against a later full file.
- Keep NPPES history separate from hospital-directory observations.
If an incremental file is missed or applied twice, the warehouse can drift. Idempotent release tracking is mandatory.
Refresh delivery report
Provide more than a replacement CSV.
Source: Example Health System
Previous collection: 2025-09-10
Current collection: 2026-07-21
Previous profiles: 4,812
Current profiles: 4,905
Added: 132
Changed: 407
Missing/unconfirmed: 39
Unchanged: 4,327
Fetch exceptions: 3
Unknown templates: 0Example values are fictional. The structure is what matters.
Frequently asked questions
Should hospital provider directories always be fully recollected?
Not always, but full snapshots are often the clearest approach when a source lacks a reliable public change feed. A hybrid method can reduce work if discovery remains complete and change signals have been tested.
How frequently should physician directory data be refreshed?
Base the cadence on the business use, source volatility, cost of staleness, and access constraints. A recurring consulting engagement may need a fresh six- or twelve-month snapshot, while an active product may require a source-specific scheduled cadence.
Can sitemap lastmod drive incremental updates?
Only after validating that the publisher maintains it consistently and that the sitemap contains the complete in-scope URL set. Continue monitoring full discovery because lastmod does not represent removed URLs reliably in every implementation.
When should a missing profile be marked inactive?
Use an explicit confirmation rule. Missing from one run is not enough. Confirm across successful snapshots, source status, redirects, permanent not-found responses, or manual review depending on the project's risk.
Should refreshes overwrite old records?
No. Keep immutable source observations or snapshot history and promote an approved current release. Historical records are required to explain when a value changed and to repair matching or normalization decisions.
Next step
Read How to Build a Multi-Source Healthcare Provider Data Pipeline for the full architecture. For recurring source collection, review Orzaen's Healthcare Provider Data Collection offer and three-year provider-directory delivery case study.

