Short answer
A reliable healthcare provider data pipeline uses one adapter per source and a shared downstream model. Each adapter discovers public profile URLs, captures raw source responses, parses source-specific fields, and emits versioned records. The shared pipeline then normalizes fields, resolves entities conservatively, runs source-aware QA, and delivers records with source URLs and collection timestamps. Raw collection, normalization, and analytical conclusions should remain separate layers.
Why a single universal scraper fails
Hospital and healthcare directories can look similar to a user while being structurally different to a pipeline.
One source may use:
- Server-rendered profile pages
- Client-rendered search results
- Page-number pagination
- Cursor-based API pagination
- A public XML sitemap
- JSON-LD embedded in profile pages
- Separate location and provider endpoints
- Search forms that require specialty and geography parameters
The provider fields also vary. A hospital profile may publish department and faculty title; a practice site may publish only name, specialty, phone, and location.
Trying to force every source through one parser creates fragile selectors and unclear field semantics. The stable pattern is:
Source adapter -> Raw observation -> Parsed source record
-> Shared normalization -> QA -> DeliveryStage 0: access and scope review
Before engineering begins, review:
- Public accessibility
- Source terms and permitted use
robots.txtdirectives- Rate expectations and technical constraints
- Requested fields and intended use
- Privacy and data sensitivity
- Whether authentication, patient portals, or restricted systems are out of scope
The Robots Exclusion Protocol is standardized in RFC 9309 (opens in a new tab). Robots directives are one part of a broader source review; they are not a substitute for terms, permission, privacy, or legal analysis.
For provider-directory projects, define the data boundary explicitly: public professional provider information only, with no patient, clinical, appointment, or claims data unless a separate authorized workflow clearly permits it.
Stage 1: source registry
Do not store source configuration only in code. A source registry makes scope and operations visible.
source_id: example_health_system
source_type: hospital_directory
base_url: https://example.org
adapter_version: 3.2.0
discovery_mode: sitemap_and_listing
required_fields:
- provider_name_raw
- source_url
optional_fields:
- credentials_raw
- specialty_raw
- organization_name_raw
- department_raw
- locations_raw
refresh_mode: full_snapshotOperational fields can include last successful collection, expected record range, access notes, parser owner, and whether a source is currently paused.
Stage 2: URL discovery
Separate discovery from profile parsing. Discovery should output a deduplicated manifest:
source_id
profile_url
discovered_from
discovered_at
discovery_run_idPossible public discovery paths include:
- Directory listing pages
- Specialty or location indexes
- XML sitemaps
- Links between practitioner and location pages
- Public structured data
- Permitted public endpoints used by the directory itself
Sitemaps can help, but they are not guaranteed to contain every profile. The Sitemaps protocol (opens in a new tab) defines URL and optional lastmod elements; it does not promise that a site's lastmod is complete or precise enough to drive your refresh logic.
Store the complete manifest for every run. A sudden drop from 8,000 discovered URLs to 600 should stop or quarantine the pipeline before it overwrites a valid snapshot.
Stage 3: raw collection
The raw layer should be append-only for the duration required by the project and source constraints.
Recommended metadata:
- Source ID
- URL
- Fetch timestamp
- HTTP status
- Content type
- Content hash
- Adapter version
- Retry count
- Response body or permitted raw snapshot reference
- Error classification
Do not overwrite the previous response before parsing the new one successfully. If a source redesign breaks the parser, the raw snapshot lets you repair the transformation without collecting the page again.
Idempotent fetch records
A stable raw key can be:
(collection_run_id, canonical_url)If the job restarts, it can skip successful URLs and retry failed ones without duplicating the run.
Stage 4: source adapters
Use a narrow adapter contract.
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime
from typing import Iterable, Mapping, Protocol
@dataclass(frozen=True)
class RawPage:
source_id: str
url: str
collected_at: datetime
body: str
@dataclass(frozen=True)
class SourceProviderRecord:
source_id: str
source_url: str
collected_at: datetime
provider_name_raw: str
credentials_raw: tuple[str, ...] = ()
specialties_raw: tuple[str, ...] = ()
organization_name_raw: str | None = None
department_raw: str | None = None
locations_raw: tuple[Mapping[str, str], ...] = ()
extra: Mapping[str, object] = field(default_factory=dict)
class ProviderSourceAdapter(Protocol):
source_id: str
version: str
def discover(self) -> Iterable[str]: ...
def parse(self, page: RawPage) -> SourceProviderRecord: ...The adapter emits source values, not final analytical truth. For example, it should emit specialties_raw=("Heart Failure",) rather than silently mapping the provider to a custom cardiology hierarchy.
Stage 5: normalization
Normalization belongs after source parsing because it changes over time.
import re
NON_DIGITS = re.compile(r"\D+")
def normalize_us_phone(value: str | None) -> str | None:
if not value:
return None
digits = NON_DIGITS.sub("", value)
if len(digits) == 11 and digits.startswith("1"):
digits = digits[1:]
if len(digits) != 10:
return None
return f"+1{digits}"In production, return both a normalized value and a reason when normalization fails. Keep the raw phone field so a reviewer can distinguish a missing phone from an unusual extension or international number.
Other normalization tasks include:
- State and country codes
- Postal codes
- Credentials
- Specialty taxonomy
- Organization names
- Address components
- URL canonicalization
Every normalized field should have a rule version.
Stage 6: provider and organization resolution
Entity resolution is the highest-risk transformation because it can merge different people or split the same provider.
Use strong identifiers first:
- NPI published directly by the source
- Approved exact crosswalk
- Exact name plus several aligned attributes
- Probabilistic candidate requiring review
Do not automatically assign an NPI from name similarity alone. CMS notes that the NPI is an identifier and that its issuance does not validate licence or credential status. Read NPPES vs Hospital Provider Directories for the source distinction.
A match record should contain:
source_record_id
candidate_provider_id
match_rule
matched_fields
conflicting_fields
confidence_band
review_statusThis lets the system improve matching rules without losing the original decision trail.
Stage 7: relationship modeling
A provider is not a single flat row.
Provider
├── has source observations
├── performs roles for organizations
├── practices at locations
├── is presented under specialties
└── may use several public contact pointsThe HL7 FHIR `PractitionerRole` resource (opens in a new tab) models roles, services, specialties, organizations, and locations around a practitioner. Even a non-FHIR warehouse should preserve those many-to-many relationships.
Our Provider Data Schema Design article provides the tables.
Stage 8: source-aware QA
Global checks are not enough. Track quality by source and run.
Discovery checks
- Profile URL count compared with previous runs
- Duplicate URL rate
- Unexpected path patterns
- Pagination completeness
Fetch checks
- Success, redirect, block, timeout, and not-found rates
- Response-size distribution
- Content-type changes
- Repeated identical bodies across many URLs
Parse checks
- Required-field coverage
- Optional-field coverage by source
- New or missing page templates
- Empty names or invalid URLs
- Parser errors by template
Dataset checks
- Duplicate provider-location relationships
- Suspicious NPI matches
- Invalid state and postal-code values
- Added, changed, and missing records versus the last run
- Source URL and timestamp coverage
Do not publish or overwrite the current dataset when a structural alert remains unresolved.
Stage 9: change detection
Create a canonical comparison payload that excludes volatile fields such as collection time.
import hashlib
import json
from collections.abc import Mapping
def stable_record_hash(payload: Mapping[str, object]) -> str:
encoded = json.dumps(
payload,
sort_keys=True,
separators=(",", ":"),
ensure_ascii=False,
).encode("utf-8")
return hashlib.sha256(encoded).hexdigest()Compare hashes only after canonicalization. A different order of specialties should not create a false change if order has no meaning.
The refresh pipeline should classify:
- Added source record
- Changed source record
- Unchanged source record
- Missing from current discovery
- Explicitly removed or inactive when the source says so
"Missing" is not the same as "confirmed removed." The page may be temporarily unavailable or the discovery logic may have failed. See Provider Directory Refresh and Change Detection.
Stage 10: delivery
Choose the output based on the consumer.
CSV or Excel
Best for analyst review and one-time source batches. Deliver one file per source plus an optional combined file.
JSON
Useful when profiles contain repeated specialties, locations, affiliations, or nested source details.
Database tables
Best for recurring refreshes, applications, and change history.
API
Useful when downstream systems need controlled query access. The API should expose freshness and provenance, not just normalized fields.
Operating the pipeline
Production systems need:
- Scheduled or approved manual runs
- Concurrency limits by source
- Retries with bounded backoff
- Structured logs
- Run-level metrics
- Parser-version tracking
- Alerting on count and field-coverage anomalies
- Resume support
- Raw and normalized retention rules
- Clear source pause and disable controls
The goal is not maximum request speed. The goal is a reproducible dataset produced within the approved access and delivery constraints.
Frequently asked questions
Should every source use the same parser?
No. Use a shared adapter contract, but keep source-specific discovery and parsing logic separate. This isolates failures and allows each source to evolve without destabilizing the entire pipeline.
Should raw HTML be stored?
It can improve reproducibility and parser repair when storage is permitted and appropriate. Define retention, security, source, and privacy constraints. At minimum, preserve raw field values, source URL, collection time, and parser version.
Can NPPES be used as the master provider table?
It can be a strong identity input, but do not overwrite source-specific hospital or directory observations with NPPES values. Store NPPES as its own source and link records under explicit matching rules.
How do you detect a directory redesign?
Monitor discovery count, response size, page-template signatures, parser errors, and field coverage by source. A sharp change in several metrics should quarantine the run before delivery.
Is a missing provider profile a deletion?
Not automatically. It may indicate a directory change, transient failure, broken discovery path, redirected page, or actual removal. Use a confirmation rule before converting missing into inactive.
Next step
See the Healthcare Provider Data Collection offer for project scoping, or review the Psychology Today extraction case study for an example involving more than 406,000 public therapist profiles.

