Short answer
Validate scraped data at four levels before delivery: structural validity, semantic validity, coverage, and reconciliation. Check the schema and types, test whether values make sense for the source, compare output with the expected entity universe, and account for every input or discovered URL. Preserve source URLs, timestamps, raw values, parser versions, and exception records so the dataset can be audited and corrected.
A file that opens successfully is not necessarily a valid dataset. Ten thousand clean rows may still be a failed delivery if the crawl was expected to contain twelve thousand.
This article owns the pre-delivery QA model. For the complete collection architecture around it, read How to Build a Scalable Web Scraping Pipeline.
The four validation layers
| Layer | Question | Example failure |
|---|---|---|
| Structural | Does the record have the required shape and types? | price_amount contains "Call for price" |
| Semantic | Does the value make sense for this field and source? | Longitude is outside -180 to 180 |
| Coverage | Did the crawl include the expected entities and partitions? | One state or category is completely absent |
| Reconciliation | Can every input or frontier item be accounted for? | Retried records shifted out of input order |
All four are necessary. Schema checks cannot detect an omitted category, and a row-count check cannot detect values mapped into the wrong columns.
Define acceptance before collection
Quality rules should be part of the data contract.
For every field, document:
- Definition
- Type
- Required or optional status
- Accepted formats or values
- Source location
- Normalization rule
- Null representation
- Uniqueness expectation
- Whether several values are allowed
For the dataset, document:
- Expected source or input set
- Expected partitions, such as state, category, alphabet, or date
- Primary source key
- Duplicate policy
- Required provenance
- Delivery format and encoding
- Exception-handling rule
- Sampling and acceptance process
Do not promise "100% accurate data" without defining a truth set, measurement method, sample design, and error treatment. Report testable metrics instead.
1. Structural validation
Structural checks verify the record's shape.
Typical rules include:
- Required keys exist.
- Values have the expected type.
- Strings meet length or pattern constraints.
- Numbers fall within declared bounds.
- Enumerations contain only allowed values.
- Arrays and nested objects follow the agreed structure.
- Extra or unexpected fields are handled deliberately.
JSON Schema validation (opens in a new tab) defines assertion keywords for types, required fields, numeric ranges, string patterns, array uniqueness, object properties, and related constraints.
A small product schema could look like:
{
"type": "object",
"required": ["source_id", "source_url", "collected_at", "product_key"],
"properties": {
"source_id": {"type": "string", "minLength": 1},
"source_url": {"type": "string", "format": "uri"},
"collected_at": {"type": "string", "format": "date-time"},
"product_key": {"type": "string", "minLength": 1},
"price_amount": {"type": ["number", "null"], "minimum": 0},
"currency": {"type": ["string", "null"], "pattern": "^[A-Z]{3}$"}
}
}Schema validation should run before file delivery and, ideally, immediately after each record is normalized so bad records enter an exception path early.
2. Semantic validation
Semantic checks test meaning, not just format.
Examples:
- A
sale_priceshould not exceed the source'slist_priceunless the source semantics allow it. - Latitude and longitude should correspond to the expected geography.
- A provider specialty should not be parsed from the page's navigation menu.
- A product's SKU should follow the pattern observed for that source family.
- A property bedroom count should not be parsed from the review score.
- A social-profile follower count should map back to the requested handle.
These rules are source-aware. A missing fax number can be normal for one provider directory and a regression for another source where fax coverage is usually high.
Useful semantic techniques include:
- Range and set checks
- Cross-field consistency
- Source-specific regular expressions
- Reference-data lookups
- Geographic containment
- Date ordering
- Identifier checksum or format checks
- Controlled normalization dictionaries
Keep raw and normalized values separate. When a normalization rule changes, the dataset can be reprocessed without another crawl.
3. Coverage validation
Coverage answers: did we collect the intended universe?
Build a count funnel for each source and partition:
seeds
-> listing pages discovered
-> entity URLs discovered
-> entity requests attempted
-> successful responses
-> parsed observations
-> valid canonical records
-> delivered recordsStore the count and the identities that moved between stages. A gap of 500 should produce an exception list, not only a lower final total.
Coverage needs a denominator
Possible denominators include:
- Client-supplied identifiers
- URLs in approved sitemaps
- URLs reached from every approved category branch
- Every geographic partition in a source map
- Every row in a previous snapshot plus newly discovered records
- A source-provided total count, after testing its reliability
The denominator itself may be imperfect. Document how it was derived and which source areas cannot be measured exactly.
Watch distribution, not only totals
An expected total can conceal missing segments.
Check records by:
- Source
- Category
- Geography
- Entity type
- Pagination partition
- Status
- Parser version
- Collection hour or worker batch
A total of 200,000 can look plausible even if one large state is missing and another partition was duplicated.
Use Web Scraping Pagination Patterns to connect pages, cursors, infinite scroll, and segmented searches to an explicit coverage denominator.
4. Reconciliation validation
Reconciliation accounts for every requested or discovered item.
For input enrichment, deliver statuses such as:
| Status | Meaning |
|---|---|
matched | Input mapped to a valid source record |
not_found | Source completed normally but no approved match was found |
invalid_input | Input could not be processed under the contract |
source_unavailable | Source could not be evaluated during the run |
retry_exhausted | Temporary failures exceeded the retry policy |
ambiguous | More than one plausible match requires review |
Do not drop failed inputs and return only successful rows. The consumer needs to distinguish "not found" from "not attempted" and "temporarily failed."
In Orzaen's Instagram bulk enrichment result, more than 400,000 handles were processed in 24 hours while preserving the input order. Batching and retries had to reconcile results with the original list; otherwise a complete-looking output could attach fields to the wrong handle.
For long-running collection, the crawl-resumption guide shows how terminal task state and idempotent writes keep this reconciliation intact after a restart.
Uniqueness and duplicate validation
Duplicates arise at several levels.
Request duplicates
The same canonical URL is scheduled more than once because tracking parameters, fragments, redirects, or discovery paths differ.
Source-record duplicates
The same entity appears on several pages or locations within one source.
Canonical-entity duplicates
Two source records refer to the same real entity after normalization or matching.
Exact output duplicates
The delivery contains identical rows because batches were appended twice or a retry was not idempotent.
Use different keys for different questions:
request key = canonicalized URL or request identity
source key = (source_id, source_entity_id)
observation key = (source key, collected_at or cycle)
canonical key = approved entity-resolution keyNever deduplicate only on a person's or company's name when names can legitimately repeat.
Missing values need explicit meanings
An empty cell can mean:
- The source published no value.
- The page did not load correctly.
- The parser failed.
- The field does not apply.
- Access was not approved.
- The source record was removed.
- Enrichment produced no confident match.
Represent these states separately where they affect decisions.
For example:
| Value | Status | Interpretation |
|---|---|---|
null | not_published | Page was valid; field was absent |
null | parser_error | Required label existed but parsing failed |
null | source_unavailable | Page could not be evaluated |
null | not_applicable | Field does not apply to this entity type |
This prevents missingness from being mistaken for a real-world fact.
Validate source and collection provenance
Every delivered record should be traceable to its source observation.
At minimum retain:
- Source identifier
- Source URL or request identity
- Collection timestamp
- Raw source key
- Parser or mapping version
- Raw source value for normalized fields
- Run or collection-cycle identifier
When records are combined across sources, retain field-level provenance if different sources contribute different columns.
This enables three necessary questions:
- Where did this value come from?
- When was it observed?
- Which code and rule produced the delivered form?
Test the export, not only the database
A valid internal table can still become a broken CSV or Excel file.
Check:
- Encoding and delimiters
- Header order and spelling
- Row count after export
- Quoting of commas, newlines, and quotes
- Spreadsheet row or cell limitations
- Date and identifier formatting
- Leading zeros
- Large integers and scientific notation
- Formula-like text in spreadsheet cells
- Line endings when the consumer requires a specific format
- JSON encoding and nesting
For Excel delivery, identifiers such as postal codes, phone numbers, and long numeric product IDs may need explicit text formatting so the spreadsheet does not alter them.
Use sampling correctly
Sampling complements automated checks; it does not replace them.
Use stratified samples that include:
- Each source or source family
- High- and low-volume partitions
- Records with missing optional fields
- Edge-case values
- Retried records
- New or changed parser versions
- First, middle, and final pagination segments
- Duplicate candidates
Compare sampled fields with the stored source observation or current page, accounting for data that may have changed after collection.
Record the sample design, date, reviewer, fields checked, errors found, and corrective action. Do not convert a tiny convenience sample into an unsupported global accuracy claim.
Build a delivery scorecard
A practical run summary can contain:
| Metric | Result |
|---|---|
| Approved sources | Count |
| Sources completed | Count and percentage |
| Entity URLs discovered | Count |
| Valid records | Count |
| Exception records | Count by class |
| Duplicate source keys | Count |
| Required-field coverage | Percentage by field and source |
| Optional-field coverage | Percentage by field and source |
| Records with provenance | Count and percentage |
| Collection window | Start and end timestamps |
| Parser versions | Version list |
The table should contain the actual run results, not fixed marketing targets.
A practical validation sequence
- Validate the input and approved source set.
- Reconcile discovered URLs with seeds and partitions.
- Validate response status, type, and signature.
- Parse into source observations.
- Run source-specific required-field checks.
- Normalize without overwriting raw fields.
- Run canonical schema and semantic checks.
- Test uniqueness at request, source, and entity levels.
- Reconcile every input or frontier item.
- Compare source and partition distributions with baselines.
- Review a stratified sample.
- Validate the final file or API payload.
- Deliver the valid dataset, exceptions, and run summary.
Common QA mistakes
Checking only whether the job crashed
A scraper can complete successfully while returning a consent page or empty selector result.
Using one accuracy percentage
Field accuracy, source coverage, record completeness, and entity matching are different measurements. Report them separately with methods.
Dropping exceptions
Removing invalid rows without reporting them makes the output look cleaner while hiding coverage loss.
Treating a missing page as a deleted entity
A timeout, access change, temporary error, or new URL may also explain absence. Report the observation and confidence separately.
Deduplicating too early
Two source observations may represent different locations, variants, roles, or collection times. Preserve source records before canonical entity resolution.
Validating before export but not after
CSV and spreadsheet transformations can alter values. Reopen and recheck the delivered artifact.
Frequently asked questions
What is the difference between data validation and data cleaning?
Validation tests whether data satisfies defined rules. Cleaning transforms values, such as standardizing phone formats or whitespace. Cleaning should not silently turn an invalid or missing observation into an assumed fact.
How do I measure scraping accuracy?
Define the field, truth source, sampling method, sample size, and error categories. Measure field correctness separately from entity coverage and matching accuracy. Without that method, one accuracy percentage is not meaningful.
Should invalid records be delivered?
Deliver valid records in the main dataset and provide an exception report when invalid or unresolved items matter to scope. Include the reason and source identity so they can be reviewed or reprocessed.
How can I validate millions of rows?
Run automated checks across every row, partition, and key; use aggregates and anomaly detection for distributions; then conduct stratified human sampling. Store invalid records separately instead of stopping the entire pipeline for every exception.
Is a row-count comparison enough?
No. The same total can contain duplicated segments and omitted segments. Reconcile identities and compare distributions across sources, categories, geography, status, and parser version.
Next step
Orzaen's web scraping and data extraction service includes schema, cleaning, deduplication, validation, and delivery planning. Connect the acceptance checks to pipeline monitoring so incomplete data is blocked before export. For examples where reconciliation was central, review the Instagram bulk enrichment result, VRBO data-transformation result, and Kramp crawl result.
Sources
- JSON Schema validation specification (opens in a new tab)
- JSON Schema: object validation (opens in a new tab)
- JSON Schema: numeric validation (opens in a new tab)
- Scrapy statistics collection (opens in a new tab)
- Scrapy item pipelines (opens in a new tab)
- W3C Data on the Web Best Practices: data provenance (opens in a new tab)
- RFC 4180: Common Format and MIME Type for CSV Files (opens in a new tab)

