Short answer
Choose the data surface before choosing the scraping tool. Prefer an approved official API or export when it provides the required fields and acceptable coverage. Otherwise inspect server-rendered HTML, embedded structured data, and the page's XHR, fetch, or GraphQL responses. Use browser automation when the target genuinely depends on JavaScript execution, interaction, or browser-held state. The best source is the simplest reliable surface that meets the data contract—not automatically the most structured one.
A website is a presentation layer. The visible card, table, or profile may be assembled from several underlying responses, and those responses may not contain the same fields.
This article focuses on choosing the source surface. For queues, raw storage, validation, monitoring, and delivery, see How to Build a Scalable Web Scraping Pipeline.
The four source surfaces
| Surface | What it is | Strength | Limitation |
|---|---|---|---|
| Official API or export | A documented interface, feed, or downloadable file | Clearer contract and structured output | May omit fields, limit coverage, or require a plan and credentials |
| Structured page response | JSON, GraphQL, or other response used by the public page | Efficient and close to the application's own data flow | Can be undocumented and may change without notice |
| Server-rendered HTML | The document returned for the page URL | Simple, auditable, and often sufficient | Markup and selectors vary; some visible content may be absent |
| Browser-rendered state | DOM and network state after JavaScript and interaction | Handles genuine browser workflows | Highest runtime and maintenance cost |
These surfaces can coexist. A product page may expose a title in HTML, price and inventory in GraphQL, and delivery estimates only after a location is selected in the browser.
Start with a field-source matrix
Do not decide "we are using the API" until the API has been compared with the required schema.
| Required field | Official API | Page JSON | HTML | Browser interaction |
|---|---|---|---|---|
| Product ID | Yes | Yes | Sometimes | Yes |
| Variant price | Limited | Yes | Sometimes | Yes |
| Availability by location | No | After location request | No | Yes |
| Description | Yes | Yes | Yes | Yes |
| Breadcrumb path | No | Sometimes | Yes | Yes |
For each field, record:
- Where the value exists
- Whether it is raw or computed
- Required parameters or state
- Pagination and coverage behavior
- Stability and documentation
- Rate or usage restrictions
- Source identifier
- Freshness
This prevents a common surprise: the team builds around the cleanest endpoint and later discovers that two business-critical fields exist only on the page.
When an official API is the best source
An official API or export is usually the first option to evaluate because it may offer:
- Structured fields
- Explicit identifiers
- Documented authentication
- Defined pagination
- Published limits
- Fewer layout dependencies
It is the best source when its documented use permits the project and it supplies the required records, fields, and freshness.
It is not automatically complete. Ask:
- Does it include every entity shown on the website?
- Are historical, inactive, regional, or variant records excluded?
- Are important display fields absent?
- Does it return the same relationship the buyer needs?
- Are there query or result caps?
- What are the allowed storage and reuse terms?
Ahmad's delivered work includes medical-license API development, Bol.com API integration for EAN/GTIN data, JustCall bulk API extraction, and database-backed scraping projects. That combination matters because API collection still needs schema mapping, pagination, retries, reconciliation, and delivery controls.
When structured page responses are useful
Modern pages often request JSON or GraphQL after the document loads. These responses can contain the records used to render the interface.
Use the browser's developer tools to inspect:
- Request URL and method
- Query parameters or request body
- Response content type
- Pagination cursor or page token
- Stable entity identifiers
- Required headers or session context
- Error and rate-limit responses
- Whether the request is documented or intended for external use
The Chrome DevTools Network panel (opens in a new tab) records requests while a page is open. Playwright can observe page requests and responses through its network APIs (opens in a new tab).
Structured does not mean stable. An undocumented endpoint may change field names, request signatures, nesting, or pagination without notice. Store the parser version and response signature, and monitor field coverage.
GraphQL needs its own coverage model
GraphQL can return precisely selected fields, but large collection still requires:
- A discovery strategy for entity IDs or URLs
- Query and variable versioning
- Pagination or cursor handling
- Partial-error inspection
- Duplicate prevention
- Response preservation where appropriate
- Validation against expected entities
Orzaen's Kramp extraction result combined recursive HTML discovery with GraphQL request and response storage. Its published figures—more than 1.25 million HTML pages, 5 million GraphQL responses, and 864,000 validated URLs—show why GraphQL extraction and crawl coverage were separate concerns.
When server-rendered HTML is the best source
HTML is often the most direct source for public directories, articles, product pages, practice websites, and older server-rendered applications.
Choose HTML when:
- Required fields are present in the initial response.
- URLs are discoverable through links, sitemaps, or a known list.
- The page is readable without browser interaction.
- Source context such as headings, labels, and breadcrumb paths matters.
- A structured response does not improve completeness or reliability.
HTML parsing can be more stable when selectors rely on semantics rather than presentation-only classes:
- Labeled definition lists
- Table headers
- Link relationships
- Stable attributes
- Structured data blocks
- Heading and section relationships
No selector is immune to change. Keep representative fixtures and verify required-field coverage on every run.
When browser automation is necessary
Use a browser when the data or approved workflow depends on browser behavior that cannot be obtained reliably from a simpler surface.
Examples include:
- JavaScript renders records with no usable underlying response.
- A search form must be completed to establish the query state.
- The source uses frames or complex interaction.
- The page requires an approved authenticated session.
- Location or consent state materially changes the response.
- Interaction generates a download or request that cannot otherwise be reproduced reliably.
Playwright and Selenium both automate browsers. Their role is covered in Requests vs Playwright vs Selenium.
Browser automation should use conditions tied to the data, not only page-load events. A page can be "loaded" while the result request is still pending, and it can remain network-active because of analytics even after the data is ready.
How to inspect a website before implementation
1. Define the exact record
Write down the required fields, entity key, expected source set, geography, date, and delivery format. Source research without a data contract tends to optimize the wrong response.
2. Check approved first-party options
Look for:
- API documentation
- Data downloads
- RSS, XML, CSV, or JSON feeds
- Public sitemaps
- Developer portals
- Bulk exports provided to account holders
The Sitemaps protocol (opens in a new tab) defines a format for listing URLs and optional metadata. A sitemap can help discovery, but it does not prove that every business entity is present or that lastmod is trustworthy.
3. Inspect the document response
Disable assumptions created by the rendered page. Fetch or view the original response and search for known values. Check for:
- Visible text
- JSON-LD
- Serialized application state
- Entity IDs
- Canonical URLs
- Pagination links
4. Record network activity
Reload the page with the Network panel open. Filter to Fetch/XHR, then interact with the page one step at a time. Identify which request corresponds to search, pagination, a detail view, or a variant change.
5. Compare fields and coverage
Sample several entity types, not one convenient record. Include records with missing fields, multiple locations or variants, different categories, and later pagination pages.
6. Test repeatability
Run the request again under a new session or process. Determine which values are stable, which expire, and which depend on browser state.
7. Review operational fit
Measure response size, latency, error behavior, limits, and maintenance risk. A technically reachable response may still be a poor production source.
Do not confuse visibility with authority
Several sources may publish different versions of a fact.
For example:
- A manufacturer API may identify the canonical product.
- A retailer page may publish the current selling price.
- A delivery request may publish location-specific availability.
The correct dataset can retain all three observations with source and collection time. It should not silently overwrite them into one supposedly authoritative value unless the business rule is explicit.
This is especially important in multi-source directories and provider data. A registry identifier, hospital relationship, professional-directory profile, and practice location answer different questions.
Pagination changes by surface
| Pattern | Where it appears | Required state |
|---|---|---|
| Numbered links | HTML | Page number or URL |
next link | HTML or API | Next URL |
| Offset and limit | API or JSON | Numeric offset |
| Cursor | API, JSON, or GraphQL | Opaque next-page token |
| Infinite scroll | Browser UI backed by network calls | Scroll or underlying cursor request |
| Geographic search cap | Search UI or API | Partition geography and reconcile overlap |
Never infer completion only because the last returned page was short. Record next-page signals, repeated cursors, duplicate IDs, and the expected partition set.
In Orzaen's VRBO USA extraction result, the collection traversed state, city, and subarea levels to work within a platform result cap. The architecture used a URL queue before detail extraction and delivered 201,717 property records. The lesson is that source coverage can require an explicit partition strategy, not just a faster page loop.
The complete web-scraping pagination guide covers page numbers, offsets, cursors, infinite scroll, segment caps, and termination rules.
Reliability and maintenance trade-offs
| Surface | Common change | Detection signal |
|---|---|---|
| Official API | Version, field, auth, or quota change | Contract error, schema difference, deprecation notice |
| Structured page response | Endpoint, request body, nesting, or signature change | Status shift, response hash, required-path coverage |
| HTML | DOM structure, label, or template change | Selector failure, field coverage, fixture test |
| Browser workflow | Element, timing, frame, state, or interaction change | Step-level timeout, screenshot, network milestone failure |
No source eliminates maintenance. APIs tend to offer clearer contracts; HTML often provides the public presentation context; structured page responses can be efficient but undocumented; browsers reproduce the richest workflow at the highest operating cost.
A source-selection scorecard
Score every candidate surface from 1 to 5 for:
- Field coverage
- Entity coverage
- Permitted use
- Identifier quality
- Pagination clarity
- Stability
- Freshness
- Runtime cost
- Maintenance cost
- Auditability
Reject any source that fails a hard requirement, even if its total score is high. An API with excellent stability but no required price field cannot be the only source for a price-monitoring dataset.
Common mistakes
Choosing a library before inspecting the source
This produces browser-heavy systems for static pages or HTTP-only systems that never receive browser-generated data.
Assuming JSON is more correct than HTML
It may be cleaner but incomplete, cached differently, or intended for another interface component. Compare actual fields and entities.
Capturing one request and skipping session tests
A copied request may rely on expiring cookies, tokens, or browser context. Test a clean start and define renewal behavior.
Ignoring partial GraphQL errors
A GraphQL response can contain both data and errors. A 200 status alone is not a successful record.
Treating browser DOM as raw evidence
The DOM after scripts and interactions is one derived state. Preserve the relevant source response, URL, timestamp, and workflow context where appropriate.
Frequently asked questions
Is an API always better than web scraping?
No. An approved official API is often preferable, but it may not include the required fields, entities, relationships, or freshness. Compare it with the data contract instead of assuming equivalence with the website.
Is calling an undocumented JSON endpoint still web scraping?
It is web data collection from a structured response used by a website. The operational and permission review still matters, and undocumented responses can change without a public compatibility promise.
How do I know whether a page is JavaScript-rendered?
Compare the original HTML response with the rendered page and inspect network activity during load. If the target values are absent from the document but appear after an XHR, fetch, or GraphQL response, JavaScript is assembling the visible state.
Should I scrape HTML or JSON-LD?
Use the source that meets the required field and coverage contract. JSON-LD can provide clean entity fields, while HTML may contain additional source context. In some projects the correct parser combines both and records their origins.
Can a browser be used only for login?
Sometimes. A browser can establish an approved session before an HTTP stage, but the design must handle session expiry, token binding, source rules, and renewal. Do not assume copied browser state is permanent or appropriate for every source.
Next step
For a source review that compares HTML, structured responses, browser state, and delivery requirements, see Orzaen's web scraping and data extraction service. If the browser is needed only for discovery or session state, continue with the JavaScript hybrid-extraction guide. The Kramp GraphQL result and VRBO USA extraction result show two different source and coverage architectures.
Sources
- Chrome DevTools Network panel (opens in a new tab)
- Playwright network documentation (opens in a new tab)
- Requests documentation (opens in a new tab)
- Selenium WebDriver documentation (opens in a new tab)
- Scrapy overview (opens in a new tab)
- Sitemaps protocol (opens in a new tab)
- GraphQL specification (opens in a new tab)
- RFC 9309: Robots Exclusion Protocol (opens in a new tab)

