Healthcare Provider Data Collection — Custom physician and provider roster datasets for healthcare research, consulting, analytics, and health-tech projects.Learn more
OrzaenOrzaen

How Production Scrapers Handle 403, 429 and Temporary Failures

Engineering2026-07-1812 minHira Arif

A source-aware method for classifying forbidden responses, rate limits, timeouts, server errors, expired sessions, and safe retries without hiding data loss.

Short answer

A production scraper should not solve every failed request by retrying immediately or changing IP addresses. It should classify the failure first. A 429 normally asks the client to reduce its request rate; a 403 says the server understood the request but refuses it; 5xx responses and timeouts may be temporary; and an HTTP 200 can still contain an expired-session page or incomplete data.

Each class needs a separate retry budget, pacing rule, stop condition, and alert.

Start with failure classification

The retry decision depends on what failed.

SignalLikely classFirst responseRetry automatically?
429 Too Many RequestsRate limitHonor Retry-After, reduce concurrency or rateUsually, after waiting and within a budget
403 ForbiddenAccess refusal, session, policy, or request mismatchInspect response, session, and approved access pathNot blindly
401 UnauthorizedMissing or expired authenticationRefresh approved credentials or stopOnly after controlled refresh
408, timeout, connection resetNetwork or server timingBack off and retry a limited number of timesOften
500, 502, 503, 504Upstream or server failureBack off; honor Retry-After if providedOften, within a budget
200 with login or challenge pageContent-level failureMark the response invalid; review sessionNot as a successful record
200 with empty dataCould be valid end, changed query, block, or parser issueCompare expected state and response shapeOnly after classification
Invalid JSON or truncated bodyResponse corruption or wrong surfacePreserve evidence and retry selectivelySometimes

The same status can have different causes across sources. Classification should include the status, headers, final URL, content type, response signature, request attempt, session identity, and source-specific expectations.

What a 429 response means

RFC 6585 defines 429 Too Many Requests for rate limiting. The response may include Retry-After, which tells the client how long to wait before sending another request.

On 429:

  1. Stop increasing traffic.
  2. Parse Retry-After when present.
  3. Pause the affected source, credential, session, or worker scope.
  4. Reduce the request rate or concurrency.
  5. Retry only after the wait period and within a bounded attempt budget.
  6. Record the event as a capacity or policy signal.

Do not retry a 429 immediately through ten workers. That turns one rate-limit response into a burst and can extend the restriction.

What a 403 response means

RFC 9110 describes 403 Forbidden as a refusal to fulfill the request. It does not tell you one universal cause.

Possible causes include:

  • The source does not permit the requested access
  • Authentication or a session is missing
  • The session no longer matches the request
  • The URL or method is not available to that client
  • A required first-party flow was skipped
  • A gateway or security layer rejected the request
  • The response is a policy or geographic restriction

Repeatedly changing headers or proxies without understanding the response is not reliable diagnosis. Preserve the response, check the final URL and content, compare it with a known valid request, and review whether the collection method is permitted.

If the source clearly refuses automated access or the approved access path is unavailable, stop that source and escalate it rather than treating refusal as a technical puzzle to defeat.

Temporary server and network failures

Timeouts, connection resets, and some 5xx responses can succeed later. Use exponential backoff with jitter so workers do not all retry at the same instant.

A simple schedule might increase the delay after each attempt, while adding a small random component:

text
delay = min(max_delay, base_delay × 2^attempt) + jitter

The formula is less important than the controls around it:

  • Maximum attempts per task
  • Maximum total elapsed retry time
  • Different budgets by method and status
  • No automatic replay of unsafe side effects
  • Per-source circuit breaking during widespread failure
  • Durable failed-task state after the budget is exhausted

Retries should improve recovery, not erase evidence that the source was unavailable.

Retry scope matters

A failure can belong to different scopes.

ScopeExampleCorrect control
RequestOne connection resetRetry that task
SessionExpired cookie or tokenRefresh or replace that session
IP or routeSource-specific network rejectionQuarantine and review that route
CredentialAPI quota or revoked keyPause that credential; do not rotate blindly
SourceWidespread 503 or changed access policyOpen a circuit for the source
ParserResponses succeed but required fields disappearStop delivery and investigate data quality

Without scope, a system may discard a healthy session because one request timed out, or flood a failing source because each worker sees only its own error.

Use a per-source rate controller

Global concurrency is not enough. Different domains, endpoints, credentials, and query types can have different limits.

A rate controller should consider:

  • Concurrent requests in flight
  • Minimum delay between requests
  • Recent latency
  • 429 frequency
  • 5xx frequency
  • Retry-After
  • Response size and page cost
  • Published source limits or API quota
  • Time-of-day or scheduled batch rules when applicable

Scrapy's AutoThrottle is one example of latency-aware throttling. A custom system may use token buckets, leaky buckets, queues, or worker semaphores. The mechanism should be observable and configurable per source.

Validate the response body before success

Many scraping failures return 200 OK.

Check:

  • Expected content type
  • Final URL and redirect chain
  • Required response keys or page signatures
  • Login, consent, challenge, or error text
  • Minimum and maximum plausible response size
  • Record count and required-field coverage
  • GraphQL errors
  • Whether the first and last record identifiers changed as expected

A browser can also render a valid interface while one background request failed. Conversely, an API request can return partial data with a successful status. Success belongs to the expected data contract, not the transport alone.

Design a retry ledger

Every task should retain enough state to answer:

  • What was attempted?
  • When and from which source/session scope?
  • What response or exception occurred?
  • Was raw evidence stored?
  • How many attempts have run?
  • When is the next attempt eligible?
  • Why was the task marked terminal?

A useful state model is:

text
pending → in_progress → completed
                    ↘ retry_wait → in_progress
                    ↘ failed_terminal
                    ↘ excluded

Do not use one boolean such as done. It cannot distinguish success, exhaustion, source refusal, deliberate exclusion, or a task waiting for later retry.

The crawl-resumption guide explains how this state survives process restarts.

Add a circuit breaker for source-wide failures

When a high proportion of requests for one source fail, continuing the full queue can waste resources and worsen a rate limit.

A circuit breaker can:

  1. Pause new work for the affected source.
  2. Allow current requests to finish or cancel safely.
  3. Wait for a defined cooling period.
  4. Send a small number of probe requests.
  5. Reopen gradually after valid responses return.

The alert should include the source, failure class, affected queue size, first and last occurrence, and a sample response reference. “Scraper failed” is not enough for diagnosis.

Where proxy rotation fits

Proxy rotation is one possible network layer, not the universal response to 403 or 429.

Use it only when:

  • The source and project permit the collection method
  • IP distribution is operationally justified
  • Session affinity is understood
  • Health and cost are measured
  • A failed route can be quarantined

The full proxy rotation guide covers pools, sticky sessions, health scoring, and cost control without turning proxies into an access-control bypass claim.

What to measure

Track failures by source and class:

  • Successful valid responses
  • 403, 429, and 5xx rates
  • Timeout and connection-error rates
  • Retry attempts per completed task
  • Tasks exhausting their retry budget
  • Retry-After duration
  • Session refresh success
  • Response-signature mismatches
  • Cost and latency per valid record
  • Remaining failed or deferred tasks at delivery

These metrics should feed the broader web-scraping pipeline monitoring system.

Common retry mistakes

Retrying every exception

Parser bugs, invalid URLs, source refusal, and missing required state may never improve with time. Classify them instead.

Using the same backoff for every source

Per-source limits and response patterns differ. Make the policy configurable and conservative.

Losing the original failure after a later success

Keep attempt history. Frequent recovery can still reveal an unstable source, route, or session.

Marking a challenge page as a valid document

Validate content signatures and fields before parsing the response as source data.

Hiding exhausted tasks from delivery

Reconcile discovered, completed, failed, excluded, and deferred tasks. A clean CSV can still be incomplete.

Frequently asked questions

Should a scraper retry a 429 response?

Usually yes, but only after honoring Retry-After when present and reducing the load that caused the limit. Use a bounded retry budget and pause the correct source, credential, or session scope.

Should a scraper retry a 403 with a new proxy?

Not automatically. A 403 can represent source policy, missing authorization, invalid state, geographic restriction, or a gateway decision. Diagnose the response and confirm the approved access path before changing network infrastructure.

What is the difference between backoff and throttling?

Throttling controls the normal request rate. Backoff increases the wait after failures. A production system normally needs both: conservative steady-state pacing and slower retries after temporary errors.

How many retries should a scraper use?

There is no universal number. The budget depends on failure class, source cost, task urgency, and whether the request is safe to repeat. Store the attempts and expose tasks that exhaust the budget.

Can a response be blocked even if it returns 200?

Yes. Login pages, challenges, consent pages, empty fallback payloads, and partial application responses can all return 200. Validate expected content and data signals before marking the task successful.

Next step

Use the scalable web-scraping pipeline to place retries inside the full architecture, then add pipeline monitoring so rate limits and exhausted tasks cannot disappear inside an otherwise successful run. For source-specific production work, review Orzaen's Web Scraping & Data Extraction service.

Sources

Tags

Web ScrapingHTTP 429RetriesRate Limiting

Need this fixed?

Have the same system problem?

Share what is manual, messy, broken, or disconnected. We’ll review the cleanest next step.

Get System Review