Gigflow docs
Engineering

ADR-0016: Edge browser retries for blocked pages

Retry a page that a bot defense blocks on an edge device, and never again from the data-centre browsers.

Status: Proposed.

Date: 2026-09-29

Context

Browser workers run Chrome in the in-cluster Browserless fleet. Since #297 it is native headful Chrome with a consistent fingerprint. Its traffic leaves from the cluster's data-centre IP address. In the production run of the browser stack E2E, the Cloudflare challenge page stopped that address in 2 of 2 runs. The same browser passed the page from a home connection.

Edge devices (the browser-edge package) run Browserless on home connections and register with the browserless gateway through frps. A session with pool=edge goes to an edge with headroom, and otherwise to the in-cluster fleet. The gateway sends a session only to an edge on the fleet's Browserless version.

Edge capacity is small: about 1.25 GiB of memory per session, so a laptop gives about four sessions. Commercial residential proxies cost $2 to $12 per GB. The planned volume, 10 sessions at a time around the clock, is about 0.5 to 1.8 TB per month. That makes such proxies too expensive for all traffic.

Configuration discovery detects block pages with getBlockReason in @gigflow/harvest: a blocked HTTP status on a page without links, or a known challenge marker from Cloudflare, DataDome, Imperva, or HUMAN. It records the outcome blocked. The harvest workers do not detect blocks. vacancy-harvest sets pool=edge for all its sessions.

No popular, maintained open-source library identifies every vendor's block pages. is-antibot knows 38 vendors, but has one maintainer and about 5,000 downloads per month. It missed an Akamai "Access Denied" page and three Cloudflare block pages that site owners configure. Crawlee is popular, but detects only a Cloudflare challenge frame, Incapsula, and status codes. Vendor signatures change, and Engine cannot maintain them.

A test on 2026-09-29 with plain requests found 8 block responses from Cloudflare (challenge and block pages), Akamai, DataDome, and one unknown vendor. All 8 had HTTP status 403. All normal pages, including pages behind Cloudflare and Akamai, had HTTP status 200.

On 2026-09-29, production had 68 discovery runs, 8 of them blocked (12%). All of them ran before block detection and before native Chrome. Harvest records no block data. The blocked share is not known yet.

ADR-0006 applies: each worker owns one queue, and payload schemas live in @gigflow/engine-jobs.

Decision

  1. Data-centre browsers first. A browser job runs on the in-cluster fleet unless its payload or a stored route asks for an edge.

  2. Strict edge pool. The gateway accepts pool=edge-only. It dials only edges that are healthy, on the fleet's version, and with headroom. When no such edge exists, it refuses the session with HTTP 503 and counts browserless_gateway_rejected_total{reason="no_edge"}. It never falls back to the in-cluster fleet. pool=edge keeps its current meaning.

  3. Pool per session. @gigflow/browser takes a pool option for each session and adds it to the connection query. BROWSERLESS_QUERY stays the worker default.

  4. Requeue on a block. When a job on the in-cluster fleet ends as blocked, its processor submits the same job again to its own queue, with browserPool: "edge" in the payload. Such a job connects with pool=edge-only. The field is optional in each payload schema that allows it. Bizzy payloads do not allow it.

  5. Wait for edge capacity. When the gateway refuses an edge-only session, the processor defers the job with deferJob from @gigflow/worker. A deferral does not use an attempt. The delay grows from 1 to 10 minutes. After 24 hours without an edge session, the job ends as blocked with the reason edge_unavailable.

  6. No loop. A job that is blocked on an edge ends as blocked, with both attempts recorded. It is not submitted again.

  7. Remember the route. Engine stores each host that needed an edge, with an expiry of 30 days. Discovery and harvest jobs for a stored host start on an edge and skip the data-centre attempt. After the expiry, the host returns to the data-centre browsers.

  8. Vendor-independent block detection. A page is blocked when at least one of these signals is present:

    • an HTTP status in BLOCKED_STATUSES of @gigflow/browser;
    • the header cf-mitigated: challenge, which Cloudflare documents for all its challenge pages;
    • none of the content that the job expects, such as job or career links in discovery and items in a harvest;
    • a generic phrase in the title or text, such as "access denied", "security check", "verify you are human", or "unusual traffic".

    Discovery, configuration-harvest, and vacancy-harvest use the same check. Engine adds no vendor rules. For reports, a block records the server header and the names of known vendor headers such as cf-ray, akamai-grn, and x-datadome. @gigflow/browser returns the headers of the main document response for this.

  9. Reserve the edges. vacancy-harvest stops setting pool=edge for all sessions. Edge capacity serves blocked work only.

Consequences

  • Edge devices serve only the work that the data-centre address cannot do. One edge device is enough while the blocked work stays below its session count.
  • A blocked job takes longer: one blocked data-centre attempt, then a wait for edge capacity. A stored route removes the blocked attempt.
  • With no edge online, blocked work waits up to 24 hours and then ends as blocked. It never goes back to the data-centre address.
  • New signals: refused edge-only sessions, deferred edge jobs (jobs_deferred_total), and blocked pages per worker and pool. An alert fires when edge jobs are deferred for more than one hour while no edge is online.
  • Discovery and harvest payload schemas change without breaking: browserPool is optional.
  • Bizzy account-connect keeps using the in-cluster fleet only.

Implementation steps

Each step adds an E2E test in the owning package and a verification report.

  1. Vendor-independent block detection and metrics in discovery and both harvest workers, without a routing change. Measure the blocked share for one week.
  2. pool=edge-only in the gateway and the pool session option in @gigflow/browser.
  3. Requeue, deferral, and the loop guard in configuration discovery.
  4. Stored host routes, and requeue in both harvest workers.
  5. vacancy-harvest back to the in-cluster fleet by default, and the alert.

Alternatives

  • Commercial residential proxies for all traffic. About $1,000 to $7,000 per month at the planned volume.
  • ISP (static residential) proxies. Priced per IP with traffic included: about $20 to $120 per month for 10 to 20 addresses. Not tested yet. They can add capacity for blocked work later through the existing PROXY_SERVER support.
  • Vendor signature libraries (is-antibot, Crawlee). They name the vendor, but miss common block pages, and their signatures need updates that Engine cannot maintain. The routing decision does not need the vendor.
  • A paid unblocking service (for example Zyte, Scrapfly, or Bright Data Web Unlocker) for pages that stay blocked after the edge retry. The provider maintains the detection, and Engine pays per request. This can be added later for that remainder.
  • A separate edge queue per worker. This conflicts with ADR-0006. A deferral gives the same waiting behavior in the worker's own queue.
  • A laptop as a proxy for the in-cluster browsers. All page traffic would cross the home connection twice, and every session would share one household address. Edge devices keep the browser, its address, and its traffic on the same machine.

On this page