"If you see a CAPTCHA or similar test, just wait for it to get solved automatically." That works, but the model decides for itself how long to wait and how to tell a solve is still running, from a screenshot alone.
Captcha telemetry gives you the signal directly. This cookbook wires it into a @onkernel/browser-loop agent so that browser actions are held while a solve is outstanding, and the agent is told the outcome in terms it can act on.
The design follows one split:
Telemetry decides what to say. The live page decides whether to interrupt the agent at all. A solver task succeeding and a challenge clearing are different facts, so the gate never reports one as the other.
What the events tell you
Three rules from Correlate captcha tasks and challenges shape everything below:
task_idis the only join. Pair a start with its result on it.challenge_idgroups tasks for one visible challenge and is present only when Kernel tracked the widget — tasks without one can never be joined to acaptcha_challenge_result, even when one is emitted for the same page.- Delivery is best-effort and unordered. A start can arrive after its result, and any event can be absent. Nothing may depend on arrival order, and every wait needs a deadline.
- Fall back to the page. When you need a challenge-level outcome and don’t have one, use the available task results and the current page state.
Setup
KERNEL_API_KEY and a provider key for whichever model LOOP_MODEL points at (anthropic:claude-sonnet-5 by default, so ANTHROPIC_API_KEY):
Build the gate
The four snippets below are one file,captcha-gate.ts, in order.
1
Read the events
Three maps, one per thing the telemetry can tell you, all keyed so nothing depends on arrival order.
joinable is the important one: it holds only the challenge_ids that actually appeared on a task event, which are the only ones a challenge result may be attributed to.A task record created by a result already has a status, so a captcha_solve_started that arrives afterwards finds a closed task and leaves it closed.captcha-gate.ts
2
Decide when to hold
A task counts as open while it has no terminal status and is still inside its deadline — that deadline is what stops a missing
captcha_solve_result from holding the agent forever. until is the only waiting primitive, and it always takes a timeout.captcha-gate.ts
3
Resolve an outcome
resolve settles what it can, waits a bounded interval for a challenge-level result only when a task actually carried a challenge_id, then reads the page and picks the best available source: an attributable challenge result, then any challenge result (labelled as a page observation rather than a join), then task results, then the page alone.holding and pending are separate on purpose. pending covers a terminal outcome that lands while no action is in flight — without it, a challenge result arriving between tool calls is recorded and never told to anyone.captcha-gate.ts
4
Wire it to the agent
harness.on("tool_call", …) is awaited before the tool is dispatched, so returning from it late holds the action and returning { block: true, reason } replaces it with a message the model reads. That is cheaper than aborting the turn and it stops the action before it reaches the page.The page probe is a read-only Playwright snippet. Its result is what decides whether the agent is interrupted at all — a cleared page just resumes with no extra turn.captcha-gate.ts
The complete captcha-gate.ts
The complete captcha-gate.ts
captcha-gate.ts
What the agent is told
Every verdict pairs a telemetry claim with the page state it was checked against:
Durations come straight from each event’s
duration_ms, which is authoritative; don’t compute them from event timestamps.
Limits
TASK_SETTLE_MSandCHALLENGE_GRACE_MSare the safety net. Events can be absent, so both waits are bounded and the gate falls through to the page rather than stalling. RaiseTASK_SETTLE_MSif your solves routinely run longer than 20s.- Challenge results aren’t emitted for every widget type. Turnstile, for instance, reports task events only, so the task-plus-page path is the one that runs there.
- One episode at a time. The gate resets after each verdict; overlapping visible challenges from the same provider aren’t split apart, and the telemetry docs note Kernel’s own event model can’t always attribute a result in that case either.
- The page probe is best-effort. It reads the DOM for known widget markers and a response token. A site that renders its challenge somewhere those selectors miss will fall back to the telemetry verdict alone.
Next steps
- Telemetry Categories — the full captcha event schema and correlation rules
- Stream Telemetry — resuming a dropped stream, filtering by category
- Stealth mode — what the automatic captcha solver covers
- Playwright with Computer Use Fallback — holding and redirecting an agent mid-run for a different reason