App Proxy

Why AI Coding Agents Die Mid-Task: Long Streams Killed by the Proxy Path

When long agent runs break but short prompts never do, the thing being killed is the open stream, not the route. Claude Code, Codex CLI, and Cursor Agent push model output as it is generated over SSE or chunked transfer, holding one HTTP connection open for anywhere from three to thirty minutes. Any hop that decides the connection has been idle too long closes it, and you see frozen output, connection reset, or stream disconnected before completion.

This is a different failure from MCP servers that will not connect, where the handshake or tool listing never completes. Here the stream was working and died mid-flight. Base proxy setup lives in Claude Code proxy setup, Codex CLI proxy setup, and Cursor proxy setup and is not repeated below.

The one-sentence check

Before reading further, ask the only question that matters: do short prompts always work? If "explain this file" succeeds every time but a thirty-minute refactor dies at minute twelve, you have a long-connection problem and this page is for you. If even a one-line prompt fails, that is a path problem—go back to the base proxy setup above, because no amount of stream tuning will fix a route that never opened.

Confirm the symptom before touching anything

Three signals together are enough to classify this as a long-connection kill:

  • Short requests never fail. "Explain this file" or a one-function edit completes every time.
  • Failures land mid-run. Hundreds of lines already streamed, several tool calls deep, then nothing.
  • The break has a rhythm. Time three failures from task start. If the seconds cluster near 60, 120, 300, or 600, a timer is firing—this is not packet loss.

Error text varies by tool and tells you little on its own. Claude Code usually surfaces API Error: Connection error or simply stops producing output; Codex CLI prints stream disconnected before completion followed by the request URL; Cursor Agent leaves the last tool call spinning. All three mean the same thing: the stream ended early. Which layer ended it still has to be found.

Every hop that can kill a stream

Hop How it kills Where to look
The agent itself Request or stream-idle timeout expires and it gives up API_TIMEOUT_MS, stream_idle_timeout_ms
Local network stack Wi-Fi switch, sleep/wake, another VPN grabbing the route Does the failure timestamp match a network change? See offline after sleep
Clash Verge core Auto node switch, subscription reload, or config restart drops every live connection Connection age plus switch entries in the log
Protocol layer Multiplexed session reclaimed; QUIC migration fails after a network change smux and transport settings on the node
Exit node server Server-side idle timeout or connection-count reclaim Compare a second node in the same datacenter
Provider relay Session-table aging; relays that bill per connection recycle idle ones Only comparable across lines and providers
Model API Long gaps between token batches trip an upstream HTTP timeout Reproduces on a clean network too

That last row is real and not your fault. An OpenAI maintainer explained on the Codex tracker that newer models sometimes take unusually long to produce the next batch of tokens, and that pause alone can exceed an HTTP timeout and disconnect the stream. Even a perfectly clean proxy path leaves some baseline failure rate on very long runs, which retries and smaller tasks mitigate rather than fix. Budget for it: design tasks so a single failure does not discard an hour of work.

Isolate the break with the connections view and one curl

Keep the connections view open during a long run and filter by host (anthropic.com, chatgpt.com, or your own gateway). Four columns answer the question:

  1. Connection age at failure. Round and repeatable across runs means a timer. Different every time points at line quality or the upstream.
  2. Whether download bytes are still climbing. Bytes frozen while the row survives means data stopped arriving. The row vanishing means something at or below the proxy closed it.
  3. A new connection appearing. A fresh row to the same host while the old one disappears is a group re-selecting a node. That stream is already dead.
  4. Which node it used. Three failures across three different nodes means auto-switching is the first thing to disable.

To reproduce without burning API credits, hold a stream open against an endpoint that trickles bytes. This one drips one byte per second for 600 seconds through the local mixed port:

curl -N -v --max-time 700 \
  -x http://127.0.0.1:7897 \
  "https://httpbin.org/drip?duration=600&numbytes=600&delay=0" \
  -o /dev/null

Reading it is mechanical. Surviving the full 600 seconds proves the path can hold a ten-minute stream. Dying at second 60 or 300 across repeated runs hands you the exact timeout value, with no agent involved. Rerun it without -x, or against a different node, and the comparison names the layer. Use the mixed port shown in your settings—see mixed port. Public test endpoints throttle, so read the result as lifetime only, never throughput.

Quick decision tree

Once the curl result is in, the next move is determined:

  • Survives 600s with -x, dies without it — your local router or ISP is killing idle sockets; raise client timeouts and move on.
  • Dies at the same second every run, regardless of node — a fixed timer in the agent or the upstream; adjust API_TIMEOUT_MS / stream_idle_timeout_ms upward.
  • Dies at random seconds, different node each time — auto-switching is reaping live connections; pin the exit (next section).
  • Dies at the same second only under Clash Verge — a core setting such as keep-alive or a scheduled probe; check the config reload logs.

Auto-switching groups are the top cause

A url-test group re-tests its members on interval and may re-select. The moment it does, live connections through the old node are dropped—the long task was not killed by a bad network, it was killed by your own config. load-balance is worse for this: it spreads new connections across exits, so a reconnect lands somewhere the session does not exist.

Pinning agent domains to one manual exit is the single highest-value change here:

proxy-groups:
  - name: AI-Stable
    type: select          # manual only, never auto-reselects
    proxies:
      - HK-01
      - JP-01

  - name: Auto
    type: url-test
    interval: 600         # keep auto for everything else, but not aggressive
    tolerance: 100
    lazy: true
    proxies:
      - HK-01
      - JP-01

rules:
  - DOMAIN-SUFFIX,anthropic.com,AI-Stable
  - DOMAIN-SUFFIX,claude.ai,AI-Stable
  - DOMAIN-SUFFIX,chatgpt.com,AI-Stable
  - DOMAIN-SUFFIX,openai.com,AI-Stable
  - DOMAIN-SUFFIX,cursor.sh,AI-Stable

Two habits go with it: do not refresh the subscription and do not edit config while a long task is running. A reload rebuilds the whole connection table, which is indistinguishable from unplugging the cable. Update first, then run.

On desktop you can also probe more often so intermediate devices do not treat the socket as idle:

# mihomo core, top level of the config file
keep-alive-idle: 15       # seconds idle before probing
keep-alive-interval: 15   # seconds between probes
disable-keep-alive: false

This helps against session-table aging and does nothing against a server enforcing a hard 60-second read timeout, since keepalive probes are not payload. On Android the core force-disables TCP keepalive to save power, so the setting has no effect there.

Node and protocol choices that hurt long streams

  • Be careful with multiplexing. smux and similar layers pack several streams into one underlying connection to save handshakes. When that carrier is reclaimed, every stream on it dies together. Web browsing never notices; a thirty-minute agent run does. Disable mux on the node you pin for agents, or shorten its keepalive.
  • QUIC and UDP transports are more fragile across network changes. They support connection migration in principle, but home NATs and corporate networks age UDP sessions far more aggressively than TCP, so moving from Wi-Fi to a hotspot breaks them harder. For long runs a plain TCP transport is usually the calmer choice.
  • Relays and multi-hop lines add reclaimers. Every extra hop is another device that may enforce an idle timeout. A direct exit in the same datacenter often beats an "optimized relay" on long tasks even when the relay's latency number looks better.

Tun mode versus system proxy matters here for one reason only: Tun captures at the network layer, so a GUI editor or a newly opened shell cannot silently miss HTTPS_PROXY. That removes a whole class of false diagnoses—see global proxy for AI tools. Tun does not extend any timeout. If runs still die at the same second under Tun, the timer lives further upstream.

Client-side timeouts, and why smaller tasks beat bigger numbers

Tune the tools after the path is clean, not before.

Claude Code reads environment variables and the env block in settings.json. API_TIMEOUT_MS defaults to 600000 ms; the documentation calls out raising it when requests time out on slow networks or through a proxy, with 2147483647 as the ceiling—above that the underlying timer overflows and requests fail immediately.

// ~/.claude/settings.json
{
  "env": {
    "API_TIMEOUT_MS": "1200000",
    "BASH_DEFAULT_TIMEOUT_MS": "300000"
  }
}

Codex CLI keeps its three knobs inside a provider block in ~/.codex/config.toml. They are only valid there; older top-level equivalents are ignored.

# ~/.codex/config.toml
[model_providers.openai]
name = "OpenAI"
stream_idle_timeout_ms = 600000   # default 300000 (5 minutes)
stream_max_retries = 10           # default 5, hard cap 100
request_max_retries = 4           # default 4, hard cap 100

Raising stream_max_retries genuinely works as damage control—one user on the Codex tracker reported that a task which had always failed completed once retries were set to 10, typically reconnecting on the seventh or eighth attempt. The cost is real too: every disconnect now burns several reconnect attempts, and total wall-clock time on long jobs grows noticeably. Treat it as a bandage. Recent Codex builds also ship codex doctor, which checks version, config, proxy environment, network reachability, WebSocket transport, and the provider endpoint without changing anything.

The engineering fix outperforms every parameter: split "refactor this module" into five independently verifiable steps that each finish in a few minutes. Failure probability rises with duration, and a run that died with half the files rewritten costs more to clean up than the failure itself. Long Codex sessions carry an extra hazard—automatic context compaction kicks in once a thread is already large, and a failure there can wedge the session until you start a new one. Smaller tasks keep you out of that region entirely.

What does not help

  • Retrying the same long task. A repeatable break at the same second means retrying only buys the identical failure again. Measure the second first.
  • Chasing lower latency. Round-trip time on one handshake says nothing about whether a connection survives fifteen minutes. Use latency tests to eliminate dead nodes, not to rank stability.
  • Resetting config or reinstalling the client. Short tasks working proves the config is fine. These steps fix "cannot connect at all," not "dies halfway."
  • Cranking every timeout to the maximum. Longer timeouts tolerate slowness; they cannot stop another party from closing the socket, and they turn fast failures into twenty-minute waits.
  • Flipping to global mode mid-run. Mode changes reset connections. Run comparisons with the curl above instead of experimenting on a live task.

Pre-flight checklist for a long run

Run this in order before launching anything that should take more than five minutes:

  1. Confirm short prompts work—if not, stop here and fix the route.
  2. Pin AI domains to a single manual exit; disable url-test interval on that path.
  3. Run the 600-second curl through the same mixed port; it must survive.
  4. Raise API_TIMEOUT_MS / stream_idle_timeout_ms above the observed kill second.
  5. Break the task into steps each under a few minutes; keep checkpoints.

Slow and disconnected are separate problems—for throughput see speed optimization, and for requests that never leave the machine see proxy on but sites will not open. Client builds: download center.

Long runs need egress that survives long connections

A streaming response holds one HTTP connection open for minutes. Install Clash Verge Rev from the download center, pin AI domains to a fixed node, then rerun the task.