When long agent runs break but short prompts never do, the thing being killed is the open stream, not the route. Claude Code, Codex CLI, and Cursor Agent push model output as it is generated over SSE or chunked transfer, holding one HTTP connection open for anywhere from three to thirty minutes. Any hop that decides the connection has been idle too long closes it, and you see frozen output, connection reset, or stream disconnected before completion.
This is a different failure from MCP servers that will not connect, where the handshake or tool listing never completes. Here the stream was working and died mid-flight. Base proxy setup lives in Claude Code proxy setup, Codex CLI proxy setup, and Cursor proxy setup and is not repeated below.
The one-sentence check
Before reading further, ask the only question that matters: do short prompts always work? If "explain this file" succeeds every time but a thirty-minute refactor dies at minute twelve, you have a long-connection problem and this page is for you. If even a one-line prompt fails, that is a path problem—go back to the base proxy setup above, because no amount of stream tuning will fix a route that never opened.
Confirm the symptom before touching anything
Three signals together are enough to classify this as a long-connection kill:
- Short requests never fail. "Explain this file" or a one-function edit completes every time.
- Failures land mid-run. Hundreds of lines already streamed, several tool calls deep, then nothing.
- The break has a rhythm. Time three failures from task start. If the seconds cluster near 60, 120, 300, or 600, a timer is firing—this is not packet loss.
Error text varies by tool and tells you little on its own. Claude Code usually surfaces API Error: Connection error or simply stops producing output; Codex CLI prints stream disconnected before completion followed by the request URL; Cursor Agent leaves the last tool call spinning. All three mean the same thing: the stream ended early. Which layer ended it still has to be found.
Every hop that can kill a stream
| Hop | How it kills | Where to look |
|---|---|---|
| The agent itself | Request or stream-idle timeout expires and it gives up |
API_TIMEOUT_MS, stream_idle_timeout_ms
|
| Local network stack | Wi-Fi switch, sleep/wake, another VPN grabbing the route | Does the failure timestamp match a network change? See offline after sleep |
| Clash Verge core | Auto node switch, subscription reload, or config restart drops every live connection | Connection age plus switch entries in the log |
| Protocol layer | Multiplexed session reclaimed; QUIC migration fails after a network change |
smux and transport settings on the node |
| Exit node server | Server-side idle timeout or connection-count reclaim | Compare a second node in the same datacenter |
| Provider relay | Session-table aging; relays that bill per connection recycle idle ones | Only comparable across lines and providers |
| Model API | Long gaps between token batches trip an upstream HTTP timeout | Reproduces on a clean network too |
That last row is real and not your fault. An OpenAI maintainer explained on the Codex tracker that newer models sometimes take unusually long to produce the next batch of tokens, and that pause alone can exceed an HTTP timeout and disconnect the stream. Even a perfectly clean proxy path leaves some baseline failure rate on very long runs, which retries and smaller tasks mitigate rather than fix. Budget for it: design tasks so a single failure does not discard an hour of work.
Isolate the break with the connections view and one curl
Keep the connections view open during a long run and filter by host (anthropic.com, chatgpt.com, or your own gateway). Four columns answer the question:
- Connection age at failure. Round and repeatable across runs means a timer. Different every time points at line quality or the upstream.
- Whether download bytes are still climbing. Bytes frozen while the row survives means data stopped arriving. The row vanishing means something at or below the proxy closed it.
- A new connection appearing. A fresh row to the same host while the old one disappears is a group re-selecting a node. That stream is already dead.
- Which node it used. Three failures across three different nodes means auto-switching is the first thing to disable.
To reproduce without burning API credits, hold a stream open against an endpoint that trickles bytes. This one drips one byte per second for 600 seconds through the local mixed port:
curl -N -v --max-time 700 \
-x http://127.0.0.1:7897 \
"https://httpbin.org/drip?duration=600&numbytes=600&delay=0" \
-o /dev/null
Reading it is mechanical. Surviving the full 600 seconds proves the path can hold a ten-minute stream. Dying at second 60 or 300 across repeated runs hands you the exact timeout value, with no agent involved. Rerun it without -x, or against a different node, and the comparison names the layer. Use the mixed port shown in your settings—see mixed port. Public test endpoints throttle, so read the result as lifetime only, never throughput.
Quick decision tree
Once the curl result is in, the next move is determined:
- Survives 600s with
-x, dies without it — your local router or ISP is killing idle sockets; raise client timeouts and move on. - Dies at the same second every run, regardless of node — a fixed timer in the agent or the upstream; adjust
API_TIMEOUT_MS/stream_idle_timeout_msupward. - Dies at random seconds, different node each time — auto-switching is reaping live connections; pin the exit (next section).
- Dies at the same second only under Clash Verge — a core setting such as keep-alive or a scheduled probe; check the config reload logs.
Auto-switching groups are the top cause
A url-test group re-tests its members on interval and may re-select. The moment it does, live connections through the old node are dropped—the long task was not killed by a bad network, it was killed by your own config. load-balance is worse for this: it spreads new connections across exits, so a reconnect lands somewhere the session does not exist.
Pinning agent domains to one manual exit is the single highest-value change here:
proxy-groups:
- name: AI-Stable
type: select # manual only, never auto-reselects
proxies:
- HK-01
- JP-01
- name: Auto
type: url-test
interval: 600 # keep auto for everything else, but not aggressive
tolerance: 100
lazy: true
proxies:
- HK-01
- JP-01
rules:
- DOMAIN-SUFFIX,anthropic.com,AI-Stable
- DOMAIN-SUFFIX,claude.ai,AI-Stable
- DOMAIN-SUFFIX,chatgpt.com,AI-Stable
- DOMAIN-SUFFIX,openai.com,AI-Stable
- DOMAIN-SUFFIX,cursor.sh,AI-Stable
Two habits go with it: do not refresh the subscription and do not edit config while a long task is running. A reload rebuilds the whole connection table, which is indistinguishable from unplugging the cable. Update first, then run.
On desktop you can also probe more often so intermediate devices do not treat the socket as idle:
# mihomo core, top level of the config file
keep-alive-idle: 15 # seconds idle before probing
keep-alive-interval: 15 # seconds between probes
disable-keep-alive: false
This helps against session-table aging and does nothing against a server enforcing a hard 60-second read timeout, since keepalive probes are not payload. On Android the core force-disables TCP keepalive to save power, so the setting has no effect there.
Node and protocol choices that hurt long streams
-
Be careful with multiplexing. smux and similar layers pack several streams into one underlying connection to save handshakes. When that carrier is reclaimed, every stream on it dies together. Web browsing never notices; a thirty-minute agent run does. Disable mux on the node you pin for agents, or shorten its
keepalive. - QUIC and UDP transports are more fragile across network changes. They support connection migration in principle, but home NATs and corporate networks age UDP sessions far more aggressively than TCP, so moving from Wi-Fi to a hotspot breaks them harder. For long runs a plain TCP transport is usually the calmer choice.
- Relays and multi-hop lines add reclaimers. Every extra hop is another device that may enforce an idle timeout. A direct exit in the same datacenter often beats an "optimized relay" on long tasks even when the relay's latency number looks better.
Tun mode versus system proxy matters here for one reason only: Tun captures at the network layer, so a GUI editor or a newly opened shell cannot silently miss HTTPS_PROXY. That removes a whole class of false diagnoses—see global proxy for AI tools. Tun does not extend any timeout. If runs still die at the same second under Tun, the timer lives further upstream.
Client-side timeouts, and why smaller tasks beat bigger numbers
Tune the tools after the path is clean, not before.
Claude Code reads environment variables and the env block in settings.json. API_TIMEOUT_MS defaults to 600000 ms; the documentation calls out raising it when requests time out on slow networks or through a proxy, with 2147483647 as the ceiling—above that the underlying timer overflows and requests fail immediately.
// ~/.claude/settings.json
{
"env": {
"API_TIMEOUT_MS": "1200000",
"BASH_DEFAULT_TIMEOUT_MS": "300000"
}
}
Codex CLI keeps its three knobs inside a provider block in ~/.codex/config.toml. They are only valid there; older top-level equivalents are ignored.
# ~/.codex/config.toml
[model_providers.openai]
name = "OpenAI"
stream_idle_timeout_ms = 600000 # default 300000 (5 minutes)
stream_max_retries = 10 # default 5, hard cap 100
request_max_retries = 4 # default 4, hard cap 100
Raising stream_max_retries genuinely works as damage control—one user on the Codex tracker reported that a task which had always failed completed once retries were set to 10, typically reconnecting on the seventh or eighth attempt. The cost is real too: every disconnect now burns several reconnect attempts, and total wall-clock time on long jobs grows noticeably. Treat it as a bandage. Recent Codex builds also ship codex doctor, which checks version, config, proxy environment, network reachability, WebSocket transport, and the provider endpoint without changing anything.
The engineering fix outperforms every parameter: split "refactor this module" into five independently verifiable steps that each finish in a few minutes. Failure probability rises with duration, and a run that died with half the files rewritten costs more to clean up than the failure itself. Long Codex sessions carry an extra hazard—automatic context compaction kicks in once a thread is already large, and a failure there can wedge the session until you start a new one. Smaller tasks keep you out of that region entirely.
What does not help
- Retrying the same long task. A repeatable break at the same second means retrying only buys the identical failure again. Measure the second first.
- Chasing lower latency. Round-trip time on one handshake says nothing about whether a connection survives fifteen minutes. Use latency tests to eliminate dead nodes, not to rank stability.
- Resetting config or reinstalling the client. Short tasks working proves the config is fine. These steps fix "cannot connect at all," not "dies halfway."
- Cranking every timeout to the maximum. Longer timeouts tolerate slowness; they cannot stop another party from closing the socket, and they turn fast failures into twenty-minute waits.
- Flipping to global mode mid-run. Mode changes reset connections. Run comparisons with the curl above instead of experimenting on a live task.
Pre-flight checklist for a long run
Run this in order before launching anything that should take more than five minutes:
- Confirm short prompts work—if not, stop here and fix the route.
- Pin AI domains to a single manual exit; disable url-test interval on that path.
- Run the 600-second curl through the same mixed port; it must survive.
- Raise
API_TIMEOUT_MS/stream_idle_timeout_msabove the observed kill second. - Break the task into steps each under a few minutes; keep checkpoints.
Slow and disconnected are separate problems—for throughput see speed optimization, and for requests that never leave the machine see proxy on but sites will not open. Client builds: download center.