Diagnose Timeouts One Hop at a Time
Use bounded probes, cumulative timing fields and controlled route comparisons to narrow a timeout without inventing a cause.
On this page
A timeout tells us that a waiting condition exceeded a limit. It does not identify the component responsible. The missing event could be a DNS answer, a completed connection, a TLS handshake, response headers or another part of a streamed body. Start by naming what the caller was waiting for and which layer imposed the limit.
Keep the first probe small and read-only. Choose one resource that represents a useful path, record the client location and preserve the normal identity checks. Avoid a loop of uncontrolled retries: it can increase load while erasing the evidence that the first attempt supplied.
Write the route and the budget
List the client, any explicit proxy, the public edge and the application hop. Record each relevant timeout and its owner. A browser limit, an edge upstream-read limit and an application deadline can describe different waiting conditions even when all are expressed in seconds.
For a direct HTTPS probe, curl's connection timeout covers the connection phase, including resolution and protocol handshakes. Its total-time limit bounds the overall transfer. The following command is a template for a public read-only resource. It performs one request and preserves certificate verification.
curl --noproxy '*' --silent --show-error \
--connect-timeout 5 --max-time 15 --output /dev/null \
--write-out 'code=%{http_code} dns=%{time_namelookup} connect=%{time_connect} tls=%{time_appconnect} first=%{time_starttransfer} total=%{time_total}\n' \
https://example.com/Replace the URL only with an intended endpoint. The direct-route template bypasses an inherited proxy; where a proxy is required, use that intended route and document it. The five- and fifteen-second values are illustrative budgets, not performance targets for Nedan or universal recommended limits.
Record the command's exit code and error output. An HTTP error response and a transfer timeout are different observations. A numerical status alone does not explain an incomplete transfer, and a zero-looking field should not be turned into a made-up phase duration.
Read cumulative timings carefully
The lookup, connection, handshake and first-response timing fields are measured from the beginning of the operation. They are not independent durations to add together. On a simple fresh direct TCP/TLS connection, differences between successive milestones can help describe where time accumulated.
Those differences require the stated assumptions. Connection reuse, a proxy, redirects, alternative protocols and concurrent address attempts can change the interpretation. Record which protocol was used and avoid labeling the entire first-byte interval application execution time. The interval can include work before the request reaches the application and delay after it responds.
Keep a timing transcript as evidence for one attempt. A slow lookup milestone suggests a resolution question to investigate; it does not prove the authoritative server was slow. A long gap after TLS suggests a later boundary; it does not distinguish an edge queue from an upstream query by itself.
Change one boundary at a time
If resolution is suspect, compare a controlled address override that preserves the hostname, as described in the TLS note. A successful overridden request narrows the question, but the public DNS route still needs its own explanation. Keep both attempts in the notebook.
If the edge-to-application route is suspect, inspect the private listener from the intended host or network. Use the same useful resource when possible. A local health route can be successful while a normal article route fails, so record the exact path and method being compared.
Never open a private application port publicly just to make the comparison easier. That changes the proxy trust boundary. A diagnostic procedure should preserve the route constraints whose behavior it is trying to understand.
Interpret proxy limits by their actual meaning
For nginx, proxy_read_timeout governs the interval between successive upstream reads, rather than an absolute deadline for the complete response. A slowly progressing stream can therefore behave differently from a silent upstream. Increasing that setting does not by itself create a coherent end-to-end request budget.
Review the application's deadline and the caller's total budget alongside the proxy setting. Decide which component should stop work first and what response the caller can observe. A server continuing expensive work after the caller has abandoned the request can turn retries into overlapping workload.
A retry policy needs a separate decision about replay safety. A failed GET probe and an uncertain write do not have the same consequences. Keep retries bounded, use backoff and require an appropriate idempotency design for operations whose repeated execution could change state.
Correlate without collecting everything
Use a bounded request identifier generated or sanitized at the trusted edge when joining restricted logs. Record the time range, route, status and milestone needed for the investigation. Avoid logging authorization fields, session cookies or arbitrary request bodies merely because a verbose mode is available.
Compare timestamps with awareness of clock and logging differences. A client timing record and a server wall-clock event are not interchangeable measurements. If the clocks or boundaries are uncertain, preserve that uncertainty instead of reporting a precise queue delay unsupported by the available evidence.
Finish with the smallest supported finding and the next discriminating probe. For example, the selected public request verified TLS and then timed out, while the permitted local application request succeeded. That narrows the route under examination; it does not establish the root cause without further edge and upstream observations.
References
Tool definitions: curl manual. Proxy timeout semantics: nginx proxy module. The budgets, investigation sequence and examples are editorial methods, not measured production results.
Return to the request-path notebook to attach each result to its boundary.