Module I · Loading
The browser as operating system
The standard introduction to a browser is a diagram with an arrow: URL in one side, rendered page out the other. We will spend nine modules taking that arrow apart. But the arrow is the wrong shape to start with, because it suggests a pipe — bytes flowing through a fixed channel. The browser is not a pipe. It is a small operating system, and the page is a program it agrees to run.
1.1 Why “operating system” and not “renderer”
Call the browser a renderer and you have described one of its subsystems while hiding the rest. Rendering — turning markup and styles into pixels — is genuinely hard and gets three modules of its own. But it is not the part that makes the browser interesting as a piece of systems software. The interesting part is everything around the rendering: the browser accepts a program written by someone it has never met, transmitted over a network it does not control, and runs that program on your machine, with access to your screen, your input, and a curated slice of your hardware — without letting it read your filesystem, your other tabs, or your bank session.
That sentence is a description of an operating system. An OS runs untrusted programs and keeps them from harming the machine or each other. Substitute “web page” for “program” and “origin” for “process” and you have the browser's job description. Once you see it this way, a long list of browser behaviors that otherwise look like trivia — why scripts block rendering, why a slow tab can stutter an animation in another, why a page cannot read a file you did not hand it, why the camera asks permission — stop being trivia and become the predictable consequences of a few systems-design decisions.
A web page is untrusted code from a stranger. Everything the browser does is shaped by the problem of running it safely and fast.
1.2 The subsystems
A browser is conventionally decomposed into a handful of major components. The decomposition below has been stable since the mid-2000s — the famous “How Browsers Work” primer used it — even as the implementations underneath changed completely. The labels are worth memorizing because every later module lives inside one of these boxes.
- The user interface (the browser shell, or “chrome”)
- Everything that is not the page: the address bar, the back and forward buttons, tabs, bookmarks, the menus. None of it is defined by any web specification; it is pure convention, evolved by browsers imitating each other. The word “chrome” meant this long before it was a product name.
- The browser engine
- The coordinator between the shell and the rendering engine. It marshals actions: you click back, the shell tells the browser engine, the browser engine drives the rendering engine to a previous state. In a modern multiprocess browser this is the privileged core — hold that thought, it becomes the kernel of our analogy.
- The rendering engine
- The subsystem that turns a document into a picture: parse HTML into a tree, parse CSS, compute styles, lay out boxes, paint pixels. Blink (Chrome, Edge, Opera, and every other Chromium browser), WebKit (Safari), and Gecko (Firefox) are the three that matter in 2026. Modules 3 through 5 are entirely about this box.
- Networking
- Fetches bytes. DNS, connections, TLS, HTTP, caching policy. A platform-independent interface over wildly platform-specific guts. This is the I/O subsystem; Module 2 is its tour.
- The JavaScript engine
- Parses and executes JavaScript — V8 in Chrome, JavaScriptCore in Safari, SpiderMonkey in Firefox. Distinct from the rendering engine it sits next to, though they are deeply entangled, because the script can rewrite the document mid-render. Module 6.
- The UI backend
- Draws the primitive widgets — checkboxes, drop-downs, the window itself — through a generic interface implemented with the host OS's own drawing calls. This is the seam where the browser touches the real operating system underneath it.
- Data storage
- The persistence layer: cookies, Web Storage, IndexedDB, the HTTP cache, the origin-private filesystem. The browser is a small database engine you never sign up for. Module 8.
Seven boxes. The rest of this module makes a single claim about them: arrange them under the pressure of running untrusted code at sixty frames a second, and what you get is not a media player. It is an operating system.
1.3 The mapping
Here is the analogy stated literally, term by term. Read the right column as the thing you already understand from an operating systems course, and the left as the browser's version of it.
| Browser | Operating system |
|---|---|
| The browser process / engine | The kernel and its privileged services — the trusted code that everything else must ask for resources. |
| A renderer process | A userland process: unprivileged, sandboxed, unable to touch hardware directly. |
| A tab | A running application. |
| A page's DOM + JavaScript | A program and its address space. |
| The event loop | The scheduler — decides what runs next on a thread that can only do one thing at a time. |
| The origin (scheme + host + port) | The process / permission boundary — the unit of isolation and access control. |
Web APIs (fetch, geolocation, WebGPU, clipboard) | The syscall interface — the controlled gate through which a program reaches the outside world. |
| The networking layer | The I/O subsystem and device drivers. |
| Storage (cookies, IndexedDB, cache) | The filesystem. |
| The sandbox | Memory protection and privilege separation — the wall between untrusted code and everything valuable. |
| The permission prompt | Access control — the user as the authority granting a capability. |
The rest of the series is, in effect, a walk down this table. Module 2 is I/O. Modules 3 through 5 are the program being loaded into memory and prepared to display. Module 6 is the scheduler. Module 7 is the process boundary and access control. Module 8 is the filesystem. Module 9 is the syscall surface, expanding year over year. The single most useful thing this module can give you is the habit of reaching for the right column whenever the left column surprises you.
1.4 From one process to many
The analogy was not always this clean. Early browsers were single-process programs: the shell, the rendering of every tab, and every page's JavaScript all ran in one address space on, effectively, one thread of control. This had the property every systems programmer will wince at — no isolation. A bug on one page could corrupt another. A script that spun forever froze the whole browser. A crash anywhere took down every tab you had open. In OS terms, it was a machine with no memory protection, running every application in ring zero.
- 1990s–2000s
- Single-process browsers. One tab's runaway script or fatal bug crashes or hangs them all. Isolation between pages is, at best, advisory.
- 2008
- Chrome ships with a multiprocess architecture: a privileged browser process plus separate, sandboxed renderer processes. A crashing page now takes down one tab, not the session — the “Aw, Snap” page instead of a dead browser. This is the move from a single-tasking monitor to a real multitasking OS.
- 2017
- Firefox completes its multiprocess rollout (project “Electrolysis”), splitting chrome from content after years as a single-process design.
- 2018
- The Spectre and Meltdown speculative-execution attacks change the threat model overnight. They show that code in one process can read memory it was never granted, by timing the CPU's guesses. Chrome responds by turning on Site Isolation by default: each site gets its own renderer process, so a cross-origin secret is not even in the address space the attacker's script can probe.
- 2021
- Firefox ships its equivalent, project “Fission.” Per-site process isolation is now table stakes across the major engines.
Notice the shape of that history. It is the same arc an operating system traveled decades earlier: from a single program in charge of the machine, to cooperative multitasking, to preemptive multitasking with hardware-enforced memory protection between mutually distrustful processes. The browser walked the road again, compressed into about a decade, under the same pressure that drove the original: you cannot let one untrusted program read or wreck another's memory.
1.5 The kernel and its userland
A modern Chromium browser does not run two kinds of process — it runs several, and the division of labor is the analogy made concrete:
- The browser process is the privileged core. It owns the window and the shell, it talks to the real OS, and it is the only process trusted to touch the network and the disk directly. It is the kernel.
- Renderer processes run the actual web content: the rendering engine and the JavaScript engine for a given site. They are sandboxed — stripped of OS privileges — and cannot open a socket or read a file themselves. They are userland.
- The GPU process mediates access to graphics hardware so that a compromised renderer cannot drive the GPU driver directly. A device driver behind a guard.
- The network service and other utility processes isolate specific risky jobs — parsing untrusted data formats, handling the network stack — away from both the kernel and the content.
The crucial detail for a programmer: a renderer process cannot do anything dangerous on its own. When the page calls fetch(), the renderer does not open a connection. It sends a message — an inter-process call — to the browser process, which performs the network operation on its behalf and ships the bytes back. This is exactly a syscall: an unprivileged process asking the privileged kernel to do the thing it is not allowed to do itself, through a narrow, checkable interface. Every capability the web platform exposes is a door of this kind, and Module 7 is about who is allowed through which door.
1.6 The sandbox is the whole point
It is worth dwelling on the sandbox, because it is the single feature that most justifies the operating-system framing and the one most invisible to authors. A renderer process runs with its operating-system privileges deliberately reduced to almost nothing: on Linux through seccomp-bpf syscall filtering, on Windows through restricted tokens and job objects, on macOS through the Seatbelt sandbox. The renderer can compute, and it can talk to the browser process over IPC. It cannot, by itself, read your home directory, open a network socket, start a program, or look at another tab's memory.
Think about what this buys you. You navigate to a page. The page ships you JavaScript. That JavaScript is, from your computer's point of view, an arbitrary program written by a stranger, now executing on your hardware. On any other software platform, “run an arbitrary program from the internet” is the definition of catastrophe. On the web it is the normal case, survived billions of times a day, because the program runs inside a sandbox inside an unprivileged process behind a syscall gate. The browser made “execute untrusted code on sight” routine. That is an operating-system achievement, not a rendering one.
The sandbox is why the web could become a platform. Without it, every link would be a download you ran without asking.
1.7 Every user agent is a different machine
The phrase “the browser” is a convenient lie. There is no the browser. There is a population of user agents, each a different machine: a different rendering engine with different bugs, a different process and memory budget, a different network underneath, a different set of capabilities granted or denied, a different threat posture. The cast from The Web for Programmers is the human-readable index of that variation, and it maps onto the systems concepts in this module precisely.
1.8 The lab: meet Demo Company
From here on the series carries a running example so that none of this stays abstract. democompany.com is a deliberately tiny site whose every page is engineered to add one new idea and to print its own list of HTTP requests directly in the content. Its home page makes the smallest possible point: a document and a favicon.
/index.html
/favicon.icoTwo requests. One is the program; the other is a resource the shell asks for on its own, without being told, because browsers fetch a site's icon by convention. Already the operating-system framing earns its keep: that second request did not come from anything the author wrote. It came from the user agent acting on its own policy. Throughout this series you will keep meeting requests, behaviors, and decisions that no line of HTML asked for — they are the OS doing its job around your program.
Open the site. Open DevTools to the Network panel. Reload. Watch two requests appear, and notice that the page tells you to expect exactly those two. Over the coming modules each new Demo Company page adds one resource — an image, a stylesheet, a font, a script, a frame, a remote and then a hostile resource — and the waterfall grows in a way you will be able to predict and explain rather than merely observe.
1.9 What you now have
One mental model, stated as a table and defended with history: the browser is the operating system of the web, the page is an untrusted program, and every subsequent module is a tour of one OS subsystem. You have the component decomposition — shell, browser engine, rendering engine, networking, JavaScript engine, UI backend, storage — and the reason those components are arranged the way they are: the relentless requirement to run a stranger's code safely and at speed. You have the multiprocess architecture and the reason it exists, the sandbox and why it is the platform's foundation, and the cast as a reminder that “the browser” is a population, not a thing.
Next we follow the first subsystem in motion. You press Enter on a URL. Before a single byte of HTML exists in memory, the browser parses the address, decides what kind of thing it points at, finds the server, negotiates a secure connection, and asks for the document. Module 2 is the I/O subsystem from address bar to first byte.
Stop picturing a pipe. Picture a kernel that just agreed to run a program it has never seen.
Module II · Loading
From URL to first byte
Module 1 ended with a renderer process that cannot open a network socket, handing a request up to the privileged browser process to perform. This module is what the browser process does with it. You press Enter; some milliseconds later the first byte of an HTML document lands in memory. Between those two events is the entire I/O subsystem of the web — address parsing, name resolution, connection setup, encryption, and the request itself. None of it is visible in your HTML, and all of it is happening on your behalf.
2.1 What you type is not what you mean
Start before the URL exists. You type into the address bar — the omnibox — which is not a URL field. It is a guesser. Type democompany.com and it must decide: is this an address to navigate to, or words to search for? It applies heuristics — does it look like a hostname, is there a dot, is it a known site — and either constructs a URL or hands your text to your default search engine as a query. The same keystrokes can become a navigation or a search depending on a judgment the browser makes before any request leaves your machine.
Once it commits to navigation, the browser normalizes what you gave it into a real URL: it supplies a scheme if you omitted one (increasingly https by default), lowercases the host, encodes illegal characters, and resolves the result against the URL standard. Only then does the loading machinery have something to work with.
2.2 Parsing, schemes, and DNS — covered elsewhere
2.5 Opening the connection
With an IP address in hand, the browser establishes a connection — and on the modern web that means a secure connection, because plain http is now the exception and browsers actively upgrade to and prefer https. Setting one up is a sequence of round trips, and round trips are the currency of web latency: each one is a full there-and-back across the physical distance to the server, and you cannot go faster than the speed of light lets you.
Three things deserve emphasis. First, the TLS handshake is where the browser authenticates the server — it checks that the certificate is valid, unexpired, and issued for this host by a certificate authority it trusts. That check is what the padlock means, and it is the difference between talking to Demo Company and talking to someone impersonating it. Second, connections are expensive to build and so are reused: HTTP/2 multiplexes many requests over a single connection, which is why Module 4's flood of subresources does not pay this cost again per file. Third, HTTP/3 changes the picture by running over QUIC on top of UDP, folding the transport and cryptographic handshakes together and eliminating a class of stalling that plagued earlier versions. The diagram is the worst case; the modern web works hard to do better.
- 1997
- HTTP/1.1 standardized: persistent connections, but one request at a time per connection, in order.
- 2015
- HTTP/2: many requests multiplexed over one connection, headers compressed. The per-file connection tax largely disappears.
- 2022
- HTTP/3 standardized over QUIC/UDP: handshakes merged, head-of-line blocking at the transport removed, faster recovery on lossy networks — the mobile and distant user case.
2.6 The request
Now the browser sends the request. For Demo Company's home page it is, in essence, this — a request line, then headers, then (for a navigation) an empty body:
GET /index.html HTTP/2
Host: www.democompany.com
User-Agent: Mozilla/5.0 (…)
Accept: text/html,application/xhtml+xml,…
Accept-Language: en-US,en;q=0.9
Accept-Encoding: gzip, br, zstd
Cookie: (whatever this origin has stored)Almost none of that came from your HTML. The author wrote a link with an href; the browser supplied the method, the protocol version, the Host, an identifying User-Agent, the content types and languages it is willing to accept, the compression schemes it understands, and any cookies this origin previously set. This is the user agent acting as an agent — representing you and your preferences to the server, including preferences you never consciously expressed. The Accept-Language header is how a site can greet the distant user in their own language without asking; the Cookie header is how it recognizes a returning visitor, and the subject of Module 8.
2.7 The response, and the first byte
The server answers with a status line, its own headers, and the body — the bytes we have been waiting for:
HTTP/2 200 OK
Content-Type: text/html; charset=utf-8
Content-Length: 1024
Cache-Control: max-age=3600
…
<!doctype html>
<html lang="en"> … the document begins here …The status code is the server's verdict, and the browser branches on it. 200 means here is your document. A 3xx is a redirect — the browser transparently issues a fresh request to the new location, which is why one navigation can become two or three round trips before any content arrives, and why redirect chains are a real performance cost. 404 means the server has no such resource (you will meet it deliberately in Module 7). 5xx means the server failed. The browser handles each without bothering the user with the distinction unless it must.
One header outranks the others for what happens next: Content-Type. It tells the browser what kind of thing the body is, and therefore which subsystem to hand it to. text/html goes to the HTML parser of Module 3. text/css, image/webp, application/javascript each route elsewhere. The browser will sniff the bytes if the type is missing or implausible, which is a security hazard — a server can send X-Content-Type-Options: nosniff to forbid the guessing. The type, not the file extension, is the authority. A file called data.txt served as text/html is parsed as HTML.
The browser does not trust the URL to say what a resource is. It trusts the server's declared content type, and then only cautiously.
2.8 The cheapest request is the one you skip
Before any of the above runs, the browser asks a prior question: do I already have this? The HTTP cache (Module 8) can satisfy a request with zero network at all, or with a cheap conditional request — “send the body only if it changed since the version I have,” to which the server can answer 304 Not Modified with no body. The entire round-trip apparatus of this module is something the browser spends real effort to avoid. That is the first lesson of web performance and the reason Module 1 asked you to toggle “Disable cache” and watch the waterfall change: you were turning the I/O subsystem's most important optimization off and on.
2.9 The lab: Demo Company's minimal request set
Everything in this module happens, in miniature, when you load Demo Company's home page. The page tells you to expect exactly two requests:
/index.html ← the navigation: scheme, host, path, the whole handshake
/favicon.ico ← the browser asks for this on its own, by conventionTrace the first one against this module: the omnibox decided it was a navigation, the URL parsed into scheme and host and path, DNS resolved www.democompany.com, a secure connection opened across a few round trips, a GET /index.html went out carrying headers you never wrote, and a 200 OK came back marked text/html — which is the cue for Module 3. The second request, the favicon, is the one to dwell on: no element in the document requested it. The user agent did, because fetching a site's icon is part of its behavior, not the author's instructions. The I/O subsystem has a will of its own.
2.10 What you now have
The path from a keystroke to the first byte: the omnibox's navigate-or-search decision, normalization into a parsed URL, scheme dispatch choosing which machinery runs, DNS turning a name into an address, a TCP/TLS or QUIC handshake opening an authenticated encrypted channel across costly round trips, a request carrying headers the browser supplied on your behalf, and a response whose status code and content type decide what happens next. You also have the framing that ties it to Module 1: a renderer asked, the browser process did all of this, and the result is bytes — not yet a page.
Those bytes are text/html, which means they go to the parser. Module 3 is where a stream of bytes becomes the DOM: the tree that every other subsystem in the series reads from and writes to. The I/O subsystem delivered the program. Now the browser has to load it into memory and make sense of it.
The network gives the browser a stream of bytes. Everything after this is the browser deciding what those bytes mean.
Module III · Viewing
Parsing HTML into the DOM
Module 2 delivered a stream of bytes marked text/html. A stream of bytes is not a page, not a tree, not anything you can style or script. This module is the transformation that makes it into something: the parser that turns bytes into the Document Object Model, the in-memory tree that every later subsystem in this series reads from and writes to. It is the single most consequential data structure in the browser, and building it correctly from input that is frequently broken is one of the hardest things the platform does.
3.1 Four things, not one
“Parsing HTML” sounds like one step. It is a short pipeline of distinct transformations, each with a different input and output, and conflating them is the source of most confusion about how pages load.
3.2 Bytes are not characters
The network hands over bytes, and a byte is not a character until you know the encoding. text/html; charset=utf-8 in the response header is the browser's first clue; a byte-order mark at the start of the stream is another; a <meta charset> inside the document is a third. The last one creates a small chicken-and-egg problem — the declaration of how to decode the bytes is itself inside the bytes — which browsers solve by pre-scanning the first chunk for the charset before committing. The modern default, and the right answer, is UTF-8. Get this wrong and you get mojibake: the famous garbled characters that mean someone decoded text with the wrong alphabet.
This is why <meta charset="utf-8"> belongs as early in the document as possible, and why the bytes-versus-characters distinction is not pedantry: it is the first place a page can be silently wrong before a single tag has been recognized.
3.3 A detour through parsing theory
To see why the HTML parser is unusual, recall what an ordinary parser is. Parsing means turning a flat input into a structured tree according to a grammar. The work splits in two: a tokenizer (or lexer) groups characters into meaningful units — tokens — and a parser assembles tokens into a tree by matching grammar rules. A compiler does exactly this with source code: tokenize 2 + 3 into a number, an operator, a number; then build an expression tree.
Languages amenable to this have a context-free grammar: a finite set of rules that say how symbols combine, independent of surrounding context. CSS and JavaScript are close enough to this that their parsers can be generated mechanically from a grammar. The classic tools — Lex and Yacc, Flex and Bison — exist precisely to turn a grammar into a working parser. If HTML were context-free, the browser could do the same and this module would be short.
3.4 Why HTML breaks the rules
HTML is not context-free, and cannot be parsed by a generated parser, for three reasons that compound.
It is forgiving by design. You may omit the <html>, <head>, and <body> tags entirely and still get a correct document with all three elements — the parser inserts them. You may leave a <li> or a <p> unclosed and the parser closes it for you when the context demands. This forgiveness is the main reason HTML succeeded: it lowered the cost of authoring to near zero, and a generation of authors learned it by trial and error rather than by reading a grammar.
It must tolerate decades of broken pages. Browsers compete on rendering the existing web, and the existing web is full of malformed markup that nonetheless “worked” because some browser once guessed well. Any new parser has to make the same guesses, so the error handling is not an afterthought — it is a specification.
It can rewrite its own input. A <script> running mid-parse can call document.write() and inject new characters into the stream the parser is still reading. The source is not fixed during parsing; the act of parsing can change what is being parsed. No ordinary parser model survives that.
The resolution, hard-won, is that the HTML Standard now specifies one exact parsing algorithm — every state, every transition, every error-recovery action — so that all conformant browsers build the identical tree from identical bytes, valid or not. The forgiveness was made deterministic.
- 1990s–2000s
- Each browser invents its own HTML error handling. The same broken page yields different trees in different browsers; “works in one, breaks in another” is the daily reality.
- 2008–2014
- HTML5 specifies the parsing algorithm in full, including precise error recovery. Compatibility stops being folklore and becomes a written contract.
- Today
- The algorithm lives in the WHATWG HTML Standard as a living document. Every engine implements the same state machine; the tree is portable even when the markup is wrong.
HTML's parser is lenient on purpose — and the leniency is specified to the comma, so that being lenient is also being identical everywhere.
3.5 The tokenizer is a state machine
The first half of the parser is the tokenizer, and it is a finite state machine: it reads one character at a time, and what it does depends on which state it is in. In the data state, ordinary text accumulates into character tokens. A < moves it into a tag-open state; a letter after that starts a tag name; whitespace inside a tag begins the attribute machinery; a > emits the finished tag token and returns to the data state. The full specification defines roughly eighty states to handle comments, doctypes, character references, raw-text elements like <script>, and every edge case; the spine of it is small enough to draw.
3.6 Tree construction and the art of fixing soup
Tokens are still flat. The second half of the parser — tree construction — consumes the token stream and builds the DOM, maintaining a stack of currently open elements and a notion of “insertion mode” that changes as it goes (before head, in head, in body, in table, and so on). A start-tag token pushes an element onto the stack and into the tree; an end-tag pops it. When the markup violates expectations — a stray </p> with no open paragraph, a <table> with text loose inside it, mis-nested bold and italic tags — the algorithm has named, specified recovery procedures with wonderful names like “foster parenting” and the “adoption agency algorithm.” The result is that a tag soup of broken markup still produces a clean, well-formed tree.
Consider a small, valid document and the tree it yields:
<!doctype html>
<html lang="en">
<head><title>Demo Company</title></head>
<body>
<h1>Demo Company</h1>
<p>Fake home page.</p>
</body>
</html>doctype as a child to style. The tree is a clean object model, not a copy of the text.3.7 The DOM is the runtime, not the source
This is the conceptual payload of the module. The DOM is not your HTML file. It is a live object model built from your HTML file, and from that point on the file is irrelevant — everything operates on the tree. CSS selectors match against DOM nodes, not against source text. Layout reads the DOM. document.querySelector queries the DOM. When a script adds an element, it adds a DOM node, and there is no corresponding text anywhere; “view source” still shows the original bytes while the page on screen reflects a tree that has moved far beyond them.
The DOM also has children downstream of itself. The browser derives the accessibility tree — the structure a screen reader navigates — from the DOM, mapping elements and ARIA attributes to roles and names. It is a parallel tree computed from the same source, and it is how the blind user from the cast experiences the document the parser just built. A malformed tree, or a tree of meaningless <span>s, produces a malformed or meaningless accessibility tree.
You author text. The browser runs a tree. Confusing the two is the root of a surprising number of bugs.
3.8 Parsing is incremental — and interruptible
The parser does not wait for the whole document. It streams, building the tree from bytes as they arrive off the network, which is why pages can start displaying before they have fully downloaded. But the stream can be stalled, and by one thing in particular: a plain <script> with no async or defer attribute. Because that script might call document.write() and alter the very input the parser is consuming, the parser must stop, hand control to the JavaScript engine, let the script run to completion, and only then resume. If the script itself has to be fetched from the network first, parsing halts until it arrives.
To soften this, browsers run a preload scanner: a lightweight look-ahead that races down the raw bytes while the main parser is blocked, spotting src and href references and kicking off their downloads early. It is a speculative optimization layered on top of a fundamentally sequential algorithm. These two facts — that scripts block parsing and that a scanner mitigates it — set up the whole of Module 4, where a page reveals itself to be a tree of dependencies, and Module 6, where blocking gets explained through the event loop. Demo Company's critical-render-path page exists precisely to let you watch a synchronous script freeze the parser.
3.9 The lab: source versus tree
Open Demo Company's home page and do two things side by side. Use “View Source” to see the bytes the server sent — the input to this module. Then open DevTools' Elements panel to see the DOM — the output. On a clean page they look almost identical, which is the point of authoring valid markup. The lesson lands harder on a broken page: paste malformed HTML into a file — omit closing tags, nest things wrongly — load it, and compare. The Elements panel shows you the tree the parser repaired into existence, often with tags you never wrote, sitting exactly where the algorithm's recovery rules put them.
3.10 What you now have
The pipeline from bytes to DOM: decode bytes into characters under a known encoding; tokenize characters into tags, text, comments, and doctypes with a state machine; construct a tree from those tokens with a specified, error-tolerant algorithm that makes even broken markup yield an identical, well-formed tree across browsers. And the central idea: the DOM is the runtime artifact — not your source text but a live object model, parent to the accessibility tree, target of every selector and script, the thing the rest of the browser actually operates on.
But the DOM the parser builds is rarely complete on its own. It is full of references — to stylesheets, images, fonts, scripts, frames — each of which is another resource the browser must go and fetch. A page is not a file; it is a dependency tree. Module 4 is how the browser discovers, prioritizes, and loads everything the document points at, with Demo Company's resource pages adding one dependency type at a time.
The parser's output is a tree of nodes — and a list of things the page still needs. Next we go and get them.
Module IV · Viewing
The resource model
Module 3 ended with a DOM and a list: the tree the parser built, plus all the things it referenced and does not yet have. A real page is rarely one file. It is an HTML entry point that points at stylesheets, scripts, images, fonts, and embedded documents — each of which is a separate request, some of which point at further resources of their own. A page is a dependency tree, and loading it is a resolution problem. This module is how the browser discovers, prioritizes, and fetches everything a document depends on, and why the order it does so determines how fast the page feels.
4.1 A page is not a file
This is the single shift in thinking the module asks for. The index.html you author is the root of a tree, not the whole thing. Demo Company makes the tree visible by building it one branch at a time: the home page needs only itself and a favicon; the image page adds a .webp; the CSS page adds a stylesheet; the font page adds a font discovered inside that stylesheet; the JavaScript page adds a local script and a third-party one from a CDN; the iframe page adds an entire embedded document with its own dependencies. Each page prints its own request list, so you can read the dependency tree off the screen.
4.2 Discovery: finding the dependencies
The browser learns of most dependencies while parsing, as the tokenizer and tree builder of Module 3 encounter <link>, <script>, <img>, and <iframe> elements. But discovery is not uniform. A font referenced by @font-face or a background image set in url() is invisible until the CSS itself has been fetched and parsed — the browser has to load one resource to find out it needs another. That is the amber layer in Figure 4.1, and it is why a web font often starts downloading conspicuously late.
Worse, discovery competes with the parser-blocking problem from Module 3: a synchronous <script> stops the parser, which would stop discovery of everything below it in the document. The browser's answer is the preload scanner — a secondary, lightweight pass that races ahead through the raw bytes while the main parser is blocked, finds src and href references, and starts fetching them early. Discovery order and network order are therefore not the same thing, and a good deal of browser engineering exists to make the network busy as early as possible.
4.3 The waterfall
Plot every request against time and you get the waterfall — the most important diagnostic view in web performance, and the one in your DevTools Network panel. Each bar is one resource; its horizontal position is when it started and how long it took; its vertical stacking shows what happened in parallel and what had to wait.
Two structural facts jump out of any waterfall. Parallelism is limited — over HTTP/1.1 a browser opens only a handful of connections per host, so requests queue; HTTP/2 and HTTP/3 multiplex many over one connection and largely dissolve that queue (Module 2). And chains are expensive — a resource that depends on another cannot start until the first finishes, so a three-deep chain (HTML → CSS → font) costs three sequential round trips no amount of bandwidth can collapse. The art of loading performance is turning chains into parallel fetches.
4.4 Render-blocking, async, and defer
Not all dependencies are equal in how they interrupt the page. Two behaviors dominate, and controlling them is most of what front-end performance work consists of.
CSS is render-blocking by default. The browser will not paint until it has the CSS, because painting first and restyling second would flash unstyled content at the user. So a stylesheet in the head holds back the first paint of the entire page — which is usually what you want, and is why the critical CSS should be small and fast.
Scripts choose their blocking behavior. A plain <script> is parser-blocking and ordered; the modern attributes change that:
async fetches alongside parsing and runs whenever it lands; defer fetches alongside parsing but runs in order after the document is built. For a script that touches the DOM, defer is almost always the right default. (Module-type scripts defer by nature.)4.5 Priorities, limits, and hints
The browser does not fetch the dependency tree in document order; it fetches by priority. Render-blocking CSS and the fonts needed for visible text are high priority; an image far down the page or an async analytics script is low. The browser assigns these priorities automatically from the kind of resource and where it sits, and you can nudge them: the fetchpriority attribute raises or lowers a specific request, and the resource hints — preconnect to warm up a connection, preload to start a critical fetch early, prefetch to grab something for the next navigation — let the author correct the browser's guesses. These are how you fix the “font discovered late” problem: preload it so it does not wait for the CSS.
The browser is a scheduler for the network. Priorities decide what the user sees first; hints are how the author argues with the defaults.
4.6 First party, third party: what another host costs
Every waterfall so far came from one host. Put a request on a second host — a library from a CDN, an analytics tag, a stock photo — and its bar grows a prefix before the first byte: look up the name (DNS), open a connection (TCP), and agree on encryption (TLS 1.3 needs one more round trip; HTTP/3 folds the connection and the encryption together, as Module 2 showed). Your own host paid those costs once, for the page itself. Each new host pays them again.
The hints from 4.5 exist for this. <link rel="preconnect" href="https://cdn.example"> does the lookup, connection, and encryption early, for a host you know the page will need; rel="dns-prefetch" does only the lookup. Add crossorigin when the files will be fetched with CORS, as fonts are, or the browser opens a second connection anyway. Hints save the setup, not the download, and a preconnect to a host the page never uses is a connection wasted.
The old case for a shared CDN was the cache: load the popular library from the popular address, and visitors probably have it already. That case is closed. Browsers now keep a separate HTTP cache for each top-level site (Safari since 2013, Chrome since version 86, Firefox since 85), so a copy fetched while visiting someone else's site is not reused on yours. Served from your own host, the same file costs the same bytes and no new connection.
Which requests count as third party? An origin is a URL's scheme, host, and port. A site is its scheme and registrable domain, the part a person can register, as the Public Suffix List decides. static.democompany.com is a different origin from www.democompany.com but the same site: usually a second connection, yet still the first party. www.someothercompany.com is another site, a true third party. Try any pair in the URL Explorer’s Compare tab.
A third party also costs control. Its speed, its uptime, and the file at its address tomorrow are not yours. For any other origin, the browser shows your own scripts only how long a file took — not its size and not its DNS, connection, and download phases — unless that server sends Timing-Allow-Origin. Demo Company's page 5 lists the ZingGrid script's time beside a size marked “hidden” in its request log. The film Everything You Embed follows what goes wrong next.
Every new host is a toll: a name, a connection, a handshake, and some trust. Pay it on purpose.
4.7 The critical rendering path
Out of the whole dependency tree, only a subset is required before the browser can show anything: the HTML, the render-blocking CSS, and any parser-blocking JavaScript ahead of the visible content. That subset is the critical rendering path, and shortening it — less blocking CSS, deferred scripts, preloaded essentials — is the most direct lever on how quickly a page paints. Everything outside the critical path (below-the-fold images, async scripts, non-essential fonts) can arrive afterward without delaying first paint. Demo Company's critical-render-path page exists to let you watch a single synchronous script sit squarely in that path and hold the whole page hostage until it loads and runs.
This is the bridge to Module 5. Once the critical resources are in — the DOM from Module 3 and the CSSOM built from the render-blocking CSS — the browser has everything it needs to compute styles, lay out boxes, and paint. The resource model decides when the pixel pipeline can start.
4.8 When dependencies fail
The network is not reliable terrain, and a dependency tree is only as robust as its handling of missing branches. A resource can be slow, can 404, can hang, can be blocked. A well-built page degrades rather than breaks: a failed image leaves its alt text and reserved space rather than collapsing the layout; a font that does not arrive falls back to a system face via font-display; a failed enhancement script leaves working HTML behind it. This is progressive enhancement seen from the loader's side — build the critical path so that the non-critical failing does not take the page down with it.
4.9 The lab: watch the tree grow
Walk Demo Company's resource pages in order with the Network panel open, and read each new request against this module:
image.html + cat.webp a leaf: an image referenced straight from the DOM
css.html + main.css render-blocking; holds first paint
font.html + font.css → font a chain: the font is named inside the CSS
javascript.html + local-script.js + zinggrid.min.js local and third-party (cdn.zinggrid.com) scripts
iframe.html + embedded document a whole subtree with its own resourcesSort the waterfall by start time and watch the parallelism; sort by the dependency and watch the chains. Find the font on the font page and confirm it begins only after font.css finishes — then imagine the preload that would fix it. You are looking at the resource model's own ledger.
4.10 What you now have
The page as a dependency tree rather than a file; discovery through parsing and the preload scanner racing ahead of blocked parsing; the waterfall as the view that exposes parallelism and chains; the blocking behaviors that matter most — render-blocking CSS and the plain/async/defer choice for scripts; priorities and hints as the controls; the price of every new host, and the difference between a second origin and a second site; and the critical rendering path as the subset that gates first paint. The unifying idea: the browser is a scheduler for the network, and how it orders the fetch decides how fast the page feels.
With the critical resources in hand — the DOM and the CSSOM — the browser can finally turn structure and style into something on screen. Module 5 is the pixel pipeline: computing styles, laying out boxes, painting, and compositing layers on the GPU, and the reason some changes are cheap and others force the whole machine to run again.
You wrote one file. The browser fetched a forest. Now it has to draw it.
Module V · Viewing
Style, layout, paint, composite
The browser now holds two trees: the DOM from Module 3 and the CSSOM built from the render-blocking CSS of Module 4. Neither is pixels. This module is the pipeline that turns structure and style into a lit screen — four stages, each with a distinct job, each with a distinct cost. Understanding the stages is what separates “my animation is janky” from knowing exactly which stage is running sixty times a second and why.
5.1 Two trees, one picture
The journey from the two trees to the screen is a fixed sequence: compute the styles, compute the geometry, draw the pixels, assemble the layers. The names are style, layout, paint, and composite. The order is not negotiable, because each stage consumes the output of the previous one — you cannot place a box before you know its font size, and you cannot paint it before you know where it goes.
5.2 Style: resolving the cascade
The style stage walks the DOM and, for every element, computes its computed style — the final value of every CSS property after the cascade, specificity, inheritance, and custom-property resolution have all been applied. Selector matching happens here: the engine decides which of the CSSOM's rules apply to which nodes. The output is the render tree (WebKit and Blink call it the layout tree; Gecko, the frame tree) — a tree of the elements that will actually be drawn. It is not the DOM. Elements with display: none are absent entirely; pseudo-elements like ::before are present though they exist in no markup; visibility: hidden elements are present but will paint nothing. The render tree is the DOM filtered and annotated for display.
5.3 Layout: computing geometry
Layout — historically called reflow in Gecko — takes the render tree and computes the exact position and size of every box: the box model, normal flow, flex, grid, the recent fragmentation work. Its output is a geometry tree where every node knows its rectangle in the page. Layout is the expensive stage, and it is expensive because it is global-ish: changing the width of one element can change the position of its siblings, its children, and everything after it in flow. One small write can force the browser to recompute the geometry of a large part of the page.
The classic performance trap lives here: forced synchronous layout, also called layout thrashing. The browser batches your style and layout changes to run them once per frame — but if your JavaScript writes a style and then reads a geometry property like offsetWidth, the browser must flush layout immediately to answer the read, because the answer depends on the write. Do that in a loop — write, read, write, read — and you force dozens of synchronous layouts in a single frame, turning a smooth interaction into a stall. The fix is to batch: read all the geometry first, then write all the changes. This is a Module 6 concern as much as a Module 5 one, because it is about when on the main thread the work happens.
5.4 Paint and composite
Paint turns the laid-out boxes into draw commands: fill this rectangle, stroke this border, render this text run, draw this shadow. The browser does not usually paint straight to the screen; it records lists of paint operations, often grouped into separate layers for parts of the page that can be handled independently. Rasterization — actually producing the pixels from those commands — frequently happens on the GPU.
Composite is the final assembly. The page is divided into layers; each is rasterized; the compositor places them in the right order with their transforms and opacity and hands the result to the screen. The decisive fact is that the compositor runs on its own thread and the GPU, separate from the main thread. Once a layer has been painted, the compositor can move it, scale it, rotate it, or fade it without repainting and without re-running layout — it is just re-placing an existing texture. That is the entire reason the next diagram matters.
5.5 What a change costs
Every visual change re-enters the pipeline, but not always at the top. The stage where a change enters determines its cost: a change that forces layout is expensive because paint and composite must follow; a change that only composites is nearly free because it skips the main thread's heavy stages entirely.
transform and opacity” is the most repeated advice in front-end performance. Those two properties enter the pipeline at the bottom, run only on the compositor and GPU, and so can hit sixty frames a second while the main thread is busy. Animating width or top forces the whole machine to run every frame.Cheap pixels come from the bottom of the pipeline. If you can express a change as a transform or an opacity, the GPU will do it while the main thread sleeps.
Promoting an element to its own compositor layer — with will-change or a 3D transform — is how you tell the browser “this will move; keep it ready.” It is powerful and easy to overuse: every layer costs memory, and a page that promotes everything trades smooth animation for exhausted GPU memory. Layer promotion is a hint to spend, not a free win.
5.6 Layout shift: when geometry changes late
Module 4 showed resources arriving over time. Some of them change geometry when they land, and that re-runs layout — visibly. An image with no declared dimensions takes up no space until it loads, then suddenly claims its size and shoves everything below it downward. A web font that arrives after first paint may have different metrics than the fallback, so the text re-flows the instant it swaps. The user, mid-read or mid-tap, watches the page jump. This is layout shift, measured as Cumulative Layout Shift, and it is the pipeline's most user-hostile behavior.
width and height (or aspect-ratio) on images so the box is sized from the start, and font-display plus matched fallback metrics so a font swap does not re-flow. Demo Company's building image and web-font pages are this diagram, live.5.7 The lab: watch the pipeline run
DevTools exposes every stage. In the Rendering tab, turn on “Paint flashing” to see green flashes wherever the browser repaints, and “Layer borders” to see which elements have been promoted to their own compositor layer. In the Performance panel, record an interaction and read the main-thread track: purple is layout, green is paint, and you can spot a forced synchronous layout as a tall layout spike wedged inside your scripting time. Load Demo Company's web-font page and watch the text repaint — or shift — the moment the font swaps in.
5.8 What you now have
The four-stage pipeline from two trees to a screen: style resolves the cascade into a render tree; layout computes geometry and is the expensive, global stage where forced synchronous layout lurks; paint records draw commands into layers; composite assembles those layers on a separate thread and the GPU. The actionable core is the cost model — a change that triggers layout drags the whole pipeline behind it, while a transform or opacity change rides the compositor nearly for free — and its corollary, layout shift, where late-arriving resources re-run layout in the user's face.
Every heavy stage in this module — style, layout, paint — runs on the main thread, and that thread does one thing at a time. It also runs your JavaScript. Module 6 is the scheduler that decides how rendering and scripting share that single thread: the event loop, the sixteen-millisecond frame budget, the task and microtask queues, and requestAnimationFrame — the OS scheduler from Module 1, finally in full.
The pipeline is fixed; the cost is yours to choose. You pick which stage your changes re-enter, and the user feels the difference at sixty frames a second.
Module VI · Running
The JavaScript engine and the event loop
Module 1 promised that the event loop was the browser's scheduler, and Module 5 left a thread doing style, layout, and paint — the same thread that runs your code. Now we collect on both. This module is how a single thread runs a page's JavaScript and renders the page at sixty frames a second without doing two things at once, why a long function freezes everything, and how the platform lets you escape that when you must. It is the scheduler from Module 1, in full.
6.1 JavaScript is the browser's userland
The JavaScript engine — V8 in Chrome, JavaScriptCore in Safari, SpiderMonkey in Firefox — parses your source, compiles it to bytecode, and just-in-time compiles the hot paths to machine code, managing a heap of objects and a call stack of running functions, reclaiming memory with a garbage collector. That machinery is a course of its own. For this series the engine is the userland process from Module 1: the place a page's untrusted program actually executes. The interesting question is not how the engine compiles a function, but how the browser decides when to run it, given that it has exactly one main thread to run it on.
6.2 One thing at a time: run to completion
The defining rule of the model is run to completion: once a piece of JavaScript starts, it runs to its end before the browser will do anything else — no other script, no event handler, no rendering, no input. The call stack fills and empties for that one task, uninterrupted. This is a feature: you never have to worry that another thread mutated your variable halfway through a function, because there is no other thread. It is also the source of every “the page froze” experience, because if your function does not finish quickly, nothing else gets a turn.
JavaScript is single-threaded and runs to completion. Everything good and everything painful about the browser's responsiveness follows from that one sentence.
6.3 The event loop
If a task runs to completion and then the thread is free, what runs next? The event loop answers: it repeatedly takes one task from the task queue, runs it to completion, then drains the entire microtask queue, then — at most once per frame — updates the rendering. Then it goes back for the next task. That cycle, running forever, is the whole life of a page.
6.4 Tasks versus microtasks
The two queues are not interchangeable, and the difference trips up even experienced developers. A task (or macrotask) is a unit of work the loop picks up one at a time: a timer firing, a click handler, a network response callback. A microtask is a follow-up scheduled by the currently running code: a promise's .then, the continuation after an await, a queueMicrotask callback. After every task — and after the call stack empties — the loop drains the microtask queue completely before doing anything else, including before taking the next task or rendering.
console.log('A');
setTimeout(() => console.log('B'), 0);
Promise.resolve().then(() => console.log('C'));
console.log('D');
A and D run synchronously. When the stack empties, the loop drains microtasks — C — before it will touch the task queue, so B, despite its zero delay, runs last. The practical warning hides in the word completely: because draining microtasks runs any microtasks they themselves schedule, an endless chain of promises can starve the loop and block rendering forever, even though no single function runs long. The loop is not stuck in your code; it is stuck honoring your microtasks.
6.5 Rendering is a guest on the loop
Here is where Module 5 connects. The rendering pipeline — style, layout, paint — is not a separate clock. It is step 4 of the loop, run between tasks, after microtasks, and at most once per display refresh. requestAnimationFrame is the hook into that step: a callback registered with it runs immediately before style and layout, which is exactly where visual updates belong, so they happen once per frame in sync with the pipeline rather than at some arbitrary moment that forces an extra layout. If nothing changed, or it is not yet time for a frame, the browser simply skips the rendering step and moves on.
The consequence is blunt: rendering only happens when the loop reaches step 4, and the loop only reaches step 4 when your task and its microtasks have finished. While your JavaScript runs, the page cannot repaint. The smooth-scrolling, sixty-frame-a-second page and the frozen, unresponsive one are running the identical loop; the difference is entirely whether step 2 keeps finishing in time for step 4.
6.6 The frame budget
A 60Hz display refreshes every 16.7 milliseconds, and a 120Hz display every 8.3. That interval is the entire budget for one turn of the loop that wants to produce a frame: the task, all its microtasks, the animation-frame callbacks, and style, layout, and paint must all fit. Go over, and the frame is late — the browser shows the previous frame again, and the user sees a stutter. A function that runs for forty milliseconds does not just take forty milliseconds; it costs two or three dropped frames and ignores every tap and keystroke that arrives while it runs.
6.7 Escaping the single thread
If the main thread can only do one thing at a time and rendering waits on it, heavy work has to get out of its way. There are two moves.
Yield. Break a long job into chunks and hand control back to the loop between them, so rendering and input get their turns. Historically this meant scattering work across setTimeout calls; modern platforms add purpose-built tools — scheduler.yield(), isInputPending() — that let a task voluntarily pause when something more urgent is waiting. Cooperative, but effective.
Leave the thread entirely. A Web Worker is a real operating-system thread running JavaScript in parallel with the main thread. It has no access to the DOM — it cannot touch the page — and it communicates by passing messages, but it can crunch numbers, parse data, or run a heavy algorithm without stealing a single frame from the main thread. This is genuine parallelism, and it is the same idea Module 5 already showed at work: transform and opacity animations stay smooth during a busy main thread precisely because the compositor runs them on another thread. The platform's answer to a single-threaded loop is more threads, walled off from the DOM for safety.
6.8 A cooperative scheduler
Step back to the Module 1 analogy and the picture sharpens. The event loop is a cooperative, non-preemptive scheduler: a task runs until it voluntarily finishes, and nothing can forcibly interrupt it. That is exactly the multitasking model early operating systems used before preemption — and exactly why a single misbehaving program could hang the whole machine, the same failure mode Module 1 traced in single-process browsers. The browser solved the process version with isolation; it has not made the main thread preemptive, because run-to-completion is what keeps the DOM free of data races. Web Workers are the preemptive escape hatch: real threads the operating system can schedule in parallel, kept away from shared mutable state by design. The browser is an OS that chose cooperative scheduling for its userland and quarantined the dangerous parallel work behind a message queue.
The main thread will never interrupt your code — which is a promise to you and a loaded gun pointed at the user. Finish quickly, or yield.
6.9 The lab: see the loop block
Demo Company's critical-render-path page loads a synchronous script that blocks — the loading-side version of this module's lesson. Open it with the Performance panel recording and watch the main-thread track: a long block where the script runs, no frames produced inside it, a red corner flagging it as a long task. Then run the ordering snippet from section 6.4 in the console and confirm the output is A D C B — microtasks before the next task, every time.
6.10 What you now have
The runtime model in full: one main thread, run-to-completion, and an event loop that takes one task, drains all microtasks, and renders at most once per frame. The task-versus-microtask asymmetry and its starvation risk; rendering as a guest that only runs when the loop reaches it; the 16.7-millisecond frame budget and the long task that blows past it; and the two escapes — yielding cooperatively and offloading to Web Workers, the preemptive threads kept away from the DOM. And the framing that unifies it with Module 1: the loop is a cooperative scheduler, the same design, and the same hazard, as the operating systems that came before preemption.
But notice what the loop does not know. It runs every task with equal authority — your code, the third-party CDN script from Module 4, and the malicious script you have not met yet, all on the same thread, with the same access to the same DOM. The event loop has no concept of trust. Deciding which code may do what — read which data, call which API, reach which origin — is a different subsystem entirely. Module 7 is the security model: the permission boundary that the loop, on its own, does not enforce.
The loop runs every script as an equal. The browser's hardest job is making sure equal access to the thread does not mean equal access to everything.
Module VII · Running
The security model
Module 6 left a thread that runs every script with equal authority and no notion of trust — your code, the third-party CDN script from Module 4, and code that means you harm, all on the same loop with the same reach into the same page. Something has to decide who may do what. That something is the security model, and it is the browser at its most operating-system-like: an access-control system separating mutually distrustful code, running on a machine where the code arrives unbidden from strangers. This is the module where Demo Company's threat pages stop being curiosities and become the syllabus.
7.1 The origin is the unit of trust
The browser's entire security model is built on one identifier: the origin, the triple of scheme, host, and port. Module 1 called the origin the platform's process and permission boundary; this module is what that boundary actually does. Two pieces of content share trust if and only if they share an origin. Same origin means full mutual access; different origin means walls, by default. Everything else — the same-origin policy, CORS, cookie scoping, storage partitioning — is elaboration on that one line.
https and http are different origins, which is one more reason the modern web is HTTPS-only.7.2 What the policy actually restricts
The same-origin policy is precise about what it stops, and the precision is where the danger and the defenses both live. It restricts reading across origins, not sending and not embedding. A page may freely send a request to another origin and freely embed another origin's resources — but it may not read the response or inspect the embedded content. That asymmetry is the most important table in web security.
src does not stay at arm's length — it executes inside your origin with every power your own code has. That is the Module 4 CDN line, and it is why third parties are a trust decision, not a convenience.The web lets you embed anything and read almost nothing — but a script you embed becomes you. Choosing what to embed is choosing whom to trust completely.
7.3 CORS: the controlled exception
Plenty of legitimate work needs to read across origins — a front end on one origin calling an API on another. Cross-Origin Resource Sharing is the opt-in that allows it. The server that owns the data chooses to return an Access-Control-Allow-Origin header naming who may read its responses; for anything beyond simple requests, the browser first sends a preflight OPTIONS request to ask permission before sending the real one. The key insight, and a common misunderstanding: CORS does not protect the server — the server can always be reached. It protects the user's data on other origins by default-denying the read, and lets the data's owner grant exceptions. CORS is the browser enforcing one server's policy on behalf of a user talking to another.
7.4 Cookies and credentials
A cookie set by an origin is attached automatically to future requests to that origin — the mechanism that makes you stay logged in (Module 2's Cookie header). That automatic attachment is powerful and dangerous, so cookies carry flags that are pure security controls: HttpOnly hides the cookie from JavaScript entirely, so that even a script running in your origin cannot read your session token; Secure sends it only over HTTPS; SameSite controls whether it rides along on requests triggered by other sites, the central defense against request forgery. Third-party cookies — set by an origin embedded across many sites — are the classic cross-site tracking mechanism, and their long deprecation is one of the platform's biggest privacy shifts.
7.5 The attacks the architecture permits
Three classic attack classes follow directly from the model above. We describe them as a defender must understand them — what they are and why the architecture allows them — not as recipes.
Cross-site scripting (XSS) is the bottom row of Figure 7.2 turned against you. If an attacker can get their script to run in your origin — by smuggling it through unsanitized user input that your page reflects back as markup — that script has the full access any in-origin code has: it can read the DOM, read non-HttpOnly cookies, and make same-origin requests as the user. The injected script is trusted not because it is trustworthy but because it is in the origin. This is why HttpOnly on a session cookie matters so much: it removes the session token from what an injected script can reach.
Cross-site request forgery (CSRF) weaponizes the send-but-not-read half of the asymmetry. A page cannot read another origin's response, but it can send a request — and the browser will attach that origin's cookies. So a malicious page can cause the user's browser to fire a state-changing request to a site they are logged into, riding their session, even though the attacker never sees the reply. The defense is SameSite cookies, which withhold the credential from cross-site requests, plus anti-forgery tokens the attacker cannot guess.
Clickjacking abuses cross-origin embedding: an attacker frames your real site invisibly over their own bait, so the user's click lands on your page without their knowledge. The defense is telling the browser who may frame you — the frame-ancestors directive of Content-Security-Policy, historically X-Frame-Options.
7.6 Defense in depth
No single mechanism secures a page; the web's security is a stack, each layer assuming the others might fail. The author opts into most of it through HTTP response headers and element attributes — the browser enforces what the server declares.
7.7 The lab: Demo Company's threat arc
The threat pages now read as a single argument. Walk them with the Network and Console panels open:
javascript.html third-party CDN script runs in your origin: supply-chain trust (use SRI + CSP)
resources.html cross-origin image embeds fine; pixels not readable (canvas taint)
resources-keylog.html tracking script in-origin, so it CAN read your input; SOP won't help
resources-hacked.html malicious script the XSS endpoint: in-origin code with full accessThe throughline: every one of these arrives via the same ordinary <script> or resource reference from Module 4, and the event loop of Module 6 runs each with identical authority. The same-origin policy guards the boundary between origins, but does nothing about hostile code you invited inside yours. The real defenses are upstream: do not load code you do not trust; constrain what can load with CSP and SRI; and keep the secrets that matter (HttpOnly session tokens) out of JavaScript's reach entirely, so that even a breach inside the origin cannot take them.
7.8 What you now have
The model whole: the origin as the unit of trust; the same-origin policy restricting cross-origin reads while permitting sends and embeds; the dangerous exception that an embedded script runs inside your origin with full power; CORS as the owner's opt-in to cross-origin reads; cookies and their security flags; the three attack classes the architecture permits — XSS, CSRF, clickjacking — and why each follows from the model; and defense in depth, the stack of TLS, SOP and CORS, CSP and SRI, cookie flags, and the sandbox and Site Isolation from Module 1. The unifying idea: the event loop runs every script equally, so a separate subsystem must decide what equal access to the thread is allowed to reach — and that subsystem is keyed entirely on the origin.
The origin turns out to scope more than scripts and requests. The data a page stores — cookies, but also Web Storage, IndexedDB, the cache — is partitioned by origin too, so the security boundary extends into persistence. Module 8 is the browser as a filesystem: what a page can keep, where, for how long, and under whose isolation — plus the service worker, a programmable proxy that sits between the page and the network, and the hardest naming problem in computing.
The same-origin policy guards the door between origins. It cannot help with the code you carried inside. Security is mostly about what you choose to let in.
Module VIII · Running
State, storage, and caching
A page is ephemeral: reload it and Module 3 builds a fresh DOM from scratch, with no memory of the last visit. Yet sites remember you — your login, your draft, your offline data. The memory does not live in the page; it lives in the browser, scoped to the origin, surviving reloads, tabs, and restarts. This module is the browser as a filesystem and a database you never provisioned: what a page can store, where, for how long, under whose isolation — the service worker that turns the browser into a programmable proxy, and the problem the whole field agrees is hardest.
8.1 The browser remembers, by origin
Module 7 established the origin as the unit of trust, and storage inherits that boundary completely. Every storage mechanism is partitioned by origin: https://www.democompany.com cannot read what https://api.democompany.com stored, any more than it could read its DOM. The browser is, in effect, a multi-tenant filesystem where each origin gets its own private directory it cannot escape. What differs between the mechanisms is everything else — how much they hold, how you reach them, whether they cost network, and whether a script can read them — and choosing wrong is a performance or security bug waiting to happen.
sessionStorage is localStorage scoped to a single tab; all of them are partitioned by origin.8.2 Picking the right drawer
The differences map to clear advice. Cookies are for the small piece of state the server needs on every request — chiefly the session identifier — and Module 7's flags (HttpOnly, SameSite, Secure) are mandatory there. They are a terrible general-purpose store: four kilobytes, strings only, and a network tax on every single request. Web Storage (localStorage / sessionStorage) is the easy key-value drawer for small, non-sensitive preferences — but it is synchronous, so it stalls the thread, and it is readable by any script in the origin, so a successful XSS reads it wholesale. Never put a token or secret in localStorage; that is precisely the data an injected script wants, and unlike an HttpOnly cookie, it is right there.
For real data, IndexedDB is the client-side database: asynchronous, transactional, able to hold structured objects and large volumes, queryable by index. The Cache API stores Request/Response pairs and is the service worker's pantry, the subject of the next section. And the Origin Private File System gives an origin real file-like storage with high-performance access, enough that databases such as SQLite now run in the browser compiled to WebAssembly, reading and writing OPFS as their disk. The browser has quietly grown a complete storage stack — key-value, document database, and filesystem — all behind the origin wall.
Cookies for what the server needs; Web Storage for small preferences; IndexedDB and OPFS for real data. Secrets go in none of the ones a script can read.
8.3 The HTTP cache: the request you skip
Separate from all of that — and often confused with the Cache API — is the HTTP cache, the automatic store the browser manages on its own, the one Modules 2 and 4 kept gesturing at. When the browser fetches a resource, the server's Cache-Control header tells it whether and how long the response may be reused. While a cached copy is fresh, the browser serves it with no network at all — the fastest possible request, because it never happens. When it goes stale, the browser does not simply re-download; it revalidates.
ETag; the server replies 304 Not Modified with no body when nothing changed — a tiny round trip instead of a full download — or 200 with fresh bytes when it did. This is the “Disable cache” checkbox from Module 1, explained.8.4 Compression: fewer bytes on the wire
The cache decides whether a request happens. Compression decides how big it is when it does. Every request from a browser says which compression formats it can unpack, and the server picks one and says which it used:
GET /article.html
Accept-Encoding: gzip, deflate, br, zstd what the browser can unpack
200 OK
Content-Type: text/html; charset=utf-8 what it is
Content-Encoding: br how it was packed for the trip
Vary: Accept-Encoding caches: the answer depends on that headergzip is the old standby every client understands. Brotli (br) usually packs text smaller. zstd is the newest, accepted by Chrome and Firefox, and some CDNs already prefer it: Cloudflare sends Demo Company's pages to Chrome as zstd. Note the two headers: Content-Type says what the bytes are, Content-Encoding says how they were packed. The browser unpacks first and then treats the result as its type, so your scripts never see the compressed form.
Compression is for text: HTML, CSS, JavaScript, JSON, SVG. A typical page of HTML shrinks to a third of its size or less. Images, video, audio and WOFF2 fonts are already compressed by their own formats; packing them again costs the server work and saves almost nothing, so servers leave them alone. The bytes you save on images come from choosing the format, the dimensions and the quality, not from Content-Encoding.
Two more headers matter once caches sit in between. Vary: Accept-Encoding tells any cache on the way that the same URL has several bodies, one per encoding, so it must not hand a Brotli body to a client that never asked for one. Cache-Control: no-transform asks the proxies and CDNs on the way not to change the body at all, recompression included.
DevTools knows both numbers for every request: the bytes transferred and the resource's size once unpacked (in Chrome, turn on “Big request rows” to see both in the Size column, or hover it). When the first is much smaller than the second, compression did its job; when they match on a text file, the server sent it raw and someone should turn compression on.
8.5 The service worker: a proxy you install
The HTTP cache is automatic and the browser owns its rules. The service worker hands those rules to the page. It is a worker in the Module 6 sense — a background script with no DOM access, running off the main thread — but with a special power: once registered, it sits between the page and the network and intercepts every request the page makes. For each one it decides: answer from the Cache API, go to the network, do both, or synthesize a response from nothing. It is a programmable proxy the origin installs onto the user's machine, and it is what makes a web page work offline.
The service worker has a lifecycle — register, install (typically pre-caching the app shell into the Cache API), activate, then run, handling fetch events until the browser stops it — and it is restricted to secure contexts: HTTPS only, because a proxy over your requests is far too powerful to hand to an attacker on an open network. It is, in the Module 1 framing, a background daemon the origin installs, sandboxed and event-driven, mediating the page's I/O. Offline-capable, installable web apps — the subject of Module 9 — are built on exactly this.
8.6 Cache invalidation: the hard problem
Caching is free speed until the cached thing changes, and then it is a bug. There is a famous line — only two hard problems in computer science: cache invalidation and naming things — and this module sits squarely on the first. Every cache layer above poses the same question: how do you know the copy you saved is still correct? Serve stale content and users see yesterday's prices; revalidate too eagerly and you have thrown away the speed the cache was for.
The web's practical answers are worth knowing because you will reach for them constantly. Content hashing bakes a hash of the file's contents into its name — main.9f3c1a.css — so a changed file is a new URL the cache has never seen, and you can cache the old one forever because it can never go stale; this is why build tools fingerprint assets. Conditional revalidation with ETag, from Figure 8.2, lets the server cheaply confirm a copy is still good. Service-worker versioning names each cache generation and deletes the old one on activation. Stale-while- revalidate serves the stale copy instantly and refreshes it in the background, trading a moment of staleness for speed. Each is a different point on the same trade-off between fast and fresh.
8.7 The lab: read the browser's memory
DevTools' Application panel is the filesystem browser for this module: Storage shows cookies, Local and Session Storage, and IndexedDB for the current origin; Cache Storage shows what any service worker has stashed; the Service Workers pane shows registration and lets you simulate going offline. In the Network panel, the Size column tells the caching story directly — “(from disk cache)” or “(from memory cache)” for served-without-network, or a small transfer with a 304 status for a revalidation.
8.8 What you now have
The browser as a persistence layer, all of it walled off by origin: cookies for the small state the server needs every request; Web Storage for small, non-secret preferences; IndexedDB and OPFS for real databases and files, asynchronous so they do not stall the thread; the automatic HTTP cache with its fresh/stale/revalidate logic and the cheap 304; compression, which shrinks the text that does cross the wire; and the service worker, a programmable proxy installed onto the machine, intercepting requests and serving from the Cache API to make pages fast and offline-capable. Plus the field's hardest problem, cache invalidation, and the handful of strategies — content hashing, ETags, versioned caches, stale-while-revalidate — that tame it, and the privacy stakes that come with anything a browser can remember.
Storage and the service worker are not just conveniences; they are the foundation of an idea. A site that installs a proxy, caches its own shell, stores its own data, and works offline is no longer quite a document — it is an application, indistinguishable in use from one the operating system installed. Module 9 closes the series on the browser as a full platform: installable apps, the capability and permission model for hardware and sensors, DevTools as the browser observing itself, extensions, and the long argument between the web and native — then hands you back to authoring.
A page forgets everything on reload. The origin forgets nothing unless told to. The browser is the disk the web never had to ask permission to format.
Module IX · Modern · Finale
The browser as a platform
Eight modules built the browser as an operating system: an I/O subsystem, a parser, a resource loader, a render pipeline, a scheduler, a security model, a filesystem. This final module steps back to ask what that adds up to — a platform that rivals the native one beneath it — and then asks the harder question the technology alone cannot answer: who controls it, who profits from it, and what it costs the open web that the most-used runtime on Earth is also one of the most contested. We end on the politics, because for a working developer the politics are not separate from the engineering. They decide what you are allowed to build.
9.1 When a site becomes an app
Module 8 left a site that installs a service-worker proxy, caches its own shell, stores its own data, and works offline. Add a web app manifest — a small file naming the app, its icons, and how it should launch — and the browser will offer to install it: an icon on the home screen or dock, a window with no address bar, a cold start that loads from the cache. This is a Progressive Web App, and the word that matters is progressive: the same code is a normal web page for a first-time visitor and a full application for someone who installed it. There is no separate build, no store submission, no gatekeeper. The document and the app are the same artifact at different points on a continuum.
The web app is not a different kind of thing from the web page. It is the same page, after the browser agreed to treat it like software.
9.2 Capabilities and the permission gate
An application needs to reach hardware: camera, microphone, location, notifications, USB, Bluetooth, the filesystem. The web exposes these through APIs, and Module 1 called those APIs the platform's syscall surface. But unlike a syscall, each one is gated — twice. A powerful capability requires a secure context (HTTPS, the Module 2 padlock) and explicit, revocable user permission, requested at the moment of use and scoped to the origin.
9.3 The browser watching itself, and the user's hooks
Two more platform surfaces deserve a mention before the politics. DevTools — the panels every lab in this series has used — is the browser instrumented to observe its own subsystems: the network waterfall (Module 4), the render and main-thread tracks (Modules 5 and 6), the storage and cache inspectors (Module 8), the security and origin views (Module 7). It is the rare platform that ships a complete X-ray of itself to every user. Extensions are the other hook: user-installed code that can read and rewrite pages, the one place where the person, not the site, programs the browser — powerful enough to block ads and trackers, and powerful enough to be dangerous, which is why their permissions and review are themselves a running controversy.
9.4 The false binary: app versus site
Here the module turns. We have a platform that installs, works offline, reaches hardware, and persists data — a platform, by any technical measure, capable of the things “apps” do. So why do we still speak of apps and websites as different categories of thing? The honest answer is that the distinction is mostly commercial and political, not technical. It is maintained because a great deal of money depends on it. Understanding the browser as a platform means understanding the economics that keep the platform from being treated as one.
9.5 The app store as tollbooth
A native app reaches an iPhone through one door: Apple's App Store, which historically took a commission of fifteen to thirty percent on sales and required Apple's own payment system, while forbidding apps from even mentioning cheaper options elsewhere — the “anti-steering” rule. Critics call this rent: a tax on commerce that flows through a device the user already bought, enforced by control of the only distribution channel. Apple maintains the commission pays for a curated, secure, private marketplace and the platform it is built on. Courts and regulators have spent years between these positions. The long Epic Games litigation forced Apple, in the United States, to let developers link to outside payment options; Apple's attempt to keep charging a high commission on those external purchases was found to defeat the purpose, and as this is written the permissible fee is still being fought over up to the Supreme Court. Google, facing a parallel case, settled and dropped its Play Store commission. The numbers move; the structural question does not. Whoever controls distribution controls the economics of everything distributed.
The web has no such door. A web app is published by putting it on a server, and it takes no commission and asks no permission. This is precisely why the web is a threat to the store model — and why the next fight matters so much.
9.6 The iOS engine monopoly
For most of the iPhone's life, a single rule shaped the entire web on it: every browser on iOS, including Chrome and Firefox, was required to use Apple's own WebKit engine underneath. The Chrome on an iPhone was Chrome's interface wrapped around Apple's renderer. The consequence is the quiet center of the whole story: if the only engine on the most lucrative mobile platform is one the platform owner controls, then the platform owner sets the ceiling on what web apps can do there. Hold that ceiling low — ship web capabilities slowly, leave install and notification and hardware support thin — and web apps stay a weak substitute for native ones, which keeps developers in the App Store, which protects the commission. The engine rule and the store economics are the same lever.
Apple's stated rationale is security, privacy, performance, and battery life: a browser engine is exposed to hostile content and deep in the system, so Apple argues it must control the one engine on iOS. Open-web advocates answer that the secure, interoperable alternative already exists — it is the web and web apps — and that no other platform imposes such a ban. The European Union's Digital Markets Act sided with the advocates: since 2024 Apple has been required to permit alternative engines on iOS in the EU. Yet in practice almost none have shipped, because the terms attached — a separate EU-only app, locally based engineering, missing system APIs, restrictive contracts — have kept the door technically open and practically shut, which advocates and regulators argue is non-compliance by other means. Japan's smartphone law set its own deadline; the UK is investigating; Apple is appealing. The outcome is genuinely undecided.
The phrase to sit with is if it breaks. If real engine choice arrives on iOS — a true Blink or Gecko browser, with the full web platform behind it — then web apps could match native ones where they currently cannot, the case for the store as the only door weakens, and the thirty percent looks less inevitable. That is why this is fought so hard. A browser-engine rule sounds like an obscure technical restriction. It is one of the most consequential economic policies on the internet.
9.7 The monoculture problem
There is a second monopoly, and the open web is on the wrong side of this one too. Collapse every browser to its rendering engine and the field is tiny. As of 2026, roughly four in five web sessions run on Chromium's Blink engine — Chrome, Edge, Opera, Brave, Samsung Internet, and the rest share it. Most of the remainder runs on Apple's WebKit, largely because iOS forces it. Mozilla's Gecko, the last fully independent engine, is under three percent and has declined for years.
We have been here before, and the history is the warning. The first browser war ended around 2001 with Internet Explorer holding some ninety percent of the market — and then years of stagnation, because a dominant engine with no competition has little reason to improve, and the web ossified around one vendor's quirks and bugs. The standards-based revival that followed, and the engine diversity that drove it, is much of why the modern web exists at all. A monoculture is fragile and slow by nature: when one engine defines the web, “works in Chrome” quietly replaces “follows the standard,” the standard becomes whatever that engine does, and the power to evolve the platform concentrates in one company's roadmap. That an independent engine like Gecko is endangered is not a Mozilla problem; it is a problem for everyone who wants the web to remain a commons rather than a product.
9.8 Two comfortable fallacies
The app-versus-web hierarchy rests on two widely held beliefs that do not survive contact with this series. Both are worth dismantling carefully, because both contain a grain of truth.
“Native apps are safer and more private.” Often the reverse is true. A web page runs inside the sandbox and origin isolation of Modules 1 and 7: no installation, no standing access to your files or other apps, hardware reached only through the prompted, revocable gate of Figure 9.1, and no durable device identifier by default. A native app is installed software with broad device reach, frequently bundling third-party tracking SDKs the user never sees, behind permissions granted once and rarely revisited. App-store review and platform hardening are real and do catch things — the grain of truth — but the intuition that the sandboxed, ephemeral, per-origin web is the looser security model has the architecture backwards. The web's containment is stronger precisely because it assumes every page is hostile.
“Apps are faster because the web is slow tech.” Mostly the gap is not the rendering engine; it is the network and the packaging. A native app ships its assets pre-installed and loads them from local storage; a web page has historically fetched them over the network on demand (Module 4). But that is exactly the gap a service worker and the cache close (Module 8): a well-built PWA loads from local storage too, and the render pipeline of Module 5 is the same pipeline either way — on iOS, often literally the same WebKit. Where native genuinely wins is heavy sustained computation, low-latency graphics, and the deepest hardware integration — real and fair advantages. But the everyday “the app just feels faster” is usually an engineering and caching story, sharpened by an artificial capability ceiling, not a verdict on the web as a technology.
The web is not the insecure, slow option that the app makes safe and fast. That hierarchy is mostly a story told by whoever owns the store.
9.9 The browser is everywhere, often in disguise
The final twist undoes the binary completely: a great many of the “native apps” on your machine are browsers, wearing a costume. The browser engine turned out to be such a capable application runtime that it has been embedded nearly everywhere.
The editor a developer writes in, the chat client a team lives in, the music app, the desktop client of the social network — a remarkable share of them are Chromium in a frame, web technology shipped as a download. Capacitor and Cordova do the same on mobile, wrapping a web app in just enough native shell to enter the App Store. So the app-versus-web binary collapses twice over: web apps can do what apps do, and a huge fraction of apps are web apps already, paying the store's toll for the costume rather than the substance. The browser did not lose to native. It is inside native, doing the work, frequently unacknowledged.
9.10 The cast, and what the platform owes them
This series began with a cast — the people and programs on the other end of every request — and they are the reason the politics matter as much as the pipeline. A platform controlled by one engine and one store is not just an economic problem; it is a problem for everyone the cast represents.
9.11 The series ends here
Nine modules, one claim: the browser is the operating system of the web, and every web technology is a program running on it. We traced a URL through the I/O subsystem to the first byte; watched a parser turn bytes into the DOM; followed a page's dependency tree through the loader and its render-blocking choices; ran the pixel pipeline of style, layout, paint, and composite; met the single-threaded event loop that schedules all of it against a sixteen-millisecond budget; walked the origin-based security model that decides who may touch what; opened the storage and caching layers that make the browser a filesystem; and arrived here, at the platform and the contest over who controls it. The operating-system frame held the whole way down, because it is not a metaphor. It is what the thing is.
Where this series dissected the runtime, its companions build on top of it. HTML for Programmers is the language the parser consumes; CSS for Programmers is the system the render pipeline executes; The Web for Programmers is the protocols and the history beneath all of it. The platform-first ethos they share has a sharper edge now that you have seen the machine: writing for the platform rather than around it is not only cleaner engineering, it is a small vote for the web staying a platform anyone can build on without a gatekeeper's permission. You now know what the browser is doing when it runs what you wrote. Build for it accordingly.
The browser is the operating system of the web — the most widely used, least understood, and most contested platform we have. Understanding it is the beginning of defending it.