An APIs series
The DOM
Every tag you type becomes an object you can script. That is the whole idea: the Document Object Model is a bidirectional map between markup and a live tree of objects, so anything you can write in HTML and CSS by hand, you can read and rewrite with code. Frameworks wrap it, rename it, and diff against it — but underneath, every one of them ends up calling the methods in this series.
Step 01
Hello, DOM
Type a tag, get an object. Every <p> you write becomes an HTMLParagraphElement the browser builds for you, and the map runs both ways: change the object and the page changes, change the page and the object reflects it. The DOM is that map — which means if you can type it in HTML and CSS, you can code it. This step is first contact: find the object behind a tag, look at it, and change it.
Learning Objectives
- Describe the DOM as a bidirectional map between markup and live objects
- Explain what the parse tree is made of: element, text, and comment nodes
- Select an element by id and by CSS selector from
document - Inspect a DOM object with
console.dir()and mutate it withtextContent
The tree is the key
The browser does not keep your markup as text. It parses it into a tree, and hands you an API over that tree. Every individual item in it is a node: elements are nodes, text is a node, comments are nodes — and, the classic surprise, so is the whitespace between your tags. Given this markup:
the tree under the <p> has three children: the text node "The DOM is ", the <em> element (with its own text child "very"), and the text node " powerful!". Elements you indent across several lines pick up extra whitespace-only text children the same way. Nothing about the DOM makes sense until you see the tree this way, and step 3 will make the surprise concrete. For the full story of how the browser gets from bytes to tree, see the browser series.
First contact
document is your entry point into the tree. The two selection calls you will use most, and one mutation, are the whole demo:
console.dir() matters here more than it ever will again: it shows the element as the object it is — hundreds of properties, most of them mapped straight from things you already know how to type in markup. That view is the bidirectional map made visible.
Live demo
Open your DevTools console first — the two selection buttons deliberately log with console.dir() so you can expand the real object and scroll its properties. Then change the heading's text and note that no “render” call follows: mutating the object is updating the page.
Next Steps
In Step 2: Selection, we work through the whole selection toolbox — the one move every DOM program starts with:
- Climb the method ladder from legacy named access to
querySelectorAll - See a live
HTMLCollectionchange under you while a staticNodeListholds still - Use one selector language for both CSS and JavaScript
- Experiment freely in a selector playground
Step 02
Selection
Nearly everything you will ever do with the DOM is one move, repeated:
Select the node(s) of interest, manipulate them, and repeat, until the page is in its new state.
This step is the first half of that move. The API has accumulated selection methods for thirty years, so we climb them as a ladder — oldest to newest — and end on the two you will actually reach for.
Learning Objectives
- Recognize legacy named access (
document.formsand friends) in older code - Select by id with
getElementById, the long-standing workhorse - Distinguish a live
HTMLCollectionfrom a staticNodeList - Use any CSS selector via
querySelector/querySelectorAll, scoped to any element
The method ladder
Legacy named access
The oldest layer: the document exposes collections like forms and images, indexable by name. It still works, and you should recognize it in code you inherit — but do not build on it. Names collide, and the collections cover only a few element types.
getElementById
The workhorse since DOM Level 1, and still the fastest, most direct call when you have named a thing with an id. One element or null — no collection to unwrap.
getElementsBy* — and the live-collection trap
These return an HTMLCollection, and it is live: it is a view of the tree, not a copy. Add a matching element and the collection you are holding grows; remove elements while looping over the collection and items shift under your index. That behavior has produced decades of off-by-one loop bugs.
querySelector / querySelectorAll
The modern pair takes any CSS selector — the same language you already use in stylesheets. Learn selectors once, use them everywhere: that is the payoff of the whole select-and-effect design. querySelectorAll returns a static NodeList, a snapshot that holds still no matter what happens to the tree afterward. And both methods run on any element, not just document, so you can scope a query to the subtree you care about instead of searching the whole page.
Live demo
One shared sample DOM — three forms, three paragraphs, some nested <strong> elements — and every rung of the ladder run against it. Whatever a button selects flashes highlighted for two seconds, so you can see the set each method returns, not just its length. Finish in section 5, the selector playground: type any CSS selector, press Enter, and see what matches — including what happens when the selector is invalid (the call throws, and the demo catches it).
Next Steps
Querying is not the only way to reach a node. In Step 3: Walking the tree, we move by relationship instead of by selector:
- Walk with
children,parentNode, and the sibling properties - See why every relationship comes in a node flavor and an element flavor
- Walk upward with
closest() - Traverse an entire subtree depth-first
Step 03
Walking the tree
A selector jumps you straight to a node. But once you are standing on one, the tree offers a second way to move: by relationship. Every node knows its parent, its children, and its neighbors on either side, and those connections are properties you can follow. Walking is how you make relative moves — and the properties come in two flavors, because of those whitespace text nodes from step 1.
Learning Objectives
- Move between nodes with parent, child, and sibling properties
- Explain why
firstChildandfirstElementChilddiffer, and when it bites - Walk upward to a matching ancestor with
closest() - Traverse a subtree depth-first with
firstChild/nextSibling
Two flavors of every relationship
Indent your markup like a civilized person and every element grows whitespace-only text children. So the tree API splits each relationship in two: the node properties see everything — elements, text, comments — and the element properties skip to tags only.
When you are working with page structure — which is most of the time — use the element flavor and the whitespace never troubles you. Reach for the node flavor only when text nodes are the point, and then check nodeType: 1 is an element, 3 is text, 8 is a comment.
Walking up, and walking everything
parentNode climbs one level. closest() climbs with a selector — from the current element upward, returning the first ancestor (or self) that matches:
That one call is the backbone of event delegation in the Events series: from whatever small thing was clicked, walk up to the component that owns it. And with just firstChild and nextSibling you can visit an entire subtree — the depth-first walk every serializer, sanitizer, and framework diff is built on:
Live demo
A small nested tree — a list inside a div inside a div — with each traversal one button press. Whatever a walk reaches gets marked in the sample, so you can check your prediction against the tree. Pay attention to section 2: firstChild versus firstElementChild on the same element, and the childNodes dump where the whitespace text nodes finally show themselves. Then run the depth-first walk and read the indented output against the markup.
Walk or query?
Both reach nodes; they answer different questions. A query answers “find me the nodes matching this description, wherever they are.” A walk answers “from the node I am holding, give me its neighbor.” Inside a delegated event handler, holding the element that was clicked, closest() and children are the natural moves — re-querying the whole document to find where you already are is the roundabout version.
Next Steps
You can reach any node in the tree. In Step 4: Attributes and mappings, we start changing what we find:
- Read and write attributes with
getAttribute/setAttribute - Meet the property mapping — and the renamed reserved words like
classNameandhtmlFor - Flip visual state with
classList - Attach your own data with
data-*and read it back throughdataset
Step 04
Attributes and mappings
The DOM's founding promise is that markup maps to objects — and the fine print is that the map runs in both directions. Every attribute you type becomes a property you can read, and every property you set changes the state the markup declared. This step is about that mapping: where the names shift, where strings become types, and where the mapping crosses over into CSS.
Learning Objectives
- Distinguish
getAttribute()/setAttribute()from the mapped property, and know which one holds the live state - Predict the property name for any attribute: camelCase for multiword, remapped for reserved words (
className,htmlFor) - Manage classes with
classListinstead of string-mashingclassName - Attach your own metadata with
data-*attributes and read it back throughdataset - Bridge into CSS: computed styles, custom properties, and
matchMedia
Two doors to the same data
Take a checkbox written in markup:
The checked attribute is a piece of text in the document — getAttribute('checked') returns a string, and always will, because attributes are inherently strings. The checked property on the mapped object is a boolean, and it is the live one: it changes when the user clicks, while the attribute keeps recording only the initial state. The demos in this series always read the property side:
And because the map is bound both ways, assignment works too — set $('#fCheck').checked = false and the box unticks on screen. No render call, no sync step. The object is the element.
The property names follow two rules worth memorizing. Multiword attributes camelCase: tabindex becomes el.tabIndex, maxlength becomes el.maxLength. And attributes whose names were already reserved words in JavaScript got remapped: class becomes el.className, and a label's for becomes el.htmlFor.
Live demo: the mapping at work in forms
Forms are where the mapping earns its keep — and where it has the most conveniences layered on. The demo reads and writes typed properties, walks form.elements, accesses a field by its name with elements['title'], and finishes with the constraint-validation API: checkValidity() asks quietly, reportValidity() asks out loud with the browser's own bubble, and validationMessage tells you what it said.
The class attribute, done right
className is the whole class attribute as one string, which makes adding or removing a single class an exercise in string surgery. classList is the same data as a set, with the verbs you actually want:
Why does flipping one class name matter so much? Because a class is the handle CSS already holds. Change the class and every rule written against it applies at once — colors, weights, animations — without JavaScript touching a single style property. Design stays in the stylesheet; script just moves the state.
Crossing into CSS
The mapping has a style-facing side too, and it comes with a trap: el.style only reflects inline styles. The truth — after the cascade, the classes, and the media queries have all voted — comes from getComputedStyle():
Custom properties are the cleanest JS-to-CSS channel going the other way: set one variable at the root and every rule using it updates. The demo's fourth section closes the loop with matchMedia, which hands you media queries as objects that fire change events — the same event model from the Events series, now reporting on dark mode and viewport width.
One deliberate lesson hides in the demo's third button: replace('fancy', 'pretty') swaps in a class that no stylesheet defines, and the styling simply vanishes. Class names are only handles — they do nothing until CSS gives them meaning.
Applied: data attributes
The data-* family is the sanctioned place for your own metadata — ids, roles, flags — readable in script as el.dataset.role and in CSS as [data-role="developer"]. The mini-demo below plays both sides: dataset reads and writes the values, while attribute selectors like [data-theme="dark"] restyle the box the moment the value changes. One attribute, one source of truth, two consumers.
Next Steps
In Step 5: Modification, we move from state to substance — changing what an element contains:
- Set text safely with
textContent, and know exactly wheninnerHTMLis a hazard - Swap nodes with
replaceWith - Discover that appending an existing node moves it
- Redact a document with a single class flip
Step 05
Modification
Selection found the node; now change it. Modification is the most common DOM work there is — new text, a different link target, a flipped class, a swapped list item — and nearly all of it comes down to choosing the right one of a handful of tools. The choice matters more than it looks: one of these tools parses whatever you hand it as live markup.
Learning Objectives
- Change content with
textContent, and state precisely why it is the safe default - Explain the injection hazard
innerHTMLcarries with untrusted input - Add, toggle, and remove attributes; write metadata through
dataset - Swap a node with
replaceWithand move one withprepend
Text or markup — decide on purpose
Both of these replace an element's contents. They are not interchangeable:
textContent treats the string as text, tags and all — the <strong> above appears on screen as angle brackets. innerHTML hands the string to the HTML parser, which builds real nodes from it. That is exactly what you want for markup you wrote, and exactly what an attacker wants for markup they wrote.
Attributes, classes, styles
The state-changing tools from step 4 are half of modification in practice. The demo exercises all of them against a single link and paragraph:
The ranking from last step holds under modification: reach for the class first and let CSS carry the design; write style directly only for values computed at runtime — and even then, consider setting a CSS custom property instead, as the demo's third button does.
Structural surgery
Content and state changes leave the tree's shape alone. Two more calls change the shape itself:
The second line is the one that surprises people. There is no copy: a node lives in exactly one place in the tree, so inserting a node that is already attached moves it. Watch item C jump to the front of the list in the demo — one call, no clone, no removal step.
Live demo
All four families in one page: text versus HTML, attributes and dataset, classes and styles, and the structural moves. Every action reports to its own live-region log.
Applied: redacted text
A whole feature in one class flip. The stylesheet defines what redaction looks like — .redacted .blk paints marked spans black-on-black — and the script's entire job is classList.add('redacted') and classList.remove('redacted') on the container. This is the flip-a-class principle at full strength: design in CSS, one bit of state in JS.
Next Steps
In Step 6: Creation, we stop editing what the parser built and start growing the tree ourselves:
- Build elements the long way with
createElementandtextContent - Batch a hundred insertions into one with
DocumentFragment - Stamp out repeating structure from a
<template>withcloneNode - Type markup into place with
insertAdjacentHTML
Step 06
Creation
So far every node existed because the parser read it out of your markup. Now the script grows the tree itself. There is a long way and several shortcuts, and this step takes them in that order on purpose: the long way teaches you what a node actually is, and then each shortcut earns its place by solving a problem the long way makes obvious.
Learning Objectives
- Build an element with
createElement, fill it safely, and attach it withappend - Batch many insertions into one live-tree touch with
DocumentFragment - Stamp repeating structure from a
<template>withcloneNode(true) - Place typed markup precisely with
insertAdjacentHTMLand its position keywords
The long way, first
Everything in the demo's board goes through one factory function, and it is worth reading slowly because there is no magic left in it:
Create, configure, connect. Each createElement call makes a detached node — a real element that is simply not in the tree yet — and append wires the pieces together. Note where the user-facing strings go in: textContent, so the values are data, never markup. Verbose? Yes. But every shortcut below is sugar over exactly these moves, and when a shortcut misbehaves, this is the level you debug at.
Batching: one touch, not one hundred
Appending into the live tree makes the browser account for the change — do it in a loop and you pay per iteration. A DocumentFragment is a weightless off-tree container: fill it at leisure, insert it once, and the fragment dissolves, leaving its children behind.
Templates: structure in markup, values in script
When the structure repeats, writing it in JavaScript buries markup in code. The <template> element flips that: inert HTML that the browser parses but neither renders nor runs, waiting to be cloned.
Structure lives where structure belongs, and script only fills in the blanks — through textContent, so user input stays inert. Notice too how the cards' Remove buttons work: one delegated listener on the board, using e.target.closest('button[data-action="remove"]'). That is event delegation paying off exactly as promised — cards created a moment ago are handled by a listener attached before they existed.
If you can type it, you can make it
Sometimes the honest shortcut is to write markup as a string. insertAdjacentHTML is the precise version of that idea: it parses a string and places the result at one of four positions relative to an element — beforebegin, afterbegin, beforeend, afterend.
Live demo
All three techniques against one card board: single cards the long way, a hundred at once through a fragment, template stamping with your own title and body, and insertAdjacent* placing paragraphs around the board itself.
Applied
Spoiler toggle. A disclosure widget from almost nothing: content starts with the hidden attribute, and one button flips hidden and mirrors the state into aria-expanded so assistive tech hears what sighted users see. Before building this yourself, ask whether native <details>/<summary> already does the job — the DOM version is for when you need control the native element doesn't give.
Roster builder. Three series' worth of APIs in thirty lines: a form submit event, FormData to read the fields, a template clone filled via textContent, and append onto the roster grid. This is the shape of most small DOM applications you will ever write.
Next Steps
In Step 7: Destruction, the tree shrinks:
- Remove nodes with
remove()and empty containers wholesale - See why deletion loops run bottom-up
- Clean up listeners with
AbortControllerso removed nodes don't leak
Step 07
Destruction
Half of putting a page in a new state is taking the old state away. Removal looks like the easy third of the add–edit–delete trio, and the API is genuinely small — but it hides the two classic traps of DOM programming: deleting from a collection while you iterate it, and removing a node while something else still holds a reference to it. This step covers the small API and both traps.
Learning Objectives
- Remove elements with
remove()and read the olderremoveChild()pattern in existing code - Swap nodes with
replaceWith()and empty containers efficiently - Explain why deletion loops over live collections run bottom-up
- Clean up event listeners — by delegation, and by
AbortController
Taking a node out
The modern call is the obvious one: the element removes itself.
The original DOM had no such method. Removal was a parental act: you found the parent and asked it to disown the child, and the call handed the removed node back in case you wanted to keep it.
You will read removeChild in nearly every codebase older than a few years, so know it — but write remove() and its sibling replaceWith(), which put the verb on the node you are actually acting on. A removed node is not destroyed, by the way: if you kept a reference, it is merely detached, and you can append it somewhere else. Moving a node is removing and re-inserting it.
Emptying a container
To clear a list wholesale, the blunt instrument is fine, and it is what the demo's Clear button does:
If you loop instead — say, removing only some children — remember what step 2 taught about live collections: children re-indexes itself as you delete, so a forward loop skips every other element. The classic fix is to run the loop from the end, where deletions cannot shift what you have not visited yet:
Live demo
Add items, remove them one at a time, clear the lot. Note that the per-item Remove buttons carry no listeners of their own — one delegated listener on the container (step 4 of the Events series) handles every button that will ever exist in it:
Destruction is more than nodes
Removing an element does not remove what points at it. A listener registered on window that closes over your element, an interval that touches it, a Map that keys on it — any of these keeps the “removed” node alive and the work firing. This is the leak everyone writes at least once. Two habits prevent it. Delegation is the first: listeners live on the container, so removing children removes nothing that needs cleanup. The second is AbortController, which the demo's second section uses — pass a signal when you listen, and one abort() detaches everything registered with it:
Next Steps
You can now select, walk, read, write, create, and destroy — the full manipulation vocabulary. In Step 8: Geometry and observers, we ask the tree where things actually are:
- Measure with
getBoundingClientRect()and theoffset/client/scrollfamilies - Understand what reading layout costs
- Let
IntersectionObserveranswer “is it visible?” without polling
Step 08
Geometry and observers
So far the tree has been structure: what exists and what it contains. But every element also occupies space, and the DOM will tell you exactly where and how much — several ways, each measuring a subtly different box. This step covers the measuring APIs, what asking for a measurement costs, and the modern inversion: instead of measuring on every scroll event to find out whether something is visible, ask the browser to tell you when it becomes so.
Learning Objectives
- Measure an element with
getBoundingClientRect()and say what its numbers are relative to - Distinguish the
offset*,client*, andscroll*property families - Explain why reading layout can force the browser to do work, and why reads and writes should not interleave
- Use
IntersectionObserverfor lazy loading and infinite lists
The rectangle
getBoundingClientRect() answers the question you most often mean: where is this element on my screen right now? Its numbers are relative to the viewport — scroll the page and top changes — they are fractional, and they reflect CSS transforms. It is the rendered truth, after everything the engine has done to the box.
The three families
The older measurement properties come in threes, and the prefixes are the whole lesson — each family measures a different box:
offsetWidth/offsetHeight— the border box: content, padding, and border. The element's full footprint, as an integer.clientWidth/clientHeight— the padding box: content plus padding, minus borders and minus any scrollbar. The space available inside.scrollWidth/scrollHeight— the size the content would need: everything in the element, including the parts scrolled out of view. WhenscrollHeightexceedsclientHeight, there is something to scroll to.scrollLeft/scrollTop— how far it currently is scrolled. These two are writable: assign them to scroll programmatically.
The demo's second section puts all eight on one scrollable container whose child is deliberately too big for it — scroll around and watch which numbers move (only scrollLeft/scrollTop) and which are facts about the boxes (all the rest).
The inversion: observers
The traditional way to know whether an element is visible was to listen to scroll — a firehose event, as Events step 7 showed — and measure rectangles in the handler: the expensive read, performed at the highest possible frequency. IntersectionObserver is the platform's answer to that whole category of code. You declare what you care about; the browser, which already knows where everything is, calls you only when the answer changes:
The demo uses that for simulated lazy loading — twelve placeholder cards that “load” as they approach the viewport — and a second observer for an infinite list: a sentinel element sits at the bottom of a scroll container, and whenever the sentinel comes into view, more items are appended and the sentinel moves back to the end. No scroll handler anywhere in the file.
The frame below is deliberately shorter than the demo — the point is the scrolling. Scroll inside it and watch the counter as cards load.
Applied: the ellipsis that isn't DOM at all
A closing calibration. To truncate overflowing text with an ellipsis, you could measure widths and slice strings — or you could let CSS text-overflow: ellipsis do it, and use the DOM only for what CSS cannot: the width control that drives the demo, and a title attribute holding the full text. Drag the number and the platform does the truncation.
That division of labor — declare what you can, script only the remainder — is the recurring theme of this series, and the right instinct to carry into the finale.
Next Steps
In Step 9: Beyond the raw DOM, the honest closing argument:
- What jQuery solved, and why its wins moved into the platform
- Why the raw DOM is the performance floor a library cannot beat
- What frameworks actually buy — and what they cost
Step 09
Beyond the raw DOM
Eight steps of raw DOM invite an obvious question: why does almost nobody write applications this way? The answer is worth getting exactly right, because the popular versions of it are wrong in both directions — the DOM is neither too hard to use directly nor made faster by wrapping it. This closing step is the honest accounting: what the libraries over the DOM actually solved, what they solve now, and how to decide when you genuinely need one.
Learning Objectives
- Translate between jQuery idioms and their modern DOM equivalents
- Explain why the raw DOM is the performance floor, and what a virtual DOM actually is
- Distinguish developer experience (DX) from user experience (UX) in framework claims
- State a defensible rule for when to adopt a library
jQuery: the library that won so hard it disappeared
In 2006 the DOM you have been learning did not reliably exist. Browsers disagreed on event binding, on selection, on almost everything, and DOM code meant writing each operation two or three ways. jQuery wrapped the mess in one terse, chainable API — $('.fancy').on('click', fn) — and became the most successful JavaScript library ever shipped. Its ideas were so good the platform absorbed them: querySelector is jQuery's CSS-selector selection, standardized. Which is why the demo below is a jQuery comparison that loads no jQuery at all — every action is the vanilla equivalent of the jQuery one-liner shown beside it, at nearly the same length:
jQuery still runs a staggering share of the deployed web, and there is no shame in reading or maintaining it. But its reason for existing — papering over browser disagreement — is gone. Browsers converged; the holes it plugged are plugged by the platform.
The performance floor
Every framework — React, Vue, Svelte, all of them — eventually updates the page by calling the same methods you have spent eight steps learning. There is no other door into the renderer. It follows that an abstraction over the DOM cannot be faster than the DOM: used properly, the raw API is the performance floor, and every layer above it adds bytes to download, parse, and execute.
So what is a “virtual DOM”? A diffing strategy, not a faster tree. The framework keeps a model of the tree in memory, lets your components update the model freely, computes the difference from last time, and applies the difference to the real DOM in one batch. If your own code would have thrashed layout — dozens of components each poking the tree, interleaving the reads and writes step 8 warned about, inside a 16.7ms frame budget — the batching can rescue you. That is real, and it is also not speed. It is protection from a mess, purchased with overhead.
So — framework or not?
Beware of solving problems you do not have. Coordinating thousands of components across a fifty-person team is a real problem — Facebook has it; your four-page course portfolio does not, and a framework there buys complexity, worse startup, and too often the accessibility failures of developers who skipped the markup underneath. Meanwhile ecosystem, hiring, and existing code are legitimately good reasons to adopt one, and none of them are performance.
So the rule this course teaches: use a library once you understand it, understand the problems it solves, and actually have those problems. You will use frameworks in your career — that is close to certain, and it is fine. Learning the platform first is not an argument against them. It is what makes you good at them: able to tell what the abstraction is doing, able to fix it when it leaks, and able to notice when it is not needed at all.
Next Steps
The whole series has been one move, practiced until it is reflex: select the nodes of interest, manipulate them, and repeat, until the page is in its new state. Everything from getElementById to IntersectionObserver was a refinement of that move — and every framework you meet from here on is someone else's opinion about how to organize it.
- If you arrived here directly, the Events series is this series' other half — nothing above runs without it
- The natural next series is Web Components, where disciplined DOM plus custom events become the platform's own component model — the framework ideas, standardized
