CSE 134B Web Client Languages
  1. Structure
  2. HTML
  3. Principles
  4. HTML as a language

Lesson 01 · Principles

HTML as a language

HTML is "the Latin of the web." Learn it well and everything built on it (CSS, the DOM, accessibility, every framework you will ever use) starts to make more sense. Most developers treat it as the warm-up before the real work begins. This lesson treats it as what it is: a language, with a grammar, a history, a parser that never refuses your input, and a tree that is the thing that actually runs.

Learning objectives

  • Explain what markup is, and what a markup language has to define
  • Place HTML among the web's three foundations: URL, HTTP and HTML
  • Describe the contract between a strict author and a lenient parser, and why only one half of it is enforced
  • Read the tree the parser builds, and tell it apart from the text you typed

What markup is

Markup is information added to a text to describe its structure, and sometimes its format. A marked-up document therefore holds two things: the data, and information about the data. That second layer is everywhere once you look for it. A Markdown file wraps a word in asterisks to make it bold. A Rich Text Format file saved from a word processor buries one short sentence under lines of control words for fonts, colours and margins. An editor's red pen uses proofreader's marks: a caret for insert, a triple underline for capitalize. Nearly every file you save carries markup; the editor just hides it.

A markup language is a formal agreement about that layer. To be a language rather than a habit, it has to specify four things:

  • what markup is allowed, and where;
  • what markup is required;
  • how markup is told apart from the text it describes;
  • what each piece of markup means.

HTML answers all four. Tags are told apart from text by angle brackets; the specification lists which elements exist, where each may appear, which are required, and what each means. Most people say "tag" for everything, but the distinction is worth having: a tag is the bracketed text you type, and an element is what the tags delimit, the thing that ends up in the browser's tree.

Markup languages tend to fall into two camps. Presentational markup says how something looks. Semantic markup says what something means. HTML has carried both for most of its life, and the tension between them explains much of its history. The smallest example of the split is a pair of elements that look identical in every browser:

<b>Use me, I am bold</b> <strong>No, use me, I am strong</strong>

Both render in bold. The difference shows up the moment you apply a look of your own. Restyle strong as italic and the rule reads sensibly: important text is now shown in italics. Restyle b as italic and the stylesheet now says "make bold things not bold," which is nonsense. <strong> names a meaning, strong importance, and leaves the look to CSS. <b> started life naming a look, and the look is the one thing CSS will want to change. (Modern HTML has since given <b> a narrow meaning of its own, text drawn attention to without extra importance; lesson 6 covers the text-level elements.) The habit to build now is to choose the element for what the content is, then style it.

The course pairs "the Latin of the web" with a second motto: HTML is "the English of the web," widely spoken but not precisely, and understood anyway. This series is about the gap between the two.

The orchestration layer

The web rests on three separate specifications, all of them created by Tim Berners-Lee between 1989 and 1991, all of them still doing their original jobs:

  1. URLs name a resource, anywhere, with one uniform scheme (the URLs section takes them apart).
  2. HTTP moves it: a request for a URL and a response carrying the bytes (the film Three Things In, Three Things Out walks through one message).
  3. HTML expresses it, and the pages it builds are full of more URLs.

The three are separable on purpose, so each could evolve on its own, but they are tightly interlocked. An HTML document is text containing URLs; each URL is an instruction to make an HTTP request; each response may be more HTML. The href, src, srcset and action attributes all carry URLs. The method on a form is an HTTP verb. The resources a page loads, the requests it triggers and the navigation it offers are all written in its markup. HTML is the layer that orchestrates the other two.

So HTML is not scaffolding for the "real" code. Even the most modern JavaScript application ends up producing HTML; CSS binds to it and JavaScript manipulates it. A chemist who said atoms do not matter would not stay a chemist for long. HTML is the very matter of the web, and the atoms are still atoms however many frameworks are stacked on them.

Programmers skip it anyway, and the web shows it. According to the HTTP Archive's Web Almanac 2024, 29% of the elements on the pages it measured are <div>s, an element that means nothing at all. The specification is blunt about this, asking authors to "view the div element as an element of last resort, for when no other element is suitable." Think of <div> as the "um" of HTML, the filler a speaker reaches for when the right word does not come. A sentence full of "um" still communicates, roughly, and a page full of <div>s still renders. Either way the listener does the work the speaker skipped, and some listeners cannot.

A short history in swings

HTML's present shape is the result of a pendulum. It has swung from strict to loose and back more than once, and each swing left something behind that you still write today.

Strict: SGML and the DTD. HTML descends from SGML, the Standard Generalized Markup Language, an ISO standard from 1986 that grew out of IBM's earlier GML. SGML is a metalanguage: you use it to define other languages. The definition lives in a Document Type Definition, or DTD, which declares the allowed elements, their attributes and how they may nest. In object-oriented terms the DTD is a class and each document an instance that conforms to it. Checking that a document conforms is called validation. Berners-Lee's first HTML borrowed SGML's look, angle brackets and all, but was not written as a DTD; HTML 2.0 was the first version formally defined as an SGML application.

Loose: Mosaic and the browser wars. NCSA Mosaic arrived in 1993 and made the web easy to use, and it displayed images in the page alongside the text. Ease of use plus images started what the course slides call a "rapid descent into presentational thinking." Netscape commercialized the idea in 1994, Microsoft's Internet Explorer followed in 1995, and the browser wars began. Vendors shipped new elements faster than any standards body could write them down, many purely visual: <font>, <center>, <marquee>, <blink>. Neither <marquee> nor <blink> was ever standardized, but in January 1997 the W3C's HTML 3.2 accepted some of the rest, <font> and <center> among them, and for a decade HTML was mostly layout: tables nested in tables, and markup that said how a page looked and nothing about what it was.Pages of the late 1990s often carried a badge: "This page best viewed in Netscape Navigator," or in Internet Explorer. It was an honest admission that the author had written for one browser's dialect and given up on the others. Bolder sites read the User-Agent header and served each browser different markup. That is what a language without an agreed parser looks like in practice.

Strict again: the failed XML generation. HTML 4.0 (1997, revised as 4.01 in 1999) began pushing presentation back out to CSS. XHTML 1.0 (January 2000) went further and reformulated HTML as XML: every tag closed, every attribute quoted, everything lowercase. The discipline was sound; the delivery was not. A true XML parser must refuse to display a document with any well-formedness error, so in a page served as XML, one unclosed tag replaced the whole page with an error screen. Most sites that wrote XHTML served it as text/html, where the forgiving HTML parser read it, which defeated the point. The W3C's planned XHTML 2.0 broke compatibility with existing HTML altogether, and in 2004 Apple, Mozilla and Opera formed the WHATWG to evolve HTML instead. XHTML 2.0 was abandoned in 2009. Its most visible residue is the trailing slash in <br />. In a document served as text/html that slash is ignored: it neither helps nor hurts on a void element like <br>, and on <div /> it closes nothing, so the div stays open.

Loose, but defined: HTML5 and the living standard. The WHATWG's work became HTML5, which made a pragmatic choice: instead of demanding correct input, it specified exactly how a browser must parse any input, including broken input, into one defined tree. The course's author calls this consistent "soup parse" the single largest win of the HTML5 effort: predictability even for junk markup. HTML5 also made many tags optional again: <html>, <head> and <body> can be left out of a valid document, case does not matter, and many attribute values need no quotes. Not every HTML5 idea shipped. The specification once described a document outline algorithm that would rank headings by how deeply their sections were nested; browsers never implemented it, and the specification has since dropped it. HTML5 became a W3C Recommendation in October 2014, and in 2019 the W3C handed HTML to the WHATWG. There is now one HTML, the living standard, with no version number.

Strict, loose, strict, loose. The pendulum explains why you meet <center> in old code, <br /> in tutorials and missing <body> tags in perfectly valid pages. The story of where all three foundations came from, starting at CERN in 1989, is told in the film Vague but Exciting.

The asymmetric contract

The HTML specification gives authors and browsers two different sets of obligations.

  • Authors must write conforming documents. Every element in a permitted place, every required attribute present, every value drawn from its allowed grammar. The rules are precise.
  • Browsers must accept anything and build a tree from it, in exactly one defined way. There is no input for which the parser is allowed to fail.

This asymmetry keeps a 1995 page rendering today while giving new code a definition of correct. The catch is that nothing enforces the strict half for you. The lenient half runs in every browser; the strict half runs only when an author chooses to validate. A compiler refuses a broken program. A browser shows the broken page anyway, so the error surfaces later, as a CSS rule that does not match or a script that finds nothing, three layers away from its cause.

The parser is a pipeline: bytes, then tokens, then a tree. The bytes are decoded into characters; a tokenizer, a state machine with roughly eighty states, turns them into tokens (start tags, end tags, text, comments, a doctype); tree construction consumes the tokens, tracks which elements are open, and applies a recovery rule whenever a token arrives where it is not allowed. Every malformed input has a rule. An unknown element such as <foo> becomes an HTMLUnknownElement in the tree, and its content renders like any other text. A stray end tag is usually dropped. A missing structural element is supplied.

You can watch that last rule with five lines:

<!doctype html> <title>Five lines</title> <h1>A heading</h1> <p>A paragraph. <p>Another one. <!DOCTYPE html> <html> <head> <title>Five lines</title> </head> <body> <h1>A heading</h1> <p>A paragraph.</p> <p>Another one.</p> </body> </html>

You typed no <html>, <head> or <body> and closed neither paragraph. The parser invented all three elements, put the title in the head and the rest in the body, and closed the first paragraph when the second opened, exactly as the specification intends. Now take it to the extreme, a file with no tags at all:

Open bare-text.html on its own, then open DevTools and look at the Elements panel. There is a full document there, html, head and body, with your text in the body, built from a file that contains not one angle bracket.

Two words describe the strict side, and this series keeps them apart. A well-formed document follows the surface grammar: tags balance, quotes match, nothing overlaps. A valid document is well-formed and follows the rules for which element may appear where, with which attributes and values. <p><div>hello</div></p> is well-formed and invalid, because a paragraph may only contain phrasing content. The parser repairs it by closing the paragraph before the div, and the final </p> becomes a second, empty paragraph. Nothing tells you. A validator would. Lesson 4 teaches content models, validity and the validator properly; for now it is enough to know that "the browser displayed it" and "it is correct" are different claims.

The tree is the contract

Here is the idea the rest of this series stands on: the text you typed is not what runs. What runs is the tree the parser built from the text you typed. Your CSS selectors match against the tree. Your JavaScript queries the tree. The browser fetches the URLs it finds in the tree. The accessibility tree that screen readers use is computed from it. When the source and the tree disagree, every one of those readers believes the tree.

Two tools show the two sides. View Source shows the text, as the server sent it. The DevTools Elements panel shows the tree, live, after the parser's repairs and any script's changes; it is not a pretty-printed copy of your file. When a page does not behave the way its source says it should, look at the tree first.

The tree also contains things you never think of as content. Indentation is text. Put a paragraph on its own line inside the body and the newline and spaces before it become a text node of their own:

<body> <p>Hello <em>web</em></p>
The tree the parser buildsbody contains a whitespace text node and a p element; p contains the text "Hello " and an em element; em contains the text "web".body#text "\n "p#text "Hello "em › #text "web"
What runs is the tree, not the text. The indentation before <p> is a text node too; DevTools usually hides it, the DOM does not.

That whitespace node is why firstChild on the body returns text rather than the paragraph: a small example of the general rule to reason about the tree, not the source.

The repairs get larger when the markup is wrong. The demo below holds seven pieces of broken HTML, each shown as source and as rendered, with a button at the end that prints the DOM the browser actually built.

Open malformed.html on its own and look closely at two of the mistakes. Mistake 5 is a table written without a <tbody>. The parser inserts one, so the rows are not children of the table, and a selector such as table > tr matches nothing. Mistake 6 puts text directly inside a <table>, where text is not allowed. The parser fosters it out: the text is moved before the table in the tree, though the source puts it inside. Both repairs are defined and reasonable. Neither is what the author wrote, and neither produces a warning.

Why it matters now

If the parser always produces something, why care? Because of who is writing the markup now. The framework era industrialized <div> output: when you write components, the framework emits the HTML, and the default element in a great many components is a <div>. Generated markup is industrializing it further. Language models learned HTML from the web as it is, which means largely from that same <div>-heavy output, and absent careful constraints they write it back the same way, at enormous volume. The course slides put the risk plainly: "If we skip learning markup and move to programmatic markup development we often just automate nonsense creation." A generator you cannot judge is a generator you cannot correct.

And pixels were never the only output. The rendering engine is the most forgiving reader of your tree; the others extract meaning, and they extract it from the elements you chose. Screen readers and other assistive technology work from the accessibility tree, which is computed from the DOM, so a meaningless DOM produces a meaningless accessibility tree (Accessibility follows that thread through the whole course). Search engines read headings and landmarks to decide what a page is about. AI agents and reader modes walk the tree to find the content and the controls, and a <div> with a click handler does not announce itself as a button the way a <button> does.

Seen that way, choosing an element is API design. The elements you write are the public surface every one of those readers consumes. <nav> versus <aside> is an interface decision; so is <button> versus <a>. Later in the course you will define elements of your own, such as a <product-card> (custom element names must contain a hyphen), and the discipline is the same one this lesson starts: pick the element that says what the thing is.

Try it: X-ray vision

The slides for this course ask you to "hone your markup X-ray vision": to look at a page and see the structure under the pixels. Four exercises, each about ten minutes:

  1. Open bare-text.html. Use View Source, then DevTools › Elements. Write down every element in the Elements panel that is not in the source.
  2. Open malformed.html and press "Show the DOM the browser built." Find the <tbody> the parser inserted in Mistake 5 and the text it moved in Mistake 6. Then find both again in DevTools, without the button.
  3. Pick a real site you use every day. Before you open DevTools, label its regions on a screenshot (header, navigation, main content, sidebars, footer, headings) and name the element you would use for each. Then compare your guesses with the Elements panel, and count the <div>s standing in for something with a better name.
  4. Run the same site through the Nu HTML Checker at validator.w3.org/nu/ and count the errors. Pick one error and work out what the parser did with that markup instead.

The point of the third exercise is the order: structure first, markup second. When you can describe what a page is before you write a tag, choosing the tags becomes the easy part.

Next steps

Lesson 02, The document, opens the file itself: the skeleton, the doctype and the rendering mode it selects, character encoding, and the elements the parser inserts when you leave them out. Several HTML topics are taught elsewhere on the site, in the sections that own them: links, forms, images and the rest. The list below points at each one, and it fills in as those pages go live.