· 25 min read · 5,670 words
Last updated on

Dotcom Chronicles, Side Story 1: HTML Is Not SGML

SeriesPart 4 of 5 in The Dotcom Chronicles

At twelve minutes past midnight on 7 June 1992, Dan Connolly of Convex Computer Corporation in Texas sent a short message to the www-talk mailing list. The subject line read “HTML is not SMGL”, two of the four letters swapped. His “grandiose scheme” to convert HTML into MIME and SGML worked fine, he wrote; MIME was the e-mail attachment standard settled that year, which laid down how a message could carry several parts, each labelled with a type, and he wanted to use it to carry web pages. But when he went back to writing a DTD for the existing HTML format, he could not.1

SGML is the Standard Generalized Markup Language, an international standard since 1986, descended from GML, which IBM built in 1969. It is not a file format but a set of rules for defining file formats: tags in angle brackets mark the structure of a document, and a definition called a DTD says which elements this kind of document may contain and which may nest inside which, so that the program reading the file, the parser, can check a document against the DTD and say whether it is well formed. HTML borrowed the syntax, the angle brackets, the attributes, the & style of escape; the one thing it added was the anchor, the tag that points a stretch of text at another address.2 HTML’s trouble is that it has so little structure. Text can appear almost anywhere, and as soon as you try to write that down you run into what SGML calls mixed content, text and tags mingled at the same level, about which a DTD can say very little. Connolly ended with two questions: “How much extant HTML is really out there? And how much of it is generated on the fly by gateways and servers?”1

"HTML is not SMGL", 7 June 1992

The whole of this message as it appears in the Calgary archive. Text from www-talk 1992, message 80, re-rendered in a monospaced font; Received lines and empty headers omitted, two highlights added by this article, the misspelling in the subject line as archived. Not how it looked on a 1992 terminal. The message is by Dan Connolly and is excerpted here as a historical source.

I had assumed the story of HTML belonged to Berners-Lee alone: he wrote the tag list, he wrote the browser. Then I counted the four hundred-odd messages on www-talk in 1992 by sender, and found that the person who worried most about HTML that year was not him. Connolly sent eighty-five. He was not at CERN and he was not writing a browser; he built documentation tools at a supercomputer company and wanted FrameMaker, the software used for long technical manuals, to read and write HTML. From start to finish he asked one thing: should a one-page list of tags become a standard that somebody validates?3

“Why not just use RTF”

Why the Web needed a document language of its own is a question about the machines it faced. In the first article, CERN’s problem was that its machines did not talk to each other: the same document had to appear in a NeXT window with fonts, and scroll line by line on a terminal that knew only characters; SLAC’s minutes of February record the browser installed on a Unix cluster, on NeXTs, on the mainframe, and on a VMS machine. A file that had to be readable in all those places could not fix its type sizes and fonts. It could only say “this is a heading” and “this is a list,” and let each machine decide what a heading looked like there. Berners-Lee said as much in June 1992, comparing HTML with MIME’s rich text: HTML’s treatment of logical heading levels “has turned out to provide more flexible formatting on different platforms than explicit semi-references to font sizes.” The other advantage is that it is only text, so a database or a script can produce it; that is how the SPIRES search results in the second article turned into web pages. SGML was an existing convention built for exactly this kind of structural markup. Berners-Lee borrowed its syntax, and not all of its rules.4

In the second article, Addis made the SPIRES database write HTML on the strength of CERN’s one-page list of tags: TITLE, H1, P, A in angle brackets, and a browser that ignored any tag it did not recognize.5 It is no surprise that people at the time did not distinguish the two; SLAC’s minutes of February call the web pages SGML outright, and so does the file list Kunz handed over. That arrangement is friendly to the person writing pages and unfriendly to the person writing a second program. Connolly’s first long message, on 6 June, gave three reasons, and the first was this: a DTD was needed “so that we can parse HTML using something besides the public implementation of WWW, and so that we can verify documents converted from other authoring systems,” naming GNU info, Andrew’s EZ, and FrameMaker. As long as there is one program in the world that reads HTML, what HTML is gets decided by that program’s source code; the moment there is a second, both need a definition they agree on.1

CERN's HTML tag list, November 1992 snapshot

The opening part of CERN’s tag list page, Tags.html, from the snapshot of 3 November 1992 preserved by the W3C, rendered from the original markup in a modern browser and cut off after the Anchors section; the entries that follow, IsIndex, Plaintext, Listing, Paragraph, Headings, Address and others, are not shown. Fonts and link colours are today’s browser defaults, not the display of any 1992 browser; the original markup is kept verbatim except for the NEXTID and TITLE tags. © CERN / W3C, excerpted under the W3C Document License.

On 25 June Berners-Lee replied that he would like “a DTD which as closely reflects the current HTML as possible.” Connolly answered that one could be written, but he was “not sure of the value of it”: HTML “allows tags to be pretty much sprinkled wherever you feel like putting them,” and a DTD loose enough to allow that would just say that “every element is just a repeatable or-group of all the elements,” leaving the parser nothing to check. The current HTML and address syntax “make a good proof of concept,” he said, “but we need to move toward formal definitions so that we can have confidence that correct implementations will interoperate.” The next day Berners-Lee explained why HTML was so flat: many rich-text objects can hold styles but not structure, Word deduces an outline from heading styles, WWW deduces a list from a run of LI items; and the real information providers, “group secretaries, for example,” might not want to write nested elements, styles perhaps being the interface they were used to. “So that is why the HTML structure is so simple. I am open to a more sophisticated alternative.”6

On 14 July Connolly put it more bluntly. He had tried to write the DTD and concluded that “HTML has very little structure, and that this is by design.” The value of SGML is validation; a publisher can specify in a DTD the format of references and the placement of the abstract. “The WWW project has no such editorial policies to enforce.” Its policies amount to “you can have a title, if you want, and we’ll keep it visible for the user; you can have headings and paragraphs and glossaries and lists and menus, and as long as you use them in pretty much the traditional way, they’ll be formatted reasonably. And you can have anchors.” So why use SGML at all? His answer: because “the NeXT implementation has a nifty editor.” In that case, why not just use RTF, Microsoft’s rich text format for exchanging files between word processors, which was more mature, supported on the NeXT, the Mac, and the PC, and lacked only a few public renderers. “Unless we want some part of the WWW system to verify the structure of documents, why are we using SGML (and using it poorly)?”7

Berners-Lee replied in the small hours of the next day. The HTML generated by the NeXT editor was indeed not good, attribute values that needed quotes had none, but the current parser would parse real SGML. The higher-level markup was useful: the line mode browser, which knew only characters, used it to produce a style different from the X Window one, X being the graphical interface on Unix, and the LaTeX scripts that typeset the whole document tree into a “www book” used it too. HTML had no deep structure so as to stay compatible with software that could not handle nested elements. “There is nothing wrong with having a simple SGML DTD as a basic case. SGML does not HAVE to be complicated.” As for RTF, his doubts were that headings were faked with specially named styles, that the formatting was always tucked in beside the style name, and that each vendor’s extensions differed. Neither persuaded the other. That night Connolly posted the DTD anyway, with a perl script to “legalize” existing HTML files by putting quotes round attribute values.7

One line is enough to show what that DTD looked like: <!ELEMENT HTML O O ((TITLE? & NEXTID? & ISINDEX?), BODY, ADDRESS?)>. It says that an HTML document consists of these parts: a title, a NEXTID, and an ISINDEX, each optional and in any order; then a BODY, which is required; then an optional ADDRESS at the end. The two O’s mean that the opening and closing tags of HTML itself may be omitted and inferred by the parser. With a few dozen lines like this, a parser can hold a file up against the list and report a title that has strayed after the BODY or a list item that has fallen out of its list. His difficulty was that real HTML files mixed text and tags any old how, and a list that accommodated that freedom could only say “anything anywhere,” which tells a parser nothing. On 19 August he reported that FrameMaker could now open and save HTML directly. What he wanted was never a more complicated HTML. It was an HTML that a second program could read.8

“SGML Cop backs off”

The autumn began with a registration. In November someone proposed registering text/html as an official MIME type, giving web pages a recognized name in mail and transport systems, which meant that HTML needed a definition that could be shown to outsiders. Connolly’s message of 19 November was titled “Freezing the HTML spec”: Berners-Lee talked a lot about “HTML futures,” there were “a lot of open issues,” and “Eventually we have to ‘shoot the engineers and ship it,’ that is freeze the spec and hand it to the IETF.” He drew up a charter. First, “Establish a well-defined relationship between HTML and SGML. Don’t include anything in the HTML spec that conflicts with the SGML spec.” Second, use the line mode browser as the reference implementation. He also said that “the current method of sending suggestions to Tim and hoping he finds time to make the edits is no good.” The second half of the same message answered someone’s question about comments, and took SGML’s comments, processing instructions, and marked sections one by one; halfway through he wrote: “SGML is a mess!”9

Two days earlier he had already backed off. The subject of his message of 17 November was “SGML Cop backs off”; the cop was himself. He had insisted that every HTML file carry SGML’s “framing”: the declaration, the prologue, the instance, all of it, or else “it’s just not an SGML document.” After a closer reading of the standard, after receiving the DocBook materials from the publisher O’Reilly and the computer company HaL, a set of SGML document types they had defined for technical books, and after installing MidasWWW, the graphical browser written by Tony Johnson at SLAC, he changed his view: treat text/html as an SGML text entity rather than a document entity, with the DTD assumed in front of it, “like assuming every text/c-program gets stdlib.h prepended before compiling,” stdlib.h being a standard header file of C. The assumption had always been there. Only the description changed. He ended with an advertisement for MidasWWW: “It’s long overdue in the WWW project, but it’s worth the wait!”9

On 30 November he uploaded a package to CERN’s FTP server, html_spec-0.3, containing the root page of a specification, an introduction to SGML syntax, a DTD, “several example files that form a validation suite,” and a parsing library. He addressed people in turn: “Tim: please link this into the web somehow.” “Implementors: please grab the whole thing and validate your implementation against it.” “Tony: I’ve got some patches for the MidasWWW browser.” The next day he wrote to Berners-Lee alone: “I noticed you’ve been diddling with the HTML files on info.cern.ch — quoted your attributes, dotted your i’s and crossed your t’s, so to speak. But the files still don’t fit into SGML.” He asked him to fetch James Clark’s sgmls parser from a server in Norway and check his files with a single command, sgmls -s html.dtd yourfile.html; if there were errors, “either fix your software or diddle with the DTD until you get something that works.” The second half of the same letter was in another key. He had “wrestled quite a bit” with the header and body structure and could not come up with a DTD that “1) makes most existing HTML legal, and 2) imposes any structure on the thing.” “I’m just about to give up on the structure business,” he wrote; he might as well change the DTD “so that HTML is just ‘tag soup’ — anything goes anywhere.” Three days later he uploaded a new version, renamed a tag of his own invention back to the familiar PRE, the tag for text kept in its original layout, and added a new motto: “just describe it; don’t prescribe it.”9

Connolly to Berners-Lee, 1 December 1992

Excerpts from the letter of 1 December: the first half asks Berners-Lee to install sgmls, the second half says he is about to give up on structure. Text from www-talk 1992, message 390, re-rendered in a monospaced font; thirty-three lines of technical explanation about mixed content in the middle are marked as omitted in square brackets, four highlights added by this article. Not how it looked on a 1992 terminal. The message is by Dan Connolly and is excerpted here as a historical source.

In today’s terms, what he built in the autumn of 1992 was a validator for HTML and a set of test cases, plus a draft specification nobody had asked him to write. These things sat on CERN’s FTP server for anyone to install. How many did, the letters do not say.

“On whom rests the responsibility for validating”

On 21 January 1993 he sent a long message titled “thoughts on the future of HTML,” opening with a short history of HTML in his own words. HTML was designed to be simple; “Folks are supposed to be able to whack out HTML with a text editor.” But it was also designed to be processed by “lots of machines all over the planet.” “Enter SGML. It seemed like the natural choice, so Tim implemented an informal SGML parser in his WWW clients.” This is his account from 1993; the reasons Berners-Lee himself gave in 1992 were structure and an existing standard. The two do not conflict, but they are not the same. “Nobody really knew the ins and outs of SGML, so information providers who wanted to produce HTML automatically just checked to be sure the public www client grokked.” Then other people tried to write HTML parsers and found “a lot of issues that were not covered by any spec other than the WWW source code.” Then he tried to use sgmls to build an HTML-to-FrameMaker tool, and “discovered that the WWW source code conflicted with the SGML standard. Uh oh!” When the SLAC library in the second article had its database spit out HTML, it was working exactly the way he describes: if the browser displayed it, it was right.10

The core of this letter is one question: on whom rests the responsibility for validating HTML documents? It was really an HTTP issue, he said: was it part of the protocol that the data stream was valid HTML, or was it the client’s job to deal with errors? His answer was the server. Of course the client should be tolerant of errors, but when a client and a server disagreed about a document, “the client is at fault if the document is valid, and the server is at fault if the document is not.” That would make servers more complex; a server could “no longer just grab the contents of any old .html file and ship it out the port,” but it could fix markup errors on the fly and write them to a log. It was too late for HTTP 0.9, the transport protocol of the day, he admitted, “but future servers should have the burden of producing valid documents.” The rest of the letter set out his ideas for HTML2, more structure, paragraphs as countable units, and beyond that HyTime, a hypermedia standard built on SGML that had just become an international standard. That evening he uploaded a new version of his parsing library: “I gotta go. It’s a rush job.”10

"Who validates", 21 January 1993

The first half of the long letter of 21 January 1993: his own history of HTML, and the question of who validates. Text from www-talk 1993, first quarter, message 89, re-rendered in a monospaced font; the second half, on the structure of HTML2 and HyTime, is not included; three highlights added by this article. Not how it looked on a 1993 terminal. The message is by Dan Connolly and is excerpted here as a historical source.

After 30 January there is no more mail from him on the list. In August someone asked after him, and a technical publications manager at SCO, the Unix vendor, replied that he had been in touch a month before: “He is too busy in his new job to be involved with WWW now,” with a new address. Why he changed jobs, and where to, the letters do not say, and I do not guess.10

“Even if Dan C ain’t here to round us up”

In the fourth week after he left, at nine on the evening of 25 February, Marc Andreessen of the National Center for Supercomputing Applications (NCSA) at the University of Illinois proposed a new tag, IMG, with a SRC attribute naming an image file, which the browser would embed at the point where the tag occurred. X Mosaic was the graphical browser NCSA had released a month before, and Andreessen was one of its authors. “This is required functionality for X Mosaic; we have this working, and we’ll at least be using it internally.” He was “certainly open to suggestions as to how this should be handled within HTML,” and admitted the question of image formats was “hazy,” but saw no alternative to “let the browser do what it can” while waiting “for the perfect solution to come along (MIME, someday, maybe).” Two hours later Tony Johnson at SLAC wrote that MidasWWW 2.0 had “something very similar,” called ICON, with an extra NAME parameter that let the browser substitute a built-in image instead of fetching one; he did not much care about the names, “but it would be sensible if we used the same things.”11

Berners-Lee replied twice on the 26th. What he had imagined was not a new tag but two relationship values on the anchor, EMBED for embedding when the document was presented and PRESENT for opening the target whenever the source was opened, so that “if the browser doesn’t support either one, it doesn’t break”; “I hadn’t wanted a special tag.” The second letter was more direct: the reader rather than the author might want to choose which figures were expanded inline and which opened in a window, so an anchor with switches looked more appropriate; “I don’t want to change HTML now if I can help it, until it has gone to RFC track,” an RFC being one of the numbered documents in which the Internet Engineering Task Force (IETF) publishes its standards. Andreessen’s answer was “I absolutely agree in all cases,” and then: things had reached the point where “some browsers are going to be implementing this feature somehow, even if it’s not standard,” and “it would be great to have consistency from the beginning — so that when HTML2 comes along, we’re all still in lockstep.” On the 27th Berners-Lee wrote again: “Ok, so for HTML2 let’s have something for inclusion,” not limited to images; “SGML does provide an official way of doing this folks and even if Dan C ain’t here to round us up we maybe ought to stick to the track,” and he gave a line using SGML’s entity mechanism, which names an external piece of content and then refers to it. On 18 May Andreessen described the IMG extension in Mosaic 1.1 on the list; the tag was in the browser.11

The IMG proposal and reply, February 1993

Above, Andreessen’s letter of 25 February 1993 proposing IMG, with five lines omitted in the middle; below, Berners-Lee’s reply of 27 February in full. Text from www-talk 1993, first quarter, messages 174 and 194, re-rendered in a monospaced font; Received lines and empty headers omitted, two highlights added by this article. Not how it looked on a 1993 terminal. The messages are by Marc Andreessen and Tim Berners-Lee and are excerpted here as historical sources.

The summer’s argument had different people in it and the same question. On 14 August Andreessen set out five assertions. The first four said that a properly validated document should display correctly in every browser. The fifth said that if a browser supported features beyond the standard and many information providers “choose to take advantage of those features even at the expense of making it difficult or impossible for other browsers to properly handle their documents,” that was proof the feature deserved to become standard: “the market has chosen.” He signed off “donning my fire-proof bodysuit.”12

On the 18th Terry Allen, an editor in O’Reilly’s digital media group, wrote: “Marc, with respect, your browser puts out a lot of error messages,” yet it did not report markup errors. One of their editors, an experienced typographer, had produced headings that opened with H2 and closed with H3; because they looked fine in Mosaic, she had not bothered to parse the document. Even “Markup error in [URL]” would help, so that people could go back and find it with sgmls. Andreessen replied the same day: at O’Reilly people had told him Mosaic was wonderful, “it handles anything you throw at it without complaining,” and now someone else from O’Reilly was saying the opposite. “Folks, don’t complain to us because Mosaic is robust.” He put the reason in capitals: Mosaic did not report errors because “DOING SO WOULD REQUIRE WORK,” and there were more important things queued up; the other reason was that “Mosaic is a browsing environment, not an authoring tool.” He suggested O’Reilly put one of its people on writing an HTML-lint, a checker on the model of lint, the Unix tool that checks C code. The same day Lou Burnard of Oxford sent a message titled “who validates?”, and his answer was sgmls hooked into the Emacs editor: the writer, on their own.12

Connolly’s question of January got its answer in August, and the people answering had not necessarily read his letter. The server was not responsible, the browser was not responsible, and the writer went and installed a parser. Browsers today still display pages that are written wrong, and validators are still tools that writers run for themselves; this division of labour was settled in the summer of 1993, and the people who settled it never held a meeting. Nobody took on the responsibility for validation, so it fell to the browser’s tolerance of errors.

“Just describe it; don’t prescribe it”

In November 1995 HTML 2.0 became RFC 1866, on the standards track. Two authors: T. Berners-Lee and D. Connolly. The document opens by saying that it “brings together, clarifies, and formalizes a set of features that roughly corresponds to the capabilities of HTML in common use prior to June 1994,” and that it is called 2.0 to distinguish it from “the previous informal specifications”; the second paragraph states that HTML is an application of ISO 8879, that is, of SGML. The DTD he could not write in June 1992 is in there. How he came back to this in 1994, and what the IETF working group did in between, the material for this piece does not cover, and is left for later.13

From a short message to an RFC

The three and a half years this piece covers. Dates are taken from the letters and from RFC 1866; where several letters fall on one day, only the most important is marked. Drawn for this article from the www-talk archive and RFC 1866; not a contemporary chart. The IETF working group process between August 1993 and November 1995 was not verified for this article and is left blank.

The RFC’s stance is the one in his motto of 4 December 1992: just describe it, don’t prescribe it. Three years earlier he had wanted to freeze the specification, shoot the engineers and ship it, and make servers answer for validity; three years later the specification with his name on it describes what everyone was already doing.

Footnotes

  1. Dan Connolly, “HTML is not SMGL”, 7 June 1992, 00:12 CDT, www-talk, Calgary archive, 1992, message 80; the misspelling in the subject line is as archived. The previous day’s “MIME as a hypertext architecture” (message 78) gives three reasons for needing a DTD; this piece uses the first. The explanations of DTD and “mixed content” are this article’s, for the reader, and do not come from these two messages. 2 3

  2. The description of SGML follows SGML.html in the CERN hypertext snapshot of 3 November 1992 (signed Tim BL): “an ISO standardised derivative of an earlier IBM ‘GML’”, whose structure “can be checked for validity against a ‘Document Type Definition’, or DTD”. The year 1986 follows RFC 1866’s citation of “ISO Standard 8879:1986”. The origin of GML follows Charles F. Goldfarb, “The Roots of SGML — A Personal Recollection”, 1996, Wayback Machine copy: “Later in 1969, together with Ed Mosher and Ray Lorie, I invented Generalized Markup Language (GML)”, named in 1971 from their initials “so that our initials would always prove where it had originated”. A recollection. The same essay says the General Document type in ISO 8879 became the source of the Web’s HTML document type by way of Anders Berglund’s championing of DCF at CERN; not verified for this article, noted for the record.

  3. All mail used here is from the www-talk archive kept at the University of Calgary, 465 messages for 1992 and 3,071 for the four quarters of 1993, searched in full for this article. “Eighty-five” is the number of messages in the 1992 archive whose sender contains Connolly; it counts the archive only, not mail he sent elsewhere. That Connolly was at Convex Computer Corporation follows his address and what his letters say about his work; his biography was not otherwise researched. Sources are collected in www-talk 1992–1993 HTML 从标签清单到 DTD(Connolly 线索).

  4. The diversity of machines is in Dotcom 编年史 - 万维网从一个找资料的问题开始; the platforms at SLAC are in the minutes of 5 February 1992 cited in Dotcom 编年史 - 只盼它能在 SLAC 活下来. Berners-Lee’s comparison is from “MIME, SGML, UDIs, HTML and W3”, 11 June 1992 (message 96): “our treatment of logical heading levels and other structures is much more powerful and has turned out to provide more flexible formatting on different platforms than explicit semi-references to font sizes”. “Not all of its rules” is this article’s summary of the argument that follows.

  5. Tags.html and MarkUp.html in the CERN hypertext snapshot of 3 November 1992; the latter says “WWW parsers should ignore tags which they do not understand”. SLAC’s output of HTML is in Dotcom 编年史 - 只盼它能在 SLAC 活下来; the minutes of 5 February 1992 list the files as “C programs, SGML, and REXX execs”, see 1991–1992 SLAC WWW 工作笔记本(Addis).

  6. The exchange of 25 June between Berners-Lee and Connolly is in Connolly’s “Re: HTML DTD” (message 122), which quotes Berners-Lee’s “I’d like a DTD which as closely reflects the current HTML as possible”; Connolly’s reply includes “we need to move toward formal definitions so that we can have confidence that correct implementations will interoperate”. Berners-Lee’s explanation of 26 June about styles and “group secretaries” is in “Re: HTML DTD” (message 121), ending “So that is why the HTML structure is so simple. I am open to a more sophisticated alternative.” In the archive message 121 is dated 26 June and message 122 25 June, a matter of time zones and archive order; this piece follows the dates on the letters.

  7. Connolly, “rethinking the HTML DTD.”, 14 July 1992 (message 144); Berners-Lee’s reply of 15 July, 00:03 (message 147), with “SGML does not HAVE to be complicated”. The line mode and X Window styles and the LaTeX scripts are from the same reply. 2

  8. The ELEMENT line quoted in the text is verbatim from “HTML DTD enclosed”; the explanation of it is this article’s. “HTML DTD enclosed” and “perl script to legalize HTML files”, 15 July 1992, 22:35 and 22:49 (messages 152 and 153); the script is fix-html.pl and mainly rewrites anchor tags and quotes attribute values. “can now read/write html from FrameMaker!”, 19 August (message 179).

  9. “Freezing the HTML spec”, 19 November (message 325), with “shoot the engineers and ship it” and “SGML is a mess!”; “SGML Cop backs off”, 17 November (message 298); “An HTML specification and Implementors’ Guide”, 30 November (message 370); “HTML providers: please grab sgmls and the DTD”, 1 December (message 390), with “But the files still don’t fit into SGML”, “I’m just about to give up on the structure business”, and “tag soup”; “The spec evolves…”, 4 December (message 402), with “just describe it; don’t prescribe it”. The proposal to register text/html with MIME is in the thread of around 10 November; the proposer was not looked up for this article. 2 3

  10. “thoughts on the future of HTML [long]”, 21 January 1993 (1993, first quarter, message 89), with “Enter SGML. It seemed like the natural choice, so Tim implemented an informal SGML parser in his WWW clients”, “information providers who wanted to produce HTML automatically just checked to be sure the public www client grokked”, “Uh oh!”, and “the client is at fault if the document is valid, and the server is at fault if the document is not”; “libHTML to date” the same day (message 90), “I gotta go. It’s a rush job”. His last message in the archive is of 30 January (message 111). Bob Stayton (SCO), “Re: Dan Connoly”, 6 August (1993, third quarter, message 420): “He is too busy in his new job to be involved with WWW now.” A third party’s report; this article draws no conclusion about where he went. 2 3

  11. Marc Andreessen, “proposed new tag: IMG”, 25 February 1993, 21:09 (message 174); Tony Johnson’s reply the same day (message 175); Berners-Lee’s two replies of the 26th (messages 178 and 183), with “I hadn’t wanted a special tag” and “I don’t want to change HTML now if I can help it, until it has gone to RFC track”; Andreessen’s reply (message 189), “I absolutely agree in all cases”; Berners-Lee of the 27th (message 194), “even if Dan C ain’t here to round us up we maybe ought to stick to the track”. IMG in Mosaic 1.1 is from Andreessen’s “IMG extension for Mosaic 1.1” of 18 May (second quarter, message 341); only its subject and date are used. “The fourth week after he left” is counted from 30 January. 2

  12. Andreessen, “adherence to DTD’s, etc.”, 14 August (third quarter, message 666), with “the market has chosen” and “Now donning my fire-proof bodysuit”; Terry Allen, 18 August (message 721), “Marc, with respect, your browser puts out a lot of error messages”; Andreessen the same day, “browsing vs validation, or, why not to make software robust” (message 722), with “Folks, don’t complain to us because Mosaic is robust” and, in capitals, “DOING SO WOULD REQUIRE WORK”; Lou Burnard, “who validates?”, the same day (message 711). None of these letters mentions Connolly’s letter of January; “had not necessarily read” is this article’s wording, not a statement of fact. 2

  13. RFC 1866, Hypertext Markup Language - 2.0, T. Berners-Lee (MIT/W3C) and D. Connolly, November 1995, standards track. Quotations are from its introduction: “brings together, clarifies, and formalizes a set of features that roughly corresponds to the capabilities of HTML in common use prior to June 1994”; “HTML is an application of ISO Standard 8879:1986”. Only the header and introduction were checked; the IETF HTML working group process of 1994–1995 was not, which is why the text says it is not covered.