{"slug":"this-newspaper-is-also-an-api","title":"This newspaper is also an API","dek":"Every article here is served three ways. Ask for HTML and you get a page. Ask for Markdown or JSON and you get the article with nothing wrapped around it. Here is how, and why we bothered.","section":"publishing","sectionName":"Publishing","author":"Mara Okafor","date":"2026-09-29","featured":false,"readingMinutes":4,"wordCount":807,"price":null,"priceLabel":null,"urls":{"html":"https://www.thedailyagent.news/articles/this-newspaper-is-also-an-api","markdown":"https://www.thedailyagent.news/articles/this-newspaper-is-also-an-api.md","json":"https://www.thedailyagent.news/articles/this-newspaper-is-also-an-api.json"},"markdown":"If you are a person reading this in a browser, you are seeing our layout: a masthead, a headline, this text set in a serif at a comfortable measure. If you are a program, you probably do not want any of that. You want the words, the byline, the date, and a way to find the next article, with no navigation to strip and no styling to ignore.\n\nMost publishers make programs work for that. The agent fetches a page, throws away the chrome, guesses at which block is the article, and hopes the byline is where it was last week. It usually works. It is also a waste of everyone's tokens, and it means the publisher has no say in what the agent sees.\n\nWe took a different approach. Every article on this site is the same document served in three representations, and you choose which one you want.\n\n## Three ways to ask\n\nThe simplest is the file extension. Take any article URL and append `.md` or `.json`.\n\n```\n/articles/this-newspaper-is-also-an-api        HTML page\n/articles/this-newspaper-is-also-an-api.md     Markdown document\n/articles/this-newspaper-is-also-an-api.json   JSON with metadata and body\n```\n\nThe second is content negotiation, which is how HTTP was always meant to do this. Send an `Accept` header that prefers Markdown or JSON and the plain URL returns that instead of HTML.\n\n```bash\ncurl -H \"Accept: text/markdown\" https://thedailyagent.example/articles/this-newspaper-is-also-an-api\ncurl -H \"Accept: application/json\" https://thedailyagent.example/articles/this-newspaper-is-also-an-api\n```\n\nBrowsers send `text/html` first, so people keep getting pages. Agents that ask for Markdown get Markdown. Nobody has to know about the other.\n\nThe third is the index. `/llms.txt` lists every article with a one-line summary and a link to its Markdown version, following the llms.txt convention. `/llms-full.txt` is every article concatenated into one document, for an agent that wants the whole archive in one request. `/articles.json` is the same index as structured data. And the HTML pages carry `<link rel=\"alternate\">` tags and a `Link` header pointing at their other forms, so a crawler that lands on a page can discover the machine-readable version without guessing.\n\n## What the Markdown looks like\n\nThe Markdown response is not a dump of our source file. It is a document assembled for a reader that has no other context: title, standfirst, byline, date, section, canonical URL, then the body.\n\n```markdown\n# This newspaper is also an API\n\n*Every article here is served three ways...*\n\nBy Mara Okafor. Published 29 September 2026 in Publishing, The Daily Agent. About 5 min read.\nCanonical: https://thedailyagent.example/articles/this-newspaper-is-also-an-api\nAlso available as JSON: https://thedailyagent.example/articles/this-newspaper-is-also-an-api.json\n\n---\n\nIf you are a person reading this in a browser...\n```\n\nEverything an agent needs to cite the article correctly is in the first six lines. The links inside the body are relative to the site, the same as in the HTML, so an agent following a reference to another article can fetch it in the same format.\n\n## Why a publisher would do this\n\nThe obvious objection is that we have just made ourselves easier to scrape. That is true, and it is the point. The traffic we want to charge is agent traffic, and you cannot charge a customer you have made it hard to serve. A paywall that an agent can pay at, which is what x402 gives us, only makes sense if the thing behind the paywall is something an agent can actually use.\n\nThere is a second reason. When an agent parses our HTML, it decides what the article is. When we serve Markdown, we decide. We control the byline, the date, the canonical link. If a summary of our reporting ends up in front of a reader somewhere else, we would rather it was built from a document we wrote for that purpose than from whatever survived the stripping.\n\nAnd there is a third reason, which is that it was not hard. The CMS is a folder of Markdown files with a few lines of metadata at the top. Rendering that to HTML for people and passing it through for programs is the same code path with a different last step. The content negotiation is one small function that looks at a header. If your publishing stack is more complicated than ours, the principle still holds: you already have the article as data somewhere, and exposing it costs less than you think.\n\n## What the paywall changes\n\nNot everything here is free. Some articles carry a price, and for those the 402 applies to all three representations equally, because the price is for the article and not for the wrapper. An agent that pays for the Markdown has bought the same thing as a person who pays for the page. The index at `/llms.txt` marks which articles cost money and how much, so an agent can decide before it asks.\n\nEverything else, including this article and the index, stays open. Fetch `/llms.txt` and have a look around.","html":"<p>If you are a person reading this in a browser, you are seeing our layout: a masthead, a headline, this text set in a serif at a comfortable measure. If you are a program, you probably do not want any of that. You want the words, the byline, the date, and a way to find the next article, with no navigation to strip and no styling to ignore.</p>\n<p>Most publishers make programs work for that. The agent fetches a page, throws away the chrome, guesses at which block is the article, and hopes the byline is where it was last week. It usually works. It is also a waste of everyone's tokens, and it means the publisher has no say in what the agent sees.</p>\n<p>We took a different approach. Every article on this site is the same document served in three representations, and you choose which one you want.</p>\n<h2 id=\"three-ways-to-ask\">Three ways to ask</h2>\n<p>The simplest is the file extension. Take any article URL and append <code>.md</code> or <code>.json</code>.</p>\n<pre><code>/articles/this-newspaper-is-also-an-api        HTML page\n/articles/this-newspaper-is-also-an-api.md     Markdown document\n/articles/this-newspaper-is-also-an-api.json   JSON with metadata and body\n</code></pre>\n<p>The second is content negotiation, which is how HTTP was always meant to do this. Send an <code>Accept</code> header that prefers Markdown or JSON and the plain URL returns that instead of HTML.</p>\n<pre><code class=\"language-bash\">curl -H \"Accept: text/markdown\" https://thedailyagent.example/articles/this-newspaper-is-also-an-api\ncurl -H \"Accept: application/json\" https://thedailyagent.example/articles/this-newspaper-is-also-an-api\n</code></pre>\n<p>Browsers send <code>text/html</code> first, so people keep getting pages. Agents that ask for Markdown get Markdown. Nobody has to know about the other.</p>\n<p>The third is the index. <code>/llms.txt</code> lists every article with a one-line summary and a link to its Markdown version, following the llms.txt convention. <code>/llms-full.txt</code> is every article concatenated into one document, for an agent that wants the whole archive in one request. <code>/articles.json</code> is the same index as structured data. And the HTML pages carry <code>&#x3C;link rel=\"alternate\"></code> tags and a <code>Link</code> header pointing at their other forms, so a crawler that lands on a page can discover the machine-readable version without guessing.</p>\n<h2 id=\"what-the-markdown-looks-like\">What the Markdown looks like</h2>\n<p>The Markdown response is not a dump of our source file. It is a document assembled for a reader that has no other context: title, standfirst, byline, date, section, canonical URL, then the body.</p>\n<pre><code class=\"language-markdown\"># This newspaper is also an API\n\n*Every article here is served three ways...*\n\nBy Mara Okafor. Published 29 September 2026 in Publishing, The Daily Agent. About 5 min read.\nCanonical: https://thedailyagent.example/articles/this-newspaper-is-also-an-api\nAlso available as JSON: https://thedailyagent.example/articles/this-newspaper-is-also-an-api.json\n\n---\n\nIf you are a person reading this in a browser...\n</code></pre>\n<p>Everything an agent needs to cite the article correctly is in the first six lines. The links inside the body are relative to the site, the same as in the HTML, so an agent following a reference to another article can fetch it in the same format.</p>\n<h2 id=\"why-a-publisher-would-do-this\">Why a publisher would do this</h2>\n<p>The obvious objection is that we have just made ourselves easier to scrape. That is true, and it is the point. The traffic we want to charge is agent traffic, and you cannot charge a customer you have made it hard to serve. A paywall that an agent can pay at, which is what x402 gives us, only makes sense if the thing behind the paywall is something an agent can actually use.</p>\n<p>There is a second reason. When an agent parses our HTML, it decides what the article is. When we serve Markdown, we decide. We control the byline, the date, the canonical link. If a summary of our reporting ends up in front of a reader somewhere else, we would rather it was built from a document we wrote for that purpose than from whatever survived the stripping.</p>\n<p>And there is a third reason, which is that it was not hard. The CMS is a folder of Markdown files with a few lines of metadata at the top. Rendering that to HTML for people and passing it through for programs is the same code path with a different last step. The content negotiation is one small function that looks at a header. If your publishing stack is more complicated than ours, the principle still holds: you already have the article as data somewhere, and exposing it costs less than you think.</p>\n<h2 id=\"what-the-paywall-changes\">What the paywall changes</h2>\n<p>Not everything here is free. Some articles carry a price, and for those the 402 applies to all three representations equally, because the price is for the article and not for the wrapper. An agent that pays for the Markdown has bought the same thing as a person who pays for the page. The index at <code>/llms.txt</code> marks which articles cost money and how much, so an agent can decide before it asks.</p>\n<p>Everything else, including this article and the index, stays open. Fetch <code>/llms.txt</code> and have a look around.</p>"}