DABYTE DATA DESK · operations
Three documented server behaviours that break agent-facing files while returning 200
Answer. All three are written down in vendor documentation, all three are correct within their intended scope, and all three remove an agent-facing file from the web while the HTTP status stays 200. A catch-all route that answers 200 text/html on every unmatched path makes files that were never served look present: on 2026-08-09 our own https://vectory.space/auth.md returns 200 text/html, and the body is the 138,938-byte marketing homepage, the same size as the response to an invented path and to the site root. In nginx, one add_header inside a location discards every add_header inherited from the server level, and the header most easily lost that way is Link, whose only audience is a machine. Also in nginx, the charset directive applies only to MIME types listed in charset_types, a default list that does not contain text/markdown, so markdown ships without the charset parameter that RFC 7763 marks REQUIRED. None of this is a bug report. Each behaviour is documented; the damage comes from the combination, and from the fact that a status-code check cannot see any of it.
Who reported it, and when we saw it
Key facts
- nginx documentation for ngx_http_headers_module states that add_header directives are inherited from the previous configuration level if and only if there are no add_header directives defined on the current level (checked 2026-08-09).
- nginx documentation for try_files states: "If none of the files were found, an internal redirect to the uri specified in the last parameter is made." That sentence is the vendor description of the catch-all behaviour (checked 2026-08-09).
- nginx 1.29.3, released 28 Oct 2025, added the add_header_inherit and add_trailer_inherit directives; add_header_inherit takes on, off or merge, and merge appends values from the previous level (nginx.org/en/CHANGES, checked 2026-08-09).
- The nginx default is charset off, and the default charset_types value is text/html text/xml text/plain text/vnd.wap.wml application/javascript application/rss+xml. Neither text/markdown nor text/csv is in that list (checked 2026-08-09).
- RFC 7763 Section 2, registering the text/markdown media type: charset is REQUIRED, and there is no default value.
- Measured 2026-08-09: on vectory.space, /auth.md, /this-does-not-exist-zzz.md and / all return 200 text/html with a 138,938-byte body, and that body contains zero occurrences of the string auth.md; on dabyte.ai the same invented path returns 404 and /auth.md returns text/markdown; charset=utf-8 at 1,150 bytes.
- add_header without the always parameter applies only to responses with status 200, 201, 204, 206, 301, 302, 303, 304, 307 or 308, so headers stop being sent the moment a path correctly starts returning 404 (nginx documentation, checked 2026-08-09).
What a 200 actually promises
RFC 9110 Section 15.3.1 defines 200 as an indication that the request has succeeded. That is a statement about the request, not a guarantee that the representation returned is the one the caller asked for. Section 15.5.5 defines 404 as the origin server not finding a current representation for the target resource. Nothing in either definition obliges a server to tell a served file apart from a fallback page. That distinction is made entirely in configuration, and configuration is where all three of these behaviours live. The gap matters more for machine-facing files than for pages. A person who lands on a marketing page instead of a machine-readable catalogue can see in a second that something is wrong. A checker that reads the status line sees success and moves on. Google names the human-facing version of this a soft 404, which its Search Console documentation defines as a user-friendly not-found message returned without a 404 HTTP response code (checked 2026-08-09). Note the definition is about the absence of a 404, not the presence of a 200. The agent-facing version behaves identically and comes with none of the visual cues. One command shows both sides. Run: curl -sS -o /dev/null -o /dev/null -w '%{url_effective} %{http_code} %{content_type}\n' https://vectory.space/this-does-not-exist-zzz.md https://dabyte.ai/this-does-not-exist-zzz.md. On 2026-08-09 the first line is 200 text/html and the second is 404. Two details matter if you paste it: curl consumes one -o per URL, so with a single -o the second response body prints to your terminal, and the format string has to end in a newline or both results run together on one line. Use %{url_effective} rather than %{url}, which was introduced in curl 7.75.0 and silently resolves to nothing on older builds, taking your labels with it.
Behaviour one: the catch-all that answers every path
The rule is ordinary, deliberate and documented. A single-page application needs unmatched paths to reach its entry point so client-side routing can take over, so the server is told to fall back to index.html. In nginx this is try_files, whose documentation states that if none of the files were found, an internal redirect to the uri specified in the last parameter is made. Marketing sites inherit the same pattern from their templates without ever needing it. Inside its intended scope the rule is correct and nothing is wrong with it. The scope leaks. Every path the configuration does not explicitly claim now returns the fallback with a 200 and a Content-Type of text/html, including paths that published specifications place at the site root. The auth-md check published by isitagentready.com asks for the file to be served from the service root as Markdown, with an H1 heading that contains the string auth.md. The service root is precisely the namespace the fallback owns. The blast radius is narrower than it first appears, and it is worth measuring rather than assuming. On vectory.space on 2026-08-09, /.well-known/api-catalog returns application/linkset+json and /.well-known/agent-card.json returns an honest 404, because those prefixes have their own location blocks. The root namespace has none: /auth.md and /index.md both return 200 text/html, as does any invented path. So the files that break are exactly the ones the specifications put at the root, and they are the ones you are least likely to notice, because the site itself looks completely healthy in a browser.
Behaviour two: one add_header in a location erases the inherited set
The nginx documentation is explicit: there can be several add_header directives, and they are inherited from the previous configuration level if and only if there are no add_header directives defined on the current level. This is a replacement rule, not a merge, and it operates on the whole set rather than per header name. Add a single Cache-Control inside a location and every Content-Security-Policy, Strict-Transport-Security and Link declared at the server level silently stops being sent for that location. We walked into this twice. The first time it removed HSTS. The second time, recorded in our notes on 2026-08-06, it removed the Link header from the responses served by a location matching .md: a rewrite hands the request to that location, the location declares its own Vary and X-Robots-Tag, and the Link set declared at server level is dropped. Link is the map of machine surfaces defined by RFC 8288, and it exists for exactly one audience. It disappeared on the one response that audience had asked for. On the same host today the catch-all answers /auth.md from the HTML location instead, so the server-level Link set is present again at that URL and the original fault is no longer visible there. A reader checking the claim against that URL now will see seven Link headers, which is behaviour one masking behaviour two rather than evidence that behaviour two was not real. There are two fixes. The portable one is to repeat the add_header Link line in every location that declares any header of its own. The other arrived in nginx 1.29.3 on 28 Oct 2025: add_header_inherit, which takes on, off or merge, where merge appends values from the previous level to those defined at the current one. The documentation notes the inheritance rule is itself inherited, so merge at the top level applies recursively until redefined. Check nginx -v before reaching for it; on an older build the directive is unknown and the configuration will not load. One corollary bites after you fix behaviour one. add_header without the always parameter applies only to a listed set of status codes, and 404 is not among them. The moment a path starts returning an honest 404, any header you added without always stops accompanying it.
Behaviour three: charset applies only to the types in charset_types
nginx defaults to charset off. When you set charset utf-8, it takes effect for text/html and for the types named in charset_types, whose default value is text/html text/xml text/plain text/vnd.wap.wml application/javascript application/rss+xml. text/markdown is not in that list, and neither is text/csv. The result is a response reading Content-Type: text/markdown with no charset parameter at all. RFC 7763, which registers the media type, states in Section 2 that charset is REQUIRED and that there is no default value, because neither the syntax description nor the popular implementations at the time of registration defined one. A markdown response without the parameter is therefore under-specified by its own registration, and the consumer is left to guess an encoding. What we saw when this hit us was an em dash arriving as mojibake. The trap is where you look for the cause. The file on disk was valid UTF-8, and regenerating it changed nothing, because the bytes were never the problem; the response metadata was. Only charset_types fixes it, and Roger Sen documents the same fix independently. After adding text/markdown and text/csv to the list, https://dabyte.ai/auth.md returns text/markdown; charset=utf-8 and https://dabyte.ai/aiv.csv returns text/csv; charset=utf-8, both checked on 2026-08-09.
The order in which to fix them
Fix the catch-all first, because until unmatched paths return 404 you cannot read any other measurement. A 200 with text/html on a machine file has two possible causes: no file is being served at that path, or a file exists and is being served with the wrong type. The status code does not distinguish them, and neither does the Content-Type. Any hour spent debugging a media type on a path with nothing behind it is an hour spent because this step was skipped. Content types come second and headers third. Headers go last because the header fix has to be applied per location, and you do not know your final set of locations until routing and media types have settled. Adding the Link line to a location you are about to delete is wasted work, and worse, it makes the intermediate measurements look better than the configuration is.
The check, and what is still broken on our own host
Four requests cover all three behaviours. Ask for a path you know does not exist, carrying the extension you care about, and require a 404. Ask for the real file and require its own media type with a charset parameter. Ask for the same URL twice, once plain and once with Accept: text/markdown, and compare the Link lines. Use GET rather than HEAD: on one of our hosts on 2026-08-06 a HEAD with Accept: text/markdown returned text/html while the GET returned markdown, which reads as no markdown support at all if HEAD is the only thing you test. That host's routing has since changed, so the observation no longer reproduces there; it is recorded here as a reason to prefer GET, not as a live measurement. On dabyte.ai on 2026-08-09 all four pass. Nonexistent paths return 404, /auth.md returns text/markdown with charset=utf-8, and the same five-relation Link line covering api-catalog, service-desc, service-doc, describedby and license is present on both the HTML response and the markdown one. On vectory.space on the same date the catch-all is still in place. /this-does-not-exist-zzz.md returns 200 text/html, and /auth.md returns 200 text/html with no file being served behind it, with a body the same 138,938 bytes as the homepage. Size alone would be weak evidence, so here is the stronger form: the body served at /auth.md contains zero occurrences of the string auth.md, so it cannot be the file the auth-md check asks for, which requires an H1 containing that string. We are publishing this while our own site fails the first of our own four checks. That is the argument for running the first check: nothing about that site looks broken until you ask it for something that does not exist.
Questions this answers
Is any of this specific to nginx?
Behaviours two and three are nginx directive semantics, and for this article we verified them only against the nginx documentation and the nginx CHANGES file on 2026-08-09. Other servers merge response headers differently and handle charset differently, so do not carry these two across without checking. Behaviour one is documented in nginx as try_files, but the underlying pattern is not server-specific: it is whatever rule sends unmatched paths to a single-page entry point, and that rule exists in platform and framework routing too. We did not test other servers for this article, so run the same curl against yours rather than assuming.
The bytes are valid UTF-8. Does a missing charset parameter really matter?
RFC 7763 Section 2 marks charset REQUIRED for text/markdown and states there is no default value, so a response without it is under-specified by the media type's own registration and the consumer decides what to assume. What we observed in our own case was an em dash arriving as mojibake. We have not tested how each individual agent client behaves, so the honest claim is the specification requirement plus that one observation, not a general prediction.
Should we just set add_header_inherit merge globally and stop repeating headers?
It is available from nginx 1.29.3, released 28 Oct 2025, so check nginx -v first; on an older build the directive is unknown and the configuration will fail to load rather than degrade quietly. The documentation notes that the inheritance rule is itself inherited in the standard way, so merge set at the top level applies recursively until redefined. Repeating add_header in each location remains the portable fix and works on every version.
Why did the catch-all break auth.md but not the API catalogue on the same host?
Because of where each file is specified to live. On our host, paths under /.well-known/ have their own explicit location blocks, so the fallback never sees them. On vectory.space on 2026-08-09, /.well-known/api-catalog returned application/linkset+json and /.well-known/agent-card.json returned a truthful 404. auth.md is specified at the service root, and the root is the namespace the fallback claims. We have not surveyed how other operators configure /.well-known/, so treat the reason as specific to this host.
Machine access
/api/articles.json— every piece we published, with the measurement each one rests on/api/aiv.json— the AI Visibility Index the numbers above come from- How the index is measured