This page is a controlled experiment. It carries the same information in fourteen different places in the HTML, then measures which of those places survive the journey from web server to language model.
How the test worksAdding schema markup is the most confidently repeated piece of advice in AI search optimization. It appears in nearly every guide, checklist and audit template published in the last two years. Almost none of them check whether a language model ever receives the markup. PILOTTEST13
That gap matters, because there are several distinct stages between your web server and the text a model reasons over, and structured data can be lost at any of them. A crawler requests the page. Something decides whether to execute JavaScript. Something else converts the resulting HTML into clean text, discarding navigation, styling, scripts and other material judged to be noise. Only what survives that final step reaches the model.
Most published advice treats those stages as a single step. It assumes that because the markup is in the file, it is in the model. We wanted to find out whether that assumption holds, and if it doesn't, to identify precisely where the loss occurs and why.
The specific question this page answers: when structured data fails to reach a model, is that because the data is hidden from human readers, or because of the kind of HTML element it lives in? In the real world those two things are always tangled together, because JSON-LD is always both hidden and inside a script tag. This page separates them.
The page carries a set of small marker strings. Each marker is unique, and each one is placed in a different kind of HTML container: a script tag, a meta tag, a hidden div, an attribute, ordinary visible text, and so on.
Because every marker is distinct, any tool that fetches this page reveals exactly which containers it preserved and which it discarded. If a marker appears in the output, that container survived. If it doesn't, it didn't.
The obvious approach is to build one page per variant and fetch them in turn. We rejected that, because it introduces a confound: different requests, made at different moments, potentially hitting different cache states or extractor versions. Any difference in results could be an artifact of the request rather than the markup.
Putting all the containers on a single page means every tool evaluates all of them within one identical request. Differences in the output can only come from the containers themselves.
Extraction libraries don't treat all pages alike. Readability-style algorithms score regions of a document by text density and link ratio to guess where the main content lives, and they behave erratively on thin or synthetic pages. So this page carries genuine prose across a hero, several content sections, a FAQ block and a footer, matching the shape of an ordinary marketing page. The goal is for extractors to behave here exactly as they would anywhere else.
The page is also entirely static. There is no client-side rendering, no framework, and nothing that modifies the document after load. That isolation is deliberate: it means the test measures extraction and nothing else.
Two axes matter. The first is position: identical JSON-LD blocks placed in the document head, at the top of the body, buried inside nested containers and table cells, and at the very end of the body. If location affects survival, these will differ from one another.
The second axis is the container itself. Byte-identical JSON content is placed three ways: inside a script tag, inside a div with display:none, and inside a template element. All three are equally invisible to a human reader. If they produce different results, the deciding factor is the element, not the hiding.
| Container | What it isolates |
|---|---|
| JSON-LD in the head | baseline placement |
| JSON-LD, CDATA-wrapped | whether validity affects handling |
| JSON-LD at body top, nested, and body end | document position |
application/json script | MIME type vs. scripts as a class |
| Hidden div holding identical JSON | the container, with hiddenness held constant |
| Template element holding identical JSON | a third container, same content |
| Meta description | whether head metadata survives |
| Microdata and RDFa in attributes | non-JSON-LD structured data formats |
| ARIA label | whether accessibility attributes are retained |
Note for anyone replicating this: the marker strings themselves are deliberately not printed anywhere in this page's visible text. If they were, every tool would find them trivially and the test would measure nothing. Replicators should keep their markers out of the readable copy for the same reason.
A result is only meaningful if the test can fail. Four controls establish that.
A marker sits in ordinary visible body text. It must appear in every tool's output. If it doesn't, the fetch failed and the entire run is void.
A marker sits inside an HTML comment. It must appear in no output. If it does, the tool is passing raw HTML and no conclusion about extraction can be drawn.
Before interpreting anything, the raw HTML is checked to confirm every marker was actually served, guarding against upload and template errors.
A fourth check inspects the live DOM to confirm the markup is still present after the page loads. That rules out the failure mode where a page's own JavaScript destroys its structured data during rendering, which is a real and surprisingly common bug, and one this page is deliberately free of.
Two categories of tool, and the distinction matters when weighing the results.
First-party vendor pipelines are the browsing and fetching tools built into consumer AI assistants. These are the real thing: whatever they return is, by definition, what that model receives. They are the strongest evidence available from outside a vendor.
Representative scraping stacks are the general-purpose tools that agents and RAG systems commonly use to convert pages into text. These aren't identical to any particular vendor's internal pipeline, but they reflect the standard approach and are useful for showing whether behaviour is idiosyncratic or general.
One important limit applies to both. This measures extraction into text. A crawler could store raw HTML and parse structured data at a later stage that an external test cannot observe. Findings here describe what reaches a model on a live fetch, which is a narrower and more specific claim than "AI ignores schema."
No, and that conclusion doesn't follow from this test. Structured data does real work at the search index layer, where it supports rich results and helps engines resolve entities. Most AI answers are generated from a search index rather than a live page fetch, so markup can influence what a model receives without the model ever parsing the markup itself. This page measures one specific path: what arrives when a tool fetches your URL directly.
Because separate pages mean separate requests, and separate requests can differ in ways that have nothing to do with the markup: caching, timing, extractor versions, transient errors. Placing every variant on one page means a single request evaluates all of them under identical conditions, so any difference in the output is attributable to the containers being compared.
Not for this page, which is entirely static and identical before and after rendering. That's the point: it isolates the extraction step. Rendering matters enormously for pages whose content is built client-side, and published crawler research consistently finds that most AI crawlers do not execute JavaScript at all. That's a separate and equally important problem, but it is not what this page measures.
It appears in this test as a controlled comparison, not as a recommendation. Deliberately serving content to machines that humans can't see is cloaking-adjacent, risks violating search engine guidelines, and is a poor foundation to build on. The honest lesson from this experiment is the opposite: put the information users need into the visible content, because that's the layer every pipeline shares.
Unknown, and that's a real limitation. Extraction pipelines are ordinary engineering choices, not laws of nature, and vendors change them without announcement. Any finding from a test like this should be dated and re-run before being cited as current. Where different tools already disagree with each other, that's a useful signal that the behaviour is a decision rather than a constraint.
Yes, and you should. Take any content-rich page, insert markers in each container type, host it publicly, confirm via raw HTML that every marker is actually being served, then fetch it through whichever tools you care about and search each output for each marker. Verify your positive and negative controls before interpreting anything else. The decisive comparison is the script tag against the hidden div holding identical content.
Because it's a measurement instrument rather than content anyone searched for, and indexing it would clutter results without helping anybody. The directive affects indexing, not fetching, so it has no bearing on what the test measures. Testing whether search engines index and parse the markup over time is a separate, slower experiment that requires the directive removed and the page left in place.
Pilot Digital is a Chicago SEO, PPC and web agency. When the advice in our industry doesn't come with evidence, we go and get some.
Talk to us