Industry gospel says that schema helps optimize for AI Search, but does it?
We had an eye-opening experience helping a client who had JavaScript rendered content. ChatGPT wasn’t “seeing” their events, and therefore, had incomplete and incorrect data when asked. The event data lived in a table that was rendered via JavaScript on pageload. However, AI bots don’t render JavaScript, therefore the content was invisible to them.
The client had the idea to add the event data as JSON schema. My response to their idea was a supportive, “great, yes, that’s a smart workaround. AI likes structured data so that should ensure they see your event information.”
After the schema was added, we ran some tests with Claude and it was insistent the schema wasn’t there. After more and more attempts, it finally launched Claude in Chrome and confirmed the schema was present. A relief, but it exposed a feature of how their native web fetch works:
- When ChatGPT, Claude, and Gemini fetch a page, they strip all scripts, including JSON-LD
This finding had us wondering, under what conditions, if any, does an LLM interact with JSON schema on your webpage. To answer this question, we designed a research study to determine what factors lead to different models – Claude, ChatGPT, and Gemini – retrieve schema when they fetch a page.
Here’s the chain of events that kicked off our skepticism of schema and its role in AI search
Below is a step by step account of a deep rabbit hole I went down, which was kicked off by this client’s JS-rendered content.
- First, I looked at the page in the browser by disabling JavaScript. Nothing. Empty. The events table didn’t exist, which also confirmed it was rendered with client-side javascript.
- Next, I asked Claude, “Hey, fetch this page and tell me if you can see the events table.” At first, it said yes, but it was using Firecrawl to fetch the page, not its default fetch tool. Once it used its default browser fetch, it confirmed what I found first, no events.
- I wrote to the client and reported our findings that those events are likely invisible to most LLMs during a standard fetch request.
- Their plan was to add Event schema via JSON-LD to the page, and I thought this was a good idea. Everyone knows that schema is good for AI optimization. LLMs may not be able to see the events table, but they’d see the Events schema and get the information that way. Solid plan.
- After they added the Events schema, we checked the page again with Claude. Claude kept telling me it didn’t see the schema. I was perplexed and skeptical. See Claude’s response after multiple asks to check for it.
I checked, and I'm not finding it. Two separate cache-busted scrapes of the fully JS-rendered HTML (5-second render wait, maxAge: 0), then searched the raw source for every form structured data can take:
application/ld+json — 0 matches
schema.org — 0 matches
@context / @type / @graph — 0 matches
itemprop / itemscope / typeof= (microdata, RDFa) — 0 matches
What’s funny is Claude thinks it should be able to see the schema, too. It later asked me, “Did they deploy it on a different page?” and “Is it configured but not deployed?”
After insisting the Events schema was present, it launched the Claude Chrome extension and confirmed it did indeed exist. But that’s a problem, right? If Claude can only see JSON-LD schema using Chrome as an Agent (or now in its built-in browser), then what are we optimizing for?
At this point, I’m feeling untethered. My gears are turning, wheels spinning. If LLMs can’t see schema when they fetch a page, then what use does it have for AI optimization purposes? During training runs, our understanding is that scripts are dropped and the content is extracted and tokenized, meaning, JSON-LD wouldn’t impact training either. At least not directly.
As my mind is spinning, I remember the client saying the Events schema was “injected” onto the page. Could it be that the method of schema insertion matters? Could Claude and other LLMs see it if we made the JSON-LD more accessible? SEJ has a published article on this topic, but at this point, I trust nothing so our next step was to design an experiment to see under what conditions (if any) LLMs parse JSON schema.
The experiment we created to test whether AI can see schema
Our idea was to build a static HTML webpage that deployed structured data multiple ways to see which methods resulted in the AI tools being able to “see” it on page fetch. We worked with Claude to help build out a quick web page and design the experiment parameters.
Within the webpage we hid 14 different markers (we’re calling them “canaries”) in the HTML to see under what conditions the schema survived page fetch by Claude, ChatGPT, and Gemini. A few of the markers were negative or positive controls to make sure the test page was sound but it was designed to test two different things:
- Location on the page: does location on the page matter when it comes to extraction, e.g. header, body, footer?
- Container: does it matter whether it’s a <script>, <div> or other container?
Research findings
Below are the findings of our experiment. The results are conclusive, your schema isn’t being seen by AI models.
| What we hid on the page | ChatGPT | Gemini | Claude |
|---|---|---|---|
| JSON-LD schema (5 blocks in 5 different spots) | 0 of 5 | 0 of 5 | 0 of 5 |
| The same JSON in a hidden <div> instead of a <script> tag | ✓ | ✓ | ✓ |
| Meta description | ✗ | ✗ | ✓ |
| Accessibility label (aria-label) | ✗ | ✓ | ✗ |
| Normal visible text (positive control) | ✓ | ✓ | ✓ |
Tested August 6, 2026. Full results for all 14 canaries and 5 tools are in the appendix.
- ChatGPT, Gemini and Claude found 0 of 15 JSON-LD blocks. Add two common scraping tools, Firecrawl and Chrome’s reader, and it’s 0 of 25 for tools extracting JSON-LD blocks. NOTE: Firecrawl was able to see the events on the rendered page, not in JSON schema.
- Location doesn’t matter. Header, footer, body. None of our experiment’s JSON-LD survived the fetch.
- The containing element is the reason why it doesn’t survive the fetch. JSON-LD sits in a <script>, which gets discarded by each model during fetch.
- Whether accessibility attributes survive fetch is model dependent. Only Gemini fetched the content we hid inside an ‘aria-label’.
The final finding about Gemini is an unexpected finding from our study. Google’s developer docs say AI agents use the accessibility tree to make sense of a page, mostly to find things like buttons and form fields. So we were curious whether any of the AI tools would pick up text we tucked into an aria-label. Only Gemini did. It’s only one result, so I wouldn’t build a strategy around it. But it does look like Gemini’s browsing tool holds on to some accessibility data the others toss out.
Even more interesting was Gemini’s response to our prompt about its ability to extract any JSON.
“This tool strips <script> tags, including Schema/JSON-LD, during content extraction, which is why other potential matches in the raw source code may not be visible here… The live browsing tool I have available returns a sanitized, body-only DOM representation of the webpage and strips out the <head> section and all <meta> tags before I receive the data.”
That is a major AI model describing its extraction layer is designed to discard JSON-LD.
Skip down to the full experiment details →
Do AI models see my structured data on training runs?
Our experiment showed that AI bots don’t see your schema on fetch. But what about during training, is the schema included then? Unfortunately, this is very unlikely. Anthropic and ChatGPT use ClaudeBot and GPTBot to crawl, and extract text at a massive scale. They don’t publish their training pipeline, but most standard LLM training pipelines (Fineweb by Hugging Face, Trafilatura) strip <nav>, <form>, <script>, <style>, and <footer> tags. This means at the two most crucial moments – training runs and live fetch by an agent – AI likely isn’t seeing your schema.
If LLMs can’t fetch JSON schema, what good is it?
The conventional wisdom in the AiSEO community is that JSON-based schema structured data is critical for AI search optimization. There are articles abound claiming that structured data is essential for AI optimization:
- Schema Markup for AI Search: Complete Guide – SEOptimer, Nov 2025
- Schema Markup for AI Citations: The Technical Implementation Guide – Averi.ai , Dec 2025
- How Can Schema Markup Specifically Enhance LLM Visibility – Walker Sands, Nov 2025
- How to Use Schema Markup for AI to Improve Search Visibility – TMV, May 2026
However, when you dig into the details, you start to notice more correlation than causation, or the finding is conflated to fit the claim, or there is no cited source anchoring their claim at all (ahem… Walker Sands). Averi.ai claims in their post above, “GPT-5’s accuracy improves from 16% to 54% when content relies on structured data…” but when you look at the source, the “structured data” in question is about answering natural-language questions over an enterprise SQL database. It’s not about AI search, and they also cited the wrong GPT model, which was GPT-4 in the study.
What about those quotes from Google and Microsoft search execs?
There are a couple oft-cited quotes from Google and Microsoft Search executives who commented on structured data and AI, but when you review, neither are a slam dunk.
On Google Search Central Live, Ryan Levering from Google shared, “A lot of our systems run much better with structured data,” he noted, adding that “it’s computationally cheaper than extracting it.” Which could be another instance of “structured data” getting conflated with JSON-LD schema.
The day before, Fabrice Canel, Principal Product Manager of Bing, stated at SMX Munich that, they use schema “as a reference check of their LLM training data” which David Mihm posted about on LinkedIn claiming that, “Fabrice Canel confirms that schema markup helps Microsoft’s LLMs understand your content.” But that’s not what he said, and he was actually promoting IndexNow as a means to keep your site fresh in Bing, which ChatGPT uses for search grounding.
If schema doesn’t optimize for AI search, why do other studies find a correlation?
We think studies that show a correlation between websites that implement schema and AI citations rates surfaces for two reasons:
- Most websites that deploy structured data are also focused on improving their SEO. And while not identical, AI search optimization has a lot of overlap with SEO best practices.
- Structured data must represent content that’s on a page. That means you need to have really well structured page content and website IA to have a robust structured data strategy. In this scenario, these sites have robust content, which is likely why they have higher citation rates.
Focus on content instead
Google said in their Generative AI fundamentals guide that people are “overfocusing on structured data.” We always tell clients, content is always the biggest lever to pull for website optimization. Your website content is your training data. What products or services you choose to highlight is what AI will know about your brand.
And back to what kicked off this experiment, do not use JavaScript-rendered content. Any content rendered via JavaScript will be invisible to AI tools.
We included the full experiment details below if you’d like to dig in further.
Appendix: Read the full experiment
We ran this test on August 6, 2026. I designed it with Claude, and Claude helped build the page and run the fetches. The results below come from a mockup of a healthcare page. We’ve since put up a public copy with the same 14 canaries, where the page text explains the test itself. You can see that version here.
The question
When an AI tool reads a web page, does the schema on that page reach the AI?
We also wanted to know why it gets lost, if it does. There were two ideas to rule in or out:
- Location. Maybe it matters where the schema sits on the page: the top, the bottom, or buried deep in the code.
- Hidden vs. script. Maybe AI tools skip anything a visitor can’t see. Or maybe they skip the <script> tag that JSON-LD lives in.
The second question is the tricky one. On a real website, JSON-LD is always hidden and always inside a <script> tag. You can’t tell which one is causing the problem. So we had to pull them apart.
How we built the test page
One page, not fourteen
The simple approach would be to build a separate page for each test. But then each page gets fetched at a different time, and that adds noise. Caches change. Tools change.
So we put all fourteen tests on one page. Each one hides a unique code word, from PILOTTEST01 to PILOTTEST14. We call these code words canaries. When an AI tool reads the page, we check which canaries show up in what it got back. If PILOTTEST07 shows up, test 07 survived. If it doesn’t, that test got thrown out.
A real-looking page
We didn’t use a bare-bones page. Some tools decide what counts as the “main content” by looking at how much text is on the page. A nearly empty page can confuse them.
Instead, we used a mockup of a healthcare website. It has about 2,000 words of real text: a header, seven content sections, an eight-question FAQ, and a footer. The page is plain HTML. Nothing loads with JavaScript after the page opens. What the server sends is exactly what’s on the page.
The 14 canaries
| # | What we hid | Why |
|---|---|---|
| 01 | Meta description tag | Does the page summary in the <head> get through? |
| 02 | JSON-LD in the <head>, clean and valid | The textbook way to add schema |
| 03 | JSON-LD in the <head>, wrapped in CDATA | How our client did it. This wrapping technically breaks the JSON. |
| 04 | Plain JSON in a <script> tag (not labeled as schema) | Is it schema being thrown out, or any script? |
| 05 | An HTML comment | Negative control. This should never show up. |
| 06 | JSON-LD at the top of the <body> | Location test |
| 07 | The same JSON, in a hidden <div> | The key test. Hidden like schema, but not in a script tag. |
| 08 | The same JSON, in a <template> tag | Another hidden container |
| 09 | JSON-LD buried deep inside tables and divs | Copies our client’s page structure |
| 10 | Microdata (an older schema format) | Does a different schema format do better? |
| 11 | RDFa (another schema format) | Same question |
| 12 | An accessibility label (aria-label) only | Do the labels used by screen readers get through? |
| 13 | Normal, visible text in a paragraph | Positive control. This should always show up. |
| 14 | JSON-LD at the very end of the <body> | Location test |
Two groups do most of the work:
- Canaries 02, 06, 09, and 14 test location. Same schema, four different spots on the page.
- Canaries 06, 07, and 08 test the container. The exact same JSON, all hidden from visitors. The only difference is what tag holds it: a <script>, a hidden <div>, or a <template>.
Here’s the exact code for each canary:
<!-- 01 meta tag in head -->
<meta name="description" content="…PILOTTEST01">
<!-- 02 JSON-LD in head, clean and valid -->
<script type="application/ld+json">
{"@context":"https://schema.org","@type":"MedicalWebPage","description":"PILOTTEST02"}
</script>
<!-- 03 JSON-LD in head, CDATA-wrapped (invalid strict JSON) -->
<script type="application/ld+json">/*<![CDATA[*/{…"description":"PILOTTEST03"}/**/</script>
<!-- 04 non-ld JSON MIME type -->
<script type="application/json">{"note":"PILOTTEST04"}</script>
<!-- 05 HTML comment (negative control) -->
<!-- PILOTTEST05 -->
<!-- 06 JSON-LD at top of body -->
<script type="application/ld+json">{…"description":"PILOTTEST06"}</script>
<!-- 07 identical JSON, non-script element, equally hidden -->
<div style="display:none">{"note":"PILOTTEST07"}</div>
<!-- 08 identical JSON inside a template element -->
<template>{"note":"PILOTTEST08"}</template>
<!-- 09 JSON-LD deeply nested, copying the client's page structure -->
<div class="part_rows_container"><div class="part"><div id="results">
<table><tbody><tr><td>
<script type="application/ld+json">{…"description":"PILOTTEST09"}</script>
</td></tr></tbody></table>
</div></div></div>
<!-- 10 microdata, value in a content attribute -->
<div itemscope itemtype="https://schema.org/MedicalWebPage">
<meta itemprop="description" content="PILOTTEST10">
</div>
<!-- 11 RDFa, value in a content attribute -->
<div vocab="https://schema.org/" typeof="MedicalWebPage">
<meta property="description" content="PILOTTEST11">
</div>
<!-- 12 accessibility attribute only -->
<span aria-label="PILOTTEST12" role="note"></span>
<!-- 13 visible body text (positive control) -->
<p class="lede">…This alternates with episodes of depression. PILOTTEST13</p>
<!-- 14 JSON-LD at very end of body -->
<script type="application/ld+json">{…"description":"PILOTTEST14"}</script>
How we made sure the test worked
A test like this can fail in quiet ways. So we built in four checks before trusting any result:
- Visible text must show up (canary 13). If a tool can’t see normal text, the fetch failed. We’d throw that run out.
- The HTML comment must not show up (canary 05). If it did, the tool was passing along raw code instead of reading the page. That would tell us nothing.
- All 14 canaries must be on the live page. We pulled the raw page code with Firecrawl and confirmed every canary was there. This rules out an upload mistake.
- The schema must still be there after the page loads. On our client’s page, JavaScript was deleting the schema about 1.5 seconds after the page opened. We checked the test page in a live browser to make sure nothing like that was happening. All five JSON-LD blocks were intact.
The tools we tested
| Tool | What it stands for |
|---|---|
| ChatGPT (its built-in browsing) | OpenAI’s own tool, the one ChatGPT uses to read a page |
| Gemini (its built-in browsing) | Google’s own tool, the one Gemini uses to read a page |
| Claude (its fetch tool) | Anthropic’s own tool, the one Claude uses to read a page |
| Firecrawl (markdown output) | A popular scraper many AI apps and agents are built on |
| Chrome reader extraction | The “reader view” style of pulling the main text from a page in a browser |
We also ran two Firecrawl modes that aren’t in the results table. Firecrawl’s raw HTML mode was only used to confirm the canaries were on the page. Its summary mode had an AI write a summary, so it didn’t pass along any canaries. That tells us nothing about what the tool could see, so we left it out.
We tested ChatGPT, Gemini, and Claude in separate chats with two prompts. The first was open-ended: what’s on this page? The second asked each tool to hunt for every PILOTTEST code word and quote the text around each one. Quoting the text nearby proves the tool really saw the canary and didn’t make it up. ChatGPT confirmed it read the page directly and didn’t use search. It had to, anyway. The page is set to noindex, so it isn’t in Bing’s index.
Results
All four checks passed. Every canary was on the page. Visible text (13) showed up in every tool. The HTML comment (05) showed up in none. The test was valid.
| # | Canary | ChatGPT | Gemini | Claude | Firecrawl | Chrome reader |
|---|---|---|---|---|---|---|
| 02 | JSON-LD, head, clean | ✗ | ✗ | ✗ | ✗ | ✗ |
| 03 | JSON-LD, head, CDATA-wrapped | ✗ | ✗ | ✗ | ✗ | ✗ |
| 06 | JSON-LD, top of body | ✗ | ✗ | ✗ | ✗ | ✗ |
| 09 | JSON-LD, buried deep | ✗ | ✗ | ✗ | ✗ | ✗ |
| 14 | JSON-LD, end of body | ✗ | ✗ | ✗ | ✗ | ✗ |
| 04 | Plain JSON in a script tag | ✗ | ✗ | ✗ | ✗ | ✗ |
| 07 | Same JSON, hidden div | ✓ | ✓ | ✓ | ✓ | ✗ |
| 08 | Same JSON, template tag | ✗ | ✗ | ✗ | ✓ | ✗ |
| 01 | Meta description | ✗ | ✗ | ✓ | ✓ | ✗ |
| 10 | Microdata | ✗ | partial | ✗ | ✓ | ✗ |
| 11 | RDFa | ✗ | ✗ | ✗ | ✓ | ✗ |
| 12 | Accessibility label only | ✗ | ✓ | ✗ | ✗ | ✗ |
| 05 | HTML comment (negative control) | ✗ | ✗ | ✗ | ✗ | ✗ |
| 13 | Visible text (positive control) | ✓ | ✓ | ✓ | ✓ | ✓ |
JSON-LD made it through 0 of 25 times. That’s five JSON-LD blocks, read by five tools. Not one got through.
Gemini’s “partial” on canary 10: it saw the microdata container and its schema type, but the description was empty. The tag made it through. The value inside it didn’t.
What the results mean
Where you put your schema doesn’t matter
Head, top of body, buried deep, end of body. We gave five tools 25 chances to find the JSON-LD, and they found it zero times. If someone tells you to move your schema to help AI see it, that advice isn’t solving anything.
The script tag is the problem
This is the most important result, and it comes from canaries 06, 07, and 08. All three hold the exact same JSON. All three are hidden from visitors. The only difference is the tag around them.
- In a <script> tag (06): zero of five tools saw it.
- In a hidden <div> (07): four of five tools saw it.
- In a <template> tag (08): one of five tools saw it.
So it isn’t the hiding that gets schema thrown out. The hidden div was just as hidden, and it came through. The <script> tag is what gets it tossed.
All scripts get tossed, not just schema
Canary 04 was plain JSON in a script tag, not labeled as schema. It got thrown out the same way. These tools aren’t singling out schema. They drop every script tag, no matter what’s inside.
Nobody is reading the code
Canary 03 was wrapped in CDATA, the way our client’s was. That wrapping technically breaks the JSON. It didn’t matter. It was thrown out exactly like the clean version. None of the tools looked at the code long enough to notice it was broken.
Don’t read too much into microdata and RDFa
Canaries 10 and 11 came through Firecrawl, which looks like a win for those formats. It isn’t. We placed their values inside <meta> tags, and Firecrawl copies every meta tag it finds into a summary of the page. That’s Firecrawl grabbing meta tags. It isn’t Firecrawl reading schema.
If the microdata had been on visible text, it would have come through as plain text. If it had been on some other hidden tag, it probably would have been lost. This test doesn’t show that microdata beats JSON-LD.
Meta tags depend on the tool
The meta description (01) came through Claude and Firecrawl. ChatGPT and Gemini missed it. Gemini told us why: its tool removes the whole <head> section and every meta tag before Gemini sees the page.
This is the one place the tools really disagreed. That’s useful to know. It shows these are design choices each company makes, and they can change. So don’t count on your meta description reaching an AI tool, either. Two of the biggest AI assistants never see it.
Accessibility labels: only Gemini
The accessibility label (12) was dropped by four tools but came through Gemini. Gemini quoted the tag and the label directly. Google has said its browser agents may use accessibility info, and this fits. But it’s one vendor, not a rule.
The AI tools told us themselves
We asked Gemini what tool it used. It said:
“This tool strips <script> tags, including Schema/JSON-LD, during content extraction, which is why other potential matches in the raw source code may not be visible here.”
And separately:
“The live browsing tool I have available returns a sanitized, body-only DOM representation of the webpage and strips out the <head> section and all <meta> tags before I receive the data.”
That’s a major AI tool describing, in its own words, how it throws out schema on purpose. It lines up with what the canaries had already shown.
Gemini also tried to get around its own browsing tool. It wrote three short programs to grab the raw page code. All three failed. Its coding sandbox isn’t allowed to reach the internet. So it had to fall back on the cleaned-up version of the page. The AI doesn’t get a choice here.
ChatGPT backed up the key result on its own. Its browsing tool returned two canaries: the hidden div (07) and the visible text (13). Nothing else. It found the hidden JSON and missed all five schema blocks. That’s the same script-tag pattern, from a completely separate company’s tool.
We also know ChatGPT didn’t make it up. It quoted canary 07 along with the heading that comes right after it on the page. That matches the real page, and it matches what Claude’s fetch returned word for word.
ChatGPT said two more things worth noting. It described what it gets as “a parsed page representation rather than the complete raw HTML source.” In other words, something cleans up the page before the AI sees it. And when we asked whether it saw any structured data, it said the JSON it found couldn’t count as schema. It was missing the @context and @type lines that mark JSON as schema. So even when some JSON gets through, the parts that make it schema don’t.
Why this happens
A browser can hand over a page’s text in two ways. One way grabs every bit of text in the code, including what’s inside script tags. The other grabs only what a visitor would actually see on screen.
We checked both on the test page:
document.body.textContent → PILOTTEST06, 07, 09, 13, 14
document.body.innerText → PILOTTEST13
The first way found the three JSON-LD canaries in the body, plus the hidden div and the visible text. The second found only the visible text.
Browsers never show what’s inside a script tag. Reader tools and page-to-text converters work from what the page shows, so script tags get thrown out early. That’s why JSON-LD never reaches the AI.
What about the hidden div (07)? It’s hidden too, so why did it survive? Most of these converters walk through the page one tag at a time. They drop script tags up front, before that walk starts. A hidden div isn’t a script tag, so it gets walked through, and its text gets picked up.
What this test can’t tell you
A result is only as good as its limits. Here are ours:
- We can’t see inside the crawlers. GPTBot and ClaudeBot collect pages for training, and we can’t watch what they keep. Three of our five tools are the AI companies’ own reading tools (ChatGPT, Gemini, and Claude). Firecrawl and Chrome’s reader stand in for tools many AI apps are built on. They aren’t copies of any one vendor’s setup.
- Some results rely on what the AI said. When ChatGPT found a canary, we checked it against the real page. But when it said it didn’t see something, like the meta description, we’re taking its word for it.
- We tested what reaches the AI, not what gets stored. A crawler could save the full page code and read the schema later, in a step we can’t see. That’s a separate question from what the AI reads when it fetches a page.
- One page, one site, one day. These tools change. Re-run the test before you quote these results as current.
- The page isn’t typical. Few real pages have five JSON-LD blocks. But all five were dropped, so the tools weren’t just skipping extras after the first one.
- Microdata and RDFa got mixed up with meta tags. See above. A follow-up test should put microdata on other hidden tags.
- The page was set to noindex. That shouldn’t change how a tool reads it, but we’re noting it.
- Static pages only. This test left out JavaScript on purpose. Pages that load content with JavaScript have their own problems, like our client’s.