Your best ideas and deepest thinking are invisible to the LLMs. You’re deeply confused. You’re convinced your whitepaper with proprietary data and rich insights is exactly the kind of material that should be a go-to citation in AI search.
Yet it’s not.
The problem is that most likely your work lives behind a form, saved as a PDF.
That’s great for building your CRM list, but it’s not doing anything for your discoverability in the new world of zero-click search.
The Bots Don’t Want to Read a PDF
PDFs are built for print. To an LLM, they look like a set of instructions for drawing shapes on a page. There’s no map of what’s a headline, what’s a caption, and what’s a data table. A parser has to reconstruct that structure from scratch, and complex PDF layouts routinely break it, especially once you add embedded images or multi-column design.
Some bots skip PDFs altogether. OpenAI’s GPTBot, for example, only reads static HTML. If your content isn’t already sitting there in clean markup when the bot arrives, it doesn’t exist to the model.
OtterlyAI ran a controlled experiment publishing matched HTML and Markdown versions of the same pages, with equal internal linking to both. Over 14 days, the HTML pages picked up AI crawler traffic and citations. The alternate-format versions got zero citations. If a format as close to plain text as Markdown gets shut out, an image-heavy, gated PDF doesn’t stand a chance.
Gating Makes It Worse
The instinct to gate your best material runs counter to AI visibility goals.
Gated PDFs create a second layer of friction on top of the format challenge. Unlike a normal HTML landing page, a PDF behind a form typically has none of the structured markup that helps a crawler understand what it’s looking at.
With AI Overviews now appearing on roughly 1 in 5 Google searches, and citation in those overviews associated with organic click-through rates nearly double the baseline, it could mean the difference between securing a lead or handing it to a competitor. Your competitor’s ungated, HTML blog post is winning the citation, and not because the content is better.
The Fix
Don’t let your best content sit in a digital vault.
Review every PDF asset that contains original data, research, or a proprietary framework.
Convert to structured HTML with real headers, paragraph tags, and a logical hierarchy a parser can walk through.
Publish the core findings, the stats, and the citable claims in the open. Save the appendix, raw dataset, or interactive tool for the gated version, if you still want one.
Add schema markup to every converted page. Give the crawler a shortcut to understanding what the page is and what claims it’s making.
Read the full playbook for any brand or leader who wants to stop being invisible and start being the authority here.


