JOURNALDIGITAL
11 AUG 2026/ 10 MIN/ Melih Yiğit, Dijital Pazarlama Uzmanı - Kurucu

ChatGPT’s in-house index includes small sites: Why the H1 and first 200 characters matter

New research finds small sites in ChatGPT’s in-house index and frequent use of the H1 and page opening in snippets.

Coral, black, and bone editorial cover with CHATGPT SEARCH and small and large web cards entering one retrieval aperture

A new reverse-engineering study of ChatGPT web retrieval indicates that OpenAI's in-house search index includes small sites and publishers without licensing deals in free, instant answers. The same research found that citation candidates from this index often use the page H1 and roughly the first 200 characters of the rendered body.

This does not establish “the first 200 characters” as a new ranking factor. The study is not official OpenAI ranking documentation; it is a data-based snapshot of product behaviour observed in July 2026. Free, paid, instant, and thinking modes can use different retrieval routes. The “result_source” label that exposed those routes was removed on July 21, making some technical details harder to observe directly. Even so, the findings offer a practical audit framework for smaller sites that should not assume AI search visibility is limited to large licensed publishers.

What did the study find?

Resoneo researchers examined 1,249 ChatGPT answers that used search. The sample contained 682 instant answers and 567 thinking-mode answers across free, paid, signed-in, and signed-out sessions, multiple countries, and several accounts. Network logs exposed more than 27,000 pages. The team used historical “result_source” labels, request formats, and returned snippets to distinguish retrieval routes.

In stable samples of free instant answers, an in-house index labelled “labrador” appeared frequently. Large licensed publishers and smaller unlicensed sites could travel through the same result route and arrive in a similar output format. “Same route” does not mean equal ranking, equal citation probability, or equal visibility. It means the small site was not technically excluded from that in-house corpus.

Paid thinking mode showed a different mix. In a sample of about 16,407 search results, roughly 74.6% came through a “bright” route associated with Google result retrieval, while about 23.6% came from the in-house “labrador” index. These percentages are not permanent product quotas. They describe the chosen accounts, countries, prompts, and July 2026 observation window.

Text-free diagram showing large publishers and small-site cards moving through one retrieval tunnel into a shared result format
Fark Studio illustration of small sites entering an in-house ChatGPT retrieval route.Source: Fark Studio

This text-free Fark Studio illustration shows large publishers and small sites entering the same in-house retrieval tunnel. A uniform card format does not claim equal ranking or citation probability, and the visual is not a ChatGPT interface.

An independent technical review supports the variability. Suganthan Mohanadasan documented small Italian sites arriving through “labrador” in free-tier tests from Italy, while different accounts, countries, and cohorts did not always show the same route. The durable takeaway is not one percentage: ChatGPT does not behave like one fixed web search system, and the distribution observed in one account should not be generalised to every user.

Why do the H1 and first 200 characters matter?

Resoneo compared 534 cited pages in detail. Of those pages, 463 had an H1, and 387 snippets contained the H1 text. That is an overlap of roughly 83.6% among pages with an H1. Snippets from the in-house index averaged about 202 characters and appeared to come primarily from the start of the rendered body rather than the meta description.

The finding exposes a small but consequential page-design issue. If a cookie notice, repeated navigation text, campaign ribbon, unhelpful accessibility label, or very long image alt text precedes the H1, part of a narrow snippet budget can be consumed before the subject sentence appears. About one in seven pages in the sample had no H1. That is not merely a ChatGPT issue; it is a weak information architecture signal for screen readers, search engines, and human scanning.

Text-free comparison of cluttered and clean page openings passing through a narrow extraction window
Fark Studio illustration of the H1 and opening paragraph in citation context.Source: Fark Studio

The illustration compares cluttered page-top elements with a clean H1 and introduction as they pass through a narrow extraction window. The window is not a product interface and does not promise a fixed character limit.

The observed 202-character average must not become a new magic number. The research reports snippet length, while OpenAI has not said it is a ranking or citation rule. A good page does not shorten its opening only for a bot. Its heading and introduction are edited so a person can quickly understand what the page provides.

What is confirmed and what remains uncertain?

OpenAI's official crawler documentation explains that OAI-SearchBot crawls sites for search results, GPTBot can be used for model training, and ChatGPT-User may visit a page following a user-triggered request. Site owners can manage these bots separately through robots.txt and security controls. The official page does not document an index called “labrador”, its ranking signals, or a snippet-generation formula.

The in-house index and measured distributions must therefore be presented as reverse-engineering evidence. The “result_source” label disappeared on July 21, and the system may have changed since the study. A page appearing in a result set does not guarantee that the answer will cite it. Being crawled, indexed, retrieved, and presented as a visible source are four separate stages.

Freshness is uncertain too. Resoneo observed that some snippets appeared frozen at crawl time and that about 13% in one dataset were more than a month behind. That figure is not a universal freshness level for the web. Teams responsible for news, pricing, stock, or regulatory pages should track their own examples with a dated record.

What should brands in Türkiye do now?

1. Verify OAI-SearchBot access. Review robots.txt, CDN, WAF, and bot-protection rules using the official user-agent and IP-verification guidance together. Permission in robots.txt does not guarantee the security layer passes the request.

2. Use one descriptive H1. Name the topic, product, or question in natural language. A logo, decorative slogan, or several competing H1 elements weakens the page's first context.

3. Clean up the content before the H1. Inspect cookie copy, hidden navigation repeats, campaign banners, and misconfigured alt text in rendered DOM order. An element that is not visually prominent can still be read first in the source order.

4. Make the opening paragraph a useful summary. State the solution, scope, and important limitation in the first two or three sentences. Do not stuff keywords or lead with a generic paragraph full of company names.

5. Do not rely on the meta description alone. The study indicates that in-house snippets often use the body opening. The meta description still matters for search presentation, but it is not a hidden summary that repairs weak on-page text.

6. Match structured data to visible content. Organization, Article, Product, or FAQ markup must agree with the information a visitor can see. Schema does not repair a missing H1 or an unsupported claim.

7. Build evidence of authority for a small site. Make the author, date, sources, company identity, expert review, and update history visible. Entry into an in-house index does not guarantee selection as a trusted answer source.

8. Test free and paid modes separately. Run the same query set from signed-in and signed-out Türkiye sessions and, where possible, different account cohorts. Even if the route label is gone, record the cited page, quoted passage, and date.

9. Preserve server logs. Verified OpenAI bot requests, crawled URL, status code, and date are the strongest operational evidence when visibility changes. Never accept a user-agent string alone as verification.

10. Audit on a schedule. Recheck the same 20 to 30 commercial and informational queries each month. Product interfaces and provider mixes change quickly, so one day's screenshots should not become a permanent strategy.

Text-free audit loop connecting bot access, page cleanup, country and account cohorts, dated evidence, and retesting
Recurring technical audit loop for AI search visibility. Fark Studio illustration.Source: Fark Studio

This Fark Studio audit loop connects bot access, page-top cleanup, account and country cohorts, dated evidence, and repeat testing. It is not an official OpenAI optimisation diagram.

What should teams avoid?

Do not pack a keyword list into the first 200 characters. Do not rewrite the H1 solely for ChatGPT while leaving a vague opening for users. Do not interpret inclusion of small sites as proof that authority no longer matters. A lack of a licensing deal may not block technical access, but trust, original information, freshness, and query fit can still affect selection.

Do not present one third-party study as an official OpenAI ranking rule. Likewise, do not claim that allowing the official crawler guarantees visibility. Technical access can be necessary without being sufficient. Measurement should connect “did the bot visit?” to “which claim from which URL appeared in what answer context?”

Fark Studio perspective

The study identifies a credible opportunity for smaller brands: ChatGPT's in-house route does not appear limited to large licensed publishers. The opportunity is not a shortcut. It rewards clean information architecture and verifiable, original information. SEO and web design should review the page opening together so the visual hierarchy and rendered reading order do not conflict.

Digital marketing teams should also avoid reducing AI visibility to a single tool's mention count. Query set, account type, country, date, cited URL, and business outcome belong in the same record. If you want to assess access, page structure, and evidence quality for both ChatGPT and classic search, plan an AI-search visibility audit with Fark Studio.

Sources

Resoneo, Inside ChatGPT's in-house search index, July 2026; a reverse-engineering study of 1,249 answers and more than 27,000 pages, with no exact publication day stated.

Search Engine Journal, ChatGPT's search index serves small sites too, data shows, August 11, 2026; industry summary of the research.

Suganthan Mohanadasan, How ChatGPT picks sources, Part 2, published July 14 and updated July 22, 2026; independent network testing across account and country cohorts.

OpenAI Developers, Overview of OpenAI crawlers, accessed August 12, 2026; current official documentation for OAI-SearchBot, GPTBot, and ChatGPT-User.

02 — Next

Keep reading.

START A PROJECT

Let's talk about your next difference.

Schedule a free consultationinfo@farkworks.com