Transform your existing website content into five coordinated, machine-readable assets designed to help AI crawlers, large language models, enterprise retrieval systems, and intelligent agents understand your digital presence more clearly.
Traditional websites distribute important information across pages, navigation structures, metadata, sitemaps, and structured markup. That can make it difficult for automated systems to identify your most authoritative content, understand relationships between pages, and retrieve the right information efficiently.
The ThatWare AI platform converts your existing website into a coordinated set of machine-readable resources, giving AI systems a clearer path to your content, entities, priorities, and source relationships.
Learn more about AI readinessEstablish Your Website’s AI Guidance Layer
ai.txt acts as the foundational entry point within the ThatWare AI Stack. It provides a concise, machine-readable reference that helps organize how your website presents its AI-facing resources.
Give automated systems one location from which to identify your website’s AI-ready resources.
Organize website-level instructions and important source references in one file.
Start improving machine readability without deploying the complete five-file stack.
# Example: ThatWare AI
# AI Guidance File
site: https://www.example.com
name: Example Company
description: Example website description
for AI readiness and digital growth.
resources:
- llms.txt: https://www.example.com/llms.txt
- ai-manifesto.json: https://www.example.com/ai-manifesto.json
- vector-feed.xml: https://www.example.com/vector-feed.xml
- semantic-sitemap.xml: https://www.example.com/semantic-sitemap.xml
version: 1.0
last_updated: 2025-05-20
Curate the Content LLMs Should Understand First
llms.txt provides a structured, Markdown-based overview of a website and its most important resources. It can present priority URLs with descriptions that explain what each resource contains and why it matters.
The proposed llms.txt specification is intended to provide curated LLM context and coexist with existing web standards such as sitemaps and robots.txt. Adoption may vary between AI platforms.
Highlight the pages and resources that best represent your organization.
Explain what important URLs contain rather than presenting an unannotated URL list.
Guide retrieval systems toward primary product, service, or knowledge resources.
# Example: llms.txt
site: https://www.example.com
content:
- homepage
- products
- services
- documentation
priority: high
format: machine-readable
version: 1.0
Define Your Brand, Content, and Governance in Structured JSON
ai-manifesto.json provides a structured representation of the organization behind the website, its primary entities, its content architecture, and its preferred source hierarchy.
This should be presented as a ThatWare AI structured asset, not as a universally adopted web standard.
Present your organization, products, services, and content categories in a standardized structure.
Communicate source ownership, preferred references, and content policies.
Provide JSON-formatted information that can be processed by enterprise applications and data pipelines
{
"name": "Example Company",
"type": "Organization",
"description": "Example AI-ready
organization",
"entities": [
"Example Product",
"Example Service"
],
"ai_ready": true,
"version": "1.0"
}
Prepare Website Content for Retrieval and Vector Workflows
vector-feed.xml organizes website content and metadata into a structured feed designed to support ingestion, chunking, embedding, retrieval, and enterprise RAG workflows.
It should be positioned as a preparation and interoperability asset—not as a guarantee that external AI providers will ingest the feed.
Reduce the amount of custom extraction needed before website content enters a retrieval pipeline.
Connect each content segment to its canonical page and stable identifier.
Help downstream systems identify content that has changed and may require reprocessing
<vector-feed>
<site>https://www.example.com</site>
<document>
<url>/about</url>
<title>About Us</title>
<priority>high</priority>
</document>
<document>
<url>/services</url>
<title>Services</title>
</document>
</vector-feed>
Map What Your Content Means—not Only Where It Lives
A traditional sitemap primarily helps systems discover website URLs and updates. semantic-sitemap.xml extends that concept by describing content types, topics, entities, and relationships within the website.
It should complement the standard sitemap.xml, not replace it. Google describes conventional sitemaps as a mechanism for informing search systems about new or updated website pages.
Show how pages, entities, topics, and content clusters relate to one another.
Expose relationships that may otherwise remain hidden inside navigation or internal links.
Help machines explore large websites by meaning, content type, and hierarchy.
<semantic-sitemap>
<entity id="company">
<name>Example Company</name>
</entity>
<relationship>
<source>company</source>
<target>services</target>
<type>provides</type>
</relationship>
<relationship>
<source>services</source>
<target>products</target>
<type>includes</type>
</relationship>
</semantic-sitemap>
Create clear, centralized resources through which automated systems can locate your most important content and AI-facing assets.
Define your organization, products, services, content categories, and canonical sources in coordinated machine-readable formats.
Prepare website content for retrieval, vectorization, internal search, assistants, and RAG applications without beginning every project with a custom scraping pipeline.
Generate coordinated website assets from a centralized platform instead of manually authoring and maintaining each file.
Connect content records with source URLs, ownership information, modification dates, and preferred references.
Apply one coordinated AI-readiness architecture across extensive content libraries and multi-site environments
ThatWare AI discovers accessible pages, content types, metadata, internal relationships, and important entities.
The platform organizes website content into source records, topic groups, entities, hierarchy, and machine-readable metadata.
ThatWare AI produces the files included in your selected plan using one coordinated content model.
Inspect coverage, file health, excluded URLs, missing information, and structural issues before deployment.
Publish the generated assets to the appropriate root locations and regenerate them when your website content changes.
Deploy a coordinated AI-readiness architecture across large websites, mobile brands, regional domains, product portfolios, and enterprise knowledge environments.
Organize the machine-readable layer that helps crawlers, LLMs, retrieval systems, and agentic AI understand what your site contains, where it matters, and how it relates.