Sites that speak for themselves: an agent endpoint for every small website

Sites that speak for themselves: an agent endpoint for every small website. A sketch of one small web page standing on its index, sending glowing ripples out to an agent core, a lens and a speech bubble.

The idea is a small open-source tool that lets a small business or personal website answer for itself: AI agents, the site’s own search box and a chat window, all served from one index built when the site is published. I keep coming back to it because the reader a small site was built for is changing.

For twenty years a small website had one kind of reader to plan for: a person who arrived from a search result and read the page. People now ask an assistant, and the assistant visits the site on their behalf. HUMAN Security reports that AI-driven traffic grew 187% between January and December 2025, and that traffic from agents acting on the web grew 7,851% year over year. The report gives no absolute volumes, so the second figure says more about speed than size. The direction is clear either way.

When an assistant visits a site today, it guesses. It downloads a page or two, strips away the menus and footers as best it can, and hopes it picked the right page. The owner has no say in what gets read and never learns what was asked. I keep wondering why the site can’t answer for itself. The big platforms have started to make that possible for their own customers. This is an idea for everyone else.

Who this is for

Small business websites and personal websites: a physiotherapy clinic, a family restaurant, a wedding photographer, an independent consultant, a local accounting firm, a writer’s portfolio, or the home page of an open-source project. These sites are small, usually a few dozen pages. Many were built once by a freelancer or from a template and are updated a few times a year. Nobody on staff is “the web team”.

They are also the sites where a wrong answer costs the most, because the questions are practical:

  • Is the clinic open on Saturday?
  • Does the restaurant have a vegetarian menu?
  • Does this consultant work with companies my size?
  • What does the photographer charge for a half day?

The answers are on the site. Whether the assistant finds them is luck. Today, when a person asks an assistant, the assistant downloads a few pages and guesses which parts matter, so the answer may be wrong or out of date. With a site that answers, the assistant asks the site, the site returns the right passages, and the answer links back to the page.

Person asks an assistant → Assistant downloads a few pages: Today; Assistant downloads a few pages → Assistant guesses which parts matter; Assistant guesses which parts matter → Answer may be wrong or out of date; Person asks an assistant → Assistant asks the site: With a site that answers; Assistant asks the site → Site returns the right passages; Site returns the right passages → Answer links back to the pagePerson asks an assistantAssistant downloads a few pagesAssistant guesses which parts matterAnswer may be wrong or out of dateAssistant asks the siteSite returns the right passagesAnswer links back to the pageTodayWith a site that answers

A large company can close this gap. It has a platform that builds the feature in, a budget for a hosted service, or engineers to wire something up. The small site has none of those.

What’s out there today

What exists today falls into four groups, and each leaves a small site owner without something.

Platforms that build it in. Every published GitBook site now includes a Model Context Protocol (MCP) server. Every Shopify store has a storefront MCP endpoint, and WordPress 6.9 shipped an official MCP adapter. That is excellent if your site lives on one of those platforms, and no help if it doesn’t.

Hosted services. Cloudflare’s AI Search will crawl your site, index it and put an MCP endpoint in front of it. kapa.ai and similar products offer a hosted MCP server connected to your documentation. They work, and they tie the site to one vendor or to a monthly bill.

Open-source “site to MCP” tools. SiteMCP fetches an entire site and serves it as an MCP server. Others crawl sites into a database for retrieval. Nearly all of them are built for the reader, not the owner: a developer points one at someone else’s documentation and runs it on their own laptop for their own coding assistant.

Proposed standards. NLWeb, from Microsoft, is an open project for giving a website a natural-language interface. It is a protocol with reference code, not a product you switch on. WebMCP lets a page offer tools to an agent inside the browser. It is in a Chrome trial, and no mainstream agent uses it yet. llms.txt is a file that describes a site for language models, and Google has said it does not support it.

OptionBuilt forWhat a small site owner is left without
Platform built-insCustomers of that platformAnything at all, if the site is hosted elsewhere
Hosted servicesTeams with a budgetControl over where the index lives and what it costs
Open-source crawlersA developer reading other people’s sitesA public endpoint the owner runs for their own site
Proposed standardsThe whole web, eventuallySomething that works this year

The closest things

Two projects come closest, and neither does the whole job.

The first has nothing to do with AI. Pagefind is a search library for static sites. It runs right after the site is built, reads the finished pages, and writes its search index as ordinary files next to the site. There is no search server and no account. The index is a by-product of publishing, and that is the model I’d borrow.

The second is an open-source project that turns a static site into an MCP-searchable knowledge base. It builds an index when the site is built and serves it to agents from Cloudflare. It proves the idea works. It is also narrow by design: it reads Markdown source files through adapters for particular site generators, it matches keywords only, and it runs on Cloudflare alone.

What’s missing is the general version: any small site, on any host, with search that understands meaning as well as keywords, and more than one way to ask.

What it could look like

Two small pieces, with one index between them.

  • An indexer that runs when the site is published. It reads the finished pages, keeps the real content, drops the menus and footers, and writes the site index. It follows the rules the site already publishes for search engines.
  • A small answering service, always available. It loads that index and answers questions from it. It holds no other data, and there is nothing to administer.

One index, built at publish time, serves three kinds of visitor: AI agents, the site’s own search box and a chat window.

When the site is published: Finished pages, Indexer, Site index; Always available: Answering service; Finished pages → Indexer; Indexer → Site index; Site index → Answering service; Answering service → AI agents; Answering service → The site's own search box; Answering service → Chat windowWhen the site is publishedAlways availableFinished pagesIndexerSite indexAnswering serviceAI agentsThe site's own search boxChat window

From the owner’s side there should be almost nothing to do. Whoever maintains the site sets it up once. After that, every time the site is published the index is rebuilt, so the answers change when the pages change. There is no second copy of the content to keep in sync, no dashboard to log into, and nothing new to write.

The answering service has three doors:

  • For AI agents: an MCP endpoint with two abilities, search the site and read a page in full.
  • For visitors: a search endpoint that the site’s own search box can call, using the same index.
  • For conversation, later: a chat endpoint that answers in sentences from the site’s content and links to the pages it used.

A few principles I’d hold to:

  • One site, one index. No accounts and no dashboard.
  • Every answer points back to a page on the site.
  • Search works by meaning and by keyword together, so “opening hours” finds the page that says “we’re open from nine”.
  • It runs on any host, cheaply enough that a personal site can afford it.
  • It is open source, so a small owner isn’t trading one vendor for another.

Here is what a visit looks like once the site can answer. A person asks their AI assistant whether a studio teaches beginners. The assistant searches the site and gets matching passages with page links, reads the full page, and answers with a link to the page.

Person → AI assistant: Does this studio teach beginners?; AI assistant → The site: Search the site; The site → AI assistant: Matching passages with page links; AI assistant → The site: Read the full page; The site → AI assistant: Page content; AI assistant → Person: Answer, with a link to the pageThe siteAI assistantPersonDoes this studio teach beginners?Search the siteMatching passages with page linksRead the full pagePage contentAnswer, with a link to the page

What the owner gets

The owner gets five things, the first of them a say in what agents read.

  • A say in what agents read. The index holds what the owner chose to publish, cleaned up once, not whatever a scraper managed to salvage. If the price list page is the truth, the assistant quotes the price list page.
  • A view of the questions. The owner can see what agents and visitors ask, and which questions found nothing. If ten people a week ask whether the clinic takes a certain insurer and the site never says, the owner now knows what to add.
  • A way back to the site. Because every answer carries a link, the person who asked can click through to book, call or buy.
  • A better search box as a side effect. The same index that serves agents serves the site’s own visitors.
  • Independence. The index is a file the owner keeps, and the service can move to another host with it.

What it is not

The point is that the site’s own content becomes easy to ask, so it helps to say what this is not.

  • Not a chatbot product. Chat is the last of the three doors, not the point.
  • Not a ranking trick. It does not make an assistant choose your site. It makes the answer right once the assistant gets there.
  • Not a second website to maintain. The index is built from the pages that already exist.
  • Not for private content. It covers only what the site already shows the public.

What I’m unsure about

Discovery is the question I can’t answer yet, and it isn’t the only one.

  • Will agents find it? There is no settled way for an assistant to discover that a site has its own endpoint. A proposal for a discovery file exists, and it is still under review. Until that settles, an owner has to register the endpoint by hand or advertise it on the site. I wrote about this gap in MCP is in its Yahoo era.
  • Will hosting companies absorb it? If every host ships this as a checkbox, a separate tool matters less. I think a portable, open version still earns its place, but that is a bet.
  • Messy pages. The indexer can only be as good as the pages it reads. A site with clear structure will index well. One built from tangled templates won’t, and I’d rather not ask owners to configure their way out of that.
  • Stale answers. The index is rebuilt when the site is published. A site nobody has touched in two years will give two-year-old answers, and an assistant will repeat them with confidence. Answers probably need to carry the date the page last changed.
  • Questions are personal. People tell an assistant things they would never type into a search box. If the owner can see the questions, the tool has to be careful about what it keeps and for how long.
  • Who pays for chat? Search costs the owner almost nothing. Chat spends money on a language model with every question, and a public chat box invites abuse. It needs firm limits before it is safe to switch on.

Where to start

I’d start with the indexer and the agent endpoint on one small static site, and add a door at a time, each step reusing the index from the step before.

  1. Build the indexer and the agent endpoint, and prove them on one small static site.
  2. Add the search endpoint, and replace that site’s search box with it.
  3. Add chat, with the limits in place first.
  4. Only then look at larger sites, which need a different kind of indexer.

A website has always been written for people to read. The next reader is an assistant asking on someone’s behalf, and a small site should be able to answer it in its own words.

If you’d like to build this, with me or without me, or to help fund it, I’d love to hear from you.

I’m Amir Pournasserian. I build AI and platform systems for a living, maintain FluentCMS and YeSvelte, and write here about what I find along the way.