Skip to content
Case Study

Building Breezy AI Search: Semantic Indexing, Retrieval, and RAG in Production

Breezy Team ·Sep 26, 2026 ·8 MIN READ
Discover finds and reads every page Index meaning as numbers, kept up to date Retrieve keyword + vector, blended ranking Answer grounded, cited, streamed revisit on a schedule, skip what has not changed the visitor’s question “not covered” = content gap
Breezy AI Search
Product
Breezy AI Search
Built by
A lean team at WhooshPro, working with AI agents
Timeline
About six months from first commit to launch

TL;DR: We built Breezy AI Search in about six months: a crawler that keeps a site’s content current, an index that stores meaning as embeddings, retrieval that blends keyword and vector matching, and answers grounded in the site’s own pages. This is how it came together, and the principles we built it on.

The brief we set ourselves

The goal was narrow on purpose. A site owner pastes one script tag. From then on, a visitor can ask a question in their own words and get an answer built from that site’s content, with a link back to the page it came from. If the site doesn’t cover the question, the answer says so.

That one sentence hides four stages: crawling, semantic indexing, retrieval and generation. We explain how those work in general in How AI Search Works: From Semantic Indexing to RAG. This piece is the other half: how we put them to work for real sites.

From prototype to production

Before the production build there was a prototype that proved the basic loop: crawl a site, index it, search it. Production meant everything a demo leaves out: sites built in every imaginable way, content that keeps changing, answers that have to be trusted, and running it all economically. Launch came after 33 releases.

We built it as a lean team working with AI agents, with a written spec behind every change of any size. There were nearly 500 before launch.

At a glanceFrom first commit to launch
6
Months from first commit to launch
33
Tagged releases before launch
3,400+
Commits before launch
~500
Written specs behind them
Month 1
Production build starts. A working prototype already proved crawl, index and search.
Month 2
Site discovery. Breezy finds a site’s pages by following links, with no sitemap needed.
Month 4
Inline citations. Underlined phrases in an answer link to the page they came from.
Month 5
Reliability hardening and streaming answers.
Month 6
Follow-up questions and answers in the visitor’s own language.
Launch
1.0 and the first public trials.
After launch
Crawler reliability, index integrity, and first scans that draw on sitemaps as well as links.

The design in one table

StageWhat we chose
DiscoveryEach page is read the way a visitor sees it, with its main content picked out.
IndexingPages are broken into chunks that keep ideas together, and each chunk becomes an embedding of its meaning.
RetrievalKeyword and vector matching blended into one ranking, with checks that know when nothing is a good match.
AnswersThe best few passages go to the model under one of three trust levels, with links verified in code and answers streamed to the visitor.
FreshnessA scheduler keeps each site current, skipping pages that haven’t changed.

The crawler: built for the open web

Every site is built differently, so the crawler reads each page the way a visitor sees it and picks out the content that matters. What gets indexed is the substance of the page, whatever kind of page it is.

A new site’s first pages become searchable within seconds, so a visitor can search while the rest of the site is still being read. We also keep crawls lean, which makes them more than twice as fast on some sites.

2x+
Faster crawls on some sites

The open web is unpredictable, so the crawler is built to recover from the ordinary hiccups of real sites. It retries in a way that gives a site time to recover, reports failure reasons that say what really happened, draws on sitemaps as well as links, and revisits known pages even when nothing links to them.

Semantic indexing: keeping embeddings in step with the site

Each chunk becomes an embedding, a numeric representation of what it means. That is what lets a question find a page that never uses the same words.

The index is built to stay in step with the site. Pages that haven’t changed are skipped, so a recrawl only does work when there is something new. When a page changes, its old chunks are replaced. A page is removed only once it is confirmed gone, and it counts as indexed only when it is genuinely searchable.

Retrieval: keywords and meaning, together

Vector search is excellent at meaning and less suited to exact things: a product code, a person’s name, a policy number. So retrieval blends keyword matching with vector similarity into one ranking.

Vector search also always returns the nearest thing it has, even for a question the site never covers. So “we don’t know” can’t come from retrieval alone. We tune retrieval against real questions so that specific questions still find their answer, and honesty is enforced in the answer step.

RAG: grounding is a product decision

Retrieval-augmented generation hands the retrieved passages to a language model along with the question. The rule is simple to say: answer from those passages. Making it hold took several decisions.

Answers in the site’s own voice

Answers speak as the site itself, in a brand-voice persona, and never talk about their own machinery. Every passage the model sees carries its page title and exact address, so it can send a visitor straight to the right page, such as a contact page. Giving the model what it needs to answer matters as much as ranking the passages well.

Three trust levels

Site owners choose how tightly answers stay tied to their content. Each level is a different set of instructions to the model.

3
Trust levels for answers
Strict

Answers only from the site’s content, and declines when the retrieved content doesn’t cover the question. The default.

Balanced

Mostly the site’s content, may add one clearly general point, still prefers declining to inventing specifics.

Creative

Conversational and freer with general knowledge, and rarely declines.

If the setting is ever missing or unrecognised, it falls back to Strict, so grounding is never loosened by accident.

Rules that always hold, enforced in code

A prompt can ask the model to cite its sources, but it can’t guarantee it. So links are verified in code: an answer can only link to pages that were actually retrieved. Site owners can shape the voice and focus of answers with their own instructions, and the citation rules always apply on top.

Honest answers, and the owner’s own words

When the content can’t answer a question, the model declines in a recognisable way, and we log that question as a content gap, which is how a site owner learns what their pages are missing. And a site owner can write the answer to a specific question, in which case we return their words as written.

Follow-up questions

Follow-ups are search that remembers, not a chatbot. A question like “what about the enterprise tier?” only makes sense in context, so Breezy carries the conversation forward and searches on what the visitor actually means.

Every visitor, in their own language

Visitors ask in their own language and get answers in it, even when the site’s content was written in another. The widget’s own interface speaks many languages too, so every visitor feels at home.

Speed: answers that arrive as they are written

A visitor shouldn’t wait on the slowest step. Results appear right away while the AI answer streams in as it is written, so the page feels responsive even while the model is still working. Streaming an answer that is checked as it goes takes care, and we test changes side by side, keeping the ones that make answers faster without making them worse.

Reliability principles we built in

Across the build, one theme kept returning: a system people can trust is one that makes its own health visible.

  • Say what really happened. Setup and crawl errors surface with the real reason, so they can be acted on.
  • Check health before spending effort. A crawl checks its dependencies first and stops with one clear message if something is wrong.
  • Give problems another chance. Spaced and whole-crawl retries smooth over the ordinary hiccups of real sites.

Five things we would tell anyone building AI search

01

Make health visible. Record the real reason when something can’t be done, check before you spend effort, and treat a job that stops as an event worth noticing.

02

Give the model what it needs. A perfect ranking is wasted if the text handed to the model is missing something the answer needs, like a link.

03

Decide where “I don’t know” lives. Vector search always returns something, so build the honest refusal into the answer step.

04

Enforce in code what must always be true. Prompts set style and tone. Code guards correctness, like which links an answer may contain.

05

Test side by side, and be willing to cut. More is not automatically better. Keep what measurably helps and let go of the rest.

If you’d rather not build and maintain this yourself, that is what Breezy AI Search is for. It runs this pipeline on your own site’s content and, by default, answers only from it, with a trust setting if you want to give answers more room.

What’s next

Launch was the starting point. We keep shipping improvements with the same discipline, and there is much more planned: new capabilities for site owners, and more Breezy products on the horizon.

Nicholas on the Build

“A demo only has to work once. A product has to work on every site, every day, including the sites we have never seen. Getting there meant a step into the unknown: building with an agentic workflow, a first for us.

It worked because we already knew our stack deeply, which let us direct the agents well. What I value most is the time it gave back. Less manual coding and troubleshooting, more architecture and thinking, and room for our other projects.”

— Nicholas Chua, Technology Director, WhooshPro

Don’t just read about search.
Try what we built.

Every guide here comes from building Breezy AI Search itself — add it to your own site and see the difference firsthand.

Simple pricing
Growth
For individuals getting started
US$29/mo
Professional Popular
For growing teams
US$49/mo
Scale
For businesses at scale
US$149/mo
See full pricing →