Oct 2026

A free RAG chatbot: a free-model router, retries and a decision judge

Paper cards travel along a conveyor toward a gate of two white posts; the crumpled ones fall to the floor and only one, flat with an orange edge, gets through.

This site has an assistant, the sparkles button at the bottom right, that answers questions about my work using only what is on the site and in my CV, and cites where each claim comes from. It costs zero a month: the model that answers is free and the function that calls it runs on Vercel’s free plan.

Free comes with a catch, of course. Free models fail more often, take longer and sometimes don’t do what you ask. This post explains how it is built and, above all, the two real failures I saw in production on day one and how I fixed them. I will keep it updated with usage data.

The whole flow

browser POST /api/ask (Go) origin · token · per-IP limits topic gate Mercury · optional openrouter/free /ask/en.json numbered sources cleanup + heuristic judge: Mercury Decide rejected or failed: retry (max. 3) answer with [n] citations
Every question goes through the Go function, the free router and the judge. If the judge rejects the text or something fails, another answer is requested, which may come from another model.

There are only a few pieces:

  • The sources: a static JSON file that Astro generates at build time.
  • A Go function: it runs on Vercel and stores nothing.
  • Two OpenRouter models: one that writes (openrouter/free) and one that decides (inception/mercury-decide:free).

The sources: RAG without a vector database

The whole site fits in a model’s context. The build generates /ask/es.json and /ask/en.json with numbered sources: the intro, each project, my career path, how I work, the CV, the case studies and an excerpt of every post. About 15,000 tokens in total. No chunking or similarity search needed: everything goes along with every question.

// Knowledge renders the sources as the cacheable system block.
func Knowledge(sources []Source) string {
	var b strings.Builder
	b.WriteString("<sources>\n")
	for _, s := range sources {
		fmt.Fprintf(&b, "<source id=\"%d\" title=%q url=%q>\n%s\n</source>\n", s.ID, s.Title, s.URL, strings.TrimSpace(s.Text))
	}
	b.WriteString("</sources>")
	return b.String()
}

The instructions are short and strict:

  • Answer only from the sources, citing each claim with [n].
  • Use two to five sentences, in the visitor’s language.
  • Talk about me in the third person.
  • Don’t invent figures or internal company details.
  • Decline anything that isn’t about me in one sentence.

The function reads the citations in the answer and returns only the cited sources, so the browser can render them as links.

Building the sources with the site has an advantage I didn’t expect: the assistant is never out of date. If I publish a post, the next deployment already knows about it.

The free model and its two failures

openrouter/free is a router: it sends each request to one of the free models available at that moment. It is the simplest way to pay nothing, but it means every question may be answered by a different model, each with its own quirks.

On day one in production I saw two failures.

1. «The assistant is not available right now». Some models took too long or returned an error, and the function gave up on the first failure.

2. The reasoning instead of the answer. To the question «How does he work with teams?», asked on the Spanish site, one model returned this, in English:

The user is asking “How does Pelayo work with teams?”. I need to answer based only on the provided sources. Let me look for information about how Pelayo works with teams. Looking through the sources, I find relevant information in source 10…

That was its reasoning, with no tags, placed in the answer field. Cut off, too, because it had used up the token limit thinking.

First defence: ask for no reasoning, and clean up

OpenRouter has a parameter so that models that keep their reasoning separate don’t return it. You also have to allow more tokens, because reasoning counts towards max_tokens even when it isn’t returned:

payload, err := json.Marshal(map[string]any{
	"model": model,
	// Reasoning models spend tokens thinking before they answer; leave room for both.
	"max_tokens": 1500,
	// Keep the thinking out of the answer.
	"reasoning": map[string]any{"exclude": true},
	"messages":  msgs,
})

That isn’t enough for models that write their reasoning inside the text. For those there are two filters:

  • A regular expression removes <think>…</think> blocks, even cut-off ones.
  • A heuristic, Leaked, looks for how a model starts thinking out loud. It checks whether the text starts with phrases like «The user», «Okay,», «Let me» or «El usuario», or says «the user is asking», «according to the rules» or «source 10» anywhere. A good answer cites [10]; it never says «source 10».

The heuristic has tests with the real leaked text and with good answers, to avoid false positives. But it is a list of phrases, and models are very imaginative.

Second defence: retry with another model

Since the router picks a different model each time, what fails with one usually works with the next. The function retries up to three times while there is time left:

for try := 1; ; try++ {
	text, err = cfg.Complete(ctx, system, turns)
	if err == nil {
		if _, err = cfg.judge(ctx, req, text, try); err == nil {
			break
		}
	}
	if errors.Is(err, ErrQuota) || try == maxTries || ctx.Err() != nil {
		return response{}, fmt.Errorf("model (try %d): %w", try, err)
	}
	if dl, ok := ctx.Deadline(); ok && time.Until(dl) < 6*time.Second {
		return response{}, fmt.Errorf("model (try %d, out of time): %w", try, err)
	}
	log.Printf("ask: try %d: %v", try, err)
}

Three details matter:

  • Spent quota: it isn’t retried, because it won’t come back in two seconds.
  • Time: the whole question has 45 seconds, and each model call 20. With less than 6 left, no new attempt starts.
  • Vercel: the function needs maxDuration: 60 in vercel.json, or Vercel cuts it off sooner.

The judge: a decision model, not another chat

The heuristic catches what I have already seen. For what I haven’t, I use a judge: before showing an answer, another model decides whether it really is one.

I could use another chat model with a prompt like «is this an answer? Reply yes or no». But then I would have to parse free text, and the judge could fail in the same ways as the model it judges. Instead I use Mercury Decide (inception/mercury-decide:free), a decision model. It doesn’t write text: it takes a state and typed questions, and returns a choice, a score or a yes/no with its probability, in a bit under a second. On OpenRouter it is called through its own API, System One (POST /api/v1/systemone), not chat completions.

This is the request the function makes:

{
  "model": "inception/mercury-decide:free",
  "state": {
    "visitor_question": "How does he work with teams?",
    "candidate_reply": "For Pelayo, leading means making things easy [10]…"
  },
  "questions": {
    "q": {
      "type": "noul",
      "instructions": "Is candidate_reply a final reply to the visitor, written in English, that could be shown as is?",
      "criteria": {
        "true": "A finished reply in English addressed to the visitor: it answers the question, citing sources as [n], or says plainly that the site does not cover it, or politely declines.",
        "false": "Not a reply: it thinks out loud about the question, the sources or the rules (\"The user is asking…\", \"Looking at source 10…\", \"I need to…\"), is a draft or plan, is cut off, or is not in English."
      }
    }
  }
}

noul is System One’s yes/no type, and the criteria describe each side with examples. The answer is a number:

{ "answers": { "q": { "type": "noul", "noul": 0.93 } } }

In Go it looks like this:

func (cfg Config) judge(ctx context.Context, req request, text string, try int) (string, error) {
	if cfg.Judge == nil {
		return "", nil
	}
	p, err := cfg.Judge(ctx, req.Question, text, req.Lang)
	if err != nil {
		log.Printf("ask: judge error (try %d), answer shown unchecked: %v", try, err)
		return fmt.Sprintf("judge%d=error", try), nil
	}
	if p < minAnswer { // 0.5
		log.Printf("ask: judge rejected p=%.2f (try %d)", p, try)
		return fmt.Sprintf("judge%d=%.2f rejected", try, p), fmt.Errorf("judge: not an answer (p=%.2f)", p)
	}
	log.Printf("ask: judge ok p=%.2f (try %d)", p, try)
	return fmt.Sprintf("judge%d=%.2f", try, p), nil
}

There are two design decisions:

  • It fails open: if Mercury doesn’t answer, the reply is shown anyway. The judge improves quality, but it can never take the assistant down.
  • A rejection is just another error: the retry loop already knows what to do with it. No separate path is needed.

The topic gate, optional

The same model can answer another question, this time before the writing model is called: «is this question about Pelayo or his site?». The state also carries the earlier questions, so a follow-up like «and that?» isn’t taken as off-topic. Below a probability of 0.15, the function answers «I can only answer questions about Pelayo…» directly, without spending a call on the large model. The threshold is low on purpose: I’d rather let an odd question through than turn away a good one.

It is off by default, and the ASK_GATE=on environment variable turns it on. The reason is quotas.

Quotas, in numbers

Limit Value
OpenRouter free models, per minute 20 requests
OpenRouter free models, per day 50, or 1,000 once you have bought at least $10 of credits
My limit per visitor 8 questions every 10 minutes and 40 a day
My limit per function instance 300 questions an hour
Question 500 characters; the previous 4 are sent as context

Mercury is a :free model, so I assume its calls count against the same quota (I’ll confirm it with the usage data). Without the gate, a question costs at least two requests (answer and judge), and up to six with retries. With the gate, one more. At 50 a day that covers few questions; at 1,000, plenty. Buying $10 of credits once, which aren’t spent if you only use free models, is what turns this into something usable.

So that nobody drains the quota, the function also has the same protections as the contact form:

  • an origin check;
  • a signed token that is requested before asking and is rejected if used too quickly;
  • in-memory per-IP limits.

When the quota runs out, visitors see «the assistant has used up its free answers for today» instead of a generic error.

How to tell whether the judge works

Every verdict is recorded in two ways:

  • In Vercel’s logs: lines like ask: judge ok p=0.94 (try 1) or ask: judge rejected p=0.12 (try 1).
  • In the X-Ask-Checks header: judge1=0.94, or judge1=0.12 rejected judge2=0.91 when a second attempt was needed. You can see it in the browser’s Network tab.

What is left to measure

I am writing this with the system freshly built. The interesting questions will be answered by a few days of production data:

  • How many answers Mercury rejects.
  • How many of those the heuristic would have let through.
  • How often a third attempt is needed.
  • Whether the 0.5 threshold is in the right place.

When I have them, I will add them here.

I think the underlying idea is more general than this assistant. A cheap, unreliable model, plus a fast decision model watching it, plus retries, makes a system that is much more reliable than any of its parts. And the judge doesn’t need to be clever: it only needs to tell an answer from something that isn’t one.