7 October 20262 min read

Building a RAG assistant on Amazon Bedrock — and measuring whether it works

A finance-policy assistant that cites its sources, refuses rather than guesses, and is scored against a fixed test set before anyone trusts it.

By Srini Vankeepuram · Architecture, Engineering & Data Leader

Everyone can demo a chatbot. The harder question, and the one a CFO or a risk team will ask, is: how do you know it gives the right answer, and what does it do when it doesn't know? So I built a small retrieval-augmented generation (RAG) assistant to answer exactly that, and measured it.

It answers questions about a fictional company's finance policies: approval limits, expenses, travel, supplier payments, month-end close. Every policy and person in it is made up. The code and results are public on GitHub.

How it works

  • Chunk: twelve policies are split by section, so every answer can point to the exact clause.
  • Embed and retrieve: each section becomes a vector with Amazon Titan Text Embeddings V2; a question finds the three closest sections by meaning.
  • Decide: if even the best match is weak, the assistant refuses with "I can't find that in our finance policies" instead of guessing.
  • Generate: Claude on Amazon Bedrock answers only from those sections, citing each one, at temperature 0.
  • Guard: questions about named people's personal data, or predictions such as share prices, are declined before they reach the model.

Measure before you trust it

I wrote a fixed test set first: 20 questions with a known source and a known fact in the answer, and 3 that should be refused. The same set runs in a no-cost practice mode (word matching, no AI model) and on Bedrock, so every change to prompts, models or chunking is judged on evidence, not impressions.

MeasurePractice modeAmazon Bedrock
Right policy section retrieved20/2020/20
Answer contains the expected fact18/2019/20
Out-of-scope questions refused3/33/3
Results, 7 October 2026 (Bedrock: Titan Text Embeddings V2 and Claude Haiku 4.5, eu-west-1). Sources: Code, test set and full run output on GitHub.

What the misses taught me

Practice mode missed two questions that need reasoning, not word matching: it couldn't work out that £12,000 is over a £10,000 limit, or that a taxi is an expense. Claude answered both correctly. That is precisely the value an LLM adds on top of retrieval.

Bedrock's one miss wasn't a wrong answer. Asked whether an order can be split to avoid CFO approval, Claude said "No, you cannot split an order" and explained the 30-day rule correctly. But my test looked for the policy's exact words, "must not be split". Exact-phrase checks are cheap and strict, and they undercount good answers. The next step is a meaning-based check, an LLM as judge, with a person reviewing a sample of its verdicts.

What I'd tell a team starting out

  • Write the test set before the prompt. Otherwise you tune until the demo looks good.
  • Make refusing a feature. A confident wrong answer about an approval limit is worse than no answer.
  • Cite everything, so the finance team can check the answer against the policy in seconds.
  • Start cheap. An in-memory index and a small model cost pennies; scale the infrastructure once the numbers justify it.
  • Plan for the boring parts: a new AWS account's request quotas throttled my first run, so the code now retries with back-off and paces itself.
The question isn't whether the AI can answer. It's whether you can show when it's right, and what it does when it doesn't know.

Contact

Let's talk.

I'm based in London, UK. Whether it's a leadership role, an advisory engagement or a transformation that needs shaping — my inbox is open.

srini.vankee@gmail.com