Thoughts

Plumbing is easy. Filters are hard.

Anyone can pipe a model to a user now. Knowing the water is clean is the real work.

Varun Singla9 min read

The first AI feature I ever shipped took five days.

We wanted a model to read the invoices merchants upload when they dispute a chargeback. So we tried one. We fed a vision model ten or fifteen invoices, asked it to pull out the merchant, the amount, the date and ten or so other fields, and to tell us whether the invoice matched the disputed transaction. It got them right. Confident, well reasoned, fast. Everyone in the room relaxed.

So we built the feature around it. Prompt, upload flow, a gate that asked merchants to re-upload when the model said no. Five days from demo to production. I was proud of that number.

Then we opened the tap.

Real merchants, real uploads. Hand-drawn invoices. Proper invoices where the quantity times the rate did not add up to the total. Forged stamps and signatures. Bills in three or four languages. GST numbers that were wrong or made up. The model kept answering with the same confidence it had shown in the demo. And it kept saying yes. Roughly three in every four invoices it approved were wrong.

What saved us was a boring design choice. The model sorted the pile, but a human in operations still looked at every decision before money moved. So the cost was ops time, not cash. Without that layer, this story would have a very different ending.

Nothing had changed. Same model, same prompt. The only difference was who chose the inputs. In the demo, we did. In production, the world does.

That gap is where most AI products fail. Not in the demo. After it. This piece is about the one thing that closes it, which is evals, and the framework I now use to think about them. I learned it by getting it wrong first.

Pipes and filters

I explain AI products to my teams as plumbing.

The model is a giant tank. It holds everything on the internet, including the wrong parts, and it serves all of it at the same pressure. It does not know your customer.

Your app is the pipes. It takes the input, builds the prompt, calls the model, calls your tools, and shows the answer. Most of the work of building an AI product is laying pipe.

The filter is evals. It sits between the pipes and the user and catches confident nonsense before anyone drinks it.

An AI product drawn as plumbing: a tank (the model) feeds pipes (your app) through a filter (evals) to a clean glass (the user). A dashed bypass skips the filter and fills a brown glass.NO FILTER. STRAIGHT TO THE USER.

The model

Everything on the internet. The wrong parts too, at the same pressure.

Your app

The pipes. Input, prompt, model call, tools, output.

The filter

Evals. Catches confident nonsense before anyone drinks it.

The user

Clean water. An answer you can trust.

The dashed line over the top is the bypass. Most demos ship on it. The brown glass is what their users drink.

Every demo you have ever seen ran on the bypass. The dashed line over the top, straight from the tank to the glass. Every product you actually trust has the filter. Your spam folder has one. Maps has one. Spotify has one. You never see them, but someone watches their false positives and false negatives every week.

A chatbot answer usually has none. You are the filter. That is fine for a homework assistant. It is not fine for anything that touches money.

Five stages, and where projects die

Every AI product I have built walks the same five stages.

  1. 1ProblemFind the leak
  2. 2PromptOpen the tap
  3. 3PipelineLay the pipes
  4. 4ProofFit the filter
  5. 5ProductionSupply the city

Most projects die here. A prompt that works on five examples, and no filter.

Problem is finding the leak: a real cost, paid by real people, that a model could remove. Prompt is opening the tap: the first version that works on five examples. Pipeline is laying the pipes: the gate, the override, the logging, the tools around the model. Proof is fitting the filter: the test set and the threshold that tell you whether it works. Production is supplying the city: rollout, dashboards, rollback.

Most projects die between stage two and stage four. A prompt that works on five examples, and no way to know if it works on the sixth. The question that decides whether you have a product or a demo is simple. How do you know it works?

Two stories from my own work. The first one is embarrassing, which is why it is first.

Story one: the bouncer who waved the fakes through

You already know how this one ends. Here is the fuller version, because the mistakes are in the details.

When a merchant disputes a chargeback, they upload evidence. A receipt, a bill, sometimes a photo of the goods. Someone in operations opened every image by hand and decided whether it matched the disputed transaction. Merchants waited days to hear that their upload was useless.

The fix was obvious. A vision model reads the image, extracts merchant, amount, date and a dozen other fields, compares them to the dispute, and either accepts or asks for a re-upload with a reason. I wrote the prompt myself, tested it on those ten or fifteen invoices, and handed it to engineering as the spec.

We also picked the model for reasons that had nothing to do with benchmarks. It was the only capable model hosted in India at the time. Data sovereignty and a fast security sign-off mattered more than a few points on a leaderboard. Constraints choose your model. The filter decides whether it works.

The pipeline was good. A soft gate at upload time that said "this does not look right, re-upload or proceed anyway." Confidence shown on every field. An override one tap away. Low-value invoices first, high-value disputes only once the numbers held. Humans still decided. The AI sorted the pile.

Five days later we went live.

~3 in 4

invoices the model approved on day one that were actually wrong.

0

evals. No test set, no metric, no agreed threshold.

1

human queue behind the model. It caught every one of them.

~2%

bad invoices getting through after about six weeks of fixing one failure class at a time.

Think of a bouncer checking IDs at a door. Some IDs are fake. Now imagine a bouncer who waves through three out of four fakes with a smile. That was us. The only reason it stayed a bad week instead of a bad quarter is that a second person stood behind the bouncer and checked every ID again.

There were no evals because evals were not on anyone's radar. The model's own confidence score was the improvised gate. The human queue behind it caught what the model let through while we fixed the prompt, and I am grateful we kept it. But we built the whole pipe before anyone knew whether the decisions coming out of it were right.

The six weeks that followed were the real education. Read the failed cases by hand. Group them. Hand-drawn invoices. Arithmetic that did not add up. Forged stamps. Regional languages. Bad GST numbers. Fix the biggest class, re-run, repeat. The prompt changed many times. The model never changed once.

The filter I should have built first

If I were doing it again, four things would exist before launch. None of them is complicated.

A golden set. About thirty real invoices with agreed answers. Every prompt change runs against all thirty before it ships.

Two kinds of wrong, named. A real invoice rejected costs operations time and annoys a merchant. A fake invoice accepted costs money. Which one you tune for is a product decision, and it should be made on purpose.

False positive

The bouncer turns away a regular. Ops opens the case anyway. The merchant waits. Cheap, but it adds up.

False negative

The bouncer lets in a fake ID. The dispute is lost on bad evidence. Expensive, and nobody notices until later.

A threshold, agreed before launch. "Ship when bad invoices getting through are under two percent on the set." Written down, so nobody can move it later.

And the habit of iterating by failure class. Group the failures. Fix the biggest group. Re-run. Repeat. Your spam folder does exactly this. Now you do too.

After this product we built an internal prompt-evaluation tool so the team never had to iterate blind again. That tool is the ancestor of everything in the next story.

Story two: an assistant with a bank account

The second product is a voice-first assistant inside the merchant app. It is wired to more than sixty live tools and APIs across payments, settlements, disputes and devices. It understands the question, decides which system to call, calls it, and speaks the answer. It acts.

The problem it solves is the one every merchant has at nine at night: a question about their money that only a live system can answer. Where, why, and when. A chatbot that cannot look it up just apologises in a nicer voice. An assistant that cannot act is a very expensive FAQ page.

We built it in-house because the voice is a commodity. The tools and the merchant's live context are the product.

A voice assistant is fun until it can see your account and act on it. Then you need crash-test dummies before you let real passengers in. Sixty-plus tools means sixty-plus new ways to be wrong, and they fall into four classes.

  1. 01

    Wrong tool

    Asked about a refund, checks settlement status. Fluent, confident, wrong question.
  2. 02

    Right tool, wrong arguments

    Looks up yesterday when the merchant said today. Or a half-heard transaction id.
  3. 03

    No tool at all

    Guesses a number instead of fetching it. The most dangerous, and the most fluent.
  4. 04

    Personality drift

    Too casual with an angry merchant. Apologising for nothing. Character is a promise too.

The third one is why tool calls get checked exactly, not by a judge. A wrong settlement amount is a P0.

This time we built the filter before the pipe carried real traffic. Here is how.

The golden set came from real calls. We did not write test cases at a desk. We pulled verbatim lines from merchants' historical support calls. Vernacular. Multilingual, often within one sentence. Background noise. One sentence carrying two or three intents. Ten different ways of asking the same thing. From those recordings we built more than five hundred golden cases, each one a multi-turn conversation, not a single prompt.

Two kinds of accuracy, checked separately. For every turn in every case we recorded two things: which tool should be called with which arguments, and what the merchant should hear back. Tool accuracy is checked exactly. Response accuracy is checked against the expected answer.

A third model as the judge. The assistant runs the conversation. A separate model then reads the transcript and scores both halves, the tool calls and the spoken output, against what the golden case expected. The judge is never the same prompt as the assistant.

Not every case is equal. Each golden case carries a criticality. Anything that touches money is P0. General information is P1. A P0 miss counts for far more than a P1 miss, so the score reflects what a real failure would cost, not just how many cases passed.

A threshold, then a loop. We agreed a score the assistant had to reach before it could talk to merchants. Then we ran the loop: run the set, read the failures, change the prompt, run again. Same model throughout. Only the prompt moved.

500+

golden cases, built from verbatim lines in real merchant calls.

~0.6

score of the starting prompt on the set.

~1.0

score now live with merchants. Same model, better prompt.

Passing the golden set got the assistant to the door. Four habits keep it safe once it is live.

It ran in shadow mode first. For a while it only drafted answers that a human agent read before replying. It did not speak to a merchant until the drafts were good.

We watch every metric in pairs. Resolution rate alongside satisfaction. Speed alongside tool accuracy. On its own, a fast call looks like a great call. Often it just means the merchant gave up.

Every tool has its own dashboard. If one API starts failing or slowing down, we see which one, not just that "the assistant got worse."

Every prompt change goes to a small slice of merchants first, and rolling it back takes minutes. The filter does not stop at launch.

How do you know it works?

That is the whole question. Not "does the demo look good." Not "which model is best." How do you know it works, on the inputs you did not choose, next week, after the prompt changed?

If you can answer that with a set, a threshold and a weekly sample, you have a product. If you cannot, you have a very convincing demo.

Build the filter first. I did not, and it cost me six weeks and a lot of angry merchants. You do not have to.

Keep going

The chat is an AI trained on my own notes. It knows this essay and the work behind it.