Skip to content
Mayank Khanvilkar

Selected work

AI Quality

An evaluation framework that tells you whether an AI agent is right

The problem

Nobody could prove an AI agent’s answers were correct, and every model update was a gamble.

Context

AI agents were being trusted with customer queries and business tasks, but “it seems fine” was the only test.

What I built

  1. A written rubric for each task, defining exactly what a correct result looks like and which sources it must come from.
  2. A solved reference answer for every test case.
  3. A scored regression run against every new model release, so improvements and regressions show up immediately.
  4. Test cases calibrated across several models, adjusting ambiguity and the number of steps until failures pointed to real weaknesses rather than random noise, while every task stayed objectively gradable.

Outcome to confirm

Agents were iterated against these criteria until performance was reliable. Tier 1 support tickets fell by 60% and customer satisfaction rose by 15%.

Trade-off worth naming

Writing rubrics and reference answers takes real time up front. It’s the only way to know an agent is right, rather than hope.

Have a similar problem? Book a consultation call.

Book a consultation call