Operate and improve
See what every answer rested on
Find out how an agent behaves on the cases nobody thought to write a test for.
Read, per agent, whether its answers came from your knowledge, a live tool call, or the model's own general knowledge - and what the engine corrected on the way.
What this changes for your team.
Test drives and saved cases tell you how an agent handles the situations you imagined. The Health page tells you how it handled the rest. Every settled production turn is filed under exactly what it rested on, and the corrections the engine made mid-turn are counted beside it - so a knowledge base with a hole in it, or a release that made grounding worse, shows up as a number instead of a hunch.
How it works in practice.
- 01
Every settled turn records what its answer rested on and what the engine had to correct, as facts rather than transcripts.
- 02
Open Health on an agent to read the sources it drew on, the repairs it needed, and the population each rate is over.
- 03
Cut the same counts by the setup revision each turn ran on to see whether the last release moved them.
What you can plan around.
The behaviour you can design against, stated concretely.
One narrow row per settled turn is written inside the same transaction that settles the turn, so the record cannot disagree with the run it describes.
That row holds tool names, counts, tokens and a source category - never message content.
Repair rates are counted over grounded turns rather than over every turn, so runs that ended before a model answered never dilute the number.
The per-revision view counts each turn under the revision it actually executed, so a window spanning a release never averages a fix together with the bug it fixed.
Bring one real process
See how Yekar.AI fits the way you work.
Start with a job your team already owns, plus the tools and decisions around it.
Talk to us