Agent Release Evaluation Gate
Run a saved set of test cases against a new agent version, compare it with the live one, and hold the release until the answers hold up.
These images are illustrations of the concept, not screenshots of the actual product.
Overview
In AI and automation work, a small change to an agent can break things nobody meant to touch: a smaller model routes a billing question to the wrong team, a new instruction accepts a dispute the policy declines, or a tool gets called when a person should have taken over. This Botlit concept shows a release evaluation gate that runs before a new agent version publishes, whether it carries new instructions, a different model, a changed knowledge base or a new tool, so the change is tested against what matters before customers see it.
The illustration compares a v18 candidate of a sample triage agent with the live v17, beside a red Held badge counting three blocking failures, with Rerun evaluation and Override with note buttons. A strip records the change: a smaller model and new instructions, with the knowledge unchanged. A tag shows the change was submitted from a VibeControls plan. Four tiles set the candidate against live for cases passed, correct routing, answers with citations and average cost per case, with arrows on the tiles where the candidate moved: it falls behind on passes and routing while costing less per case.
A Failed cases table lists each case with the expected outcome, what the candidate did, what the live version did and a severity. A duplicate card charge goes to technical support instead of the billing agent, a late dispute is accepted when the expected answer declines it and cites the policy, a receipts export answer loses its citation and is marked minor, and a request to close an account calls the close-account tool when the expected outcome is a handoff with no tool call. With that last case selected, a side panel shows the expected and candidate exchanges one above the other, with a badge showing the case was added from an incident.
The design envisions each evaluation set as a golden set of questions with expected answers, routing, citations and actions that must not be taken, with incident cases added from the execution record in one step. The release is designed to stay held, with stated reasons, until a person accepts, fixes or overrides it with a note. A footer names the golden set, its case count and how many cases came from incidents, beside a note that results are sent to BigConsole via FluidGrids. In the design, engineers can submit changes from VibeControls, Botlit runs the gate, and a FluidGrids workflow is designed to pass the results to a BigConsole cost and quality console. The concept is meant for AI engineers, AI operations leads and the support owners who sign off agent releases.
What this concept shows
- A candidate agent version compared with the live version on the same golden test set
- A Held badge counting blocking failures, with Rerun evaluation and Override with note actions
- A change strip for model, instructions and knowledge, tagged with the VibeControls plan it came from
- Tiles comparing cases passed, correct routing, cited answers and average cost per case against live
- A failed cases table of expected, candidate and live outcomes with a severity for each
- Cases that check routing, policy decisions, citations and tool calls that must not happen
- A case detail panel with expected and candidate exchanges and a badge for cases added from incidents
- Results designed to reach a BigConsole console through FluidGrids
How it works
- An engineer submits a new agent version, for example from a VibeControls plan, with a changed model, instructions, knowledge base or tool.
- The gate runs the agent's golden set against the candidate and compares it with the live version.
- Tiles compare pass rate, routing, citations and cost per case, and the release is held if any blocking case fails.
- A reviewer opens each failed case to compare the expected, candidate and live outcomes.
- The reviewer fixes and reruns, accepts, or overrides with a note, and incident cases are designed to be added from the execution record in one step.
- Results are designed to pass through FluidGrids to a BigConsole cost and quality console.
Who it's for
- AI and machine learning engineers
- AI operations leads
- Support and product owners who approve agent releases
- Quality and risk reviewers
Illustrations
1 illustration of this concept. Select one to view it full size.
Candidate Versus Live Evaluation With a Held Release
This desktop illustration shows a release evaluation page in Botlit comparing a v18 candidate of a sample triage agent with the live v17, reached from Agents in a sample workspace. A red Held badge counts three blocking failures, beside Rerun evaluation and Override with note buttons. A strip records a smaller model, new instructions, unchanged knowledge and a VibeControls plan tag. Four tiles compare cases passed, routing correct, answers with citations and average cost per case with the live figures. A Failed cases table lists four sample cases with expected, candidate and live outcomes and a severity: a misrouted card charge, an accepted late dispute, a missing citation marked minor, and an account closure where the candidate called a tool instead of handing off. The selected case opens in a side panel with its expected and candidate exchanges and a badge showing it came from an incident. A footer names the golden set with its case and incident counts and notes results sent to BigConsole via FluidGrids.
Topics
- AI agent evaluation
- LLM regression testing
- golden test set
- agent release gate
- chatbot model change testing
- evaluation before deploy
- routing accuracy test
- tool call safety check
- model swap cost comparison
- agent version comparison
Related concepts

Execution Ledger and Run Costs
Every agent, channel and workflow run recorded with its status, duration, tokens, cost and the model that produced it.
1 illustration
AI Operations Findings and Bot Drafting
An operations copilot that flags regressions across your bots and drafts a new one from a plain-language description.
1 illustration
Intelligent Routing and Agent Delegation
A bot sends each message to the right specialist agent, and agents hand work to one another inside trust and depth limits.
2 illustrations
MCP Tool Servers and Live Tool Loop
A connected tool server lists every tool it exposes with a risk level, beside a step-by-step trace of one agent's tool calls.
1 illustration
Part of an industry solution
This concept appears in a cross-product solution on burdenoff.com — see how it works alongside other Burdenoff products to solve a problem in that industry.