Agentarium
An instrumented lab where AI agents test conversational AI
A testing agent runs multi-turn conversations against the assistants under test, and Agentarium records every experiment: transcripts, assessments, flags, and notes. Your team sees at a glance whether anything has changed.
From the team delivering production AI at
Assistants change quietly, and customers notice first
A prompt tweak, a knowledge-base update, or a platform release can change how an assistant answers overnight. Manual spot checks only catch the failures someone thought to look for; the answer an assistant invents at 2am goes unnoticed until a customer quotes it back. The teams that trust their assistants are the ones that test them the way they test software: continuously, against a baseline, with the evidence kept.
Your agent brings the science, Agentarium keeps the record
Agentarium doesn't decide what good means. It supplies the specimen chamber, the instruments, and the lab notebook, then guarantees every experiment is recorded and attributable. The science lives in the testing agent's prompts.
-
Define
An experiment carries the criteria and personas the testing agent evaluates against. They're authored by your team, not baked into the harness, so the testing method can evolve without a redeploy.
-
Converse
The agent drives real multi-turn conversations through MCP tools, against draft assistants by default. Every turn is recorded automatically: text, citations, latency, even platform errors.
-
Review
Assessments, flags, and notebook entries land in the dashboard. Your team browses transcripts, compares runs against a baseline, and triages whatever a human needs to see.
From "did anything break overnight?" to the exact turn that broke
The dashboard is built around one journey: notice the change, trace it to a run, and read the conversation that proves it. These screens are from a walkthrough dataset for a fictional company; select one to see it full size.
-
The morning check
The overview opens on the answer, not on settings: a week of runs, a live count of what needs review, and each experiment's recent results as a colour strip. A regression is visible before you click anything.
-
Regressions, spelled out
Compare any run against its baseline and the story is one line long: a persona went pass to fail. Every outcome links straight to the conversation behind it, so the next click is the evidence itself.
-
Evidence, not summaries
Full transcripts with the agent's scorecard as plain pass and fail rows. Turns, citations, latency, and platform errors are all kept, with raw payloads one click away for engineers.
Built so the record can be trusted
-
The harness judges nothing
Criteria and scoring rubrics live outside the system, in the testing agent's prompts. What counts as good stays yours to define.
-
Nothing goes unrecorded
An agent cannot hold a conversation without leaving a transcript behind. Even platform errors are kept as first-class turns.
-
Production is protected
Runs hit draft environments unless targeting live is a deliberate act. The assistant your customers talk to is never disturbed by accident.
-
Everything is attributable
Each run, flag, and note is stamped with the person who operated the agent, so the chain of evidence holds up when it matters.
Find anything, and take the data anywhere
Full-text search spans every recorded conversation, so you can find exactly where an assistant said the wrong thing. CSV and JSON exports, raw payloads included, hand the same data to spreadsheets and LLM-based analysis. And it all works on a phone, for the flag you check from a corridor.
Catch it, trace it, prove it, fix it, before a customer ever sees it
From "did anything break overnight?" to the exact turn where an assistant went wrong, Agentarium keeps the full chain of evidence in one place: recorded by agents, reviewed by people.
-
Talk to us
A 30-minute call about how your assistants are tested today. You'll leave with a clear next step.
-
Customer Assistants
How we design, build, and run the conversational assistants a harness like this keeps honest.
-
Customer stories
Conversational AI in production, including student support at scale for the University of Auckland.