Skip to content
ElementX

Agentarium

An instrumented lab where AI agents test conversational AI

A testing agent runs multi-turn conversations against the assistants under test, and Agentarium records every experiment: transcripts, assessments, flags, and notes. Your team sees at a glance whether anything has changed.

From the team delivering production AI at

  • University of Auckland
  • Mott MacDonald
  • Southern Cross
  • New Zealand Defence Force
  • The Warehouse Group
  • Deutsche Telekom

Assistants change quietly, and customers notice first

A prompt tweak, a knowledge-base update, or a platform release can change how an assistant answers overnight. Manual spot checks only catch the failures someone thought to look for; the answer an assistant invents at 2am goes unnoticed until a customer quotes it back. The teams that trust their assistants are the ones that test them the way they test software: continuously, against a baseline, with the evidence kept.

Your agent brings the science, Agentarium keeps the record

Agentarium doesn't decide what good means. It supplies the specimen chamber, the instruments, and the lab notebook, then guarantees every experiment is recorded and attributable. The science lives in the testing agent's prompts.

  1. Define

    An experiment carries the criteria and personas the testing agent evaluates against. They're authored by your team, not baked into the harness, so the testing method can evolve without a redeploy.

  2. Converse

    The agent drives real multi-turn conversations through MCP tools, against draft assistants by default. Every turn is recorded automatically: text, citations, latency, even platform errors.

  3. Review

    Assessments, flags, and notebook entries land in the dashboard. Your team browses transcripts, compares runs against a baseline, and triages whatever a human needs to see.

From "did anything break overnight?" to the exact turn that broke

The dashboard is built around one journey: notice the change, trace it to a run, and read the conversation that proves it. These screens are from a walkthrough dataset for a fictional company; select one to see it full size.

The overview: a week of runs and suite health, at a glance

Agentarium overview dashboard showing a week of runs, a review count, and suite health strips for each experiment

Baseline comparison: the persona that went pass to fail

Baseline comparison for a customer support experiment, showing one persona changing from pass to fail between runs

The transcript behind the failure, scorecard and flags included

A failing conversation transcript with the testing agent's assessment and flags shown beneath the turns

The review queue: every unresolved flag in one place

The review queue listing every unresolved flag with its experiment, conversation, and a resolve action
  • The morning check

    The overview opens on the answer, not on settings: a week of runs, a live count of what needs review, and each experiment's recent results as a colour strip. A regression is visible before you click anything.

  • Regressions, spelled out

    Compare any run against its baseline and the story is one line long: a persona went pass to fail. Every outcome links straight to the conversation behind it, so the next click is the evidence itself.

  • Evidence, not summaries

    Full transcripts with the agent's scorecard as plain pass and fail rows. Turns, citations, latency, and platform errors are all kept, with raw payloads one click away for engineers.

Built so the record can be trusted

  • The harness judges nothing

    Criteria and scoring rubrics live outside the system, in the testing agent's prompts. What counts as good stays yours to define.

  • Nothing goes unrecorded

    An agent cannot hold a conversation without leaving a transcript behind. Even platform errors are kept as first-class turns.

  • Production is protected

    Runs hit draft environments unless targeting live is a deliberate act. The assistant your customers talk to is never disturbed by accident.

  • Everything is attributable

    Each run, flag, and note is stamped with the person who operated the agent, so the chain of evidence holds up when it matters.

Find anything, and take the data anywhere

Full-text search spans every recorded conversation, so you can find exactly where an assistant said the wrong thing. CSV and JSON exports, raw payloads included, hand the same data to spreadsheets and LLM-based analysis. And it all works on a phone, for the flag you check from a corridor.