
Elastic Principal Data Scientist Susan Chang used a recent QCon AI presentation to outline how the company moved from scattered, ad hoc testing for agentic AI products to a shared evaluation framework designed to work across teams, products, and workloads. Her talk focused on the practical challenges of tracing, grading, and regression-testing AI agents in production, especially when those agents span cybersecurity, enterprise chat, and retrieval-augmented generation workflows.
From siloed tests to a shared evaluation framework
Chang said Elastic’s early approach mirrored what many organizations do when they first ship AI agents: individual teams built their own datasets, metrics, and evaluators. That worked initially, but it also created duplicated effort, inconsistent storage, and different ways of measuring quality across products.
Over time, the company consolidated those efforts into a reusable framework. The goal was not to eliminate team ownership, but to make common building blocks available so developers could spin up production-grade evaluations more quickly and with less reinvention.
That shift mattered because Elastic’s agent portfolio covers very different use cases. Chang described internal security tools that analyze logs to help discover attacks, as well as enterprise chatbot systems that answer questions using proprietary data already stored in Elastic.
Different agents, different failure modes
The talk highlighted how evaluation criteria change depending on the agent. For a cybersecurity workflow, the team cares about whether the model identifies attacks correctly without hallucinating false incidents or unrelated MITRE tactics. For an enterprise chatbot, the focus shifts to factuality, response relevance, completeness, and whether answers stay grounded in documentation.
Elastic also has agents that generate ES|QL queries, which adds another layer of testing. Those systems need to produce valid syntax, and the team can use tool calls to check or format outputs when rules-based checks are more appropriate than model grading alone.
Chang’s broader point was that agent evaluations cannot be one-size-fits-all. The same underlying framework can support multiple products, but the datasets and success criteria still need domain-specific input.
Tracing as the foundation for evaluation
Before a team can evaluate an agent well, Chang argued, it needs detailed tracing. Elastic has used LangSmith and Phoenix heavily, alongside its own observability stack, to capture what an agent is doing step by step rather than only inspecting the final output.
That matters because a wrong answer does not always reveal where the failure happened. An agent might query the wrong database, call the wrong tool, or take a bad deterministic step long before the final response is generated. Without traces, teams can see the symptom but not the cause.
Chang said the company values trace-level detail because it enables both debugging and evaluation. In some cases, the team can assess whether the agent made the right tool call, not just whether it produced the right final response.
She also pointed to workflow features such as adding bad examples back into a dataset after user feedback. That allows teams to continuously expand their test coverage with real failures instead of relying only on synthetic scenarios.
What the shared system actually reuses
Elastic’s shared framework standardizes several elements across teams:
- Dataset import, despite different input formats across use cases
- Shared schemas for evaluation data
- Trace-based evaluators for latency, token usage, performance, and tool calls
- Shared RAG evaluators
- Custom evaluators for team-specific needs
Chang said the framework combines code-based checks, LLM-as-a-judge scoring, and trace-based metrics in a single run. Results are then stored back into Elastic or other tools depending on the workflow.
One notable implementation detail was the move from Python-based evaluation work to TypeScript. Chang said the production agents at Elastic were already written in TypeScript, while much of the initial evaluation work came from the data science side in Python. Rather than keep the evaluation stack separate from production, the team worked with engineers to translate more of the tooling into TypeScript.
That transition was supported by Playwright and a customized internal runner called Scout. Chang said Scout can load datasets, run the TypeScript-based agents, collect outputs, and execute evaluations both in shared environments and locally for developers.
LLM-as-a-judge and programmatic checks
Chang spent a large part of the talk on the trade-offs between LLM-as-a-judge and deterministic, rules-based evaluation. Elastic found value in model-based grading for open-ended tasks such as tone, coherence, style, and other ambiguous outputs where strict rules are hard to define.
But she also described the limits of relying on an LLM to grade another LLM. A judge model may not be granular enough for structured outputs such as JSON, YAML, or query syntax. It can also produce inconsistent results across runs, and it may miss internal details like product IDs or other factual constraints.
That is where rules-based evaluation helps. Programmatic checks can verify syntax, entity extraction, or whether a required value appears in the response. Chang said those checks are fast, cheap, and deterministic, making them useful for the parts of the problem where there is a clear right answer.
Her recommendation was not to choose one or the other, but to combine both. The shared framework at Elastic uses LLM-as-a-judge where ambiguity is expected and programmatic rules where precision is essential.
What could not be abstracted away
Chang was clear that no shared framework can replace domain expertise. Elastic could standardize the mechanics of evaluation, but it could not automate bespoke dataset creation. Security analysts and researchers still had to define realistic attack scenarios for the cybersecurity agents, and product teams still had to define positive and negative behaviors for their own workflows.
She also said calibration remains a human responsibility. If an evaluator does not align with human judgment or target behavior, the scores become noise. In practice, that means teams still need to tune metrics, review outcomes, and refine what “good” means for each agent.
Lessons for teams just starting with AI agents
Chang’s closing advice was practical rather than theoretical. Start small, she said, even if the first evaluation set contains only 20 to 50 records. Early-stage teams do not need to overbuild abstractions before they have real usage patterns to learn from.
She also urged teams to involve domain experts early, especially when the agent interacts with specialized systems or high-stakes data. In Elastic’s case, that meant working closely with security analysts and other internal experts rather than asking model outputs alone to define the right behavior.
Finally, she said consolidation should happen over time, not after teams have drifted too far apart. Elastic’s journey shows how quickly multiple tools and evaluation styles can spread across an organization once more than one agent reaches production. A shared framework can reduce friction later, but only if teams are willing to standardize what they can while preserving the flexibility to handle product-specific needs.
Source: Original report
Was this helpful?
Explore more: Software Testing Services More AI & Automation Tech News
Last Modified: October 6, 2026 at 10:33 pm
0 views

