Skip to main content
Blog_Building_Agents_Backwards_from_Evaluation.jpg
All Things AI

Building Agents Backwards from Evaluation

By Carter Church & Gabriel Bernadett-Shapiro

Executive Summary

  • Evaluating a security agent begins with defining the work it should perform and the evidence required to consider that work complete and accurate. Domain experts establish reference answers and grading criteria. 
  • A successful approach must combine evaluations of complete runs alongside tests of individual competencies to both demonstrate whether a system is useful and help engineers identify what exactly to improve.
  • Every material change to a model, prompt, tool, or workflow needs a measured comparison with the previous system. Operational failures supply new test cases, and fresh cases test whether improvements generalize.

In Part 1 of LLMs in the SOC, SentinelLABS examined the gap between cybersecurity benchmark scores and operational usefulness. In the months that have followed, we’ve observed those same design gaps in the flood of agentic systems hitting the market. This fail-fast development trend treats evaluation as a process that emerges after the proof-of-concept. In other words, teams build the system, demonstrate that it can work, and defer the harder task of establishing how reliably it performs and where it fails for a later date. 

This makes sense in the new world of rapid prototyping and AI experimentation. However, these systems are often brittle or only work until something in the environment changes or fails to generalize. This raises the urgent question, "How do I fix or improve this one behavior?" inside of a monolithic, non-deterministic pipeline.

For a system that will support security decisions, we believe the most important design choice is to make its performance measurable from the start. That means building the agent so that engineers can identify where an investigation goes wrong, test a correction, and verify that the change improves the complete workflow.

Consider a grounded example: an investigation agent that reviews an alert, queries the SIEM, and returns a verdict that matches an experienced analyst's. That looks like a success. On closer inspection, however, the query used the wrong device identifier and retrieved unrelated events. The verdict may be correct, but the records gathered during the investigation do not support the decision.

How would the team building that agent discover the mistake? And after fixing it, how would they know that the next model or harness update had not brought it back?

Start With the Work the Agent Should Complete

When building agents, attention often goes to the visible machinery: which model, which prompts, which tools, and so on. Those matter, but they're also some of the easier parts to modify. The part that's hardest to retrofit is the ability to answer two questions: is the system effective? And if not, what exactly do I change to make it better?

The industry has converged on this from several directions. Frontier lab engineering guidance observes that teams get surprisingly far on manual testing and intuition, but once an agent is in production, building without evaluations breaks down. Users report that the agent "feels worse" after a change, engineering scrambles to find out whether the claim has any validity, and it becomes difficult to tell real regressions from noise or to test changes against realistic scenarios before shipping [2]. For some the answer is what they label eval-driven development, scoped tests at every stage, before and during implementation, the way test-driven development writes tests before the code exists [3].

What that guidance usually leaves implicit is that evaluation discipline is an architectural decision. If you want to measure components individually with rigor, the agent has to be built out of individually exercisable competencies.

Before we choose a model or write a prompt, we need to define the work we expect the agent to do. For an investigation agent, that means reviewing an alert, gathering the evidence needed to understand what happened, and reaching a conclusion that the evidence supports. When it cannot resolve the case, it should give an analyst a clear account of what it found and what still needs investigation.

On Evaluations

"The first principle is that you must not fool yourself, and you are the easiest person to fool." – Richard Feynman

An evaluation is a way to measure a model's performance against a given scenario. More formally, every evaluation decomposes into a task (the point of indecision the system must reason through), a scenario (the environment wrapped around that task), and rubrics (how completion gets judged). This may seem simple on the surface, but LLM evaluations are an emerging discipline that every industry is struggling with. As the Pydantic team puts it, "Anyone who claims to know exactly how your evals should be defined can safely be ignored" [4]. A good eval set asks "what will a team use this for?" and ensures that the system gives them sufficient answers, and that it does so consistently.

End-to-End Evaluation: Does the System Perform?

An end-to-end evaluation runs the whole system on a realistic workload with grades as the final outcome. For an investigation agent, the natural frame is the human baseline it augments. A SOC analyst starts the day with a queue. Over a shift they work some number of alerts, dispositioning them into buckets (benign, suspicious, escalate, respond), and they end the day with a measurable amount of work left over. One metric that matches is burndown: of everything the system started with, how much still needs human attention? Alongside it sit the metrics that keep this burndown honest, such as disposition accuracy versus expert ground truth, time-to-verdict, stability, consistency, and cost per alert.

End-to-end metrics like these are often seen as the only metrics that matter. They tell a compelling story, they demonstrate that a system is worth building, and they're exactly the narrative that users want to hear. They also catch emergent failures such as compounding errors, context loss, and mis-sequenced steps that only show up when every component runs together against realistic inputs [2,6].

However, end-to-end metrics can aggregate away the information an engineer needs. When the overall performance is bad, or when it degrades after a change, it doesn’t provide any context about what went wrong. A regression in burndown could indicate one of many problems, meanwhile, all we can measure here is that suddenly we aren't performing. In a monolithic system, the only answer is to start reading traces and hunting for sources of failure.

Competency Evaluation: What Exactly Do We Fix?

To tackle that question we use the competency evaluation (elsewhere called component-level [6]). A competency evaluation isolates one skill and tests it directly, with its own specific inputs, definition of success, and a wide variety of different cases. These evaluations are helpful because they tell you whether a single part of a larger system is functional and capable across many scenarios. The inverse is also true, which is to say that it tells you which exact skills are underperforming.

End-to-End and Competency Evaluations: What Each Reveals


End-to-end evaluation

Competency evaluation

Question answered

Does the system deliver value?

Which capability is weak, and how?

Unit under test

The whole system on realistic workloads

One severable skill in isolation

Example metric

Alert burndown, disposition accuracy, escalation precision

Query correctness rate against a specific source's schema

Primary audience

Leadership and the teams relying on the system

The engineers improving the system

Detects

Emergent, compounding, integration failures

Localized skill deficits, regressions in one capability

Fails to provide

Any indication of what to change

Assurance that the assembled whole works

A mature program runs both continuously, and many teams add additional layers. For example, trajectory or trace evaluation, which grades the path the agent took rather than only the endpoints [2,6]. These evaluations are also critical, but this article specifically serves to spotlight the two defined above.

A Real-World Example

Take our Purple AI® Agentic Investigation agent as an example. When invoked, the system runs the investigation in the same way an analyst would. It pulls the surrounding evidence for an alert, enriches that evidence with additional data, forms and submits queries against the organizations telemetry, reasons over what comes back, and lands a verdict with the findings to support it. The tempting way to grade that answer is obvious, does the verdict match what an expert would have decided?

Very tempting, and also very dangerous. Suppose an agent agrees with the expert 95% of the time. That sounds strong, but the agent could just be pattern-matching on alert characteristics instead of actually retrieving evidence and building intermediate conclusions. A correct verdict reached for the wrong reasons is a failure, and only having a verdict-level score doesn’t tell us how we arrived at a conclusion. This is the same blind spot SentinelLABS identified in Part 1 - the unit of evaluation was more than often the answer to a question, as opposed to the path taken to reach it.

Decomposing the Problem

So, how do you evaluate competencies in this example? Start from the job you're trying to solve and decompose. On inspection, investigation hinges on many competencies. The agent has to form hypotheses worth testing given alert data and state from previous steps. It has to translate those hypotheses into correct and efficient queries against each data source, and it has to understand the schema of that source and what the fields it's filtering on actually mean. It has to interpret result sets, pick the next pivot, and know when the evidence is sufficient to stop. 

Each of these is a distinct competency that may be evaluated on its own. 

An (extremely abbreviated) tree:

Investigation 

├─ Hypothesis generation

├─ SIEM query construction (per-source eval suites)

├─ Investigation efficiency (tool-call budget)

├─ Adherence to organizational policy

├─ Confidence & evidence sufficiency judgment

└─ etc.

Now our evaluations can become concrete. "SIEM query construction" means many cases, each specifying an investigative intent, a target data source, and ground truth for what a correct query returns, graded automatically by running the query against representative data and comparing results, and backed by experts who wrote their own queries for the same asks.

For example, consider the device-ID failure we mentioned above. A competency case could give the agent an alert with a device ID and ask it to retrieve the related authentication events from a representative SIEM dataset. The case already defines which events a correct query should return.

If the agent queries the wrong identifier, it fails the case. The query may run successfully and return plausible data, but it does not return the expected events. This turns a failure that was previously hidden inside a successful investigation into something the team can measure directly.

Once the team fixes the problem, the case stays in the evaluation suite to ensure future changes to the system are consistently measured.

Even just a few dozen well-chosen evaluations per capability, drawn from real failures, can provide a strong initial signal. Each of these competencies gets its own suite since we know they are areas where models may struggle.

Where to Invest Evaluation Effort First

Not every node in the tree deserves equal attention. When setting up your own evaluations we suggest asking two questions:

Is this actually hard for models? Some competencies are commodities today. Summarization is the canonical example. Modern LLMs summarize nearly anything competently and the difference between a good and a great summary rarely moves outcomes. Contrast that with SIEM query construction, a real technical hurdle, with objective failure (the query is wrong, slow, or returns the wrong rows), high downstream impact (a bad query corrupts the whole investigation), and known model weakness on dialect and schema-specific syntax. An area that exhibits similar challenges is where your evaluation cases should go.

Can you define a defensible standard? Some competencies rely on direct checks against known facts, while others require expert judgement about whether an agent’s decision was justified by the evidence available. Evaluation effort pays off where success is decidable. Query construction has crisp ground truth (run it, compare results). Evidence-sufficiency judgment is softer, but experts mostly agree on clear cases, so a rubric plus expert-labeled examples works. Where even experts can't agree what "good" means, an eval suite just encodes noise. Before evaluating, you must first have a clear definition of success [10,11].

A useful heuristic is to simply combine the two and prioritize competencies by impact of failure × probability of model failure × decidability of ground truth. Wherever possible, look to grade binary pass/fail per case instead of arbitrary quality scores as binary judgments are easier for experts to make consistently, easier to trend, and harder to argue with [10]. If your system is already in the wild, you can still build cases - but we recommend you prioritize observed failures over successes [2,3].

The Value of Domain Expertise

Every competency evaluation needs to answer the question "what does good look like?" The person who can say whether a generated query succeeded (not just executed, but "all things considered, this is equal to or better than what I would have written in my day job") is the person who writes those queries well.

This is why many companies are spending heavily on contracted expert annotators. They're buying, at market rates and arm's length, the domain judgment that their evaluations require. In reality, most AI engineers are not domain experts, and many teams building agents don't have these resources in abundance. This is where being a company of cyber professionals is our advantage: for nearly any skill a security agent needs, someone here already performs it at an expert level daily. That expertise is the standard every one of our evaluations is built against.

When you identify a competency, identify its experts, understand how the work actually functions, and co-author cases with them. Evaluations built this way are true to real life, benefiting generalization when it reaches the field as it was tested against reality rather than the builder's imagination of what the field might look like.

Prove It Before Production

A well-built competency tree carries a risk because it makes the evaluation effort feel finished. However, evaluation is a constant game of discovering scenarios, hill climbing performance, and monitoring for regressions. Inside SentinelOne®, once an agent has been well-evaluated, it earns the chance to prove itself alongside real security work. 

At this stage, we actively execute solutions in parallel with working analysts in our internal SOC and Wayfinder Managed Services teams. When runs are flagged for review, an expert re-investigates the alert from scratch and establishes ground truth before reading the agent's verdict, so the reference standard can never be anchored by the machine it's meant to judge. 

Every interesting divergence becomes a new case, and every fix then has to beat the case that motivated it while not regressing against the larger suite. Only after our internal teams have thoroughly vetted an agentic offering in production does it reach anyone outside SentinelOne.

Underneath all of this effort sits one standard: in security response, unmeasurable capability is unshippable capability. An agent that cannot demonstrate exactly where it is weak cannot be trusted where it is strong. Autonomy in a SOC must be earned competency by competency, evaluation by evaluation, and finally in production, in front of the people who do the job.

References

[1] G. Bernadett-Shapiro and E. Garcia Lazo. LLMs in the SOC (Part 1) | Why Benchmarks Fail Security Operations Teams. SentinelLABS, January 2026.

[2] Anthropic. Demystifying Evals for AI Agents. Anthropic Engineering, January 2026.

[3] OpenAI. Evaluation Best Practices. OpenAI API Documentation.

[4] Pydantic. Pydantic Evals. Pydantic AI Documentation.

[5] OpenAI. Advancing the Price-Performance Frontier with GPT-5.6. OpenAI, July 2026.

[6] Confident AI. LLM Agent Evaluation Metrics: Tool Calling, Task Completion, Reasoning, and Trace-Based Evals. Confident AI, 2026.

[7] R. Aleithan et al. SWE-Bench+: Enhanced Coding Benchmark for LLMs. arXiv:2410.06992, 2024.

[8] Anthropic. Writing Effective Tools for AI Agents — Using AI Agents. Anthropic Engineering, 2025.

[9] Anthropic. Building Effective AI Agents. Anthropic Research, December 2024.

[10] H. Husain. Your AI Product Needs Evals. hamel.dev, 2024.

[11] A. Szymanski et al. Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks. Proceedings of the 30th International Conference on Intelligent User Interfaces (IUI), 2025.

Third Party Disclaimer:

All third-party product names, logos, and brands mentioned in this publication are the property of their respective owners and are for identification purposes only. Use of these names, logos, and brands does not imply affiliation, endorsement, sponsorship, or association with the third-party.

Related Articles

Decorative background gradient

Subscribe

Get the Latest From the SentinelOne Blog