Introducing SecIT-bench, Cribl’s AI telemetry benchmark - og image

Introducing SecIT-bench, Cribl’s AI telemetry benchmark: A new standard for evaluating AI models in IT and security workflows

Last edited: August 17, 2026

We gave 14 AI models the same 30 incidents. Here’s what we learned.


AI is quickly making its way into SOC and SRE workflows, promising faster investigations and automated root cause analysis. But many teams are running into a frustrating reality: spending heavily on AI inference and burning through tokens, only to end up at conclusions they can't fully trust.

With a rapidly evolving landscape of models, vendors, and versions, there's still no standard way to compare what actually matters: how they reason, perform in real investigations, and handle cost across messy, incomplete telemetry. 

That’s why we’re introducing SecIT-bench, a benchmark for evaluating AI agents on real-world IT and security workflows. We tested 14 frontier models across 30 realistic incident scenarios, measuring not just whether they got the answer right, but how they investigated, handled uncertainty, and what it cost. We hope SecIT-bench gives teams a more rigorous way to evaluate AI agents for telemetry workflows.

Why we did this 

We've seen this pattern before. Statistical AIOps, correlation engines, and copilots each arrived with ambitious promises, yet teams still rely on humans to investigate the hardest incidents. The fundamental question remains: Can an agent diagnose an incident it has never seen before?

“An agent found the root cause in our demo” isn't a measurement. Neither is a model's score on a coding benchmark. Debugging a software repository and diagnosing a production incident are fundamentally different tasks: one has a defined answer you can compile and test; the other requires navigating noisy telemetry, missing signals, conflicting evidence, and incomplete context to reason backward from symptoms to cause.

A capable investigator gathers evidence, follows leads, eliminates possibilities, and builds a case for a conclusion. They also need to know when the evidence isn't sufficient to make one. Evaluating AI for these workflows means measuring not just whether the final answer is correct, but how the agent investigates, how it handles uncertainty, how consistently it performs, and what it costs to get there.

How we tested the models

To make the benchmark representative of the conditions agents face in real-world investigations, we created 30 incident scenarios across four categories

  1. Intrusion and breach

  2. Covert data egress

  3. Service errors

  4. Performance degradation

Most scenarios are simulated using local and remote infrastructure configured to reproduce realistic conditions. We generated traffic and failures intentionally, then captured the resulting telemetry as evidence for the agents to investigate. Some scenarios were intentionally benign, with no incident to find, testing whether agents could recognize when the evidence did not support a root-cause conclusion.

We evaluated 14 AI models, including leading frontier models as well as newer and open-source models. All models were configured with maximum reasoning capabilities.

Two ways to evaluate an agent

We tested each model in two different “weight classes” to separate the capabilities of the underlying model from the tools and environment surrounding it.

Every agent received the same scenario evidence and one hour to investigate and document its findings, with no token or cost limit. We ran three independent rollouts for every scenario, model, and harness combination to account for the non-deterministic nature of agentic systems. 

How we scored the results

Evaluating an incident investigation isn't as simple as checking whether an answer contains the right phrase. A good RCA needs to identify the right components, connect the evidence, and explain the causal mechanism.

We therefore evaluated every RCA using a committee of three independent LLM judges. Each judge assessed the report against a defined set of binary diagnostic conditions, with the final score determined by majority vote.

We measured agent performance across three dimensions:

  • Accuracy: How many diagnostic conditions the agent correctly satisfied

  • Cost: Accuracy points achieved per dollar spent

  • Runtime: Time required to complete the investigation

This gives teams a more complete picture than accuracy alone. A model that achieves slightly higher accuracy at dramatically higher cost may not be the best choice for running investigations at scale.

What we found

Grok-4.6 takes the top spot, but competition is fierce. 

  1. Its 80.7% score barely unseats Claude Opus 5 in our tests.

  2. When looking at the granular scenario-level data, the top four performing models—spanning four distinct providers and featuring an open-source option—are virtually tied in a statistical dead heat for first place.

  3. Different models excel in different areas, highlighting key tradeoffs among accuracy, cost, and speed.

  4. Choosing a model depends on balancing these factors to fit specific workflow requirements rather than simply picking the highest score.

Cost doesn’t always buy precision. 

  • Our data reveals a narrow 17% spread in diagnostic accuracy set against a 20x delta in investigation spend.

  • Investigation expenditures varied wildly, ranging from a mere $0.28 to upwards of $7.34 per session. 

  • While the top-performing model averaged $3.27, we observed that premium pricing does not guarantee superior precision, as the costliest setups failed to secure the highest accuracy scores. 

  • This fundamental disconnect between spend and performance is the critical insight for any procurement strategy.

Introducing SecIT-bench, Cribl’s AI telemetry benchmark: A new standard for evaluating AI models in IT and security workflows - img 1
Introducing SecIT-bench, Cribl’s AI telemetry benchmark: A new standard for evaluating AI models in IT and security workflows - img 2
Introducing SecIT-bench, Cribl’s AI telemetry benchmark: A new standard for evaluating AI models in IT and security workflows - img 3

Security breaches were the easiest for agents to navigate, whereas diagnosing performance degradation proved to be the most elusive challenge for every model we tested.

Introducing SecIT-bench, Cribl’s AI telemetry benchmark: A new standard for evaluating AI models in IT and security workflows - img 4

The harness matters, too

One of the more surprising findings was how much the surrounding agent environment affected cost efficiency.

When we compared the same model running in a lightweight harness versus a coding harness, median cost efficiency was 23% worse with the coding harness.

These results give us strong conviction that IT and security workflows warrant purpose-built agent harnesses. Software development agents are optimized for a different set of tasks, tools, and success criteria. For IT and security investigations, purpose-built tools and environments can help agents work more efficiently, reduce unnecessary steps, and ultimately deliver greater productivity gains.

In other words, it's not just about choosing the right model. It's about choosing the right model, tools, and environment for the job.

Introducing SecIT-bench, Cribl’s AI telemetry benchmark: A new standard for evaluating AI models in IT and security workflows - img 5

Why should this matter to you

The benchmark data reveals something more interesting than a capability ceiling. Frontier models are genuinely strong reasoners and investigators. What they lack isn't intelligence — it's an environment built for getting the job done efficiently.

Consider how much scaffolding software engineering has accumulated. A coding agent gets a file tree, language servers, a test suite, structured diffs, and a CI loop that tells it whether it was right. Years of tooling investment went into making code legible and accessible to a model. Investigation work has almost none of that. That's an encouraging result, not a discouraging one. Architecture problems are tractable in a way that capability problems are not.

It does mean that evaluating agentic solutions requires looking past the demo at structural capability:

  • What's the accuracy of investigation on incidents the model has never seen? A measured number, across a scenario set you can inspect.

  • What does it cost per investigation? Accuracy without cost is half a number.

  • How consistent is it across harnesses? Ask for the same incident run on multiple harnesses and see where the variance lands.

  • How does it handle benign scenarios? An agent that invents problems when nothing is wrong introduces risk at scale.

Today's SOTA models are exceptionally good at compressing the first twenty minutes of an investigation — gathering evidence, triaging leads, ruling out the obvious. Closing the loop end-to-end requires purpose-built tooling: environments that put telemetry in one place at full fidelity, and action grammars designed for investigation.

That's the work we're excited about. We published these numbers to move the conversation from hype to measurement, and because we think the gap they expose is the interesting part. We have a few things brewing at Cribl on exactly this problem.

Practitioner notes:

  1. Evaluations were conducted using maximum reasoning configurations for all models.

  2. Defining accuracy: Rather than a binary pass/fail, our accuracy metric reflects the mean percentage of diagnostic conditions met across agent reports. A 79% score indicates the agent consistently satisfied the majority of investigative requirements rather than solving 79% of total cases. We believe this provides a more granular and transparent look at model performance.

  3. Measuring runtime: Execution durations for open-weights models were skewed by external rate-limiting constraints.

  4. Cost methodology: Spend estimates are calculated using standard, non-cached token pricing provided by each model lab.

Beyond the benchmark

This isn't about creating another leaderboard. It's about giving teams a more honest way to evaluate AI agents across accuracy, cost, speed, and trust.

The benchmark will evolve as the technology does. We'll keep the leaderboard up to date, expand into new scenarios, and use these insights to shape the models and agent harnesses we're building at Cribl.

And we don't want to do it alone. Have a scenario, dataset, or idea for how AI agents should be tested in IT or security? Let's build the next benchmark together. Contact us and join our Slack community.

Cribl, the AI Platform for Telemetry, empowers enterprises to manage and analyze telemetry for both humans and agents with no lock-in, no data loss, no compromises. Trusted by organizations worldwide, including half of the Fortune 100, Cribl gives customers the choice, control, and flexibility to build what’s next.

We offer free training, certifications, and a free tier across our products. Our community Slack features Cribl engineers, partners, and customers who can answer your questions as you get started and continue to build and evolve. We also offer a variety of hands-on Sandboxes for those interested in how companies globally leverage our products for their data challenges.

More from the blog

get started

Choose how to get started

Data growth and tiering/distribution, why the need now for a data engine - what got you to 2024 isn’t going to get you to 2034.

See Cribl

See demos by use case, by yourself or with one of our team.