SecIT Bench is now part of the Fireworks Specialized Intelligence Index.
When the Cribl AI Research Lab launched SecIT Bench in August, our goal was simple: give IT and security practitioners a clearer picture of which models perform best in their domain. Today, we're taking that vision further.
We're partnering with Fireworks AI to bring SecIT Bench into the Specialized Intelligence Index as a benchmark in its Security / SRE category, giving practitioners another way to evaluate AI models against real-world domain expertise, not just general intelligence.
Built by practitioners, for practitioners
Most AI leaderboards sit at one of two extremes. Generalist rankings measure how models perform across a broad set of tasks, while individual benchmarks (SecIT Bench included) go deep on specific capabilities and workflows. Both are useful, but neither alone answers a bigger question: how well does a model perform across the range of work that matters in your domain?
The Specialized Intelligence Index takes a different approach. Its benchmarks are built by practitioners, for practitioners, and models are ranked across multiple benchmarks within a domain. That creates a broader view of specialized model performance without losing the depth that makes domain-specific testing valuable.
For IT and security teams, that matters. Choosing a model isn't just about finding the one that tops a single benchmark. Teams need to understand which models hold up across the work their SOC, SRE, and infrastructure teams actually perform, and how that performance compares with cost.
That's ultimately the question teams face when choosing models for SOC or SRE workflows: which model holds up across the full range of work we do, and at what cost?
More to come from Cribl
SecIT Bench is our first contribution, not our last. We're building new benchmarks for IT, security, and telemetry use cases and we plan to bring them to the index as they're ready, along with new research on how models perform across the domain.
As AI becomes a bigger part of IT and security operations, we believe practitioners need better ways to evaluate models based on the jobs they actually need them to do. We're excited to help build that standard alongside Fireworks and the broader community.
Stay tuned.
Explore the results
See SecIT Bench in the Specialized Intelligence Index, and visit the SecIT Bench leaderboard for our latest research.
Have a scenario, dataset, or idea for how AI agents should be tested in IT or security? Let's build the next benchmark together. Contact us and join our Slack community.








