Cribl Privacy Model lost 20 heads and got 2.8x faster

Cribl Privacy Model lost 20 heads and got 2.8x faster

Last edited: September 27, 2026

Making sensitive data detection faster

The Cribl Guard Privacy Model has a pretty focused job: find sensitive data in telemetry, and do it accurately.

But accuracy is only part of the problem. When you're scanning large volumes of logs and other machine data, the model also needs to be fast enough to keep up.

So we started asking a simple question: how much of the model could we remove without hurting detection quality? The answer turned out to be quite a lot. In our development benchmarks, we increased throughput by 2.8× and reduced median latency by 63%, while F1, one of the metrics we use to measure prediction quality, changed by just 0.06 percentage points.

Starting with cribl-privacy-2.1, the Cribl Privacy Model family uses the compressed architecture we'll walk through here.

Why make the model smaller?

The cribl-privacy-2.0 family started with open-source BERT models that we fine-tuned to recognize sensitive information in telemetry using a pattern called Named Entity Recognition or NER. Instead of generating text like a large language model, the model looks at an event and identifies specific spans that match the sensitive data types it has been trained to recognize.

The thing is, even relatively small general-purpose BERT models are built to understand a broad range of language patterns, but our privacy model doesn't need to do all of that. It has a much narrower job and a fixed set of entities to detect. That means some of the capacity inherited from the original model may not be doing much useful work anymore. If we could find and remove that unnecessary computation, we could make the model faster without sacrificing the detection quality we care about.

What we tried

We didn't start with one compression technique in mind. We tested a few different approaches and measured what each one did to both speed and quality.

Attention head pruning

Transformer models use multiple attention heads to understand relationships between different parts of an input. A simple way to think about them is that each head gives the model a slightly different way to look at the same event. That flexibility is useful, but after fine-tuning for a very specific task, not every head stays equally important. Some can be removed with little impact on the final prediction. The important part is figuring out which ones.

Feed-forward neuron pruning

Each transformer layer also contains a feed-forward network, or FFN. This makes up a large share of the model's parameters and computation. One useful way to think about FFN neurons is as learned pattern detectors. They react when the model sees something similar to what they learned during training. A general-purpose model may need a huge number of those patterns. A model focused on privacy detection in telemetry may not. That gave us another place to look for unnecessary capacity.

Layer pruning

We also looked at removing entire transformer layers. This is a straightforward way to make a model smaller, but it's also pretty blunt. A layer can contain both useful and unnecessary behavior, so removing the entire thing can hurt quality quickly. That made more targeted pruning techniques more interesting for our use case.

Knowledge distillation

Once you've made a model smaller, there's another question: can you help it recover some of the behavior it lost? that's where knowledge distillation comes in. The original model acts as the teacher, and the smaller model acts as the student. Instead of learning only from labeled training data, the student also learns from the predictions of the stronger teacher model. That extra signal helps the smaller model hold onto useful behavior even after we've reduced its size.

Mixture-of-experts research

We also explored ideas from MoEBERT, which converts dense feed-forward networks into smaller experts. We ultimately didn't use expert routing in the final architecture, but one idea from that work was especially useful: ranking neurons by how important they are to the task. That helped us make pruning decisions based on measured importance rather than just removing capacity arbitrarily.

So how many heads did we actually need?

One experiment gave us a particularly good look at how much redundancy was in the model. We reproduced the iterative attention-head pruning method from Michel, Levy, and Neubig's paper, Are Sixteen Heads Really Better than One?. Our Pro model started with eight transformer layers and four attention heads per layer, for 32 heads total.

We ranked the heads by how much they appeared to contribute to privacy detection, removed the least important one, reran the benchmark, and repeated the process. For a while, not much happened. The starting model had a relaxed span F1 score of 82.57%. After removing one head, it was 82.65%. Even after removing eight heads, F1 was still 81.95%, just 0.62 percentage points below the original model. Then we hit a wall.

With 12 heads removed, F1 was 80.22%. Removing one more dropped it to 74.91%.

Cribl Privacy Model lost 20 heads and got 2.8x faster - img 1

That was the interesting part as there wasn't a steady decline where every head we removed made the model slightly worse. There was a fairly large stretch where we could remove capacity with little impact, followed by a sharp drop. In other words, there was real redundancy in the model, but only up to a point and that's exactly what we were trying to find.

This doesn't mean attention heads are generally unnecessary but their value depends on the model and the task. For the Cribl Privacy Model, though, the experiment showed that we could remove a meaningful amount of computation before prediction quality started to fall apart

The workflow we landed on

The final workflow combines the strongest parts of the methods we tested. It removes task-specific redundancy in stages, then uses the original trained model to guide recovery.

Cribl Privacy Model lost 20 heads and got 2.8x faster - img 2

The one-epoch training step after attention pruning is important. It lets the reduced model adapt before its activations are used to score FFN neurons. Distillation comes after both structural changes. The teacher keeps the capacity of the fine-tuned model, while the student learns to reproduce its behavior with less computation.

Pro model example

The Pro model is a good example of the full process.mThe original model was based on boltuix/NeuroBERT and had eight transformer layers, 32 attention heads, and an FFN width of 1,024. The compressed version kept just 12 attention heads and reduced the FFN width to 256.

The result was what we were hoping for. F1 changed by just 0.06 percentage points, while throughput increased by 2.8× and median latency dropped by 63%. For telemetry workloads, that's a meaningful difference as privacy detection may need to run continuously across large volumes of logs and other machine data. The less work the model has to do for every event, the easier it is to apply that detection at scale without creating unnecessary processing overhead.

That's really the point of this work. We weren't trying to build the smallest model possible. We were trying to build a model that does exactly the work this task needs, and not much more.

Note: These results come from FP32 PyTorch models in our model-development environment. The measurements were made before production quantization and threshold tuning. They do not represent the performance of the pr

oduction models.

The final models

We adapted the same approach across three operating points. Each model uses task-specific attention-head and FFN pruning, but the underlying architecture, sequence length, teacher model, and degree of compression vary. The details differ, but the idea stays the same: keep the model capacity that helps us find sensitive data, and remove the parts that don't.

Cribl Privacy Model lost 20 heads and got 2.8x faster - img 3

That same philosophy extends to how we think about AI across Cribl: evaluate models based on the jobs they actually need to do, not generic benchmark scores alone. SecITBench, Cribl’s benchmark for AI in real-world IT and security workflows, takes a similar approach by evaluating how models perform on the kinds of investigations practitioners actually face, including accuracy, speed, and cost.

Want to dig deeper into how Cribl evaluates AI for IT and security? Explore our latest AI research. And if these are the kinds of problems you like working on, we’re hiring. Come help us build, test, and evaluate AI for real-world IT and security work.

Cribl, the AI Platform for Telemetry, empowers enterprises to manage and analyze telemetry for both humans and agents with no lock-in, no data loss, no compromises. Trusted by organizations worldwide, including half of the Fortune 100, Cribl gives customers the choice, control, and flexibility to build what’s next.

We offer free training, certifications, and a free tier across our products. Our community Slack features Cribl engineers, partners, and customers who can answer your questions as you get started and continue to build and evolve. We also offer a variety of hands-on Sandboxes for those interested in how companies globally leverage our products for their data challenges.

More from the blog

GET STARTED

Ready to see what Cribl can do?

Whether you’re modernizing your stack, scaling security, or building AI‑powered operations, Cribl can help you take control of your telemetry.

See

Cribl

See demos by use case, by yourself or with one of our team.

Try

Cribl

Get hands-on with a Sandbox or guided Cloud Trial.

Join

Cribl

Help us build the AI Platform for Telemetry.