Making sensitive data detection faster
The Cribl Guard Privacy Model has a pretty focused job: find sensitive data in telemetry, and do it accurately.
But accuracy is only part of the problem. When you're scanning large volumes of logs and other machine data, the model also needs to be fast enough to keep up.
So we started asking a simple question: how much of the model could we remove without hurting detection quality? The answer turned out to be quite a lot. In our development benchmarks, we increased throughput by 2.8× and reduced median latency by 63%, while F1, one of the metrics we use to measure prediction quality, changed by just 0.06 percentage points.
Starting with cribl-privacy-2.1, the Cribl Privacy Model family uses the compressed architecture we'll walk through here.
Why make the model smaller?
The cribl-privacy-2.0 family started with open-source BERT models that we fine-tuned to recognize sensitive information in telemetry using a pattern called Named Entity Recognition or NER. Instead of generating text like a large language model, the model looks at an event and identifies specific spans that match the sensitive data types it has been trained to recognize.
The thing is, even relatively small general-purpose BERT models are built to understand a broad range of language patterns, but our privacy model doesn't need to do all of that. It has a much narrower job and a fixed set of entities to detect. That means some of the capacity inherited from the original model may not be doing much useful work anymore. If we could find and remove that unnecessary computation, we could make the model faster without sacrificing the detection quality we care about.
What we tried
We didn't start with one compression technique in mind. We tested a few different approaches and measured what each one did to both speed and quality.
Attention head pruning
Transformer models use multiple attention heads to understand relationships between different parts of an input. A simple way to think about them is that each head gives the model a slightly different way to look at the same event. That flexibility is useful, but after fine-tuning for a very specific task, not every head stays equally important. Some can be removed with little impact on the final prediction. The important part is figuring out which ones.
Feed-forward neuron pruning
Each transformer layer also contains a feed-forward network, or FFN. This makes up a large share of the model's parameters and computation. One useful way to think about FFN neurons is as learned pattern detectors. They react when the model sees something similar to what they learned during training. A general-purpose model may need a huge number of those patterns. A model focused on privacy detection in telemetry may not. That gave us another place to look for unnecessary capacity.
Layer pruning
We also looked at removing entire transformer layers. This is a straightforward way to make a model smaller, but it's also pretty blunt. A layer can contain both useful and unnecessary behavior, so removing the entire thing can hurt quality quickly. That made more targeted pruning techniques more interesting for our use case.
Knowledge distillation
Once you've made a model smaller, there's another question: can you help it recover some of the behavior it lost? that's where knowledge distillation comes in. The original model acts as the teacher, and the smaller model acts as the student. Instead of learning only from labeled training data, the student also learns from the predictions of the stronger teacher model. That extra signal helps the smaller model hold onto useful behavior even after we've reduced its size.
Mixture-of-experts research
We also explored ideas from MoEBERT, which converts dense feed-forward networks into smaller experts. We ultimately didn't use expert routing in the final architecture, but one idea from that work was especially useful: ranking neurons by how important they are to the task. That helped us make pruning decisions based on measured importance rather than just removing capacity arbitrarily.
So how many heads did we actually need?
One experiment gave us a particularly good look at how much redundancy was in the model. We reproduced the iterative attention-head pruning method from Michel, Levy, and Neubig's paper, Are Sixteen Heads Really Better than One?. Our Pro model started with eight transformer layers and four attention heads per layer, for 32 heads total.
We ranked the heads by how much they appeared to contribute to privacy detection, removed the least important one, reran the benchmark, and repeated the process. For a while, not much happened. The starting model had a relaxed span F1 score of 82.57%. After removing one head, it was 82.65%. Even after removing eight heads, F1 was still 81.95%, just 0.62 percentage points below the original model. Then we hit a wall.
With 12 heads removed, F1 was 80.22%. Removing one more dropped it to 74.91%.

That was the interesting part as there wasn't a steady decline where every head we removed made the model slightly worse. There was a fairly large stretch where we could remove capacity with little impact, followed by a sharp drop. In other words, there was real redundancy in the model, but only up to a point and that's exactly what we were trying to find.
This doesn't mean attention heads are generally unnecessary but their value depends on the model and the task. For the Cribl Privacy Model, though, the experiment showed that we could remove a meaningful amount of computation before prediction quality started to fall apart
The workflow we landed on
The final workflow combines the strongest parts of the methods we tested. It removes task-specific redundancy in stages, then uses the original trained model to guide recovery.

The one-epoch training step after attention pruning is important. It lets the reduced model adapt before its activations are used to score FFN neurons. Distillation comes after both structural changes. The teacher keeps the capacity of the fine-tuned model, while the student learns to reproduce its behavior with less computation.
Pro model example
The Pro model is a good example of the full process.mThe original model was based on boltuix/NeuroBERT and had eight transformer layers, 32 attention heads, and an FFN width of 1,024. The compressed version kept just 12 attention heads and reduced the FFN width to 256.
The result was what we were hoping for. F1 changed by just 0.06 percentage points, while throughput increased by 2.8× and median latency dropped by 63%. For telemetry workloads, that's a meaningful difference as privacy detection may need to run continuously across large volumes of logs and other machine data. The less work the model has to do for every event, the easier it is to apply that detection at scale without creating unnecessary processing overhead.
That's really the point of this work. We weren't trying to build the smallest model possible. We were trying to build a model that does exactly the work this task needs, and not much more.
Note: These results come from FP32 PyTorch models in our model-development environment. The measurements were made before production quantization and threshold tuning. They do not represent the performance of the pr
oduction models.
The final models
We adapted the same approach across three operating points. Each model uses task-specific attention-head and FFN pruning, but the underlying architecture, sequence length, teacher model, and degree of compression vary. The details differ, but the idea stays the same: keep the model capacity that helps us find sensitive data, and remove the parts that don't.

That same philosophy extends to how we think about AI across Cribl: evaluate models based on the jobs they actually need to do, not generic benchmark scores alone. SecITBench, Cribl’s benchmark for AI in real-world IT and security workflows, takes a similar approach by evaluating how models perform on the kinds of investigations practitioners actually face, including accuracy, speed, and cost.
Want to dig deeper into how Cribl evaluates AI for IT and security? Explore our latest AI research. And if these are the kinds of problems you like working on, we’re hiring. Come help us build, test, and evaluate AI for real-world IT and security work.









