BlogOperations

The Variance Problem: Why the Same Alert Gets Two Different Answers

Ask a SOC manager what their worst analyst is doing right now and watch the pause. The gap between your strongest and weakest analyst is the least measured risk in security operations, and it is not a training problem.

Triad Secure ResearchOperationsPublished Sep 21, 20267 min read

The Same Alert, Two Answers

Outcome quality tracks the analyst, not the alert.

Every SOC manager can name their strongest analyst and their weakest one. Very few can say what the difference costs them, because the difference almost never becomes visible. An alert is worked once, by one person, and whatever they concluded becomes the record.

That is the uncomfortable part. You do not get a second opinion on a closed alert. If the analyst who caught it stopped three systems short of the answer, nothing in the queue says so. The alert closes, the metric improves, and the gap that produced the miss stays exactly where it was.

Consider the same alert, an unusual authentication followed by a permission change, routed to three different analysts on the same team.

AnalystTime to CloseDepth ReachedConclusion
Tier 1, 8 months in role6 minAlert fields onlyClosed as false positive
Tier 2, 3 years22 minIdentity and endpointClosed as benign admin activity
Tier 3, 7 years48 minIdentity, endpoint, role chainEscalated, valid lateral movement

Illustrative of the pattern SOC managers describe when the same alert is re-reviewed. Not a measured benchmark.

Two of those three answers are wrong, and the environment does not care which analyst was on the queue. The alert was the same. The evidence available was the same. What differed was how far each analyst went before deciding they had enough.

This is not a failure of the junior analyst. It is a failure of a system that expects every analyst to assemble the environment themselves, from memory, under time pressure.

What Actually Varies

Not skill in the abstract. Four specific, observable things.

When teams talk about analyst variance they usually reach for experience as the explanation. Experience is the proxy, not the mechanism. What experience buys is a better private model of the environment, and that model is exactly what the system fails to supply.

Four things vary between two analysts working the same alert:

1

Which systems get checked at all

One analyst pivots from the alert to identity, then to the cloud role attached to that identity. Another reads the alert fields and closes. Nothing in the queue tells either of them which systems were relevant, so the scope of the investigation is set by habit.

2

Where the analyst decides to stop

Every investigation ends when the analyst feels they have enough. That threshold is personal. It moves with workload, time of day, and how many alerts are still waiting. It is the single largest driver of variance and it is never written down.

3

What counts as sufficient grounds to close

Two analysts can look at the same evidence and disagree on whether it clears the alert. Without a shared definition of sufficient, the close decision reflects the analyst's tolerance for ambiguity rather than the state of the environment.

4

Calibration on signal quality

Judgment about which alerts are probably noise is built through repetition and is uneven across a team. A newer analyst has not seen enough of the environment to know what normal looks like in it, so their calibration defaults to the rule that fired.

"The ticket said the alert was a false positive. It did not say why the analyst believed that. When the same pattern came back a month later, a different analyst reached the opposite conclusion, and neither of them was wrong given what they had in front of them."

Notice that none of the four is about intelligence or diligence. They are all about what the analyst had available at the moment they made the call. Change what is available and the variance moves, without anyone becoming a better analyst.

The Cost of Invisible Variance

You cannot manage a gap that never shows up in a report.

Most SOC metrics are averages. Mean time to triage, alerts closed per analyst, backlog depth. Averages hide variance by construction. A team where every analyst investigates to the same depth and a team where half of them stop early can produce identical dashboards.

The cost surfaces in two places, and both arrive late.

Which Version Does the Auditor Read?

The first is review. When an incident gets examined after the fact, someone reads the investigation record. What they find is one analyst's version of events, written under time pressure, with no indication of what was not checked.

If the analyst who worked it was your strongest, the record is defensible. If it was your newest, the record is thin in ways that are obvious in hindsight and invisible at the time. You do not get to choose which one the reviewer reads. The queue chose for you, months earlier.

The question is not whether your team is good. It is whether the record of a given alert would hold up regardless of who happened to be working it.

The second is the miss. A real intrusion that closed as a false positive does not announce itself. It reappears weeks later as a larger incident, and the investigation into that incident eventually finds the earlier alert that was closed too quickly. By then the cost is measured in dwell time.

Both failures share a cause. The floor of the team, not the ceiling, determines what gets caught. Hiring a stronger senior analyst raises the ceiling. It does nothing for the alerts that land on the newest person at 2am.

Why Training Does Not Close It

Training raises the mean. Variance is a floor problem.

The standard response to analyst variance is to train it away. Run more tabletop exercises, write better runbooks, pair junior analysts with senior ones, formalize the escalation criteria.

These are good practices and they help. They do not close the gap, for three structural reasons.

Training decays against a moving environment. What an analyst learns about your cloud estate is accurate until the estate changes, which it does continuously. The knowledge that made someone effective last quarter is partly stale this quarter, and nothing tells them which parts.

Attrition resets the floor. SOC turnover means there is always someone new on the queue. However well you train, the least experienced person on your team is working real alerts within weeks, and the floor is wherever they are. You are not training a team, you are training a population that keeps changing.

And runbooks codify the known. They tell an analyst what to do for scenarios someone anticipated. The alerts that matter most are the ones nobody wrote a runbook for, and those are exactly the alerts where individual judgment, and therefore variance, dominates.

The Limit of Better People

There is a version of this problem you can hire your way out of, and it is expensive. Staff every seat with seven-year analysts and the floor rises to meet the ceiling.

Almost no one can do that, and the teams that could would rather spend the headcount elsewhere. The realistic move is not better analysts. It is giving the analysts you have the same starting point the best one would have built for themselves.

What a Floor Looks Like

Set the starting point before the investigation begins, not after.

If variance comes from each analyst assembling the environment privately, the fix is to stop asking them to. The environment is knowable. Assets, identities, permissions, exposure, and what has already been concluded about them are all facts the platform can hold.

Holding them in a graph rather than in documentation matters, because the useful questions are relational. What can this identity reach. What else touched this host. Has this pattern been cleared before, and on what grounds. Those are traversals, and they either resolve the same way every time or they are not a floor at all.

Four properties make the difference:

1

The environment is resolved before the analyst opens the alert

Assets, identities, permissions, exposure, and prior findings related to the alert are assembled from a context graph of the tenant environment, not queried by hand. The starting picture is the same whoever opens it.

2

Resolution is deterministic, not generated

The same alert against the same environment produces the same context every time. The relationships come from the graph, so they can be inspected and traced back to their source rather than taken on trust.

3

Scope is set by the environment, not by habit

If an identity in the alert holds a role that reaches production data, that reachability is present in the context whether or not the analyst thought to look for it. The floor is set before the investigation starts.

4

The decision record shows the reasoning, not just the outcome

What was surfaced, what was considered, and what was decided are captured as part of the work. A reviewer can see how a conclusion was reached without asking the analyst to reconstruct it later.

Determinism is the load-bearing property. A system that generates a plausible summary of the environment has moved the variance rather than removed it, because now the starting point changes run to run and no one can tell you why. Grounding the context in a graph of the actual environment means the analyst can follow any claim back to the record it came from.

This works without any autonomy at all. The floor is set by resolution, not by automation, and it holds whether or not a team turns on a single automated action. Where teams do want workflows to act, those run on the same grounded context and stay opt-in, under the analyst's authority.

Raise the Floor, Not the Headcount

The goal is not to replace judgment. Analysts should still decide. The point is that they should all be deciding from the same picture of the environment, so the decision reflects the alert rather than the person.

When that holds, the answer to what your worst analyst is doing right now stops being a source of anxiety. They are working from the same ground your best analyst is, and the distance between them is judgment rather than information.

Security does not fail at detection. It fails at understanding.

See how Triad Secure restructures security operations around clarity, not noise.