Cybermaniacs Human Risk Management Guides

Human Risk Scores: What They Tell You — and What They Don’t

Written by Team CM | Aug 23, 2026, 1:29:23 PM

Everyone wants a Human Risk Score.

The appeal is obvious. Give every employee a score, roll those scores up by department, aggregate the departments into an enterprise number, and put 8.7/10 or 88% on the CISO dashboard. Now the board has something simple to understand and the Human Risk Management program has something simple to report.

The problem is that human risk is not simple.

A Human Risk Score can be useful. So can an average, an index, a rate, a benchmark, a weighted score or a Key Risk Indicator. We need quantitative measures in Human Risk Management, and we should get much better at using them.

But a score is a representation of selected evidence. It is not the risk itself.

Human cyber risk emerges from the interaction between what people know, what they believe, how confident they are, what they actually do, the culture around them, the role they perform, the access they have, the technology they use, the processes they work inside, the pressure they are under, and the threats and controls surrounding all of it.

Then those conditions change.

So if an entire workforce can be summarized as 9/10: LOW HUMAN RISK, we should probably ask what disappeared on the way to that number.

Human Risk Management needs scores, but it also needs metrics, measures, rates, diagnostics, surveys, dimensions, patterns, thresholds, confidence, context and interpretation. The real goal is not to find one magical number. It is to create enough reliable evidence to understand the risk, explain what is driving it, choose an appropriate intervention and demonstrate whether anything changed.

That is how we prove Human Risk Management is working.

Quick Answer: What Is a Human Risk Score?

A Human Risk Score is a numerical value or category used to summarize selected indicators of workforce-related cybersecurity risk.

Depending on the methodology, a score might incorporate security awareness results, competency assessments, phishing and reporting behavior, security events, policy-related events, role or access information, behavioral patterns, cultural indicators, or other workforce and security evidence.

Human Risk Scores can be useful for prioritization, segmentation, trend analysis and communication. Their value, however, depends entirely on what sits underneath them.

A useful score should let you answer several basic questions: What does this number represent? Why were these measures selected? How current and reliable is the evidence? How is context handled? How confident are we in the result? Why did the score change? And what decision is the score intended to support?

That is very different from assuming one number represents the totality of human risk.

For the wider measurement architecture, see What Should a Human Risk Management Platform Actually Measure?.

Scores, Metrics, Measures, Rates and Risk Indicators Are Not the Same Thing

One reason Human Risk Management measurement gets muddled is that we often use measurement terms as though they mean the same thing.

They don't.

Term What it means in Human Risk Management Example
Event Something that happened An employee reported a suspicious email
Measure A quantitative or qualitative observation Reporting time was 11 minutes
Metric A defined measure tracked consistently Median time to report
Rate A normalized frequency or proportion 72% reporting rate
Score A value assigned or calculated from one or more measures Competency score: 82
Index A composite representation of several measures Security culture index
KPI An indicator of program or operational performance Percentage of workforce completing required learning
KRI A Key Risk Indicator intended to signal changing risk or exposure Increasing repeated risky behavior in a critical population
Signal Evidence interpreted as potentially meaningful to risk Repeated verification failures in a high-exposure role
Pattern Repeated, clustered or connected evidence One function showing similar behaviors across several data sources
Threshold A defined point at which investigation or action may be warranted Escalation when several defined risk conditions occur together
Baseline The starting condition used for comparison Initial competency or culture assessment
Benchmark An internal or external reference point Comparison with historical performance or a peer population
Diagnostic A method for understanding contributing conditions Structured culture or behavioral assessment
Outcome What happened after action or over time Reporting behavior improved after intervention

All of these can contribute to Human Risk Management, but they answer different questions.

A metric tells you something measurable. A score summarizes selected evidence. A risk indicator tells you something may be changing. A model helps you understand how different pieces of evidence might relate to the risk you are trying to manage.

That last distinction matters enormously.

Human Risk Scoring Is Not the Same as Human Risk Modeling

Scoring asks a relatively narrow question: what value should we assign to this?

Modeling asks a much bigger one: what evidence matters, how might the evidence relate, which conditions influence it, and what can we reasonably infer about risk?

A score can be one output of a model. It should not be confused with the model itself.

Imagine two employees both receive a Human Risk Score of 75.

The first has excellent cybersecurity competency, poor phishing performance, high system privilege and excellent reporting behavior. The second has lower competency, no recent phishing failures, limited access, poor reporting confidence and a recurring pattern of security workarounds.

It is entirely possible to construct a weighted scoring formula that produces 75 for both.

Operationally, however, those are two completely different stories. The exposure is different. The evidence is different. The likely causes are different. The appropriate intervention is almost certainly different.

That is why Cybermaniacs focuses on relationships between competency, psychology, behavior, culture, organizational conditions and security evidence, rather than simply generating isolated numbers in separate silos.

Our Cyber Learning Experience (CLX), for example, creates repeated learning and human-factor evidence instead of treating completion as the final outcome. Within the wider Human Resilience System, those dimensions can be considered in relation to one another so that knowledge, psychology, behavior and culture do not become four unrelated dashboards.

The value is not simply having more data.

The value is understanding the system.

Why a Single Human Risk Score Can Be Misleading

Imagine an organization's board receives the following headline:

Enterprise Human Risk Score: 8.7/10 — LOW RISK

It sounds reassuring. The problem is that nobody yet knows what has been averaged together.

Perhaps training completion is 96%. General cybersecurity knowledge is strong. Phishing reporting has improved. Most employees occupy relatively low-exposure roles.

At the same time, reporting confidence in one division may be deteriorating, a privileged workforce group may be demonstrating recurring unsafe behavior, another business function may have serious security-culture problems, and AI-related data-handling issues may be rising quickly.

The enterprise average can still look excellent.

That is what averages do: they compress variation.

Sometimes that is useful. Sometimes it hides exactly what risk management needs to see.

Averages can mask clusters, outliers, contradictory evidence and highly exposed populations inside a much larger workforce. Most importantly, the number itself rarely tells you why the result looks the way it does.

Was the risk driven by capability? Behavior? Culture? Access? Operational pressure? Technology? A change in the threat environment?

Without that context, a total Human Risk Score can create the kind of false confidence risk management is supposed to prevent.

We Already Learned This Lesson With Phishing Click Rates

Security awareness has already lived through a version of this.

For years, one of the industry's favorite measures was the phishing click rate. Lower was good. Higher was bad. Programs reported it, vendors benchmarked it and executives could understand it without needing a ten-minute explanation.

The metric is still useful. What changed was our understanding of what it could tell us.

A click does not tell you whether the simulation was unusually difficult, whether an employee was distracted, whether the message perfectly matched a normal business process, whether the person clicked and immediately reported it, or whether the employee recognized the warning signs but acted differently under pressure.

A phishing result is evidence about a particular event under a particular set of conditions.

The mistake was allowing one convenient measure to become shorthand for an entire category of risk.

We should be very careful not to replace phishing click rate with Human Risk Score and declare the measurement problem solved.

Our guide How to Measure Human Cyber Risk: Beyond Phishing Clicks and Completion Rates explores why mature Human Risk Management needs a broader data foundation of events, measures, signals, patterns and outcomes.

Traditional Security Awareness Metrics Are Still Valuable

Moving into Human Risk Management does not mean abandoning everything security-awareness practitioners have counted for years.

Training completion, participation, assessment results, quiz scores, phishing click rates, reporting rates, time to report, repeat simulation behavior, campaign reach and communications engagement all have value.

They answer useful questions. Did people receive the learning? Did they complete it? Did they demonstrate the competency being assessed? Did they report the simulation? How quickly? Are we consistently reaching the workforce?

The issue is not the metric. It is what we claim the metric proves.

A completion rate measures completion. It is not a security-culture score.

A quiz result gives us evidence about demonstrated knowledge or competency. It is not automatically evidence of secure behavior.

A phishing click records an event. It is not a diagnosis of someone's total cyber risk.

This measurement gap is one of the reasons we wrote Why Measuring Human Risk Success Is So Hard—and How HRM Solves It. The challenge is not that traditional metrics are worthless. It is that human behavior is multifaceted, and individual activity metrics cannot carry more meaning than the evidence supports.

Human Risk Management should preserve the useful traditional measurement stack and then add enough evidence around it to interpret what those measures actually mean.

Counts Tell You What Happened. Risk Management Needs to Know What It Means.

This is where the shift from security awareness reporting to Human Risk Management becomes much more substantial than changing the name of the program.

Traditional awareness reporting often follows a fairly simple pattern:

activity → count → report

Risk management has a different job. It needs enough visibility and observability to understand the environment, recognize patterns, identify relevant exposure, interpret changing conditions, decide when something crosses a threshold, and choose an appropriate response.

The process begins to look more like:

visibility → evidence → analysis → insight → decision → intervention → outcome

That is much closer to the way mature cybersecurity functions already think about technical risk.

Human Risk Management has to do the same while accounting for something considerably messier: people working inside real organizations.

That means balancing risk against the speed, goals and realities of the business. The same behavior can mean something very different depending on whether somebody is working with customer data, moving money, maintaining industrial systems, writing production code or trying to meet an impossible deadline inside a badly designed process.

Our work on Competing Priorities: When Business Speed and Security Pull in Different Directions looks at why those organizational pressures themselves can become human-risk evidence.

What Makes a Human Risk Score Useful?

The best way to make a Human Risk Score useful is to give it a job.

Perhaps its purpose is to prioritize which population requires deeper investigation. It might show directional movement in a specific risk dimension, compare a population with its own baseline, help identify where additional evidence is needed, or contribute to a threshold for intervention.

Those are all reasonable uses.

“This number represents our total human risk” is a much harder claim to defend.

A useful score starts with purpose. Before anyone discusses weighting, algorithms or visualization, ask: What decision is this score supposed to help us make?

The evidence underneath it should also be understandable. You may not need access to a vendor's proprietary formulas, but you should know whether the score represents competency, phishing behavior, observed security activity, role exposure, culture, or some combination of those things.

The evidence should be relevant and current enough for the job. Twelve irrelevant inputs do not make a sophisticated score, and precision in the display does not create precision in the evidence.

Context also needs to matter. The same security event can carry very different significance depending on role, access, workflow, controls, threat exposure and organizational conditions.

Finally, you should know why the score changed. If a workforce risk score moves from 72 to 58, was that because behavior improved? Competency increased? Exposure changed? New evidence arrived? The scoring methodology was updated?

If nobody can explain the movement, the organization has a reporting number, but not necessarily risk intelligence.

Confidence Matters More Than Decimal Places

One of the strangest things we do with risk scores is display extraordinary numerical precision while ignoring uncertainty.

Imagine an employee has a Human Risk Score of 83.7.

That looks scientific.

Now imagine the underlying evidence consists of one phishing simulation, a course completed nine months ago, a generic job-role classification and no relevant observed security events.

Do we actually know enough to call this person an 83.7?

Probably not.

A mature Human Risk Management approach should distinguish between what we observe, what we infer, and how much confidence we have in that inference.

Some findings will be supported by repeated evidence from several sources. Others will be early signals based on relatively little data. They should not look equally certain simply because the software can put both on a 100-point scale.

False precision is still false confidence.

Explainability Matters in Human Risk Management

Explainability becomes particularly important when a score influences decisions about people or workforce populations.

If a platform identifies a group as higher risk, practitioners should be able to understand the evidence behind the finding. That does not require every model to be simple or every algorithm to be public. It does require enough transparency to investigate the result responsibly.

Why are we seeing this condition? Which evidence contributes to it? How strong is that evidence? What relevant context might alter the interpretation? Are there competing explanations? What should we investigate next?

Those questions shift Human Risk Management away from surveillance and toward analysis.

The objective is not to find the risky humans.

It is to understand the conditions that create or reduce human-related risk.

How to Map Human Risk in Your Organization Like a Threat Network explores why looking at patterns across people, behavior, friction, culture and organizational structure often reveals more than ranking employees one by one.

Correlation Is Useful. It Is Not Causality.

As Human Risk Management analytics improve, organizations are going to discover more correlations in their data.

That is a good thing. It is also somewhere we need to be disciplined.

Suppose employees with lower competency scores experience more security events. That relationship is interesting and may deserve further investigation.

It does not automatically mean lower competency caused the events.

Perhaps those employees occupy roles with much higher exposure. Maybe they interact with more external parties, process more transactions, have access to more systems, or work under unusual operational pressure. Perhaps the relationship disappears when another variable is taken into account.

Correlation helps us identify patterns worth investigating. Causality requires stronger evidence.

That is why sophisticated Human Risk Management needs context, repeated measurement, intervention tracking, comparison, validation and an honest representation of confidence.

The objective is not to avoid making useful inferences. It is to avoid claiming more certainty than the evidence can support.

Human Error Is Usually a System Story

“Human error” often sounds like an explanation.

Most of the time it should be the beginning of the investigation.

An employee sent sensitive information to the wrong place. Fine. What happened next in the analysis?

What system were they using? What workflow were they following? What information did they have? What did they believe was happening? Were they under unusual time pressure? Was the process confusing? Was there a technical control? Did it work? Was there a warning? Was the warning understandable? Is the same behavior occurring elsewhere?

Those questions change the conclusion from:

A human made an error.

to:

These conditions produced this outcome, and here are the levers we may be able to change.

That is much more useful.

It is also why human-risk measurement has to account for the business environment around a person. Someone who repeatedly works around a security process may need additional competency or support. Or they may be responding rationally to a process that makes their actual job almost impossible.

The numbers alone cannot tell us which explanation is correct.

Psychology, Culture and Behavior Need to Map Together

Some human-risk evidence can be observed directly. Other dimensions require different measurement methods.

We can observe whether somebody reported a simulated phishing message. We cannot derive their confidence in challenging a senior leader from that event alone.

We can count DLP events. We cannot infer psychological safety from the count.

We can record a security workaround. We may need diagnostic or survey evidence to understand the attitudes, pressures or norms associated with it.

That means a mature Human Risk Management measurement system needs different types of evidence.

Observed and objective evidence can include security events, actions, rates, assessments and intervention outcomes. Structured diagnostics and surveys can help organizations understand dimensions such as confidence, self-efficacy, attitudes, perceived risk, social norms and cultural conditions.

The interesting part is not assigning each dimension its own score.

It is understanding how they relate.

Cybermaniacs' Cyber Learning Experience captures learning and human-factor evidence over time. The wider Human Resilience System is built around the idea that competency, psychology, behavior and culture become more useful when they can be considered together rather than counted independently.

Our work on Measuring Cyber Security Culture: NCSC-Aligned Metrics That Actually Work makes the same point from a culture perspective: activity data, behavioral evidence, perception and organizational conditions each reveal a different part of the system.

We do not need a universal formula that declares:

Knowledge = 22%. Culture = 18%. Behavior = 34%.

There is no reason to assume those relationships are identical in every organization, every workforce population or every risk condition.

The better question is whether the evidence helps explain what is happening here.

Leading and Lagging Human Risk Indicators

The distinction between leading and lagging indicators is particularly useful in Human Risk Management.

A leading human risk indicator gives us evidence that a condition may be deteriorating or improving before a negative outcome materializes. Depending on the organization and risk being managed, that might include declining competency in a critical area, reduced reporting confidence, increasing security workarounds, repeated verification failures, deteriorating cultural conditions or emerging behavioral patterns in a highly exposed workforce population.

None of those indicators guarantees an incident.

That is not their job.

Their value is that they may tell us something is moving early enough to investigate it.

A lagging indicator, by contrast, describes something that has already happened: a confirmed security incident, data-loss event, policy violation, successful compromise or realized business impact.

Good risk management needs both. If you only use lagging indicators, the organization discovers risk after it has materialized. If you only use leading indicators, you may never establish whether those signals have a meaningful relationship with the outcomes you actually care about.

This is part of the frontier of Human Risk Management. We are only beginning to understand which combinations of human, organizational and technical evidence will become the strongest early indicators of particular forms of cyber risk.

That frontier deserves good research and careful thinking.

It does not need premature certainty.

How Do You Tie Events and Behaviors to Human Risk Scores?

Start by resisting the temptation to score every event.

Imagine an employee clicks a simulated phishing email. The event itself is evidence, but it is only the beginning.

Has the person done this repeatedly? How difficult was the simulation? Does the employee normally report suspicious messages? What competency evidence exists? What role do they perform? What is the potential exposure associated with that role? Are similar patterns emerging elsewhere in the workforce? Are there relevant security events or cultural conditions?

As that evidence accumulates, the isolated event may contribute to a signal. Repeated or related signals may become a pattern. Once context and exposure are added, the pattern may indicate a risk condition. Only then does it make sense to ask whether that condition should influence a score, trigger a threshold or drive an intervention.

A useful measurement progression is:

event → measure → metric → signal → pattern → risk condition → decision → intervention → outcome

The alternative —

event → subtract five points from employee

— is simpler.

It is not necessarily smarter.

How Do You Prove a Human Risk Management Program Is Working?

This is ultimately the question practitioners need measurement to answer.

Eventually somebody asks, “What are we getting for all of this?”

Historically, security-awareness teams have answered with the things they could count: courses completed, phishing click rates, campaign participation and training hours.

Those numbers demonstrate program activity. They do not necessarily demonstrate program effectiveness.

A stronger approach begins by defining what the program is actually trying to change.

Establish a baseline. What is true before the intervention? That might involve competency, behavior, culture, risk conditions, relevant security evidence or program maturity. Cybermaniacs ASSURE supports deeper strategic human-risk baselining, culture assessment and program maturity work where an organization needs a more defensible starting point.

Then define the risk condition. “Improve awareness” is difficult to measure because it is not very specific. What behavior, capability, condition or outcome are you actually trying to affect?

Choose the evidence that will help you observe it. Record the intervention. Measure again.

Then look across the evidence.

Perhaps competency improved but behavior did not. Perhaps reporting improved while a culture indicator deteriorated. Maybe the intervention worked strongly in one workforce group but had almost no effect in another.

That is useful information.

Proving the ROI of Human Risk Management looks at why program effectiveness has to be understood through behavior, business context, culture and outcomes—not simply whether activity occurred.

How Do I Show I Am Effective as a Human Risk or Security Awareness Practitioner?

This is another place where weak metrics can undersell very good work.

A practitioner's effectiveness should not be reduced to 98% training completion.

A mature practitioner should increasingly be able to show that they understand the organization's workforce-related risks, establish meaningful baselines, identify priority populations or conditions, choose interventions based on evidence, improve relevant capability and behavior, strengthen useful cultural conditions, target resources more effectively and explain movement over time.

There is another part of effective practice that often gets overlooked: recognizing when something did not work.

Good risk management does not mean proving every intervention was successful. It means learning what is actually happening and adjusting intelligently.

That is considerably more credible than finding a different metric that makes every quarter look green.

For a more executive-focused view of the same problem, How to Measure the ROI of Security Awareness and Human Risk Programs looks at connecting behavior, readiness and response to business-relevant outcomes.

How Do You Prove Human Risk Reduction?

Carefully.

Cybersecurity contains many outcomes that are difficult to prove causally, and human risk is no exception.

Suppose phishing reporting improves after six months of intervention. That is useful evidence. If relevant competency improves at the same time, behavioral patterns move in the expected direction and related security events decline, the story becomes stronger.

That still does not necessarily justify the statement:

Our training prevented 14 breaches.

Good Human Risk Management reporting should distinguish between what was observed, what appears associated, what we reasonably believe contributed to the change and what we can actually demonstrate.

That is not weaker reporting.

It is better risk reporting.

This is one reason Why Measuring Human Risk Success Is So Hard—and How HRM Solves It remains such an important measurement question: success in security is frequently represented by something bad not happening, which makes disciplined baselining, trend analysis and triangulation across evidence particularly important.

What Should You Report to the CISO?

The CISO does not need 147 Human Risk Management metrics.

They also deserve more than one mystery number.

The CISO view should explain the important risk conditions the program can currently see, where the exposure is concentrated, whether those conditions are improving or deteriorating, what the organization is doing about them and whether there is evidence that the intervention is working.

Leading and lagging indicators can support that story. So can selected scores, rates and benchmarks.

Confidence and context should be part of the conversation too. There is a difference between “we have strong repeated evidence of this condition” and “we have an early signal we are investigating.”

The CISO needs enough abstraction to understand the system without losing the reasons behind the result.

A score can sit inside that view.

It should not replace it.

What Human Risk Metrics Should Go to the Board?

Board reporting requires even more compression, but compression does not have to mean distortion.

Boards generally need to understand the material human-related cyber risks, how those risks connect to enterprise exposure, whether the position is changing, what management is doing and where a leadership decision or investment may be required.

Compare these two reports.

Human Risk Score: 88% — GREEN

versus:

Human-related cyber risk is stable overall, but verification behavior in two high-exposure populations is deteriorating. Competency remains strong while reporting confidence has declined. Targeted interventions are underway and will be measured against the established baseline.

The second statement may have dozens of measures underneath it.

That is fine.

The numbers support the risk story.

The risk story should not exist simply to justify the number.

How to Measure the ROI of Security Awareness and Human Risk Programs goes deeper into how program reporting can move from activity measures toward behavior, readiness, response and executive outcomes.

What Is a Good Human Risk Dashboard?

A good Human Risk Management dashboard helps somebody make a decision.

That sounds obvious, but it is a useful test.

The practitioner running the program may need operational information about learning, phishing, communications, interventions and outstanding work. An analyst may need to see component data, patterns, confidence and relationships between evidence. A CISO needs a view of material conditions, exposure, trend and response. The board needs a concise risk story.

Those are different jobs.

A useful HRM reporting architecture might therefore include several layers:

Reporting layer What it should help answer
Program Operations What are we delivering and managing?
Workforce Capability What do people know, believe and demonstrate?
Behavior & Culture What patterns and conditions are shaping security behavior?
Risk Conditions Where are meaningful signals or thresholds emerging?
Security Evidence What relevant evidence exists outside the HRM platform?
Interventions & Outcomes What did we do, and what changed afterward?
Executive Risk View What matters, where is it moving, and what decision is required?

The analyst may need the pipes.

The practitioner needs the levers.

The CISO needs the system.

The board needs the risk story.

Trying to serve every audience with one giant Human Risk Score is exactly how useful detail gets lost.

Benchmarks Are Useful — Until They Become the Goal

Organizations naturally want to know whether they are doing well.

Benchmarks can provide orientation, and there is nothing wrong with using them.

The trouble starts when “better than average” becomes synonymous with “acceptable risk.”

Your organization has its own workforce, access model, threat environment, technology, culture, controls, business processes and risk appetite.

Suppose your phishing click rate is substantially better than an industry benchmark. Great.

Now suppose the remaining people who consistently fail are twelve employees with authority to initiate very large financial transfers.

The benchmark becomes much less comforting.

Risk management eventually has to answer:

Is this good or bad for us?

That is why context-specific thresholds become more useful as a Human Risk Management program matures.

Risk Thresholds Can Be More Useful Than Vanity Scores

A mature HRM program needs to know when evidence has moved far enough to warrant attention.

That may be when a pattern becomes persistent, several weak signals begin appearing together, a highly exposed population starts deteriorating, behavioral evidence conflicts with strong competency results, or an intervention fails to create the movement the organization expected.

Thresholds turn measurement into action.

They can also be tuned to the organization's actual environment rather than assuming that a universal score has the same meaning everywhere.

This is what we mean when we say Human Risk Management should be fit for purpose and fit for use.

The metric matters because of the decision it supports.

Human Risk Scores Need Time

Risk is not static, which means a snapshot is never the whole story.

A Human Risk Score today may be built from a very different body of evidence than the same score six months ago. That makes trend, recency and history important.

A mature system should help distinguish a current condition from movement, a persistent pattern from a temporary anomaly, and real improvement from a scoring artifact.

This is another reason continual learning and repeated measurement are important. CLX is designed around continual development rather than a single annual training event, creating repeated opportunities to understand how competency and related human factors move over time.

One measurement is a snapshot.

Repeated evidence begins to create a picture.

AI Makes the Single-Score Problem Even More Obvious

AI workforce risk introduces another set of variables that traditional security-awareness metrics were never designed to capture.

Organizations increasingly need to understand AI competency, trust, reliance, verification, data handling, shadow AI, decision authority, escalation, human oversight and the ways workflows themselves are changing.

What is the universal weighting for all of those things?

There isn't one.

A verification behavior that matters enormously in an AI-supported legal or financial decision may be irrelevant to a low-impact use case. Reliance means something different when AI drafts an internal meeting agenda than when it influences production code or a customer decision.

Cybermaniacs AIECM approaches AI Workforce Risk & Enablement through this broader human and organizational lens.

Our article How Do You Measure Human Risk in AI-Driven Work? explores what begins to matter when risk appears through trust, reliance, verification, escalation and real-world working behavior rather than a traditional training event.

AI makes the argument for context almost impossible to ignore.

Cognitive Security Adds Another Dimension

AI also changes the cognitive environment in which security decisions are made.

People increasingly work with convincing synthetic content, conversational systems, automated recommendations and machine-generated answers delivered with enormous apparent confidence.

That puts questions of trust, verification, attention, cognitive load and automation bias much closer to the center of cybersecurity.

The Psychological Perimeter: Human Risk, AI, and Cyber Resilience explores this intersection between cognition, behavior, AI, culture and organizational risk.

A Human Risk Score may eventually incorporate evidence associated with some of these dimensions.

But again, the important work is understanding what the evidence means in context—not simply finding another way to convert cognition into a number.

Agentic AI Expands the Model Again

Agentic systems take the measurement challenge further.

As people increasingly delegate work to systems capable of taking actions, interacting with other systems and carrying out increasingly complex tasks, organizations may need to understand appropriate reliance, delegation, oversight, escalation, intervention, override quality, decision authority and changes in human capability over time.

Cybermaniacs ARC focuses on the human and organizational readiness required for this new form of work.

This is why the Human Risk Management measurement architecture we build now needs to be extensible.

We are only scratching the surface of what will eventually be observable and measurable across human, AI and agentic work.

The right response is not to rush toward one universal score.

It is to build stronger ways of understanding the system.

Questions to Ask a Human Risk Management Vendor About Scores

If a vendor tells you it can give your CISO or board a Human Risk Score, that is not necessarily a red flag.

But it should generate some questions.

What exactly does the score represent?

Is it employee susceptibility, learning performance, security behavior, exposure, enterprise human risk or something else?

What kinds of evidence contribute to the score?

You do not necessarily need the proprietary formula, but you should understand the categories.

Why were those measures selected?

There should be a rationale beyond “we had the data.”

Are all inputs treated as equally reliable?

A recent observed behavior and an old training result should not automatically carry equivalent evidentiary weight.

How does context affect the interpretation?

Role, workflow, access, organizational pressure and existing controls can all change what an event means.

Does workforce exposure affect prioritization?

Risk without exposure can produce misleading conclusions.

Can we see the component measures?

A composite number should simplify the evidence, not bury it.

Can you explain why the score changed?

If not, operational use becomes difficult.

Does the model represent confidence or uncertainty?

Limited evidence should not look identical to a well-supported finding.

Can thresholds be tuned to our organization?

Your risk appetite is not generic.

How do you distinguish correlation from causality?

This is a particularly useful question once a vendor begins talking about advanced analytics.

Can interventions be connected to outcomes?

Otherwise the measurement loop stops halfway through.

How does the model adapt as our workforce, technology and risks change?

Human Risk Management is not a static problem.

A sophisticated-looking score is not the same thing as a sophisticated risk model.

A Better Human Risk Measurement Stack

Instead of imagining Human Risk Management as one master score, it is more useful to think about a measurement stack.

Events tell us what happened.

Measures and metrics make parts of that activity consistently observable.

Rates and scores summarize selected evidence.

Diagnostics and dimensions help us understand conditions that cannot be captured through event data alone.

Signals and patterns tell us where evidence may be becoming meaningful.

Risk indicators help us see when a condition or exposure may be moving.

Context and exposure tell us why it matters here.

Thresholds help us decide when to act.

Interventions record what we did about the condition.

Outcomes tell us what changed.

Assurance asks whether the body of evidence is strong enough to believe the risk is being managed.

A score can exist inside this system.

It should never be mistaken for the system itself.

So, Are Human Risk Scores Useful?

Yes.

Use scores. Use rates, averages and benchmarks. Use quantitative metrics and qualitative evidence. Use structured diagnostics, surveys, dimensions and indexes. Use KPIs and KRIs. Use leading indicators and lagging indicators. Use thresholds.

Just know what each one tells you.

And know what it doesn't.

A Human Risk Score should simplify evidence for a defined purpose. It should not manufacture certainty, erase context, hide contradictory evidence or transform correlation into causality.

And it should never give leadership the impression that the wonderfully complicated interaction between humans, technology, culture, work and cyber risk has finally been solved because the dashboard says 9/10.

Human Risk Management is a system.

You need to understand the levers and the pipes, what flows between them, where pressure builds and which conditions actually matter in your organization. You need enough quantitative evidence to make objective decisions, and enough context to understand what those numbers mean.

That is harder than selling a score.

It is also far more useful.

Organizations that genuinely want to get Human Risk Management right should look for partners doing the difficult work of connecting cybersecurity, behavioral science, psychology, culture, workforce data and organizational context—not simply adding another scoring layer to a SaaS product and hoping the number explains itself.

That is the frontier Cybermaniacs is building toward.

Not one perfect number.

A better way to understand the system.

Frequently Asked Questions

What is a Human Risk Score?

A Human Risk Score is a numerical value or category used to summarize selected evidence about workforce-related cybersecurity risk. Depending on the methodology, the score may incorporate competency, phishing behavior, reporting, security events, workforce exposure or other factors.

The important questions are what the score represents, how good the underlying evidence is and what decision the score is intended to support. For a wider view of the evidence that can sit underneath scoring, see What Should a Human Risk Management Platform Actually Measure?.

How is a Human Risk Score calculated?

There is no universal Human Risk Score formula.

Some scores use simple averages or weighted measures. Others combine multiple categories of learning, behavioral, security, cultural or workforce evidence. The sophistication of the score depends less on the number of inputs than on whether the selected evidence is relevant, current, explainable and fit for purpose.

Our guide How to Measure Human Cyber Risk: Beyond Phishing Clicks and Completion Rates explains how events can progressively become metrics, signals, patterns and risk conditions rather than being scored automatically.

What is the difference between a human risk metric and a Human Risk Score?

A human risk metric is a defined measure tracked consistently, such as reporting rate, time to report or a competency measure. A Human Risk Score typically combines or translates one or more measures into a numerical scale or category.

Neither is inherently better. They perform different jobs.

Why Measuring Human Risk Success Is So Hard—and How HRM Solves It explores why individual awareness metrics can become misleading when they are separated from psychology, behavior, culture and organizational context.

What is the difference between a KPI and a KRI in Human Risk Management?

A Key Performance Indicator (KPI) tells you how well an activity or program is performing. A Key Risk Indicator (KRI) gives you evidence that a risk condition or exposure may be changing.

Training completion might be an important KPI. A persistent rise in risky behavior inside a highly exposed workforce population could potentially become a KRI.

For more on connecting program measures to business outcomes, see How to Measure the ROI of Security Awareness and Human Risk Programs.

What are leading indicators of human cyber risk?

Leading indicators are measures or signals that may identify a deteriorating or improving risk condition before a negative outcome occurs.

Depending on the organization, examples might include changes in competency, reporting confidence, verification behavior, security workarounds, cultural conditions or repeated patterns within a high-exposure population.

The exact leading indicators should be chosen around the risk you are trying to understand rather than copied from a generic list.

What are lagging indicators of human cyber risk?

Lagging indicators describe outcomes that have already occurred, such as incidents, data-loss events, policy violations, successful compromise or realized business impact.

Mature Human Risk Management needs to understand relationships between leading and lagging evidence. Proving the ROI of Human Risk Management explores why behavior, readiness, operational context and outcomes need to be considered together when trying to demonstrate risk reduction.

Are phishing click rates Human Risk Scores?

No.

A phishing click rate is a behavioral rate associated with a particular simulation or set of simulations. It can provide useful Human Risk Management evidence, but it does not represent total human cyber risk.

A click becomes much more useful when the organization can interpret it alongside reporting, competency, repeated behavior, workforce exposure and relevant context.

Can one score represent an organization's total human risk?

A composite score can summarize selected evidence and may be useful for trend analysis or executive communication.

It should not be assumed to represent the full complexity of workforce-related cyber risk.

Averages can hide highly exposed populations, contradictory evidence, variation between groups and different causes of risk. The underlying evidence therefore needs to remain accessible enough to explain the story behind the score.

How do you tie employee behavior to a Human Risk Score?

Start by treating observed behavior as evidence rather than automatically converting it into points.

Ask whether the behavior repeats, whether other evidence supports it, what conditions surrounded it and how much exposure is associated with it. Repeated and connected evidence may form signals and patterns that eventually contribute to a risk condition, threshold or score.

How to Map Human Risk in Your Organization Like a Threat Network explores why relationships between behavior, culture, friction and organizational context often tell us more than isolated employee events.

How do you prove a Human Risk Management program is reducing risk?

Begin with a defined risk condition and baseline. Decide which measures will help you observe it, record the intervention and then measure what happens afterward.

The evidence becomes stronger when several relevant sources move in a consistent direction—for example, competency, behavior, reporting, culture or related security evidence.

For a deeper treatment of this problem, see Proving the ROI of Human Risk Management and How to Measure the ROI of Security Awareness and Human Risk Programs.

How do you measure the effectiveness of a security awareness program?

Security awareness effectiveness can include operational measures such as completion and participation, but should increasingly extend into competency, phishing and reporting behavior, changes over time and relevant program outcomes.

The important distinction is between proving that the program did something and providing evidence that something changed because of the program.

Why Measuring Human Risk Success Is So Hard—and How HRM Solves It looks specifically at why completion and click rates are not enough to answer that question.

How do you measure security culture?

Security culture requires a wider measurement approach than training participation or phishing results.

Useful evidence can include perceptions, psychological safety, leadership behavior, social norms, reporting confidence, observed behavior and the organizational processes surrounding employees.

Our detailed guide Measuring Cyber Security Culture: NCSC-Aligned Metrics That Actually Work looks at culture measurement across perception, behavior and organizational structure.

Why does organizational context matter when measuring human risk?

The same behavior can mean something very different depending on the environment around it.

Role, workflow, access, business pressure, leadership expectations, process friction and organizational culture can all change both the cause and consequence of a behavior.

Competing Priorities: When Business Speed and Security Pull in Different Directions explores how operational pressure itself can become an important part of the risk picture.

What Human Risk Management metrics should be reported to the CISO?

CISO reporting should focus on material risk conditions, exposure, trends, important leading and lagging indicators, the interventions underway, resulting outcomes and the strength of the underlying evidence.

Detailed operational measures should support that view rather than overwhelm it.

A useful executive report explains what is changing, why it matters, what is being done about it and how confident we are in the interpretation.

What Human Risk Management metrics should be reported to the board?

Boards generally need a concise view of material human-related cyber risks, whether exposure is increasing or decreasing, what management is doing about those risks and whether there is evidence that interventions are effective.

A Human Risk Score can support this discussion, but it should not replace an explanation of the important risk conditions underneath it.

How to Measure the ROI of Security Awareness and Human Risk Programs goes deeper into translating Human Risk Management evidence into executive and board-level outcomes.

Why does explainability matter in Human Risk Scoring?

Explainability helps practitioners understand which evidence contributed to a score, why it changed, how reliable the underlying information is and what action may be appropriate.

This becomes particularly important when scoring is used to prioritize people or workforce populations. A result should support investigation and decision-making rather than simply label somebody “high risk.”

What is the difference between correlation and causality in Human Risk Management?

Correlation tells us that two things appear to move together. Causality means one condition actually contributed to another.

Human Risk Management analytics can surface valuable correlations, but context, repeated evidence, validation, intervention outcomes and other analysis may be required before stronger causal conclusions are justified.

This is one reason a sophisticated risk model should preserve uncertainty rather than make every statistical relationship look definitive.

Should Human Risk Scores include security culture and psychology?

Relevant psychological and cultural factors can provide valuable context for behavior, particularly when they are measured through well-designed assessments, diagnostics or surveys.

They should not simply be assigned arbitrary universal weights.

Their real value comes from understanding how they relate to competency, behavior and organizational conditions. Measuring Cyber Security Culture: NCSC-Aligned Metrics That Actually Work provides one example of how these less directly observable dimensions can be measured without pretending they are simple activity counts.

How should Human Risk Scores account for AI workforce risk?

AI introduces new human-risk conditions involving trust, reliance, verification, data handling, shadow AI, escalation, human oversight and changing decision authority.

Those factors should be interpreted according to the AI use case, workforce population, business impact and surrounding controls rather than dropped indiscriminately into a generic scoring formula.

How Do You Measure Human Risk in AI-Driven Work? explores these emerging AI workforce measurement questions in more depth.

What makes a good Human Risk Score?

A good Human Risk Score has a clear purpose. It uses relevant and sufficiently current evidence, can be explained, accounts for appropriate context, represents uncertainty honestly and helps somebody make a real risk-management decision.

The sophistication of a Human Risk Score is not determined by how many inputs it contains or how many decimal places it displays.

It is determined by whether it helps the organization understand and manage risk better.