Human Risk Management does not need more numbers for the sake of having them. It needs a measurement program that helps an organization understand risk, notice meaningful change and know when to act.
For a long time, measuring the human side of cybersecurity meant working with whatever data the awareness program happened to produce. Completion rates told us whether people had taken the training. Phishing click rates gave us a behavioral signal we could track over time. Assessment scores added another useful data point, and if the systems talked to each other without a heroic spreadsheet operation in the middle, so much the better.
That was not necessarily bad measurement. It was measurement constrained by the tools, data and budgets available at the time. The problem is that Human Risk Management is now asking much larger questions of the same numbers.
Organizations want to understand where human-related cyber risk exists, why it may be surfacing differently across populations, whether culture or working conditions are affecting behavior, which interventions deserve investment, whether those interventions worked, and how human risk connects with the wider security environment. A training completion percentage was never designed to answer all of that, and a phishing average does not become enterprise risk intelligence simply because it has acquired a more expensive dashboard.
The next step in Human Risk Management is therefore not simply collecting more data or calculating more sophisticated scores. It is learning how to measure: choosing measures for a defined purpose, understanding what they can and cannot tell you, combining evidence appropriately, setting meaningful thresholds and building an analytical foundation that can turn disparate observations into a useful account of risk.
The important distinction is between having HRM metrics and having a Human Risk Management measurement program.
That distinction is now reflected in broader cybersecurity guidance as well. NIST's 2024 SP 800-55 update deliberately separates identifying and selecting information-security measures from developing a measurement program, and expands its guidance around statistical analysis, data quality, uncertainty, validation, reporting, organizational context and the management of measurement over time.
There is no universal set of Human Risk Management metrics that every organization should put on the same dashboard. A useful measurement system normally needs several different kinds of evidence because different measures have different jobs.
| Type of measurement | What it helps you understand | Typical time horizon |
|---|---|---|
| Program metrics | Whether the HRM program is reaching people and operating as intended | Ongoing operational |
| Health metrics | The underlying condition of relevant human, cultural or organizational factors | Periodic / longitudinal |
| Flow metrics | What is happening through the environment now: activities, events and behavioral signals | Frequent / continuous |
| Change metrics | Whether a specific condition targeted by an intervention is actually moving | Defined around an intervention |
| Watch metrics and thresholds | When a pattern moves far enough to deserve attention, investigation or action | Ongoing monitoring |
| Risk views | What the available evidence means when interpreted together in organizational context | Decision-dependent |
The categories can overlap, and the exact measures inside them should differ by organization. The purpose of the model is not to create five more tabs in the dashboard. It is to stop expecting a metric designed for one job to perform another.
A completion rate can be an excellent program metric. It becomes problematic when somebody quietly promotes it to a health indicator, a change measure and evidence of reduced enterprise cyber risk without doing the analytical work in between.
Human Risk Management has a vocabulary problem around scores, metrics, signals and risk. They are often discussed as if the words are interchangeable, which makes a dashboard look more conclusive than the underlying evidence warrants.
Consider a simulated phishing exercise. An employee clicking a link is an event: something observable happened. The proportion of recipients who clicked is a metric derived from a defined population of those events. A platform might normalize that and other information into a score, which can be very useful for comparison or prioritization. A conclusion about risk, however, asks a wider question about what that evidence means in context, how confident we should be in the interpretation and why it matters to the organization.
| Concept | What it represents |
|---|---|
| Event | An observable occurrence |
| Measure / metric | A defined way of representing or comparing evidence |
| Score | A summarized, normalized or combined representation of one or more measures |
| Signal | Evidence that may indicate something worth investigating |
| Risk interpretation | A reasoned conclusion about why the evidence matters in context and what decision may follow |
The distinction is not academic hair-splitting. If the thing being measured changes, the population changes, the data quality is weak or the score combines unlike measures without a defensible rationale, the resulting number can still look beautifully precise.
NIST's current measurement guidance explicitly addresses the need to develop, test and validate measures, consider quantitative and qualitative evidence, account for uncertainty and understand data quality before using measures to support decisions. That is a useful discipline for HRM, where apparently simple workforce measures can become increasingly consequential once they are used to classify populations or influence risk decisions.
A human risk score can be valuable, particularly when it helps compress complexity for prioritization or executive communication. The important thing is being able to get back underneath it. If the number changes, the practitioner should be able to investigate what contributed to that movement, determine whether it is meaningful and understand what action the evidence reasonably supports.
That is risk measurement. The score is one of its outputs.
Completion, participation, assessment performance, simulation coverage, reporting rates and campaign engagement have not suddenly become embarrassing relics of the awareness era. They remain useful because a program needs to know whether it is actually reaching the intended population and whether its machinery is working.
NIST SP 800-50 Rev. 1 continues to treat measurement and evaluation as important parts of a cybersecurity and privacy learning program, while placing them in a wider lifecycle concerned with organizational goals, behavior change, culture and continual improvement.
The problem is one of scope. If 98% of employees completed a program, that is useful evidence about delivery. If completion falls to 62% in one business unit, it may tell you something about program operations or organizational conditions that warrants investigation. Neither result, on its own, establishes how much human cyber risk exists in the population or whether a particular security behavior changed.
This is one place where teams often get stuck as they move toward HRM. The metrics they have spent years reporting are familiar, understood by leadership and easy to extract, so they acquire a gravitational pull. New measures get added around them, but the fundamental reporting model remains the same. Before long, there is an HRM dashboard containing twenty-seven numbers while the team is still struggling to answer the original question: what is happening to our risk?
The answer is rarely to delete all the old metrics. It is to give them the right job.
Some Human Risk Management questions are less about what happened this week and more about the condition the workforce or organization is in.
Does a population have the capability required for the risks it faces? Are there persistent barriers to reporting? Are people confident enough to act when something is ambiguous? Are aspects of the security culture helping or undermining the behaviors the organization wants? Are meaningful differences emerging between business functions, regions or workforce populations?
These kinds of measures have a different rhythm from operational activity data. They may be established through a baseline and revisited periodically, with the organization looking for meaningful movement rather than constant fluctuation. Their job is to help practitioners understand the underlying health of relevant parts of the human-risk system and identify where deeper investigation or intervention may be warranted.
Culture is a particularly good example. The NCSC describes cyber security culture as the collective understanding of what is normal and valued in the workplace, and stresses that every organization's culture journey will be different rather than prescribing a universal set of measures or interventions. We explore that further in NCSC Cyber Security Culture Principles: What They Are and Why They Matter, including why culture needs to be understood as an organizational system rather than reduced to a sentiment score.
Health measures can be extremely powerful, but they usually need interpretation. A low result does not automatically tell you what caused it, and a high result is not permission to stop looking. The value comes from reading these measures alongside other evidence, changes in the organization and what people actually do.
Other evidence behaves much more like a flow. Reports arrive, simulations are completed, warnings are acted upon or ignored, policy-related events occur, security systems generate telemetry, and patterns begin accumulating through ordinary work.
Modern HRM has access to far more of this information than traditional awareness teams did. Enterprise APIs, identity systems, collaboration tools, SIEM, DLP and other security platforms can potentially contribute relevant signals, although connecting systems and governing the resulting data remains non-trivial.
Our existing guide on How to Measure Human Cyber Risk goes deeper into the progression from individual events toward signals, patterns and risk. The important point for the measurement program is that flow data is useful because it helps us observe what is happening; it should not automatically be interpreted as an enduring property of the person who generated the event.
That distinction becomes more important as HRM platforms ingest more enterprise telemetry. The temptation is to collect every available event because technically we can, then infer risk from volume. A mature measurement program works in the opposite direction: start with the risk question, determine which evidence is relevant to that question and establish how it should be interpreted.
More data can create better visibility. It can also create a much larger pile of things nobody quite knows what to do with.
This is where Human Risk Management gets particularly interesting.
Suppose a baseline suggests that one workforce population understands how to report suspicious activity but is significantly less willing to do so than its peers. Perhaps qualitative evidence and culture findings suggest that people are worried about wasting the security team's time or looking foolish if the concern turns out to be harmless.
The program could choose an intervention designed to change that condition. At that point, the useful measurement question is no longer simply What metrics do we normally report? It becomes What evidence would tell us whether this particular thing moved?
That may require watching several measures together over a defined period. Program metrics can confirm that the intervention reached the population. Health measures can show whether the underlying condition is changing. Flow evidence may show whether reporting behavior changes in practice. Qualitative evidence may help explain why the numbers moved—or why they did not.
This is a different discipline from comparing this month's dashboard with last month's and looking for green arrows.
Change measurement needs a hypothesis about what the intervention is intended to influence and a sensible way to test whether the expected movement occurred. It should also allow for the uncomfortable but useful possibility that nothing changed, or that something changed for reasons unrelated to the intervention.
That is one reason a good measurement program needs enough analytical depth to move beyond raw averages. As teams mature, they may need longitudinal analysis, comparisons between populations, pattern analysis, tests of relationships between factors and appropriate treatment of uncertainty rather than relying entirely on descriptive dashboards. NIST's revised SP 800-55 guidance specifically adds material on statistical analysis, comparison, validation, uncertainty and continuous improvement for this reason.
You do not need a statistics department to begin Human Risk Management. You do need to recognize when the question you are asking has become more sophisticated than the calculation on the screen.
Another distinction we find useful with clients is separating what the organization is watching from what it is actively trying to change.
Some measures exist because the organization wants to know if something begins moving outside an expected range. Perhaps a reporting pattern drops sharply, a particular type of security event starts clustering in one population, or a long-established health indicator deteriorates beyond a threshold that warrants investigation. Those are monitoring questions, and good thresholds help the program know when normal variation becomes interesting enough to look at more closely.
Other measures deserve much more concentrated attention for a period because the program is actively attempting to influence them. If a specific intervention is underway, the team may temporarily care far more about a small set of change measures than about twenty standing indicators elsewhere in the dashboard.
This has a practical implication that gets lost in many software products: not every metric deserves the same prominence, cadence or lifespan.
A measure that matters intensely during a six-month intervention may become much less useful once the question has been answered. A health measure may remain relevant for years but only need periodic assessment. A program metric may need weekly operational attention without ever belonging in a board report. A threshold may sit quietly in the background until something moves far enough to deserve human attention.
Measurement programs should be allowed to evolve. Otherwise dashboards become rather like attics: everything anybody once thought might be useful gets kept forever, and after a while nobody remembers why half of it is there.
This is where an off-the-shelf SaaS dashboard and Human Risk Management begin to part company.
A good dashboard is enormously useful. It can make information accessible, help practitioners explore patterns, keep important measures visible and communicate complex information far more effectively than a spreadsheet full of raw events. We build dashboards too.
The mistake is assuming that the presence of the dashboard means the measurement problem has already been solved.
NIST SP 800-55 Volume 2 is particularly helpful here because it treats cybersecurity measurement as an organizational program involving scope, roles and responsibilities, measure development, data management, communication and continual review—not simply the production of metrics.
A Human Risk Management measurement program has to make similar decisions. What risk questions matter to this organization? Which measures are appropriate for them? What populations are meaningful? How reliable are the underlying sources? How frequently should something be measured? Who interprets it? What happens when it crosses a threshold? When should a measure be reviewed, replaced or retired?
The software can support all of that beautifully. It cannot make those choices fit your organization merely by arriving with default widgets already switched on.
One of the reasons we are cautious about universal HRM scorecards is that human risk is deeply contextual.
The measures that deserve attention in a global manufacturer with large operational populations may look different from those in a professional-services firm where almost everybody lives in Microsoft 365. A healthcare organization has different workflows, pressures and consequences from a retailer, transport operator or financial institution. Regulation, threat exposure, security architecture, culture, workforce composition and business priorities all change what evidence is useful.
Even the same metric can have different meaning in two organizations. A low reporting rate might reflect weak detection capability, poor confidence, a cumbersome reporting mechanism, cultural reluctance to raise concerns, low exposure during the period—or some combination of them. The number does not arrive with its own diagnosis attached.
For that reason, we think every meaningful measure should have a clear reason for existing. At minimum, the team should know what question it answers, which population it applies to, where the evidence comes from, how often it matters, what its limitations are, what decision it can support and what would cause the organization to investigate or act.
That documentation need not become a 40-page methodology manual. Its purpose is simply to stop numbers slowly losing their meaning as they travel from analyst to dashboard to executive deck.
By the time organizations come to Cybermaniacs for deeper Human Risk Management work, lack of data is often not their central problem.
They already have years of learning records. They have phishing results. They may have survey data, security events, reporting statistics, workforce information and more dashboards than anyone can reasonably be expected to love. What they lack is a measurement architecture capable of telling them which evidence matters, how different pieces should be read together and what the resulting story means for risk and action.
That is a different problem from implementing another data connector.
Sometimes the difficulty is analytical: the team has outgrown averages and needs to understand meaningful differences, patterns or change over time. Sometimes it is conceptual: several metrics have been combined into a score without enough clarity about what the score is supposed to represent. Sometimes the data itself is fine but there is no operating model for turning a finding into an intervention. Very often, the organization needs all three pieces to mature together.
Our ASSURE Human Risk Baseline exists in this territory. It combines quantitative and qualitative evidence to help organizations establish a clearer current-state view, identify meaningful differences and prioritize what deserves attention, while our wider Human Risk Management work helps clients develop the program around that evidence. The important public principle is not the proprietary methodology underneath ASSURE; it is that measurement should produce a defensible risk story that somebody can use rather than another static collection of scores.
We see the difference most clearly when clients begin asking functional or organizational questions. The CISO rarely wants to know only that the enterprise average is 74. They want to understand why one part of the organization looks different from another, whether that difference is meaningful, what may be driving it, how confident the team is in that interpretation and what the business leader responsible for that population should do next.
That is where measurement becomes management.
AI is a useful reminder that Human Risk Management metrics cannot be frozen in time.
An organization beginning enterprise AI adoption certainly needs to know whether people understand policies, approved tools and basic security expectations. Before long, however, the more interesting questions concern how AI is actually changing work: whether people are verifying outputs appropriately, what they trust or delegate, how sensitive information is being handled, where shadow usage appears, whether roles and workflows are changing, and how human oversight functions when AI systems become more agentic.
A training completion rate can tell you whether someone received AI education. It cannot tell you whether the workforce is ready to exercise good judgment in AI-enabled work.
That is why our AI Enablement & Change Management (AIECM) and Agentic Readiness & Change (ARC) work sit naturally beside Human Risk Management. The measurement program has to evolve as the nature of the work, technology and risk evolve; otherwise organizations become very good at measuring the workforce they used to have.
Cybermaniacs does not believe every customer should have the same collection of Human Risk Management metrics, nor that the goal is to compress every available signal into one master number.
We start from the questions the organization needs to answer and the decisions it needs to make. That means distinguishing program performance from underlying health, separating observable activity from interpretation, deciding what needs continual monitoring and identifying the smaller set of things the organization is deliberately trying to change. The analytical approach can then become more sophisticated as the questions demand it, rather than adding complexity simply because the platform can.
There is an important boundary here too. We can explain publicly what good measurement needs to accomplish and the kinds of analytical disciplines it requires without publishing the constructs, scoring architecture, thresholds, validation methods or decision logic that make our own models work. Those are part of the specialist capability customers are buying.
What matters for the buyer is knowing that the expertise exists behind the dashboard—and that somebody can help read the story when the data becomes more interesting than a row of green arrows.
The move from security awareness measurement into Human Risk Management does not require abandoning completion, phishing or the other measures that got the industry this far. It requires becoming more precise about what those measures are for and surrounding them with the additional evidence needed to answer more consequential questions.
A mature measurement program will usually contain relatively mundane operational metrics alongside deeper health measures, fast-moving event flows, targeted measures of change and a smaller number of thresholds worth watching closely. Some will last for years; others should disappear once they have served their purpose. Their value comes from how well they help the organization recognize meaningful conditions, interpret them in context and decide what to do.
That is a more demanding job than filling a dashboard. It requires a sound data foundation, analytical capability, governance, organizational context and practitioners who understand the limits as well as the possibilities of the evidence in front of them. Once those pieces are in place, Human Risk Management starts to produce something much more useful than another score: an increasingly defensible understanding of what is happening, why it matters and whether the organization's efforts are actually changing it.
Human Risk Management should use measures that help an organization understand program performance, underlying human and organizational conditions, behavioral and security-event flows, targeted change and conditions that warrant investigation. The exact measures should be selected for the organization's workforce, threat environment, objectives and decisions rather than copied from a universal scorecard.
A metric represents defined evidence, such as a rate, result or measured condition. A score usually summarizes, normalizes or combines one or more measures. Scores can be useful for prioritization and communication, but they do not replace the need to understand the evidence underneath them or interpret what it means for risk.
Yes. Completion and phishing measures can provide useful information about program delivery and particular behavioral events. Problems arise when they are treated as complete measures of human cyber risk rather than evidence with a specific and limited purpose.
A Human Risk Management measurement program defines what the organization needs to understand, which evidence and measures can answer those questions, how data will be collected and interpreted, what thresholds or changes matter, who owns the resulting decisions and how measures will be reviewed as the program evolves.
No. Organizations differ in workforce composition, technology, threats, culture, business model and risk priorities. Common measures can provide useful comparison, but the overall measurement system should be fit for the organization and the decisions it needs to make.
The appropriate cadence depends on the purpose of the measure. Operational and flow measures may warrant frequent monitoring, while underlying health measures may be assessed periodically. Change measures should generally follow the lifecycle of the intervention they are intended to evaluate rather than being monitored forever.
Not necessarily. Individual-level scoring is only one possible use of human-risk evidence and may be inappropriate or unnecessary for many decisions. Organizations should choose the level of measurement—individual, team, function, population or enterprise—that is proportionate and useful for the risk question being addressed.