We have developed an oddly comforting habit in AI governance. Whenever we encounter a problem we cannot quite solve technically, legally, or organizationally, we put a human in the loop. The model might hallucinate, so a human should check it. The agent can take consequential action, so a human should approve it. The reasoning is opaque, so a human should interpret it. Accountability is murky, so a human should own the decision. There is residual risk somewhere in the process? Excellent. Put a human there too.
At some point, human-in-the-loop stops describing a thoughtfully designed control and starts describing the place we put everything we have not yet figured out.
There is a famous I Love Lucy scene in which Lucy and Ethel are working at a chocolate factory and the conveyor belt keeps getting faster. At first, the job is perfectly reasonable: inspect the chocolates, wrap them, keep the line moving. Then throughput starts exceeding human capacity and the entire control system degenerates spectacularly. Chocolates go into mouths, hats, uniforms — anywhere that keeps them from visibly piling up while the conveyor continues cheerfully on its way.

A great deal of enterprise AI oversight is in danger of becoming the digital version of that conveyor belt.
The human is technically still there. The approval may even be logged. But if we keep increasing the speed, volume, complexity, and opacity of the work while handing every unresolved exception to the same finite human attention system, we should not be terribly surprised when “meaningful oversight” eventually becomes a very expensive way of recording who clicked Approve.
This is one of the reasons we believe Human Risk Management has to evolve toward Human Resilience Engineering: deliberately designing the human and organizational capacity required to keep complex digital and AI-enabled systems working safely when the world does not behave exactly as the process diagram expected. And, ultimately, toward Human Resilience Systems that connect that capacity to the wider technology, governance, culture, risk, and learning environment rather than leaving resilience to individual heroics.
HITL is useful. The problem is what we are asking it to carry.
There is nothing inherently wrong with human-in-the-loop design. Some AI systems absolutely should require human judgment before consequential actions are taken, and serious governance frameworks are considerably more nuanced about this than the casual shorthand sometimes suggests.
The EU AI Act, for example, does not simply say “insert person here.” Its requirements for human oversight of high-risk AI include enabling the people responsible for oversight to understand relevant system capabilities and limitations, recognize the possibility of over-reliance, interpret outputs, disregard or override them where appropriate, and intervene in or stop the system. Deployers are also expected to assign oversight to people with the necessary competence, training, authority, and support. That is a much richer conception of human oversight than an approval button at the end of a workflow. EU AI Act — Regulation (EU) 2024/1689
NIST reaches a similar conclusion from a risk-management perspective. Its AI Risk Management Framework calls for organizations to define and differentiate roles in human-AI configurations, establish proficiency around AI system performance and trustworthiness, and actually assess the processes used for human oversight. Its human-AI interaction guidance goes further still: simply assigning a human expert to an AI-supported decision does not mean that person is equipped to perform an oversight or governance function, particularly where they played no role in the system’s development and may have limited visibility into its behavior. NIST AI Risk Management Framework NIST guidance on AI risk management and human-AI interaction
So the problem is not that governance has chosen humans as part of the control environment. Humans are extraordinarily good at things machines still struggle with: contextual judgment, ambiguity, novel situations, ethical considerations, competing priorities, tacit knowledge, and noticing that something is strange even when it is difficult to articulate exactly why.
The problem appears when we start treating that human adaptability as an unmetered utility.
Human attention is not infinitely scalable infrastructure
This becomes especially obvious when we move from occasional AI-assisted decisions toward agentic work.
A person reviewing one unusual recommendation may be able to inspect the evidence, think carefully about the context, challenge the output, and decide whether intervention is warranted. Give the same person several hundred routine agent actions to supervise and the nature of the job changes. They are no longer performing considered oversight of exceptional decisions; they are monitoring a high-throughput automated system and waiting for the rare moment when something important looks different.
Human-factors research has been warning us about this class of problem for decades. Automation bias — excessive reliance on automated decision support — is not simply a character flaw in inattentive operators. Research has repeatedly connected it to the design of the task itself, including cognitive load and the complexity involved in independently verifying an automated recommendation. One systematic review found that higher verification complexity was particularly associated with automation bias, while more recent HITL research has continued to flag the problem that highly accurate automation can eventually make humans worse at spotting the unusual occasions when it fails. Systematic review of automation bias and verification complexity 2026 systematic review of Human-in-the-Loop AI
That should make us wary of designing agentic systems whose safety case effectively reads:
The AI will perform the work at machine speed, and the human will catch anything weird.
There are several assumptions hiding inside that sentence. The human has to know what “weird” looks like. They need enough information to distinguish unusual-but-correct from subtly dangerous. They need enough time to investigate. They need relevant domain expertise, and some understanding of the AI’s limitations, and a realistic intervention mechanism. After 199 perfectly good outputs, they also need to bring the same quality of attention to number 200.
The technology may scale almost frictionlessly. Human vigilance does not.
This is precisely the kind of problem we are beginning to examine with clients as they move from general AI adoption into more consequential human-agent workflows. One of the most revealing questions we can ask is not, “Is there human approval?” but “What, exactly, are you asking the human to do here?” Once you unpack the word approval, you often find several different jobs hiding underneath it: verify factual accuracy, notice a security problem, interpret uncertainty, evaluate business judgment, detect a policy exception, assume accountability, and know when something needs to be escalated.
Calling all of that a “human-in-the-loop control” does not make it one control. It makes it a bundle of assumptions about human performance.
This is where Human Resilience Engineering changes the question
We are borrowing here from a discipline with a much longer history than enterprise AI. Resilience engineering emerged from safety science and the study of complex socio-technical systems, where researchers became increasingly uncomfortable with the idea that safety could be understood simply by counting failures and eliminating human variability.
Complex systems work, much of the time, precisely because people adapt. They make small adjustments around incomplete information, conflicting goals, resource limitations, unexpected events, and differences between the work imagined in a process document and the work actually required to keep the organization functioning. Resilience engineering studies and strengthens that adaptive capacity rather than treating every departure from an idealized process as an error to be stamped out. A common formulation describes resilient performance through the capacities to anticipate, monitor, respond, and learn. Safety-II and Resilience Engineering in a Nutshell
That lineage matters because we are not claiming to have invented resilience engineering. What Cybermaniacs is developing through Human Resilience Engineering (HRE) is an application of that systems logic to the human layer of cyber, AI, and digital risk.
Our working definition is:
Human Resilience Engineering is the deliberate design, measurement, and strengthening of the human and organizational capacities that allow people to anticipate, monitor, respond, adapt, and learn as digital and AI-enabled conditions change.
That is subtly but importantly different from asking how we prevent people from making mistakes.
Human Risk Management remains essential because we need to understand where human-related risk exists, what is driving it, which populations are exposed, and which interventions are likely to change the outcome. But a purely deficit-oriented conception of risk eventually hits a ceiling. Complex organizations do not become resilient because every person follows every procedure perfectly. They become resilient because people, technology, and organizational systems can continue producing good outcomes under conditions that were not completely anticipated in advance.
This is the conceptual move from managing human risk toward engineering human resilience.
Our broader idea of a Human Resilience System follows naturally from that. A Human Resilience System is the socio-technical architecture through which an organization can sense changes in human risk, understand the conditions producing them, maintain the capabilities people need to respond, intervene at the right level, and learn from what actually happens. That means connecting human-risk intelligence with competency, psychology, behavior, culture, work design, governance, technical signals, learning, escalation, interventions, and outcomes rather than treating each as an unrelated program.
You can see some of that direction already in the way we describe AI Workforce Risk Management and our work on Agentic Readiness & Change. The practical question is increasingly not simply whether employees know how to use AI safely, but whether the surrounding human-agent system has enough adaptive capacity to remain safe when tools, workloads, roles, and conditions keep changing.
A Human Resilience Engineering test for human oversight
If we apply an HRE lens to HITL, the standard becomes considerably harder than “a person approves this step.”
A credible human control needs to survive contact with actual work.
| HRE condition | What we should be asking |
|---|---|
| Purpose | What specific failure, ambiguity, or decision is the human expected to detect or resolve? |
| Observability | Can the person see enough context, provenance, system state, and evidence to make that judgment? |
| Capability | Does the person have the domain knowledge, AI understanding, and judgment required for this particular oversight task? |
| Capacity | Is the volume, frequency, complexity, and timing of review compatible with sustained human attention? |
| Authority | Can the person meaningfully challenge, override, stop, or redirect the system without needing three more approvals to do it? |
| Conditions | Do incentives, workload, culture, and workflow friction support the behavior the control assumes? |
| Escalation | Where does genuine uncertainty go when the reviewer cannot confidently resolve it? |
| Learning | Are overrides, near misses, unusual outputs, workarounds, and successful adaptations feeding back into the system? |
| Adaptation | As the AI, workflow, and human expertise change, who determines whether the control still works? |
None of those questions is particularly exotic. The unsettling part is how often the word human is allowed to stand in for all of them.
NIST's human-centered AI work is useful here because it similarly argues for examining AI in terms of the activities people and systems perform together rather than treating the model in isolation. That shift toward the architecture of the human-AI task is exactly what we need if oversight is going to become something we can genuinely design and measure. NIST AI Use Taxonomy: A Human-Centered Approach
The human should not become the universal exception handler
There is a deeper architectural problem here too. If every risk that cannot be fully eliminated technically is handed to a person, the human layer becomes the universal exception handler for the entire AI stack.
That is a terrible use of human capability.
Some risks should be constrained technically. Some actions should never reach the approval queue because permissions make them impossible in the first place. Some high-volume decisions should be sampled, monitored, or routed according to consequence rather than individually reviewed. Some outputs need provenance and evidence designed into the interface so verification is fast rather than forensic. Some tasks should remain human-led because the judgment involved is genuinely difficult to delegate. Others should become mostly autonomous because human approval adds ritual rather than safety.
This is why good Human Resilience Engineering cannot be another name for “train the humans better.” Sometimes the correct HRE intervention is to remove unnecessary cognitive work from the human altogether.
That idea can initially sound odd coming from a Human Risk Management company, but it is central to the discipline we think HRM has to become. If the evidence tells us that people repeatedly fail a control because the workflow requires an unrealistic level of vigilance, the answer is not automatically another learning module about the importance of vigilance. We need to be willing to say that the control itself is badly engineered.
Our work on cognitive offloading and AI explores the related question from the other direction: offloading cognitive work can be extremely useful, but it becomes risky when people lose the situational awareness or practiced judgment required to recognize when automated work has gone wrong. The design problem is therefore not “human or AI?” It is deciding which cognitive work belongs where, and how the combined system preserves the capability to adapt when conditions change.
Work-as-imagined versus work-as-done is about to become an AI governance problem
One of resilience engineering's most useful distinctions is between work-as-imagined and work-as-done. The former is how an organization believes a process works: the policy, workflow diagram, control matrix, responsibility model. The latter is how people actually make the thing work amid time pressure, awkward systems, incomplete information, competing goals, and all the other inconvenient features of reality. Research in resilience engineering treats the gap between the two as something worth understanding rather than automatically interpreting it as worker non-compliance.
AI is going to make that gap fascinating.
Work-as-imagined may say that an employee reviews every significant agent output before execution. Work-as-done may reveal that the agent is accurate enough that detailed review feels wasteful, the queue is growing faster than the employee can examine it, the explanations are not particularly useful, and experienced employees have developed an informal pattern for deciding which outputs deserve real scrutiny.
Traditional compliance thinking might look at that and conclude that people need to be reminded to follow the process.
Human Resilience Engineering asks a more useful set of questions. What adaptation have people made? Why did they need to make it? Does that adaptation increase risk, reduce risk, or perhaps do both under different conditions? What information are experienced workers using that the formal control failed to capture? Can the system be redesigned so the useful adaptation becomes part of the operating model rather than remaining invisible tribal knowledge?
Working with clients on human risk has taught us repeatedly that what people actually do is often diagnostic information about the system. Workarounds, shortcuts, unusual verification habits, reluctance to adopt a tool, or excessive confidence in one can all tell us something that a policy-compliance metric cannot.
That becomes even more valuable as human and non-human work become intertwined.
Resilient oversight has to be able to anticipate, monitor, respond, and learn
The four classic resilience capabilities give us a more useful way to think about AI oversight than the binary question of whether a human is present.
Anticipation means understanding where the human-agent combination is likely to become brittle before an incident proves it. Which tasks are becoming more autonomous? Where is human expertise decaying because the system normally performs the work? Where could volume overwhelm meaningful review? Which roles are inheriting accountability without equivalent visibility or authority?
Monitoring means seeing whether the conditions assumed by the control remain true. That includes more than agent telemetry. Human Risk Management can contribute evidence about competency, confidence, verification behavior, escalation, culture, workload conditions, and the different ways populations are adapting as AI becomes normal.
Response requires real intervention capacity. The employee has to know what to do, possess the authority to do it, and have a usable route for situations that exceed their expertise. This is where our work on verification behavior as a security control becomes relevant: telling somebody to verify is easy; designing proportionate, role-specific verification into the actual handoff between human and machine is harder and far more useful.
And learning means that the control evolves. If people regularly override a particular recommendation, that is evidence. If almost nobody ever overrides anything, that is evidence too — although perhaps not the evidence the dashboard initially suggests. Near misses, strange outputs, successful interventions, workarounds, questions, and escalation patterns should all make the next version of the system better.
Otherwise, we are not operating a resilient control. We are repeatedly asking the same human to compensate for the same unexamined weakness.
The goal is not humans in every loop. It is resilience across the system.
This may be the most important distinction.
A Human Resilience System does not maximize human involvement. It puts human capability where it contributes meaningful adaptive capacity and supports that capability with the right technology, information, authority, culture, and learning mechanisms.
Sometimes that means a consequential decision needs careful human judgment. Sometimes it means an agent should act autonomously inside tightly engineered boundaries while humans monitor higher-order conditions. Sometimes it means multiple people need to contribute different forms of expertise. Sometimes the correct design is to stop asking a human to approve something that no human could realistically verify at the speed and scale required.
The objective is not to preserve human activity for sentimental reasons. Nor is it to automate away the irritating unpredictability of people. Complex socio-technical systems need both reliable technical constraints and human adaptability. The engineering challenge is deciding how those capacities should interact.
That is also why we think Human Resilience Engineering eventually reaches beyond AI governance. Cybersecurity has spent years framing people predominantly in deficit terms: error rates, failure to comply, phishing susceptibility, risky behavior. Human Risk Management improved that picture by asking what conditions are producing risk and what interventions can change it. Human Resilience Engineering adds another question: what capabilities must we deliberately preserve and strengthen so the organization can continue to function safely when the environment changes faster than its controls?
In an AI-enabled organization, that is not a philosophical distinction. It is an operating requirement.
Where Cybermaniacs is taking this
Our work with customers across Human Risk Management, AI workforce readiness, culture, learning, and now agentic change is pushing us toward a much more integrated model of the human layer.
We already know that competency alone does not explain behavior. Psychology matters. Culture matters. Organizational conditions matter. Work design matters. A technically capable employee can still make a poor decision under pressure; a less experienced employee can produce an excellent outcome because the surrounding system gives them strong cues, clear escalation, usable guardrails, and enough psychological safety to stop when something looks wrong.
AI adds new variables to that system — reliance, influence, cognitive offloading, delegation, intervention, override quality, changing skills, and new forms of human-machine interdependence — but the underlying lesson is familiar. Human performance is contextual.
That is why our Agentic Readiness & Change work does not start and end with agent literacy. We look at roles, workflows, human oversight, governance, capability, behavior, and the organizational conditions in which people will actually work with agents. And it is why the broader Cybermaniacs direction is increasingly about building Human Resilience Systems rather than simply producing more interventions at the point where somebody already appears risky.
There is a satisfying irony here. AI was supposed to remove a great deal of mundane human work. We should probably resist designing its governance in a way that gives all of that time back as an endless queue of approval requests.
Lucy did not need another memo explaining the importance of wrapping chocolates correctly. Somebody needed to slow down the conveyor belt, redesign the job, or change the system around it.
Our AI controls deserve at least that much thought.
Frequently Asked Questions
What does human-in-the-loop mean in AI?
Human-in-the-loop, or HITL, describes an AI or automated workflow in which a person participates at some point in the decision or action process. Depending on the use case, the person may review an output, approve an action, correct the system, provide feedback, intervene in an exception, or retain final decision authority.
Why is human-in-the-loop not automatically an effective control?
Human presence does not guarantee meaningful oversight. The person also needs sufficient information, capability, attention, authority, and time to perform the intended oversight task. High workload, complex verification, over-reliance on accurate automation, weak escalation, poor interfaces, or incentives that discourage intervention can all reduce the effectiveness of a HITL control.
What is Human Resilience Engineering?
Human Resilience Engineering (HRE) is Cybermaniacs' developing approach to deliberately designing, measuring, and strengthening the human and organizational capacities required to anticipate, monitor, respond, adapt, and learn as digital and AI-enabled conditions change. It draws on established resilience-engineering and human-risk principles while applying them specifically to cyber, AI, and workforce risk.
What is a Human Resilience System?
A Human Resilience System is the broader socio-technical architecture that connects human-risk intelligence, capability, behavior, psychology, culture, work design, governance, technical signals, interventions, escalation, and learning so that organizations can adapt safely as risk and working conditions change.
How is Human Resilience Engineering different from Human Risk Management?
Human Risk Management focuses on identifying, understanding, prioritizing, and reducing human-related risk. Human Resilience Engineering extends that perspective by asking what adaptive capacities the human and organizational system needs in order to continue producing safe outcomes under expected and unexpected conditions. The two are complementary: HRM helps identify and manage risk, while HRE deliberately builds capacity for resilient performance.
What makes human oversight of AI effective?
Effective AI oversight requires a clearly defined purpose, adequate visibility into the system and decision context, appropriate competency, manageable workload, authority to intervene, supportive cultural and organizational conditions, usable escalation paths, and mechanisms for learning from overrides, failures, near misses, and successful adaptations.
Does Human Resilience Engineering mean putting more humans into AI workflows?
No. In some cases, better resilience engineering may reduce unnecessary human approval. The objective is to place human capability where judgment and adaptation add value while using technical controls, automation, risk-based routing, and better system design where they can provide more reliable protection.