The Human Holding the Loop
Does human presence equal human oversight?
Does human presence equal human oversight?
We woke up to some interesting news this morning. Dario Amodei and Sam Altman, the CEOs of Anthropic and OpenAI, are publicly discussing the fact that we may need to put the brakes on the development of AI, specifically the development of frontier models. At the same time, Evan Hubinger, who leads alignment research at Anthropic (but was speaking on his own behalf), has said that he believes there is a greater than 10 percent chance that AI could eventually kill all humans within the next ten years.
Of course, the way this is being reported in news outlets and around the internet emphasizes the “AI is going to kill us all in ten years” concept, and that’s not exactly what was said. These are estimates of risk, not predictions. But the fact that leaders at two of the companies working at the leading edge of AI are saying that development may need to slow down and that there are dangers we need to be aware of is something all of us should be paying attention to.
There have always been people who decry the creation of AI and espouse theories that it’s going to take over and destroy us. On the other side are those who completely embrace AI, or who believe that it may be dangerous but feel there’s nothing they can do about it, so they simply brush it off. Basically, either panic or shrug.
But here’s the thing: it doesn’t really matter if the probability of meaningful loss of human control is 10 percent or 1 percent or some much smaller number. When the potential consequence is catastrophic, even a relatively small probability deserves serious attention. As AI becomes more capable and more autonomous, humans have to remain in control of what we’ve created.
The conversation we’re hearing now exposes the need for this kind of oversight and thought at the AI development level, and many people have also talked about the need for conversation and regulation at the government level. But I think the solution needs to begin much smaller than that, at the level of the individual. And I happen to have an interesting and somewhat funny example of how a human needs to maintain control of AI.
I have a couple of YouTube channels, and I sometimes use Ask Studio, YouTube’s AI, to help me understand my analytics. Recently I asked it to give me analytics for my most recent video. It gave me four different metrics and then compared them against the typical metrics for those same four measurements. The problem was that the numbers for the current video and the typical numbers were exactly the same.
Something didn’t seem right, so I questioned it. The AI gave me a confident explanation that the video just happened to be performing exactly at the typical level across every measure. It was early in the analytics, it explained, and for several reasons it was entirely possible that all four numbers simply matched.
That still didn’t make sense, so I asked it for evidence. Was it assuming this, or was it using specific numbers? What were the specific numbers it was using? It gave me a longer and even more confident answer and told me that it was looking at the data and extrapolating the numbers from the information it had.
So I said, okay, show me the data.
And then it admitted that it didn’t actually have the data. It was extrapolating from what it knew about how the analytics worked. It had come to the wrong conclusion and then kept working hard at trying to sound convincing about why it was the right conclusion.
That’s a funny example and an utterly low-stakes one, but I think it demonstrates a serious lesson. We often hear that humans must stay “in the loop” in order to maintain control over AI. But what does being in the loop actually mean? I was in the loop from the beginning to the end of that conversation, but had I simply accepted what the AI told me, I would have accepted a wrong answer. The fact that I was there wasn’t enough.
Human oversight is more than human presence. It’s human authority. Or, to continue with the metaphor, it’s not enough to keep a human in the loop. The human must keep hold of the loop. That means the human needs to understand enough to question, verify when necessary, override when appropriate, and stop the process when required.
Almost all of us at this point are using AI in some fashion. Even if we don’t deliberately use it in our personal work, we encounter it when we’re searching for products and reviews on Amazon or searching for information on Google. So we all have some power to determine how we want to use AI. The question really isn’t whether we’re using AI. The question is: what are we delegating to AI? And there’s a big difference between delegating work and delegating judgment.
In practical terms, that means that when we use AI, we need to question outputs that don’t make sense. We need to verify claims when the consequences of acting on those claims matter. We need to understand the limitations of the system we’re using and not let confident language convince us of an incorrect answer. This becomes even more important when we’re using AI agents that can actually take actions on our behalf. We need to know what permissions an AI agent has before allowing it to act, and we need to preserve the ability to reverse, override, or stop what AI is doing.
The YouTube example is a tiny version of keeping that kind of control. The stakes didn’t matter in that case, but building the habit did, so that the habit exists when the stakes do matter. The solution begins with the individual, and then it scales upward step by step. It can’t reasonably depend on starting at the top and working its way down.
As individuals, when we start creating the good habit of staying in control of the loop, we can carry the same habit into our careers as employees. The organizations we work in probably have their own AI strategies, and we may not be able to directly influence those strategies. But we certainly can influence the way we use AI in our own work. The questions become even more consequential in this case. What information am I giving the AI? What am I asking it to produce? Am I letting it influence decisions? Have I verified what I’m putting my name on or what I’m producing? Would I know if the output were wrong? Do I have enough knowledge of the subject to be working effectively with the AI in the first place? Delegating a task to AI doesn’t delegate responsibility for its outcome, or at least it shouldn’t. That responsibility remains with the human being.
And when we’re working within organizations, particularly with data that belongs to our company or its customers, I think our responsibility begins to extend beyond ourselves. We have some responsibility for our organization and our peers. You don’t need authority to become part of the solution. People without formal authority shape organizational behavior every single day. While we’re developing good habits around AI use and verification, we should be modeling those habits for our colleagues. We should be asking whether a particular piece of output was verified. When we come up with an effective way of checking AI-generated work, we should share it. When we see something in an AI-automated process that suggests decisions are being made without meaningful human intervention, we should raise the concern.
And if we become champions of working productively with AI, our voices will carry more weight when we’re also willing to say, stop, this is not a correct use of AI. This is something that needs to be reconsidered or guardrailed. We need to help hold the line between reflexive AI adoption and reflexive AI rejection. Being informal leaders is one way we can truly become part of the solution, regardless of our role.
For those in formal leadership, the responsibility grows even greater. Managers, directors, and organizational executives control things that determine how their employees work. They determine workflows and responsibilities. They establish expectations and timelines. In the case of formal leadership, keeping hold of the loop means deliberately designing how human oversight is going to work when AI becomes part of those processes.
Leaders constantly have to be asking themselves where AI can assist, where it can recommend, where it can act, and where a human must approve an action. There are also decisions that should never be delegated, and there needs to be a clear process for what happens if something goes wrong. Accountability needs to be baked into the workflow and into the human-AI collaboration. Importantly, the person charged with oversight of the process must have the knowledge, time, and authority necessary to exercise it. Formal leaders also need to create an environment in which employees feel safe questioning AI systems without being treated as obstacles to innovation. Again, not blind AI rejection, but not blind AI adoption either.
The organization as a whole then creates the institutional infrastructure that makes all of this work. Individual good judgment can’t carry the entire burden. Organizations need systems, policies, and guardrails that preserve human control. They need to name the humans who are accountable, create boundaries around data use and permissions, determine when human approval is necessary, test new AI products before deployment, monitor them during deployment, and audit their output afterward. They need escalation procedures and mechanisms for reporting unexpected behavior. And they need both the ability and the strength of will to interrupt or shut down systems if something isn’t working correctly. The organization isn’t just inserting humans into workflows. It needs to design workflows in which humans remain capable and expected to exercise judgment over artificial intelligence.
That distinction between simply having a human present and giving that human meaningful control is beginning to appear in AI regulation as well. The European Union’s AI Act, for example, requires human oversight of high-risk AI systems to include understanding the system’s capabilities and limitations, recognizing the risk of over-relying on its output, being able to override or reverse its decisions, and being able to interrupt the system when necessary.
And that brings us back to the actual builders of AI. The same principles we’ve discussed here apply to the builders, but with vastly greater stakes. OpenAI, Anthropic, and other frontier developers have to ask the same questions of themselves, but at scale. How much autonomy should a system have? What capabilities require additional safeguards? How is dangerous behavior tested before release? Can the system circumvent its safeguards? Who decides when a model is safe enough to deploy? And who checks on that decision-maker?
These companies also need to find a way to balance their natural commercial competition with an appropriate focus on safety so that competitive pressure doesn’t force everyone to move too quickly. It’s actually good news that the CEOs of Anthropic and OpenAI are voluntarily stepping up and saying that they think the creation of frontier models needs to slow down. But voluntary restraint at the company level isn’t the only mechanism available. We’ve already seen some governments regulate AI use, and others are in active discussions about doing so. The problem is that poorly designed regulation can inhibit innovation. On the other hand, too little regulation can leave significant risks governed by companies that have powerful competitive incentives to move forward despite those risks.
Most individuals won’t have direct influence over when and how governments regulate AI, although we can make our voices heard through elected representatives and public-comment processes. What matters at this level is that regulation preserve the same principle we’ve been applying all along: meaningful human control.
There are different authorities at each of these levels and different stakes. The responsibility is shared, but it isn’t equal. The greater the authority over how AI is developed or used, the greater the responsibility for maintaining meaningful human control.
So back to the interesting news we heard this morning. Maybe AI does eventually present a significant risk to humanity. Or maybe the estimates we’re hearing now will eventually prove wildly overinflated. But we don’t need to be sure one way or the other before we start acting responsibly. There are two tempting reactions to extraordinary warnings. One is panic: assume we’re doomed, live in fear, and deny the use of something that can greatly advance humanity. The other is to shrug and say, “There’s nothing I can do about it.”
I don’t believe either one is the appropriate response. I think there’s a third response, and that’s to practice keeping hold of the loop.
An individual can question, an employee can verify, an informal leader can influence, a formal leader can establish norms, an organization can build safeguards, and an AI company can design for control. Finally, a government can establish boundaries where voluntary ones are insufficient.
We may not individually control what happens at the frontier of artificial intelligence, but every one of us who uses it is helping establish what human-AI collaboration can look like. We can surrender our judgment a little bit at a time because AI is fast, capable, and often right. Or we can become part of using and building something better.
Because being the human in the loop is not enough. The human needs to keep hold of it.
Published 13 September 2026