Who Is Minding the Bot? Human Oversight and AI
10.7.2026

Since the launch of ChatGPT in 2023, the use of artificial intelligence to do research, take over customer service functions and streamline tasks promises to radically change how people interact with the world. At the same time, we’ve also seen some of the perils of using AI uncritically: hallucinations in legal briefs, uninformed chatbots that help prepare a case pro se, students using AI to draft essays, or in even more life-threatening situations such as recent reports of people following directions from chatbots to do harm to themselves and self-driving cars causing accidents.[1]
It’s become clear that AI is not yet at the point where it can be trusted to always do the right thing. Any person or company using AI needs to be aware of the risks. Even more important, there needs to be a reliable framework for its use that includes human oversight. New laws, regulations, and organizational policies are increasingly requiring that actual human beings oversee the AI programs they put into place. The EU AI Act, for example, employs human oversight as a necessary component of responsible AI governance. A term that has become popular in referring to the use of human oversight in AI systems is human-in-the-loop, suggesting that people are integral to any system using AI.
Yet organizations often struggle to put effective human oversight into practice. Simply placing a human somewhere in a workflow does not necessarily reduce risk or satisfy the underlying purpose of preventing bad outcomes or harm. Meaningful human oversight requires a person who can identify problems, exercise independent judgment and intervene before harm or injury occurs.
AI Incidents on the Rise
According to the AI Incident Database, which collects and indexes various types of incidents that cause harm due to AI, the number of such reports have doubled each year between 2022 and 2025. By October 2025, the number had already increased beyond the previous year.[2]
Nor have regulations kept pace with AI use. This is for three main reasons. First is the speed at which AI has been adopted. According to Tom Wheeler, a visiting fellow at the Brookings Institute’s Center for Technology Innovation, “existing rules are insufficiently agile to deal with the velocity of AI development.”[3] Second is the technology’s multifaceted nature. Varying types, responses to prompts and unexpected applications make AI technology difficult to categorize neatly. Guidelines for attorneys using AI for research are going to be different from laws to prevent children from using bots to plan self-harm. Finally, the regulatory landscape is complex and fragmented. This is evident in the U.S.’s state-by-state approach and more so on a global level with major jurisdictions like the U.S. and EU employing different regulatory strategies.[4]
Still, regardless of the industry, regulation or framework, one piece of guidance is relatively common, and that is human oversight. The goal is straightforward: To ensure that an AI solution, an application that uses AI to address specific problems, such as a retail company deploying an AI-recommendation tool that analyzes customer purchase history to suggest products, operates safely and as expected, a human should be part of the decision-making process at critical points to monitor outputs. Human feedback and/or intervention is vital. However, questions tend to arise around when and how exactly human oversight ought to be implemented as a governance mechanism.
We must therefore explore how to meaningfully implement human oversight rather than superficially document instances of what is merely acceptable from a compliance standpoint in order to mitigate harm and liability.
Defining Human-in-the-Loop
While some prefer terms like human-in-the-loop, human-over-the-loop, human-out-of-the-loop, etc., such terms can sometimes be used inaccurately and/or interchangeably. They all refer to:
“AI systems that include human feedback or intervention as part of their operation. In these systems, humans may provide guidance, correct errors, or make final decisions to improve the accuracy and reliability of AI. This approach combines the strengths of both humans and machines, ensuring that complex or high-stakes tasks benefit from human judgment and oversight.”[5]
Human-in-the-loop is mentioned by name or alluded to in almost every playbook providing guidance to groups using AI, especially in the context of systems that have the potential to cause material harm. To summarize a few areas where we see this concept:
- New York City’s Local 144 prohibits employers or employment agencies from using an automated employment decision tool to make an employment decision unless the tool is audited for bias annually. The employer publishes a public summary of the audit and provides certain notices to applicants and employees who are subject to screening by the tool. According to labor and employment attorneys at Littler, the “threshold question is whether an employer is using an automated employment decision tool to make a covered employment decision and thus subject to the audit and notice requirements of the law,” where a “covered employment decision” is one that involves hiring or promotion.[6] Here, while human-in the loop is not explicitly prescribed, it does serve as a tool to counteract bias.
- The EU AI Act, Art. XIV (“Human Oversight”) requires that high-risk AI systems be designed so humans can effectively monitor, understand and override them. Its goal is to prevent harm to health, safety and fundamental rights by keeping algorithms under meaningful control. The AI system must also enable overseers to understand AI’s limits, recognize bad outputs and safely intervene or stop the system before harm occurs.
- The National Institute of Standards and Technology AI Risk Management Framework recognizes that humans need to be more involved in some decisions made by AI than others: “Human-AI configurations can span from fully autonomous to fully manual. AI systems can autonomously make decisions, defer decision-making to a human expert, or be used by a human decision maker as an additional opinion. Some AI systems may not require human oversight, such as models used to improve video compression. Other systems may specifically require human oversight.”[7]
This, of course, is a non-exhaustive list. Organizations like the International Organization for Standardization and the International Electrotechnical Commission’s international standard for creating an AI management system, the Organization for Economic Cooperation and Development’s AI Principles and states like Colorado also make recommendations for some aspects of human involvement. However, it does effectively demonstrate the range of ways in which human-in-the-loop is a vital part of AI oversight.
Ultimately, a simple litmus test should leverage an ethical, commonsense approach to compliance: If a system is reasonably deemed capable of materially harming an individual, a higher level of human involvement should be required in an organization’s governance process. The next question then is what is meant by “material harm”? And then, what is meant by “proportional human involvement”?
Less Than Effective Examples of Human Oversight
Across industries and different types of systems, a number of people and organizations are not quite getting compliance right when it comes to human oversight, as shown by examples mentioned above and below. It could be that those in charge may not be doing enough to anticipate and prevent harm. This is not to suggest that these entities are unethical, only that when laws and regulations are overly burdensome or poorly tailored, the process can turn into a bureaucratic checklist rather than meaningful reforms. This approach can lead to a hyperfocus on parts of the process and miss the bigger picture or lose sight of the AI tool’s broader purpose. The hallucination of legal citations illustrates this problem: A user may focus narrowly on whether AI has produced citations that appear responsive to a research request, rather than whether those authorities actually exist and support the proposition for which they are being cited. In doing so, the user loses sight of the broader purpose of the tool, which should be to assist – rather than replace – the attorney’s independent research, judgment and responsibility for the final work product.
Meaningful human oversight can be especially difficult to design, as it – as a control – is susceptible to a range of well-documented cognitive biases. For example, a reviewer looking at a 50-page AI-generated report of a generated model might assume that the report’s conclusions are likely accurate because the report was produced by a sophisticated system. Similarly, when multiple reviewers are involved and there is no clarity about who should be checking what, each may assume someone else has already performed a detailed review. In both cases, oversight exists on paper but may be ineffective in practice because of user bias or assumptions. Collectively, blind spots like these can reduce reviewers from active decision-makers to passive monitors, undermining the very safeguards that human oversight is intended to provide. Essentially, the diligence an individual applies may diminish when that person believes someone or something else is also responsible for the output, especially when it’s done by a nonhuman tool seen as infallible. This varies depending on the extent to which one perceives authority or expertise. In other words, whether it is a group project or an AI output, the need one feels to contribute diminishes since some individuals might consciously or unconsciously believe that “the rest of the group won’t allow us to fail, so I can do a little less” or “the likelihood of the AI getting it wrong is probably low, so I don’t need to review too closely.”
Weaknesses in the System: The Example of Michael Cohen
Shortly after the first mainstream generative AI models became prevalent, instances of misuse by attorneys began to appear. Arguably the most famous of these cases involves Michael Cohen, President Trump’s former lawyer, and Cohen’s counsel. In 2023, Cohen was in the process of securing an early end to a 2018 court-ordered supervised release after serving for charges ranging from campaign finance fraud to tax evasion. Cohen passed along AI-generated, hallucinated case citations to David Schwartz, his attorney at the time, who then submitted them to the judge as part of their filings, who subsequently discovered the hallucinations.
There are a few notable takeaways from this incident. One comes from Cohen blaming Schwartz for failing to check the validity of his citations before submitting them to the judge after stating that as a non-lawyer (he was disbarred) he has “not kept up with emerging trends (and related risks) in legal technology and did not realize that Google Bard was a generative text service that, like ChatGPT, could show citations and descriptions that looked real but actually were not.” He “understood it to be a super-charged search engine and had repeatedly used it in other contexts to successfully find accurate information online.”[8] This is an effective articulation of automation bias, to which Cohen had succumbed.
However, the failure was not just Cohen’s. Schwartz, as the attorney, had an ethical obligation as the responsible party for conducting an appropriate review of the information Cohen gave him. He acknowledged that he had failed in his duty, stating, “I am fully aware that I bear the responsibility for any submission on my letterhead and the inaccuracies contained in this filing are completely unacceptable. … I sincerely apologize to the court for not checking these cases personally before submitting them to the court.”[9] Ultimately, U.S. District Judge Jesse M. Furman determined that the false citations were not submitted in bad faith, so he did not impose sanctions. He did, however, in a March 2024 ruling, refer to the incident as “embarrassing.” Cohen’s other representative in this matter, E. Danya Perry, stated that “even a quick read” of the citations “should have raised an eyebrow.”[10] Another key takeaway is that at this point, courts had already been seeing an increase in case filings riddled with AI hallucinations; this just happened to be particularly high-profile. And, while one might assume that media attention would lead to the elimination of such carelessness, since then, lawyers and law firms continue to make these kinds of mistakes despite the document-review resources at their disposal.[11]
Failures of human oversight can be observed across a variety of areas besides the practice of law. The hallucination problem, for example, can be seen in multiple professional consulting instances. In 2025, Deloitte offered the Australian government a partial refund “for a report that was littered with AI-hallucinated quotes and references to nonexistent research.”[12] This year, the consulting firm EY had to withdraw a study it produced that “was used by EY consultants in Canada to market their cybersecurity business, used made-up data, misattributed citations and referenced a McKinsey report that does not exist, online researchers discovered.”[13] KPMG, one of the largest consulting firms in the world, was also recently accused of “publishing a report on the future of AI with 40 of 45 citations that appeared to be at least partially AI hallucinations,” resulting in KPMG having to take it down.[14]
When AI Goes Really Wrong: Self-Driving Cars and Failures of Human-in-the-Loop
Among the most serious examples of ineffective human oversight is the death of Elaine Herzberg in 2018 by a self-driving car. In that case, while Uber was test-driving autonomous vehicles in Arizona, a backup driver who was seated in the driver’s seat expressly to take action in case of AI failure was allegedly streaming an episode of “The Voice” when the vehicle struck and killed Herzberg, a pedestrian. Publicly reported evidence included dashcam footage showing the driver “looking down, away from the road, for several seconds immediately before the crash, while the car was traveling at 39 mph,” and records from the streaming service Hulu. One report showed that from the start of the trip up until the accident, the backup driver was on her phone 34% of the time.[15] The driver, originally charged with homicide, pled guilty to a charge of endangerment in July 2023.
The investigation also found that Uber’s system failed to include the potential of a pedestrian crossing outside of a crosswalk and that the company had removed a requirement for there to be two backup drivers in each car in the months leading up to the accident. Drivers were also assigned the same monotonous route throughout their usual eight-hour shifts, making them more likely to seek distraction. Uber was not charged but pledged to make changes to address these issues.[16]
Meaningful and Ethical Design Choices
It is unlikely that those individuals and organizations held accountable for these human oversight failures intended to deceive or cause harm. Instead, the reality is that AI and its potential to do harm are still being understood. Any regulations that are imposed, whether by governing bodies or within companies themselves, need to include an appropriate level of human oversight. While the use of AI at this scale is evolving, there are steps that can be taken to reduce the potential for harm.
When assessing whether an AI system has sufficient potential for harm, regulations like the EU AI Act can get quite prescriptive. For instance, the act’s Annex III lists high-risk use cases including biometrics, critical infrastructure, employment, law enforcement, immigration and the administration of democratic processes. Borrowing from the EU AI Act and other relevant frameworks and regulations can help formulate a practical test on whether meaningful oversight should be implemented when a given AI system (1) has the potential to produce an inaccurate result or unintended consequence, recommendation or decision that (2) could reasonably result in significant harm to an individual, an organization, or the public, where “harm” includes:
- Physical injury
- Financial loss
- Denial of employment or educational opportunities
- Adverse legal consequences
- Discrimination
- Privacy or cybersecurity harms
- Reputational damage
- Other decisions that materially affect a person’s rights, safety or well-being
Even simpler is to ask: What is the worst that could happen if a given AI system fails? And if the answer is that someone could suffer material harm, meaningful human oversight should be embedded into the system’s lifecycle at the point where it would have optimal impact. But what is meant by meaningful human oversight?
In order to have meaningful human oversight with the right level of proportional human involvement, the person has the information, authority, ability and opportunity to independently evaluate an AI system’s output and intervene before material harm occurs. Put simply, meaningful human oversight requires more than a human in the process – it requires a human who can meaningfully influence the outcome and who is properly empowered to do so. In order to enable such human oversight, we must also consider social and cultural concepts that can influence whether intervention will be effective, as we’ve seen with the Michael Cohen case.
Going back to the case of Uber’s self-driving car hitting a pedestrian, the role of smartphones is one area that needs to be addressed. In a scenario where a person is tasked with serving as a backup driver for a self-driving vehicle that can potentially cause physical harm to other drivers or pedestrians, a responsible use of human-in-the-loop would require that phones not be used during test sessions or at least restricting access to certain apps so that emergency calls can still be made. Apps that may lead an individual to engage actively and for extended periods of time should be blocked from access, such as streaming apps and mobile games. Backup drivers for self-driving cars should be expected to assume as much responsibility as if they were driving the cars themselves.
In the case of legal case hallucinations, someone seeking representation who trusts legal counsel to practice competently (an expectation that is also codified and required under the ABA Model Rules of Professional Conduct), which includes reviewing AI decisions and actions, can end up losing a case, resulting in significant damages or incarceration. This is a material harm and therefore requires that there be formalized processes to ensure that filings have been properly reviewed.
In the case of United States v. Cohen, Schwartz said he thought drafts of the papers to be submitted to the court were reviewed by Cohen’s other counsel in the matter. If faced with a similar scenario, attorneys should know to verify such an assumption, rather than diffusing responsibility and carelessly filing case citations that are handed to them. In the consulting industry, clients similarly trust reputable firms to produce reliable information when making critical decisions that have impacts on their business reputation and financial condition, and therefore facts and footnotes should be properly vetted before publication.
Conclusion: Questions To Consider for AI Governance
It is easy to compile a list of mistakes and criticize those who committed them. However, as attorneys, we have a responsibility to learn the right lessons when things do go wrong and use them to enhance our processes to meaningfully address governance controls like human-in-the-loop. Effective governance of AI systems and risk management relies on the understanding that the processes that support them are iterative and constantly improving based on what we learn in taking actions, and what we reasonably can anticipate before taking those actions. That is why ethical reasoning is critical to deploying effective and sustainable governance processes. Before deploying or approving the use of an AI system, legal and compliance teams should consider:
- What is the reasonably foreseeable harm if the system makes a mistake?
- Who is accountable for reviewing high-risk decisions?
- Does the reviewer have sufficient expertise to identify errors or inappropriate recommendations?
- Does the reviewer have authority to override, modify or reject the AI-generated output before harm can be done?
- Are responsibilities clearly assigned, or is there a risk that multiple reviewers assume someone else performed the review?
- Are reviewers subject to distractions, incentives or workflow pressures that could undermine effective oversight?
- Can the organization demonstrate that meaningful review occurred if challenged by a regulator, customer, court or auditor?
- Is the oversight process calibrated to the severity of the potential harm?
- Does the reviewer have sufficient information and context to identify and correct potentially biased or discriminatory outputs, rather than simply reinforcing them through human approval?
As AI continues to evolve and becomes more integrated into our work and our lives, human oversight is critical.
Ultimately, effective compliance is not simply about whether humans appear somewhere in the workflow; it is about whether those people have the information, authority, ability and opportunity to prevent harm before it occurs. It’s about checking biases and assumptions, critically reviewing what AI is telling us and implementing safeguards to prevent real harm. If we truly want to fully harness all that AI has to offer, it’s going to mean doing the opposite of letting AI do all the work. It will mean instituting human-led policies and practices, ones that rest on those qualities that make us human: empathy, ethics, and ingenuity.
Matthew Lowe is an associate general counsel for a government contractor. He is a fellow of information privacy with the IAPP and lectures at the University of Connecticut School of Law on data privacy, cybersecurity law and AI governance. He has also taught these subjects at Cornell University and the University of Massachusetts Amherst. He is a member of the New York State Bar Association’s Committee on Technology and the Legal Profession and the Committee on Communications and Publications.
Endnotes:
[1] John Sanford, Why AI Companions and Young People Can Make for a Dangerous Mix, Stanford Medicine, Aug. 27, 2025, https://med.stanford.edu/news/insights/2025/08/ai-chatbots-kids-teens-artificial-intelligence.html.
[2] Harry Booth, What the Numbers Show About AI’s Harms, Time, Jan. 19, 2026, https://time.com/7346091/ai-harm-risk/.
[3] Tom Wheeler, The Three Challenges of AI Regulation, Brookings, June 15, 2023, https://www.brookings.edu/articles/the-three-challenges-of-ai-regulation/.https://www.brookings.edu/articles/the-three-challenges-of-ai-regulation/.
[4] Id.
[5] What Is Human-in-the-Loop, Stanford HAI, https://hai.stanford.edu/ai-definitions/what-is-human-in-the-loop.
[6] Jim Paretti et al., New York City Adopts Final Regulations on Use of AI in Hiring and Promotion, Extends Enforcement Date to July 5, 2023, Littler Mendelson, April 13, 2023, https://www.littler.com/news-analysis/asap/new-york-city-adopts-final-regulations-use-ai-hiring-and-promotion-extends.
[7]. Artificial Intelligence Act, Regulation (EU) 2024/1689, art. 14(4)(b), 2024 O.J. (L. 1689) 1.
[8] The Associated Press, Michael Cohen Says He Unwittingly Sent AI-Generated Fake Legal Cases to His Attorney, NPR, Dec. 30, 2023, https://www.npr.org/2023/12/30/1222273745/michael-cohen-ai-fake-legal-cases.
[9] Megan Brand, Michael Cohen Says He Unknowingly Submitted Fake AI-Generated Legal Cases to Lawyer, NBC News, Dec. 29, 2023, https://www.nbcnews.com/politics/politics-news/michael-cohen-says-unknowingly-submitted-fake-ai-generated-legal-cases-rcna131631.
[10] United States v. Cohen, 724 F. Supp. 3d 251 (2024), https://www.nysd.uscourts.gov/sites/default/files/2024-03/18cr602%20Cohen%20Opinion.pdf.
[11] Chris Stokel-Walker, AI Is Flooding the Courts With More Cases, More Filings, and More Fake Citations, Fast Company, May 11, 2026, https://www.fastcompany.com/91539168/ai-is-flooding-the-courts-with-more-cases-more-filings-and-more-fake-citations.
[12] Kyle Orland, Deloitte Will Refund Australian Government for AI Hallucination-Filled Report, Ars Technica, Oct. 6, 2025, https://arstechnica.com/ai/2025/10/deloitte-will-refund-australian-government-for-ai-hallucination-filled-report/.
[13] Stephen Foley, EY Retracts Study After Researchers Discover AI Hallucinations, Financial Review, May 17, 2026, https://www.afr.com/world/north-america/ey-retracts-study-after-researchers-discover-ai-hallucinations-20260517-p5zxss.
[14] Will McCurdy, KPMG Allegedly Published AI Report Filled with Hallucinations, PC Mag, June 15, 2026, https://www.pcmag.com/news/kpmg-allegedly-published-ai-report-filled-with-hallucinations.
[15] Rory Cellan-Jones, Uber’s Self-Driving Operator Charged Over Fatal Crash, BBC, Sept. 16, 2020, https://www.bbc.com/news/technology-54175359.
[16] Backup Driver of Self-Driving Uber Car Takes Plea Deal for Fatal Crash in Tempe, 12News, July 28, 2023, https://www.12news.com/article/news/local/valley/driver-self-driving-uber-car-takes-plea-deal-fatal-crash-in-tempe/75-d75e0710-d757-4640-9b53-5867c05ea041.




