AI Hallucinated Legal Evidence: ChatGPT Fabricated Witness Testimony in Murder Appeal
New Mexico Supreme Court Fines Attorney $5,000 After Unverified AI-Generated Evidence Entered a Murder Conviction Appeal
AI hallucinated legal evidence in a murder-conviction appeal, and the attorney who trusted the output without adequately validating it was held in direct contempt and sanctioned $5,000.
On September 9, 2026, the New Mexico Supreme Court issued a dispositional order in State v. Sandoval, involving attorney Stephen D. Aarons and the appeal of Oscar Renee Sandoval. The court found that the appellate brief contained false testimony attributed to fabricated witnesses, along with other factual and legal misrepresentations.
The court used striking language: the filing contained “false testimony from wholly fabricated witnesses.”
Aarons acknowledged using ChatGPT while preparing the appellate brief. According to Reuters’ September 11, 2026 reporting, he had provided trial materials to ChatGPT and expected the technology to help produce what he described as a reliable summary.
ChatGPT created fictitious witness testimony that did not exist anywhere in the factual details or trial record. The attorney failed to catch those fabrications before submitting the brief. Consequently, an AI hallucination moved from a generative AI system into an official filing before a state supreme court.
That progression makes this case an important AI literacy lesson for lawyers, executives, compliance leaders, auditors, healthcare professionals, HR teams, investigators, and anyone using generative AI to summarize authoritative documents.
The central question reaches far beyond one attorney:
What happens when we trust an AI-generated summary more than we verify the evidence behind it?
What Happened in the New Mexico ChatGPT Murder Appeal?
The case arose from the appeal of Oscar Renee Sandoval’s murder conviction. Attorney Stephen D. Aarons represented Sandoval in the appeal before the New Mexico Supreme Court. During preparation of the appellate brief, Aarons used ChatGPT to assist with summarizing trial materials.
The resulting brief contained 4 fabricated witnesses not present in the underlying record.
According to the New Mexico Supreme Court’s September 9 order, Aarons admitted that the filing included false testimony from four wholly fabricated witnesses:
- Officer Michelle Amarillo,
- Officer Sanchez,
- Manal Al-Jibury,
- Teresa Marquez.
Furthermore, the court identified false testimony attributed to actual individuals, including statements concerning alleged threats and descriptions of the shooter’s clothing and appearance.
The filing also misrepresented legal authority. Therefore, the AI failure extended beyond an inaccurate sentence or minor summarization error. The brief introduced unsupported factual claims into a criminal appeal.
For a proceeding involving a murder conviction and a life sentence, factual accuracy carried extraordinary importance.
How Did ChatGPT Fabricate Witness Testimony From Murder Trial Records?
Aarons acknowledged using ChatGPT during preparation of the brief. According to Reuters, he supplied a computer-generated transcript and other case materials to the AI system and expected it to create a reliable summary of the trial.
However, generative AI does not guarantee that every statement in a summary originated in the documents provided to it. Large language models generate language based on patterns and probabilities. As a result, they can produce plausible statements that sound consistent with source material even when those statements lack evidentiary support.
In this case, the danger became concrete.
- The system generated supposed witness testimony. Some of those witnesses did not exist.
- Because the attorney did not adequately trace the generated claims back to the original trial record, the fabricated material survived the review process and entered the appellate brief.
- The trusted documents were real. The generated summary contained fiction.
- The human reviewer failed to identify the difference before filing.
What Did the New Mexico Supreme Court Do About the AI-Generated False Testimony?
The New Mexico Supreme Court found Stephen Aarons in direct contempt of court. It also:
- Fined him $5,000.
- Referred him for disciplinary review.
- Barred him from appearing before the court while the review remains pending.
- Struck the faulty briefs and appointed a public defender to continue Sandoval’s appeal.
The appeal itself remains pending.
Therefore, the consequences extended considerably beyond a monetary sanction. The failure affected the attorney, the court proceedings, the appellate record, and the representation of the defendant.
AI Hallucinated Evidence Is More Serious Than a Fake Legal Citation
Generative AI errors in legal practice have already produced nonexistent cases, fictitious citations, inaccurate quotations, and misstated precedent.
AI Hallucinated evidence failures can cause serious harm.
However, fabricated testimony creates an additional level of concern because testimony relates directly to the factual record of a case.
- AI invents a legal citation, the system creates nonexistent authority.
- AI invents witness testimony, the system can create a false version of what supposedly happened.
- That distinction matters enormously in litigation.
The American Bar Association has warned lawyers about generative AI risks in litigation workflows and has emphasized that attorneys remain responsible for the work they submit, even when AI helps create it.
The ABA has also discussed the lawyer’s duty of candor when using generative AI, reinforcing the need to verify authorities, factual representations, and consequential material before filing.
The New Mexico case now provides a vivid example of why those controls matter for factual evidence as well as legal citations.
AI Hallucination Risk: Why Trusted Documents Do Not Guarantee a Trustworthy AI Summary
One of the most important lessons from this case involves a common assumption about generative AI.
Providing trusted documents does not guarantee a factual AI summary. Assuming it does, as this story illustrates has created serious risk and impact.
A reliable source does not automatically produce a reliable AI-generated summary.
Generative AI can merge concepts, infer unsupported details, confuse speakers, misattribute statements, introduce outside information, or generate plausible facts that never appeared in the source.
Therefore, professionals who work with legal records, medical information, audits, investigations, compliance evidence, financial documents, or regulatory materials need to verify the transformation from source document to generated output.
The original evidence may remain perfectly accurate while the AI-generated interpretation becomes unreliable.
That gap creates one of the most important AI governance risks facing organizations today.
KEY AI Lesson in AI Literacy: Verify Important Documents Before You Trust the AI Summary
Generative AI can help professionals organize, summarize, compare, explain, and draft information quickly. Those capabilities can create enormous productivity gains.
However, AI literacy requires recognizing that fluent output does not equal verified output.
AI literacy means understanding both what AI can do and what you must still do yourself.
A professional reviewer should approach consequential AI-generated material with one fundamental question:
Where did this claim come from?
If the answer cannot be traced to credible evidence, the claim should not move forward merely because the AI expressed it confidently.
Verification becomes especially important when an AI-generated statement could influence a legal filing, patient decision, audit finding, compliance determination, investigation, financial conclusion, regulatory disclosure, or executive decision.
In those settings, trust must follow verification.
How Do You Verify AI-Generated Summaries and Important Documentation?
A strong verification process begins with the original evidence rather than the AI response.
First, identify every material factual claim in the generated output. A material claim includes any statement that could influence the conclusion, recommendation, decision, finding, or legal argument.
Next, locate the original source supporting that statement.
For example, if an AI summary says a witness testified that a suspect wore a white shirt, the reviewer should locate that testimony in the transcript. If the statement does not appear there, the AI-generated claim cannot serve as evidence.
Then, verify quotations word for word against the source. Names, dates, dollar amounts, diagnoses, legal citations, numerical findings, and other high-consequence facts deserve the same treatment.
Additionally, reviewers should distinguish among three different categories of AI output: facts extracted directly from the source, conclusions inferred from those facts, and language generated by the model.
Those categories can look remarkably similar on the page, although they carry very different evidentiary value.
The ABA’s guidance on using AI wisely in litigation supports this broader principle: professionals need checkpoints that force verification before AI-assisted work becomes final work.
Finally, the person approving the document should understand that accountability transfers with approval. Once someone signs, submits, publishes, or acts on AI-generated work, the practical consequences no longer belong solely to the software.
The AI Verification Test: Can You Trace Every Important Claim to Evidence?
Organizations can make AI literacy practical by adopting a simple verification standard:

What Happens When Professionals Forget to Verify AI-Generated Information?
The New Mexico case demonstrates how quickly an AI error can become a real-world consequence.
- A fabricated statement inside a chatbot remains an AI output.
- A fabricated statement copied into an official document becomes something else.
- It can become a court representation, audit finding, medical summary, investigative conclusion, compliance determination, regulatory statement, financial recommendation, or executive briefing.
- At that stage, the hallucination begins impacting people and institutions.
The possible consequences can include professional sanctions, disciplinary investigations, damaged credibility, incorrect decisions, rework, delayed proceedings, financial loss, legal exposure, reputational harm, and harm to the people whose lives depend on the accuracy of the information.
In the Sandoval appeal, the New Mexico Supreme Court imposed a $5,000 sanction, held the attorney in direct contempt, referred the matter for disciplinary consideration, struck the briefing, and appointed new counsel.
Those are human consequences arising from an AI-generated failure that passed through human review.
Human in the Loop AI Governance Requires Critical Thinking
“Human in the Loop” has become a familiar phrase in responsible AI and AI governance.
However, organizations should examine what that human actually does.
A reviewer adds value when that person questions unexpected conclusions, validates material facts, follows citations, checks original evidence, challenges unsupported statements, and refuses to approve information that cannot be substantiated.
Critical thinking transforms human participation into human oversight.
Without those behaviors, an approval step can provide the appearance of governance without delivering the protection organizations expect from it.
The New Mexico case illustrates that distinction clearly. A human participated in the workflow. Yet fabricated information reached the court.

Therefore, organizations need to define human review by the actions a reviewer must perform rather than by the simple presence of a person somewhere in the process.
Human Trust Without AI Validation is Legal, Moral, Governance or Enterprise Risk
The legal profession provides an unusually visible example because court filings create public records and judges can impose sanctions.
However, the same AI failure pattern can occur quietly throughout an organization.
- Imagine an AI system summarizing hundreds of pages from an internal investigation. A manager reads the concise summary and relies on it when deciding whether to terminate an employee.
- Consider an AI system reviewing audit evidence and generating an executive summary that introduces a fact the auditors never established.
- Consider a healthcare workflow that condenses clinical documentation but attributes a condition or medication to the wrong portion of the record.
Alternatively, imagine an AI-generated compliance report that states a control was tested when the underlying documentation never supports that conclusion.

AI Controls for High-Stakes Documents
Organizations should build consequential AI workflows around evidence traceability.
- Ground every claim. Limit AI to approved sources, but never treat grounding as verification.
- Link the evidence. Add claim-level citations so reviewers can inspect the exact supporting source.
- Label the content. Identify material as directly quoted, source-summarized, or AI-generated analysis.
- Require real verification. An approval click does not prove that anyone checked the facts.
- Assign accountability. Final approvers must know what was verified, what remains uncertain, and what responsibility they accept.
AI can accelerate documentation. Humans remain accountable for its accuracy.
AI Literacy Requires Healthy Skepticism
Most workplace AI training still concentrates heavily on productivity.
Employees learn how to prompt more effectively, summarize documents faster, draft emails, generate reports, analyze information, and automate repetitive work.
Those capabilities matter.
Responsible AI literacy requires verification literacy.
AI can generate false information with complete confidence.
Challenge the output. Trace the source. Verify the facts.
People need to recognize hallucinations, distinguish evidence from generated language, locate original sources, test claims against those sources, identify missing context, assess uncertainty, and recognize when the stakes require deeper review.
- Furthermore, employees need permission to slow down when verification matters.
- Productivity gains disappear quickly when inaccurate AI-generated work creates sanctions, investigations, rework, reputational damage, or harmful decisions.
- Speed creates value only when accuracy travels with it.
Real AI Governance Formula: Human in the Loop + Critical Thinking + Verifiable Evidence
The New Mexico Supreme Court case offers a simple but powerful AI governance lesson.
Human in the Loop + Critical Thinking + Verifiable Evidence
- The human brings professional judgment.
- Critical thinking challenges the generated output.
- Verifiable evidence establishes whether the claim deserves to be trusted.
- Together, those elements create meaningful oversight.
Organizations that remove one of them create avoidable vulnerability.
- A human without critical thinking can become an approval mechanism.
- Critical thinking without access to original evidence can become educated guessing.
- Evidence without accountability can remain unchecked.
Effective AI governance connects all three.
What Should Lawyers and Leaders Learn From the ChatGPT Fabricated Evidence Case?
The broader lesson reaches far beyond ChatGPT, one attorney, or one murder appeal.
Generative AI will continue entering workflows that depend on accurate records.
Organizations will use it to summarize thousands of pages faster than humans can read them, identify relevant information across large document collections, prepare briefings, draft reports, and support increasingly complex decisions.
- Those capabilities make verification more important, not less.
- AI can accelerate the movement of information through an organization.
- However an unsupported claim can also travel farther and faster with greater damages than it could before.
The goal should remain straightforward:
Use AI to accelerate understanding without allowing it to manufacture evidence.
The New Mexico case shows what can happen when that boundary disappears.
ChatGPT generated information unsupported by the murder-trial record. The attorney failed to catch the fabricated testimony before filing the brief. The New Mexico Supreme Court discovered the problem and imposed real consequences.

Five Legal AI Rules
- Protect the record. Use approved tools and never expose privileged information.
- Trace every claim. Connect each fact to the exact source, page, and line.
- Open every case. Confirm the citation, quotation, holding, and current status.
- Challenge confident answers. Fluent language can still contain fabricated information.
- Document human review. Record who verified the work, how they checked it, and which sources they used.
Other Legal AI Resources
Use these credible guides to strengthen legal AI literacy, verification, confidentiality, and responsible use:
- Tech Law Boot Camp 2026: AI Chatbots - Negligence & the Mental Health Harms | LinkedIn
- LexisNexis:
- Protect Legal Citation Integrity: Learn why AI invents cases, quotations, and facts. Always open every linked authority and confirm the language, context, and holding.
- Responsible Legal AI Tips Review accuracy, confidentiality, professional responsibility, and workflow controls before introducing AI into legal work.
- Choosing Legal AI : Evaluate whether a legal AI tool uses authoritative sources, provides verifiable citations, protects client data, and supports human review.
- American Bar Association: Responsible AI Implementation: Protect privilege, understand where client data goes, use approved platforms, and maintain attorney oversight.
- UIC School of Law: AI Best Practices: Explore responsible AI research, ethical use, source validation, and practical guidance for lawyers and legal students.
- Loyola University Chicago: Generative AI in Legal Practice: Examine how lawyers can use generative AI while addressing hallucinations, confidentiality, bias, and professional duties.
- The Law Society: Generative AI Essentials Understand AI capabilities, technology risks, data protection, accountability, and responsible adoption.
- Berkeley Study: Women Lawyers and AI Access: See how training and access helped close the gender gap in legal AI use while reinforcing the need to verify output and protect confidential information.
- Dawn C. Simmons: Follow practical guidance on AI governance, human-centered automation, service management, and accountable enterprise AI.
- WomenAILabs:
- DAUGHTER Equity Analyzer TM : Evidence-based DAUGHTER reviews with user perspectives, remediation, and retesting.
- Global Student Talent Advantage Tool: WomenAILabs™ career strategist helping international students build proof, timing, and distinction.
- Global Seat at the Table Tool: Benchmarking and action platform for women’s leadership and decision authority. It helps companies, governments, universities, and institutions measure where women are represented, where they hold real power, how performance compares with peers, and what actions will accelerate progress.


