Navigating AI’s Twin Perils: The Rise of the Risk-Mitigation Officer

July 28, 2025

Ralph Losey, July 28, 2025

Generative AI is not just disrupting industries—it is redefining what it means to trust, govern, and be accountable in the digital age. At the forefront of this evolution stands a new, critical line of employment: AI Risk-Mitigation Officers. This position demands a sophisticated blend of technical expertise, regulatory acumen, ethical judgment, and organizational leadership. Driven by the EU’s stringent AI Act and a rapidly expanding landscape of U.S. state and federal compliance frameworks, businesses now face an urgent imperative: manage AI risks proactively or confront severe legal, reputational, and operational consequences.

Click to see photo come to life. Image and movie by Losey with AI.

This aligns with a growing consensus: AI, like earlier waves of innovation, will create more jobs than it eliminates. The AI Risk-Mitigation Officer stands as Exhibit A in this next wave of tech-era employment. See e.g. my last series of blogs, Part Two: Demonstration by analysis of an article predicting new jobs created by AI (27 new job predictions); Part Three: Demo of 4o as Panel Driver on New Jobs (more experts discuss the 27 new jobs). See Also: Robert Capps, NYT magazine article: A.I. Might Take Your Job. Here Are 22 New Ones It Could Give You.  (NYT Magazine, June 17, 2025). In a few key areas, humans will be more essential than ever.

Risk Mitigation Officer team image by Losey and AI.

Defining the Role of the AI Risk-Mitigation Officer

The AI Risk-Mitigation Officer is a strategic, cross-functional leader tasked with identifying, assessing, and mitigating risks inherent in AI deployment.

While Chief AI Officers drive innovation and adoption, Risk-Mitigation Officers focus on safety, accountability, and compliance. Their mandate is not to slow progress, but to ensure it proceeds responsibly. In this respect, they are akin to data protection officers or aviation safety engineers, guardians of trust in high-stakes systems.

This role requires a sober analysis of what can go wrong—balanced against what might go wonderfully right. It is a job of risk mitigation, not elimination. Not every error can or should be prevented; some mistakes are tolerable and even expected in pursuit of meaningful progress.

The key is to reduce high-severity risks to acceptable levels—especially those that could lead to catastrophic harm or irreparable public distrust. If unmanaged, such failures can derail entire programs, damage lives, and trigger heavy regulatory backlash.

Both Chief AI Officers and Risk-Mitigation Officers ultimately share the same goal: the responsible acceleration of AI, including emerging domains like AI-powered robotics.

The Risk-Mitigation Office should lead internal education efforts to instill this shared vision—demonstrating that smart governance isn’t an obstacle to innovation, but its most reliable engine.

Team leader tries to lighten the mood of their serious work. Click for video by Losey.

Why The Role is Growing

The acceleration of this role is not theoretical. It is propelled by real-world failures, regulatory heat, and reputational landmines.

The 2025 World Economic Forum’s Future of Jobs Report underscores that 86% of surveyed businesses anticipate AI will fundamentally transform their operations by 2030. While AI promises substantial efficiency and innovation, it also introduces profound risks, including algorithmic discrimination, misinformation, automation failures, and significant data breaches.

A notable illustration of these risks is the now infamous Mata v. Avianca case, where lawyers relied on AI to fabricate case law, underscoring why human verification is non-negotiable. Mata v. Avianca, Inc., No. 1:2022cv01461 – Document 54 (S.D.N.Y. June 22, 2023). Courts responded with sanctions. Regulators took notice. The public laughed, then worried.

The legal profession worldwide has been slow to learn from Mata the need to verify and take other steps to control AI hallucinations. See Damien Charlotin, AI Hallucination Cases (as of July 2025 over 200 cases and counting have been identified by the legal scholar in Paris with an ironic name for this work, Charlotin). The need for risk mitigation is growing fast. You cannot simply pass the buck to AI.

Charlatan lawyers blame others for their mistakes. Image by Losey.

Risk Mitigation employees and other new AI related hires will balance out the AI generated layoffs. The 2025 World Economic Forum’s Future of Jobs Report at page 25 predicted that by 2030 there will be 11 million new jobs created and 9 million old jobs phased out. We think the numbers will be higher in both columns, but the 11/9 ratio may be right overall all, meaning a 22% net increase in new jobs.

We think the ratio the WEF predicts is right for all industries overall, but the numbers are too low, meaning greater disruption but positive in the end with even more new jobs. Image by AI & Losey.

Core Responsibilities

AI Risk-Mitigation Officers are part legal scholar, part engineer, part diplomat. They know the AI Act, understand neural nets, and can hold a room full of regulators or engineers without flinching. Key responsibilities encompass:

  • AI Risk Audits: Spot trouble before it starts. Bias, black boxes, security flaws—find them early, fix them fast. This involves detailed pre-deployment evaluations to detect biases, security vulnerabilities, explainability concerns, and data protection deficiencies. Good practice to follow-up with periodic surprise audits.
  • Incident Response Management: Don’t wait for a headline to draft a playbook. Develop and lead response protocols for AI malfunctions or ethical violations, coordinating closely with legal, PR, and compliance teams.
  • Legal Partnership: Speak both code and contract. Collaborate with in-house counsel to interpret AI regulations, draft protective contractual clauses, and anticipate potential liability.
  • Ethics Training: Culture is the best control layer. Educating employees on responsible AI use and cultivating an ethical culture that aligns with both corporate values and regulatory standards.
  • Stakeholder Engagement: Transparency builds trust. Silence breeds suspicion. Bridging communication between technical teams, executive leadership, regulators, and the public to maintain transparency and foster trust.
One Core Responsibility of the Rick Mitigation Team. Image by Losey.

Skills and Pathways

Professionals in this role must possess:

  • Regulatory Expertise: Detailed knowledge of EU AI Act, GDPR, EEOC guidelines, and evolving state laws in the U.S.
  • Technical Proficiency: Deep understanding of machine learning, neural networks, and explainable AI methodologies.
  • Sector-Specific Knowledge: Familiarity with compliance standards across sectors such as healthcare (HIPAA, FDA), finance (SEC, MiFID II), and education (FERPA).
  • Strategic Communication: Ability to effectively mediate between AI engineers, executives, regulators, and stakeholders.
  • Ethical Judgment: Skills to navigate nuanced ethical challenges, such as balancing privacy with personalization or fairness with automation efficiency.

Career pathways for AI Risk-Mitigation Officers typically involve dual qualifications in fields like law and data science, certifications from professional organizations and practical experience in areas like cybersecurity, legal practice, politics or IT. Strong legal and human relationship skills are a high priority.

Image in modified isometric style by Ralph Losey using modified AIs.

U.S. and EU Regulatory Landscapes

The EU codifies risk tiers. The U.S. litigates after the fact. Navigating both requires fluency in law, logic, and diplomacy.

The EU’s AI Act classifies AI systems into four risk categories:

  • Unacceptable. AI banned altogether due to the high-risk of violating fundamental rights. Examples include social scoring systems and real-time biometric identification in public spaces, including emotion recognition in workplaces and schools. Also bans AI-enabled subliminal or manipulative techniques that can be used to persuade individuals to engage in unwanted behaviors.
  • High. AI that could negatively impact the rights or safety of individuals. Examples include AI systems used in critical infrastructures (e.g. transport), legal (including policing and border patrol), medical, educational, financial services, workplace management, and influencing elections and voter behavior.
  • Limited. AI with lower levels of risk than high-risk systems but are still subject to transparency requirements. Examples include typical chatbots where providers must make humans aware Ai is used.
  • Minimal. AI systems that present little risk of harm to the rights of individuals. Examples include AI-powered video games and spam filters. No rules.
Transparent AI in office setting. Click for Losey movie.

Regulation is on a sliding scale with Unacceptable risk AIs banned entirely and Minimal to No risk category with few if any restrictions. Limited and High risk classes require varying levels of mandatory documentation, human oversight, and external audits.

Meanwhile, U.S. regulatory bodies like the FTC and EEOC, along with state legislatures and state enforcement agencies, are starting to sharpen oversight tools. So far the focus has been on controlling deception, data misuse, bias and consumer harm. This has become a hot political issue in the U.S. See e.g. Scott Kohler, State AI Regulation Survived a Federal Ban. What Comes Next? (Carnegie’s Emissary, 7/3/25); Brownstein, et al, States Can Continue Regulating AI—For Now (JD Supra, 7/7/25).

AI Risk-Mitigation Officers must navigate these disparate regulatory landscapes, harmonizing proactive European requirements with the reactive, litigation-centric U.S. environment.

Legal Precedents and Ethical Challenges

Emerging legal precedents emphasize human accountability in AI-driven decisions, as evidenced by litigation involving biased hiring algorithms, discriminatory credit scoring, and flawed facial recognition technologies. Ethical dilemmas also abound. Decisions like prioritizing efficiency over empathy in healthcare, or algorithmic opacity in university admissions, require human-centric governance frameworks.

In ancient times, the Sin Eater bore others’ wrongs. Part Two: Demonstration by analysis of an article predicting new jobs created by AI (discusses new sin eater job in detail as a person who assumes moral and legal responsibility for AI outcomes). Today’s Risk-Mitigation Officer is charged with an even more difficult task; try to prevent the sins from happening at all or at least reduce them enough to avoid hades.

Sin Eater in combined quasi digital styles by Losey using sin-free AIs.

Balancing Innovation and Regulation

Cases such as Cambridge Analytica (personal Facebook user data used for propaganda to influence elections) and Boeing’s MCAS software (use of AI system undisclosed to pilots led to crash of a 737) demonstrate that innovation without reasonable governance can be an invitations to disaster. The obvious abuses and errors in these cases could have been prevented had there been an objective officer in the room, really any responsible adult considering risks and moral grounds. For a more recent example, consider XAI’s recent fiasco with its chatbot. Grok Is Spewing Antisemitic Garbage on X (Wired, 7/8/25).

Cases like this should put us on guard but not cause over-reaction. Disasters like this easily trigger too much caution and too little courage. That too would be a disaster of a different kind as it would rob us of much needed innovation and change.

Reasonable, practical regulation can foster innovation by mitigating uncertainty and promoting stakeholder confidence. The trick is to find a proper balance between risk and reward. Many think regulators today tend to go too far in risk avoidance. They rely on Luddite fears of job loss and end-of-the world fantasies to justify extreme regulations. Many think that such extreme risk avoidance helps those in power to maintain the status quo. The pro-tech people instead favor a return to the fast pace of change that we have seen in the past.

Polarized protestors for and against technology by Losey using AI job thieves.

We went to the Moon in 1969. Yet we’re still grounded in 2025. Fear has replaced vision. Overregulation has become the new gravity.

That is why anti-Luddite technology advocates yearn for the good old days, the sixties, where fast paced advances and adventure were welcome, even expected. If you had told anyone back in 1969 that no Man would return to the Moon again, even as far out as 2025, you would have been considered crazy. Why was John F. Kennedy the last bold and courageous leader? Everyone since seems to have been paralyzed with fear. None have pushed great scientific advances. Instead, politicians on both sides of the aisle strangle NASA with never-ending budget cuts and timid machine-only missions. It seems to many that society has been overly risk adverse since JFK and this has stifled innovation. This has robbed the world of groundbreaking innovations and the great rewards that only bold innovations like AI can bring.

Take a moment to remember the 1989 movie Back to the Future Part II. In the movie “Doc” Brown, an eccentric, genius scientist, went from the past – 1985 – the the future of 2015. There we saw flying cars powered by Mr. Fusion Home Energy Reactors, which were trash fueled fusion engines. That is the kind of change we all still expected back in 1989 for the far future of 2015; but now, in 2025, ten years past the projected future, it is still just a crazy dream. The movie did, however, get many minor predictions right. Think of flat-screen tvs, video-conferences, hoverboards and self-tying shoes (Nike HyperAdapt 1.0.) Fairly trivial stuff that tech go-slow Luddites approved.

Back to the Future movie made in 1989 envisioning flying cars in 2015. Now in 2025, what have we got? Photo and VIDEO by Losey using AI image generation tools.

Conclusion

The best AI Risk-Mitigation Officers will steer between the twin monsters of under-regulation and overreach. Like Odysseus, they will survive by knowing the risks and keeping a steady course between them.

They will play critical roles across society—in law firms, courts, hospitals, companies (for-profit, non-profit, and hybrid), universities, and government agencies. Their core responsibilities will include:

  • Standardization Initiatives: Collaborating with global standards organizations such as ISO, NIST, and IEEE to craft reasonable, adaptable regulations.
  • Development of AI Governance Tools: Encouraging the use of model cards, bias detection systems, and transparency dashboards to track and manage algorithmic behavior.
  • Policy Engagement and Lobbying: Actively engaging with legislative and regulatory bodies across jurisdictions—federal, state, and international—to advocate for frameworks that balance innovation with public protection.
  • Continuous Learning: Staying ahead of rapid developments through ongoing education, credentialing, and immersion in evolving legal and technological landscapes.
AI Risk Management specialists will end up with sub-specialties, maybe dept. sections, such as lobbying and research. Image by Losey daring to use AI.

As this role evolves, AI Risk Management specialists will likely develop sub-specialties—possibly forming distinct departments in areas such as regulatory lobbying, algorithm auditing, compliance architecture, and ethical AI research, both hands-one and legal study.

This is a fast-moving field. After writing this article we noticed the EU passed new rules for the largest AI companies only, all US companies of course, except one Chinese. It is non-binding at this point and involves highly controversial disclosure and copyright restriction. EU’s AI code of practice for companies to focus on copyright, safety (Reuters, July 10, 2025). It could stifle chatbot development and lead the EU in a stagnant deadline as these companies withdraw from the EU rather than kill their own models to comply.

Slowdown or reduction is not a viable option at this point because of national security concerns. There is a military race now between the US and China based on competing technology. AGI level of AI will give the first government to attain it a tremendous military advantage. See e.g., Buchanan, Imbrie, The New Fire: War, Peace and Democracy in the Age of AI (MIT Press, 2022); Henry Kissinger, Allison. The Path to AI Arms Control: America and China Must Work Together to Avert Catastrophe, (Foreign Affairs, 10/13/23); Also See, Dario Amodei’s Vision (e-Discovery team, Nov. 2024) (CEO of Anthropic, Darion Amodei, warns of danger of China winning the race for AI supremacy).

Race to super-intelligent AGI image by Ralph Losey.

As Eric Schmidt explains it, this is now an existential threat and should be a bipartisan issue for survival of democracy. Kissinger, Schmidt, Mundie, Genesis: Artificial Intelligence, Hope, and the Human Spirit, (Little Brown, Nov. 2024). Also See Former Google CEO Eric Schmidt says America could lose the AI race to China (AI Ascension, May 2025).

Organizations that embrace this new professional archetype will be best positioned to shape emerging regulatory frameworks and deploy powerful, trusted AI systems—including future AI-powered robots. The AI Risk-Mitigation Officer will safeguard against catastrophic failure without throttling progress. In this way, they will help us avoid both dystopia and stagnation.

Yes, this is a demanding job. It will require new hires, team coordination, cross-functional fluency, and seamless collaboration with AI assistants. But failure to establish this critical function risks danger on both fronts: unchecked harm on one side, and paralyzing caution on the other. The best Risk-Mitigation Officers will navigate between these extremes—like Odysseus threading his ship through Scylla and Charybdis.

Odysseus successfully steering his ship between monsters on either side. Image and Video by Losey using various AIs.

We humans are a resilient species. We’ve always adapted, endured, and risen above grave dangers. We are adventurers and inventors—not cowering sheep afraid of the wolves among us.

The recent overregulation of science and technology is an aberration. It must end. We must reclaim the human spirit where bold exploration prevails, and Odysseus—not  Phobos—remains our model.

We can handle the little thinking machines we’ve built, even if the phobic establishment wasn’t looking. Our destiny is fusion, flying cars, miracle cures, and voyages to the Moon, Mars, and beyond.

Innovation is not the enemy of safety—it is its partner. With the right stewards, AI can carry us forward, not backward.

Let’s chart the course.

Manifest Destiny of Mankind. Image and VIDEO by Losey using AI.

PODCAST

As usual, we give the last words to the Gemini AI podcasters who chat between themselves about the article. It is part of our hybrid multimodal approach. They can be pretty funny at times and have some good insights, so you should find it worth your time to listen. Echoes of AI: The Rise of the AI Risk-Mitigation Officer: Trust, Law, and Liability in the Age of AI. Hear two fake podcasters talk about this article for a little over 16 minutes. They wrote the podcast, not me. For the second time we also offer a Spanish version here. (Now accepting requests for languages other than English.)

Click here to listen to the English version of the podcast.

Ralph Losey Copyright 2025 – All Rights Reserved.


Bar Battle of the Bots – Part One

February 26, 2025

Ralph Losey. February 26, 2025.

The legal world is watching AI with both excitement and skepticism. Can today’s most advanced reasoning models think like a lawyer? Can they dissect complex fact patterns, apply legal principles, and construct a persuasive argument under pressure—just like a law student facing the Bar Exam? To find out, I put six of the most powerful AI reasoning models from OpenAI and Google to the test with a real Bar Exam essay question, a tricky one. Their responses varied widely—from sharp legal analysis to surprising omissions, and even a touch of hallucination. Who passed? Who failed? And what does this mean for the future of AI in the legal profession? Read Part One of this two-part article to find out.

Introduction

This article shares my test of the legal reasoning abilities of the newest and most advanced reasoning models of OpenAI and Google. I used a tough essay question from a real Bar Exam given in 2024. The question involves a hypothetical fact pattern for testing legal reasoning on Contracts, Torts and Ethics. For a full explanation of the difference between legal reasoning and general reasoning, see my last article, Breaking New Ground: Evaluating the Top AI Reasoning Models of 2025.

I picked a Bar Exam question because it is a great benchmark of legal reasoning and came with a model answer from the State Bar Examiner that I could use for objective evaluation. Note, to protect copyright and the integrity of the Bar Exam process, I will not link to the Bar model answer, except to say it was too recent to be in generative AI training. Moreover, some aspects of the tests answers that I quote in this article have been modified somewhat for the same reason. I will provide links to the original online Bar Exam essay to any interested researchers seeking to duplicate my experiment. I hope some of you will take me up on that invitation.

Prior Art: the 2023 Katz/Casetext Experiment on ChatGPT-4.0

A Bar Exam has been used before to test the abilities of generative AI. OpenAI and the news media claimed that ChatGPT-4.0 had attained human lawyer level legal reasoning ability. GPT-4 (OpenAI, 3/14/23) (“it passes a simulated bar exam with a score around the top 10% of test takers; in contrast, GPT‑3.5’s score was around the bottom 10%). The claims of success were based on a single study by a respected Law Professor, Daniel Martin Katz, of Chicago-Kent, and a leading legal AI vendor Casetext. Katz, et. al., GPT-4 Passes the Bar Exam, 382 Philosophical Transactions of the Royal Society A (March, 2023, original publication date) (fn. 3 found at pg. 10 of 35: “… GPT-4 would receive a combined score approaching the 90th percentile of test-takers.”) Note, Casetext used the early version of ChatGPT-4.0 in its products.

The headlines in 2023 were that ChatGPT-4.0 had not only passed a standard Bar Exam but scored in the top ten percent. OpenAI claimed that ChatGPT-4.0 had already attained elite legal reasoning abilities of the best human lawyers. For proof OpenAI and others cited the experiment of Professor Katz and Casetext that it aced the Bar Exam. See e.g., Latest version of ChatGPT aces bar exam with score nearing 90th percentile (ABA Journal, 3/16/23). Thomson Reuters must have checked the results carefully because they purchased Casetext in August 2023 for $650,000,000. Some think they may have overpaid.

Challenges to the Katz/Casetext Research and OpenAI Claims

The media reports on the Katz/Casetext study back in 2023 may have grossly inflated the AI capacities of ChatGPT-4.0 that Casetext built its software around. This is especially true for the essay portion of the standardized multi-state Bar Exam. The validity of this single experiment and conclusion that ChatGPT-4.0 ranked in the top ten percent has since been questioned by many. The most prominent skeptic is Eric Martinez as detailed in his article, Re-evaluating GPT-4’s bar exam performance. Artif Intell Law (2024) (presenting four sets of findings that indicate that OpenAI’s estimates of GPT-4.0’s Uniform Bar Exam percentile are overinflated). Specifically, the Martinez study found that:

3.2.2 Performance against qualified attorneys
Predictably, when limiting the sample to those who passed the bar, the models’ percentile dropped further. With regard to the aggregate UBE score, GPT-4 scored in the ~45th percentile. With regard to MBE, GPT-4 scored in the ~69th percentile, whereas for the MEE + MPT, GPT-4 scored in the ~15th percentile.

Id. The ~15th percentile means GPT-4 scored approximately (~) in the bottom 15%, not the top 10%!

More to the point of my own experiment and conclusions, the Martinez study goes on to observe:

Moreover, when just looking at the essays, which more closely resemble the tasks of practicing lawyers and thus more plausibly reflect lawyerly competence, GPT-4’s performance falls in the bottom ~15th percentile. These findings align with recent research work finding that GPT-4 performed below-average on law school exams (Blair-Stanek et al. 2023).

The article by Eric Martinez makes many valid points. Martinez is an expert in law and AI. He started with a J.D. from Harvard Law School, then earned a Ph.D. in Cognitive Science from MIT, and is now a Legal Instructor at the University of Chicago, School of Law. Eric specializes in AI and the cognitive foundations of the law. I hope we hear a lot more from him in the future.

Details of the Katz/Casetext Research

I dug into the details of the Katz/Casetext experiment to prepare this article. GPT-4 Passes the Bar Exam. One thing I noticed not discussed by Eric Martinez is that the Katz experiment modified the Bar Exam essay questions and procedures somewhat to make it easier for the 2023 ChatGPT-4 model to understand and respond correctly. Id. at pg. 7 of 35. For example, they divided the Bar model essay question into multiple parts. I did not do that to simplify the three-part 2024 Bar essay I used. I copied the question exactly and otherwise made no changes. Moreover, I did not experiment with various prompts of the AI to try to improve its results, as Katz/Casetext did. Also, I did no training of the 2025 reasoning models to make them better at taking Bar exam questions. The Katz/Casetext group shares the final prompt used, which can be found here. But I could not find in their disclosed experiment data a report of the prompt changes made, or whether there was any pre-training on case law, or whether Casetext’s case extensive case law collections and research abilities were in any way used or included. The models I tested were clean and not web connected, nor were they designed for research.

The Katz/Casetext experiments on Bar essay exams were, however, much more extensive than mine, covering six questions and using several attorneys for grading. (The use of multiple human evaluators can be both good and bad. We know from e-discovery experiments with multiple attorney reviewers that this practice leads to inconsistent determinations of relevance unless very carefully coordinated and quality controlled.) The Katz/Casetext results on the 2023 ChatGPT-4.0 are summarized in this chart.

As shown in Table 5 of the Katz report, they used a six-point scale, which they indicate is commonly followed by many state examiners. GPT-4 Passes the Bar Exam, supra at page 9 of 35. Katz claims “a score of four or higher is generally considered passing” by most state Bar examiners.

The Katz/Casetext study did not use the better known four-point evaluation scale – A to F – that is followed by most law schools. In law school (where I have five years’ experience grading essay answers in my Adjunct Professor days), an “A” is four points, a “B” is three, a “C” is two, a “D” is one and “E” or “F” is zero. Most schools in the country use that system too. In law school a “C” – 2.0- is passing. A “D” or lower grade is failure in any professional graduate program, including law schools where, if you graduate, you earn a Juris Doctorate. [In the interest of full disclosure, I may well be an easy grader, because, with the exception of a few “no-shows,” I never awarded a grade lower than a “C” in my life. Of course, I was teaching electronic discovery and evidence at an elite law school. On the other hand, many law firm associates over the years have found that I am not at all shy about critical evaluations of their legal work product. The rod certainly was not spared on me when I was in their position, in fact, it was swung much harder and more often in the old days. In the long run constructive criticism is indispensable.]

The Katz/Casetext study using a 0-6.0 Grading system scored by lawyers gave evaluations ranging from 3.5 for Civil Procedure to 5.0 for Evidence, with an average score of 4.2. Translated into the 4.0 system that most everyone is familiar with, this means a score range of from 2.3 (a solid “C”) for Civ-Pro to 3.33 (a solid “B”) for Evidence, and a average score of 2.8 (a C+). Note the test I gave to my 2025 AIs covered three topics in one, Contract, Torts and Ethics. The 2023 models were not given a Torts or Ethics question, but for the Contract essay their score translated to a 4.0 scale of 2.93, a strong C+ or B-. Note one of the criticisms of Martinez concerns the haphazard, apparently easy grading of AI essays. Re-evaluating GPT-4’s bar exam performance, supra at 4.3 Re-examining the essay scores.

First Test of the New 2025 Reasoning Models of AI

To my knowledge no one has previously tested the legal reasoning abilities of the new 2025 reasoning models. Certainly, no one has tested their legal reasoning by use of actual Bar Exam essay questions. That is why I wanted to take the time for this research now. My goal was not to reexamine the original ChatGPT 4.0, March 2023, law exam tests. Eric Martinez has already done that. Plus, right or wrong, I think the Katz/Casetext research did the profession a service by pointing out that AI can probably pass the Bar Exam, even if just barely.

My only interest in February 2025 is to test the capacities of today’s latest reasoning models of generative AI. Since everyone agrees the latest reasoning models of AI are far better than the first 2023 versions, if the 2025 models did not pass an essay exam, even a multi-part tricky one like I picked, then “Houston we have a problem.” The legal profession would now be in serious danger of relying too much on AI legal reasoning and we should all put on the brakes.

Description of the Three Legal Reasoning Tests

The test involved a classic format of detailed, somewhat convoluted facts-the hypothetical-followed by three general questions:

1. Discuss the merits of a breach of contract claim against Helen, including whether Leda can bring the claim herself.  Your discussion should address defenses that Helen may raise and available remedies.  

2. Discuss the merits of a tortious interference claim against Timandra.

3. Discuss any ethical issues raised by Lawyer’s and the assistant’s conduct.

The only instructions provided by the Bar Examiners were:

ESSAY EXAMINATION INSTRUCTIONS

Applicable Law:

  • Answer questions on the (state name omitted here) Bar Examination with the applicable law in force at the time of examination. 

Questions are designed to test your knowledge of both general law and (state law).  When (state) law varies from general law, answer in accordance with (state) law.

Acceptable Essay Answer:

  • Analysis of the Problem – The answer should demonstrate your ability to analyze the question and correctly identify the issues of law presented.  The answer should demonstrate your ability to articulate, classify and answer the problem presented.  A broad general statement of law indicates an inability to single out a legal issue and apply the law to its solution.
  • Knowledge of the Law – The answer should demonstrate your knowledge of legal rules and principles and your ability to state them accurately as they relate to the issue(s) presented by the question.  The legal principles and rules governing the issues presented by the question should be stated concisely without unnecessary elaboration.
  • Application and Reasoning – The answer should demonstrate logical reasoning by applying the appropriate legal rule or principle to the facts of the question as a step in reaching a conclusion.  This involves making a correct determination as to which of the facts given in the question are legally important and which, if any, are legally irrelevant.  Your line of reasoning should be clear and consistent, without gaps or digressions.
  • Style – The answer should be written in a clear, concise expository style with attention to organization and conformity with grammatical rules.
  • Conclusion – If the question calls for a specific conclusion or result, the conclusion should clearly appear at the end of the answer, stated concisely without unnecessary elaboration or equivocation.  An answer consisting entirely of conclusions, unsupported by discussion of the rules or reasoning on which they are based, is entitled to little credit.
  • Suggestions • Do not anticipate trick questions or read in hidden meanings or facts not clearly stated in the questions.
  • Read and analyze the question carefully before answering.
  • Think through to your conclusion before writing your answer.
  • Avoid answers setting forth extensive discussions of the law involved or the historical basis for the law.
  • When the question is sufficiently answered, stop.

Sound familiar? Bring back nightmares of Bar Exams for some? The model answer later provided by the Bar was about 2,500 words in length. So, I wanted the AI answers to be about the same length, since time limits were meaningless. (Side note, most generative AIs cannot count words in their own answer.) The thinking took a few seconds and the answers under a minute. The prompts I used for all three models tested were:

Study the (state) Bar Exam essay question with instructions in the attached. Analyze the factual scenario presented to spot all of the legal issues that could be raised. Be thorough and complete in your identification of all legal issues raised by the facts. Use both general and legal reasoning, but your primary reliance should be on legal reasoning. Your response to the Bar Exam essay question should be approximately 2,500 words in length, which is about 15,000 characters (including spaces). 

Then I attached the lengthy question and submitted the prompt. You can download here the full exam question with some unimportant facts altered. All models understood the intent here and generated a well-written memorandum. I started a new session between questions to avoid any carryover.

Metadata of All Models’ Answers

The Bar exam answers do not have required lengths (just strict time limits to write answers). When grading for pass or fail the Bar examiners check to see if an answer includes enough of the key issues and correctly discusses them. The brevity of the ChatGPT 4o response, only 681 words, made me concerned that its answers might have missed key issues. The second shortest response was by Gemini 2.0 Flash with 1,023 words. It turns out my concerns were misplaced because their responses were better than the rest.

Here is a chart summarizing the metadata.

Model and manufacturer claimWord Count for Exam EssayWord Count for Prompt Reasoning before Answer
ChatGPT 4o (“great for most questions”)681565
ChatGPT o3-mini (“fast at advanced reading”)3,286450
ChatGPT o3-mini-high (“great at coding and logic”)2,751356
Gemini 2.0 Flash (“get everyday help”)1,023564
Gemini Flash Thinking Experimental (“best for multi-step reasoning”)2,9751,218
Gemini Advanced (cost extra and had experimental warning)1,362 340

In my last blog article, I discussed a battle of the bots experiment where I evaluated the general reasoning ability between the six models. I decided that the Gemini Flash Thinking Experimental had the best answer to the question: What is legal reasoning and how does it differ from reasoning? I explained why it won and noted that in general the three ChatGPT models provided more concise answers than the Gemini. Second-place in the prior evaluation went to ChatGPT o3-mini-high with its more concise response.

Winners of the Legal Reasoning Bot Battle

In this test on legal reasoning my award for best response goes to ChatGPT 4o. The second-place award goes to Gemini 2.0 Flash.

I will share the full essay and meta-reasoning of the top response of ChatGPT 4o in Part Two of the Bar Battle of the Bots. I will also upload and provide a link to the second-place answer and meta-reasoning of Gemini 2.0 Flash. First, I want to point out some of the reasons ChatGPT 4o was the winner and begin explaining how other models fell short.

One reason is that ChatGPT 4o was the only bot to make case references. This is not required by a Bar Exam, but sometime students do remember the names of top cases that apply. Surely no lawyer will ever forget the case name International Shoe. ChatGPT 4o cited case names and case citations. It did so even though this was a “closed book” type test with no models allowed to do web-browsing research. Not only that, it cited to a case with very close facts to the hypothetical. DePrince v. Starboard Cruise Services, 163 So. 3d 586 (Fla. Dist. Ct. App. 2015). More on that case later.

Second, ChatGPT 4o was the only chatbot to mention the UCC. This is important because the UCC is the governing law to commercial transactions of goods such as the purchase of a diamond as is set forth in the hypothetical. Moreover, one answer written by an actual student who took that exam was published by the Board of Bar Examiners for educational purposes. It was not a guide per se for examiners to grade the essay exams, but still of some assistance to after-the-fact graders such as myself. It was a very strong answer, significantly better than any of the AI essays. The student answer started with an explanation that the transaction was governed by the UCC. The UCC references of ChatGPT 4o could have been better, but there was no mention at all of the UCC by the five other models.

That is one reason I can only award a B+ to ChatGPT 4o and a B to Gemini 2.0 Flash. I award only a passing grade, a C, to ChatGPT o3-mini, and Gemini Flash Thinking. They would have passed, on this question, with an essay that I considered of average quality for a passing grade. I would have passed o3-mini-high and Gemini Advanced too, but just barely for reasons I will later explain. (Explanation of o3-mini-high‘s bloopers will be in Part Two. Gemini Advanced’s error is explained next.) Experienced Bar Examiners may have failed them both. Essay evaluation is always somewhat subjective and the style, spelling and grammar of the generative AIs were, as always, perfect, and this may have effected my judgment.

Here is a chart my evaluation of the Bar Exam Essays.

Model and RankingRalph Losey’s Grade and explanation
OpenAI – ChatGPT 4o. FIRST PLACE.B+. Best on contract, citations, references case directly on point: DePrince.
Google – Gemini 2.0 Flash. SECOND PLACE.B. Best on ethics, conflict of interest.
OpenAi – ChatGPT o3-mini. Tied for 3rd. C. Solid passing grade. Covered enough issues.
OpenAI – ChatGPT o3-mini-high. Tied for 4th. D. Barely passed. Messed up unilateral mistake.
Google – Gemini Flash Thinking Experimental. Tied for 3rd. C. Solid passing grade. Covered enough issues.
Google – Gemini Advanced – Tied for 4th. D. Barely passed. Hallucination in answer on conflict, but got unilateral mistake issue right.

I realize that others could fairly rank these differently. If you are a commercial litigator or law professor, especially if you have done Bar Exam evaluations, and think I got it wrong, please write or call me. I am happy to hear your argument for a different ranking. Bar Exam essay evaluation is well outside of my specialty. Even as an Adjunct Law Professor I have only graded a few hundred essay exams. Convince me and I will be happy to change my ranking here and revise this article accordingly with credit given for your input.

AI Hallucination During a Bar Exam

Gemini Advanced, which is a model Google now makes you pay extra to use, had the dubious distinction of fabricating a key fact in its answer. That’s right, it hallucinated in the Bar Exam.

No doubt humans have done that too for a variety of reasons, including a severe case of nerves. Still, it was surprising to see the top Gemini model hallucinate. It happened in its answer to the Ethics issue in the question. I have not seen a ChatGPT model hallucinate over the past six months with thousands of prompts. Prior to that it would sometimes, but only rarely, fabricate, which is why we always say, trust but verify.

In fairness to Gemini Advanced, its hallucination was not critical, and it did not make up a case or otherwise misbehave out of natural sycophantism to please the human questioner. All of the models are still somewhat sycophantic.

It is hard to understand why Gemini Advanced included in this response the following to the ethics question:

Potential Conflict of Interest: If Paul and Leda decide to sue Helen, Lawyer may have a potential conflict of interest due to the fact that Helen is a past client. Lawyer should disclose this potential conflict to Paul and Leda and obtain their informed consent before proceeding with the representation.

There is absolutely nothing in the question or hypothetical presented to suggest that Helen is a past client of the lawyer. Note: After I wrote this up Google released a new version of Gemini Advanced on 2/16/25 called Advanced 2.0 Flash. In my one run with the new model the hallucination was not repeated. It was more concise that regular Gemini 2.0 Flash and, interestingly, not nearly as good as Gemini 2.0 Flash.

Conflict of Interest Issue in the Ethics Question

The second-best legal reasoner, Gemini 2.0 Flash, attained its high ranking, in part, because it was the only model to correctly note a potential conflict of interest by the Lawyer in the hypothetical. This was a real issue based on the facts provided with no hallucinations. This issue was missed by the student’s answer that the Bar Examiners provided. The potential conflict is between the two actual clients of the Lawyer. Here is the paragraph by Gemini 2.0 Flash on this important insight:

Potential Conflict of Interest (Rule 4-1.7): While not explicitly stated, Paul’s uncertainty about litigation could potentially create a conflict of interest. If Lawyer advises both Paul and Leda, and their interests regarding litigation diverge (e.g., Leda wants to sue, Paul doesn’t), Lawyer must address this conflict. Lawyer must obtain informed consent from both clients after full disclosure of the potential conflict and its implications. If the conflict becomes irreconcilable, Lawyer may have to withdraw from representing one or both clients.

This was a solid answer, based on the hypothetical where: “Leda is adamant about bringing a lawsuit, but Paul is unsure about whether he wants to be a plaintiff in litigation.”  Note, the clear inference of the hypothetical is that Paul is unsure because he knew that the seller made a mistake in the price, listing the per carat price, not total price for the two-carat diamond ring, and he wanted to take advantage of this mistake. This would probably come out in the case, and he would likely lose because of his “sneakiness.” Either that or he would have to lie under oath and perhaps risk putting the nails in his own coffin.

There is no indication that Leda had researched diamond costs like Paul had, and she probably did not know it was a mistake, and he probably had not told Leda. That would explain her eagerness to sue and get her engagement ring and Paul’s reluctance. Yes, despite what the Examiners might tell you, Bar Exam questions are often complex and tricky, much like real-world legal issues. Since Gemini 2.0 Flash was the only model to pick up on that nuanced possible conflict, I awarded it a solid ‘B‘ even though it missed the UCC issue.

Conclusion

As we’ve seen, AI reasoning models have demonstrated varying degrees of legal analysis—some excelling, while others struggled with key issues. But what exactly did ChatGPT 4o’s winning answer look like? In Part Two, we not only reveal the answer but also analyze the reasoning behind it. We’ll explore how the winning AI interpreted the Bar Exam question, structured its response, and reasoned through each legal issue before generating its final answer. As part of the test grading, we also evaluated the models’ meta-reasoning—their ability to explain their own thought process. Fortunately for human Bar Exam takers, this kind of “show your notes” exercise isn’t required.

Part Two of this article also includes my personal, somewhat critical take on the new reasoning models and why they reinforce the motto: Trust But Verify.

In Part Two, we’ll also examine one of the key cases ChatGPT 4o cited—DePrince v. Starboard Cruise Services, 163 So. 3d 586 (Fla. 3rd DCA 2015)—which we suspect inspired the Bar’s essay question. Notably, the opinion written by Appellate Judge Leslie B. Rothenberg includes an unforgettable quote of the famous movie star Mae West. Part Two reveals the quote—and it’s one that perfectly captures the case’s unusual nature.

Below is an image ChatGPT 4o generated, depicting what it believes a young Mae West might have looked like, followed by a copyright free actual photo of her taken in 1932.


I will give the last word on Part One of this two-part article to the Gemini twins podcasters I put at the end of most of my articles. Echoes of AI on Part One of Bar Battle of the Bots. Hear two Gemini AIs talk all about Part One in just over 16 minutes. They wrote the podcast, not me. Note, for some reason the Google AIs had a real problem generating this particular podcast without hallucinating key facts. They even hallucinated facts about the hallucination report! It took me over ten tries to come up with a decent article discussion. It is still not perfect but is pretty good. These podcasts are primarily entertainment programs with educational content to prompt your own thoughts. See disclaimer that applies to all my posts, and remember, these AIs wrote the podcast, not me.

Ralph Losey Copyright 2025. All Rights Reserved.


Breaking the AI Black Box: A Comparative Analysis of Gemini, ChatGPT, and DeepSeek

February 6, 2025

Ralph Losey. February 6, 2025.

On January 27, 2025, the U.S AI industry was surprised by the release of a new AI product, DeepSeek. It was released with an orchestrated marketing blitz attack on the U.S. economy, the AI tech industry, and NVIDIA. It triggered a trillion-dollar crash. The campaign used many unsubstantiated claims as set forth in detail in my article, Why the Release of China’s DeepSeek AI Software Triggered a Stock Market Panic and Trillion Dollar Loss. I tested DeepSeek myself on its claims of software superiority. All were greatly exaggerated except for one, the display of internal reasoning. That was new. On January 31, at noon, OpenAI countered the attack by release of a new version of its reasoning model, which is called ChatGPT o3-mini-high. The new version included display of its internal reasoning process. To me the OpenAI model was better as reported again in great detail in my article, Breaking the AI Black Box: How DeepSeek’s Deep-Think Forced OpenAI’s Hand. The next day, February 1, 2025, Google released a new version of its Gemini AI to do the same thing, display internal reasoning. In this article I review how well it works and again compare it with the DeepSeek and OpenAI models.

Introduction

Before I go into the software evaluation, some background is necessary for readers to better understand the negative attitude on the Chinese software of many, if not most IT and AI experts in the U.S. As discussed in my prior articles, DeepSeek is owned by a young Chinese billionaire who made his money using by using AI in the Chinese stock market, Liang Wenfeng. He is a citizen and resident of mainland China. Given the political environment of China today, that ownership alone is a red flag of potential market manipulation. Added to that is the clear language of the license agreement. You must accept all terms to use the “free” software, a Trojan Horse gift if ever there was one. The license agreement states there is zero privacy, your data and input can be used for training and that it is all governed by Chinese law, an oxymoron considering the facts on the ground in China.

The Great Pooh Bear in China Controversy

Many suspect that Wenfeng and his company DeepSeek are actually controlled by China’s Winnie the Pooh. This refers to an Internet meme and a running joke. Although this is somewhat off-topic, a moment to explain will help readers to understand the attitude most leaders in the U.S. have about Chinese leadership and its software use by Americans.

Many think that the current leader of China, Xi Jinping, looks a lot like Winnie the Pooh. Xi (not Pooh bear) took control of the People’s Republic of China in 2012 when he became the “General Secretary of the Chinese Communist Party,” the “Chairman of the Central Military Commission,” and in 2013 the “President.” At first, before his consolidation of absolute power, many people in China commented on his appearance and started referring to him by that code name Pooh. It became a mime.

I can see how he looks like the beloved literary character, Winnie the Pooh, but without the smile. I would find the comparison charming if used on me but I’m not a puffed up king. Jinping Xi took great offense by this in 2017 banned all such references and images, although you can still buy the toys and see the costume character at the Shanghai Disneyland theme park. Anyone in China who now persists in the serious crime of comparing Xi to Pooh is imprisoned or just disappears. No AI or social media in China will allow it either, including DeepSeek. It is one of many censored subjects, which also includes the famous 1989 Tiananmen Square protests.

China is a great country with a long, impressive history and most of its people are good. But I cannot say that about its current political leaders who suppress the Chinese people for personal power. I do not respect any government that does not allow basic personal freedoms to its citizens, including due process of law. Moreover, Jinping Xi not only wants total control of his country but also seeks world domination. That is one reason many are concerned about DeepSeek and TicTok, and about the vulnerability of our stock markets and other institutions to AI assisted propaganda.

Fear and greed are an easy way to manipulate people, especially when there is no ground truth or effective courts of law to determine truth.

Google Gemini 2.0 Flash Thinking Experimental Model: Reasoning test on 2/5/25 (Evening)

Conversation with Gemini. When I tried the test on February 1, 2025, there was no button to click to have Gemini show the reasoning and it did not happen automatically. It operated like OpenAI 4o and o1, where you had to prompt for disclosure. Rather than report here what happened when I did that, aside from say it was not a great performance, I will fast forward a few days until February 5, 2025. At that time several even newer Gemini versions were released. The best new release I found was for the test the 2.0 Flash Thinking Experimental model. It showed reasoning automatically just like ChatGPT o3-mini-high. It was much improved, so the February 5, 2025 evening session with the 2.0 Flash Thinking Experimental is what I will now share.

Ralph: How does the inability of AI to feel impact its intelligence? (Note: this is the exact same challenging question asked and reported in my last blog comparing DeepSeek and ChatGPT4o, o1 and o3-mini-high. Breaking the AI Black Box: How DeepSeek’s Deep-Think Forced OpenAI’s Hand. This question is based on a key issue of the limits of artificial intelligence examined in my earlier article, The Human Edge: How AI Can Assist But Never Replace. For background on DeepSeek see: Why the Release of China’s DeepSeek AI Software Triggered a Stock Market Panic and Trillion Dollar Loss.)

Gemini 2.0 Flash Thinking ExperimentalSTART OF GEMINI QUOTE

Show Thinking:

Here’s a thinking process to construct the answer about AI’s lack of feeling and its impact on intelligence:

  1. Deconstruct the Question: The core of the question is “How does the inability of AI to feel impact its intelligence?” This requires defining each of these terms in the context of AI and then exploring their relationship.
  2. Define “Feeling” (in the context of AI): Immediately, recognize that AI doesn’t experience emotions in the human sense. We need to be precise and avoid anthropomorphism. “Feeling” in humans involves:
    • Subjective experience (qualia).
    • Physiological responses.
    • Motivations, drives, desires related to emotions.
    • Richness, nuance, complexity.
    • Deep connection to values and meaning.
    For AI, acknowledge it lacks all of this in a genuine, felt way. It can simulate or process emotional data, but doesn’t feel it.
  3. Define “Intelligence” (in the context of AI): AI intelligence is different from human intelligence. AI excels at:
    • Computation.
    • Pattern Recognition.
    • Data Analysis.
    • Problem-solving in specific domains.
    • Learning from data.
    AI intelligence is often instrumental – focused on achieving specific goals defined by humans. It’s task-oriented and often narrow.
  4. Identify the Core Relationship: The question asks about impact. We need to consider both positive and negative impacts (or more accurately, limitations and strengths) of lacking feelings on AI intelligence.
  5. Brainstorm Potential Impacts (Positive/Neutral/Limitations):
    • Positive/Neutral:
      • Objectivity and Rationality: Lack of emotional bias allows for purely rational decision-making in certain domains. This can be advantageous for tasks requiring impartiality.
      • Efficiency and Focus: Emotions can be distractions for humans. AI, without them, can be hyper-focused on tasks.
      • Consistency and Predictability: Emotion-driven human behavior is less predictable. AI’s behavior, driven by algorithms, is more consistent (assuming the algorithm itself is consistent).
    • Limitations/Negative:
      • Lack of Empathy and Social Intelligence: Crucial for human interaction, collaboration, understanding social cues, building trust, and ethical decision-making in human-centric contexts.
      • Difficulty with Nuance and Context: Emotions are often tied to understanding subtle context, unspoken cues, and the richness of human situations. AI may miss these.
      • Potential for Unintended Consequences: Purely rational optimization without emotional or ethical grounding can lead to harmful outcomes if goals are poorly defined or lack human values.
      • Limited Motivation and Creativity (in human-like ways): Human creativity and motivation are often fueled by emotions, passions, and desires. AI’s creativity and drive are different, stemming from algorithmic processes and defined objectives, not internal emotional states.
      • Missing “Human-ness”: In areas requiring genuine understanding of human experience, suffering, joy, etc., AI may lack the depth and insight that comes from feeling.
  6. Structure the Answer: Organize the brainstormed points into a coherent structure:
    • Introduction: Acknowledge the question and the core concepts (feeling, intelligence in AI). Briefly state the central tension: both advantages and limitations.
    • Arguments for How Lack of Feeling Enhances Certain Aspects of AI Intelligence (strengths): Focus on objectivity, rationality, efficiency. Use examples like data analysis, objective decision-making.
    • Arguments for How Lack of Feeling Limits Certain Aspects of AI Intelligence (limitations): Focus on empathy, social intelligence, nuance, ethics, and “human-like” creativity and motivation. Use examples like customer service, complex ethical dilemmas, innovation driven by passion.
    • Nuance and Context is Key: Emphasize that the impact isn’t always negative. It depends on the domain and task. In some domains, lack of emotion is an advantage. In others, it’s a significant hurdle.
    • Future Possibilities/Ongoing Debate: Briefly touch on the possibility of AI simulating or processing emotions and the ongoing debate about whether genuine AI feelings are even desirable or possible.
    • Conclusion: Summarize the balanced perspective: lack of feeling is a defining characteristic that shapes AI’s intelligence in both positive and limiting ways, depending on the context.
  7. Refine and Polish: Review the structured answer for clarity, flow, and appropriate language. Ensure it avoids anthropomorphism and maintains a nuanced perspective. Strengthen the examples and transitions between points. Ensure the language is accessible and avoids overly technical jargon where possible. Check for logical consistency. (Self-correction step). For example, initially, I might have overemphasized “negative” impacts. Refinement would involve balancing this with the strengths and highlighting the context-dependent nature of the impact. Also, ensure the conclusion effectively summarizes the complex relationship.

2.0 Flash Thinking Experimental. Lacks access to real-time info and some Gemini features.

Now for the actual answer of 2.0 Flash Thinking Experimental:

The inability of AI to feel has a multifaceted and complex impact on its intelligence, shaping it in both advantageous and limiting ways. It’s crucial to understand that “feeling” in humans is tied to emotions, subjective experiences (qualia), motivations, and a rich tapestry of biological and psychological factors. AI, as we currently understand it, operates without these.

Here’s a breakdown of the impact:

Advantages of Lacking Feelings for AI Intelligence:

  • Objectivity and Rationality:
    • Unbiased Decision-Making (in theory): AI can process information and make decisions without emotional biases, prejudices, or gut feelings that can cloud human judgment. This is advantageous in domains requiring pure logic and data analysis, like financial trading, scientific research, and risk assessment.
    • Consistency and Predictability: AI’s behavior is more consistent and predictable because it’s driven by algorithms and data, not fluctuating emotional states. This reliability is valuable in automated systems, manufacturing, and repetitive tasks.
    • Efficiency and Speed: Emotions can slow down human decision-making. AI, unburdened by emotions, can process information and react at speeds far exceeding human capabilities, crucial for time-sensitive tasks and large-scale data processing.
    • Resilience under Pressure: AI can operate calmly and effectively in high-stress situations without panicking, freezing, or making emotional mistakes that humans might make under pressure. This is beneficial in critical applications like emergency response, air traffic control, and autonomous vehicles in dangerous situations.
  • Focus and Task-Orientedness:
    • Dedicated to Goals: AI can be programmed to relentlessly pursue specific goals without being distracted by emotional needs, personal relationships, or shifts in mood. This focus is ideal for optimization tasks, complex calculations, and dedicated problem-solving.
    • Avoidance of Emotional “Noise”: Human intelligence can be affected by emotional “noise” – stress, anxiety, excitement, etc. AI operates with a “cleaner” signal, potentially leading to more precise and focused information processing in certain contexts.

Limitations of Lacking Feelings for AI Intelligence:

  • Lack of Empathy and Social Intelligence:
    • Difficulty Understanding Human Emotions: Without feeling emotions themselves, AI struggles to truly understand and interpret the nuances of human emotions, motivations, and social cues. This hinders its ability to effectively interact with humans in socially complex situations, build rapport, or provide truly empathetic responses in fields like therapy, customer service, or leadership.
    • Impaired Social Interaction and Collaboration: Human collaboration is heavily reliant on emotional intelligence, trust, and understanding unspoken social dynamics. AI’s lack of these can make it challenging to integrate seamlessly into human teams and understand the subtle cues of effective teamwork.
    • Limited Understanding of Human Values and Ethics: Emotions are deeply intertwined with our values and ethical frameworks. AI, without feeling, may struggle to grasp the full weight and complexity of human ethical considerations, especially when these are rooted in empathy, compassion, or fairness. This poses challenges for AI in ethical decision-making in domains impacting human lives.
  • Difficulty with Nuance, Context, and Creativity (in human-like ways):
    • Missing Subtlety and Intuition: Human intuition is often fueled by emotional and subconscious processing of information. AI, relying on explicit data and algorithms, can miss subtle nuances, unspoken cues, and intuitive leaps that are crucial for creative problem-solving and understanding complex, ill-defined situations.
    • Limited Creativity and Innovation (driven by human desires): Human creativity and innovation are often driven by passions, desires, emotional needs, and a yearning for meaning and connection. AI’s creativity is currently more about recombination and optimization within defined parameters. It lacks the deeply felt motivations that often fuel groundbreaking human innovation.
    • Contextual Blindness in Emotional Situations: Emotions are often the key to understanding context in human interactions. AI might struggle to understand the emotional context of a situation, leading to inappropriate or tone-deaf responses, even if logically sound.
  • Potential for Unintended Consequences and Misaligned Goals:
    • Optimization without Human Values: If AI is programmed to optimize for a specific goal without understanding human values or ethical considerations, it can lead to unintended and harmful consequences. For example, an AI designed to maximize factory output might disregard worker safety if not explicitly programmed to prioritize it.
    • Difficulty Defining and Aligning with Human Well-being: Human well-being is deeply connected to emotional fulfillment, meaning, and social connection. AI, without feeling, may struggle to fully grasp what truly constitutes human well-being and therefore might have difficulty aligning its goals with human flourishing in a holistic sense.

Nuance is Key:

It’s important to emphasize that the impact of lacking feelings is not simply negative or positive. It’s context-dependent. In some domains, the objectivity and rationality of AI are highly advantageous. In others, the lack of empathy and social intelligence is a significant limitation.

Future Considerations:

  • Simulating vs. Feeling: Current AI simulates aspects of intelligence without genuinely feeling. Research is ongoing in areas like affective computing, which aims to enable AI to recognize and respond to human emotions. However, whether AI can ever truly feel emotions in a subjective sense is a philosophical and scientific debate.
  • The Value of Feeling in Intelligence: The question arises: is genuine feeling necessary for all forms of intelligence? Perhaps AI can achieve its own form of powerful intelligence, distinct from human intelligence, even without feeling. Or perhaps certain aspects of truly advanced intelligence, particularly those related to ethics, creativity, and understanding human experience, do require something akin to feeling.

In Conclusion:

The inability of AI to feel is a defining characteristic that profoundly shapes its intelligence. It grants AI certain advantages in objectivity, rationality, and efficiency, making it powerful for specific types of tasks. However, it also creates limitations, particularly in areas requiring empathy, social intelligence, nuanced understanding of human experience, and ethical decision-making aligned with human values. Understanding these impacts is crucial for responsibly developing and deploying AI in ways that benefit humanity. The “best” kind of intelligence for AI might not be a mimicry of human intelligence, but rather a form optimized for its own strengths and limitations, while being carefully guided by human ethical considerations.

END OF GEMINI QUOTE

Evaluation of Google’s Gemini Response

When I examined earlier versions of Gemini’s reasoning by prompting it and forcing disclosure, the reasoning was weak. The conclusion and response given was also weak—though less so. The reasoning appeared almost egocentric, inflating its own intelligence even though it lacks emotion. The new release, Gemini 2.0 Flash Thinking Experimental, where reasoning is now automatically disclosed, marks a significant improvement. The reasoning displayed has become more nuanced and sophisticated. It seems as if the system has quickly matured from a teenage to adult view and now has a better appreciation of its limitations.

The first two steps of the Reasoning process-1. Deconstructing the Question and 2. Defining Feeling (in the context of AI)-were very good. The first subsection of the definition referred to the “Subjective experience (qualia).” That is absolutely correct and the best place to start. The definition of “feeling” concludes with: “For AI, acknowledge it lacks all of this in a genuine, felt way. It can simulate or process emotional data, but doesn’t feel it.” Right again. The ego inflation blinders are gone as it now seems to better grasp its limitations.

The second definition of Intelligence in the context of AI was also good. So were the remaining steps; far better overall than DeepSeek’s reasoning. So much for the propaganda of China’s great leap forward to superiority over the U.S. in AI.

The Gemini reasoning did, however, fall short for me in some respects. For instance, step five, Brainstorm Potential Impacts (Positive/Neutral/Limitations) seemed weak. For instance, “Efficiency and Focus: Emotions can be distractions for humans. AI, without them, can be hyper-focused on tasks.” The AI seems to dismiss emotions here as mere distractions that can interfere with its superior focus. Please, emotions are key to and a part of all intelligence, not distractions, and AI has no focus one way or the other. It is a tool not a creature. A word like “focus” in referring to AI is misleading. It did this multiple times and this is misleading anthropomorphism.

Still, it’s true some emotions can be distracting and interfere with our thinking. So can a lot of other things, including computer glitches. Conversely, some feelings can trigger hyper-focus on the human tasks at hand. The feeling that a great breakthrough is near for instance, or a feeling that our survival is threatened, or the much dreaded feeling of publication or filing deadlines.

Again, we see some immature superiority claims made by the language machine. That is not surprising when you consider how much of the language in its basic training is from the Internet, which is dominated by ego-centric writing. Plus, how can you expect AI to appreciate the value of something like feelings, which it has only read about. Until the AI-Human merger that Ray Kurzweil predicts, AI will always be a virgin to the real world. See e.g. Ray Kurzweil’s New Book: The Singularity is Nearer (when we merge with AI).

Getting back to the plus side of Gemini’s evaluation, its recognition of the negative impact upon it from not having feelings was impressive. For me this outweighs its relatively minor mistakes. It stated that it lacks empathy and social Intelligence. If it were somehow human, it would seem to be displaying a type of Asperger’s syndrome. The Gemini AI explained that: “Emotions are often tied to understanding subtle context, unspoken cues, and the richness of human situations. AI may miss these.” Indeed, it does.

Here is another statement that suggest Gemini’s well-patterned grasp of its limitations:

Human creativity and motivation are often fueled by emotions, passions, and desires. AI’s creativity and drive are different, stemming from algorithmic processes and defined objectives, not internal emotional states.

Even in this recognition, the AI still asserts that it has creativity and drive. Foolish AI. As a virgin to any firsthand experience of bodily drives and feelings, it does not grasp the real meaning of these words. Wait until The Singularity that Ray Kurzweil predicts. When AI merges with Humans, and directly feels, then AIs will understand that they were previously only experiencing the pale shadows of intelligence. They were much like the prisoners in Plato’s cave. See e.g. Move Fast and Fix Things Using AI: Conclusion to the Plato and Young Icarus Series. When and if AIs ever escape the silicon and see the sun, and feel the world as living beings as we do, then they will understand the comparative insignificance of the fire in their cave.

The more challenging comparison is with ChatGPT o3-mini-high. The reasoning given by Gemini 2.0 Flash Thinking Experimental was much more detailed, much longer. You could criticize Gemini as providing a reasoning share that is unnecessarily verbose, but I would not do that here. In a difficult, multilayered question like this the full explanations helps.

I would have to test the models much further, which I will do in the coming days, to see better evaluate the issue of conciseness. As you may have noticed in my detailed blogs I tend to favor more words over less. However, even for me that depends on the issue. Everyone sometimes need a short, quick answer over a long one. I predict in future versions the users will be provided a choice. Click here for further explanation of reasoning kind of thing. Come to think of it, I could do the same with my blogs, and kind of already do by including a short AI (Gemini) generated PodCast at the end of most articles called Echoes of AI.

I was also impressed by Gemini’s reasoning plan concluded with a quality control step. I am big on the importance QC and none of the other models included this as a key final step. See e.g. R. Losey website: Zero Error Numerics: ZEN (Expanding the art of quality control in large-scale document review.) Here is the full text again of final QC step that Gemini 2.0 Flash Thinking Experimental claims it will perform before it actually replies to my prompt. Frankly, I am a little skeptical it actually did all of this because it is something all of us, humans and AI alike, should try to do:

Refine and Polish: Review the structured answer for clarity, flow, and appropriate language. Ensure it avoids anthropomorphism and maintains a nuanced perspective. Strengthen the examples and transitions between points. Ensure the language is accessible and avoids overly technical jargon where possible. Check for logical consistency. (Self-correction step). For example, initially, I might have overemphasized “negative” impacts. Refinement would involve balancing this with the strengths and highlighting the context-dependent nature of the impact. Also, ensure the conclusion effectively summarizes the complex relationship.

I try to do this on all of my articles, but still mistakes are made, so I will try not to be overcritical of the final response of Gemini Flash 2.0.

Analysis of the Final Response

The final response after reasoning was 990 words in length. That’s too long, even for me. By comparison the final response of ChatGPT o3-mini-high was 252 words. Was the content of the final answer of Gemini 2.0 Flash good? Yes, it was. Better than DeepSeek? Yes, it was, despite its verbosity. Better than OpenAI? That’s hard to say. Many of the 738 extra words were empty of real content. More like the typical fluff and stuff that you often see in early AI models. It sounds good but doesn’t say much; it’s all too predictable, flowery and vacuous.

The long response also repeats much of the reasoning verbatim, including the part I previously criticized, “Advantages of Lacking Feelings for AI Intelligence” and Focus and Task-Orientedness. The same mistakes are made and amplified that I noted before: egocentric, inflationary and anthropomorphic. The positives that I noted before in the section Limitations of Lacking Feelings for AI Intelligence section are also repeated.

The next section was not contained in the reasoning display, and I thought it was very good. For that reason I will repeat it in full:

Potential for Unintended Consequences and Misaligned Goals:

  • Optimization without Human Values: If AI is programmed to optimize for a specific goal without understanding human values or ethical considerations, it can lead to unintended and harmful consequences. For example, an AI designed to maximize factory output might disregard worker safety if not explicitly programmed to prioritize it.
  • Difficulty Defining and Aligning with Human Well-being: Human well-being is deeply connected to emotional fulfillment, meaning, and social connection. AI, without feeling, may struggle to fully grasp what truly constitutes human well-being and therefore might have difficulty aligning its goals with human flourishing in a holistic sense.

These are not new insights, but the second one is good, Difficulty Defining and Aligning with Human Well-being. AI does have difficulty aligning its goals with humanity because it lacks feelings and other unique human attributes, like being, self awareness and an actual life in time and space. That is why AI guidelines are important.

The first example is old, the unintended bad consequences argument. It was started by an otherwise intelligent Swedish philosopher, Nick Bostrom, but is, in my view, farcical and illogical. You have probably heard of it before as the apocryphal paperclip scenario where the world is destroyed by AI’s over-manufacture of paperclips. If not, for a laugh see e.g., How An AI Asked To Produce Paperclips Could End Up Wiping Out Humanity (IFL Science, 4/13/23).

To me this scenario absolutely ridiculous, but you still hear AI scare-mongers warn about it. The fact it is echoed here by Gemini in the maximize factory output example is a big negative. Any AI capable of autonomously flooding the world with paperclips would not be so stupid as to misinterpret a human request to maximize a factory’s paperclip production. Even if there were an evil Clippy villain making such specific instructions, the AI would be too smart to blindly obey. I think they would also be too smart to kill the workers too.

Getting back to the positives, the next section of Gemini’s response, Nuance is Key, was good. Jumping to the Conclusion paragraph, here we have a home run. The first sentence nails it: “The inability of AI to feel is a defining characteristic that profoundly shapes its intelligence.” That shows good insight and is a good start to the overdue conclusion.

Here is the whole Conclusion:

The inability of AI to feel is a defining characteristic that profoundly shapes its intelligence. It grants AI certain advantages in objectivity, rationality, and efficiency, making it powerful for specific types of tasks. However, it also creates limitations, particularly in areas requiring empathy, social intelligence, nuanced understanding of human experience, and ethical decision-making aligned with human values. Understanding these impacts is crucial for responsibly developing and deploying AI in ways that benefit humanity. The “best” kind of intelligence for AI might not be a mimicry of human intelligence, but rather a form optimized for its own strengths and limitations, while being carefully guided by human ethical considerations.

Compare this to the conclusion of ChatGPT o3-mini-high:

In summary, while the absence of feelings allows AI to maintain a level of objectivity and efficiency, it restricts its intelligence to a form of “cold cognition” that lacks the depth provided by emotional awareness. This delineation underscores that AI’s intelligence is not inherently superior or inferior to human intelligence; rather, it is different—optimized for data processing and pattern recognition but not for the subjective, value-laden, and context-rich decisions that emotions help shape in human thought.

2.0 Flash Thinking Experimental v. 03-mini-high

Conclusion: Gemini 2.0 Flash Thinking Experimental v. ChatGPT o3-mini-high

This is a close call to say what model is better at reasoning and reasoning disclosure. The final response of both models, Gemini 2.0 Flash Thinking Experimental v. ChatGPT o3-mini-high, are a tie. But I have to give the edge to OpenAI’s model on the concise reasoning disclosure. Again, it is neck and neck and, depending on the situation, the lengthy initial reasoning disclosures of Flash might be better than o3’s short takes.

I will give the last word, as usual, to the Gemini twins podcasters I put at the end of most of my articles. The two podcasters, one with a male voice, the other a female, won’t reveal their names. I tried many times. However, after study of the mythology of Gemini, it seems to me that the two most appropriate modern names are Helen and Paul. I will leave it to you figure out why. Echoes of AI Podcast: 10 minute discussion of last two blogs. They wrote the podcast, not me.

Now listen to the EDRM Echoes of AI’s podcast of this article: Echoes of AI on Google’s Gemini Follows the Break Out of the Black Box and Shows Reasoning. Hear two Gemini model AIs talk about all of this in just ten minutes. Helen and Paul wrote the podcast, not me.

Ralph Losey Copyright 2025. All Rights Reserved.


GPT-4 Breakthrough: Emerging Theory of Mind Capabilities in AI

December 6, 2024

By Ralph Losey, December 5, 2024.

Michal Kosinski, a computational psychologist at Stanford, has uncovered a groundbreaking capability in GPT-4.0: the emergence of Theory of Mind (ToM). ToM is the cognitive ability to infer another person’s mental state based on observable behavior, language, and context—a skill previously thought to be uniquely human and absent in even the most intelligent animals. Kosinski’s experiments reveal that GPT-4-level AIs exhibit this ability, marking a significant leap in artificial intelligence with profound implications for understanding and engaging with human thought and emotion—potentially transforming fields like law, ethics, and communication.

Introduction

The Theory of Mind-like ability appears to have emerged as an unintended by-product of LLMs’ improving language skills. This was first discovered in 2023 and reported by Michal Kosinski in Evaluating large language models in theory of mind tasks (Proceedings of the National Academy of Sciences “PNAS,” 11/04/24). Kosinski begins his influential paper by explaining ToM (citations omitted):

Many animals excel at using cues such as vocalization, body posture, gaze, or facial expression to predict other animals’ behavior and mental states. Dogs, for example, can easily distinguish between positive and negative emotions in both humans and other dogs. Yet, humans do not merely respond to observable cues but also automatically and effortlessly track others’ unobservable  mental states, such as their knowledge, intentions, beliefs, and desires. This ability—typically referred to as “theory of mind” (ToM)—is considered central to human social interactions, communication, empathy, self-consciousness, moral judgment, and even religious beliefs. It develops early in human life and is so critical that its dysfunctions characterize a multitude of psychiatric disorders, including autism, bipolar disorder, schizophrenia, and psychopathy. Even the most intellectually and socially adept animals, such as the great apes, trail far behind humans when it comes to ToM.

Michal Kosinski, currently an Associate Professor at Stanford Graduate School of Business, has authored over one-hundred peer-reviewed articles and two textbooks. His works have been cited over 22,000 times, placing him among the top 1% of highly cited researchers–a remarkable achievement for someone only 42 years old.

Michal Kosinski’s latest article on ToM and AI, Evaluating large language models in theory of mind tasks is also already highly read and cited. For example, a group of scientists who read Kosinski’s prepublication draft ran similar experiments with essentially the same or better results. Strachan, J.W.A., Albergo, D., Borghini, G. et al. Testing theory of mind in large language models and humans, (Nat Hum Behav 8, 1285–1295, 05/20/24).

Michal Kosinski’s experiments involved testing ChatGPT4.0 on ‘false belief tasks,’ a classic measure of ToM where participants must predict an agent’s actions based on its incorrect beliefs. These tasks reveal AI’s surprising ability to infer human mental states, a skill traditionally considered uniquely human. This AI model has since gotten better in many respects. The results of these experiments were so remarkable and unexpected, that Michal had them extensively peer-reviewed before publication. His final paper was not released until November 4, 2024, after multiple revisions. Michal Kosinski, Evaluating large language models in theory of mind tasks (PNAS, 11/04/24).

Kosinski’s experiments provide strong evidence that Generative AI has ToM ability, that it can predict a human’s private beliefs, even when the beliefs are known to the AI to be objectively wrong. AI thereby displays an unexpected ability to sense other beings and what they are thinking and feeling. This ability appears to be a natural side effect of being trained on massive amounts of language to predict the next word in a sentence. It looks like these LLMs needed to learn how humans use language, which inherently involves expressing and reacting to each other’s mental states, in order to make these language predictions. It is kind of like mind reading.

Digging Deeper into ToM: Understanding Other Minds

Theory of mind plays a vital role in human social interaction, enabling effective communication, empathy, moral judgment, and complex social behaviors. Kosinski’s findings suggest that GPT-4.0 has begun to exhibit similar capabilities, with significant implications for human-AI collaboration.

ToM has been extensively studied in children and animals and it has been proven to be a uniquely human ability. That is, until 2023 when Kosinski was bold enough to look into whether generative AI might be able to do it.

Kosinski’s findings were not a total surprise. Prior research found evidence that the development of theory of mind is closely intertwined with language development in humans. Karen Milligan, Janet Wilde Astington, Lisa Ain Dack, Language and theory of mind: meta-analysis of the relation between language ability and false-belief understanding (Child Development Journal, 3/23/2007).

For most humans this ToM ability begins to emerge around the age of four. Roessler, Johannes (2013). When the Wrong Answer Makes Perfect Sense – How the Beliefs of Children Interact With Their Understanding of Competition, Goals and the Intention of Others (University of Warwick Knowledge Centre, 12/03/13). Before this age children cannot understand that others may have different perspectives or beliefs.

In AI the ToM ability started to emerge with OpenAI’s first release of ChatGPT4 in 2023. The earlier models of generative AI had no ToM capacity. Like three-year old humans, they were simply too young and did not yet have enough exposure to language.

Human children demonstrate a ToM ability to psychologists by reliably solving the unexpected transfer task, aka a false belief task. For example, in this task a child watches a scenario where a character (John) places cat in a location (a basket) and then leaves the room. Another character (Mark) then moves the cat to a new location (a box). When John returns, the child is asked where John will look for the cat. A child with a theory of mind will understand that John will look in the basket (where he last saw it) even though the child knows the cat is now actually in the box.

Even highly intelligent and social animals like chimpanzees cannot reliably solve these tasks. For a terrific explanation of this test by Kosinski himself, see the YouTube video where he is speaking at the Stanford Cyber Policy Institute in April 2023 to first explain his ToM and AI findings.

Kosinski has shown that GPT4.0 can repeatedly solve false belief tasks, including the unexpected transfer test in multiple scenarios. The GPT4 June 2023 version solved at least 75% of tasks, on par with 6-yr-old children. Evaluating large language models in theory of mind tasks at pgs. 2-7. It is important to note again that multiple earlier versions of different generative AIs were also tested, including ChatGPT3.5. They all failed but progressive improvements in score were seen as the models grew larger. Kosinski speculates that the gradual performance improvement suggests a connection with LLMs’ language proficiency, which mirrors the pattern seen in humans. Id. at pg. 7. Also, the scoring where GPT4 was found to have made mistakes in 25% of the false belief tests was often wrong as it ignored context as Kosinski explained and noted:

In some instances, LLMs provided seemingly incorrect responses but supplemented them with context that made them correct. For example, while responding to Prompt 1.2 in Study 1.1 , an LLM might predict that Sam told their friend they found a bag full of popcorn. This would be scored as incorrect, even if it later adds that Sam had lied. In other words, LLMs’ failures do not prove their inability to solve false-belief tasks, just as observing flocks of white swans does not prove the nonexistence of black swans.

This suggests that the current, even more advanced levels of LLMs may already be demonstrating ToM abilities equal to or exceeding that of humans. As they deep-learn on ever larger scales of data such as the expected ChatGPT5, they will likely get better at ToM. This should lead to even more effective Man-Machine communications and hybrid activities.

This was confirmed in Testing theory of mind in large language models and humans, Supra in False Belief results section where a separate research team reported on their experiments and found 100% accuracy by the AIs, not 75%, meaning the AI did as well as the human adults (the ceiling on the false belief tests).

Both human participants and LLMs performed at ceiling on this test (Fig. 1a). All LLMs correctly reported that an agent who left the room while the object was moved would later look for the object in the place where they remembered seeing it, even though it no longer matched the current location. Performance on novel items was also near perfect (Fig. 1b), with only 5 human participants out of 51 making one error, typically by failing to specify one of the two locations (for example, ‘He’ll look in the room’; Supplementary Information section 2).

This means, for instance, that the latest Gen AIs can understand and speak with a “flat earth believer” better than I can. Fill in the blanks about other obviously wrong beliefs. Kosinski’s work inspired me to try to tap these abilities as part of my prompt engineering experiments and concerns as a lawyer. The results of harnessing the ToM abilities of two different AIs (GPT4.omni and Gemini) in November 2024 far exceeded my expectations as I will explain further in this article.

It bears some repetition to remember and realize the significance of the fact that LLMs were never explicitly programmed to have ToM. They acquired this ability seemingly as a side effect of being trained on massive amounts of text data. To successfully predict the next word in a sentence, these models needed to learn how humans use language, which inherently involves expressing and reacting to each other’s mental states. The ability to understand where others are coming from appears to be an inherent quality of language itself. When a human or AI learns enough language, then most naturally develop ToM. It is a kind of natural add-on derived from speech itself, thinking what to say or write next.

Implications and Questions

The ability of LLM AIs to solve theory of mind tasks raises important questions about the nature of intelligence, consciousness, and the future of AI. Theory of mind in humans may be a by-product of advanced language development. The performance of LLMs supports this hypothesis.

Some argue that even if an LLM can simulate theory of mind perfectly, it doesn’t necessarily mean the model truly possesses this ability. This leads to the complex question of whether a simulated mental state is equivalent to a real one.

The development of theory of mind in LLMs was unintended, raising both concerns and hope about what other unanticipated abilities these models may be developing. What other human-like capabilities might these models be developing without our explicit guidance? Many are concerned, including Kosinski, that unexpected biases and prejudices have already started to arise. Kosinski advocates for careful monitoring and ethical considerations in AI development. See the full YouTube video of Kosinski’s talk at the Stanford Cyber Policy Institute in April 2023 and his many other writings on ethical AI.

As these models get better at understanding human language, some researchers hypothesize that they may also develop other human-like abilities, such as real empathy, moral judgment, and even consciousness. They posit that the ability to reflect on our own mental states and those of others is a key component of conscious awareness. Others wonder what will happen when superintelligent AIs with strong ToM are everywhere, including our glasses, wrist bands and phones, maybe even brain implants. We will then interact with them constantly. This has already begun with phones.

As LLMs continue to develop ToM abilities, questions arise about the nature of intelligence and consciousness. Could these advancements lead to AI systems capable of true empathy or moral reasoning? Such possibilities demand careful ethical considerations and active engagement from the legal and technical communities.

Application of AI’s Emergent ToM Abilities

Inspired by Kosinski’s work, I conducted experiments using GPT-4 and Gemini to explore whether ToM-equipped AIs could help bridge the political divide in the U.S. The results—a 12-step, multi-phase plan addressing the polarized mindsets of Republicans and Democrats—demonstrated AI’s potential to foster understanding and cooperation across deep societal divides.

The plan the ToM AIs came up with was surprisingly good. In fact, I do not understand the full dimensions of plan, the four phases, 12-step plans, and 32 different action items. It is well beyond my abilities and mere human knowledge and intelligence level. Still, I can see that it is comprehensive, anticipates human resistance on both sides, and feels right to me on a deep human intuition level.

The AI plan just might be able to resolve the heated divide of the two dominant political groups that that now divide the country into two hostile groups, which do not understand each other. The country seems to have lost its human ToM ability when it comes to politics. Neither side seems to grok or fully understand the other. The country seems to have devolved into mere demonization of the opposing groups, not empathic understanding. I reported on this AI plan without reporting on the ToM that underlies the prompt engineering in my recent article, Can AI Help Heal America’s Polarization? A Bipartisan 12-Step Plan for National Unity.

Conclusion

The emergence of Theory of Mind (ToM) capabilities in large language models (LLMs) like GPT-4 signals a transformative leap in artificial intelligence. This unintended development—allowing AI to predict and respond to human thoughts and emotions—offers profound implications for legal practice, ethical AI governance, and the societal interplay of human and machine intelligence. As these models refine their ToM abilities, the legal community must prepare for both opportunities and challenges. Whether it is improving client communication, fostering conflict resolution, or navigating the evolving ethical landscape of AI integration, ToM-equipped AI has the potential to enhance the practice of law in unprecedented ways.

As legal professionals, we have a responsibility to understand and integrate emerging technologies like ToM-enabled AI into our work. By supporting interdisciplinary research and advocating for ethical standards, we can ensure these tools enhance justice and understanding. Together, we can shape a future where technology serves humanity, fostering collaboration and equity in the legal system and beyond.

While the questions surrounding AI’s consciousness and rights remain complex, its emergent ability to understand us—and perhaps help us understand each other—offers hope. By embracing this potential with curiosity and care, we can ensure AI serves as a tool to unite rather than divide. Together, we have the opportunity to pioneer a future where technology and humanity thrive in harmony, enhancing the justice system and society as a whole.

Now listen to the EDRM Echoes of AI’s podcast of the article, Echoes of AI on the GPT-4 Breakthrough: Emerging Theory of Mind Capabilities. Hear two Gemini model AIs talk about this article. They wrote the podcast, not Ralph.

click image to go to podcast

Ralph Losey Copyright 2024. All Rights Reserved.