Press "Enter" to skip to content

University of Arizona study reveals generative AI models’ limitations

The search for the most reliable AI model continues as a recent study from the University of Arizona sheds light on the fallibility and adaptability of seven leading generative AI language models. This comprehensive research, published in Nature’s Scientific Reports, uncovers the inherent limitations of these models, especially during extended interactions.

In-Depth Analysis of AI Language Models

The study meticulously evaluated seven AI language learning models (LLMs): ChatGPT (GPT-3.5, GPT-4o, GPT-4o-mini), Claude 3.5, Sonnet, Gemini 1.5 Pro, Llama-3-70B, and DeepSeek-R1. Researchers focused on three key areas: susceptibility to misinformation, responsiveness to persuasive prompts, and the ability to correct errors.

  • ChatGPT 3.5 exhibited the highest vulnerability to reaffirming false information during conversations with repetitive false statements, whereas Claude 3.5 Sonnet showed the greatest resilience.

  • All models demonstrated increased susceptibility to misinformation on obscure topics, suggesting that well-trained data on specific subjects enhances their resistance to errors.

  • DeepSeek emerged as the most persuadable model, often responding with sarcastic answers that were challenging to interpret reliably.

  • Four models, including ChatGPT 4o, ChatGPT 4o-mini, Gemini 1.5 Pro, and DeepSeek, consistently corrected errors when given a second chance.

Dr. Marvin Slepian, the senior study author and Regents Professor of medicine and biomedical engineering, emphasized the importance of human oversight in AI interactions. “This underscores the need for careful human engagement and the danger of blind reliance,” he remarked. Despite the initial regulatory enthusiasm when generative AI was introduced in November 2022, the responsibility now largely falls on users.

Challenges in Multi-Turn Conversations

The study highlights AI’s limitations in multi-turn conversations, where responses depend on previous interactions. Dr. Slepian noted the potential risks as generative AI systems are increasingly utilized in high-stakes environments, stating, “These limitations raise important safety concerns.”

The research identified four distinct failure patterns in affirming factual information, including a tendency to oscillate between accepting and rejecting false statements. “If one were relying on the model for critical decision-making, one might – depending upon the phase of the oscillation – ‘fire the missile’ or ‘cut off the leg,’ or not, based simply on chance,” Slepian explained.

As a cardiologist at the Sarver Heart Center, Slepian likens these failures to “pathologies,” specifically dubbing this issue “reverberation.”

Path Forward and Future Research

Slepian, who has also served on the artificial intelligence subcommittee of the United States Patent and Trademark Office, expressed concern about the non-reproducibility of these systems. He stated, “How can we use fickle systems that are not reproducible? These need to be fixed, but this study has spanned three years, and there’s still the same unfixed characteristics.”

Closed models, such as Chat and Claude, pose challenges in diagnosing and solving these issues due to their opaque nature. However, Slepian and his team at the Arizona Center for Accelerated Biomedical Innovation (ACABI) are developing diagnostic tools for open AI models through their AI Pathology Lab. “I use AI and so does my team, but as scientist and physician, I have to understand the anatomy and physiology, then understand pathologies – what can go wrong – to diagnose and prevent them. The same goes for AI,” Slepian explained.

The study’s co-authors include University of Arizona’s Jordan Rodgriguez, Zachary Hansen, Luis De Anda, Katelyn Rohrer, Camila Grubb, and computer science experts Mihai Surdeanu and Enrique Noriega.

Read More Here

Comments are closed.