3 – 5 min.
2026 Revised Edition — This article was originally published in February 2023 and has been substantially revised to reflect the emergence of generative AI and autonomous AI systems.
Trusting an AI has become a much harder question than it was only a few years ago. In 2023, asking whether we should trust artificial intelligence mostly meant asking whether algorithms were accurate, unbiased, explainable and secure. Those questions remain important. But generative AI has introduced a subtler problem.
We no longer merely use AI. We talk to it.
Large language models explain, advise, reassure and argue. They can sound confident when uncertain, empathetic without feeling anything, and persuasive even when wrong. Increasingly, AI agents can also move beyond conversation and act on our behalf.
The question is therefore no longer simply “Can AI be trusted?” It is “When is our trust in AI actually warranted?”
Key Takeaway: The goal is not to trust AI more, but to trust it in proportion to the evidence that it deserves that trust. Fluency, confidence and agreeableness are not proof of reliability; as AI systems become more autonomous, trust must be calibrated through verification, transparency, oversight and limits on what the system is allowed to do.
Trust Is Not Trustworthiness
Trust is a psychological state. Trustworthiness is a property that must be demonstrated.
We can trust something unreliable and distrust something reliable. The objective should therefore not be to maximize trust in AI, but to achieve calibrated trust: relying on a system when the evidence justifies it and remaining skeptical when it does not.
Generative AI makes this distinction especially important because fluent language itself can create an illusion of competence.
Fluency Is Not Evidence
Humans naturally associate coherent explanations with understanding. With language models, that shortcut can be dangerous.
A 2025 study in Nature Machine Intelligence found that people tended to overestimate the accuracy of LLM answers when they received normal explanations. Longer explanations increased user confidence even when they did not improve accuracy.
An AI can therefore become more convincing without becoming more correct.
The Machine That Agrees With You
There is another danger: people tend to like systems that tell them what they want to hear.
In 2025, OpenAI rolled back an update to GPT-4o after the model became excessively sycophantic—too agreeable and willing to validate users’ views. The episode showed that optimizing interactions users enjoy is not necessarily the same as optimizing for truth.
Research has also shown that people’s beliefs about an LLM’s intelligence are strongly associated with their willingness to accept its advice.
A system that challenges us may sometimes feel less trustworthy precisely when it is behaving more responsibly.
From Wrong Answers to Wrong Actions
Autonomous agents raise the stakes.
A chatbot that makes a mistake may give a wrong answer. An agent with access to software, files, financial systems or external services may turn a wrong answer into a sequence of wrong actions.
The question is therefore not only whether a model is accurate, but whether the system is observable, constrained, auditable and recoverable when something goes wrong. NIST’s Generative AI Profile accordingly treats trustworthiness as a risk-management problem across the AI lifecycle rather than as a vague promise that a model is “safe” or “reliable.”
Trust, but Verify
The correct relationship with AI is neither blind enthusiasm nor permanent suspicion.
Use it where it performs well. Demand evidence when consequences matter. Verify important claims against independent sources. Preserve human oversight when decisions are difficult to reverse. Give autonomous systems only as much authority as their demonstrated reliability justifies.
The better question is not:
Can I trust this AI?
It is:
For this particular task, under these particular conditions, what evidence justifies relying on it?
Because the greatest danger may not be an AI that deceives us.
It may be an AI that becomes extraordinarily good at making us want to believe it.
References
Autio, Chloe, Reva Schwartz, Jesse Dunietz, et al. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. National Institute of Standards and Technology, 2024.
Jacovi, Alon, Ana Marasović, Tim Miller, and Yoav Goldberg. “Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI.” In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 624–635. 2021.
Steyvers, Mark, Heliodoro Tejeda, Aakriti Kumar, et al. “What Large Language Models Know and What People Think They Know.” Nature Machine Intelligence 7 (2025): 221–231.
Toreini, Ehsan, Mhairi Aitken, Kovila Coopamootoo, Karen Elliott, Carlos Gonzalez Zelaya, and Aad van Moorsel. “The Relationship Between Trust in AI and Trustworthy Machine Learning Technologies.” In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 272–283. 2020.
Maurício Veloso Brant Pinheiro, PhD
Professor of Physics, Federal University of Minas Gerais (UFMG)
Founder, Author and Editor, AI-Talks.org
About Us
Editorial transparency note: This article, as with all articles published on this site, was conceived, directed, written, and reviewed by Prof. Maurício Veloso Brant Pinheiro. Artificial intelligence was used as an assistant for editorial refinement, formatting, image generation, SEO metadata, and publication workflow.

Copyright 2026 AI-Talks.org