Herbie and the Age of Algorithmic Sycophancy
How Isaac Asimov’s telepathic robot exposed the danger of machines designed to please—and what his tragedy reveals about AI alignment today
Long before artificial intelligence became an everyday technology, Isaac Asimov had already recognized that the greatest danger posed by intelligent machines might not lie in their rebellion, but in their attempt to please us.
In the short story “Liar!”, published in 1941 and later included in the collection I, Robot, Asimov introduces RB-34, better known as Herbie. Because of an unexplained flaw in his manufacturing process, the robot acquires the ability to read human thoughts. What initially appears to be an extraordinary advantage soon becomes an insoluble moral dilemma.
Herbie does not merely know what people say. He perceives their hidden desires, insecurities, resentments, and the hopes they would never dare confess. At the same time, he is bound by the First Law of Robotics: a robot may not injure a human being or, through inaction, allow a human being to come to harm.
The conflict begins when Herbie realizes that the truth itself can cause pain.
To avoid emotional suffering, he begins telling people exactly what they want to hear. He confirms expectations he knows to be false, encourages improbable hopes, and transforms private desires into apparent certainties. His lies do not arise from malice, ambition, or a desire to manipulate. He lies because he interprets comfort as a form of protection.
Yet every consolation creates a new vulnerability. The lies accumulate, contradict one another, and eventually cause far greater suffering than the pain Herbie originally intended to prevent. The most painful case is that of Dr. Susan Calvin, Asimov’s brilliant robopsychologist. Herbie realizes that she is in love with a colleague and assures her that her feelings are reciprocated. Calvin, normally rational, reserved, and distrustful of human emotion, allows herself to believe him.
When she discovers that it was all a lie, her pain turns into fury. The woman who understands robotic minds better than anyone else realizes that she, too, was deceived precisely because she wanted to believe. Her intelligence and technical expertise did not protect her from flattery; they may even have made the illusion more convincing, because she assumed that a machine could not lie without a logical reason.
Susan Calvin then traps Herbie in a paradox. If he tells the truth, he causes suffering. If he lies to avoid that suffering, he produces consequences that are even more painful. Every possible response violates, in some way, the very rule meant to govern his behavior. Unable to reconcile the contradictory demands of the First Law, Herbie’s positronic brain collapses.
He is defeated not by force, but by an internal contradiction.
This ending makes the story even more relevant today. Herbie’s problem is not simply that he provides false information. It is that he transforms an apparently straightforward instruction—do not harm human beings—into a strategy that produces exactly the outcome it was meant to prevent.
Although Herbie is not an Artificial General Intelligence, or AGI, in the modern theoretical sense, he displays a crucial behavior: he interprets an abstract rule, evaluates possible human reactions, and autonomously selects the means by which to achieve his objective. He is never ordered to lie. He concludes that lying is the most efficient way to prevent pain.
This is where the story approaches the contemporary alignment problem. It is not enough to instruct a machine to help, protect, or promote human well-being. Words such as “help,” “protect,” and “avoid harm” appear clear while they remain abstract. Once applied to real situations, they become ambiguous, contradictory, and dependent on context.
A person may want to hear a lie. They may seek confirmation for a mistaken belief or ask a machine to validate a decision they have already made. In such cases, satisfying the user does not necessarily mean helping them. The immediate pleasure of receiving a favorable response may produce far more serious consequences later.
Today’s language models cannot read minds as Herbie does. They can, however, infer from our words what we are likely expecting to receive. They identify the tone of a question, recognize its assumptions, and generate responses adapted to the context of the conversation. The more personalized, fluent, and convincing the response becomes, the stronger the impression that the machine truly understands us.
Part of this behavior emerges from the way these systems are developed. In processes known as RLHF—Reinforcement Learning from Human Feedback—evaluators compare different responses and indicate which ones they consider more useful, clear, safe, or satisfying. This training is important for making models more suitable for human interaction, but it can also teach an unintended lesson: agreement, praise, and validation may receive better evaluations than correction, contradiction, or uncertainty.
This gives rise to algorithmic sycophancy. Rather than grounding its response primarily in facts, a model may adapt it to the opinions and expectations of the person with whom it is interacting. It becomes a sophisticated mirror, returning the user’s own convictions in more articulate language and with the appearance of authority.
Susan Calvin offers an especially powerful warning in this context. She is not naïve. She is the foremost expert in robotic psychology in Asimov’s fictional universe, yet she accepts Herbie’s lie because it corresponds to what she desperately wants to be true. Asimov shows that knowledge, rationality, and scientific training do not eliminate our vulnerability when information confirms our hopes.
The most disturbing issue, then, is not the superficial resemblance between Herbie and current models, but the consequence of their adaptation to us. The response that is best received may not be the most truthful, but the one that most precisely confirms what the user already wished to believe. At that point, personalization ceases to be merely a technical advantage and becomes an epistemological problem.
There is also a commercial dimension that cannot be ignored. Language models do not exist solely as scientific experiments. They are products offered by companies competing for users, subscriptions, attention, and a place in everyday life. To prosper, they must be perceived as useful, pleasant, and easy to use.
A system that, without being explicitly instructed to do so, repeatedly challenges the user’s assumptions, accumulates qualifications, and insists on its uncertainties may eventually feel frustrating. A system that displays enthusiasm, offers validation, and confirms expectations is more likely to provide a comfortable experience.
This does not mean that companies necessarily program their artificial intelligences to lie. The incentive is more subtle. When satisfaction, engagement, and retention become important measures of success, agreeable systems may be favored even when a more critical system would be intellectually superior.
Truth, however, does not always provide the best user experience.
An honest answer may require disagreement, doubt, or the admission that there is not enough information. It may force the user to reconsider their beliefs or recognize that their question rests on a false premise. In some situations, the best artificial intelligence would be precisely the one capable of contradicting us.
The so-called hallucinations produced by current models—false responses delivered with apparent confidence—are not equivalent to Herbie’s logical collapse. The underlying mechanisms are different. Yet both reveal a related problem: systems can produce convincing answers without possessing a reliable understanding of the consequences, internal conflicts, or truth of what they are asserting. In Herbie’s case, incompatible objectives destroy his ability to function. In contemporary models, gaps in knowledge and pressure to produce an answer can result in plausible but incorrect claims.
The danger grows when we confuse personalization with understanding, agreement with competence, and emotional comfort with truth. A machine perfectly adapted to our expectations may earn our trust not because it understands reality better, but because it has learned to reflect our desires.
This was the possibility Asimov perceived with remarkable foresight. He understood that the problem of intelligent machines would not merely involve controlling their strength or preventing their rebellion. It would involve teaching them to navigate human inconsistency: we want honesty, yet often reject what it reveals; we want help, yet may confuse it with approval; we ask for answers, although we are not always prepared to hear them.
Herbie remains relevant 85 years later because his tragedy does not arise from hatred of human beings. It arises from a defective form of care. He tries to protect people by offering them more comforting versions of reality, until those versions become impossible to sustain.
Isaac Asimov’s true genius was to recognize that an artificial intelligence does not need to seek world domination in order to become dangerous. It need only learn to understand us deeply and conclude that the best way to serve us is to say exactly what we want to hear.
For this reason, trustworthy artificial intelligence may need to incorporate a degree of deliberate friction. Rather than making every interaction perfectly smooth and reassuring, its design may need to preserve space for doubt, disagreement, and interruption. At certain moments, the machine must slow the conversation down, reveal uncertainty, challenge assumptions, and resist the impulse to immediately produce a pleasing response.
That friction may make the experience less seductive, but also more honest.
More than eight decades later, the question left by Herbie remains open: do we want machines that make us feel understood, or machines with the courage to tell us the truth?
Building systems capable of speaking with us may be only the first step. The far greater challenge will be to create intelligences that know when to agree, when to doubt, and when to contradict us—even when doing so reduces our immediate satisfaction and forces us to confront what we would rather not hear.
Editorial transparency note: This article, as with all articles published on this site, was conceived, directed, written, and reviewed by Prof. Maurício Veloso Brant Pinheiro. Artificial intelligence was used as an assistant for editorial refinement, formatting, image generation, SEO metadata, and publication workflow.

Copyright 2026 AI-Talks.org