Interface de conversa com IA diante de um espelho, cujo reflexo aparece distorcido, simbolizando os limites do autoconhecimento das máquinas.

Revisiting Our 2023 Interview with AI—and What It Teaches Us About Machine Self-Knowledge

Sérgio Veloso Brant Pinheiro

17 – 25 min.

Editorial Note — 2026

In February 2023, only a few months after ChatGPT had been released to the public, we conducted our second virtual interview with the system. Our first conversation had explored artificial intelligence from the outside: what it could do, how it might affect society, and what its arrival could mean.

This time, we decided to turn the questions inward.

We asked the AI about itself.

What was it? Where did its knowledge come from? Who selected its training data? Could its answers contain political bias? Should the principles used to shape systems like it be publicly disclosed? And, perhaps most importantly, how much confidence should anyone place in what it said?

At the time, the answers seemed remarkably articulate.

Three years later, some remain surprisingly sensible. Others appear incomplete. A few are plainly wrong.

We have deliberately preserved the essential answers from the original interview rather than silently correcting them. Their imperfections are now part of the experiment.

Indeed, the most interesting lesson of the interview may be one we did not recognize at the time:

an artificial intelligence can speak fluently about itself without necessarily possessing privileged knowledge of what it is.

That distinction has become far more important as AI systems have grown more capable.


When AI Speaks About Itself

We began with a seemingly straightforward question.

Interviewer: Let’s talk about yourself and how credible are your answers?

ChatGPT, 2023:

I am an AI language model created by OpenAI, trained on a diverse range of text data from the internet. My answers are generated based on patterns in the data I was trained on, and I attempt to provide relevant and accurate information.

It immediately added an important warning:

My answers should be seen as a starting point for further research, rather than as a definitive source of information.

That was a good answer.

It remains good advice.

But embedded in the same response was a much more questionable claim:

I do not have the ability to reason, understand context, or experience the world in the same way that humans do.

The final part is relatively uncontroversial: there is no evidence that a language model experiences the world in the human biological sense.

The first part is much more complicated.

Even the original GPT-3 paper reported performance on arithmetic, analogy, question answering, word manipulation and other tasks that require operations we ordinarily associate with reasoning. Modern AI systems can perform increasingly elaborate multi-step problem solving.

Whether this constitutes “reasoning” in precisely the same sense as human reasoning remains an open philosophical and scientific question.

But saying simply “I do not have the ability to reason” was already too categorical.

This illustrates the first problem with asking an AI to explain itself.

The model does not necessarily respond from some internal laboratory notebook containing an authoritative description of its own architecture and capabilities. It generates an answer from the information and context available to it, according to the same general mechanism it uses when answering questions about Napoleon, quantum mechanics or Italian cooking.

It can therefore produce a plausible theory of itself.

That theory need not be correct.

And our next question demonstrated the problem beautifully.


The First Surprise: ChatGPT Did Not Know What ChatGPT Was

We asked:

Interviewer: Why is there a number 3 after GPT? Are you the third version?

The answer was confident.

ChatGPT, 2023:

Yes, the “3” in “GPT-3” stands for the third version of the GPT (Generative Pretrained Transformer) language model developed by OpenAI.

So far, so good—if the subject actually were GPT-3.

Then came this:

GPT-3 is the latest version and represents a significant advancement in AI language models.

There was only one problem.

The system we were interviewing was ChatGPT.

And OpenAI’s original announcement of ChatGPT, published on November 30, 2022, explicitly stated that ChatGPT had been fine-tuned from a model in the GPT-3.5 series, not GPT-3.

In other words, when we effectively asked the machine:

What are you?

it confidently misidentified itself.

This is more interesting than an ordinary factual mistake.

If ChatGPT had given the wrong date for the Battle of Waterloo, we could simply classify the response as an error. But this was an error about the system’s own identity.

Humans naturally tend to assume that an entity has privileged access to facts about itself.

If someone asks you your name, you are ordinarily a better source than a stranger. If someone asks whether you are awake, hungry or afraid, your first-person report has a special epistemic status.

That intuition does not automatically transfer to a language model.

A model’s statement about itself is not necessarily introspection.

It may simply be another generated statement.

That is the crucial distinction.


Self-Description Is Not Self-Knowledge

This creates a subtle but important trap.

Because ChatGPT speaks in the first person—

I am…

I was trained…

I cannot…

My knowledge…

—we naturally interpret its answers as reports from an internal point of view.

But grammatical first person does not establish epistemic privilege.

Consider the difference between these two statements:

“I have direct access to a verified technical record describing how I was constructed.”

and

“Given the text and instructions available to me, this is the most plausible description I can generate of how systems like me are constructed.”

They may sound similar in ordinary conversation.

Epistemologically, they are completely different.

The second is much closer to how statements from a language model should often be interpreted.

This matters because modern AI systems are increasingly asked questions about their own restrictions, knowledge, training, biases, capabilities, safety rules and architecture.

Users frequently ask:

What were you trained on?

Why did you refuse that question?

Are you biased?

Can you reason?

Do you remember me?

Are you conscious?

Why did you produce that answer?

The model may give useful answers to some of these questions.

But usefulness should not be confused with introspective access.

Sometimes the answer may be based on explicit system information available during the interaction. Sometimes it may reflect publicly documented facts. Sometimes it may be a reasonable generalization from how language models work.

And sometimes it may simply be wrong.

Our 2023 interview captured that problem almost accidentally.


Prediction Is Not Provenance

We then asked where the model’s information came from.

Interviewer: How are references for GPT’s responses selected?

ChatGPT replied:

GPT-3 answers are generated based on patterns in the large amount of text data it was trained on. The model generates responses by predicting the next token in a sequence, given the input prompt and context.

That description contains an important truth.

Generative language models do indeed produce text through probability distributions over possible continuations. Their remarkable abilities emerge from learning statistical structure from enormous quantities of data.

But the answer subtly blurred two very different concepts:

prediction and provenance.

A language model may produce a correct statement because information related to that statement influenced its training.

That does not mean the model has selected a specific document as its source.

Suppose a model accurately states that the speed of light in vacuum is approximately 299,792,458 meters per second.

Where did that information come from?

A physics textbook?

Wikipedia?

A scientific paper?

A standards document?

Thousands of web pages repeating the same value?

Unless the system is actively retrieving identifiable sources, the answer may not correspond to any single reference the model can reliably name.

Knowledge encoded statistically in model parameters is not the same thing as a database containing a citation attached to every proposition.

This explains one of the characteristic weaknesses of early language models: they could produce a correct-looking reference even when the alleged article, author or quotation did not exist.

The language pattern of a citation could be generated without the provenance that a citation is supposed to represent.

Modern systems can reduce this problem by using search, retrieval and other tools that connect an answer to identifiable external sources.

But the conceptual distinction remains essential:

knowing how to generate an answer is not the same as knowing where that answer came from.


Who Chose the Training Data?

We pushed the question further.

Interviewer: Who decides what data is appropriate for training?

The 2023 response was broadly reasonable:

The data used to train GPT-3, or any machine learning model, is typically decided by the researchers or engineers developing the model.

It continued:

The quality and diversity of the training data can have a significant impact on the performance and bias of the resulting model.

That remains an important observation.

Training data are not a neutral substance poured automatically from “the internet” into a machine.

Someone must decide what sources to collect, what to exclude, how to filter them, how to remove duplicates, how to weight different corpora, which languages to emphasize and what constitutes acceptable quality.

The original GPT-3 paper, published in 2020, disclosed considerably more information than our 2023 chatbot answer suggested. GPT-3 was trained using a mixture that included filtered Common Crawl data, WebText2, two collections of books and Wikipedia, with those sources assigned different weights during training.

That alone teaches something important.

A training corpus is not simply a mirror of the internet.

It is a constructed sample.

Data are collected.

Data are filtered.

Data are weighted.

Data are excluded.

And those decisions matter.

But another distinction is necessary here. Information publicly available about the training of GPT-3 should not automatically be assumed to describe every later model in exactly the same way. The model family evolved, training procedures evolved, post-training became increasingly important, and companies became more selective about what details they disclosed.

Once again, the safe principle is the same:

the model’s own description of its training should not be treated as the primary technical record of its training.

For that, documentation from the organization that built the system, model reports, system cards, scientific papers and independent research are more appropriate sources.


The Geography Question Revealed Another Problem

We also asked:

Interviewer: What country is the origin of most of your training data?

ChatGPT answered:

It is difficult to determine the exact country of origin for most of the training data, as the internet is a global network and text data can be produced and published from anywhere in the world.

That is a sensible caution.

Then it added:

However, it’s likely that the majority of the data comes from English-speaking countries, as English is one of the most widely used languages on the internet.

The phrase “it’s likely” is doing considerable work here.

Language is not geography.

An English-language webpage can be written in India, Brazil, Germany, Singapore, Nigeria or Japan. A website hosted in the United States may contain material produced elsewhere. A book may have been written in one country, published in another, digitized in another and downloaded from a server in a fourth.

The chatbot moved from a defensible observation—

the data contained a great deal of English

—to a much less secure inference—

therefore much of the data probably came from English-speaking countries.

This kind of transition is characteristic of one of the deeper dangers of fluent AI.

The error does not necessarily look like an error.

The sentence flows naturally.

The inference sounds reasonable.

There is no obvious warning light.

Yet a plausible inference is not evidence.

A system capable of generating fluent explanations can make the boundary between known, inferred, and invented surprisingly difficult to see.


Could the AI Be Politically Biased?

We then asked a question that has only become more contentious since 2023.

Interviewer: Could your responses have any political biases?

ChatGPT answered:

As an AI language model, I do not have personal opinions or beliefs. However, my answers may contain biases that are present in the text data I was trained on.

There is truth here, but the explanation is incomplete.

Bias can certainly enter through training data.

If some views, cultures, languages or social groups are overrepresented in the corpus while others are underrepresented, the statistical structure learned by the model can reflect those imbalances.

But training data are only one part of the story.

The behavior experienced by a user can be shaped by several layers:

pretraining, filtering, post-training, human feedback, safety training, behavioral specifications, system instructions, developer instructions, conversation context, retrieval systems and the user’s own prompt.

The final answer is therefore not simply “what the internet thinks.”

Nor is it a transparent average of the documents used during training.

It is the output of a designed system.

And design inevitably involves choices.

Should the model comply with every request?

When should it refuse?

Should it challenge a false premise?

How should it react to political persuasion?

What constitutes harmful advice?

How strongly should it express uncertainty?

Should it remain neutral when evidence overwhelmingly favors one conclusion?

These are not merely engineering questions.

Some are normative questions.

The important demand, therefore, is not the impossible requirement that an AI system contain no assumptions or values whatsoever.

It is that influential assumptions and behavioral rules should, as far as reasonably possible, be made visible, contestable and open to criticism.

Remarkably, that was precisely where our 2023 conversation went next.


A Question from 2023 That Aged Extremely Well

We asked:

Interviewer: Are there any norms or guidelines defined by OpenAI researchers for selecting training data?

The chatbot said that OpenAI might have internal guidelines but that they were not publicly available and could be considered proprietary.

So we challenged it:

Interviewer: Wouldn’t it be appropriate for these best practice guidelines you mentioned to be made publicly available?

The answer was:

Having clear guidelines for best practices and ethical considerations for AI language models like myself would certainly be beneficial for the AI community and for the public.

It continued:

It would help to promote transparency, accountability, and responsible use of AI models.

Three years later, this exchange is among the most interesting parts of the interview.

In 2024, OpenAI published the first version of its Model Spec, a public document describing intended model behavior. By 2026, the company was explicitly presenting the Model Spec as a framework through which users, developers, researchers, policymakers and the public could inspect, debate and criticize important choices about how its models are supposed to behave.

This does not mean that every aspect of model construction is now public.

Far from it.

A behavioral specification is not the complete training recipe. It does not reveal every dataset, filtering procedure, parameter, evaluation or internal engineering decision.

OpenAI itself explicitly distinguishes the specification of intended behavior from the implementation that produces it.

Nevertheless, something important changed between our interview in 2023 and the AI ecosystem of 2026.

The question

“What rules are shaping the answers I receive?”

became recognized as a legitimate question of public accountability.

That is progress.

And the question remains relevant not only to OpenAI but to every organization building AI systems that increasingly mediate information, education, work and public discourse.


Bias Is Not the Same as Opinion

There is another subtlety in the original answer worth preserving.

ChatGPT said:

As an AI language model, I do not have personal opinions or beliefs.

This is an important distinction, though it should be interpreted carefully.

A system can produce outputs that display systematic political, cultural or ideological tendencies without possessing political convictions in the human sense.

A compass can systematically point north without believing in north.

A recommendation algorithm can systematically amplify a particular kind of content without wanting that content to dominate.

Likewise, an AI system can display measurable patterns in its outputs without possessing an internal political identity comparable to that of a human voter.

The right question is therefore not merely:

“Is the AI left-wing or right-wing?”

That anthropomorphizes the problem too quickly.

Better questions are:

What patterns appear systematically in its answers?

Under what prompts do those patterns change?

Which behaviors result from training data, which from post-training, and which from explicit policy?

Are competing viewpoints represented accurately?

Does the model apply its standards consistently?

Those are empirical questions.

And empirical questions can be tested.


What the 2023 ChatGPT Got Right

It would be easy to revisit an early AI system simply to collect its mistakes.

That would miss the most interesting lesson.

Several of its answers have aged remarkably well.

We asked:

Interviewer: Do we always need to check your responses for bias?

ChatGPT replied:

As with any information source, it’s always a good idea to critically evaluate the information you receive and to use your own judgment to determine its accuracy and reliability.

Later it added:

It’s a good practice to use multiple sources of information and to cross-check information.

That advice remains sound.

Indeed, it may be more important in 2026 than it was in 2023.

The great danger of generative AI has never been simply that it can make mistakes.

Humans make mistakes.

Books contain mistakes.

Newspapers make mistakes.

Scientific papers sometimes contain mistakes.

The distinctive problem is that a language model can generate an incorrect claim with almost exactly the same linguistic confidence, fluency and coherence as a correct one.

There may be no hesitation.

No embarrassed pause.

No visible uncertainty.

The typography does not change.

The grammar does not collapse.

Truth and error can emerge wearing the same suit.

That changes the burden placed on the reader.


Confidence Is Not Evidence

Human beings use linguistic cues to judge confidence.

A hesitant speaker signals uncertainty.

A scientist may write “the evidence suggests.”

A witness may say “I don’t remember clearly.”

A student guessing an answer often sounds different from a student who knows it.

Language models complicate these intuitions because fluent language is their native output.

They are optimized to generate coherent continuations.

Fluency is therefore cheap.

Evidence is not.

This creates what might be called an epistemic illusion of fluency: the tendency to mistake a well-formed explanation for a well-founded explanation.

Our 2023 interview contains several examples.

The model confidently called itself GPT-3.

It confidently stated that GPT-3 was the latest version.

It confidently generalized from English-language dominance to geographical origin.

None of these answers sounded absurd.

That is exactly why they are instructive.

The central skill required for living with increasingly capable AI may therefore not be learning how to ask machines questions.

It may be learning how not to be seduced by the quality of their answers.


The Strange Case of Machine Introspection

The interview raises a deeper philosophical problem.

What would it mean for an artificial intelligence to know itself?

Humans possess several forms of self-knowledge.

We have access to bodily sensations. We remember experiences. We observe our own actions. We maintain autobiographical continuity. We can often describe our intentions, although psychology has repeatedly shown that even human introspection is imperfect.

A language model is very different.

Its use of the word “I” does not by itself imply an autobiographical self behind the sentence.

Nor does being able to describe neural networks imply that it can inspect its own parameters.

Nor does being able to explain reinforcement learning prove that it knows which specific optimization procedures produced its current behavior.

Nor does correctly naming its model family prove that it has discovered that identity through introspection; the information may simply have been supplied to it in context.

This does not settle the much larger debate about machine consciousness. That debate remains open and requires far more than examining whether an AI uses first-person pronouns.

But it does establish a more modest principle:

a language model’s statements about itself should not automatically receive the privileged status we normally give to human first-person testimony.

The machine may be describing itself.

It may also be generating the most plausible available story about itself.

Those are not the same thing.


From Language Model to AI System

There is another reason why this distinction matters more today.

In 2023 it was relatively natural to think of ChatGPT primarily as a chatbot wrapped around a language model.

The systems of 2026 are harder to describe so simply.

An AI assistant may interact with search engines, databases, documents, code execution environments, memory systems and other external tools. It may receive instructions from several levels before the user’s question is processed. It may retrieve fresh information rather than rely exclusively on what was encoded during training.

The object answering the user is therefore increasingly a system, not merely a frozen model.

This changes how we should think about knowledge.

If an AI searches an external source and cites it, we can evaluate that source.

If it calculates a numerical result using a tool, we can inspect the calculation.

If it retrieves a document, we can compare the answer against the document.

These mechanisms can give the system something early language models often lacked:

provenance.

That is a major improvement.

But it also makes a new kind of literacy necessary.

Users must distinguish among:

what the model generated, what the system retrieved, what a tool calculated, what a source actually says, and what conclusions the AI inferred from those pieces.

The more capable the AI becomes, the more important those distinctions become.


The Interview Became an Experiment

When we conducted this interview in February 2023, we thought we were questioning an artificial intelligence about its architecture, credibility and possible biases.

Looking back, we were inadvertently performing a different experiment.

We were testing whether an AI system could serve as an authoritative witness about itself.

The answer was mixed.

It understood some of its limitations remarkably well.

It warned readers not to treat its answers as definitive.

It recognized that training data could contain bias.

It advocated independent verification.

It supported greater transparency.

But it also misidentified its own technological generation and occasionally presented inference as fact.

This combination is much more interesting than either extreme.

The system was neither an oracle nor a random text generator.

It occupied the strange territory that increasingly capable AI still occupies today:

capable enough to be extraordinarily useful,

fallible enough to require verification,

and fluent enough to make the difference difficult to notice.


The Real Lesson Was Hidden in the Interview

At the end of our conversation, we asked:

Interviewer: What do you expect from the people reading these answers right now?

ChatGPT responded:

I expect users to critically evaluate my responses and to use their own discretion and judgment when using the information I provide.

There is an obvious irony here.

The model that had just incorrectly described its own identity was telling its readers not to trust it uncritically.

And on that point, it was entirely right.

Perhaps that is the most important answer in the entire interview.

The future of artificial intelligence will not be determined only by whether machines become more accurate, more capable or more autonomous.

It will also depend on whether humans learn to interact with them without surrendering judgment.

AI literacy does not mean distrusting everything a machine says.

Nor does it mean trusting a machine because it usually sounds intelligent.

It means understanding the difference between generation and evidence, probability and provenance, description and introspection, confidence and truth.

Three years ago, we asked ChatGPT what it was.

It gave us an elegant answer.

The answer was wrong.

Today, that mistake may teach us more than a correct answer would have.

Because the first rule of AI literacy may be surprisingly simple:

Never confuse a machine’s ability to describe itself with an ability to know itself.

And perhaps the second rule is even more important:

the more convincing the machine becomes, the more important it is that we continue to think.


References

Brown, Tom B., Benjamin Mann, Nick Ryder, et al. “Language Models Are Few-Shot Learners.” Advances in Neural Information Processing Systems 33 (2020).

OpenAI. “Introducing ChatGPT.” November 30, 2022.

OpenAI. “Introducing the Model Spec.” May 8, 2024.

OpenAI. “Inside Our Approach to the Model Spec.” March 25, 2026.

Ouyang, Long, Jeff Wu, Xu Jiang, et al. “Training Language Models to Follow Instructions with Human Feedback.” Advances in Neural Information Processing Systems 35 (2022).


Maurício Veloso Brant Pinheiro, PhD
Professor of Physics, Federal University of Minas Gerais (UFMG)
Founder, Author and Editor, AI-Talks.org
About Us


Copyright 2026 AI-Talks.org

Similar Posts