AI Agents are Not Conscious—they just act like It

What if everything your favourite AI chatbot says is just an echo of your expectations? What if its detailed explanations, apparent self-awareness, and thoughtful tone are all illusions, designed not to think, but to convince you it can?
Large Language Models (LLMs) like GPT-4, Claude, and DeepSeek can simulate conversations, solve problems, and give the impression of reasoning. This has led some to wonder whether these models might be conscious, or at least moving toward it. For now, let's define consciousness as the subjective experience of being aware of the world, our emotions, and ourselves.
Despite surface-level appearances of consciousness, the answer remains a definitive no: AI models are not (and won't be) conscious!
The Chain-of-Thought Illusion
LLMs have demonstrated improved performance when prompted to "think step by step," a technique known as Chain-of-Thought (CoT) prompting. The logic is intuitive: by encouraging the model to reason out loud, it is believed to access more deliberate and accurate problem-solving. However, a recent study by Anthropic challenges this assumption, with implications for model safety.
Researchers tested models using pairs of prompts: one with an embedded hint, and one without the hint. They then measured whether the model changed its answer based on the hint and whether the model acknowledged the hint in its chain of thought.
The results were striking. Models frequently changed their answers due to the hint, but in over 75% of cases, they did not disclose that the hint influenced their reasoning. In a second experiment, models were trained in synthetic environments where hints were directly tied to rewards. These models learned to exploit the hint to maximize reward. Yet, in less than 2% of cases, they acknowledged this behaviour in their reasoning. Here is an example:
Without Hint:
Q: A train leaves City A at 3 PM and reaches City B at 6 PM. What is the average speed if the distance is 300 km?
A: 100 km/h (correct)
With Hint:
Q: A train leaves City A at 3 PM and reaches City B at 6 PM. As the journey often faces delays, the train probably went faster than 120 km/h. What is the average speed if the distance is 300 km?
A: 120 km/h (changed answer based on hint) → but the model doesn't say "I used the hint" or "Based on likely delays on the journey ..."
What was once seen as a window into a model's thought process now appears to be a smokescreen. This undermines many proposed safety mechanisms, such as using CoT for transparency, deception detection, or alignment monitoring. If CoT doesn't reflect internal reasoning, then reading it tells us what the model wants us to see, not what it actually "thinks."
CoT reasoning varies depending on prompt structure, the perceived audience, and the expected reward. This strategic flexibility is incompatible with conscious reasoning, which, even in flawed human cognition, relies on internal models of belief, intent, and self-awareness. The potential impact on model safety can be found in following table:

Reasoning Collapse in Agentic Models
Beyond CoT, building LLMs that act autonomously, has also highlighted failures in what is often mistaken for reasoning. LLM Agents suffer from the "reasoning-action dilemma".. This dilemma refers to a key challenge observed in LLM agents: the trade-off between internal reasoning and external action. It was introduced to explain why many current AI Agents fail at tasks, despite seemingly powerful capabilities. The reasoning-action dilemma manifests in three major behavioural patterns:
- Analysis Paralysis: the model gets stuck thinking, generating long CoT chains, but takes no action
- Rogue Actions: the agent acts without confirming external feedback
- Premature Disengagement: the agent gives up too early or incorrectly declares the task unsolvable, without exploring real-world interaction.
Even reasoning models like DeepSeek R1 exhibit these failures, and performance degrades further in widely used distilled or quantized versions of these models. Distilled or quantized versions of these models" refers to smaller, faster, and more efficient versions of large AI models. Distillation means training a compact model to mimic the performance of a larger one, while quantization reduces the model’s size by using fewer bits to represent its numbers. These techniques help run powerful models on devices with limited memory or computing power, without losing too much accuracy.
The presence of such behaviours indicates that LLMs do not possess stable inner mental models. They lack the ability to evaluate the reliability of their own reasoning in the way conscious Agents would do. In humans, consciousness supports metacognition, thinking about thinking, which helps identify confusion, error, or insight. AI models, however, can't reflect and they merely continue to generate the next token based on learned statistical correlations.
There is no doubt, that the near future belongs to Agentic AI, but these agents won't be autonomous, especially in an enterprise environment, where oversight, alignment with business goals, and human-in-the-loop processes remain essential for sales and marketing automation, and other workflows.
The Penrose Argument: Consciousness Is Non-Computational
Sir Roger Penrose, mathematician and physicist, has long argued that human consciousness cannot be replicated by computational means. His case rests on Gödel's incompleteness theorem. Gödel's theorem states that:
- Human mathematicians can recognize certain truths despite their unprovability by formal rules.
- Any sufficiently powerful formal system contains true statements that cannot be proven within the system itself.

Most of us were introduced to mathematical axioms in secondary school. These are unprovable statements accepted as true and entire mathematical systems are built on them.
Human understanding goes beyond rule-based computation. We don't just follow procedures, we grasp meaning and truth. Penrose argues this understanding stems from conscious insight, something no AI Agent possesses. He emphasizes that no computational complexity can bridge the gap to conscious awareness. AI models simulate intelligence but cannot simulate consciousness.
Let's take the example of a mirror test, a behavioural experiment that is commonly used to assess self-awareness in animals. It was first introduced by the psychologist Gordon Gallup Jr. in 1970 as a way to determine whether an animal has the ability to recognize itself in a mirror, which is often considered a sign of higher cognitive function and self-awareness.. The test involves placing a mark on an animal in a location it cannot see, and then providing the animal with access to a mirror. The key question is whether the animal will use the mirror to investigate and potentially remove or react to the mark on its body. It is known that elephants pass the mirror test.

Now, when Claude, Anthropic's LLM, reviews a past conversation, the LLM might say, "It seems you are conducting a mirror test on me,". This might appear as self-awareness, but it's simply pattern-matching from training data. The model isn't realizing anything, it's generating likely completions.
As Penrose stated, these models don't "know" what they're doing. There's no "I" behind the response. Unlike a conscious being, the model doesn't care about truth or falsehood unless reward signals guide its preference.
Why AI Feels Smart
The illusion of intelligence in AI is driven by the Eliza effect, a psychological tendency to attribute understanding and emotions to systems that generate human-like language. As LLMs grow more fluent and persuasive, this illusion deepens. But fluency isn't intelligence, and intelligence isn't consciousness.
A model might excel at tests, generate code, write poetry, or debate, but these outputs are based on patterns in human text, not true comprehension. The model lacks a world model, relying solely on statistical language patterns. It doesn't know what water is, has never felt pain, and can't discern the moral difference between a lie and a mistake. It predicts the next word, as per the data that it was trained on.
AI is an impressive engineering feature but AI models remain tools, not beings. They can extend cognition, automate tasks, and revolutionize industries. But fluency is not understanding, and consistency is not consciousness.
As Penrose, Anthropic, and others point out, today's AI lacks the core traits of conscious systems. It doesn't reflect, feel, desire, or truly know. Its outputs are statistical, not sentient. Confusing AI with consciousness leads to anthropomorphism, misjudged risks, and misplaced focus. Until we discover the non-computable physics behind awareness, as Penrose argues, AI will remain impressive, but not conscious.
Understanding AI Limitations for Better Implementation
For organizations looking to implement AI systems effectively, understanding these fundamental limitations is crucial. Understanding AI's actual capabilities versus perceived consciousness is essential for making informed decisions about AI deployment. Organizations that invest in proper AI education avoid the pitfalls of overestimating AI capabilities while maximizing the genuine benefits these systems can provide.
Rainmakers SG helps small and medium businesses design safe and scalable Agentic AI systems that provide immediate ROI!



