Why My Cat Still Beats AI Agents at Reasoning and Planning

Artificial General Intelligence (AGI)
The term AGI is widely used but has many interpretations. OpenAI defines AGI as AI systems that outperform humans in most economically valuable tasks. OpenAI's CEO, Sam Altman, recently suggested AGI could emerge during DJ Trump's term. He also revealed that OpenAI's ambitions go beyond AGI, aiming for superintelligence: AI capable of accelerating science and prosperity far beyond human limits. Altman believes this could be achieved within a few thousand days. Elon Musk predicts that AI could surpass humans by 2026.
In contrast, Yann LeCun, Meta's Chief AI Scientist, argues that Large Language Models (LLMs) lack core abilities like reasoning, planning, and understanding intuitive physics. LeCun stresses that understanding of intuitive physics is far more complex than learning language, hence AGI will require entirely new AI architectures and this takes time. Sir Roger Penrose, Nobel Prize for Physics winner in 2020, argues that human consciousness cannot be replicated by computational means."
My Cat vs. ChatGPT
The amount of information that humans and animals gather via vision and touch significantly outscores the information that they receive through text. Cats learn by interacting physically with the world through observation, trial and error, and feedback. They develop spatial awareness, motor skills, understand cause-effect, adapt, and plan actions like stalking prey. Their memory is long-term, grounded in direct, real-world experience.
LLMs like GPT-4 lack real-world experience. They are trained on large amounts of text via self-supervised learning. In self-supervised learning the LLM learns the structure of language by predicting masked words based on context. For every prompt, the LLM outputs the probability of each word of the English dictionary being next. Assuming that the English dictionary contains 50,000 words, the LLM outputs 50,000 (word, probability) pairs during every inference cycle. When generating text, the LLM typically selects the word with the highest probability, adds it to the original prompt and repeats the process. A pre-trained LLM is not ready for deployment and several post training activities are required to make it safe and have it aligned to human values. Self-supervised learning proves highly effective for text, making it the dominant method for pre-training nearly all modern LLMs.

However, applying the same self-supervised method to video fails. Masking and predicting video frames proves infeasible due to the continuous nature of video, where infinite outcomes exist for any frame sequence. Accurate prediction requires understanding physics and causality, far beyond text prediction. As a result, video prediction will demand a fundamentally different approach, one that goes beyond next-word prediction and addresses continuous, high-dimensional uncertainty.

Below table contrasts LLMs with humans and animals, highlighting key differences in learning, reasoning, adaptability, and planning.

LLMs don’t reason, they just guess at scale!
OpenAI and DeepSeek have released their so-called reasoning models, with OpenAI launching o3 in April 2025, and DeepSeek releasing their updated R1-0528 model in May 2025, despite knowing that LLMs cannot truly reason. Their operation is inherently statistical and yet, they may give the illusion of reasoning due to:
- Memorizing billions of examples
- Chain-of-Thought prompting guiding them toward logical outputs
- Fine-tuning to avoid wrong answers
What appears as reasoning is actually brute-force generation of outputs based on probabilistic patterns. A human or another AI system subsequently selects the best answer. This way of reasoning is inefficient, costly, and fragile process.
In contrast, human reasoning is grounded in internal mental models of the world. Humans simulate possible actions, predict outcomes, set subgoals, and adapt dynamically to achieve long-term objectives. Our reasoning is abstract, hierarchical, and shaped by real-world constraints.
Given the prompt "If it rains, the ground is...", an LLM predicts "wet" based on statistical patterns in text. However, it lacks any real understanding of rain, water, or ground. In contrast, a human knows rain produces water, water soaks the ground, and as a result, the ground becomes wet, demonstrating causal understanding, not just correlation.
Agentic AI: The New Hype
In most workflows a single prompt generates a single response. This technique is often referred to as zero-shot prompting. Agentic workflows plan, execute, reflect and revise, leading to better results. Agentic wrapping improves model performance significantly. Andrew Ng considers 4 design patterns for Agentic AI systems:
- Reflection: LLM evaluates its own output, for example, LLM agents review and refine the code they previously generated.
- Tool Use: LLM relies on external tools including web search, prompt augmentation (RAG), calculators and code execution to complete a task
- Planning/Reasoning: LLM plans a sequence of tasks, re-routes if failure occurs and recovers iteratively
- Multi-Agent Collaboration: multiple LLMs collaborate, debate and correct each other
AI Agents and agentic AI systems refer to fully autonomous tools built on LLMs designed to execute complex tasks with minimal human oversight across enterprise workflows including sales automation and marketing automation. However, significant challenges persist and the gap with human cognition remains significant. True, complex, general-purpose and fully autonomous agentic systems are still out of reach.
Memory modules like vector databases in RAG setups offer shallow recall, not deep understanding. Planning is typically rule-based or hard-coded rather than emerging naturally.

RAG, AI Agents and Enterprise AI
Many business leaders I've spoken with admit their AI efforts focus almost entirely on internal use cases. The most common application is Retrieval-Augmented Generation (RAG), essentially enterprise-level semantic search that enhances AI accuracy by incorporating external knowledge sources. However, proving a clear ROI, especially for complex RAG systems, remains challenging.
To minimize risk, enterprises prefer to limit AI's influence on core operations. They are likely to keep autonomous features tightly controlled, enforcing strict guardrails. As a result, Agentic AI systems will mainly be used internally. In an enterprise context, predictability trumps autonomy, hence we will not see fully autonomous agents in the near term.
The key message is simple: LLMs excel at automating repetitive, text-based tasks, summarizing documents, drafting emails, powering chatbots, but they lack flexibility. They struggle with unexpected scenarios or tasks requiring real reasoning, such as complex legal cases or supply chain disruptions.
Organizations implementing large-scale AI deployments increasingly require specialized AI Consulting to navigate the complex trade-offs between performance, cost, and environmental impact while ensuring their AI strategies align with both business objectives and sustainability commitments.
The Reality Check for AI Implementation
Understanding these limitations becomes crucial for organizations developing effective AI strategies. Companies must balance the hype around AI capabilities with practical implementation realities. Another concern companies are dealing with is to manage shadow AI use. AI limitations and shadow use can be managed via AI training, and an AI policy.
For organizations evaluating AI implementation, understanding AI's fundamental limitations helps in making informed model selection decisions, whether considering traditional models or alternative options such as DeepSeek.
Rainmakers SG helps small and medium businesses design safe and scalable Agentic AI systems that provide immediate ROI!



