
Welcome back, AI prodigies!
In todayâs sunday special:
đ The Prelude
đ„ Infants Have Physical Intuition
đ From Predicting to Problem-Solving
đ Words to Worlds Is AIâs Next Frontier?
đ Key Takeaway
Read time: 7 minutes
đ©ș PULSE CHECK
As AI gains physical form, how do you feel about the future?
đ Key Terms
AI Agents: Software Programs that analyze, arrange, and automate on your behalf without you lifting a finger.
Large Language Models (LLMs): AI Models pre-trained on vast amounts of data to generate human-like text.
đ THE PRELUDE
As you scoot between the booth and the table, your elbow knocks over a cup, sending it tumbling over the tableâs edge. You grit your teeth, bracing for the impact. It lands with a barely audible âclink!,â completely intact. Your split-second reaction to the falling cup wasnât planned. It was an instinctive response.
Believe it or not, humans experience millions of distinct social interactions throughout their lives. Over time, the human brain compresses them into a mental model equipped with predictive coding that continuously anticipates reactions to specific actions.
Today, AIâs ability to grasp reality is severely limited because it doesnât possess lived experience. To achieve superintelligence, it must directly interact with the underlying laws that govern the physical world, whether itâs gravity, friction, or collision, to understand the forces and flows of the physical universe and supercharge everything from humanoid robots to scientific discoveries.
đ„ INFANTS HAVE PHYSICAL INTUITION
⊿ 1ïžâŁ Do Babies Know More Than We Think?
In 1991, Canadian American research psychologist Renée Baillargeon tested whether young infants understand object permanence: an object continues to exist even when it can no longer be seen.
Since young infants canât explain what theyâre thinking, she measured how long they looked at different events. This method of measurement is rooted in the fundamental belief that babies tend to look longer at events that are inconsistent with what they understand should happen.
⊿ 2ïžâŁ The Carrot Event, Explained.
In one experiment, 3.5-month-old infants watched a short carrot and a tall carrot slide along a track. The trackâs center was hidden by a screen with a large window in its upper half. The short carrot was shorter than the large windowâs lower edge, and it expectedly didnât appear in the large window when passing behind it. The tall carrot was taller than the large windowâs lower edge, and it unexpectedly didnât appear in the large window when passing behind it.
The 3.5-month-old infants looked significantly longer at the tall carrot event than at the short carrot event, suggesting that they held a mental model of the existence, height, and trajectory of each carrot behind the screen and expected the tall carrot to appear in the large window and were surprised when it didnât.
đ FROM PREDICTING TO PROBLEM-SOLVING
⊿ 3ïžâŁ Language Models, Explained.
The worldâs most popular LLMs, like ChatGPT, Claude, and Gemini, are statistical systems designed to predict the probability of a sequence of tokens. In simple terms, theyâre essentially sophisticated autocomplete machines trained on the entire internet to process diverse inputs and generate plausible outputs that sound human.
For context, tokens are units of text that enable LLMs to understand humans. They bridge the gap between human input and machine output. In English, one token generally corresponds to four characters. This translates to roughly Ÿ of a word, meaning 1,000 tokens are roughly equivalent to 750 words.
WHAT COUNTS AS A TOKEN?
đ Words: âthe,â ârun,â and âapple.â
đ Parts of Words: â-un,â â-ing,â and â-tion.â
đ Characters: âa,â âX,â and â9.â
đ Punctuation: â.â or â?â or â!â
đ Special Symbols: â+,â â=,â and â%.â
To train an LLM, itâs shown millions of sequences of text and is prompted to predict the next token. This foundational training method is referred to as autoregressive next-token prediction. For example, when given: âI feel anxious when I speak in front of a {BLANK}!â LLMs ask themselves, given the tokens so far, whatâs the most likely next token? In this instance, it might predict: â{CROWD}!â
Each time the LLM makes an inaccurate prediction, it adjusts billions of weights, or numerical parameters, that help it recognize which patterns of tokens are more important for making better predictions in the future. These numerical parameters control how tokens relate to each other within the Neural Network (NN), which mimics the human brain by processing diverse inputs through hundreds of transformer layers comprised of millions of ânodes.â The core components of a transformer layer include:
đ Attention Weights: Measure how much attention each token assigns to every other token within a sequence of text. Consider the following sentence: âMiami, coined the âMagic City,â has beautiful white-sand beaches!â In this case, âbeachesâ would assign more attention to âMiamiâ because theyâre closely related.
đ Self-Attention Mechanisms: Clarify the meaning of each token within a sequence of text to capture relevant context. Consider the following sentence: âThe cat chased the mouse!â In this case, it would determine that âcatâ is important because itâs doing the chasing and âmouseâ is important because itâs being chased. In other words, âchasedâ establishes the contextual relationship between âcatâ and âmouse.â
âïž Feed-Forward Networks (FFNs): Individually refine the meaning of each token within a sequence of text after the relevant context has been captured by overlaying it with real-world dynamics. In this case, âcatâ carries the context of hunter and âmouseâ carries the context of prey, meaning âchasedâ carries the real-world dynamics of danger and urgency in a life-or-death struggle for survival.
⊿ 4ïžâŁ Reasoning Models, Explained.
LRMs are designed to spend more time thinking before responding. They achieve this by leveraging Test-Time Compute (TTC), which dedicates additional compute to AI inference: everything that happens after you enter your prompt. Imagine asking ChatGPT: âSummarize this article in three bullet points!â AI inference is the reasoning ChatGPT performs to generate that bulleted summary. The core components of AI inference include:
đ Retrieval-Augmented Generation (RAG) to retrieve up-to-date facts directly related to a prompt. Itâs exactly like a student taking an open-book exam. For example, a hospitality agent designed for booking hotels will leverage RAG to access internal policy documents for rules on late check-outs.
đ WORDS TO WORLDS IS AIâS NEXT FRONTIER?
⊿ 5ïžâŁ World Models, Explained.
On Sept. 13th, 2024, World Labs raised a $230 million Series B funding round at a $1 billion post-money valuation to advance LWMs, which attempt to emulate the way the human brain builds mental models of everyday life. Dr. Fei-Fei Li, widely recognized as the âgodmother of AI,â described world models and spatial intelligence as AIâs next frontier. On Feb. 18th, 2026, World Labs secured a $200 million strategic investment from Autodesk as part of a larger $1 billion growth-stage funding round.
Imagine asking ChatGPT: âThe dog fetched the {BLANK}!â In this instance, it might predict: â{TENNIS BALL}!â ChatGPT excels at understanding and generating tokens. It can tell you what happens, but not why or how it happens. ChatGPT canât understand, simulate, and predict the dynamics of the physical world like we do. If we toss a tennis ball into the air, we know gravity will pull it back down. This kind of physical intuition is something we naturally develop by observing how gravity, friction, and collision influence the physical world around us.
⊿ 6ïžâŁ Physical Intuition, Explained.
We rely on our physical intuition every day to predict how the physical world might react to our actions. For example, knowing how much force to use when opening a car door. Itâs not some magic ability. Itâs our brain constantly building a mental model of how the physical world works based on our lived experience.
So, what if language models and reasoning models could also autonomously experience the forces and flows of the physical universe? Thatâs where world models come in. They aim to help ChatGPT, Claude, and Gemini evolve beyond text, images, audio, video, and code by learning to uncover the underlying laws that govern the physical universe.
đ KEY TAKEAWAY
âA four-year-old child has seen 50x more data than the biggest LLMs,â said Yann LeCun, who famously served as Metaâs Chief AI Scientist. âText is simply too low bandwidth and too scarce a modality to learn how the world works.â As young infants, we begin to build mental models of how the physical world functions. In contrast, LLMs predict tokens, and LRMS reason about and act on tokens. Without experiencing the underlying laws that govern the physical universe, both struggle to truly navigate and comprehend unseen, real-world environments.
đ FINAL NOTE
FEEDBACK
How would you rate todayâs email?
â€ïž Todayâs Featured Reply
âItâs hands down the best non-technical guide to AI. Kudos to the team!â
REFER & EARN
đ Your Friends Learn, You Earn!
{{rp_personalized_text}}
Share your unique referral link: {{rp_refer_url}}

