
Welcome back, AI prodigies!
In todayâs sunday special:
đ The Prelude
đȘ Why Donât We Do Whatâs Good for Us?
đ The 2008 Housing Crisis
đ€ Human Misalignment Undermines AI Alignment
đ Key Takeaway
Read time: 7 minutes
đ©ș PULSE CHECK
Will AI ever truly align with human values?
đ Key Terms
AI Alignment: The pursuit of ensuring that AI Models behave consistently with human intentions, values, and goals.
Chain-of-Thought (CoT): The intermediate reasoning steps that provide a partial view of how AI Models arrive at answers.
đ THE PRELUDE
We know what we should do. We just donât always do it. We tell ourselves we want to sleep earlier, eat healthier, or exercise regularly. Even so, we find ourselves endlessly pulled elsewhere: we avoid, we delay, we stall. We might regret it, but we often repeat it.
This disconnect between what we say and what we do forces us to ask: âWhat do we really want?â Our brain says one thing, while our behavior says another. The mismatch between intention and implementation points to at least one of two realities:
đŹ REALITY #1: We arenât always sure what we truly want.
đ REALITY #2: Our long-term principles arenât strong enough to resist our short-term pleasures.
If we canât always act on our own principles, how can we teach AI to do it, especially when we donât always agree on what those principles are? In other words, how can we expect AI to align with such fractured, unstable, and conflicting principles?
đȘ WHY DONâT WE DO WHATâS GOOD FOR US?
⊿ 1ïžâŁ The Origins of Short-Term Thinking
Throughout the course of human history, life was short, and livelihood was fragile. For example, hunter-gatherers had an average lifespan of 33, nearly 40 years less than todayâs average lifespan. Ultimately, survival was precious: no injury was minor, no illness was mild, no harvest was plentiful, and no hunt was bountiful. During that time, planning too far ahead was a risky proposition because staying alive required constant vigilance. Today, success often depends on planning for the future. For example, the average retirement age in the U.S. is about 65, with around 60% of Americans funneling money into retirement accounts, such as a 401(k).
Despite these societal shifts across generations of human evolution, our brains remain hardwired for immediacy, nudging us to prioritize short-term gratification over long-term goals. This evolutionary trait explains why our behaviors often betray our stated intentions: we snack on junk food instead of eating healthier or scroll through social media instead of sleeping earlier. What once helped us survive now pulls us toward momentary pleasures.
⊿ 2ïžâŁ The Mental Stress of Holding Conflicting Beliefs
In 1957, American social psychologist Leon Festinger proposed the theory of âCognitive Dissonanceâ: the psychological discomfort experienced when a person holds two or more conflicting beliefs. Crucially, he discovered that people typically donât resolve this tension by changing their behavior to align with their beliefs. Instead, they unconsciously change their beliefs to justify their behavior.
For instance, we know that scrolling through social media for hours can reduce our attention span and cripple our sleep quality. Still, we might rationalize our behavior: âEveryone does it!â or âItâs how I connect with friends!â In simple terms, our beliefs mutate to preserve our behavior.
⊿ 3ïžâŁ The Popular Myth of Being âUnmotivatedâ
Have you ever lied to your dentist about how often you floss? Well, âliedâ is a strong word. Letâs say, âfudged a little.â If youâre like most people, you think about how your dentist might react if you told the truth. Maybe you worry that your dentist will judge you for not taking better care of your teeth. Maybe you worry that your dentist will lecture you about proper flossing technique. Giving the ârightâ answer allows you to avoid unpleasantness and go on with your day.
Why wouldnât you be willing to do something thatâs good for you? Why donât you floss every day? When a person isnât doing the ârightâ thing, we tend to label them as unmotivated. In reality, the notion of being unmotivated lies in the frustrating, demoralizing, and completely human state of ambivalence: the all-too-common experience of being âon the fenceâ about something. For example, while eating healthier and exercising regularly would improve your overall well-being, you might have to give up late-night snacks and the comfort they bring. In other words, people can be motivated to improve their lives while simultaneously being motivated to avoid the discomfort that improvement requires.
đ THE 2008 HOUSING CRISIS
⊿ 4ïžâŁ The Subprime Mortgage Meltdown
Even when we act in alignment with our values, beliefs, or goals, things can still go terribly wrong. The 2008 Housing Crisis is a textbook example of how misaligned incentives can lead to catastrophic outcomes. So, how did predatory lending in a relatively small portion of the U.S. home mortgage market trigger the most severe financial crisis in America since the Great Depression? Each interested party involved acted ârationallyâ within their own silo, without considering how their own incentives affected the entire System. To understand how Systematic Misalignment compounded, we must examine the five major players:
đ Mortgage Originators: They gave out subprime mortgages: high-interest-rate home loans offered to risky borrowers with poor credit scores. They didnât care whether the risky borrowers could actually repay the high-interest-rate home loans because they immediately sold them to investment banks on Wall Street. Their incentive was: âI get paid when I issue the high-interest-rate home loan, not when it gets repaid!â
đ Investment Banks: They bought thousands of high-interest-rate home loans given to risky borrowers and repackaged them into a new investment product called Mortgage-Backed Securities (MBS), which allowed institutional investors to bet on whether the mortgages within MBS would be paid off. Essentially, investment banks made money by purchasing high-interest-rate home loans, bundling them, then selling those bundles to institutional investors. They relied on credit rating agencies to give MBS favorable credit ratings. Their incentive was: âAs long as the investment product is given a favorable credit rating, we can sell it at a premium to institutional investors and make a profit!â
đ Credit Rating Agencies: They independently assessed the credit risks of MBS, but the investment banks bundling them also paid them a hefty fee for credit ratings. To keep the investment banks happy, they often gave overly favorable credit ratings to keep collecting these hefty fees. Their incentive was: âIf we donât rate these favorably, the banks will just go to a competing credit rating agency!â
đ Institutional Investors: These included pension funds, which manage retirement savings, and insurance companies, which invest the premiums collected from policymakers. Both are heavily regulated and required to invest in âsafe assetsâ to meet expected pension payments and anticipated insurance claims. Because MBS were given favorable credit ratings from the credit rating agencies, institutional investors bought them. Their incentive was: âItâs rated highly, so it must be safe, and weâre required to hold high-rated assets!â
đ Regulators and Policymakers: They were reluctant to interfere with the booming U.S. housing market. After all, homeownership is a cornerstone of the âAmerican Dream,â a belief that anyone can achieve success through hard work. Their incentive was: âLet the U.S. housing market work itself out. Itâs helping Americans become homeowners!â
đ€ HUMAN MISALIGNMENT UNDERMINES AI ALIGNMENT
⊿ 5ïžâŁ Itâs Just Billions of Interconnected Weights
The worldâs leading frontier AI firms, including OpenAI, Anthropic, and Google DeepMind, are actively investigating CoT to monitor how language models like ChatGPT, Claude, and Gemini transform inputs into outputs.
For context, a language model is often referred to as a âblack boxâ because we canât confidently trace the internal steps that lead to a specific output. In other words, itâs not always clear why language models generate the responses they do. Opening up the âblack boxâ doesnât help either because language models âthinkâ before generating responses, which appears as a series of dense mathematical vectors called neural activations that are extremely complex: â{0 = 0.83, 1 = 0.15, 2 = 0.47}.â These neural activations happen across multiple layers within a language model:
đ Initial Layers: Syntax/Punctuation/Vocabulary
đ Middle Layers: Context/Content/Concepts
đ Deeper Layers: Tone/Intent/Impact
CoT encourages language models to produce intermediate reasoning steps before generating a specific output. In other words, CoT offers a rare glimpse into how language models âthinkâ in a way that we can easily understand. For example, Anthropic recently developed âAIM,â which leverages a Cross-Layer Transcoder (CLT) to identify computational âcircuits,â revealing parts of the pathway that render prompts into specific outputs. CLT essentially acts as a âcomputational translator,â grouping a language modelâs internal workings into human-understandable representations.
⊿ 6ïžâŁ The Core Dilemma: Whose Alignment?
The most unsettling aspect of todayâs language models isnât that they work; itâs that we donât fully understand how they work. If we can understand how language models choose to generate a specific output, we can theoretically build safer and more steerable versions of ChatGPT, Claude, and Gemini that align with human values. But how can we expect to build âaligned AIâ when we donât even have âaligned humans?â
This creates a Recursive Trap: misaligned humans train AIâs next frontier to mirror, mutate, and multiply their misalignment. The future of AI Alignment requires a collective willingness to confront human misalignment. This means acknowledging that the very human values we want AIâs next frontier to reflect are often unstable, fragmented, and conflicting.
đ KEY TAKEAWAY
We routinely struggle to align our intentions with our actions. AI Alignment inherits this challenge: how can we expect AI to align with human values that are constantly divided, disputed, and debated? Until we decode the forces that fuel human misalignment, âaligned AIâ will remain more myth than mission.
đ FINAL NOTE
FEEDBACK
How would you rate todayâs email?
â€ïž Todayâs Featured Reply
âItâs very well done and has a nice, digestible tone. The breakdown of NVIDIA providing up to $105 billion in financial guarantees for leases tied to AI compute infrastructure was spot on.â
REFER & EARN
đ Your Friends Learn, You Earn!
{{rp_personalized_text}}
Share your unique referral link: {{rp_refer_url}}

