Welcome back, AI prodigies!

In today’s sunday special:

  • 📜 The Prelude

  • đŸȘž Why Don’t We Do What’s Good for Us?

  • 🏠 The 2008 Housing Crisis

  • đŸ€– Human Misalignment Undermines AI Alignment

  • 🔑 Key Takeaway

Read time: 7 minutes

đŸ©ș PULSE CHECK

Login or Subscribe to participate

🎓 Key Terms

  • AI Alignment: The pursuit of ensuring that AI Models behave consistently with human intentions, values, and goals.

  • Chain-of-Thought (CoT): The intermediate reasoning steps that provide a partial view of how AI Models arrive at answers.

📜 THE PRELUDE

We know what we should do. We just don’t always do it. We tell ourselves we want to sleep earlier, eat healthier, or exercise regularly. Even so, we find ourselves endlessly pulled elsewhere: we avoid, we delay, we stall. We might regret it, but we often repeat it.

This disconnect between what we say and what we do forces us to ask: “What do we really want?” Our brain says one thing, while our behavior says another. The mismatch between intention and implementation points to at least one of two realities:

  1. 💬 REALITY #1: We aren’t always sure what we truly want.

  2. 💭 REALITY #2: Our long-term principles aren’t strong enough to resist our short-term pleasures.

If we can’t always act on our own principles, how can we teach AI to do it, especially when we don’t always agree on what those principles are? In other words, how can we expect AI to align with such fractured, unstable, and conflicting principles?

đŸȘž WHY DON’T WE DO WHAT’S GOOD FOR US?

⊿ 1ïžâƒŁ The Origins of Short-Term Thinking

Throughout the course of human history, life was short, and livelihood was fragile. For example, hunter-gatherers had an average lifespan of 33, nearly 40 years less than today’s average lifespan. Ultimately, survival was precious: no injury was minor, no illness was mild, no harvest was plentiful, and no hunt was bountiful. During that time, planning too far ahead was a risky proposition because staying alive required constant vigilance. Today, success often depends on planning for the future. For example, the average retirement age in the U.S. is about 65, with around 60% of Americans funneling money into retirement accounts, such as a 401(k).

Despite these societal shifts across generations of human evolution, our brains remain hardwired for immediacy, nudging us to prioritize short-term gratification over long-term goals. This evolutionary trait explains why our behaviors often betray our stated intentions: we snack on junk food instead of eating healthier or scroll through social media instead of sleeping earlier. What once helped us survive now pulls us toward momentary pleasures.

⊿ 2ïžâƒŁ The Mental Stress of Holding Conflicting Beliefs

In 1957, American social psychologist Leon Festinger proposed the theory of “Cognitive Dissonance”: the psychological discomfort experienced when a person holds two or more conflicting beliefs. Crucially, he discovered that people typically don’t resolve this tension by changing their behavior to align with their beliefs. Instead, they unconsciously change their beliefs to justify their behavior.

For instance, we know that scrolling through social media for hours can reduce our attention span and cripple our sleep quality. Still, we might rationalize our behavior: “Everyone does it!” or “It’s how I connect with friends!” In simple terms, our beliefs mutate to preserve our behavior.

⊿ 3ïžâƒŁ The Popular Myth of Being “Unmotivated”

Have you ever lied to your dentist about how often you floss? Well, “lied” is a strong word. Let’s say, “fudged a little.” If you’re like most people, you think about how your dentist might react if you told the truth. Maybe you worry that your dentist will judge you for not taking better care of your teeth. Maybe you worry that your dentist will lecture you about proper flossing technique. Giving the “right” answer allows you to avoid unpleasantness and go on with your day.

Why wouldn’t you be willing to do something that’s good for you? Why don’t you floss every day? When a person isn’t doing the “right” thing, we tend to label them as unmotivated. In reality, the notion of being unmotivated lies in the frustrating, demoralizing, and completely human state of ambivalence: the all-too-common experience of being “on the fence” about something. For example, while eating healthier and exercising regularly would improve your overall well-being, you might have to give up late-night snacks and the comfort they bring. In other words, people can be motivated to improve their lives while simultaneously being motivated to avoid the discomfort that improvement requires.

🏠 THE 2008 HOUSING CRISIS

⊿ 4ïžâƒŁ The Subprime Mortgage Meltdown

Even when we act in alignment with our values, beliefs, or goals, things can still go terribly wrong. The 2008 Housing Crisis is a textbook example of how misaligned incentives can lead to catastrophic outcomes. So, how did predatory lending in a relatively small portion of the U.S. home mortgage market trigger the most severe financial crisis in America since the Great Depression? Each interested party involved acted “rationally” within their own silo, without considering how their own incentives affected the entire System. To understand how Systematic Misalignment compounded, we must examine the five major players:

❝
  1. 🔘 Mortgage Originators: They gave out subprime mortgages: high-interest-rate home loans offered to risky borrowers with poor credit scores. They didn’t care whether the risky borrowers could actually repay the high-interest-rate home loans because they immediately sold them to investment banks on Wall Street. Their incentive was: “I get paid when I issue the high-interest-rate home loan, not when it gets repaid!”

  2. 🔘 Investment Banks: They bought thousands of high-interest-rate home loans given to risky borrowers and repackaged them into a new investment product called Mortgage-Backed Securities (MBS), which allowed institutional investors to bet on whether the mortgages within MBS would be paid off. Essentially, investment banks made money by purchasing high-interest-rate home loans, bundling them, then selling those bundles to institutional investors. They relied on credit rating agencies to give MBS favorable credit ratings. Their incentive was: “As long as the investment product is given a favorable credit rating, we can sell it at a premium to institutional investors and make a profit!”

  3. 🔘 Credit Rating Agencies: They independently assessed the credit risks of MBS, but the investment banks bundling them also paid them a hefty fee for credit ratings. To keep the investment banks happy, they often gave overly favorable credit ratings to keep collecting these hefty fees. Their incentive was: “If we don’t rate these favorably, the banks will just go to a competing credit rating agency!”

  4. 🔘 Institutional Investors: These included pension funds, which manage retirement savings, and insurance companies, which invest the premiums collected from policymakers. Both are heavily regulated and required to invest in “safe assets” to meet expected pension payments and anticipated insurance claims. Because MBS were given favorable credit ratings from the credit rating agencies, institutional investors bought them. Their incentive was: “It’s rated highly, so it must be safe, and we’re required to hold high-rated assets!”

  5. 🔘 Regulators and Policymakers: They were reluctant to interfere with the booming U.S. housing market. After all, homeownership is a cornerstone of the “American Dream,” a belief that anyone can achieve success through hard work. Their incentive was: “Let the U.S. housing market work itself out. It’s helping Americans become homeowners!”

☑ FUN FACT: At the height of the 2008 Housing Crisis, lenders offered NINJA loans to risky borrowers with no job, no income, and no assets, making it possible for them to become homeowners without having to prove creditworthiness.

đŸ€– HUMAN MISALIGNMENT UNDERMINES AI ALIGNMENT

⊿ 5ïžâƒŁ It’s Just Billions of Interconnected Weights

The world’s leading frontier AI firms, including OpenAI, Anthropic, and Google DeepMind, are actively investigating CoT to monitor how language models like ChatGPT, Claude, and Gemini transform inputs into outputs.

For context, a language model is often referred to as a “black box” because we can’t confidently trace the internal steps that lead to a specific output. In other words, it’s not always clear why language models generate the responses they do. Opening up the “black box” doesn’t help either because language models “think” before generating responses, which appears as a series of dense mathematical vectors called neural activations that are extremely complex: “{0 = 0.83, 1 = 0.15, 2 = 0.47}.” These neural activations happen across multiple layers within a language model:

  1. 📕 Initial Layers: Syntax/Punctuation/Vocabulary

  2. 📗 Middle Layers: Context/Content/Concepts

  3. 📘 Deeper Layers: Tone/Intent/Impact

CoT encourages language models to produce intermediate reasoning steps before generating a specific output. In other words, CoT offers a rare glimpse into how language models “think” in a way that we can easily understand. For example, Anthropic recently developed “AIM,” which leverages a Cross-Layer Transcoder (CLT) to identify computational “circuits,” revealing parts of the pathway that render prompts into specific outputs. CLT essentially acts as a “computational translator,” grouping a language model’s internal workings into human-understandable representations.

⊿ 6ïžâƒŁ The Core Dilemma: Whose Alignment?

The most unsettling aspect of today’s language models isn’t that they work; it’s that we don’t fully understand how they work. If we can understand how language models choose to generate a specific output, we can theoretically build safer and more steerable versions of ChatGPT, Claude, and Gemini that align with human values. But how can we expect to build “aligned AI” when we don’t even have “aligned humans?”

This creates a Recursive Trap: misaligned humans train AI’s next frontier to mirror, mutate, and multiply their misalignment. The future of AI Alignment requires a collective willingness to confront human misalignment. This means acknowledging that the very human values we want AI’s next frontier to reflect are often unstable, fragmented, and conflicting.

🔑 KEY TAKEAWAY

We routinely struggle to align our intentions with our actions. AI Alignment inherits this challenge: how can we expect AI to align with human values that are constantly divided, disputed, and debated? Until we decode the forces that fuel human misalignment, “aligned AI” will remain more myth than mission.

📒 FINAL NOTE

FEEDBACK

How would you rate today’s email?

We read every reply you leave after rating!

Login or Subscribe to participate

❀ Today’s Featured Reply

❝

“It’s very well done and has a nice, digestible tone. The breakdown of NVIDIA providing up to $105 billion in financial guarantees for leases tied to AI compute infrastructure was spot on.”

-John (1ïžâƒŁ 👍 Nailed it!)
REFER & EARN

🎉 Your Friends Learn, You Earn!

{{rp_personalized_text}}

Share your unique referral link: {{rp_refer_url}}