Mind, AI & ConsciousnessDeep Dive #4Post-event recap

The Illusion of Thinking

Do LLMs "Understand"?

Date
July 24, 2025
Location
SFU Vancouver
Attendance
17 people

Central question

What does it mean to "understand"? Do LLMs have understanding or just simulate it?

Key debates

  • Do LLMs simulate understanding without experiencing it?
  • Is "functional understanding" real understanding?
  • If AI can manipulate psychology, does it need to experience emotions?

Readings

  • PaperThe Illusion of Thinking (GSM-Symbolic) · Shojaee et al. (Apple Research)
  • PaperAlphaEvolve Paper · DeepMind

Where the room landed

Critical distinction emerged: Functional understanding (AI can do this) vs Phenomenal understanding (requires consciousness). Set up the P-zombie debate.

#understanding#llm#ai#qualia

Try it yourself · 2 interactive

Walk through the experiments from this session

These widgets turn the session questions into thought experiments. They are not evidence of room consensus. State stays in your browser.

The GSM-Symbolic stress test from Apple Research. Same math problem, different surface details: can the LLM keep up?

Apple Math Reasoning Tester

From Deep Dive #4: Apple Research's The Illusion of Thinking paper revealed that LLMs show cliff-like performance degradation when math problems are modified, exposing pattern matching rather than true understanding.

You'll see an original math problem that LLMs solve easily. Then you'll see it modified (changed names, structure, or complexity). Predict: Will the LLM still solve it?

"Changing names shouldn't break logic... right?": Participant J, Deep Dive #4

How to Play
  • • 10 math problems
  • • See the original (LLM succeeds)
  • • See the modified version
  • • Predict: Will LLM solve it?
  • • Learn why it fails
What You'll Learn
  • • Where LLMs break
  • • Pattern matching vs reasoning
  • • The performance cliff
  • • Functional vs phenomenal understanding

The recap

Record status: Completed-session recap synthesized from the organizer archive and public reading list. It uses no participant quotations or attributed views.

What Counts as Understanding?

Deep Dive #4 used research on reasoning-model failures to examine the gap between an answer that looks reasoned and a process that remains brittle under changed conditions. The room treated strong performance as important evidence while resisting the jump from successful output to a claim about inner experience.

Two Meanings of Understanding

Functional understanding concerns what a system can do: transfer a concept, explain a result, detect an error, or act effectively. Phenomenal understanding concerns whether there is a felt sense of grasping the problem. A system might satisfy some functional tests without resolving the phenomenal question.

The discussion also complicated simple human-versus-machine comparisons. People rationalize, confabulate, and fail under unfamiliar representations too. The useful task is to design discriminating tests, not protect a flattering story about either kind of intelligence.

Questions That Continued

  1. Which failures reveal shallow pattern matching rather than ordinary limits?
  2. Is functional understanding enough for trust and responsibility?
  3. Can a system recognize the difference between knowledge and fabrication?

Source Boundary

This recap removes the historical participant roster and chat excerpts from the public rendering.

Public Sources

Photos coming

MAC sessions ran small (~20 people, Chatham House rules). Where event photos surface, they'll be embedded here.

You might also like

Go deeper

The MAC microsite has the interactive version

We built a separate interactive site for this series, with visualizations, thought experiments, and a glossary of consciousness terms. It's the full deep-dive experience for this session.

Explore the interactive deep-dive