Mind, AI & ConsciousnessDeep Dive #4Explainer

Do LLMs understand, or only simulate thinking?

Do LLMs "Understand"?

Date
July 24, 2025
Location
SFU Vancouver
Attendance
17 people

The explainer

Models can look thoughtful while failing brittle tests, and still persuade people they "get it." This explainer separates benchmark fluency, chain-of-thought theatre, and the harder claim that something is understood from the inside.

Fluency is a performance

A long, tidy answer can be a pattern that survived the prompt. Change names, invert the structure, or move the same logic into an unfamiliar wrapper, and some systems fall off a cliff. That is useful evidence about robustness. It is not, by itself, a window into phenomenal understanding.

Functional understanding is about what a system can do: transfer a concept, explain a result, catch an error, act effectively in a new case. Phenomenal understanding is whether there is a felt grasp. A model can pass some functional tests and leave the phenomenal question untouched. People also rationalize, confabulate, and fail under unfamiliar representations. The job is to design discriminating tests, not to protect a flattering story about either kind of mind.

Chain-of-thought is not a confession

Intermediate tokens that look like reasoning can be useful scaffolding, a style the training distribution rewards, or a way to spend test-time compute. Extra “thinking” can even make some tasks worse. Treat the trace as behaviour to evaluate, not as a diary of inner life.

Session history

Deep Dive #4 (July 24, 2025, SFU Vancouver) used Apple’s GSM-Symbolic work and inverse-scaling notes as receipts, not as a verdict that machines are empty. The room treated strong performance as important evidence while resisting the jump from successful output to a claim about experience. This public page removes the historical participant roster and chat excerpts.

Nearby questions in the series

If the word “understand” itself is crowded, continue at what it means for AI to understand. If competence can exist with the lights off, see whether intelligence can exist without consciousness. If fluent personas start to feel like they have a life of their own, the later dossier is whether AI personas are information parasites.

Further reading

The curated route is the Can Machines Understand? shelf on the MAC Library. Sibling surfaces: the Community Source Index for receipts, and mac.bc-ai.ca for the group home. Do not treat any one of those as the single reading-resource winner.

Key debates

  • Do LLMs simulate understanding without experiencing it?
  • Is "functional understanding" real understanding?
  • If AI can manipulate psychology, does it need to experience emotions?

Readings

  • PaperThe Illusion of Thinking (GSM-Symbolic) · Shojaee et al. (Apple Research)
  • PaperAlphaEvolve Paper · DeepMind

Where the room landed

Critical distinction emerged: Functional understanding (AI can do this) vs Phenomenal understanding (requires consciousness). Set up the P-zombie debate.

#understanding#llm#ai#qualia

Try it yourself · 2 interactive

Walk through the experiments from this session

These widgets turn the session questions into thought experiments. They are not evidence of room consensus. State stays in your browser.

The GSM-Symbolic stress test from Apple Research. Same math problem, different surface details: can the LLM keep up?

Apple Math Reasoning Tester

From Deep Dive #4: Apple Research's The Illusion of Thinking paper revealed that LLMs show cliff-like performance degradation when math problems are modified, exposing pattern matching rather than true understanding.

You'll see an original math problem that LLMs solve easily. Then you'll see it modified (changed names, structure, or complexity). Predict: Will the LLM still solve it?

"Changing names shouldn't break logic... right?": Participant J, Deep Dive #4

How to Play
  • • 10 math problems
  • • See the original (LLM succeeds)
  • • See the modified version
  • • Predict: Will LLM solve it?
  • • Learn why it fails
What You'll Learn
  • • Where LLMs break
  • • Pattern matching vs reasoning
  • • The performance cliff
  • • Functional vs phenomenal understanding

Photos coming

MAC sessions ran small (~20 people, Chatham House rules). Where event photos surface, they'll be embedded here.

You might also like

Go deeper

The MAC microsite has the interactive version

We built a separate interactive site for this series, with visualizations, thought experiments, and a glossary of consciousness terms. It's the full deep-dive experience for this session.

Explore the interactive deep-dive