[Book Notes] Jeff Hawkins: A Thousand Brains and Intelligence as Many Models in Motion
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
Introduction: What If There Is No Single Model at the Center of Intelligence?
When we recognize a coffee cup, it feels as though the brain consults one stable representation: this shape, this handle, this familiar use. Jeff Hawkins’s A Thousand Brains: A New Theory of Intelligence proposes a stranger architecture. Vision, touch, and other sensory streams each engage many small regions of the neocortex. These regions learn partially independent models through movement, then communicate until their interpretations agree. What enters consciousness as one object may be the settlement reached by many models.
The title is therefore an architectural claim, not a count of minds or personalities. Hawkins argues that the neocortex resembles a large community of model-building units. Each learns features at locations inside a reference frame; each can make predictions; and long-range connections let many units vote on what is present. Intelligence arises from the coordination of this distributed knowledge.
Published by Basic Books in 2021, the book moves through three scales. The first part presents a theory of the neocortex. The second asks what that theory would imply for machine intelligence and consciousness. The third turns toward false belief, existential risk, and the long-term survival of knowledge. The transitions are ambitious: a proposed neural mechanism becomes a design program for AI and then a philosophy of humanity’s future. Reading the book well requires keeping those levels connected without treating them as equally established.
This post reconstructs that argument and then asks what it offers to robotics and AI today.
1. Old Brain, New Brain: Goals and Models
Hawkins begins with an evolutionary division between older brain structures and the mammalian neocortex. The older systems regulate survival, movement, emotion, appetite, reproduction, and other action-driving functions. The neocortex learns a rich model of the world and uses that model to predict what will happen. In his account, our intelligent behavior results from a continual negotiation: older systems supply goals and values, while the neocortex supplies knowledge about how the world is structured and how a goal might be reached.
This distinction matters because intelligence does not automatically provide motivation. A world model can represent a forest fire, a scientific theory, or another person’s plan; none of those representations alone says what the organism should want. Goals enter through a larger embodied system shaped by evolution, development, culture, and present physiological needs.
The separation also explains a familiar human conflict. A person can understand that an action is harmful and still desire it. Knowledge in the neocortex does not simply overwrite older drives. Conversely, a drive cannot execute a complex plan without recruiting learned models of causes, objects, institutions, and other people.
The phrase “old brain” should be read as Hawkins’s functional simplification, not as a clean anatomical border or a claim that evolution replaced one complete brain with another. Real neural systems are deeply interconnected. The useful conceptual point is narrower: the machinery that learns a model need not be the machinery that chooses the ends for which the model is used. That point later becomes central to his discussion of AI safety.
2. Mountcastle’s Proposal: A Common Cortical Operation
The scientific starting point is a proposal associated with neuroscientist Vernon Mountcastle. Across the neocortex, different regions share a broadly similar layered anatomy. A visual region and a language-related region perform very different tasks, yet their local circuits have striking structural similarities. Mountcastle suggested that the neocortex might be built from repeated units performing a common operation, with functional differences arising largely from what each region is connected to.
Hawkins calls a small patch spanning the cortical layers a cortical column. In his usage, the column is a convenient functional unit; it need not be a sharply bounded cylinder visible in tissue. The hard question is what common operation could be general enough to support seeing, touching, planning, mathematics, and language.
The conventional hierarchical picture offers one answer for perception. Early stages detect simple features; later stages combine them into increasingly complex representations; an object appears only near the top. Hawkins retains cortical hierarchy but argues that hierarchy alone misses too much. It does not explain why the neocortex devotes so much circuitry to movement-related signals, how touch recognizes an object through a sequence of contacts, or why long-range connections link many regions laterally as well as vertically.
His alternative begins with a different unit of knowledge: a feature at a location. A patch of cortex does more than report what its associated sensor currently detects. It also represents where that sensation lies relative to the thing being sensed. Once feature and location are joined, the patch can learn the structure of a complete object over time.
3. Reference Frames: The Coordinate System Inside Knowledge
A location has meaning only inside a reference frame. “Five centimeters to the left” is incomplete until we know: left of the observer, the cup, the table, or the room? Hawkins argues that the brain organizes knowledge through many such frames. Some describe where the body is in an environment. Others describe where a finger is on a cup, where the eye is directed within a scene, or where one component sits inside a larger object.
The inspiration comes from place cells and grid cells in the hippocampal–entorhinal system. These cells participate in representations of an animal’s location and movement through an environment. Hawkins and his collaborators propose that grid-cell-like mechanisms also operate throughout the neocortex. There, they would encode sensor location relative to objects rather than body location relative to a room.
Consider touching a mug with one finger while your eyes are closed. A smooth patch alone could belong to thousands of objects. As the finger moves, the sequence might include a curved wall, an edge, empty space, and a handle. Each sensation becomes informative because it is registered at a location, and movement updates the expected location before the next input arrives. Recognition is therefore an active process:
current feature + object-relative location
↓
update the object model
↓
movement command → predicted new location → predicted feature
The claim is more powerful than saying that spatial information helps perception. Hawkins proposes that reference frames are a general format for knowledge. An object can be located in a room; a handle can be located on a cup; a joint can be located in a robot; a word can occupy a role in a sentence; an idea can be situated within a conceptual structure. Composition becomes possible because one reference frame can be placed inside another.
This generalization is also one of the theory’s largest empirical bets. Grid cells in the entorhinal system are well established. Grid-cell-like mechanisms in every cortical region and column, performing the proposed object-centered function, remain a hypothesis.
4. Movement Is Part of Perception
In this framework, sensing is inseparable from movement. The eyes saccade, the fingers explore, the head turns, and the body changes viewpoint. Even when an external object is still, our sensors move across it. The brain uses a copy of the movement command to update where it expects the sensor to be relative to the object.
This makes prediction local and concrete. If a finger moves from the rim toward the side of a familiar mug, the active model predicts both a new location and the tactile feature likely to occur there. A match strengthens the model. A mismatch creates evidence that the object, action, or current interpretation is wrong.
Hawkins’s earlier work emphasized prediction over time. A Thousand Brains adds the coordinate system that makes prediction structured. The brain is not merely guessing the next input in a sequence; it is predicting what should be sensed here, after this movement, within this model.
Three consequences follow:
- Learning is active. An intelligent system can choose movements that resolve uncertainty instead of waiting for a complete observation.
- Knowledge is action-ready. A model contains relations among possible sensations and possible movements, so perception and behavior share a representation.
- Surprise is spatially diagnostic. An error can reveal that the feature is wrong, the estimated location is wrong, or the reference frame itself is wrong.
For robotics, this is an immediate challenge to pipelines that treat perception as a static image-classification stage followed by a separate controller. A robot learns the meaning of “graspable” through viewpoint change, contact, force, slip, deformation, and recovery. Its useful model belongs to the sensor–body–world loop.
5. Why a Thousand Brains? Complete Models and Voting
If each cortical column only contributed one feature to a single representation elsewhere, the brain would need a central place where the pieces finally become an object. Hawkins proposes a more distributed arrangement: many columns can each learn a complete model of the same object from the inputs available to them.
A fingertip’s column learns a mug through touch across successive movements. Visual columns learn it through patches sampled by changing gaze. Different columns begin with ambiguous evidence, but long-range connections allow them to share candidate interpretations. Their activity converges through a process Hawkins describes as voting.
Imagine several fingers touching a cup at once. One detects a smooth curved surface, another an edge, and another the handle. No local observation is decisive. Yet the interpretations compatible with all three observations reinforce one another, while incompatible candidates lose support. Consensus can emerge quickly because many models operate in parallel.
| Single-model intuition | Thousand-brains proposal |
|---|---|
| Features are assembled into one object representation | Many local modules learn object models |
| Recognition culminates at a privileged high level | Recognition can emerge through agreement across levels and regions |
| Ambiguity is resolved mainly by further feedforward processing | Ambiguity is reduced by movement and lateral voting |
| Damage threatens a central representation | Distributed models offer redundancy and graceful degradation |
| A sensor reports a feature | A sensor-associated module represents a feature at an object-relative location |
The proposal does not remove hierarchy. Objects contain parts, scenes contain objects, and concepts contain other concepts. Hierarchy organizes these nested relations. Voting adds a heterarchical dimension: peer modules at different locations and modalities can constrain one another without sending every decision to a single apex.
The unified percept is thus less like a picture stored in one place and more like a stable agreement maintained by a network.
6. Concepts, Language, and the Expansion Beyond Physical Space
The theory begins with physical objects, where movement and location are easiest to visualize. Human intelligence, however, also handles democracy, evolution, legal systems, melodies, software, and mathematics. Hawkins argues that the neocortex reuses the same machinery for these abstract structures.
An abstract reference frame need not correspond to literal three-dimensional space. It can organize positions within a sequence, roles within a system, or relations within a conceptual domain. We mentally move through a family tree, a proof, a program, or an argument. Each step changes which features and relations should come next.
Language illustrates the compositional advantage. A word does not carry one fixed meaning independent of its surroundings. Its role depends on where it appears in a sentence, which larger construction contains it, and which conceptual model the listener currently activates. Nested reference frames could provide a common format for relating phonemes to words, words to clauses, clauses to narratives, and narratives to knowledge about the world.
This is an explanatory sketch rather than a worked-out theory of language. The book does not derive grammar, semantics, or reasoning from identified circuits. Its contribution is a proposed representational primitive: location within a structured model. The same primitive could support physical recognition, composition, analogy, and mental traversal if the relevant neural mechanisms exist.
The philosophical result is significant. Intelligence becomes less a collection of specialized tricks and more a general capacity to build models whose parts have stable relations, to move through those relations, and to combine many partial models into a coherent judgment.
7. Machine Intelligence: World Models Without Human Drives
Hawkins treats neuroscience as an engineering resource. If the neocortex implements a general learning algorithm, then artificial systems could reproduce its principles without reproducing every biological detail. The resulting machines would learn continuously through sensorimotor interaction, represent knowledge in reference frames, use sparse distributed activity, and combine the judgments of many parallel modules.
This differs from much of the deep-learning paradigm described in the book. A conventional model is often trained on a large, fixed dataset and then deployed with mostly fixed parameters. A brain-inspired agent would acquire models incrementally while acting, use movement to gather informative data, and update knowledge without a clean separation between training and operation.
Hawkins also separates intelligence from agency. The neocortex supplies models; older systems supply survival-related drives. An artificial world-modeling system therefore need not inherit hunger, reproductive competition, dominance, or fear of death. It can be highly intelligent without spontaneously developing human-like goals.
That distinction weakens one route to machine risk, but it does not settle AI safety. Goals can be designed, learned, delegated, or created by institutions surrounding a system. A machine without biological drives can still pursue a harmful objective with great competence. Hawkins’s argument is best read as an architectural correction: capability, consciousness, agency, and motivation are separate design questions. Treating them as one quantity obscures where risk enters.
His view of machine consciousness follows the same functional logic. If consciousness depends on the brain’s model of its own attention and state, a machine with comparable models might report subjective-like awareness. The book regards this as scientifically approachable, yet it does not resolve the philosophical problem of why any information processing should be accompanied by felt experience.
8. False Beliefs and the More Immediate Existential Risk
The third part of the book shifts from machine intelligence to the dangers created by human intelligence. A model-building brain does not guarantee a true model. People can hold beliefs that are internally coherent, socially reinforced, and resistant to contradictory evidence. The same capacities that enable science also enable ideology, conspiracy, and elaborate rationalization.
Hawkins distinguishes knowledge stored in brains and culture from goals shaped by older evolutionary systems. Knowledge can accumulate rapidly across generations; genetic change is slow. Human beings therefore command technologies of planetary scale while retaining drives formed under conditions of small-group competition, short horizons, and local scarcity.
This mismatch reframes existential risk. Nuclear weapons, ecological disruption, engineered pathogens, and destructive institutions do not require a malicious superintelligence. They require powerful knowledge coupled to badly coordinated goals or false beliefs. Intelligence expands the space of possible action faster than wisdom automatically expands the space of restraint.
The book’s response is an “estate plan” for humanity: protect and extend knowledge beyond the lifespan of individuals, institutions, and perhaps Earth itself. Hawkins imagines durable records and settlements beyond one planet as ways to reduce the chance that a single catastrophe erases what humanity has learned.
The proposal is intentionally long-term, but its near-term lesson is institutional. Reliable knowledge depends on systems that can detect error: open criticism, reproducibility, distributed archives, education, and the ability to revise public models. A society needs its own version of multiple-model voting, with a crucial addition—the votes must remain answerable to evidence rather than mere repetition.
9. A Contemporary Interpretation for Robotics and AI
Hawkins makes direct claims about future machine intelligence, but the following applications extend the book into today’s research landscape. They are interpretations, not results established by the book itself.
Embodied Robotics: Learn by Moving to Know
Many robotics failures arise because a visual label is mistaken for an actionable model. A robot may recognize “mug” while failing to anticipate how its viewpoint will change, where contact will occur, or how the object will move under force. Reference-frame learning suggests that perception should encode features together with their positions relative to objects, bodies, and tasks.
Active perception then becomes part of policy. The robot can move a camera, reposition a hand, or probe an uncertain surface to distinguish competing models. An action serves two purposes at once: changing the world and discovering which world the robot is in.
Modular Intelligence: Many Learners, Negotiated Belief
The cortical-column metaphor suggests systems composed of repeated learning modules with a shared communication protocol. Modules may specialize through their inputs while retaining a common ability to learn structured models. Their outputs become hypotheses that can be combined, challenged, and updated.
This architecture offers potential benefits—parallelism, local learning, multimodal integration, and robustness—but also creates a coordination problem. Voting is not magic. A practical system must specify what a hypothesis contains, how confidence is calibrated, how contradictory frames are aligned, and how a minority module with decisive evidence can overturn a confident majority.
Continual Learning: Training and Deployment as One Life
A sensorimotor agent encounters new objects, changing tools, and unfamiliar environments after deployment. The thousand-brains perspective treats this as the normal condition of intelligence. Learning should be rapid and associative, protect useful prior models, and let new modules or frames absorb novelty without globally retraining the system.
Recent Thousand Brains Project work turns these principles into an open research program built around modular sensorimotor learning, reference frames, and communication among learning modules. This is valuable evidence that the book’s ideas can generate implementable hypotheses. It is not yet evidence that the resulting systems reproduce general neocortical intelligence.
Foundation Models: Complement Rather Than Simple Replacement
Large pretrained models learn broad statistical structure and can support planning, language, and multimodal inference. Hawkins’s framework highlights what pretraining alone leaves unresolved: grounded reference frames, learning through self-directed movement, stable continual adaptation, and models tied to physical consequences.
A productive synthesis may use foundation models for prior knowledge and communication while sensorimotor modules maintain local, revisable models of bodies, objects, and tasks. The important comparison is empirical: which architecture learns faster, transfers better, recovers from surprise, and remains corrigible in an open world?
10. What the Theory Does Not Yet Establish
The book is strongest when it gives many observations a common direction. Its elegance can also tempt readers to treat a research program as a completed explanation.
First, the status of the cortical column is contested. Neuroscientists use “column” for several anatomical and functional patterns, and no universally accepted canonical unit has been shown to perform one operation everywhere in the neocortex. Hawkins explicitly uses the term as a functional convenience, but the theory still depends on some repeatable local organization doing the proposed work.
Second, evidence for grid cells in the entorhinal cortex does not by itself demonstrate grid-cell-like location codes throughout the neocortex. The broader claim produces testable predictions; it remains a hypothesis whose scope, cell types, and circuit mechanisms need direct evidence.
Third, moving from objects to abstract thought requires more than metaphor. Reference frames provide an attractive language for composition and mental movement, yet a full account must explain how particular neural populations encode relations, variables, logical operations, and linguistic structure.
Fourth, consensus does not guarantee truth. Many models can share the same bias or be driven by correlated evidence. Both brains and multi-agent machines need mechanisms for uncertainty, anomaly detection, exploration, and correction—not only agreement.
Finally, a theory of neocortical intelligence is not automatically a theory of the whole mind. Emotion, memory systems, development, social cognition, bodily regulation, and consciousness involve circuits beyond the simplified old-brain/new-brain division. The thousand-brains theory may explain an important organizing principle without being the only principle that matters.
These limits do not make the framework unproductive. They tell us what kind of framework it is: a compact generator of experiments and architectures, valuable in proportion to the precise predictions it exposes to failure.
Conclusion: Intelligence as Coordinated Model-Building
A Thousand Brains changes the unit from which we imagine intelligence. The basic picture is no longer a passive sensor feeding a central classifier. It is an embodied learner that moves through the world, locates features inside reference frames, builds multiple models, predicts the consequences of movement, and reaches a working consensus without a single model owning the whole truth.
| Hawkins’s idea | Design question it creates |
|---|---|
| Feature at a location | Does the representation preserve structure relative to objects and bodies? |
| Sensorimotor learning | Can action reduce uncertainty as well as achieve a goal? |
| Many complete models | Can knowledge remain distributed without becoming incoherent? |
| Voting across modules | How are confidence, conflict, and decisive minority evidence handled? |
| Nested reference frames | Can the same machinery compose parts, objects, scenes, and abstractions? |
| Intelligence separated from drives | Where do goals enter, and who can revise them? |
| Knowledge vulnerable to false belief | Which institutions keep collective models corrigible? |
For neuroscience, the theory asks whether a common location-based computation can truly span the neocortex. For AI, it asks whether intelligence should be built as a population of continually learning world models. For robotics, it makes movement part of knowing rather than a command issued after perception is complete.
The generative question is not whether the brain literally contains a thousand little minds. It is: What becomes possible when no single learner has to model the whole world, yet many learners can move, compare, and correct one another?
Further Reading
- Hawkins, J. (2021). A Thousand Brains: A New Theory of Intelligence. New York: Basic Books.
- Hachette / Basic Books: A Thousand Brains
- Hawkins, J., Lewis, M., Klukas, M., Purdy, S., & Ahmad, S. (2019). A Framework for Intelligence and Cortical Function Based on Grid Cells in the Neocortex. Frontiers in Neural Circuits, 12, 121.
- Numenta: The Thousand Brains Theory of Intelligence
- Clay, V., Leadholm, N., Hawkins, J., et al. (2024). The Thousand Brains Project: A New Paradigm for Sensorimotor Intelligence.
- Horton, J. C., & Adams, D. L. (2005). The Cortical Column: A Structure Without a Function. Philosophical Transactions of the Royal Society B, 360, 837–862.
