Workspace of Scenes
A Brain-Inspired Architectural Hypothesis for Artificial Intelligence – Preprint
A preprint
Workspace of Scenes: A Brain-Inspired Architectural Hypothesis for Artificial Intelligence
June 23, 2026
Abstract
This paper proposes Workspace of Scenes, a high-level conceptual architecture for interpreting human-like intelligence. Workspace of Scenes is intended as a high-level vocabulary and architectural hypothesis for describing cognition in terms of coordinated Scene-like neural population states, rather than as a full implementation specification. The central idea is that intelligence can be understood as the continuous construction, interpretation, comparison, simulation, navigation, and transformation of Scenes. In this framework, a Scene is not only a visual landscape, but a structured cognitive entity representing objects, agents, relations, contexts, goals, expectations, attention, predictions, plans, possible futures, and possible actions.
The hypothesis uses a software architecture perspective to describe cognition as interaction between Scenes. It interprets perception, memory, reasoning, imagination, planning, and action as operations over Scene-based representations organized by Reference Frames and navigable relations.
The core contribution is not the claim that the brain literally stores discrete symbolic scenes. Rather, it is the architectural hypothesis that active representations, memories, predictions, plans, goals, and candidate actions can be treated as Scene-like structures whose content is organized by Reference Frames and whose compact Scene Codes support navigation through Scene Relations.
Workspace of Scenes is compatible with world model approaches, because it treats intelligence as depending on structured internal representations of the environment, the agent, and possible future states. This paper is intended as a foundation for a broader research and implementation program. It does not present a complete implemented AI system or an established neuroscience theory, but develops a conceptual foundation for future formalization, implementation design, comparison with existing cognitive and AI architectures, and empirical evaluation.
1. Introduction
A brain-inspired architecture for Artificial Intelligence should not treat the brain as a single monolithic neural network. The brain is an organized system of interacting regions, pathways, recurrent loops, memory systems, attention mechanisms, and action-selection circuits. A useful architectural hypothesis should therefore describe not only local neural computation, but also how larger cognitive functions are coordinated.
Current Artificial Intelligence systems are usually described in terms of models, training objectives, data flows, and inference procedures. Those descriptions are powerful, but they do not directly explain how an intelligent system could organize perception, memory, prediction, attention, and action into a coherent interpretation of an ongoing situation. Workspace of Scenes approaches this question at a different level: as a software architecture problem.
At this level, the important questions are architectural: what kinds of representations are active, how they are selected, how they are related to memory, how they are updated by sensory input, how they can be transformed internally, and how they can influence action. The notion of a Scene is introduced as a candidate abstraction and common vocabulary for situated structure.
Although the paper is framed primarily as an Artificial Intelligence architectural hypothesis, it also argues that Scene-level abstraction may be useful for interpreting core mechanisms of biological intelligence. This claim is deliberately high-level: it suggests that perception, memory, prediction, attention, simulation, and action may be understood as coordinated operations over Scene-like structures.
The hypothesis is motivated by several converging lines of work. Global Workspace Theory emphasizes the availability of selected information across specialized systems. Thousand Brains Theory emphasizes distributed models organized around sensorimotor Reference Frames. Predictive-processing approaches emphasize expectation and error correction. Hippocampal and Cognitive-Map theories emphasize indexing, navigation, and relational structure. Workspace of Scenes draws architectural inspiration from these ideas.
The contribution of Workspace of Scenes is not a new low-level neural algorithm, nor a claim that the brain stores explicit symbolic scenes. Its purpose is to offer a reusable cognitive vocabulary and architectural lens, not to disclose a full implementation. Its contribution is a high-level architectural vocabulary for describing coordinated neural population activity as Scene-like structures. In this vocabulary, perception, memory, prediction, imagination, planning, and action are interpreted as role-dependent operations over structured active states, compact Scene Codes, and navigable Scene Relations. These hypotheses can later be formalized, implemented, compared with existing approaches, and evaluated empirically.
How to read this paper.
Readers interested in the core hypothesis can focus on Sections 1–4, 9, and 10. Readers interested in cognitive interpretation can continue through Sections 5–7. Readers interested in implementation and evaluation should focus on Sections 8–9 and Appendix E. The appendices are intended as illustrative extensions and implementation context, not as additional core claims.
2. Core Claim, Scope, and Assumptions
Workspace of Scenes uses Scene as an architectural abstraction for structured, situated cognitive states. A Scene may represent a visible layout, but it may also represent a social situation, an abstract relation, a remembered episode, a candidate plan, a bodily state, a prediction, a goal, or a possible future. This broader use is essential to the hypothesis.
The core claim is that many cognitive functions can be described in a unified way if active states, remembered states, predicted states, and candidate actions are all treated as structured Scene-like entities. The hypothesis further claims that these entities require both detailed active content and compact navigable references: Scene Content supports comparison, integration, and transformation, while Scene Codes and Scene Relations support retrieval, sequencing, planning, and simulation.
Workspace of Scenes is based on several architectural assumptions. They define the level of abstraction used by the hypothesis and identify the claims to be refined, implemented, or tested in future work.
First, the hypothesis assumes that cognition can be usefully described in terms of structured, situated representations called Scenes. A Scene is treated as an architectural abstraction: it may correspond to activity distributed across many neural populations, rather than to a single localized structure.
Second, it assumes that Scenes can be constructed, updated, compared, recalled, composed, transformed, and simulated.
Third, it assumes that Reference Frames are central to cognition. Spatial, temporal, bodily, social, and abstract relations are interpreted as forms of organization that allow Scenes to be navigated and related to one another.
Fourth, it assumes that perception, memory, reasoning, planning, and action are not separate symbolic modules, but interacting processes over shared Scene-like structures.
Finally, candidate biological mappings are assessed at a broad architectural level and remain open questions for comparison with neuroscience.
2.1. Minimum Criteria and Boundaries for Scenes
Because Scene is used more broadly than visual scene, the term needs explicit boundaries. In this paper, a candidate state counts as a Scene only when it satisfies the following minimum criteria:
it contains structured content rather than only an undifferentiated scalar, label, or activation level;
its content is organized by a Reference Frame or equivalent relational basis;
it can be associated with a compact Scene Code or retrievable handle;
it can participate in Scene Relations, such as similarity, transition, containment, transformation, prediction, or causal association; and
it has an operational role in comparison, integration, recall, prediction, simulation, planning, evaluation, or action.
Under this criterion, a raw sensory stream, a single reward value, an isolated motor command, a free-floating word label, or a local detector activation is not by itself a Scene. Such signals may contribute to Scene Content, bias Scene construction, select a Scene role, or become represented within a Scene, but they do not qualify as Scenes unless they are structured, situated, retrievable, relationally connected, and operationally usable.
2.2. Claim Types
The paper combines architectural claims, possible implementation choices, and biological speculation. Table 1 separates these levels so that the paper’s biological vocabulary is not read as a stronger claim than intended.
| Claim type | Role in the paper | Example |
|---|---|---|
| Core architectural claim | Required for Workspace of Scenes as proposed. | Cognition can be modeled as recurrent operations over structured Scene-like states, compact Scene Codes, and navigable Scene Relations. |
| Implementation hypothesis | One plausible way to realize the architecture in biological or artificial systems. | Scene Codes may function like compact hippocampal-style indexes that can reactivate distributed Scene Content. |
| Biological speculation | Possible neural correspondence that motivates experiments but is not required by the architecture. | Some role-dependent routing effects may be supported by thalamic gating or dendritic compartmentalization. |
3. Related Work and Positioning
Workspace of Scenes is related to several existing lines of research, but it is not intended as a replacement for any of them. Its role is to provide an architectural vocabulary that can connect ideas from brain theory, cognitive science, and Artificial Intelligence.
Global Workspace Theory and Global Neuronal Workspace models emphasize selective global availability of information across specialized systems, especially in relation to conscious access (Baars 1988; Dehaene and Changeux 2011). Workspace of Scenes is influenced by this broad coordination problem, but it does not attempt to define consciousness. Its emphasis is on the structure, role, retrieval, transformation, and evaluation of the content that becomes available for processing.
The closest brain-inspired influence is the Thousand Brains Theory and related work on cortical columns, sensorimotor inference, and grid cell-like Reference Frames in the neocortex (Hawkins et al. 2019; Hawkins 2021; Lewis et al. 2019; Clay et al. 2024; Leadholm et al. 2025). These theories emphasize that perception is not produced by a single centralized model, but by many local models that use movement, location, and voting to infer the structure of objects and situations. Workspace of Scenes is compatible with this direction, but shifts the emphasis from individual column-level models to Scene-level integration across perception, memory, planning, and action.
The hypothesis is also related to work on hippocampal mapping, grid cells, path integration, and cognitive maps (Bicanski and Burgess 2019; Klukas et al. 2020; Lewis 2021; Leadholm et al. 2021). These approaches support the idea that spatial navigation mechanisms may generalize to abstract, relational, and conceptual spaces. Workspace of Scenes uses this as motivation for treating Reference Frames and navigation as general architectural principles, not only mechanisms for physical movement.
Another related line of work concerns predictive processing, active inference, active dendrites, sparse representations, and biologically richer neuron models (Friston 2010; Hawkins and Ahmad 2016; Grewal et al. 2021; Iyer et al. 2022; Rolls 2021). Free-energy and active-inference approaches formulate perception, action, and learning through prediction and uncertainty minimization. Workspace of Scenes overlaps with these themes, but it does not begin from a single optimization principle. It instead asks what representational objects and routing relationships would be needed for prediction, mismatch, simulation, and action to operate over structured situations.
Object-file and event-file theories are relevant because they describe temporary structured representations that bind features of objects or perception-action episodes over time (Kahneman et al. 1992; Hommel 2004). Workspace of Scenes generalizes in a different direction: it treats object-like, event-like, planned, remembered, and counterfactual structures as special cases of a broader Scene abstraction, while requiring explicit Reference Frames, Scene Codes, and Scene Relations.
The term Scene also overlaps with visual-scene perception and contextual object-recognition research (Henderson and Hollingworth 1999; Oliva and Torralba 2007). Workspace of Scenes deliberately uses Scene more broadly than this literature. A visual scene is one important case, but the proposed abstraction also covers abstract, social, bodily, goal-oriented, and procedural structure when it satisfies the minimum criteria defined above.
The thalamus and cortico-thalamic loops are also relevant, especially in relation to routing, attention, deviance detection, and blackboard-like coordination (Worden et al. 2021; Varela et al. 2024). Workspace of Scenes does not depend on a specific thalamic implementation, but it treats routing and selective availability of Scene information as important architectural concerns.
In Artificial Intelligence, the hypothesis is closest in spirit to world model approaches and attempts to build systems that learn structured internal representations for prediction and planning (LeCun 2022; Dawid and LeCun 2023). Workspace of Scenes differs from these approaches by focusing first on the architectural abstraction of Scene-based organization, rather than on a specific training objective, neural network architecture, or implementation strategy.
Classical cognitive architectures and blackboard systems provide another important comparison point (Nii 1986; Newell 1990; Laird et al. 1987; Laird 2012; Anderson et al. 2004; Baars and Franklin 2009). ACT-R, Soar, LIDA, and related production-system or workspace architectures specify interacting memories, buffers, productions, goals, and action-selection mechanisms. Workspace of Scenes differs by making structured Scene Content, compact Scene Codes, Reference Frames, and navigable Scene Relations the central architectural objects, while leaving the production or control substrate open.
Table 2 summarizes the intended novelty relative to these influences. The point is not that Workspace of Scenes supersedes them, but that it combines several concerns that are often treated separately: the internal structure of active content, compact memory indexing, navigable relations among states, role-aware routing, and evaluation-guided action.
| Influence | Primary emphasis | Workspace of Scenes emphasis |
|---|---|---|
| Global Workspace / Global Neuronal Workspace | Broad availability of selected information across specialized systems, often in relation to conscious access. | Less concerned with consciousness as such; more concerned with the structured content being made available, the roles it can take, and how it can guide prediction, memory, planning, and action. |
| Thousand Brains Theory | Distributed cortical models, sensorimotor inference, location signals, and voting among many local models. | Shifts the focus from object-level columnar recognition toward Scene-level integration across perception, memory, simulation, planning, and action. |
| Predictive processing and active-dendrite models | Prediction, contextual modulation, feedback, mismatch, and richer local neural computation. | Uses prediction and mismatch as operations over Scenes, while giving an explicit architectural vocabulary for the objects being predicted, compared, and transformed. |
| Active inference / free-energy approaches | Unified optimization accounts of perception, action, attention, learning, and uncertainty. | Shares the concern with prediction and action, but frames Workspace of Scenes around Scene objects, routing, and relations rather than a single optimization principle. |
| Object files, event files, and visual-scene cognition | Temporary object/event representations, feature binding, visual scene perception, and contextual object recognition. | Generalizes from visual or object-specific representations to structured Scene-like states with explicit Reference Frames, codes, relations, and roles. |
| Hippocampal indexing and cognitive maps | Compact indexing, pattern separation/completion, relational memory, and navigation through physical or abstract spaces. | Treats indexing and navigation as a Mapping Module over Scene Codes and Scene Relations that can reactivate detailed Scene Content when needed. |
| World models and predictive AI architectures | Learning internal representations that support prediction, planning, and abstraction. | Focuses less on the training objective and more on the representational and routing architecture needed for Scene construction, retrieval, simulation, evaluation, and action selection. |
| Classical cognitive architectures and blackboard systems | Integrated memories, working buffers, productions, problem spaces, global workspaces, or shared blackboard structures. | Shares the goal of system-level coordination, but uses Scene Content, Scene Codes, Scene Relations, and Reference Frames as the primary coordination objects. |
4. Core Concepts
Workspace of Scenes assumes the standard agent-environment loop, but focuses on the internal architecture that mediates between perception and action: how incoming sensory information updates Scenes, how memory and simulation operate over them, and how Scene structure can guide behavior. Figure 1 situates the architecture within this wider loop.
Figure 2 expands the Workspace of Scenes element from Figure 1 into its principal architectural components and recurrent interactions.
4.1. Minimal Formal Model
This section gives a deliberately compact formal model. It is not meant to specify a final neural or software data structure. Its role is to define the computational objects that later sections discuss more descriptively.
A Feature Instance can be represented as:
where is a Feature Instance, is a Feature identifier, is a location or coordinate-like value within a Reference Frame, is a confidence value, and records source, role, or other contextual information when needed.
Scene Content is a sparse set of such instances:
A Scene can then be treated as:
where is Scene Content, is the Reference Frame or relational organizing basis, is an optional Scene Code or temporary Scene Code, is the current role configuration, and records status information such as active, predicted, simulated, temporary, durable, or decaying. A Scene Code is a compact reference:
A Scene Relation can be represented as:
where and are source and target Scene Codes, is the relation type, is an optional Reference Frame transform or delta, is strength or confidence, records provenance such as the operation, evidence source, or context that created or strengthened the relation, and is an optional relation Scene Code when the relation’s internal structure must be reactivated. A path Scene is a Scene whose content or Reference Frame organizes Scene Codes and Scene Relations as an ordered, temporal, causal, procedural, spatial, or otherwise navigable structure.
The primitive operations of the hypothesis can be summarized as follows:
| Operation | Schematic form | Architectural role |
|---|---|---|
| Encoding | Generate a compact Scene Code from structured content in a Reference Frame. | |
| Activation | Reactivate or construct detailed Scene Content from a Scene Code under a contextual cue, role, or sensory input. | |
| Comparison | Identify overlap, mismatch, missing content, added content, or transformation between Scenes. | |
| Integration | Combine or align Scenes under a role configuration to produce candidate Scene Content. | |
| Transformation | Express or modify Scene Content relative to another Reference Frame. | |
| Simulation | Generate a possible future, counterfactual, or internally modified Scene by following or applying a candidate Scene Relation. | |
| Evaluation | Estimate a priority value for a Scene relative to active goals, constraints, costs, significance signals, or expected value. |
Table 4 gives a compact orientation to the main terms before the detailed definitions. Some terms are architectural primitives in Workspace of Scenes, while others name recurring roles or operations over those primitives.
| Term | Role in Workspace of Scenes | Possible biological correlate | AI/software analogue |
|---|---|---|---|
| Feature | Detectable component that can participate in Scene Content. | Local detector, minicolumn, or population-level hypothesis. | Feature, predicate-like detector, learned latent factor. |
| Reference Frame | Navigable organization that situates Feature Instances and supports transformations. | Grid cell-like, location, displacement, and coordinate-like mechanisms. | Coordinate system, latent manifold, graph embedding, state space. |
| Scene | Active structured state representing a situation, object, memory, prediction, plan, relation, or possible action. | Distributed cortical activity pattern across specialized regions. | Structured latent state or world-state representation. |
| Scene Content | Detailed active Feature Instances arranged in a Reference Frame. | Sparse distributed cortical activity. | Active working representation or content graph. |
| Scene Code | Compact reference to a recorded or candidate Scene. | Hippocampal-like sparse index capable of reactivating cortical content. | Key, handle, embedding, hash-like index, memory pointer. |
| Scene Relation | Navigable link between Scene Codes, optionally with transformation or provenance. | Hippocampal–entorhinal relational structure. | Edge, transition, transform, association, memory-graph link. |
| Integration Workspace | Content-level substrate where Scene Content is activated, compared, transformed, and integrated. | Neocortex and associated recurrent cortical systems. | Working latent workspace or active state processor. |
| Relay Module | Routing, gating, timing, and role-aware coordination among systems. | Thalamus, cortico-thalamic loops, local routing and inhibition. | Router, attention/gating layer, message broker. |
| Mapping Module | Fast index and navigation system over Scene Codes and relations. | Hippocampal–entorhinal system. | Memory graph, associative index, retrieval and planning module. |
| Evaluation and Action Selection | Prioritizes candidate Scenes, internal actions, and overt actions. | Prefrontal, basal-ganglia, limbic, and neuromodulatory loops. | Policy, value, constraint, and controller systems. |
Table 5 separates primitive objects from derived Scene forms and operational roles. This hierarchy is intended to reduce ambiguity among terms such as Scene Relation, relation Scene, path Scene, translation Scene, and routing configuration Scene.
| Level | Terms | Interpretation |
|---|---|---|
| Primitive content objects | Feature, Feature Instance, Reference Frame, Scene Content, Scene | The structured active content that can be represented, situated, compared, transformed, or integrated. |
| Primitive mapping objects | Scene Code, Scene Relation | Compact references and navigable links used for retrieval, sequencing, similarity, transformation, and provenance. |
| Derived Scene forms | relation Scene, path Scene, translation Scene, routing configuration Scene | Scenes whose content represents relations, ordered structures, transformations, or reusable routing patterns. They are special cases of Scene, not separate primitive kinds. |
| Operational roles | evidence Scene, context Scene, prediction Scene, goal Scene, constraint Scene, candidate action Scene | Temporary roles that determine how active Scene Content influences a current operation. The same Scene can serve different roles in different cycles. |
| Architectural systems | Integration Workspace, Mapping Module, Relay Module, Scene Significance Evaluation, Evaluation and Action-Selection System | Components or processes that construct, route, navigate, evaluate, and select among Scenes and Scene Relations. |
4.2. Feature
A Feature is a detectable or activatable component of representation. In one biological interpretation, a Feature is the kind of content that a local cortical detector could support: something a neocortical minicolumn, or an equivalent local circuit, has learned or is evolutionarily prewired to detect from sensory or lower-order input. Features may represent simple sensory regularities, relations, actions, higher-order concepts, or other information that can participate in a Scene.
A Feature Unit is the local detector or population that realizes one such Feature hypothesis inside a cortical column. The Feature is the architectural content. The Feature Unit is a possible local implementation that detects, predicts, and reports that content under particular location and contextual conditions. Feature Units are discussed in detail in a later section.
A Feature is distinct from an object. In this architecture, object is used in its ordinary sense: a real or inferred thing in the world, such as a cup, face, hand, door, tool, or person. It is not a representational domain name or namespace. An object can be represented by one or more Scenes that capture its parts, states, relations, affordances, and contexts. Individual Features contribute to those Scenes but are not themselves the whole object.
A Feature Instance is the occurrence of a Feature at a location in a Scene’s Reference Frame. The same Feature Unit or detector may participate in representing multiple Feature Instances at different locations. This distinction allows a Scene to contain many instances of the same detected Feature without requiring a separate detector for each instance.
4.3. Reference Frame
A Reference Frame is the navigable structure that organizes Feature Instances within a Scene. It allows locations, directions, distances, and transitions to be represented and permits deltas or transformations to be calculated between related Scenes. Each Scene is assumed to have one primary Reference Frame, although related Scenes may express similar content in different Reference Frames.
The biological inspiration is the grid cell-based representation of physical location and proposals that similar mechanisms can represent higher-dimensional cognitive variables (Hawkins et al. 2019; Klukas et al. 2020). In Workspace of Scenes, a Reference Frame is therefore treated as a flexible, hierarchical, and potentially multidimensional grid-like structure. Its dimensions may be spatial, temporal, bodily, social, conceptual, or otherwise abstract.
A Reference Frame is not assumed to be a fixed coordinate system. During learning, it may acquire new distinguishable values within an existing dimension, or introduce a new dimension when recurring correlations make another variable useful for organizing represented content. For example, a perceptual Reference Frame could become more refined as new colors are distinguished among already known colors, while a multimodal or abstract Reference Frame could add a dimension when a new regular relation becomes useful. Values and dimensions may also weaken or disappear when they are no longer supported. Some dimensions may be cyclic, such as orientation, phase, or movement around a closed surface. Direction, point of view, or perspective may be represented as Reference Frame variables or as contextual inputs, but the architecture does not yet prescribe one mechanism.
A variation of existing represented content expressed in a different Reference Frame is treated as a related Scene rather than as the same Scene. An operation may transform a Scene into another Reference Frame, producing a relation that records the transformation between the resulting Scenes. Scenes can also participate in a higher-order Scene whose Reference Frame provides a common structure for relating their local Reference Frames.
The hypothesis currently assumes grid-like navigable Reference Frames because navigation is one of its central architectural principles. Whether more general graph or network structures can serve as Reference Frames while preserving the required navigation and transformation operations is left for further research.
4.4. Scene
A Scene is a structured state that can be active in the Integration Workspace. In the main biological interpretation, its content would correspond to a sparse, distributed pattern of active Feature Units, possibly implemented by neocortical minicolumns or cooperating local populations. Each participating Feature Unit represents a Feature that it has learned, or is evolutionarily prewired, to detect from sensory or lower-order input, and reports that Feature under particular location and contextual conditions. A Scene therefore combines situated outputs from many Feature Units across cortical regions into one representation.
A Scene is not limited to a visual environment. It may represent an object, situation, remembered episode, imagined possibility, abstract relation, plan, or other structure that can be organized in a Reference Frame. Any representable entity or process may have one or more associated Scenes when its internal structure is relevant. This recursive property allows a Feature within one Scene to be navigated into and represented by another Scene.
4.4.1. Scene Content
Scene Content is the distributed content of a Scene while it is active in the Integration Workspace. Architecturally, it is composed of Feature Instances. In the biological interpretation developed here, those instances may be represented by active minicolumns or local populations. A minicolumn does not merely indicate that its Feature is present or absent. It may represent multiple instances of the same Feature at different locations in the Scene’s Reference Frame. This interpretation extends the feature-at-location approach proposed in Thousand Brains Theory and related models (Hawkins et al. 2019; Lewis et al. 2019).
Consequently, Scene Content can be interpreted at a high level as a sparse collection of pairs: Feature detector, set of locations in the Scene Reference Frame.
The hypothesis assumes that grid cell-like mechanisms within, or functionally associated with, cortical columns can provide the required location codes, while leaving the precise biological representation open.
A Feature Instance may be associated with one or more other Scenes. Navigating into the Feature Instance activates an associated Scene, exposing structure that was not active in the previous Scene. Conversely, navigation to a containing or otherwise related Scene can replace or supplement the current Scene Content. The associations that support such navigation are assumed to be maintained by the Mapping Module.
4.4.2. Scene Code
Each recorded Scene is identified by a Scene Code: a sparse code generated from its distributed Scene Content, including active Feature detectors and their locations. Because locations contribute to the code, Scenes containing the same Features at different locations can receive different Scene Codes.
As an implementation hypothesis, the Scene Code may be a compact hippocampal-like index of active neocortical content. It must preserve enough structure to support approximate matching, generalization, and later reactivation of the corresponding cortical pattern. Hippocampal indexing theory provides related motivation for treating hippocampal activity as an index capable of reactivating distributed neocortical activity (Teyler and Rudy 2007). Findings and models concerning sparse coding, pattern separation, pattern completion, and compressed representations in the hippocampal system provide possible mechanisms for generating distinct codes and reconstructing related activity patterns (Bakker et al. 2008; Petrantonakis and Poirazi 2014). These results are consistent with, and motivate, the bidirectional Scene Code mechanism proposed here, but they do not establish it.
Every detected change in sensory input, or candidate state produced by an internal operation, initially receives a temporary Scene Code. This allows the architecture to treat every change as potentially important before its later consequences are known. A temporary Scene Code provides an immediately available reference to the candidate Scene and may acquire temporary relations to preceding or source Scenes. Scene Significance Evaluation and later consolidation influence whether that code and its relations are strengthened, retained, reorganized, or allowed to decay.
The architecture may also need Scene Codes to express negative constraints, such as an expected Feature that must not occur at a location. Such information differs from the simple inactivity of a detector. One possible mechanism is that comparison between a predicted Scene and incoming sensory content produces a mismatch or surprise signal that becomes part of the newly recorded Scene. Whether Scene Codes should encode such negative information directly, or refer to a separate constraint representation, remains an open design question.
4.4.3. Scene Properties
Distributed and sparse. A Scene is represented by a sparse pattern distributed across the Integration Workspace rather than by a single localized representation.
Situated. Feature Instances are interpreted at locations within a Reference Frame.
Relatively static. An active Scene represents a state at a useful processing timescale. Temporal structure is represented primarily through relations among Scenes.
Hierarchical and recursive. A Scene can contain Features associated with lower-order Scenes and can itself participate in higher-order Scenes.
Multiscale. Related Scenes may represent the same entity or process at different spatial, temporal, or conceptual scales.
Multidimensional. A Scene’s Reference Frame may contain multiple physical or abstract dimensions.
Partial. Scene Content need not completely describe what it represents. It can be extended, corrected, or combined with other active Scenes.
Navigable. Features, locations, and relations provide possible transitions within the Scene and to related Scenes.
4.4.4. Scene Relation
Scene Codes are connected by navigable Scene Relations. In this hypothesis, a Scene Relation is the primitive mapping-level link between Scenes. A relation may represent a sensory update, change of attention, transformation, composition, generalization, causal hypothesis, association, or another transition discovered during processing. An operation that produces a recorded Scene also records relations between that Scene and the source Scenes involved in the operation.
Most Scene Relations are compact mapping-level links rather than complete Scene Content. When a relation has reusable or recognizable internal structure, it may refer to a relation Scene. A relation Scene is a higher-order Scene whose content represents correspondences, deltas, constraints, or operations between one or more source and target Scenes. Its atomic elements may include Feature Instance relations such as adding, removing, moving, modifying, matching, conflicting, or substituting Feature Instances. This allows transformations and differences between Scenes to be learned, recalled, generalized, and composed without requiring every Scene Relation to be stored as a full Scene.
A path Scene is a higher-order Scene whose Reference Frame organizes references to other Scenes as an ordered or otherwise navigable pattern. Its dimensions may be temporal, sequential, spatial, procedural, causal, linguistic, or abstract. Such path Scenes can represent episodic sequences, learned procedures, routes, observed changes, possible futures, or reasoning traces without requiring the complete Scene Content of every step to be simultaneously active.
Under this interpretation, behavior does not require a separate representational primitive. A behavior can be represented as a Scene or path Scene whose Reference Frame includes time, sequence position, state, or procedural dimensions. Long-running behavior may therefore be hierarchical: a slower higher-order path Scene can organize faster sub-paths, motor fragments, or internal cognitive operations.
The ordered structure is therefore not an additional primitive next to Scenes. It is represented as a Scene with its own Scene Code, while its component Scenes retain their own Scene Codes and may be connected by Scene Relations in the Mapping Module. A sequence of states belonging to one continuing entity or situation is a path Scene whose Reference Frame tracks continuity and change along temporal or state dimensions.
A path Scene may contain Scene references and relations with different degrees of persistence. Significant states may become durable anchor Scenes, while intermediate temporary Scene Codes and relations may decay. Recalling a period may therefore require navigating from an anchor Scene and reconstructing less persistent intermediate states.
A path Scene may be replayed by successively activating its component Scenes. When a path Scene represents a learned procedure, such as a theatrical play, musical performance, or practiced movement, its replay may guide motor output. Path execution differs from passive replay because incoming sensory feedback can select alternative Scene Relations, update subsequent Scenes, or create a modified path Scene.
The main operations over this map can be described compactly. Scene binding creates or strengthens a Scene Relation. Path formation records an ordered or otherwise navigable pattern of related Scene Codes and relations as a path Scene. Path composition joins compatible path Scenes or fragments, while path recombination constructs a candidate path Scene by reusing, replacing, or rearranging existing fragments. These terms name operations over Scene Relations and path Scenes rather than separate fundamental structures. Selected component Scenes can then be activated to evaluate whether a candidate path Scene is coherent, useful, or compatible with current evidence, goals, and constraints.
4.4.5. Scene Translation
Scene Translation constructs a target Scene or path Scene from a source Scene or path Scene by applying learned relations between their Features, Reference Frames, or path Scene structures. Translation may preserve selected relational structure while changing the representational domain. Examples include translating meaning between languages, translating written instructions into actions, or translating a musical score into a performed motor sequence.
Scene Translation is therefore an operation that can use learned relation structure, rather than a separate primitive memory type. When the translation process itself becomes reusable, it may be represented by a translation Scene: a relation Scene specialized for constructing target Scene Content from source Scene Content. Such a translation Scene may encode correspondences among source and target Features, transformations between Reference Frames, and rules for preserving, replacing, adding, or suppressing Feature Instances. The result of a translation is still recorded as a target Scene or path Scene, connected to its sources and to the translation process by Scene Relations.
The atomic operations used by a translation Scene should not be interpreted as symbolic edit commands. A biologically plausible interpretation is that they correspond to internal action-like control patterns that configure routing, attention, prediction, inhibition, and reactivation in the Integration Workspace. In this view, a translation Scene can be treated as a learned internal action Scene whose execution transforms active Scene Content rather than directly moving the body. Prefrontal cognitive-control theories, basal-ganglia–prefrontal models of working-memory gating, frontal modulation of sensory processing, attention-dependent cortico-thalamic routing, and active-dendrite models provide related biological motivation for this interpretation (Miller and Cohen 2001; O’Reilly and Frank 2006; Moore and Armstrong 2003; Saalmann et al. 2012; Hawkins and Ahmad 2016).
Scene Translation differs from expressing the same Scene Content in another Reference Frame. A Reference Frame transformation changes how related content is situated, whereas Scene Translation may produce different Features and structures in the target Scene.
4.4.6. Scene Integration
Scene Integration is the simultaneous activation of two or more Scenes followed by an operation over their combined Scene Content. It is a central mechanism of the hypothesis and is intended to support association, comparison, inference, imagination, planning, and the construction of new Scenes.
The simultaneous activation creates a landscape of active Feature Instances that may overlap, reinforce one another, conflict, or provide different contexts. The Mapping Module is assumed to retain knowledge of the source Scene Codes. Whether source provenance must also be explicitly represented within the Integration Workspace, for example through contextual activity or separate Reference Frame bindings, remains an open architectural question.
Multiple Scenes may be active concurrently without being immediately integrated. Their Scene Content may remain weakly coupled when represented by largely separate neural populations or processing pathways. Concurrent Scenes may nevertheless compete for workspace resources, attention, influence on action, and durable recording. Scene Integration occurs when shared Features, Scene Relations, routing, or other mechanisms cause their content to interact.
Scene Integration is a generic process rather than one specific merging algorithm. Possible integration operations include union of compatible content, alignment and comparison, conflict resolution, extraction of common structure, contextual modulation, and composition of a new Scene. If the result is recorded, it receives a new Scene Code and relations to the integrated source Scenes.
4.4.7. Scene Role
Simultaneously active Scenes may participate in processing under different roles. A Scene role describes how a Scene influences the current operation rather than what the Scene represents. Candidate roles include:
Evidence Scene, providing current sensory or lower-order input.
Context Scene, modifying the interpretation of other Scene Content without directly determining it.
Prediction Scene, providing expected content for matching and surprise detection.
Goal Scene, representing a desired future state.
Constraint Scene, representing content or transitions that should, must, or must not occur.
Candidate action Scene, representing a possible overt action, motor sequence, or internal cognitive operation.
Scene roles determine how Scene Content contributes to integration, comparison, conflict resolution, significance evaluation, and action selection. The same Scene may assume different roles in different operations.
4.4.8. Scene Significance Evaluation
Every transient update or integration result initially receives a temporary Scene Code, but not every such code is preserved as a durable record. Scene Significance Evaluation is the consolidation-facing use of evaluative signals: it tags or prioritizes temporary Scene Codes and relations for maintenance, reactivation, strengthening, or decay.
It can use novelty, surprise, emotional relevance, arousal, goal relevance, value, cost, and prediction error, but its output concerns consolidation rather than action selection. Later events, replay, and updated evaluation may strengthen, weaken, or reinterpret earlier temporary structures once their consequences become clear. Amygdala-dependent modulation of memory consolidation provides one possible biological inspiration for this process (Roozendaal et al. 2006).
4.4.9. Operations on Scenes
The following operations define the intended behavioral surface of the Scene abstraction. Their low-level algorithms and biological realizations remain open.
Construction and updating
Construct candidate Scene Content from sensory input, recalled content, or internal simulation.
Add, update, move, or remove a Feature Instance.
Compare predicted Scene Content with incoming content and represent significant matches or mismatches.
Extract a subset of Scene Content and record it as a related Scene.
Generate a temporary Scene Code and Scene Relations for each candidate state.
Tag, strengthen, and consolidate sufficiently significant temporary Scene Codes as durable mapping structure.
Navigation and transformation
Navigate into a Feature Instance by activating one of its associated Scenes.
Navigate to a containing, preceding, succeeding, or otherwise related Scene.
Change spatial, temporal, or conceptual scale.
Transform Scene Content within its Reference Frame, for example by spatial displacement, rotation, or scaling.
Express Scene Content in another Reference Frame and record the transformation relation.
Translate a source Scene (e.g., a path Scene) into a related representation in another domain.
Comparison and integration
Compare Scenes and identify matching, missing, additional, or conflicting Feature Instances.
Simultaneously activate and integrate multiple Scenes.
Compose a new Scene from content originating in other Scenes.
Extract common structure from similar Scenes as a candidate generalization.
Resolve conflicts among simultaneously active Scene Content.
Create or strengthen Scene Relations and organize them into candidate path Scenes.
Simulation and prioritization
Simulate a possible modification or transition without treating it as current sensory input.
Construct and evaluate possible future Scenes or path Scenes.
Replay a path Scene or execute it while incorporating sensory feedback.
Compare simulated predictions with subsequently activated or observed Scene Content and update Scene Relations through internally generated predictive learning.
Prioritize among active or candidate Scenes.
Select which Scene Content remains active, influences action, or is recorded.
4.5. Integration Workspace
The Integration Workspace is the architectural component in which detailed Scene Content is activated and processed. The neocortex is used as its principal biological abstraction, without requiring every workspace operation to map exclusively or directly to neocortical tissue. The workspace is not a single centralized store. It is a distributed and dynamically changing activity pattern spanning specialized processing regions.
4.5.1. Architectural Role
The Integration Workspace mediates between perception, memory, simulation, and action. It receives bottom-up evidence derived from sensory processing and top-down Scene activation or reactivation initiated through the Mapping Module. Within the workspace, active Scene Content can be maintained, compared, transformed, integrated, simulated, and used to construct candidate Scenes.
The Mapping Module operates primarily over compact Scene Codes, whereas the Integration Workspace performs detailed operations over distributed Feature Instances. The Relay Module supports selective routing and coordination between sensory pathways, cortical regions, the Mapping Module, and other systems. Two evaluation-related processes influence this activity: the Evaluation and Action-Selection System biases which candidate Scenes, path Scenes, or actions receive priority or guide behavior, while Scene Significance Evaluation biases which temporary Scene Codes and relations are preserved for possible consolidation.
The workspace can be compared with a staging area or blackboard, but it is not assumed to be equivalent to working memory or to a global workspace theory of consciousness. Working memory, attention, and conscious access may use or observe parts of the Integration Workspace while depending on additional mechanisms.
4.5.2. Distributed Workspace Structure
Scene Content is represented as a sparse pattern distributed across specialized and hierarchically connected cortical regions. Lower-order regions contribute Features closely related to sensory structure, while higher-order regions contribute increasingly abstract, multimodal, contextual, and action-related Features. Research on hierarchical visual processing, cortical columns, sparse representations, and distributed object models provides biological motivation for this interpretation (DiCarlo et al. 2012; Hawkins et al. 2017, 2019; Rolls 2021).
A single Scene can therefore activate Feature detectors at several levels and in several modalities at once. The resulting pattern is not merely an identifier: it is the active Scene Content on which detailed operations can act. Its corresponding Scene Code provides a more compact reference for the Mapping Module.
Workspace capacity is not defined only by the number of active neurons. It also depends on routing, interference, inhibition, temporal coordination, and whether concurrent Scenes require the same processing resources. A larger or more effectively coordinated workspace may support more detailed or more numerous simultaneous Scene operations, but this remains an architectural hypothesis rather than a direct claim about brain size.
4.5.3. Scene Activation, Reactivation, and Deactivation
Scene activation makes distributed Scene Content active from sensory evidence, a Scene Code, or a combination of both. Bottom-up activation constructs a Scene candidate from incoming data. Top-down Scene reactivation restores previously associated cortical patterns or generates predicted and simulated content. Hippocampal indexing and cortical recall mechanisms provide related biological motivation for the possibility of reactivating distributed cortical representations (Teyler and Rudy 2007; Rolls 2021).
Activation can be partial. It may activate only the Features needed for the current operation, after which additional content can be filled in as evidence, context, or attention changes. A new activation may supplement currently active Scene Content, replace part of it, or remain weakly coupled as a concurrent Scene.
Active Scene Content can weaken through decay, inhibition, displacement by competing content, or loss of supporting input. Deactivation does not erase the corresponding Scene Code or its activation pathway. If a sufficiently persistent Scene Code and activation pathway remain available, the Scene may be reactivated.
4.5.4. Active Scene Landscape
The complete state of the Integration Workspace at a given time is called the Active Scene Landscape. It may contain sensory Scene candidates, recalled Scenes, predictions, contexts, goals, constraints, and simulated alternatives. These Scenes can overlap in their participating Feature detectors while differing in Feature locations, Reference Frames, Scene roles, or source Scene Codes.
The Active Scene Landscape is dynamic: its participating Features and effective relations change as input arrives and operations proceed. At a software-architecture level, it can be interpreted as a contextual dynamic graph or multigraph whose currently effective links depend on active Scene Content, routing, and Reference Frames. In the biological interpretation, Scene Codes and distributed neural activity provide the corresponding references.
The Mapping Module is assumed to retain source provenance for active Scenes. Whether and how the Integration Workspace itself distinguishes the source of overlapping Feature activity remains open. Possible mechanisms include contextual activity, Reference Frame bindings, temporal separation, and distinct routing pathways.
4.5.5. Scene Roles and Routing
Active Scenes influence processing according to their Scene roles. Evidence Scenes provide current input. Context and prediction Scenes modulate its interpretation. Goal and constraint Scenes influence evaluation, and candidate action Scenes provide possible outputs. The same Scene Content can assume different roles during different operations.
The Relay Module is proposed to help route Scene Content so that its role affects how it is processed. Cortico-thalamic routing, deviance detection, and thalamic blackboard proposals provide related biological motivation (Worden et al. 2021; Varela et al. 2024). At the local level, active dendrites and compartmental neuron models provide possible mechanisms by which contextual and predictive inputs can modulate responses to feedforward evidence (Hawkins and Ahmad 2016; Grewal et al. 2021; Iyer et al. 2022).
Routing need not make all active Scene Content equally available to every operation. It can expose selected Features, suppress irrelevant pathways, or direct the same Scene toward comparison, simulation, action selection, or recording.
4.5.6. Scene Integration
Scene integration begins when content from two or more active Scenes is routed into interacting parts of the Active Scene Landscape. Before integration, concurrent Scenes may remain weakly coupled and undergo largely independent processing. During integration, compatible Feature Instances may reinforce one another, mismatches and conflicts may become detectable, and Reference Frame transformations may align related content.
Integration does not imply simple union. It may perform comparison, contextual modulation, conflict resolution, extraction of common structure, translation, or composition. Local inhibition, recurrent activation, normalization, and top-down recall provide possible biological building blocks for these operations (Carandini and Heeger 2012; Bennett 2020; Rolls 2021).
The result of Scene integration is candidate Scene Content. It may remain transient, modify another active Scene, influence action, or receive a temporary Scene Code and candidate relations in the Mapping Module.
4.5.7. Operations Within the Workspace
The Integration Workspace provides the active substrate for the content-level portions of the operations defined in the Scene chapter, including comparison, transformation, integration, translation, and simulation of Scene Content. It does not independently perform mapping-level operations such as generating or storing Scene Codes, navigating Scene Relations, deciding persistence, or selecting actions. Those operations require interaction with the Mapping Module, Relay Module, Scene Significance Evaluation, and action-selection processes.
A coordinated sequence of workspace operations and navigations through related Scenes may be interpreted as execution of an algorithm at the architectural level. The formal model defines behavioral contracts for primitive operations while leaving their runtime and biological realization open.
4.5.8. Competition, Attention, and Conflict Resolution
Because the Integration Workspace has finite processing and routing capacity, active Scenes and Feature Instances compete for influence. Attention is interpreted as prioritization of selected Scene Content, locations, relations, or operations. It can be driven by bottom-up novelty and salience or by top-down goals and predictions. Experimental work showing distinct dynamics for bottom-up and top-down attentional control provides biological motivation for treating these influences separately (Buschman and Miller 2007).
An attention shift can therefore be interpreted as navigation within the current Scene, navigation into a sub-Scene, or navigation toward a related Scene. Higher-order cortical systems, including prefrontal systems, may bias this navigation, but the architecture does not treat attention as the responsibility of one isolated controller.
Competition may be resolved through inhibition, normalization, routing, voting among compatible representations, or evaluation by other systems. Normalization is a broadly observed neural computation that may contribute to competition among active representations (Carandini and Heeger 2012). Conflict resolution need not select a single complete Scene. It may preserve several alternatives with different confidence or priority.
Selected content can receive further processing, influence action, or be submitted for durable recording. Unselected content may remain active at lower priority, continue unconsciously, or decay.
4.5.9. Learning and Formation of Candidate Scenes
Every detected change and internally generated result can produce candidate Scene Content. The Integration Workspace supplies the detailed distributed pattern, while the Mapping Module generates the corresponding temporary Scene Code and creates or updates Scene Relations. Scene Significance Evaluation then influences how those temporary structures enter later consolidation.
Learning can modify Feature detectors, associations among simultaneously active Features, activation pathways, contextual effects, and relations between candidate Scenes. Coincidence-dependent and spike-timing-dependent plasticity provide possible local mechanisms through which the relative timing of active patterns affects synaptic change (Markram et al. 1997; Hawkins and Ahmad 2016). Such mechanisms could allow rapidly alternating or simultaneously active Scenes to form new associations, although the exact conditions under which this occurs remain open.
Prediction mismatch and surprise are particularly important candidate-learning signals. A prediction Scene can place Features in a predictive state, while incoming evidence Scenes confirm or violate those expectations. Significant matches, missing expected Features, and unexpected Features can then contribute to a new candidate Scene and modify future predictions.
4.5.10. Parallelism and Temporal Coordination
The Integration Workspace supports parallel processing when active Scenes use sufficiently separate populations, pathways, or Reference Frames. Concurrent processing does not imply complete independence: Scenes can still compete for attention, routing, action influence, and durable recording. Interaction begins when their active content overlaps or is deliberately routed together.
Temporal coordination may determine which distributed patterns are treated as belonging together. Relative spike timing, oscillations, transient bursts, inhibition, and conduction delays are possible mechanisms for separating, binding, or sequencing active content. Experimental findings relating beta and gamma bursts to working-memory processing illustrate one possible form of transient temporal coordination (Lundqvist et al. 2016). Workspace of Scenes does not require a particular oscillatory code.
Fast switching or replay of related Scenes may expose correlations that were not present in any single Scene. If their activity falls within suitable plasticity windows, new associations may form. Excessively overlapping activity can instead increase interference and competition.
4.5.11. Biological Interpretation
Workspace of Scenes interprets the neocortex as the principal biological substrate of the Integration Workspace because it provides large-scale, recurrent, hierarchical, and specialized processing. Cortical columns and minicolumns are treated as candidate Feature-integration structures, while cortico-cortical and cortico-thalamic loops contribute routing and coordination. Column-based and active-dendrite theories provide broad biological motivation for this interpretation (Bennett 2020; Hawkins et al. 2019).
Association cortical regions provide possible biological substrates for higher-order Scene Content and workspace operations. Unimodal, multimodal, prefrontal, parietal, temporal, and limbic regions may contribute different Features, contexts, goals, spatial relations, and evaluations to concurrently active Scenes. The architecture does not assign complete Scenes or fixed Scene roles to particular association cortical regions.
4.6. Relay Module
The Relay Module is the architectural component responsible for dynamic routing, gating, modulation, and coordination of information exchanged among sensory pathways, regions of the Integration Workspace, the Mapping Module, and other systems. The thalamus and cortico-thalamic loops are treated as its principal biological inspiration and candidate substrate.
The Relay Module is not a passive switchboard. A routed signal can be selected, suppressed, amplified, delayed, temporally coordinated, or expressed relative to another Reference Frame before reaching its destination. At the same time, the Relay Module is not a Scene store and does not perform the detailed semantic interpretation of Scene Content.
4.6.1. Architectural Role
The Relay Module determines which available signals can influence which operations at a given time. It helps sensory evidence reach relevant Feature detectors, recalled and simulated content reach selected workspace regions, and results reach systems that evaluate significance or produce action. It also helps maintain separation between concurrent Scenes until integration is required.
Routing is bidirectional and recurrent. Bottom-up signals can be routed into the Integration Workspace, while top-down cortical signals can modify subsequent routing and modulation. First-order and higher-order thalamic relay theories provide biological motivation for treating relay as an active part of both sensory and cortico-cortical information flow (Guillery and Sherman 2002; Sherman and Guillery 2002).
The architectural module is broader than the biological thalamus. Some routing operations may be distributed across direct cortico-cortical pathways, local inhibition, neuromodulatory systems, and other circuits. Direct pathways coexist with Relay Module pathways and are not required to pass through a central routing point.
4.6.2. Routing Channels and Routing Configuration Scenes
A routing channel is a possible directed pathway from a source representation or system to a destination. A channel specifies what kind of signal can be transmitted and how that signal can influence its destination. Several channels may carry related content under different Scene roles or toward different operations.
A routing configuration is the current pattern of enabled, suppressed, amplified, modulated, or temporally coordinated routing channels. A routing configuration may itself be represented as a routing configuration Scene. This allows routing strategies to be learned, recalled, compared, simulated, and selected using the same Scene-based architecture as other structured knowledge. Individual low-level routing adjustments do not each require a separate routing configuration Scene.
The Relay Module applies an active routing configuration. When a routing configuration is represented as a Scene, activation or reactivation of that Scene can help establish the configuration. The configuration can also be produced or influenced by several systems. Context and goals in the Integration Workspace, Scene Relations in the Mapping Module, significance signals, attention, and behavioral state may all contribute. Consequently, routing control is distributed even when the Relay Module executes the resulting configuration.
4.6.3. First-Order and Higher-Order Relay
First-order relay routes information from subcortical sensory pathways into the Integration Workspace. For example, the lateral geniculate nucleus relays retinal input toward visual cortex. The architecture does not require every sensory modality to use the same biological pathway.
Higher-order relay routes selected information between regions of the Integration Workspace through recurrent cortico-relay-cortical pathways. This provides an alternative to direct cortico-cortical communication and allows the transmitted signal to be modulated by context, attention, prediction, or behavioral state. Biological accounts distinguish first-order thalamic relays driven by ascending input from higher-order relays driven by cortical input (Guillery and Sherman 2002; Sherman and Guillery 2002).
Workspace of Scenes uses this distinction architecturally: first-order channels primarily introduce evidence, while higher-order channels primarily coordinate and redirect already processed or internally generated content. Both kinds of channels can be modulated.
4.6.4. Reference Frame and Location Transformation
The Relay Module may participate in recalculating Feature locations and transformations between Reference Frames. Examples include converting or relating eye-centered, body-centered, egocentric, allocentric, and object-relative locations. A transformation can modify the location information delivered with a Feature Instance or route the same content toward regions operating in different Reference Frames.
Reference Frame transformation is not assumed to occur entirely within the Relay Module. It may be a distributed computation involving the Integration Workspace, Mapping Module, sensory and motor signals, and relay pathways. Evidence that dorsal pulvinar activity carries eye-position information relevant to spatial coordinate transformations, and that thalamo-parietal circuits contribute to place–action coordination, provides biological motivation for including this capability (Schneider et al. 2020; Simmons et al. 2023).
The hypothesis extends this spatial capability to more abstract Scene Reference Frames.
4.6.5. Scene Role-Aware Routing
Routing helps determine how Scene Content influences the Active Scene Landscape. Evidence Scenes can be routed toward feedforward processing. Context and prediction Scenes toward modulatory pathways. Goal and constraint Scenes toward evaluation and conflict resolution, and candidate action Scenes toward action-selection interfaces.
The Relay Module does not independently determine the semantic meaning of a Scene role. Instead, it applies active routing configurations that cause routed content to have role-appropriate effects. The same Scene can therefore be routed differently when it serves as evidence, context, prediction, or a candidate action.
Scene role-aware routing may also help preserve source provenance. When overlapping Scene Content is routed through distinct channels, downstream processing can potentially distinguish otherwise similar Feature activity by its channel, timing, or modulatory context.
4.6.6. Gating, Attention, and Competition
Routing channels compete for limited transmission and processing capacity. The Relay Module can gate this competition by suppressing irrelevant channels, amplifying prioritized signals, and coordinating which workspace regions exchange information. The thalamic reticular nucleus provides a possible biological mechanism for inhibitory gating, while pulvinar activity has been associated with attention-dependent regulation of information transmission between cortical areas (Halassa et al. 2014; Saalmann et al. 2012).
Attention can influence active routing configurations from the top down, while salient or unexpected sensory input can influence them from the bottom up. When a configuration is represented as a Scene, its learned structure can help reinstate a previously useful routing strategy. The Relay Module thereby contributes to attention without owning the complete attention-selection process. Goals, significance, cortical competition, and other systems also participate.
Gating can preserve several alternative Scenes rather than selecting only one. Different alternatives may remain routed at different strengths or toward different processing regions until additional evidence or evaluation resolves the competition.
4.6.7. Prediction, Deviance, and Surprise Routing
The Relay Module can route prediction Scenes and evidence Scenes toward compatible Feature detectors and comparison processes. When evidence agrees with prediction, the active routing configuration can continue to prioritize the expected processing path. When a mismatch or unexpected signal occurs, the Relay Module can amplify or redirect deviance information toward attention, Scene Significance Evaluation, and candidate Scene formation.
Reviews and theoretical proposals describe thalamic participation in deviance detection and contextual routing (Varela et al. 2024). Workspace of Scenes additionally allows the Relay Module to initiate a local routing change when it detects deviance, although this local autonomy is not required for the core architecture.
The Relay Module does not decide the complete meaning or long-term significance of surprise. It makes relevant mismatch information available to systems that can interpret, evaluate, and record it.
4.6.8. Scene Activation and Integration Coordination
The Relay Module assists Scene activation and reactivation by selecting which regions and pathways participate in activating Scene Content. A Mapping Module request can identify the Scene Code to be reactivated, while the Relay Module helps route the resulting activation toward the appropriate Integration Workspace regions. The entorhinal interface remains part of the Mapping Module, consistent with its biological role as a principal interface between hippocampal and cortical systems.
Routing can support partial activation by activating only content relevant to the current operation. It can supplement an active Scene, keep concurrent Scenes weakly coupled, or bring selected Scenes into interaction for Scene Integration. It can also suppress activation pathways when content loses priority or conflicts with a selected routing configuration.
The Relay Module coordinates activation and reactivation but does not store Scene Codes or Scene Relations. Those remain responsibilities of the Mapping Module.
4.6.9. Temporal Coordination
Routing configurations may include temporal constraints that determine when signals are transmitted, repeated, or suppressed. Such coordination can help separate concurrent Scenes, bind distributed Feature activity, sequence partial activations, and synchronize processing across distant workspace regions.
Thalamic circuits exhibit tonic and burst response modes, and pulvinar interactions can coordinate rhythmic attention across cortical networks (Sherman and Guillery 2002; Fiebelkorn et al. 2019). These findings motivate, but do not establish, a Relay Module role in the temporal organization of Scene processing.
The architecture does not require the Relay Module to store path Scenes or function as a universal delay line. It may help execute or coordinate a path Scene whose ordered Scene Relations are maintained by the Mapping Module.
4.6.10. Biological Interpretation
The thalamus contains multiple specialized nuclei rather than one uniform relay. First-order nuclei, higher-order nuclei, the pulvinar, matrix-like pathways, and the inhibitory thalamic reticular nucleus provide possible biological substrates for different Relay Module functions. The proposed abstraction emphasizes their shared architectural contribution to selective and context-dependent information flow.
4.7. Mapping Module
The Mapping Module is a fast relational index and navigation system over Scenes. It generates and maintains compact Scene Codes, associates them with source provenance, persistence state, and reactivation pathways, stores navigable Scene Relations, and proposes candidate paths or path Scenes over related or possible Scene Codes. Its principal biological inspiration is the hippocampal–entorhinal system.
The Mapping Module operates primarily over references to Scenes rather than complete Scene Content. Detailed Scene Content is distributed through the Integration Workspace. This division allows the Mapping Module to navigate, compare, and extend many possible path Scenes without fully activating every Scene along them.
4.7.1. Architectural Role and Boundaries
The Mapping Module provides continuity between the current Scene, remembered Scenes, and possible future or counterfactual Scenes. It can retrieve a related Scene, advance through a learned transition, propose alternatives at a branch point, or construct a candidate path Scene. It can initiate selective activation or reactivation when detailed processing is required.
This active navigation role does not make the Mapping Module a central executive. Goals, constraints, sensory evidence, significance signals, and behavioral state bias its navigation. The Integration Workspace evaluates detailed Scene Content. The Relay Module coordinates routing and activation, and the Evaluation and Action-Selection System influences priority and action. The Mapping Module proposes and traverses possibilities but does not alone decide which possibility is correct, valuable, or acted upon.
4.7.2. Entorhinal Interface and Scene Code Formation
The entorhinal interface is treated as part of the Mapping Module. It mediates between distributed cortical Scene Content and hippocampal Scene Codes, and contributes navigable location and displacement structure. The Mapping Module generates a compact Scene Code from candidate Scene Content arriving through this interface. Later activation of the code can initiate reconstruction or reactivation of the associated distributed content.
Every detected change or internally generated candidate Scene receives a new temporary Scene Code, even when it closely resembles an existing Scene. Similar Scenes should nevertheless be discoverable as similar. The architecture therefore requires a balance between separating distinct experiences and preserving a usable similarity structure, whether through overlap between sparse codes, associated metadata, learned relations, intermediate representations, or some combination of these mechanisms.
Hippocampal indexing, sparse coding, pattern separation, and pattern completion provide biological motivation for this bidirectional mapping between distributed content and compact codes (Teyler and Rudy 2007; Bakker et al. 2008; Petrantonakis and Poirazi 2014; Neunuebel and Knierim 2014).
4.7.3. Scene Codes and Matching
A Scene Code is the compact sparse code associated with one Scene state and acts as the Mapping Module’s operational reference to that Scene. The code may be supported by learned connectivity and mapping-level state needed to operate on the Scene from the Mapping Module perspective. This can include Scene Relations, persistence and significance state, source provenance, and connections that support later Scene reactivation. A Scene Code is not a complete copy of Scene Content.
When new candidate Scene Content arrives, the Mapping Module generates a new temporary Scene Code and matches for similar existing codes. A similarity match immediately creates a weak temporary Scene Relation between the new and retrieved codes. Subsequent activation and evaluation can strengthen, reclassify, or discard that relation.
Approximate matching can support recognition, generalization, and retrieval from incomplete cues. Pattern separation keeps similar but distinct experiences individually addressable, while pattern completion allows a partial cue to retrieve a more complete associated Scene. Because similar Scene Codes are intended to imply related Scenes, a Scene Code is not merely an arbitrary identifier.
4.7.4. Scene Relations and the Navigable Map
A Scene Relation is the primitive mapping-level connection among Scene Codes. It describes a possible way of navigating from one recorded or candidate Scene state to another. A relation can represent an observed update, association, change of attention, Reference Frame transformation, causal hypothesis, composition, similarity, generalization, or candidate transition.
At the architectural level, a Scene Relation may contain:
source and destination Scene Codes;
relation type and direction;
a Reference Frame delta or transformation, when applicable;
strength, confidence, persistence, or estimated reliability;
provenance identifying the operation or experience that created it; and
an optional associated relation Scene Code, when the relation’s internal structure must be reactivated or reused.
These fields describe required behavior that may be realized through learned connectivity and population dynamics. Together, Scene Codes and Scene Relations form a navigable map that can be traversed without activating every referenced Scene. When a relation points to a relation Scene, mapping-level navigation can remain compact while detailed Feature Instance relations can be reactivated in the Integration Workspace. More complex temporal or procedural structures are built from this map rather than introduced as separate primitives.
In the simplified biological interpretation, Scene Relations are primarily hippocampal–entorhinal mapping structures, while the detailed content of relation Scenes is primarily neocortical workspace activity. This mirrors the proposed Scene split: Scene Codes and relation links support compact indexing and navigation in the Mapping Module, whereas Scene Content and relation Scene Content are distributed Feature Instance patterns in the Integration Workspace. The split is functional rather than anatomical: cortico-hippocampal and cortico-thalamic loops may participate in both activation and learning.
4.7.5. Navigation, Retrieval, and Scene Activation
Mapping navigation can follow explicit Scene Relations or search an implicit neighborhood of similar Scene Codes. Explicit relations support learned transitions and transformations. Similarity-based retrieval permits navigation between related Scenes for which no durable explicit relation has yet been learned.
The Mapping Module can autonomously propose the next Scene or path Scene fragment. Navigation can be biased by the current Scene, active goals and constraints, prediction error, significance, unfinished path Scenes, recent experience, and spontaneous replay. At a branch point, several candidate path Scenes may be proposed in sequence or in parallel.
A proposed Scene Code can remain at the mapping level while a path Scene is explored approximately. When detailed evaluation is required, the Mapping Module initiates partial or complete Scene activation through the Relay Module into the Integration Workspace. This division lets fast mapping-level navigation guide slower, content-rich processing. Figure 3 illustrates the distinction between navigation among Scene Codes and activation of their associated Scene Content.
4.7.6. Temporal Organization of Path Scenes
A path Scene is a higher-order Scene whose Reference Frame organizes component Scenes as positions along one or more navigable dimensions. Such path Scenes can represent episodes, procedures, observed changes, possible futures, counterfactual sequences, routes, reasoning traces, or successive states of one continuing entity or situation. These cases differ by Reference Frame and interpretation, not by additional primitive structures.
At the mapping level, the component Scene Codes may remain connected by Scene Relations. When a path Scene is activated in the Integration Workspace, its component Scenes can appear as Features or sub-Scenes organized by the path Scene’s Reference Frame. A temporal sequence, spatial route, procedure, and sequence of object states are therefore all path Scenes with different dimensions or constraints.
Significant states can become durable anchor Scenes while intermediate temporary Scene Codes and relations weaken. Recall may therefore navigate from an anchor and reconstruct less persistent portions of the path Scene.
4.7.7. Operations Over Relations and Path Scenes
Scene binding, path formation, path composition, and path recombination are operation names over the navigable map and path Scenes. Binding creates or strengthens a Scene Relation. Formation records a path Scene from related Scene Codes and relations. Composition joins compatible path Scenes or fragments. Recombination constructs a candidate path Scene using fragments that were previously unrelated, differently related, or embedded in another context.
The Mapping Module can replay an existing path Scene, reverse a navigable relation when such navigation is available, or generate a possible continuation. It can also recombine fragments into candidate future or counterfactual path Scenes. Selected Scenes are activated in the Integration Workspace to evaluate coherence, conflicts, expected consequences, and compatibility with evidence, goals, and constraints.
Hippocampal place cell sequences that sweep through alternative future routes, depict routes toward remembered goals, and vary with current goals provide biological motivation for an active candidate-proposal role (Johnson and Redish 2007; Pfeiffer and Foster 2013; Wikenheiser and Redish 2015). Preplay and awake replay further suggest that hippocampal sequences need not merely reproduce the immediately preceding experience (Dragoi and Tonegawa 2011; Jadhav et al. 2012).
4.7.8. Persistence, Significance, Forgetting, and Consolidation
Every recorded Scene requires a Scene Code in the Mapping Module. Reusable knowledge and associations may also become strongly represented in neocortical circuits, but they do not replace the code required to address a particular recorded Scene as that Scene.
Scene Significance Evaluation influences which temporary Scene Codes and relations are maintained for possible consolidation, strengthened, prioritized, or allowed to decay. Because the importance of an event may become clear only after later events, consolidation can retroactively strengthen earlier Scene Codes and relations that initially had only temporary or weak support. Forgetting can involve weakening or loss of a Scene Code, its relations, its reactivation connections, or some combination of them. A Scene whose code or reconstruction pathway has been lost may no longer be recallable as the original recorded Scene, even if related generalized knowledge remains.
Replay may strengthen Scene Codes and Scene Relations. Consolidation is therefore interpreted as a change in the strength and distribution of supporting representations, not necessarily as transfer of a complete Scene from one storage location to another.
Consolidation may also reorganize Scene structure. During replay or later reactivation, the architecture may form missing hierarchy levels, extract generalized Scenes from similar experiences, create variants of existing Scenes, or strengthen path Scenes around significant anchor Scenes. These are candidate functions for consolidation-facing operations.
4.7.9. Interfaces with Other Architectural Components
The Mapping Module and Integration Workspace have complementary roles. The Mapping Module navigates compact Scene Codes and preserves source provenance, while the Integration Workspace performs detailed operations over active Scene Content. Results produced in the workspace return as new temporary Scene Codes and candidate Scene Relations.
The Relay Module coordinates pathways used for selective activation and may assist Reference Frame transformations. The Evaluation and Action-Selection System and active goal or constraint Scenes bias which Scene Codes and relations are explored, while Scene Significance Evaluation influences which are maintained for consolidation. These influences can guide navigation without prescribing every mapping-level transition. Figure 4 shows a simplified biological interpretation of the bidirectional interaction between distributed cortical content and the hippocampal–entorhinal Mapping Module.
4.7.10. Biological Interpretation
The hippocampal-entorhinal system provides a plausible biological abstraction for the Mapping Module, but the proposed architecture does not require a one-to-one assignment of every operation to one anatomical region. Place cell ensembles are interpreted as a possible substrate for Scene Codes. Grid cell-like activity can contribute locations, displacements, and navigable relational structure. Boundary, head-direction, and related signals can constrain navigation and Reference Frames.
The dentate gyrus provides a possible mechanism for separating similar candidate Scenes, while recurrent CA3 circuitry provides possible mechanisms for pattern completion, associative retrieval, and sequence generation (Neunuebel and Knierim 2014; Guzman et al. 2016). CA1, subiculum, entorhinal cortex, and wider cortical loops may participate in comparison, output, and reactivation. Evidence of internally generated entorhinal activity during mental navigation supports extending the navigation interpretation beyond overt physical movement (Neupane et al. 2024).
4.8. Integration Column
An Integration Column is a reusable local processing circuit within the Integration Workspace. It combines a specialized family of inputs with location, context, prediction, and feedback signals, then emits a sparse set of locally supported Feature hypotheses and mismatch signals. The combined activity of many Integration Columns contributes to distributed Scene Content.
An Integration Column does not represent a complete Scene, store Scene Codes or Scene Relations, or control the Scene Navigation Loop. Its responsibility is to perform local inference and learning within its specialized Feature domain.
4.8.1. Internal Functional Systems
The architecture distinguishes three interacting functional systems within an Integration Column:
a local location system, which maintains candidate locations in the currently active Reference Frame;
a population of Feature Units, which combines specialized evidence with location and contextual input; and
a local hypothesis system, which integrates evidence over time, coordinates competition, and communicates selected hypotheses to other columns and systems.
These systems describe required functions rather than a fixed assignment to particular cortical layers. Their biological implementation may involve several layers, cell types, inhibitory circuits, and long-range connections.
4.8.2. Specialization and Local Input Integration
Each Integration Column processes a related family of inputs or Feature hypotheses. Lower-order columns may specialize in regularities closely tied to sensory input, while higher-order columns may specialize in relations, actions, multimodal patterns, or abstract concepts. Equivalent or related Features may be learned independently in multiple columns and regions. The architecture does not require a global one-to-one registry of Features.
The column combines driving feedforward evidence with modulatory location, context, prediction, and feedback signals. This allows the same local evidence to support different Feature hypotheses under different locations or contexts. One candidate dendritic and neuronal mechanism is sketched in Appendix 14.
4.8.3. Local Location and Reference Frame Processing
In the biological interpretation, each Integration Column is assumed to maintain, or have direct access to, a local grid cell-like location representation aligned with the active Scene Reference Frame. The representation can maintain several candidate locations when evidence is ambiguous and can update those locations using movement or more abstract displacement signals. Feature Units use this location input to represent Features as situated Feature Instances rather than as context-free detections.
The same local machinery may represent relative locations of compositional sub-Features, temporal positions in a local sequence, or other dimensions supplied by the active Reference Frame. This allows similar column-level processing to contribute to static composition, temporal behavior, and abstract relational structure.
Alignment between a column’s local location system and wider Scene Reference Frames may depend on input from other columns, the Relay Module, and the Mapping Module. The Integration Column performs local location-dependent processing but does not alone determine global Reference Frame transformations.
4.8.4. Local Hypotheses, Voting, and Competition
An Integration Column can support several compatible Feature Units simultaneously. When units provide incompatible explanations of the same evidence, they compete through a sparse voting process in which approximately locally supported hypotheses remain active. The value of and the scope of competition may vary with the column, input ambiguity, context, and available processing capacity.
Local voting is influenced by feedforward support, location consistency, predictions, contextual compatibility, lateral input from other columns, and inhibition. It does not determine the interpretation of a complete Scene. Instead, it contributes selected local hypotheses to wider distributed processing.
4.8.5. Prediction and Mismatch
Contextual or predictive input can prepare selected Feature Units before matching feedforward evidence arrives. A confirmed prediction can activate its supported hypothesis quickly and suppress incompatible alternatives. If expected Feature-at-Location activity is not confirmed, or incoming evidence supports incompatible units, column-level comparison and inhibition can produce an explicit local mismatch signal.
Mismatch signals contribute to surprise, attention, candidate Scene formation, and learning, but the Integration Column does not determine their overall significance.
4.8.6. Local Learning and Feature-Unit Recruitment
Learning can refine Feature detectors, location associations, contextual expectations, lateral associations, and local transitions. When existing Feature Units cannot adequately represent recurring input, the Integration Column can recruit and specialize previously uncommitted capacity as a new Feature Unit. Equivalent Features may consequently be represented redundantly in different columns.
Learning within a column changes how future Scene Content is constructed and interpreted. The Mapping Module remains responsible for generating Scene Codes and maintaining Scene Relations.
4.8.7. Coordination Between Columns
Integration Columns exchange sparse hypotheses, predictions, and contextual signals through cortical and relay pathways. Compatible hypotheses in different columns can reinforce one another, while incompatible hypotheses can remain alternatives or enter wider competition. This interaction permits distributed agreement without requiring any one column to contain the complete model of a Scene.
Long-range lateral interaction and column-level voting provide biological and theoretical motivation for this coordination (Hawkins et al. 2017, 2019). The exact voting, synchronization, and routing mechanisms remain open.
4.8.8. Biological Interpretation
The Integration Column is inspired by the repeated columnar and laminar organization of neocortex, but it is an architectural abstraction rather than a claim that one anatomical macrocolumn implements one indivisible algorithm. Local recurrent circuitry, inhibition, active dendrites, layer-specific pathways, and grid cell-like location mechanisms provide possible substrates (Bennett 2020; Hawkins and Ahmad 2016; Lewis et al. 2019).
A cautious layer-by-layer interpretation of this abstraction is provided in Appendix 15.
4.9. Feature Unit
A Feature Unit is a local population within an Integration Column that represents one learned or evolutionarily prepared Feature hypothesis. It integrates evidence with location, context, prediction, and inhibition, and reports whether its Feature is supported at one or more locations. A minicolumn or cooperating neuronal ensemble is proposed as a possible biological substrate.
A Feature Unit is locally defined rather than globally unique. Several Feature Units in different columns or regions may represent equivalent or overlapping Features while learning different contexts, modalities, scales, or associations.
A Feature Unit is intended to be relatively stable as a reusable detector, but its current participation in a Scene is configured by location, time, context, prediction, and feedback. Stability therefore means that a learned detector can remain addressable while being activated under different Reference Frame states, not that the underlying circuit is immutable. Learning, forgetting, and reassignment of local capacity can still alter the detector over time.
4.9.1. Feature-at-Location Representation
A Feature Unit does not act as a single boolean detector. Cells or assemblies within it may represent the Feature under different location and contextual hypotheses. Its active output can therefore express a set of Feature Instances.
Several location hypotheses may remain active when evidence is ambiguous or when the same Feature occurs at multiple locations. The containing Integration Column coordinates competition among incompatible hypotheses while permitting compatible instances to coexist.
4.9.2. Inputs and Output Contract
A Feature Unit can receive:
driving evidence from sensory pathways or lower-order Features;
location input from the Integration Column’s local location system;
contextual and predictive input from local or distant activity;
feedback from higher-order processing; and
inhibitory signals representing local competition.
Its architectural output identifies the represented Feature, its currently supported locations, and its operational state. This output contributes to distributed Scene Content. It does not itself generate a Scene Code.
4.9.3. Operational States
The Feature Unit may occupy several architectural states:
inactive, when it has insufficient support;
predictive, when context or location makes its Feature likely before driving evidence arrives;
active, when evidence supports the Feature at one or more locations; and
contradicted, when a prepared Feature-at-Location hypothesis is incompatible with incoming evidence.
The contradicted state contributes to a column-level mismatch signal. These states specify architectural behavior and do not require one dedicated neuronal population for each state.
4.9.4. Sparse Selection and Coactivation
Feature Units vote through their evidence, location consistency, predictions, and contextual support. The containing Integration Column uses these signals to retain approximately winning Feature Units among competing alternatives. Units that describe compatible aspects of the input may coactivate, while units offering incompatible explanations suppress or outcompete one another.
Selection remains contextual and provisional. Additional evidence or a change in Scene role, location, or Reference Frame can change which Feature Units are active.
4.9.5. Learning and Specialization
A Feature Unit learns the evidence patterns, location associations, and contexts that support its Feature. Learning can refine an existing unit or specialize previously uncommitted capacity into a new unit when recurring input is not represented adequately. Redundant units are permitted and can improve robustness or support different interpretations of equivalent evidence.
Coincidence-dependent plasticity, active dendrites, and local inhibitory competition provide possible biological mechanisms for this specialization (Markram et al. 1997; Hawkins and Ahmad 2016; Grewal et al. 2021).
4.9.6. Contribution to Scene Content
Active and predictive Feature Units contribute situated Feature hypotheses to the Active Scene Landscape. Contradicted hypotheses contribute mismatch information. Their combined sparse activity supplies the detailed cortical pattern from which the Mapping Module generates a Scene Code.
Feature Units do not retain source Scene provenance, decide persistence, or initiate mapping-level navigation. Those responsibilities remain distributed across the Mapping Module, Relay Module, Integration Workspace, and evaluation systems.
4.9.7. Biological Interpretation
The proposed Feature Unit is compatible with minicolumn-inspired models in which cells share related feedforward receptive fields while differing in contextual or predictive state. However, the architecture does not require every biological minicolumn to represent exactly one Feature or every Feature Unit to correspond to one anatomically distinct minicolumn. Appendix 14 gives one optional neuron-level abstraction that could support this behavior.
4.10. Evaluation, Significance, and Action Selection
4.10.1. Architectural Role and Boundaries
Evaluation supplies bias without acting as a central executive. In Workspace of Scenes, it supports two related but distinct processes:
Consolidation of temporary Scene Codes and relations. Scene Significance Evaluation names the consolidation-facing process.
Selection among candidate action Scenes. Evaluation and Action-Selection System names the action-facing process.
Both processes can draw on overlapping signals, but their outputs are different. Scene Significance Evaluation influences which temporary structures are maintained, strengthened, replayed, or allowed to decay. The Evaluation and Action-Selection System biases which internal or overt actions are selected. Neither process constructs complete Scene Content, stores Scene Codes, or replaces the Mapping Module’s navigation role.
4.10.2. Evaluation Signals and Scene-Level Targets
Evaluation can depend on active goal Scenes, constraint Scenes, predicted outcomes, learned preferences, threat and safety signals, homeostatic needs, novelty, uncertainty, reward, cost, and prediction error. Goal and constraint Scenes provide explicit Scene-level targets and restrictions against which current and candidate states can be compared.
Basic intrinsic costs are assumed to originate from prewired or slowly learned mechanisms outside ordinary Scene representation. Their interpreted causes, consequences, and possible resolutions can nevertheless be represented as evidence, goal, constraint, or candidate action Scenes. This separation allows stable drives to guide flexible Scene-based reasoning without requiring each drive to be stored as an ordinary Scene. Intrinsic cost and configurable objective mechanisms proposed for autonomous agents provide a related artificial-system motivation (LeCun 2022).
4.10.3. Scene Significance Evaluation
Scene Significance Evaluation is the consolidation-facing use of evaluative information. It tags or prioritizes temporary Scene Codes and relations for maintenance, reactivation, strengthening, reorganization, or decay. Novelty, surprise, emotional relevance, threat, reward, cost, uncertainty, goal relevance, and later consequences may all influence this process.
Significance is not assumed to be fully determined at the moment a Scene is first recorded. Later events, replay, or new evaluative context may reveal that an earlier temporary Scene Code or relation was important. The architecture therefore allows consolidation-facing evaluation to strengthen, weaken, or reinterpret earlier structures after the fact.
4.10.4. Candidate Actions and Selection
Candidate actions include both overt actions, such as motor sequences, and internal cognitive actions, such as selecting a Scene for activation, continuing or interrupting a path Scene, shifting attention, or requesting comparison of alternatives. Internal actions allow evaluation to guide thought and planning before an overt action is selected.
Some internal cognitive actions may operate directly on active Scene Content by changing routing, attention, inhibition, prediction, or reactivation. Translation operations can therefore be interpreted as a specialized family of internal cognitive actions when their target is the Integration Workspace rather than external effectors.
Several Scene Navigation Loops may evaluate and perform internal actions concurrently. Their candidate action Scenes can remain local to a navigation process, influence one another, or compete for limited activation and processing capacity. Overt action presents a stronger coordination requirement because several incompatible motor actions generally cannot control the same effector at once. The Evaluation and Action-Selection System therefore contributes to selecting a temporarily dominant overt action without requiring other internal navigation processes to stop.
4.10.5. Learning Feedback and Prediction Error
An Actor-Critic organization is one possible artificial or biological implementation pattern. A Critic-like process estimates immediate and expected value or cost for current, candidate, or path Scenes and can produce prediction-error signals when outcomes differ from predictions. An Actor-like process converts those estimates and active constraints into selection bias among candidate action Scenes.
After an internal or overt action, the resulting Scene can be compared with the predicted Scene or path Scene. Prediction errors and other evaluative signals can modify future action preferences, value estimates, navigation pressures, learned associations, and later consolidation of related Scene Codes and relations. This allows action selection and memory formation to adapt through ordinary operation rather than only during a separate training phase.
4.10.6. Biological Interpretation
The biological counterpart of this system is expected to be distributed. Basal-ganglia and related cortical loops may contribute gating and selection. Dopaminergic systems may contribute learning and prediction-error signals. Amygdala, hypothalamic, brainstem, and other limbic mechanisms may contribute intrinsic cost, threat, reward, and relevance signals. These correspondences are functional hypotheses rather than strict one-to-one mappings.
Dopaminergic reward-prediction-error findings and basal-ganglia action-selection models provide biological motivation for the Actor-Critic-like interpretation without establishing a strict correspondence (Schultz et al. 1997; Redgrave et al. 1999). Amygdala-dependent modulation of consolidation provides a related motivation for treating significance as partly separate from immediate action selection (Roozendaal et al. 2006).
4.11. Scene Navigation Loop
The Scene Navigation Loop is the recurrent architectural process through which perception, memory, simulation, evaluation, action, and learning continually influence one another. It connects the components introduced in Figure 2 into a self-directed agent: the current Active Scene Landscape biases navigation through Scene Codes and relations, Mapping proposes related or possible Scenes, selected candidates are activated and evaluated, and the results alter subsequent navigation.
The loop can continue without a new external stimulus because the Mapping Module can generate candidate transitions from existing Scene Codes and relations. It is not required to run without interruption or to preserve one line of thought indefinitely. Individual navigation processes can finish, pause, lose priority, or be replaced while the agent remains capable of beginning another cycle.
4.11.1. The Closed Architectural Loop
A simplified Scene Navigation Loop consists of the following recurrent stages:
Sensory input or internal activity establishes an Active Scene Landscape.
The Mapping Module identifies candidate Scene Codes using explicit Scene Relations and Scene Code similarity.
Similarity-based retrieval creates weak temporary Scene Relations that can later be evaluated.
The Mapping Module proposes one or more candidate Scenes or path Scenes.
The Relay Module coordinates selective activation of candidates in the Integration Workspace.
The Integration Workspace operations compare, integrate, transform, or simulate the active Scene Content.
Evaluation signals, active goals, and constraints assess the results and influence retention or action selection.
Results influence action, persistence, routing, and the next Mapping Module navigation step.
New temporary Scene Codes and Scene Relations become available to subsequent cycles.
No single stage owns the complete loop. Its behavior emerges from recurrent interaction among components with distinct responsibilities.
4.11.2. Sources of Navigation Pressure
Mapping navigation is biased by several forms of navigation pressure: current sensory evidence, active context Scenes, goal Scenes, constraint Scenes, prediction errors, novelty, threat, reward, expected utility, incomplete path Scenes, recent experience, and spontaneous activity. These influences change which Scene Codes and relations are likely to be explored without completely determining the next transition.
Navigation pressure can be external or internal. A sudden sensory change can redirect the loop toward explaining or responding to the event. In the absence of urgent input, unresolved goals, weak memories, learned habits, or spontaneous replay can initiate recall, imagination, planning, or exploration.
4.11.3. Candidate Generation, Activation, and Evaluation
The Mapping Module generates candidates by following explicit relations, retrieving similar Scene Codes, extending a path Scene, or recombining path Scene fragments. Candidate generation can remain at the compact mapping level until more detailed processing is needed.
Selected candidates are activated in the Integration Workspace, where their Scene Content can be compared with evidence, evaluated under different Scene roles, transformed, or integrated with other active Scenes. The Mapping Module can then continue along the same path Scene, explore an alternative branch, or construct a new candidate from the workspace result.
Evaluation is distributed. The Integration Workspace determines detailed compatibility and consequences; Scene Significance Evaluation biases consolidation using novelty, emotional relevance, and supplied evaluative signals; and action-selection processes contribute expected value, prediction error, and action-related evaluation. Mapping uses these results to bias further navigation but does not replace them.
4.11.4. Action and Continuous Learning
A loop cycle may produce overt action, internal navigation, a prediction, a recalled Scene, a modified path Scene, or no durable result. When action occurs, subsequent sensory feedback produces new candidate Scenes and closes the agent–environment portion of the loop.
Every cycle can also produce learning. New candidate states receive temporary Scene Codes, discovered similarities receive weak temporary relations, and operations produce relations to their source Scenes. Evaluation and consolidation determine which of these structures are strengthened, revised, or forgotten.
Continuous learning therefore does not require a separate training phase. It is a possible consequence of ordinary operation, although consolidation, replay, and deliberate practice can further modify what is retained.
4.11.5. Autonomous, Parallel, and Hierarchical Navigation
Several Scene Navigation Loops may operate concurrently over partly separate Scene populations, Reference Frames, modalities, or timescales. One loop may track immediate sensorimotor changes while another maintains a plan, recalls an episode, or explores an abstract problem. Higher-order loops can treat the states or path Scenes of lower-order loops as Features within their own Scenes.
Parallel loops can remain weakly coupled, exchange intermediate results, or compete for activation and action influence. One loop may become temporarily dominant for overt action without requiring all other processing to stop. This provides a possible architectural distinction between the current focus of behavior and other ongoing processing.
4.11.6. Stability, Interruption, and Termination
A self-directed navigation loop requires mechanisms that prevent unproductive continuation. Competition, inhibition, habituation, limited workspace and routing capacity, declining significance, completed goals, prediction satisfaction, fatigue, and behavioral interruption can reduce or terminate a navigation process. Conversely, unresolved conflict, surprise, threat, reward opportunity, or an unfinished goal can maintain it.
Repeated navigation through the same Scene Codes without useful change may be down-prioritized. Failure of this control could provide an architectural interpretation of repetitive cognitive looping, including rumination-like dynamics, but that interpretation remains speculative.
4.11.7. Biological Interpretation
The biological counterpart of the Scene Navigation Loop is expected to be distributed. Hippocampal and entorhinal dynamics provide possible mechanisms for retrieving and generating candidate path Scenes. Cortical systems activate and evaluate detailed content. Thalamic and cortico-thalamic systems coordinate routing. Prefrontal, basal-ganglia, limbic, and neuromodulatory systems contribute goals, value, significance, and action selection.
Hippocampal sequences representing alternative future routes, routes to remembered goals, and goal-dependent look-ahead support a possible role in candidate generation (Johnson and Redish 2007; Pfeiffer and Foster 2013; Wikenheiser and Redish 2015). Coordinated hippocampal–prefrontal replay provides related motivation for treating planning and memory-guided decision making as interactions between Mapping and wider evaluative systems rather than as operations of the hippocampus alone (Shin et al. 2019).
5. Worked Example: Catching a Falling Cup
Consider an agent that sees a cup near the edge of a table, predicts that it may fall, and reaches to catch it. The example is intentionally ordinary: it requires perception, memory, prediction, goal evaluation, action selection, and online updating without requiring language or explicit symbolic reasoning.
Evidence Scene. Visual and proprioceptive input activate Feature Instances for the cup, table edge, hand, distance, orientation, motion, and current body posture. These Features form an evidence Scene in an egocentric or task-relative Reference Frame.
Scene Code and related retrieval. The Mapping Module generates a temporary Scene Code for the current state and retrieves similar previous Scenes: cups at edges, objects slipping, successful catches, failed catches, and arm-reaching paths.
Prediction Scene. Retrieved relations and current motion cues activate a prediction Scene in which the cup moves beyond support and begins to fall. This Scene is not treated as current evidence; it is routed as an expected possible future.
Goal and constraint Scenes. A goal Scene represents preventing the cup from breaking or spilling. Constraint Scenes represent body limits, collision risks, timing, and the requirement that the selected action remain compatible with the current sensory state.
Candidate action Scenes. Several possible actions become active: ignore the event, move the hand to intercept the cup, stabilize the cup on the table, or step back. Each candidate can be represented as a path Scene with expected intermediate states and outcomes.
Evaluation and selection. The Integration Workspace compares candidate path Scenes against evidence, predicted timing, goals, and constraints. The Evaluation and Action-Selection System biases the intercepting reach if it has the best expected value under the current time pressure.
Execution and updating. As the hand moves, incoming sensory evidence creates updated evidence Scenes. Mismatches between predicted and observed cup motion can modify the active path Scene, select a corrected reach, or abandon the action if it becomes impossible.
Learning. The outcome produces new temporary Scene Codes and Scene Relations: the observed transition, selected action, prediction error, success or failure, and later significance. Repeated experience can strengthen reliable relations between cup-edge configurations, falling predictions, and effective catching actions.
The same Scene Content can have different effects depending on role. The cup-at-edge structure may serve as current evidence, as a remembered prior Scene, as a predicted future Scene, as a constraint on action, or as part of a candidate action path. This is why the architecture separates Scene Content from Scene roles and from routing: the same content need not mean the same thing operationally in every cycle.
The example also illustrates the split between detailed content and compact mapping. The agent does not need to fully reactivate every previous falling-object episode before acting. The Mapping Module can navigate compact Scene Codes and relations to propose useful candidates, while the Integration Workspace activates enough detailed content to evaluate the candidates against the current situation.
6. Functional Consequences of the Architecture
The preceding components define an architectural mechanism rather than a separate mechanism for every cognitive capability. This section summarizes how their interaction may support recognition, learning, recall, simulation, planning, and goal-directed action. It does not assume that every instance of these capabilities requires complete Scene activation or explicit Mapping navigation.
6.1. Recognition and Interpretation
Incoming evidence activates Feature Instances and produces candidate Scene Content with a temporary Scene Code. The Mapping Module can retrieve similar Scene Codes and selectively activate their associated Scene Content, supplying possible interpretations, contextual expectations, and predicted Features. The Integration Workspace compares these candidates with current evidence, including matching Features, missing expected Features, unexpected Features, and inconsistent locations.
Object recognition can be interpreted as extracting a candidate sub-Scene from a wider active Scene. Segmentation may be supported by location proximity, edge or distance discontinuities, shared movement, task relevance, or attention shifts. Once a sub-Scene is extracted and recorded, later navigation can return to it, recognize it in a new context, or use it as a component of other Scenes.
Recognition can also use a coarse-to-fine strategy. Generalized Scenes may provide efficient initial hypotheses, while more detailed Scenes are activated when the coarse match is ambiguous, insufficient, or behaviorally important. The hypothesis does not require this order in every case, but it provides one plausible way to reduce the number of detailed comparisons.
Recognition can remain provisional while competing interpretations accumulate evidence. Mismatch and surprise can redirect attention or create candidate action Scenes for active sensing, such as moving the eyes or changing viewpoint. Learned path Scenes can guide the order in which informative locations are inspected.
6.2. Structured Learning and Generalization
Learning occurs at several interacting levels. Feature Units refine detectors and contextual associations. The Integration Workspace forms candidate Scene Content. The Mapping Module creates Scene Codes while strengthening Scene Relations and recurring path Scenes. The resulting structures form a navigable and continuously updated model of experienced states and transformations.
For example, while entering a restaurant, attention may shift from the room to a table and then to individual items. Each shift can produce a related Scene or sub-Scene, while Reference Frame transformations and observed changes establish relations among them. Repeated experiences can strengthen stable relations while preserving differences between particular visits.
Generalization can emerge when similar Scene Codes are retrieved, activated, compared, and integrated. Shared structure may be recorded as a generalized Scene, while context-dependent differences remain represented by separate Scenes or relations. Generalizations can describe perceptual categories, situations, transformations, or behaviors and can themselves participate in hierarchical Scenes.
In this interpretation, a generalized Scene can be formed by extracting an intersection or recurring common structure from similar Scenes. A concrete object Scene can be formed by composing or binding the union of active Feature Instances that belong to one inferred object. Both generalized and concrete Scenes may appear at multiple levels of hierarchy, and repeated experience can insert intermediate levels between highly specific instances and broad abstractions.
6.3. Recall, Replay, and Temporal Reasoning
Temporal structure can be represented by path Scenes whose Reference Frames include time, sequence position, state, or other relevant dimensions. Recall can navigate these structures, selectively activate recorded states, reconstruct less persistent intermediate states, or use an incomplete path Scene to predict possible preceding or succeeding states.
Replay need not occur at the original rate. A path Scene such as a remembered episode, action sequence, or song may be rolled out faster, slower, partially, or in a modified order when the represented relations permit it. Recombining or replacing path Scene fragments provides a possible basis for editing memories, rehearsing actions, and constructing imagined temporal sequences.
6.4. Internal Navigation, Simulation, and Planning
Mapping navigation can retrieve related Scenes, propose alternative transitions, and construct candidate path Scenes. Selected candidates are activated in the Integration Workspace, where they can be compared, transformed, integrated, or evaluated under different Scene roles. Internal cognitive actions determine which alternatives receive further attention or simulation.
Problem solving can therefore be described as navigation from a current Scene toward a target or goal Scene, where the target may be physical, social, procedural, or abstract. Candidate paths can be explored, abandoned, recombined, or backtracked. Coarse candidate simulations may first test whether a direction is promising. Finer simulations can then verify details against evidence, constraints, and expected outcomes.
These operations provide a possible mechanism for important forms of imagination, deliberation, problem solving, and planning. A successful or otherwise significant reasoning process can itself be recorded as a reusable path Scene. The hypothesis does not claim that all thinking consists exclusively of explicit Scene navigation.
6.5. Goal-Directed Action and Continuous Adaptation
Predicted path Scenes or learned responses can become candidate action Scenes. Evaluation compares them with active goals, constraints, expected outcomes, and intrinsic costs, then biases selection among internal and overt actions. Subsequent evidence produces new candidate Scenes and prediction-error signals that update future evaluation, navigation, and action selection.
The architecture therefore forms a recurrent functional progression: recognize and interpret the current state, learn or recall relevant structure, simulate possible transitions, select an action, observe its consequences, and adapt. Different stages can overlap and operate concurrently rather than forming one fixed serial pipeline.
7. Evolutionary Hypotheses
7.1. Navigation as a Precursor to Abstract Cognition
One tentative hypothesis is that mechanisms originally supporting navigation through physical environments were later reused or extended for navigation through temporal and abstract relational spaces. Under this interpretation, capabilities for representing locations, displacements, scales, and routes could provide useful foundations for organizing relations among events, concepts, and possible actions.
A related possibility is that some computational principles used by the expanded neocortex share an earlier evolutionary origin with, or developed from principles present in the hippocampal formation. Later cortical expansion could then have supported increasingly detailed and reusable Scene Content, while hippocampal circuits remained specialized for rapid relational binding, indexing, and navigation.
The proposed evolutionary relationship remains speculative. Determining whether physical and abstract navigation share biological mechanisms or only computational principles requires further comparative and experimental research.
8. Future Implementation
A future implementation must define concrete representations and interfaces for Scene Content, Scene Codes, Scene Relations, path Scenes built from those relations, Reference Frames, routing configurations, evaluation signals, and action selection.
Implementation work must also determine how concurrent Scene processes are scheduled and coordinated, how temporary structures are strengthened or forgotten, how Reference Frame transformations are learned and applied, and how local learning contributes to stable system-level behavior. These choices may differ substantially between biologically detailed neural simulations and more abstract software implementations while preserving the same architectural responsibilities.
One intended implementation path is Angsi, an AI framework, that is described in Appendix 16.
9. Predictions and Evaluation Criteria
The hypothesis is intended to become testable at two levels. Neuroscience-facing predictions concern whether brain activity shows organization consistent with structured Scene-like content, compact indexing, navigable relations, and role-aware routing. Implementation-facing criteria concern whether a built system gains useful capabilities from the proposed architectural split between detailed Scene Content and compact Scene Codes.
9.1. Neuroscience-Facing Predictions
The following candidate observations would make the architecture more or less plausible.
Scene Code reactivation. During partial cueing of a learned multi-feature situation, hippocampal–entorhinal activity should predict reactivation of distributed cortical content corresponding to the cued Scene, not only isolated feature detectors. Similar but distinct cues should show both pattern completion toward a stored Scene and separation among overlapping Scenes.
Path Scene recall. During recall of an episode, route, procedure, or planned sequence, compact hippocampal–entorhinal sequence activity should precede or predict sequential reactivation of distributed cortical content corresponding to component Scenes. The sequence should be compressible, partial, goal-biased, or recombined when the task requires planning rather than literal replay.
Reference Frames beyond physical space. Relational reasoning, procedural planning, and abstract comparison tasks should sometimes show navigation-like dynamics in hippocampal–entorhinal or entorhinal-like systems. Tasks requiring transformation between relational frames should produce stronger frame-alignment or remapping signatures than tasks that only classify local features.
Role-aware routing. The same representational content should produce different downstream effects depending on whether it is routed as evidence, prediction, context, goal, constraint, or candidate action. Manipulating thalamic or cortico-thalamic routing should selectively impair role-dependent integration, comparison, or prediction while leaving some local feature detection intact.
Mismatch and delayed significance. Prediction mismatch should create or strengthen temporary representations of the unexpected transition. Later emotional relevance, reward, cost, or goal relevance should modulate consolidation and replay of earlier Scene Codes and relations, not only the state active at the moment of reward or error.
9.2. Implementation-Facing Evaluation
Evaluation should begin with constrained tasks that require several components to interact. A useful implementation test should make it possible to identify which architectural responsibilities were required and which assumptions failed when made concrete.
Initial experiments should test whether the system can:
construct distinct Scene Codes from changing input while still retrieving similar prior Scenes;
reactivate only the Scene Content needed for a current comparison, prediction, or action choice;
generalize across similar Scenes by extracting recurring structure while preserving episode-specific differences;
simulate alternative path Scenes and compare their expected consequences against goals and constraints;
route the same content differently when it serves as evidence, prediction, context, goal, or candidate action;
choose actions using predicted consequences and update later choices from prediction error; and
continue learning during interaction without requiring a separate offline training phase for every new relation.
These evaluation families are intentionally specified at the level of architectural behavior rather than implementation detail, so that different realizations can be compared without requiring disclosure of proprietary mechanisms.
The strongest tests should compare Workspace of Scenes implementations against simpler baselines. Candidate benchmark families include:
Rapid recombination. After learning several situations and transitions, the system must solve new tasks by recombining remembered components. Compact Scene Codes plus Scene Relation navigation should outperform flat latent-state baselines when successful behavior requires retrieving and composing prior situations without retraining.
Role switching. The same content is presented as evidence, context, prediction, goal, or constraint across trials. A role-aware implementation should change integration and action selection appropriately without relearning the content representation itself.
Reference Frame transfer. The system must recognize or act on equivalent structure across egocentric, allocentric, object-relative, temporal, or abstract frames. A Scene-based implementation should benefit from explicit Reference Frame transformation.
Path Scene planning. The system must compare alternative action or reasoning paths, backtrack at branch points, and update a path during execution when sensory feedback deviates from prediction.
Delayed significance and consolidation. The system encounters many temporary states whose importance is revealed only later. Significance-guided consolidation should preserve useful Scene Codes and relations better than uniform storage or immediate reward-only retention.
Ablation studies should remove or simplify individual architectural commitments: Scene Content without compact Scene Codes, Scene Codes without explicit Scene Relations, path Scenes without relation navigation, role-aware routing replaced by a single shared channel, Reference Frame transformations replaced by independent representations, or consolidation without Scene Significance Evaluation. Such comparisons would help determine whether the proposed components are essential, merely useful, or unnecessarily complex.
10. Limitations
Workspace of Scenes remains a conceptual and only minimally formalized architecture. It is not presented as an established cognitive or neuroscience theory, and the proposed biological correspondences are hypotheses rather than evidence of biological correctness. Important mechanisms remain open, including Scene Code formation, similarity and pattern separation, source provenance during simultaneous activation, flexible Reference Frame representation, routing control, significance evaluation, and coordination among parallel Scene Navigation Loops.
The hypothesis also does not yet establish computational scalability, learning stability, resource requirements, or advantages over existing Artificial Intelligence architectures. Its abstractions require concrete implementations, operational definitions, and empirical comparisons before their explanatory or engineering value can be assessed.
11. Conclusion
The principal contribution of Workspace of Scenes is a common Scene-based architectural vocabulary connecting perception, memory, simulation, navigation, evaluation, action, and continuous learning. It proposes that these capabilities can emerge from recurrent interaction among specialized components operating over detailed Scene Content, compact Scene Codes, and navigable Scene Relations. Whether this organization provides a useful basis for Artificial Intelligence remains a question for implementation and experiment.
A. Illustrative Extensions: Brain and Cognitive Phenomena
This appendix is an illustrative extension rather than part of the paper’s core claims. It applies the Workspace of Scenes vocabulary to selected brain, cognitive, behavioral, psychological, and artificial-intelligence phenomena as an exploratory map, not as evidence or complete explanation. The listed phenomena involve biological, developmental, social, and environmental factors outside this intentionally simplified architectural account, and the category labels are coarse, non-exclusive groupings intended only to orient the reader.
| Category | Phenomenon | Possible architectural interpretation |
|---|---|---|
Neural mechanisms |
Sparse neuronal firing in the neocortex | A Scene is activated as a sparse representation in the space of the neocortex. There are specialized regions in the neocortex comprising cortical columns that can detect and represent related Features, e.g., human faces. For example, a typical visual sensory input of a human will activate lower-order Features for basic feature detection and higher-order Features, located in different regions of the neocortex. |
Neural mechanisms |
Brain oscillations | Brain oscillations may contribute to temporal coordination, separation, and binding of active Scene Content. The architecture does not require a specific oscillatory code, but timing mechanisms could help determine which distributed Feature activity belongs to one Scene or path step. |
Neural mechanisms |
Neuromodulator | Neuromodulatory systems may provide global or regional bias over excitability, plasticity, routing, significance evaluation, and action selection. In the architectural vocabulary, they are closer to state-setting influences on Scene processing than to ordinary Scene Content. |
Neural mechanisms |
Voting in cortical columns | Column-level voting can be interpreted as local competition among Feature Units and location hypotheses. Compatible hypotheses reinforce one another through lateral and contextual support, while incompatible hypotheses compete through inhibition and sparse selection. |
Perception and prediction |
Pathway "Where" and "What" | The “What” pathway may contribute Scenes representing the internal Feature composition of an object, possibly using object-relative Reference Frames. The “Where” pathway situates a corresponding higher-order Feature within bodily, egocentric, or environment-centered Reference Frames. |
Perception and prediction |
Path integration | Path integration is treated as a candidate mechanism for updating positions within physical or abstract Reference Frames. Grid cell-based path integration models for movement-based object recognition provide one motivation for extending this idea from physical movement to navigation through Scene structure (Leadholm et al. 2021). |
Perception and prediction |
Tracking a moving object | A moving object may remain represented by the same higher-order Feature Unit and associated Scene while the locations of its Feature Instances are updated. In this simplified interpretation, routing through the thalamus may contribute to providing updated location input to the relevant cortical columns. |
Perception and prediction |
Moving object trajectory prediction | A moving object’s recent positions can be represented as a path Scene over space and time. The Mapping Module can extend learned relations from the recent path to candidate next states, while the Integration Workspace compares those predictions with incoming evidence. Persistent prediction error would weaken the current trajectory hypothesis or redirect attention toward a different explanation. |
Perception and prediction |
Representation of many objects of the same kind (e.g., leaves on the tree) at the same time | A higher-order Feature may represent the collection, while only some individual leaves are represented as repeated Feature Instances at approximate locations, constrained by capacity, salience, and task relevance. Navigating into a selected Feature Instance could activate a more detailed Scene of that leaf. |
Perception and prediction |
Recognition | Recognition can be interpreted as constructing candidate Scene Content from evidence, retrieving similar Scene Codes, and comparing activated expectations with incoming Feature Instances. See Section 6.1. |
Perception and prediction |
Recognizing an object that is partially obscured in the visual field | Partial evidence can activate a candidate Scene Code and retrieve predicted Feature Instances for the hidden parts. The Integration Workspace then compares visible evidence with retrieved expectations, while column-level voting and pattern completion preserve the best-supported interpretation until additional evidence arrives. |
Perception and prediction |
Invariance | Invariance can be interpreted as stable activation of a higher-order interpretation across changes in location, scale, orientation, deformation, sensory details, or context. The changed input first forms Feature Instances in the current Reference Frame; rotation, scaling, displacement, or deformation can then be represented as transformations between this Reference Frame and an object-relative, body-relative, or more canonical Reference Frame. If the transformed Feature relations match a recorded Scene or generalized Scene, the same higher-order object interpretation can remain active even though the concrete Scene Content and Feature locations differ. At the mapping level, each transformed observation may still correspond to a related Scene Code. Scene Relations record the transformation, while reusable Feature Units and generalized Scenes capture recurring structure that survives across variants. The lower-level Scene Content may still change; what remains invariant is the selected interpretation or generalized structure, not every active Feature Instance. |
Perception and prediction |
Symmetry detection | Symmetry detection may arise when a Scene can be transformed within its Reference Frame while preserving relevant Feature relations. The preserved structure can be recorded as a relation Scene or generalized Scene connecting the original and transformed states. |
Perception and prediction |
Surprise | An expected contextual Scene, or Scenes, is already active in the neocortex. It sets a corresponding population of neurons in specific columns to the predictive state. When the expected pattern is not met against the sensory input, as a simplification, Feature Units detecting unexpected Features or mismatches contribute a surprise or mismatch signal. That signal may participate in forming a new candidate Scene. Missing expected content is not equivalent to ordinary inactivity. It may be represented by mismatch signals, negative constraints, or inhibitory context associated with an expected Feature-at-Location. The exact mechanism remains open. |
Perception and prediction |
Unconscious predicting if the environment behaves as expected | The brain may make unconscious, parallel predictions of many currently active Scenes to detect whether something changed unexpectedly. It is important for an animal to detect such changes as fast as possible. When such an unexpected change occurs in one of the current Scenes, mismatch and significance signals can redirect attention toward that Scene. The Scene may then be updated to improve future prediction. Whether this redirection becomes conscious access depends on additional mechanisms outside the simplified architecture. |
Perception and prediction |
Unconscious detecting that something is wrong without knowing what exactly | An active prediction Scene may be contradicted by incoming evidence before the system has activated a detailed alternative explanation. The result can be a mismatch or significance signal that something is wrong, even while the specific conflicting Feature or relation remains unresolved. |
Attention and control |
Attention | Attention may be interpreted as prioritization of selected Scene Content, locations, relations, or operations. A higher-order Scene or path Scene may organize the current attentional episode, while lower-order Scenes represent items within its scope. The hierarchy can operate at several timescales, allowing slower contextual navigation to bias faster local processing. |
Attention and control |
Forgetting a thought after distraction | A distraction can replace the active thought Scene with a more urgent or salient Scene. If the previous thought was not consolidated or linked to a strong cue, it may become difficult to reactivate. |
Learning and memory |
Hebbian learning | Hebbian-like plasticity may contribute to strengthening associations among coactive Features, Feature Units, and Scene fragments. In the architecture, such local learning helps future Scene construction and integration, while the Mapping Module remains responsible for Scene Codes and navigable relations. |
Learning and memory |
Learning | Learning occurs at several interacting levels: Feature Units refine detectors and contextual associations, the Integration Workspace forms candidate Scene Content, and the Mapping Module creates or strengthens Scene Codes and Scene Relations. See Section 6.2. |
Learning and memory |
Association | Association may occur when active Scenes overlap in certain columns. That overlap may be detected by Feature similarity (same column), or location proximity (in terms of a Reference Frame, normalized for those Scenes). Such associations can be reflected in Scene Relations, making related Scenes navigable in the future. Multi-level Scene hierarchies allow associations to form among concrete episodes, generalized Scenes, object Scenes, relation Scenes, and path Scenes. |
Learning and memory |
Generalization | Generalization can emerge when similar Scene Codes are retrieved, activated, compared, and integrated. Shared structure may be recorded as a generalized Scene, while context-dependent differences remain represented by separate Scenes or relations. See Section [sec:key_functionalities_generalization]. |
Learning and memory |
Memory consolidation | Memory consolidation may involve replaying recent path Scenes during sleep or rest, especially around events marked as significant by evaluative systems. Reverse or partial replay could help strengthen relations from later consequences back to earlier candidate causes, while forward replay could make future recognition, prediction, and action selection faster in similar situations. |
Learning and memory |
Memories like movies | Movie-like memory may correspond to replaying a path Scene whose Reference Frame organizes component Scenes over time. Recall can be faster, slower, partial, reconstructed, or anchored around durable Scenes rather than an exact recording of continuous experience. See Section 6.3. |
Learning and memory |
Remember something by doing step by step similar things | Finding and reactivating a past Scene may be easier when the agent recreates a similar path Scene. Performing similar steps supplies overlapping Feature Instances, Reference Frame positions, and Scene Relations that can cue the Mapping Module toward the previous path. |
Learning and memory |
Remembering a sequence (e.g., music) | Remembering a sequence can be represented as activating a path Scene whose Reference Frame organizes notes, movements, or events by temporal or ordinal position. Anchor Scenes may cue the next portion when intermediate details are weak. |
Learning and memory |
Replaying a sequence (e.g., music) in memory | Replaying a sequence involves successive activation of component Scenes from a path Scene. The next state is predicted from learned Scene Relations, somewhat like next-item continuation at the behavioral level, but implemented here as navigation through Scene Codes and path Scenes. |
Learning and memory |
Difficulty navigating backward through memory | Durable Scene Codes may act as anchor Scenes, while temporary intermediate Scene Codes and Scene Relations may decay. Recalling a period may therefore require navigating from an available anchor and reconstructing less persistent intermediate states. Backward recall may be especially difficult if Scene Relations are directional or if reverse relations were not strengthened. This remains a hypothesis for future investigation. |
Memory systems |
Working memory | Working memory may correspond to actively maintained Scene Content, candidate path Scenes, and routing configurations that remain available for current operations. It is not identical to the whole Integration Workspace, because much workspace processing may remain outside conscious or task-focused maintenance. |
Memory systems |
Short-term memory | Short-term memory can be interpreted as recently active Scene Content and temporary Scene Codes that remain reactivatable before consolidation or decay. Distraction may replace the Active Scene Landscape before durable relations are strengthened. |
Memory systems |
Long-term memory | Long-term memory corresponds to durable Scene Codes, Scene Relations, path Scenes, and cortical reactivation pathways strengthened through consolidation and repeated use. Generalized knowledge may remain even when a particular episodic Scene Code becomes inaccessible. |
Memory systems |
Implicit memory | Implicit memory may be expressed through strengthened Feature Units, routing tendencies, action-selection biases, and Scene Relations that influence processing without requiring explicit activation of a reportable memory Scene. |
Memory systems |
Procedural memory | Procedural memory may be represented by reusable path Scenes and action Scenes whose component Scenes activate lower-level motor, attentional, or cognitive operations. Execution can proceed with limited explicit access to every intermediate step. |
Memory systems |
Episodic memory | Episodic memory may be represented by path Scenes whose Reference Frames organize component Scenes over time, place, perspective, and context. Durable anchor Scenes can support later reconstruction even when intermediate temporary Scene Codes have weakened. |
Memory systems |
Semantic memory | Semantic memory may correspond to generalized Scenes, relation Scenes, and durable Scene Relations abstracted from many particular episodes. Such structures can be reactivated without reconstructing a specific original experience. |
Memory systems |
Declarative memory | Declarative memory may correspond to durable Scene Codes, Scene Relations, and reactivation pathways that can be intentionally accessed or reported. Facts and episodes differ in their Scene structure and relations rather than requiring entirely separate representational primitives. |
Memory systems |
Lack of hippocampus causes inability to learn new things | When new sensory input is interpreted as candidate Scene Content, the Mapping Module must generate or strengthen a Scene Code and relations that can later reactivate it. If the hippocampal-entorhinal system is unable to provide this indexing and relational-binding function, new episodic Scenes may remain transient in the Integration Workspace and fail to become durable, addressable memories. Intact short-term memory is compatible with active Scene Content being maintained temporarily in cortical workspace activity. Older memories may remain available when their cortical representations and older reactivation pathways were consolidated before hippocampal damage, although the precise boundary between preserved and impaired memory remains a neuroscience question rather than a settled consequence of this architecture. |
Action and behavior |
Sensory-motor learning | Sensory-motor learning can be represented as strengthening relations among evidence Scenes, candidate action Scenes, motor path Scenes, and resulting sensory feedback. Repeated action-feedback loops refine predictions and make future behavior Scenes easier to activate or execute. |
Action and behavior |
Motor sequences | Hierarchical movement through path Scenes, where each step can initiate a lower-level path Scene or motor fragment. Learned motor routines can therefore be treated as reusable behavior Scenes rather than as a separate primitive representation. |
Action and behavior |
Learned template of motor actions | A learned motor template, such as a practiced musical phrase, may be represented as a reusable path Scene whose component Scenes activate lower-level motor fragments. The template is not a literal command script; it is a learned sequence of expected states, actions, and feedback relations. During execution, sensory feedback can confirm the current step, trigger correction, or redirect the action through an alternative Scene Relation. |
Action and behavior |
Executing activities (motor command) | Execution occurs when a selected candidate action Scene or behavior path Scene is routed toward motor systems. Incoming sensory feedback can confirm expected steps, redirect the path through an alternative relation, or create a modified path Scene for later learning. |
Conceptual representation |
Dynamic dimensionality of Scenes | The dimensionality of a Scene may change during learning, especially for abstract concepts. A Reference Frame can become more refined by adding values within an existing dimension, or broader by adding a new dimension when a recurring correlation becomes useful for organizing Scene Content. |
Conceptual representation |
Inheritance of object attributes | An abstract parent object and child object may each be represented through Scenes. The child Scene may be constructed by integrating the parent Scene with additional or overriding Features representing the child’s attributes. Repeating that integration after the parent Scene changes would allow inherited attributes to affect the child representation. Negative overrides, representing an attribute that must not occur in the child Scene, may use negative constraints or inhibitory mechanisms. Under this interpretation, an attribute is modeled as a sub-Feature within the relevant Scene. |
Cognition and reasoning |
Thinking | Thinking can be interpreted as Scene navigation, Scene activation, and Scene integration used for comparison, simulation, recombination, or planning. This may correspond to the experienced process of exploring ideas, changing levels of abstraction, connecting distant concepts, and predicting possible outcomes. See Section 6.4. |
Cognition and reasoning |
Exercising in mind before acting | Mental rehearsal can be described as simulated navigation from a current Scene through possible action Scenes and outcome Scenes, without immediate overt action. Useful rehearsals may create or strengthen Scene Relations, predicted consequences, and action-relevant Feature patterns that can later guide behavior in similar situations. |
Cognition and reasoning |
Internal simulations | Internal simulation activates candidate Scenes or path Scenes without treating them as current sensory evidence. The simulated content is held as hypothetical Scene Content, so it can be compared with goals, constraints, prior knowledge, expected consequences, and possible risks before an overt action is selected. Such simulations may operate at several granularities. A coarse simulation may only test whether a direction is promising, while a more detailed simulation may activate expected Features, missing Features, alternative object states, or successive states along a path Scene. Internal cognitive actions can steer the simulation by continuing a branch, interrupting it, changing the Reference Frame, replacing a Scene fragment, or comparing several candidate outcomes. If a simulated path produces a useful, surprising, or emotionally significant result, the resulting Scene or path Scene may become easier to reactivate later. In this way, imagination, rehearsal, planning, and counterfactual reasoning can all be interpreted as controlled activation and evaluation of Scenes that are not currently being asserted as the external world. |
Cognition and reasoning |
Multi-level hierarchical planning and processing | Just as grid cells let animals simulate routes, they let in the proposed architecture simulate conceptual transitions: "If this policy moves us toward A, what happens if we shift toward B instead?" Higher-order path Scenes can organize lower-level paths and candidate subgoals. Planning can therefore proceed by exploring coarse paths first, then activating more detailed Scenes only where the candidate path remains promising or uncertain. |
Cognition and reasoning |
Intuition | Intuition may correspond to rapid activation of a candidate Scene, relation, or action bias before the supporting path is fully reactivated in the Integration Workspace. It can therefore feel like an immediate answer while still depending on learned Scene Relations and prior evaluation. |
Cognition and reasoning |
Understanding of physics | Understanding of physical regularities can be represented as learned generalized relation Scenes and path Scenes over object states, movements, forces, and outcomes. The architecture does not require an explicit physics module, but it does require learned transformations that predict how Scenes tend to change. |
Cognition and reasoning |
Causal reasoning | Causal reasoning can be interpreted as navigation over learned or hypothesized Scene Relations in which one Scene reliably precedes, enables, prevents, or transforms another. Confirmed predictions strengthen those relations; surprise, missing effects, or alternative explanations weaken or reclassify them. |
Cognition and reasoning |
Generation of new ideas | New ideas may arise spontaneously when currently active context Scenes bias recombination of previously recorded Scenes, or deliberately through Scene integration, path recombination, and translation of structure from one domain into another. Evaluation then determines whether the generated Scene remains a transient candidate, is explored further, or is recorded as significant for later use. |
Cognition and reasoning |
Dreaming | Dreaming may involve internally generated replay, recombination, and partial activation of Scenes and path Scenes while ordinary sensory constraints and action execution are reduced. The Integration Workspace may therefore activate Scene Content that is coherent enough to experience as an unfolding situation, but weakly constrained by current external evidence. In this interpretation, dream episodes can be built from recent experiences, older memories, unresolved goals, emotional or significant Scenes, and spontaneous Mapping navigation. Because external correction is reduced, transitions between Scenes may follow associative, emotional, or structural relations rather than ordinary physical continuity. This could explain why dreams can feel locally meaningful while also containing abrupt changes of place, identity, time, scale, or causal structure. Dreaming may also provide a mode for testing and reorganizing Scene structures. Replayed or recombined path Scenes could strengthen important relations, weaken unstable ones, extract generalized Scenes, rehearse threat or social situations, or explore counterfactual outcomes without overt action. The architecture does not require every dream to have a single function; different dream fragments may reflect consolidation, prediction, emotional regulation, spontaneous activation, or unfinished navigation through active Scene Relations. |
Language and communication |
Talking | Talking can be interpreted as selecting a communicative goal Scene, activating semantic and social context Scenes, and translating them into a linguistic path Scene. Speech production then executes lower-level articulatory or motor path Scenes while auditory, somatosensory, and social feedback update the ongoing Scene and allow correction. |
Language and communication |
Language | Language can be interpreted as a family of path Scenes whose Reference Frames organize sounds, words, phrases, clauses, turns in dialogue, and discourse context. Words and constructions act as linguistic Features or Feature structures that can activate semantic Scenes, social Scenes, action Scenes, and abstract relation Scenes. Understanding language is therefore not only decoding a sequence of tokens. Incoming linguistic Features activate candidate meaning Scenes, while context, speaker intent, prior discourse, and embodied situation constrain which interpretation remains active. Production works in the opposite direction: a communicative goal Scene is translated into a linguistic path Scene, then into lower-level speech, writing, or gesture actions. Grammar can be treated as learned constraints and reusable relation Scenes over linguistic path Scenes. These constraints help determine which Features may combine, which roles they take, and how a sentence or discourse fragment maps to represented objects, actions, events, and relations. |
Language and communication |
Translating between languages | Language translation may involve mapping between language-specific path Scenes, often mediated by semantic Scenes that preserve intended meaning while allowing the surface form to change. Different strategies may vary in how strongly the source-language, semantic, and target-language Scenes are activated. Deliberate translation may rely on an explicit intermediate semantic Scene, while direct thinking in another language may construct a target-language path Scene with little or no activation of the original source-language form. |
Language and abstraction |
Recursion (in abstract concepts, language, etc.) | Recursion can be represented in both directions. Lower-order Scenes, path Scenes, or relation Scenes can be composed into a higher-order Scene; later, that higher-order Scene can treat the embedded structure as one Feature or role while Mapping can still navigate back into its internal organization. This allows language, procedures, social situations, and abstract concepts to build larger structures from smaller ones and also to unpack those structures when detail is needed. For example, words compose phrases, phrases compose clauses, plans contain subplans, and a concept can contain a relation that itself has internal roles and constraints. |
Affect and motivation |
Emotions | Emotions are not modeled as ordinary Scenes alone. They are interpreted cautiously as system-level states involving bodily and interoceptive signals, evaluative processes, hormonal and neuromodulatory influences, learned context, and active Scenes. In this architecture, Scenes can represent the interpreted object, cause, context, predicted consequences, and possible actions associated with an emotion. Hormonal and neuromodulatory systems may bias arousal, attention, routing, significance evaluation, consolidation, and action selection, while Scene structures provide part of the interpreted content to which the emotion is attached. |
Affect and motivation |
Spectrum of high-level behaviors (e.g., preferences) | Stable high-level behavioral tendencies may be represented by persistent, multidimensional context Scenes or families of Scenes that influence evaluation, prediction, and action selection across many situations. Such Scenes can be updated during life and modulated by social, bodily, emotional, and contextual input. This is a simplified architectural interpretation and would require substantial further research before it could be connected to specific psychological constructs or disorders. |
Self and social cognition |
Ego | An Ego Scene may be interpreted as a persistent, hierarchically organized Scene or family of Scenes anchoring bodily state, perspective, goals, constraints, and agent-environment relations. In this cautious interpretation, hippocampal-entorhinal mapping may be especially important for ego-centered and environment-centered continuity, while object and conceptual Scenes can still be represented primarily through neocortical workspace activity. |
Self and social cognition |
Theory of mind | Models of other agents may be constructed by translating or adapting Ego-related Scenes into Scenes representing another agent’s perspective, goals, constraints, and likely actions. This remains speculative, but it preserves the same Scene-based vocabulary used for self, action, and prediction. |
Psychological phenomena |
Deja vu | A deja vu experience could occur when the current Scene strongly matches an existing Scene Code or relational pattern, but source provenance and episodic context are weak, conflicting, or not reactivated. The system may therefore signal familiarity without reconstructing the expected memory context. |
Developmental variation |
Sensory overload | Sensory overload may be interpreted as insufficient filtering, inhibition, routing control, or significance thresholding over incoming Scene Content. Atypical sensory reactivity may include both hyper-reactivity and hypo-reactivity; overload occurs when too many evidence Scenes or Feature Instances compete for the Integration Workspace before context Scenes and routing configurations can stabilize interpretation. |
Developmental variation |
Atypical object labeling or word retrieval | In some language-development profiles, atypical object labeling or word retrieval may occur when partial evidence activates a nearby generalized Scene, object Scene, or linguistic path Scene more strongly than the intended one. This interpretation treats the error as competition among similar Scene Codes, labels, and context-dependent retrieval paths. |
Developmental variation |
Difficulties in detecting emotions of others | Detecting another person’s emotion may require integrating observed facial, bodily, linguistic, situational, and social-context Scenes into a candidate other-agent state Scene. Difficulty may arise when these translations, context integrations, or social predictions are weak, ambiguous, overburdened, or mismatched across different communication styles. |
AI safety |
AI alignment | An AI system built according to the proposed architecture could include intrinsic guardrails defined by durable constraint Scenes or Ego-related Scenes that apply behavioral constraints to learning, internal navigation, and action selection. The biological analogy is cautious: human behavioral constraints may be partly prewired, partly learned, and strongly shaped by social context. In an artificial system, corresponding constraint Scenes would need explicit design and validation rather than being assumed to emerge automatically. Such Scenes could be predefined or pretrained, difficult to modify under ordinary operation, and read-only in critical aspects. Operations that would modify or bypass them could be suppressed by core algorithms. In this interpretation, guardrails would be represented as part of the architecture’s ordinary Scene-based control structure rather than as an entirely external filter. |
B. Illustrative Example of Learning and Using a Scene Transition
Figure 5 summarizes a simplified version of the Scene-transition flow used throughout the proposal. The diagram should be read as an architectural algorithm sketch, not as a claim about a single serial neural process. Several steps may occur concurrently, partially, or at different timescales.
C. Candidate Neuron Abstraction
C.1. Architectural Abstraction and Boundaries
Workspace of Scenes can use a compartmental, context-sensitive neuron abstraction rather than a point neuron that responds only to a weighted sum of undifferentiated inputs. The proposed neuron can distinguish driving evidence, contextual input, top-down feedback, and inhibition, integrate inputs nonlinearly within dendritic segments, maintain short-lived operational states, and adapt its connections through local plasticity.
This abstraction specifies one possible implementation mechanism, not a complete biological neuron model. Individual neurons do not represent complete Scenes, explicitly identify Scene roles, or perform Scene-level operations. Feature Unit and Integration Column behavior emerges from interacting neuronal populations and circuits.
C.2. Dendritic Segments and Input Pathways
The neuron contains multiple dendritic segments, each of which can detect a learned or otherwise configured pattern of coincident synaptic input. A segment can have its own activation threshold and inhibitory input, allowing it to function as a nonlinear integration subunit rather than as a passive contribution to one global sum.
The architecture distinguishes the following functional input pathways:
proximal dendritic segments receive driving feedforward evidence and can directly contribute toward neuronal firing;
distal contextual dendritic segments receive lateral and contextual input that can prepare or bias the neuron;
distal feedback dendritic segments receive top-down input that can prepare, bias, or modulate the neuron’s response;
dendritic-segment inhibition suppresses selected dendritic patterns or compartments;
soma inhibition suppresses firing despite excitatory dendritic activity; and
axon-initial-segment inhibition suppresses or delays neuronal output.
These pathway types are explicit architectural descriptions. Biological circuits may realize them through dendritic location, synaptic connectivity, interneuron types, receptor dynamics, and other mechanisms. Artificial implementations may represent the distinctions directly.
C.3. Scene Roles and Dendritic Routing
One possible neuronal contribution to Scene role-dependent processing is that signals originating from Scenes operating under different roles reach different dendritic compartments or inhibitory pathways. Scene roles are not assumed to be literal labels available to individual neurons. Their effects emerge from how Scene Content is routed and from the timing, connectivity, and type of the resulting inputs.
In a simplified pyramidal-neuron interpretation, signals originating from evidence Scenes primarily contribute through proximal driving pathways. Context Scene signals can contribute through distal contextual pathways, while prediction and other top-down Scene signals can contribute through distal feedback pathways. Goal and constraint Scenes may influence contextual, feedback, and inhibitory pathways. Context-dependent thalamic and cortico-thalamic routing may help select which signals reach these circuits (Hawkins and Ahmad 2016; Grewal et al. 2021; Iyer et al. 2022; Varela et al. 2024).
Contextual and feedback input can alter how strongly or quickly incoming evidence affects a neuron without necessarily causing immediate output. High-priority Scene Content might receive stronger influence through more direct proximal input, perisomatic input, or selected contextual, feedback, and inhibitory dendritic pathways. However, the architecture does not require a universal mapping in which the most important Scene always targets the soma.
Evaluation-related signals, including Scene Significance Evaluation and action-selection bias, may influence temporary excitability and plasticity through contextual, feedback, or inhibitory dendritic inputs. These signals are not treated as ordinary driving evidence. The mapping between Scene roles, dendritic compartments, inhibition, significance-related input, and routing remains a hypothesis for further research.
C.4. Neuron States and Temporal Dynamics
The simplified neuron can occupy several operational states:
inactive, when it has insufficient driving or predictive support;
predictive, when one or more distal contextual or feedback segments recognize a supporting pattern and prepare the neuron to respond;
active, when driving evidence and current modulation cause the neuron to fire;
surprised, when unexpected driving evidence activates the neuron without sufficient predictive support, or when a simplified local mismatch mechanism identifies an unexpected activation;
absolute refractory, when the neuron temporarily cannot fire after a spike; and
relative refractory, when the neuron temporarily requires stronger activation to fire.
A predictive neuron can fire earlier than otherwise compatible but unprepared neurons when matching evidence arrives. Earlier firing can recruit inhibition that suppresses competing interpretations. Refractory states and recent spike timing constrain how quickly the neuron can participate again.
The surprised state is an optional simplification for an artificial implementation. Biologically, mismatch or surprise may instead emerge from population activity, relative timing, inhibitory circuits, or specialized pathways. Similarly, a contradicted Feature Unit and an explicit Integration Column mismatch are circuit-level results rather than required states of one neuron.
C.5. Coincidence Detection and Segment Activation
A dendritic segment becomes active when a sufficiently supported subset of its synaptic inputs is active within an appropriate temporal window. Segment-specific activation thresholds allow different dendritic segments to detect different patterns and degrees of coincidence. An active distal segment can place the neuron in a predictive or otherwise modulated state, while proximal segment activation provides driving support toward firing.
This abstraction permits a neuron to respond differently to the same proximal evidence under different contexts or predictions. It also permits several learned contextual or feedback patterns to influence one neuron without requiring all synaptic input to be treated identically. Active-dendrite and compartmental-neuron models provide biological and computational motivation for this interpretation (Hawkins and Ahmad 2016; Grewal et al. 2021; Iyer et al. 2022).
C.6. Inhibition and Sparse Competition
Inhibition can act on a dendritic segment, the soma, the axon initial segment, or a wider local circuit. Segment-level inhibition can suppress a particular contextual or evidence pattern without silencing every function of the neuron. Soma or axon-initial-segment inhibition can prevent or delay firing despite excitatory support.
At the population level, inhibition helps implement sparse selection, suppress incompatible hypotheses, and coordinate -winner competition among Feature Units. A neuron contributes to this process through its timing and output, but the Integration Column owns the architectural competition and mismatch behavior.
C.7. Local Plasticity and Synaptic Permanence
Synaptic connections can maintain a local permanence or persistence state describing whether an input belongs to a learned dendritic pattern. Permanence is conceptually distinct from the immediate activity of a synapse and need not be identical to an ordinary scalar weight.
Plasticity can depend on presynaptic activity, postsynaptic firing, dendritic-segment activation, relative timing, inhibition, and modulatory significance signals. Such local rules can strengthen confirmed predictions, weaken unsupported associations, and allow previously uncommitted dendritic capacity to learn new contextual or evidence patterns. Coincidence-dependent and spike-timing-dependent plasticity provide biological motivation for timing-sensitive local learning (Markram et al. 1997; Hawkins and Ahmad 2016).
The neuron does not directly learn a Scene Code or Scene Relation. Neuronal plasticity contributes to Feature Unit specialization, local contextual associations, reactivation pathways, and distributed cortical learning.
C.8. Candidate Artificial Neuron Model
One possible artificial implementation uses the following capabilities:
multiple proximal, distal contextual, and distal feedback dendritic segments;
synaptic terminals associated with a segment and functional pathway;
nonlinear segment activation with optional segment-specific thresholds;
inhibitory input at segment, soma, or output level;
inactive, predictive, active, and refractory temporal states, with an optional surprised state;
recent-activity information needed for spike timing and refractory behavior; and
local synaptic permanence and plasticity influenced by contextual, feedback, inhibitory, or significance-related dendritic input.
This initial abstraction may be simplified, reorganized, or replaced as experiments determine which neuronal mechanisms are needed to preserve the relevant architectural behavior.
C.9. Biological Interpretation
Pyramidal neurons, active dendritic branches, inhibitory interneurons, and systems that supply significance-related dendritic input provide possible biological substrates for the proposed abstraction. The architecture does not assume that all neocortical neurons use the same compartments, states, thresholds, plasticity rules, or spiking dynamics.
D. Illustrative Cortical-Layer Interpretation
This appendix gives an illustrative, layer-oriented interpretation of how a cortical column could be read through the Workspace of Scenes vocabulary. Cortical areas differ, granular and agranular cortices differ, and each layer contains multiple cell types and recurrent loops. The tables have a narrow purpose: to relate architectural functions required by an Integration Column to candidate laminar and system-level motifs.
The interpretation draws on the Thousand Brains proposal that grid cell-like location mechanisms exist throughout neocortex. In that proposal, cortical grid cells are predicted to be especially associated with layer 6, while displacement-like cells are predicted to be associated with layer 5 (Hawkins et al. 2019; Lewis et al. 2019). Workspace of Scenes uses these assignments as candidate biological mappings.
| Cortical layer | Possible Workspace of Scenes interpretation |
|---|---|
Layer 1 and apical dendritic input |
May carry top-down context, goals, predictions, attentional bias, and significance-related modulation into local cortical processing. In Workspace of Scenes terms, this layer is a candidate route by which active Scenes, goal Scenes, expected sub-Scenes, or evaluation signals bias Feature Units without directly supplying the main feedforward evidence. |
Layer 2/3 superficial recurrent and long-range circuitry |
May support stable local hypotheses, lateral agreement between columns, and sparse participation in distributed Scene Content. In this interpretation, layer 2/3 activity can contribute Feature-at-location hypotheses to wider Scene processing and can participate in column-to-column voting when multiple columns are trying to settle on compatible interpretations. |
Layer 4 feedforward input, especially in granular cortex |
May correspond to the principal local entry point for bottom-up sensory or lower-order cortical evidence. In Workspace of Scenes terms, this layer supplies candidate evidence for Feature Units and provides the material against which predictions and location-conditioned expectations can be compared. In cortical regions with weak or absent layer 4, analogous input functions may be redistributed across other circuits. |
Layer 5 deep output and action-oriented pathways |
May provide output from the column toward subcortical systems, thalamic relay pathways, motor-related systems, or higher cortical processing. In Workspace of Scenes terms, this layer is a candidate substrate for action-oriented hypotheses, displacement or transition signals, and Scene-change proposals. Hawkins and colleagues predict that layer 5 thick-tufted neurons may function as displacement cells; here that is treated as a possible implementation of local movement, transformation, or path-Scene update signals rather than as a settled fact. |
Layer 6 corticothalamic and corticocortical circuitry |
May help maintain local Reference Frame state, regulate thalamic routing, align feedforward evidence with predictions, and provide location-conditioned feedback to input layers. Hawkins and colleagues predict that cortical grid cell-like neurons may be located in layer 6 or layer 6a. Workspace of Scenes maps this possibility to the Integration Column’s local location system, while leaving open whether biological grid-like coding is truly laminar, distributed, or implemented by another mechanism. |
The remaining systems are not cortical layers, but they are important for interpreting how a column-like circuit could participate in a larger Workspace of Scenes architecture.
| Supporting system | Possible Workspace of Scenes interpretation |
|---|---|
Local inhibitory microcircuits |
May implement sparse selection, gain control, local competition, timing control, and suppression of incompatible hypotheses. In Workspace of Scenes terms, these circuits help local -winner competition among Feature Units and can contribute to mismatch or surprise signals when expected Feature-at-location activity is contradicted. |
Long-range corticocortical pathways |
May exchange selected hypotheses, context, predictions, and compatibility signals between columns and regions. In Workspace of Scenes terms, these pathways support distributed agreement about Scene Content without requiring one column to contain the whole Scene. |
Cortico-thalamic and relay loops |
May coordinate which signals are routed, amplified, suppressed, synchronized, or made available for further processing. In Workspace of Scenes terms, these loops are candidate biological substrates for Relay Module functions such as attention-dependent routing, Scene activation, Scene reactivation, and routing configuration. |
Hippocampal–entorhinal interface |
May anchor, index, and reactivate distributed cortical Scene Content through compact Scene Codes and navigable Scene Relations. In this interpretation, hippocampal–entorhinal systems can help align or reinstate wider Reference Frames, while cortical columns perform local Feature-at-location inference. |
Neuromodulatory, basal-ganglia, and limbic systems |
May supply value, novelty, salience, uncertainty, action-selection, learning-rate, and consolidation biases. In Workspace of Scenes terms, these systems influence which Scenes are stabilized, explored, acted on, or consolidated, without themselves representing the complete Scene. |
The tables should therefore be read as an honest architectural mapping: they identify plausible correspondences between laminar motifs, supporting systems, and Workspace of Scenes functions, while keeping the biological claim provisional. The proposed Integration Column can be useful as an AI design abstraction even if later neuroscience assigns these functions differently across layers, cell types, regions, or recurrent loops.
E. Angsi as an Implementation and Evaluation Framework
Angsi is an AI framework being developed alongside Workspace of Scenes. Its role is to provide a practical path from architectural theory to working experiments: runnable agents, controlled test environments, and concrete model implementations. Workspace of Scenes defines the architectural direction. Angsi is intended to test that direction by turning selected mechanisms into systems that can perceive, act, learn, predict, simulate, and adapt.
The goal is not to wrap existing Artificial Intelligence models. Angsi is intended as a framework for building and evaluating alternative brain-inspired foundations for intelligent agents. These foundations include neuron-like units, local circuits, cortical-column abstractions, sparse representations, prediction, routing, memory, scene-based cognition, internal simulation, and action selection.
E.1. Modeling Levels
Angsi is designed to support several levels of modeling under one experimental direction. At the lower level, it can be used to explore neuron-like units, local circuit motifs, inhibition, sparse activation, contextual prediction, and plasticity rules. Such models are not expected to reproduce biological neurons in full detail, but they can test which local mechanisms are useful for Feature Units, prediction, mismatch detection, and local learning.
Angsi can model Integration Columns, macrocolumn-like structures, Feature Units, local Reference Frames, and column-to-column coordination. Workspace of Scenes depends on local Feature-at-location hypotheses that can cooperate, compete, and participate in larger distributed Scenes. Angsi can also implement Workspace of Scenes agents. Such agents would construct Scene Content, form or retrieve Scene Codes, navigate Scene Relations and path Scenes, evaluate candidate actions, and continue learning during interaction with an environment. The same framework direction can therefore connect low-level mechanisms with higher-level cognitive behavior.
E.2. Research Directions
Angsi should make the theory testable through constrained but meaningful environments. Initial experiments can focus on recognition, prediction, planning, action selection, memory-guided behavior, and continuous learning. A useful test environment should reveal not only whether a behavior succeeds, but which architectural components were required and which assumptions failed when implemented.
The framework can also support comparison between modeling choices. For example, one implementation may use more biologically detailed neuron-like units, while another may use more abstract software structures that preserve the same architectural responsibilities. Comparing such variants can help determine which parts of Workspace of Scenes are essential, which can be simplified, and which should be revised.
E.3. Research Status
Angsi is a proprietary AI framework and is not currently open source. This should be considered when evaluating reproducibility and external validation. Its current role is to support implementation milestones, expose missing algorithmic details, and provide concrete demonstrations that can be reviewed, reproduced, and improved through appropriate collaboration.
In this sense, Angsi complements the present paper. The paper proposes the architecture. Angsi provides the experimental path for testing whether the architecture can become a working foundation for more general, adaptive, brain-inspired Artificial Intelligence.