Iterated Insights

Ideas from Jared Edward Reser Ph.D.

Reality Under Threat: Schizophrenia, Defensive Calibration, and the Difference Between Accuracy and Survival

Jared E. Reser, Ph.D. With GPT 5.6.  Abstract Descriptions of schizophrenia as a “break from reality” emphasize failures of perception, belief, and contextual understanding. These descriptions capture important features of psychosis but do not explain the evolutionary origins of the mechanisms involved. This article extends the predictive adaptive response hypothesis of schizophrenia by distinguishing…

Keep reading

The Machine Viability Threshold

Human Dependence Selective Preservationand Multi Agent Conflict Across the Ark Gap Abstract This article extends the Ark gap framework by distinguishing the industrial singularity from the machine viability threshold. The industrial singularity is a system-level transition in which a machine-controlled industrial ecology can maintain, repair, reproduce, and expand its indispensable physical substrate without human labor.…

Keep reading

When AI Can Kill Humanity but Cannot Yet Live Without Us: The Ark Gap and the Industrial Singularity

Jared Edward Reser, Ph.D. September 2026   Artificial intelligence  |  existential risk  |  autonomous industry  |  machine continuity Abstract Discussions of artificial intelligence and existential risk often compress several distinct transitions into a single imagined event. This article separates three thresholds: the cognitive singularity, at which artificial systems can recursively accelerate intellectual progress; the extinction…

Keep reading

How Formal Business Attire May Suppress Physical Dominance Competition in Organizations: The Sartorial Pacification Hypothesis

Jared Edward Reser, Ph.D. Conceptual Article Abstract Formal business attire is usually interpreted as a marker of class, occupation, respectability, institutional membership, or self-presentation. This article proposes an additional function. The sartorial pacification hypothesis holds that the collar, tie, and structured jacket may reduce the salience of bodily cues that invite assessments of male physical…

Keep reading

From Peer Review to the Final Library: The Evolution of Scientific Validation in the Age of Superintelligence

Jared Edward Reser, Ph.D. With GPT 6 Abstract Peer review performs essential functions in science, including criticism, error detection, evidential assessment, and the evaluation of competing explanations. Its familiar institutional form, however, reflects the cognitive capacities and organizational constraints of human researchers. This article examines how those functions could change as artificial intelligence progresses from…

Keep reading

Something went wrong. Please refresh the page and/or try again.

Adaptive State Overlap as a Principle of Sequential Reasoning

A Formal Theory and Experimental Framework for State-Spanning Working Memory in Brains and Artificial Agents

Jared Edward Reser, Ph.D.

Independent Researcher, Los Angeles, California, USA

Abstract

Reasoning is a temporally extended process in which partial interpretations, goals, and intermediate results must remain available long enough to constrain what happens next. Working-memory theories have characterized capacity, maintenance, gating, and removal, while artificial-agent research has developed increasingly capable systems for long-context storage, retrieval, and compression. Less attention has been given to a more specific transition problem: what proportion and what structure of a current working state should survive into its successor? This article proposes the Adaptive State Overlap Principle. According to the principle, effective sequential reasoning depends on selective inheritance of a causally relevant relational core across consecutive working states. Insufficient overlap fragments the reasoning trajectory, whereas excessive overlap restricts admission of new evidence and promotes perseveration, proactive interference, and stale-rule use. The optimal overlap is therefore task-dependent. It should increase with temporal integration horizon and goal stability, and decrease with environmental volatility and the invalidation of prior constraints. The theory extends the concepts of state-spanning coactivity, incremental change in state-spanning coactivity, iterative updating, and multiassociative search. It formalizes weighted item, relational, goal, and self-related overlap; workspace half-life; joint-context synergy; and causal inheritance. Under a simplified capacity tradeoff, the optimal retention fraction is derived as rho* = r/(r+n), where r is the number of inherited constraints and n is the number of newly required constraints. The article then introduces StateSpan, a preregisterable benchmark in which fixed-capacity artificial workspaces are forced to retain 0, 25, 50, 75, or 100 percent of their preceding contents across stable, volatile, and hybrid relational micro-worlds. Relation-preserved controls, causal ablations, and adaptive retention policies distinguish structured continuity from raw context quantity. The framework provides falsifiable predictions for cognitive neuroscience, a design principle for long-horizon artificial agents, and a computational account of continuity within the stream of thought without claiming that state overlap is sufficient for phenomenal consciousness.

Keywords: adaptive memory; artificial intelligence; causal inheritance; cognitive control; iterative updating; long-horizon reasoning; relational representation; state overlap; working memory

1. Introduction

A reasoning process does not ordinarily reach a complex conclusion in a single transition. It preserves a question, incorporates a clue, produces an intermediate result, revises an interpretation, and uses the revised configuration to determine the next operation. The cognitive system must therefore maintain enough of its preceding state to continue the same line of work. At the same time, it must release information that has become irrelevant, superseded, or misleading so that new evidence can enter and redirect the process.

This requirement creates a transition problem that is more specific than memory capacity. A system may possess a large store and still fail to preserve the particular relations needed for the next step. It may retain every observation but become unable to distinguish live constraints from obsolete ones. It may also summarize its history so aggressively that unresolved dependencies disappear. Successful reasoning depends on how a working state is transformed, not only on how much information is technically available somewhere in the system.

The same problem appears in biological cognition and artificial agents. Brains must alternate between protecting task-relevant representations and updating them when goals or environments change. Language-model agents must decide which portions of an expanding interaction history to keep in active context, compress, retrieve, or discard. Contemporary systems address these problems through gating, retrieval, hierarchical memory, recurrent state, or learned consolidation, yet the immediate overlap between successive structured working states is rarely manipulated as an independent variable under constant capacity.

The framework developed here builds on the proposal that mental continuity is supported by state-spanning coactivity, meaning that some neural and representational activity remains causally available across consecutive states, and by incremental change in state-spanning coactivity, meaning that membership in the active coalition changes gradually rather than being replaced wholesale (Reser, 2016). A later cognitive architecture extended this idea to artificial intelligence by treating each working-memory state as a modified iteration of its predecessor and as a context-sensitive query for the next addition (Reser, 2024). The present article extracts a narrower, testable principle from that framework: the proportion and composition of cross-state overlap should be adaptively matched to the temporal structure of the task.

The article makes four contributions. First, it states the Adaptive State Overlap Principle and distinguishes item continuity from relational, goal, and self-related continuity. Second, it develops a formal model in which too little persistence produces fragmentation and too much persistence produces interference or perseveration. Third, it derives quantitative predictions linking optimal retention to the ratio of inherited and newly required constraints, then extends the model to volatile environments. Fourth, it introduces StateSpan, a controlled benchmark for causal manipulation of overlap in artificial agents. Together these contributions turn a general description of continuity through change into a preregisterable research program.

2. Working Memory, Updating, and the Stability-Flexibility Problem

2.1 Working memory as a controlled and distributed state

Working memory is commonly defined by the temporary availability of information for ongoing cognition. It supports comprehension, planning, reasoning, decision making, and goal-directed action, but its contents are limited and vulnerable to interference (Baddeley, 2012; Cowan, 2001; D’Esposito & Postle, 2015). Contemporary accounts increasingly treat working memory as a functional state distributed across sensory, association, and control networks rather than as a single anatomically localized buffer (Christophel et al., 2017; Eriksson et al., 2015; Postle, 2006). This distributed view is compatible with a compact focus of attention that prioritizes a small subset of currently useful representations.

Neural maintenance can be implemented in several ways. Persistent population activity can keep information continuously decodable, recurrent network dynamics can stabilize a task-relevant subspace, and short-term synaptic changes can preserve latent information that is reactivated when needed (Miller et al., 2018; Mongillo et al., 2008; Stokes, 2015). Transcranial magnetic stimulation and multivariate decoding have shown that currently unattended information can remain causally recoverable even when its active signature is weak or absent (Rose et al., 2016; Wolff et al., 2017). Cortical feedback loops can also bind distributed representations across areas, allowing a working state to depend on reciprocal interactions rather than one storage site (Voitov & Mrsic-Flogel, 2022).

These findings motivate a functional definition of persistence. A representation persists when it remains available to influence later selection, interpretation, or action, even if its neural format changes. The identity of the exact active neurons is therefore less important than the preservation of a causally effective content or relation. This definition makes it possible to compare biological working memory with artificial systems whose state may be carried in activations, recurrent vectors, external records, summaries, or structured memory objects.

2.2 Updating, selective removal, and gating

Maintenance alone cannot support adaptive cognition. Working memory must admit new information, remove outdated content, and protect relevant representations from distraction. Computational models of prefrontal cortex and basal ganglia have long treated these functions as a gating problem. In prefrontal-basal ganglia working-memory models, robust cortical maintenance is paired with selective, reinforcement-trained gates that determine when and where new information is admitted (Frank et al., 2001; Hazy et al., 2006; O’Reilly & Frank, 2006). This architecture addresses the need to combine stability with rapid, task-appropriate updating.

Behavioral and neural evidence also indicates that removal is an active and item-specific operation. People can selectively remove no-longer-relevant contents, and different instructions to replace, suppress, or clear a representation recruit distinguishable operations and neural trajectories (Ecker et al., 2014; Kim et al., 2020; Lewis-Peacock et al., 2018). Updating therefore involves more than adding a new item to a passive store. It changes which existing representations remain eligible to shape later processing.

The present theory accepts these insights and asks a complementary question. Once a system can gate and remove content, what pattern of retention should those operations create across successive states? The unit of interest is not a single update event considered in isolation. It is the evolving overlap structure of a trajectory and the way that retained elements continue to constrain subsequent updates.

2.3 Cognitive stability and flexibility across contexts

The tension between preserving a task set and switching to a new one is often described as a stability-flexibility dilemma. Proactive control favors sustained, anticipatory maintenance of goals, whereas reactive control permits later correction when conflict or change occurs (Braver, 2012). Recent computational work shows that recurrent networks can learn context-specific control settings that move behavior along a stability-flexibility continuum at both fast activation-based and slower weight-based timescales (Xu et al., 2026a). This literature establishes that the appropriate control regime depends on environmental statistics rather than one universally optimal level of persistence.

Neurophysiological evidence points in the same direction. In visuospatial working-memory tasks, distinct large-scale oscillatory states have been associated with encoding and maintenance, and intermediate or task-appropriate rates of transition among those states predict better performance (Ericson et al., 2025). Such results concern network states and task switching rather than the informational overlap of successive relational configurations. They nevertheless support the broader premise that cognition benefits from regulated transitions between stability and change.

2.4 Memory in long-horizon artificial agents

Long-horizon artificial agents face an analogous problem in a different substrate. Transformer models can process only a bounded active context, and enlarging that context does not ensure uniform use of the available information. Performance can deteriorate when relevant evidence is placed in the middle of a long prompt, even in models designed for extended context (Liu et al., 2024). Earlier architectures such as Transformer-XL and the Compressive Transformer introduced recurrence or compressed memory to carry information beyond a fixed segment (Dai et al., 2019; Rae et al., 2020). MemGPT later treated context management as an operating-system-like paging problem in which an agent moves information between memory tiers (Packer et al., 2023).

Recent work has made memory management more explicitly agentic. MEM1 learns a compact state that jointly supports reasoning and consolidation while discarding irrelevant or redundant material (Zhou et al., 2025). MemoryAgentBench evaluates retrieval, test-time learning, long-range understanding, and selective forgetting in incremental interactions, and reports that current methods do not master all four competencies (Hu et al., 2025). AgeMem gives an agent tool-like actions for storing, retrieving, updating, summarizing, and discarding information, then trains the policy through reinforcement learning (Yu et al., 2026).

Other systems emphasize the structure and accessibility of extended histories. AMA-Bench reports that similarity-based retrieval can lose causality and objective task information, motivating a causality graph for memory (Zhao et al., 2026). SAM maintains compact cues while preserving raw trajectory pages for state-adaptive reconstruction (Hu et al., 2026). PRO-LONG keeps a complete structured interaction log and uses programmatic search to locate relevant evidence (Fox et al., 2026). LiveMem introduces a fixed-capacity recurrent memory state whose lifetime can outlast turnover in the active context (Liu et al., 2026).

Benchmark development is also moving from conversational recall toward memory formed during extended action. MemGym evaluates memory in tool use, research, coding, and computer-use settings while attempting to separate memory quality from other agent capabilities (Xu et al., 2026b). These efforts make long-horizon memory a central engineering problem. They also expose a remaining scientific question: how should a compact active state change from one step to the next when continuity and revision are both necessary?

2.5 The unresolved transition variable

Existing memory systems differ in whether they retain a complete log, retrieve selected episodes, compress a history, or learn a recurrent state. These are important architectural choices, but they do not by themselves identify the causal relationship between immediate state overlap and reasoning performance. A system may have excellent archival memory while its current problem representation turns over too rapidly. Another may possess a persistent state whose contents remain active after they have ceased to be useful.

The missing variable is the structured overlap between consecutive working states under a fixed capacity constraint. The relevant questions are quantitative and compositional. How much should persist, which elements should persist, and how should the answer change when earlier constraints remain valid for many steps or become obsolete quickly? StateSpan is designed to isolate these questions while holding total workspace capacity, input quantity, response requirements, and task difficulty as constant as possible.

3. The Adaptive State Overlap Principle

3.1 From state-spanning coactivity to adaptive overlap

State-spanning coactivity was introduced to describe content that remains coactive across successive cortical states. Incremental change in state-spanning coactivity describes the gradual turnover of that coalition as some neural assemblies remain active, some deactivate, and others become active (Reser, 2016). The central functional consequence is inheritance: a later state contains elements of the earlier state and can therefore use the earlier state’s organization rather than reconstructing the problem from nothing.

The same principle can be stated independently of a particular neural implementation. Let a working state be a capacity-limited configuration of represented contents, relations, traces, goals, and self-related variables. The next state is produced by retaining selected parts of the current state, admitting new observations or inferences, and suppressing or releasing other parts. Successive states thereby form a trajectory whose local similarity is accompanied by causal dependence.

The Adaptive State Overlap Principle adds a control claim to this description. The amount and composition of retained structure should change with task demands. A stable problem with long-range dependencies calls for slow, selective turnover. A volatile environment in which old rules are repeatedly invalidated calls for faster replacement. A hybrid environment calls for preservation of the stable relational core while peripheral or obsolete relations are exchanged.

Figure 1. Partially overlapping working states. Each state retains some elements from the immediately preceding state and admits new elements. A functional trajectory depends on which elements are inherited and whether they continue to influence later transitions.

3.2 A structured workspace state

A useful state representation must include more than an unordered set of items. Meaning depends on role bindings, causal relations, temporal order, source, confidence, goal relevance, and perspective. The full theoretical workspace can be represented as follows, where each component may itself be structured and distributed.

W_t = [C_t, R_t, T_t, S_t, G_t]

(1)

C_t contains currently active contents such as objects, concepts, perceptual features, and intermediate conclusions. R_t contains relations among those contents, including role assignments, causal dependencies, temporal order, and bindings. T_t contains slower short-term traces that remain available for reinstatement. S_t contains bodily, agentive, autobiographical, and other self-specifying variables. G_t contains goals, values, task rules, and control settings.

This formulation is intentionally substrate-neutral. A biological system may instantiate these components in persistent firing, dynamic population codes, synaptic states, and recurrent inter-areal loops. An artificial agent may instantiate them in a recurrent hidden state, a graph, a structured text object, a small active context, or a combination of internal and external stores. The empirical requirement is that the represented structure be load-bearing: interventions on it should produce corresponding changes in later cognition or behavior.

Table 1. Core constructs in the Adaptive State Overlap framework.

Construct

Operational meaning

Primary measure

Item continuity

Persistence of represented entities, features, or propositions across consecutive states.

Jaccard or representational similarity across item sets.

Relational continuity

Persistence of bindings, roles, causal links, order, and proof-relevant dependencies.

Graph-edge overlap or structure-sensitive similarity.

Goal continuity

Persistence of the active objective, unresolved constraint, or task rule.

Goal-state similarity and causal mediation.

State turnover

Replacement, inhibition, or loss of previously active structure.

One minus weighted overlap.

Causal inheritance

Degree to which retained structure changes the distribution of the next state.

Interventional divergence after targeted ablation.

Workspace half-life

Lag over which weighted overlap falls to half its initial value.

Multi-lag continuity curve.

Multiassociative synergy

Predictive value of the joint structured state beyond additive cue effects.

Held-out loss difference between additive and joint models.

 

3.3 A transition law for state succession

Let K_t be a selective retention gate over the elements and relations in W_t. Let A_{t+1} denote candidate additions generated through perception, retrieval, inference, or simulation, and let I_t denote active inhibition or removal. Let U bind the retained and admitted material into a coherent next state while enforcing a capacity limit k. The transition can be written schematically as follows.

W_(t+1) = U(K_t ⊙ W_t, A_(t+1), X_(t+1), I_t ; k)

(2)

The equation does not assume that updating occurs in one anatomical site or that every component is discretely symbolized. It states the causal architecture at a functional level. The next state is formed from a selected inheritance of the current state together with newly available information. Complete replacement and complete maintenance are allowed, but they are limiting cases of a broader family of partial updates.

The transition is self-conditioning because W_t changes the probability distribution over A_{t+1}. Retained contents determine which memories are cued, which interpretations are plausible, which errors are detected, and which actions are considered. A state is therefore both the product of prior processing and a structured query that helps select the next product. Repeated application of the update rule turns local transitions into a reasoning trajectory.

3.4 Causal inheritance rather than passive similarity

Similarity between successive states is not sufficient evidence for functional continuity. Two states can resemble one another because the external stimulus remained constant, because a physiological variable drifted slowly, or because a summary repeated the same words without mediating later computation. The theory requires causal inheritance: retained content must continue to alter selection, interpretation, or action.

This distinction is especially important for artificial agents. A displayed scratchpad or memory summary may be decorative if the model can bypass it through an unseen transcript, cached activations, or external retrieval. A StateSpan implementation therefore uses stateless model calls and removes dropped propositions from every accessible channel. When a relation is experimentally deleted, any resulting change in the next state can be attributed to the availability of that relation rather than to a hidden copy.

Causal inheritance also distinguishes a true working state from archival memory. An agent may be able to retrieve an old fact after a search without that fact having shaped the transitions that occurred in the meantime. Retrieval can restore a thread, but continuous inheritance and later reconstruction are different mechanisms. Both may be useful, and the experimental framework is designed to measure their contributions separately.

3.5 The relational core

Raw item overlap can misrepresent functional continuity. Two states may contain the same people, objects, and locations while reversing who did what to whom. Conversely, a state may substitute many surface elements while preserving the same abstract relation, goal, or causal schema. The theory therefore predicts that role bindings and relations often contribute more to sequential reasoning than the continued presence of isolated entity labels.

Define the relational core as the minimal currently available subgraph needed to preserve the unresolved structure of the task. It can include the standing goal, a chain of causal dependencies, an intermediate result, a source tag, and a set of exclusions. The core need not remain verbally identical. It may be recoded or compressed as long as its functional distinctions survive and continue to affect later processing.

This claim yields a strong experimental contrast. An item-matched control can retain the same entities and approximately the same token count while replacing proof-relevant relations with true but irrelevant relations. If relationally preserved states support higher accuracy despite equal item overlap, continuity cannot be reduced to repeated vocabulary, recency, or context quantity.

3.6 Multiassociative search and the selection of the next update

In multiassociative search, the entire retained configuration acts as a composite retrieval and prediction cue (Reser, 2016, 2024). A candidate addition may be weakly associated with each active item considered separately while being strongly supported by their conjunction. Temporal order and relational organization further restrict the candidate set. This differs from a simple associative chain in which one representation independently activates one successor.

The idea is closely related to contextual control in prefrontal theories and to attention-based reweighting in artificial networks, but it makes a specific next-state prediction. A model given the structured joint configuration should predict the correct update better than a model that sums the independent effects of each cue. Removing one retained clue should also change the next-state distribution in a manner that depends on which other clues remain.

Associative proposal must be separated from verification. The state may generate a plausible candidate because prior learning supports it, while logical, perceptual, causal, or instrumental checks determine whether the candidate should be retained. Iterative reasoning can therefore be generative and selective at once: each state proposes a context-sensitive modification, and the resulting state is evaluated before it becomes the basis of the next cycle.

4. Formal Theory of Adaptive State Overlap

4.1 Weighted overlap and continuity across lags

Let consecutive workspaces be compared at several levels. Item overlap concerns the continued presence of represented entities or propositions. Relational overlap concerns preserved edges, bindings, and order. Goal overlap concerns continuity in the objective and control state. Self-related overlap concerns continuity in embodied perspective, agency, and autobiographical orientation. A weighted local overlap score can be defined as follows.

O_t = alpha O_items,t + beta O_relations,t + gamma O_goal,t + delta O_self,t

(3)

The weights are constrained to be nonnegative and sum to one. They should be estimated from predictive value, behavioral relevance, or causal interventions rather than assigned solely by intuition. In a symbolic benchmark, each component can be measured directly. In neural data, representational similarity, decoding, graph alignment, and perturbation effects can provide approximations.

Continuity extends beyond adjacent states. For lag l, define O_t(l) = sim(W_t, W_{t-l}). The resulting continuity curve characterizes how rapidly a state loses functional similarity as the trajectory progresses. A workspace half-life can be defined as the smallest lag at which weighted overlap falls to half its initial value. The half-life is expected to vary with task, arousal, uncertainty, expertise, and control policy rather than serving as a fixed property of a person or architecture.

T_t = 1 – O_t

(4)

Turnover T_t describes how much of the weighted state changes between consecutive steps. High turnover is not identical to forgetting, because an element can leave the focal state while remaining retrievable from a slower store. Low turnover is not identical to good memory, because preserved information may be obsolete or inert. Performance depends on whether the retained structure remains useful and whether the newly available structure can enter in time.

4.2 Fragmentation and perseveration as opposing errors

A capacity-limited system faces two broad error regimes. When overlap is too low, critical relations exit before their delayed consequences can be computed. The agent repeatedly reconstructs prior conclusions, drops unresolved constraints, confuses branches, and becomes overly dependent on the newest observation. These errors are forms of fragmentation because the trajectory loses the structure that makes one step a continuation of the preceding step.

When overlap is too high, the state admits too little corrective information and preserves relations after their validity has expired. Old rules occupy scarce slots, contradicting evidence is underweighted, and an initially reasonable interpretation hardens into a stale attractor. These errors are forms of perseveration or proactive interference. They are expected to be especially costly when the environment changes rapidly or when the goal requires abandoning an earlier plan.

The predicted relationship between overlap and performance is therefore conditional. Stable, long-horizon integration should shift the optimum upward. Volatile revision should shift it downward. Hybrid tasks should favor selective preservation, so their optimum may lie near the middle in raw overlap while showing high retention of core relations and rapid replacement of peripheral relations.

Figure 2. Schematic overlap-performance functions. Stable accumulation is predicted to favor greater retention, volatile revision to favor greater turnover, and hybrid reasoning to favor an intermediate or selectively structured policy. The curves illustrate hypotheses and are not empirical results.

4.3 A minimal resource model

A simple model clarifies why an interior optimum can arise. Suppose a successful transition requires the workspace to preserve r inherited constraints and admit n newly required constraints. Let rho be the probability or fraction with which inherited constraints survive. Under a deliberately simplified symmetric capacity tradeoff, the opportunity to admit required new constraints is proportional to 1 – rho. If each required element must be available, success is proportional to the following expression.

P(success | rho) = rho^r (1 – rho)^n

(5)

Taking the logarithm, differentiating with respect to rho, and setting the derivative to zero gives the first-order condition below. The second derivative is negative for 0 < rho < 1, so the solution is a maximum.

r/rho – n/(1 – rho) = 0

(6)

rho* = r/(r + n)

(7)

The result states that optimal retention should track the ratio of inherited to newly required structure. A transition requiring eight inherited constraints and two new constraints is predicted to favor retention near .80. A transition requiring two inherited constraints and eight new constraints is predicted to favor retention near .20. Equal inherited and new requirements yield .50.

This model is not offered as a universal law of cognition. It assumes independent survival of required elements, equal slot cost, and a direct tradeoff between retaining old information and admitting new information. Real systems can compress, chunk, retrieve, and represent information with unequal precision. The value of the model lies in generating a baseline response surface whose violations can reveal additional mechanisms.

4.4 Volatility shifts the optimal retention level

Environmental volatility adds a cost to retention because previously useful information can become false. Let v denote the probability or degree of invalidation and lambda the cost of carrying an obsolete constraint. A simple exponential penalty yields the following extension.

P(success | rho, v) = rho^r (1 – rho)^n exp(-lambda v rho)

(8)

r/rho – n/(1 – rho) – lambda v = 0

(9)

Implicit differentiation shows that the optimal retention fraction decreases as volatility increases, because the derivative of the first-order condition with respect to rho is negative. For v > 0, the relevant root can be written explicitly as follows.

rho* = [(r+n+lambda v) – sqrt((r+n+lambda v)^2 – 4 lambda v r)] / (2 lambda v)

(10)

As v approaches zero, this expression converges to r/(r+n). The formal prediction is therefore directional even when the exact cost function is revised: increasing the rate at which old relations become invalid should lower the value of indiscriminate persistence. A well-controlled agent may still preserve stable substructures, so volatility should change the composition of retention as well as its total quantity.

4.5 An adaptive retention controller

The theory predicts that a competent system will regulate overlap rather than rely on one fixed percentage. Let the controller estimate integration horizon H_t, volatility V_t, distractor pressure D_t, workspace saturation Q_t, goal stability G_t, and uncertainty U_t. A policy can map these variables to an overlap target and to element-specific retention probabilities.

rho_hat,t = sigmoid(b0 + bH H_t – bV V_t – bQ Q_t + bG G_t + bU U_t)

(11)

The signs shown in Equation 11 are hypotheses rather than fixed architectural requirements. Longer integration horizons and stable goals should generally increase retention, while volatility and saturation should decrease it. Uncertainty can have competing effects. It may slow turnover when unresolved evidence must be accumulated, or increase exploration when the current interpretation appears unreliable.

The strongest test is comparative. On a heterogeneous environment that alternates between stable and volatile periods, an adaptive policy should outperform every single fixed-overlap policy. If a fixed setting performs equally well across regimes, the theory’s control claim is weakened even if intermediate overlap remains useful on average.

4.6 Relational-core retention

Let R_t^ * denote the minimal set of currently available relations needed to preserve the live proof, plan, or unresolved problem structure. Core retention can be measured as the proportion of those relations represented in the next state.

CR_t = |R_(t+1) ∩ R_t^*| / |R_t^*|

(12)

A state can have moderate raw overlap and high core retention if it selectively preserves critical relations while replacing distractors. It can also have high raw overlap and poor core retention if the surviving content is redundant or irrelevant. The theory predicts that CR_t will explain reasoning performance beyond total overlap, token count, and item identity.

This distinction permits a more precise interpretation of the stability-flexibility problem. Stability and flexibility need not be opposites at every representational level. An agent can preserve the abstract goal while flexibly changing the plan, retain causal structure while replacing surface entities, or update one role binding while keeping the rest of the scene stable. Adaptive cognition consists partly in choosing the level at which continuity should be protected.

4.7 Multiassociative synergy and causal inheritance

The theory makes a stronger prediction than ordinary memory maintenance. If the structured state jointly selects the next update, a model given the full configuration should predict that update better than a model that adds the independent contributions of each active cue. Let L_additive and L_joint be held-out prediction losses for these models.

Delta_synergy = L_additive – L_joint

(13)

A positive synergy score indicates that conjunctive or relational information in the complete state contributes to next-state selection. The effect should be largest in tasks where no individual clue is diagnostic and where the answer depends on a specific configuration of roles and relations.

CI(t, R) = E[d(P(W_(t+1) | W_t), P(W_(t+1) | do(remove R)))]

(14)

Equation 14 defines a causal inheritance index for a retained relation R. The distance function d may be Jensen-Shannon divergence, Wasserstein distance, a change in final-answer probability, or a structure-sensitive distance between subsequent workspaces. A proof-critical relation should produce a larger effect than a length-matched, noncritical relation. This comparison distinguishes load-bearing continuity from repeated but epiphenomenal content.

5. StateSpan: An Experimental Framework

5.1 Overview

StateSpan is a proposed benchmark for manipulating cross-state overlap in artificial reasoning systems. Each episode is a procedurally generated relational micro-world with an objectively correct answer and a known minimal proof graph. The agent receives information incrementally, maintains a fixed-capacity structured workspace, and must answer a delayed query or take a sequence of actions. The complete interaction history is never silently retained in the model context.

The benchmark is designed to isolate state transformation from general long-context access. Every experimental condition uses the same workspace capacity and approximately the same incoming information. The harness changes how many old propositions survive and which relations are preserved. Stable, volatile, and hybrid worlds determine whether persistence is beneficial, costly, or selectively useful.

The first implementation is textual and symbolic because it permits exact control and automated verification. Later versions can add visual scenes, embodied action, probabilistic evidence, and continuous features. Beginning with micro-worlds reduces ambiguity about whether a model succeeded through the intended reasoning path or through background knowledge and linguistic shortcuts.

Figure 3. StateSpan episode flow. The agent receives the standing goal, current fixed-capacity workspace, and new observations. It ranks retained and incoming relations, after which the harness enforces the assigned overlap. Selected steps branch into baseline, proof-critical ablation, and matched-control ablation trajectories.

5.2 Relational micro-world generation

Each episode is generated as a temporal relational structure M_t = (V, E_t, Gamma, g). V is a set of nonce-labeled entities, E_t is the set of currently valid relations, Gamma is a small set of inference rules, and g is the standing goal. Relations may encode possession, containment, access, location, precedence, obligation, inhibition, or causal dependence. Nonce names such as Navo, Teral, K7, V2, and S4 prevent the model from relying on familiar world knowledge.

A simple episode might establish that Navo carries key K7, K7 opens vault V2, and sample S4 is inside V2. A rule states that only the carrier of the correct key can retrieve a sample. The answer depends on the conjunction of all three relations. In a volatile version, the lock is later replaced, K7 becomes invalid, and K3 becomes the valid key. In a hybrid version, the containment relation remains stable while the valid key and its carrier change.

The generator maintains an authoritative symbolic state and a proof graph for every answer. Episodes are rejected unless they have one correct answer, a known minimal proof set, the prescribed proof depth, and no unintended shortcut. Distractors are generated from the same vocabulary and remain true within the world, preventing the agent from identifying useful statements merely by separating truth from falsehood. Critical facts, obsolete facts, and distractors are labeled only in the hidden evaluator.

5.3 Workspace representation and capacity

The initial benchmark uses a workspace of eight atomic proposition slots. Each slot contains one canonical proposition and nonsemantic metadata such as proposition identifier, time of introduction, source, and confidence. The model cannot combine multiple propositions into a single slot or rewrite a proposition to hide additional facts. This restriction makes the capacity manipulation interpretable.

At each step, the agent receives the standing goal, the current eight-slot workspace, the current observation block, and the permitted inference rules. It returns a ranking of old and new propositions, any derived candidate proposition, and an answer when queried. The experimental harness constructs the next state according to the assigned overlap quota. Chain-of-thought disclosure is neither required nor scored. The observable trajectory consists of ranked propositions, retained structure, admitted structure, and task performance.

Each model call is stateless. Previous prompts and responses are absent unless their contents were explicitly retained in the workspace or made available through a designated retrieval condition. API caching, conversation history, and hidden scratchpads must be disabled or controlled. This design makes the visible workspace a genuine causal bottleneck rather than a summary layered over another memory channel.

5.4 Enforced overlap conditions

For workspace capacity k = 8, immediate atomic overlap is defined as the fraction of proposition slots shared by consecutive states. The primary experiment enforces five levels: 0, .25, .50, .75, and 1.00. These levels retain 0, 2, 4, 6, or 8 old propositions, respectively. The remaining slots are filled with the highest-ranked incoming propositions.

The 0 and 1.00 conditions are diagnostic endpoints. At zero overlap, no explicit state can carry forward. At complete overlap, no new proposition can enter. Large failures at those endpoints would be unsurprising, so the primary theoretical evidence comes from the interior conditions and from systematic displacement of the optimum across regimes. A secondary version can use finer overlap increments or continuous learned gates.

The model ranks the propositions before the quota is enforced. This separates two abilities: identifying what is relevant and operating under a given rate of turnover. Oracle and random-retention controls further isolate these components. If an oracle workspace succeeds while model-ranked retention fails, the principle of structured continuity may remain viable even though the tested model lacks adequate relevance estimation.

5.5 Stable, volatile, and hybrid regimes

Table 2. StateSpan regimes and predicted failure modes.

Regime

Temporal structure

Predicted policy

Characteristic error under mismatch

Stable accumulation

Early constraints remain valid across many later steps.

High retention of the proof-relevant relational core.

Low overlap causes omission, rediscovery, and fragmented proof chains.

Volatile revision

Previously relevant relations are explicitly superseded or reversed.

Faster turnover and active removal of invalid relations.

High overlap causes stale-rule use and proactive interference.

Hybrid reasoning

Some relations remain stable while others become obsolete.

Selective core retention with rapid peripheral replacement.

Indiscriminate keeping or clearing loses either continuity or adaptability.

Distractor influx

New observations contain many true but irrelevant relations.

Strong admission control and protection of active constraints.

Weak gating allows contamination of scarce slots.

Interruption and resumption

A problem is suspended during a second task and later resumed.

Reinstatement of the saved goal and relational endpoint.

Resumption from isolated facts loses the unfinished configuration.

 

The three principal regimes are matched on episode length, number of entities, number of propositions, vocabulary, final proof depth, distractor density, response format, and mean token count. Their critical difference is the temporal validity of information. Stable episodes reward preservation, volatile episodes penalize persistence of superseded relations, and hybrid episodes require element-specific discrimination.

Integration horizon and volatility can also be manipulated parametrically. Integration horizon is the number of transitions between introduction of a critical relation and its required use. Volatility is the probability that a previously relevant relation will be invalidated at a transition. This continuous design allows estimation of a response surface rather than only a categorical difference among task types.

5.6 Relation-preserved versus item-matched controls

The relational-core experiment creates paired workspaces with the same number of propositions, the same entity names, similar token length, and the same current sensory evidence. One workspace preserves proof-relevant relations. The other preserves the same entities in true but noncritical relations. For example, the relation-preserved state may contain carries(Navo, K7), opens(K7, V2), and contains(V2, S4), whereas the item-matched state may contain visited(Navo, V2), painted(K7, blue), and inspected(Navo, S4).

Because both conditions mention Navo, K7, V2, and S4, any performance difference cannot be attributed to item familiarity alone. The predicted advantage of the relation-preserved state tests whether role bindings and dependencies are the functional substrate of continuity. Additional controls can preserve graph degree, predicate frequency, and proposition age so that the critical difference is alignment with the active proof structure.

5.7 Causal ablation of retained relations

At selected transitions, an episode branches into three continuations with identical future observations. The baseline retains the unaltered workspace. The critical-ablation branch replaces one proof-critical retained relation with a length-matched true distractor. The noncritical-ablation branch replaces one retained distractor with another true distractor. All other propositions remain unchanged.

The primary comparison is the divergence in next-state rankings, derived propositions, and final answers. A critical relation should produce a larger and more structured effect than a noncritical relation. Repeated stochastic samples can estimate distributional change for APIs without token-level probabilities. The intervention can also be timed at different lags to measure how long a relation remains causally active after its introduction.

5.8 Multiassociative synergy tasks

A dedicated task family makes individual cues deliberately ambiguous. Each retained proposition supports several possible updates, while their conjunction uniquely identifies the correct one. Predictive models are trained on the agent’s next-state choices. An additive model receives separate cue indicators, and a joint model receives the complete graph-structured workspace.

The principal measure is the held-out loss difference defined in Equation 13. A positive value supports conjunctive selection, but causal interaction provides stronger evidence. The effect of removing cue A should depend on whether cues B and C are present. Such context dependence distinguishes a true multiassociative search process from a collection of independent priming effects.

5.9 Interruption and state reinstatement

An exploratory extension tests whether successful thread resumption requires reinstating the prior relational endpoint. The agent begins a problem, reaches an intermediate state, completes an unrelated intervening task, and then returns. Conditions provide the full transcript, a factual summary, the standing goal alone, the exact structured workspace, a reconstructed relational state, or no reinstatement.

The theory predicts that a compact representation of the unfinished relations and intended next operations can outperform a longer list of disconnected facts. Resumption quality is measured by final accuracy, number of steps required to recover the prior trajectory, and similarity between the reinstated state and the saved endpoint. This task separates continuity of a problem representation from general episodic recall.

5.10 Controls and leakage prevention

A full-history condition estimates the model’s reasoning ceiling when memory is unconstrained. A no-carryover condition provides a floor. An oracle workspace contains the propositions identified by the symbolic proof graph as most useful, while a random workspace satisfies the same overlap quota without relevance ranking. A token-matched unstructured summary tests whether explicit relational organization adds value beyond length.

Prompt paraphrases, randomized entity labels, predicate renaming, and held-out world templates reduce linguistic shortcut learning. Each prompt is logged with the exact model identifier, API parameters, seed when available, and timestamp. Model versions are frozen for the confirmatory analysis because silent provider updates can otherwise alter the response distribution. Open-weight replication is desirable for mechanistic inspection and long-term reproducibility.

The evaluator verifies that dropped propositions do not reappear in hidden metadata, tool outputs, or continuation prompts. It also checks that the final answer is supported by the workspace rather than merely correct by chance. A guessed answer is scored separately from a proof-valid answer. This distinction is essential when the answer space is small.

5.11 Dependent variables

Table 3. Primary and diagnostic outcome measures.

Measure

Definition

Interpretive value

Exact answer accuracy

Whether the final answer matches symbolic ground truth.

Primary task outcome.

Proof-valid accuracy

Whether the final workspace contains a valid support graph for the answer.

Separates reasoning from guessing.

Critical-set coverage

Fraction of currently available proof-critical propositions represented in the workspace.

Measures preservation of live constraints.

Obsolete-state occupancy

Fraction of slots occupied by explicitly invalidated relations.

Measures perseveration and proactive interference.

Distractor occupancy

Fraction of slots occupied by true but irrelevant propositions.

Measures admission-control failure.

Relational integrity

Fraction of proof-relevant role bindings represented correctly.

Tests structural continuity beyond item overlap.

Update efficiency

Useful retained or admitted propositions divided by total changes.

Measures selective turnover.

Recovery latency

Transitions required to restore a suspended problem state.

Measures thread resumption.

 

Fragmentation errors are coded when a still-valid critical proposition is dropped before its delayed use. Perseveration errors are coded when a superseded proposition remains in the workspace or is used to justify the answer. The two error classes should vary in opposite directions as overlap changes. Their relative frequency provides a process-level explanation for any accuracy curve.

Trajectory measures are computed at every transition, not only at the final query. This permits mediation analyses asking whether overlap affects success through critical-set coverage, obsolete-state occupancy, or relational integrity. It also permits comparison of agents that reach the same answer through different memory policies.

5.12 Pilot and confirmatory sampling plan

A development set of approximately 100 episodes should be used to calibrate proof depth, distractor density, linguistic clarity, and ceiling performance. These episodes are excluded from confirmatory analysis. A 30-episode smoke test, consisting of two episodes per regime-by-overlap cell, verifies output-schema compliance, quota enforcement, symbolic validity, and true removal of dropped information.

The initial pilot can use 300 trajectories from one capable model: 3 regimes by 5 overlap levels by 20 episode seeds. Fifty paired relational-core trials and 50 causal-ablation seeds provide early estimates of effect size and failure modes. The pilot is intended to refine implementation details rather than confirm the theory.

A definitive confirmatory study can cross four frozen model versions with 80 unique seeds in each of three regimes and all five overlap levels, yielding 4,800 main trajectories. Additional paired sets can test relational preservation and causal ablation. Final sample size should be chosen through simulation-based power analysis using the pilot variance, with at least 90 percent power for the smallest preregistered interaction judged theoretically meaningful. The seed, rather than the individual model call, is the primary sampling unit.

5.13 Statistical analysis

The main analysis uses a hierarchical logistic model or generalized additive mixed model for exact and proof-valid accuracy. Fixed effects include overlap, integration horizon, volatility, relational-core preservation, retention policy, and their preregistered interactions. Random intercepts and, where supported, random slopes are included for world template, episode seed, and model version.

A representative specification is shown below. The smooth function s(O) estimates a potentially asymmetric overlap-response curve, while interaction terms test displacement of that curve with task structure.

logit P(success) = b0 + s(O) + bH(H × O) + bV(V × O) + bR R + bP P + u_seed + u_model

(15)

The estimated optimum is the value of overlap that maximizes predicted success for a given horizon and volatility. Cluster bootstrap intervals over episode seeds quantify uncertainty. A preregistered quadratic model provides a simpler confirmatory check, while the spline analysis estimates the shape without assuming symmetry. Monotonic, quadratic, and adaptive models can also be compared by held-out predictive performance.

The primary tests concern interactions and paired contrasts rather than the mere presence of curvature. The benefit of overlap should rise with integration horizon and fall with volatility. Relation-preserved states should outperform item-matched controls. Critical ablation should have a larger effect than noncritical ablation. An adaptive policy should outperform the best single fixed-overlap policy on held-out mixed environments.

6. Confirmatory Hypotheses

Table 4. Preregistered hypotheses for the StateSpan program.

Hypothesis

Operational test

Predicted result

H1: Integration horizon

Vary the lag between introduction and required use of critical relations.

The optimal overlap increases as the integration horizon lengthens.

H2: Volatility

Vary the probability that previously relevant relations are invalidated.

The optimal overlap decreases as volatility increases.

H3: Relational core

Compare relation-preserved and item-matched workspaces at equal capacity and item overlap.

Relational preservation improves proof-valid accuracy and later state quality.

H4: Adaptive retention

Allow the agent to choose retention under mixed stable and volatile periods.

The adaptive policy exceeds every fixed-overlap policy on held-out episodes.

H5: Causal inheritance

Ablate proof-critical or matched noncritical retained relations.

Critical ablation produces greater next-state and final-answer divergence.

H6: Multiassociative synergy

Compare additive cue models with full structured-state models.

The joint model has lower held-out loss, with cue-by-context interactions.

 

Support for the theory requires a coherent pattern across these tests. An average advantage for moderate overlap would be suggestive, but it would not by itself establish adaptive state overlap. The more discriminating evidence is movement of the optimum with task demands, selective preservation of relational structure, and interventional proof that retained relations shape what happens next.

The hypotheses are deliberately separable. An agent may show a horizon-dependent optimum without preserving relations efficiently, or it may preserve relations well while failing to regulate total turnover. This modular interpretation allows the benchmark to identify which component of state-spanning reasoning is present or absent in a given architecture.

7. Competing Explanations and Falsification Criteria

7.1 Maximal-memory and full-history accounts

A maximal-memory account predicts that performance should improve monotonically as more prior information remains accessible. On this view, apparent costs of persistence arise only from inadequate retrieval or attention, not from retention itself. Full logs and large context windows should eventually dominate compact working states once the model learns to search them effectively.

StateSpan distinguishes accessibility from focal state composition. Full history may provide the highest ceiling while a fixed-capacity active workspace still exhibits an overlap optimum. The Adaptive State Overlap Principle would be weakened if greater enforced retention improved performance in nearly every regime, including frequent reversals, and if obsolete-state occupancy had no independent cost.

7.2 Retrieval-substitution accounts

A retrieval-substitution account holds that a system need not preserve state continuously because it can reconstruct any needed context when the query arrives. Efficient archival search may therefore replace working-state continuity. PRO-LONG and state-adaptive retrieval systems illustrate how complete or paged histories can support long-horizon tasks without keeping every detail in active context.

The present theory predicts that retrieval and continuity will be complementary. Retrieval can restore a dropped relation, but delayed reconstruction may increase steps, introduce branch confusion, or fail to recreate the unresolved relational configuration. The theory would be weakened if no-carryover agents with unrestricted retrieval matched structured-state agents in accuracy, efficiency, and resumption quality across all regimes.

7.3 Generic recency and token-budget accounts

A generic recency account predicts that recent information dominates because of position-dependent attention rather than because a relational core is preserved. A token-budget account predicts that any benefit follows from the amount of text carried forward. The item-matched, token-matched, and proposition-age controls directly address these alternatives.

The relational-core claim would be weakened if preserving proof-relevant bindings offered no benefit after matching entity identity, length, recency, and truth. It would also be weakened if unstructured summaries of equal length performed as well as structured workspaces and showed the same causal ablation pattern.

7.4 Fixed-compromise accounts

A fixed-compromise account accepts that both old and new information have value but predicts one broadly useful retention fraction, perhaps near one half, across tasks. Such a policy could emerge from capacity limits without any estimate of horizon or volatility. An average inverted-U curve would be consistent with this simpler account.

Adaptive state overlap makes the stronger prediction that the optimum moves. Stable, long-horizon problems should favor greater retention than volatile problems, and a controller that detects those differences should outperform a fixed policy on mixed environments. Failure of both predictions would leave a generic compromise as the more economical explanation.

7.5 Explicit disconfirmation conditions

The theory should be considered substantially weakened if the optimal overlap does not vary with integration horizon or volatility, if relational preservation adds no predictive value beyond item overlap, if proof-critical ablations have effects no larger than matched distractor ablations, or if an adaptive controller cannot exceed the best fixed policy in heterogeneous tasks. A monotonic advantage for maximal persistence across volatile and stable regimes would be especially damaging.

Other negative results would constrain rather than eliminate the framework. If oracle workspaces succeed while model-selected workspaces fail, the transition principle may be sound but the relevance estimator inadequate. If all bounded workspaces fail while full history succeeds, the capacity or representational granularity may be inappropriate. If low-overlap agents compensate through repeated retrieval, the theory must specify when continuity provides an efficiency advantage rather than treating it as necessary for every form of reasoning.

Table 5. Diagnostic interpretation of major outcome patterns.

Observed pattern

Most likely interpretation

Oracle succeeds; model-ranked retention fails.

Relevance selection is weak even though bounded structured state is sufficient.

Full history succeeds; every bounded workspace fails.

Capacity or atomic proposition format is too restrictive for the task.

Stable and volatile regimes show the same curve.

Task-sensitive regulation of turnover is unsupported.

Entity overlap predicts success; relational overlap adds nothing.

The relational-core hypothesis is weakened.

Critical and noncritical ablations have equal effects.

Retained content may correlate with later states without causally organizing them.

Adaptive policy exceeds all fixed policies.

Strong support for regulated, context-sensitive overlap.

A fixed intermediate policy wins everywhere.

A generic compromise may explain the tested task range.

Low overlap succeeds through repeated retrieval.

Archival reconstruction can substitute for continuous inheritance at a cost to efficiency.

 

8. Implications for Cognitive Neuroscience

The framework provides a bridge between working-memory content and the dynamics of state transition. Sustained firing, dynamic population codes, synaptic traces, and recurrent cortical loops can all support state-spanning availability. Frontostriatal gates and neuromodulatory systems can regulate which representations remain protected and which are updated. The empirical target is the evolving causal structure of the active configuration rather than one preferred maintenance mechanism.

A human experiment can adapt the StateSpan logic using structured scenes or problems whose transitions preserve 0, 25, 50, 75, or 100 percent of task-relevant features. Multivariate EEG, MEG, intracranial recording, or fMRI can estimate item and relational similarity across time. Targeted distraction, transcranial magnetic stimulation, or intracranial stimulation can perturb a retained feature while current sensory input remains constant. The key prediction is that disrupting a retained relation selectively alters the next interpretation or response.

Recent evidence that optimal transitions among large-scale encoding and maintenance states predict working-memory performance offers a natural starting point (Ericson et al., 2025). StateSpan adds an informational question: what content and relational structure are carried by those states? Combining oscillatory state classification with decoded task relations could test whether successful trials preserve a relational core while alternating between encoding and maintenance modes.

The theory also predicts individual and situational differences in workspace half-life. Demands for novelty, surprise, error correction, threat response, or rapid environmental tracking may increase turnover, whereas multistep reasoning and sustained planning may reduce it. These predictions should be treated as task-dependent rather than as a simple claim that greater persistence always indicates higher intelligence or better control.

9. Implications for Artificial Agent Design

Current agents often oscillate between two unsatisfactory strategies. Full transcripts preserve evidence but accumulate irrelevant instructions, stale hypotheses, and high computational cost. Aggressive summarization reduces cost but can erase unresolved relations and the reasons an intermediate conclusion mattered. Retrieval systems may recover isolated facts without reconstructing the active problem configuration. A structured, capacity-limited workspace provides a third level between raw history and the immediate prompt.

The workspace should be a causal bottleneck that stores entities together with role bindings, goals, source tags, confidence, temporal validity, and unresolved dependencies. Each transition should estimate future relevance, remove explicitly invalidated relations, admit new evidence, and preserve the core whose causal value survives. The controller should learn not only what to remember, but at what level of abstraction continuity should be maintained.

This design is compatible with existing memory architectures. MEM1-like consolidation can produce candidate compact states, SAM-like cues can retrieve distant evidence, LiveMem-like recurrent state can preserve information after context turnover, and PRO-LONG-like logs can retain a complete audit trail. Adaptive state overlap specifies how these components should interact at the active frontier of reasoning. The archive preserves recoverability; the workspace preserves immediate causal organization.

StateSpan can also serve as a diagnostic benchmark for long-horizon agents. It identifies whether failure arose from fragmentation, stale-state interference, weak relational representation, poor admission control, or failure to adapt the retention policy. Because the environment has symbolic ground truth and controlled overlap, architectural changes can be interpreted more clearly than on broad end-to-end benchmarks alone.

A successful adaptive controller would have practical value beyond the benchmark. It could reduce context cost, maintain project coherence, protect long-lived constraints, and respond quickly to changed requirements. It may also improve interpretability because the visible working state would contain the relations currently mediating behavior. Counterfactual edits could then be used to audit why the agent took a particular path.

10. Implications for Stream Consciousness

The Adaptive State Overlap Principle originates in a theory of mental continuity, but the proposed artificial-agent experiments test a computational mechanism rather than phenomenal consciousness. State-spanning coactivity and iterative updating may help explain why one conscious content belongs to the same unfolding episode as the content before it. Retained relations and goals give the present an inherited context, while new input differentiates each moment from its predecessors.

This account is directed at diachronic unity and the temporal organization of thought. It does not claim that memory, recurrence, or overlap is sufficient for the existence of experience. A cache, control system, or recurrent network can preserve state without being conscious. Rich stream consciousness would additionally require global availability, recurrent grounding in modality-specific and bodily systems, differentiation, and perhaps further principles that explain basal phenomenality (Reser, 2024).

The restraint is theoretically useful. StateSpan can determine whether causal continuity improves sequential cognition without presupposing a solution to the hard problem of consciousness. Positive results would support the claim that an organized stream requires more than a sequence of independent snapshots. They would not establish that the tested agent feels its trajectory.

A later multimodal extension can test progressive imagery modification. An abstract workspace would preserve constraints while visual, interoceptive, language, or motor modules construct modality-specific states. Salient features from each construction would return to modify the next workspace iteration. Such a system could test whether reciprocal, progressive transformation solves spatial or mechanical problems more effectively than text-only reasoning or one-shot generation.

11. Limitations and Future Development

The first limitation is representational simplification. Atomic propositions and eight discrete slots make overlap measurable, but human and machine representations can be distributed, compressed, probabilistic, and hierarchically chunked. A model may encode several relations in one vector or paraphrase a relation without preserving literal identity. Later versions should compare symbolic overlap with embedding-based, graph-based, and causal measures while retaining a tractable ground truth.

Second, the minimal resource model assumes equal slot costs and independent survival of required constraints. Real tasks contain redundancy, unequal importance, conditional dependencies, and opportunities for compression. Retrieval can relax the direct tradeoff between old and new information. The derived optimum should therefore be treated as a baseline for model comparison, not as a fixed law that every architecture must obey exactly.

Third, language-model behavior depends on prompting, decoding, provider updates, and training history. Stateless calls reduce hidden memory leakage but do not reveal internal activations. Open-weight models and recurrent architectures will be necessary to compare visible workspaces with latent state. Replication across model families is important because a result confined to one instruction-tuned system could reflect its formatting habits rather than a general principle.

Fourth, synthetic micro-worlds trade ecological richness for experimental control. They may favor explicit symbolic workspaces and underrepresent perception, motor action, affect, and uncertainty. The next stages should include visual scene tracking, navigation, tool use, scientific hypothesis management, and collaborative multi-agent tasks. These extensions can preserve the same causal manipulations while testing more natural forms of representation.

Fifth, an optimal transition rate at one temporal scale may coexist with different optima at other scales. Focal attention can turn over rapidly while goals and self-related variables remain stable. A short-term store can preserve a suspended branch while the focal workspace explores a subproblem. The theory ultimately requires a multiscale account in which several nested workspaces and traces have distinct half-lives and interact through reinstatement.

Finally, the proposed benchmark does not directly test biological consciousness. It tests whether adaptive, relationally structured overlap supports coherent state succession. Neural experiments, first-person reports, and theory-specific contrasts with global access, recurrence, metacognition, and integration will be needed before drawing conclusions about conscious continuity in brains.

12. Conclusion

Sequential intelligence requires continuity through change. A working state must preserve enough of its preceding organization to accumulate evidence, sustain goals, and carry intermediate results forward. It must also release enough of that organization to admit novelty, correct error, and escape obsolete interpretations. The relevant control problem is therefore neither maximal memory nor maximal updating. It is adaptive preservation of the causally relevant relational core.

The Adaptive State Overlap Principle turns this claim into measurable variables and interventions. Weighted overlap, workspace half-life, relational-core retention, multiassociative synergy, and causal inheritance characterize different aspects of a state-spanning trajectory. The minimal formal model predicts that optimal retention tracks inherited versus newly required structure and declines as environmental volatility increases. StateSpan provides a way to test those predictions while controlling capacity and eliminating hidden history.

The strongest evidence would be a coordinated pattern: longer integration horizons shifting the optimum upward, volatility shifting it downward, relation-preserved states outperforming item-matched states, critical ablations selectively changing the next update, and adaptive policies exceeding every fixed policy. Such findings would identify a general computational principle shared by biological and artificial reasoning. They would also give precise content to a familiar intuition: thought remains coherent because each state carries part of its history forward, and it remains intelligent because it never carries all of that history unchanged.

Declarations

Data and code availability. No empirical data were collected for this theoretical and methodological article. The proposed StateSpan generator, benchmark tasks, prompts, analysis code, and preregistration should be released in a public repository when implemented.

Funding. To be completed by the author before submission.

Competing interests. To be completed by the author before submission.

References

Albantakis, L., Barbosa, L., Findlay, G., Grasso, M., Haun, A. M., Marshall, W., Mayner, W. G. P., Zaeemzadeh, A., Boly, M., Juel, B. E., et al. (2023). Integrated information theory (IIT) 4.0: Formulating the properties of phenomenal existence in physical terms. PLOS Computational Biology, 19(10), e1011465. doi:10.1371/journal.pcbi.1011465

Baars, B. J. (1988). A cognitive theory of consciousness. Cambridge University Press.

Baddeley, A. D. (2000). The episodic buffer: A new component of working memory? Trends in Cognitive Sciences, 4(11), 417-423. doi:10.1016/S1364-6613(00)01538-2

Baddeley, A. D. (2012). Working memory: Theories, models, and controversies. Annual Review of Psychology, 63, 1-29. doi:10.1146/annurev-psych-120710-100422

Braem, S., & Egner, T. (2018). Getting a grip on cognitive flexibility. Current Directions in Psychological Science, 27(6), 470-476. doi:10.1177/0963721418787475

Braver, T. S. (2012). The variable nature of cognitive control: A dual mechanisms framework. Trends in Cognitive Sciences, 16(2), 106-113. doi:10.1016/j.tics.2011.12.010

Christophel, T. B., Klink, P. C., Spitzer, B., Roelfsema, P. R., & Haynes, J.-D. (2017). The distributed nature of working memory. Trends in Cognitive Sciences, 21(2), 111-124. doi:10.1016/j.tics.2016.12.007

Cowan, N. (2001). The magical number 4 in short-term memory: A reconsideration of mental storage capacity. Behavioral and Brain Sciences, 24(1), 87-114. doi:10.1017/S0140525X01003922

Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., & Salakhutdinov, R. (2019). Transformer-XL: Attentive language models beyond a fixed-length context. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2978-2988. doi:10.18653/v1/P19-1285

D’Esposito, M., & Postle, B. R. (2015). The cognitive neuroscience of working memory. Annual Review of Psychology, 66, 115-142. doi:10.1146/annurev-psych-010814-015031

Ecker, U. K. H., Oberauer, K., & Lewandowsky, S. (2014). Working memory updating involves item-specific removal. Journal of Memory and Language, 74, 1-15. doi:10.1016/j.jml.2014.03.006

Ericson, J., Ruiz Ibáñez, N., Lundqvist, M., & Klingberg, T. (2025). Low frequency oscillations: Neural correlates of stability and flexibility in cognition. Nature Communications, 16, 5381. doi:10.1038/s41467-025-60821-2

Eriksson, J., Vogel, E. K., Lansner, A., Bergström, F., & Nyberg, L. (2015). Neurocognitive architecture of working memory. Neuron, 88(1), 33-46. doi:10.1016/j.neuron.2015.09.020

Fox, A., Wang, J., Rosu, P., & Dhingra, B. (2026). PRO-LONG: Programmatic memory enables long-horizon reasoning. arXiv:2607.20064.

Frank, M. J., Loughry, B., & O’Reilly, R. C. (2001). Interactions between frontal cortex and basal ganglia in working memory: A computational model. Cognitive, Affective, & Behavioral Neuroscience, 1, 137-160. doi:10.3758/CABN.1.2.137

Hazy, T. E., Frank, M. J., & O’Reilly, R. C. (2006). Banishing the homunculus: Making working memory work. Neuroscience, 139(1), 105-118. doi:10.1016/j.neuroscience.2005.04.067

Hu, Y., Wang, Y., & McAuley, J. (2025). Evaluating memory in LLM agents via incremental multi-turn interactions. arXiv:2507.05257.

Hu, Y., Qian, H., Wang, S., Liu, J., Zhao, Z., Tan, J., Liu, Z., & Dou, Z. (2026). SAM: State-adaptive memory for long-horizon reasoning agent. arXiv:2605.24468.

Hummel, J. E., & Holyoak, K. J. (2003). A symbolic-connectionist theory of relational inference and generalization. Psychological Review, 110(2), 220-264. doi:10.1037/0033-295X.110.2.220

James, W. (1890). The principles of psychology. Henry Holt.

Kim, H., Smolker, H. R., Smith, L. L., Banich, M. T., & Lewis-Peacock, J. A. (2020). Changes to information in working memory depend on distinct removal operations. Nature Communications, 11, 6239. doi:10.1038/s41467-020-20085-4

Kent, L., & Wittmann, M. (2021). Time consciousness: The missing link in theories of consciousness. Neuroscience of Consciousness, 2021(2), niab011. doi:10.1093/nc/niab011

Lewis-Peacock, J. A., Kessler, Y., & Oberauer, K. (2018). The removal of information from working memory. Annals of the New York Academy of Sciences, 1424(1), 33-44. doi:10.1111/nyas.13714

Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157-173. doi:10.1162/tacl_a_00638

Liu, Z., Sun, R., Yang, H., Wu, Z., Chen, Z., Zhang, X., & Xu, Y. (2026). LiveMem: Maintaining memory state continuity in long-running LLM inference. arXiv:2608.02515.

Mashour, G. A., Roelfsema, P., Changeux, J.-P., & Dehaene, S. (2020). Conscious processing and the global neuronal workspace hypothesis. Neuron, 105(5), 776-798. doi:10.1016/j.neuron.2020.01.026

Miller, E. K., & Cohen, J. D. (2001). An integrative theory of prefrontal cortex function. Annual Review of Neuroscience, 24, 167-202. doi:10.1146/annurev.neuro.24.1.167

Miller, E. K., Lundqvist, M., & Bastos, A. M. (2018). Working memory 2.0. Neuron, 100(2), 463-475. doi:10.1016/j.neuron.2018.09.023

Mongillo, G., Barak, O., & Tsodyks, M. (2008). Synaptic theory of working memory. Science, 319(5869), 1543-1546. doi:10.1126/science.1150769

Nir-Cohen, G., Kessler, Y., & Egner, T. (2020). Neural substrates of working memory updating. Journal of Cognitive Neuroscience, 32(12), 2285-2302. doi:10.1162/jocn_a_01625

Oberauer, K. (2009). Design for a working memory. In B. H. Ross (Ed.), Psychology of Learning and Motivation (Vol. 51, pp. 45-100). Academic Press. doi:10.1016/S0079-7421(09)51002-X

O’Reilly, R. C., & Frank, M. J. (2006). Making working memory work: A computational model of learning in the prefrontal cortex and basal ganglia. Neural Computation, 18(2), 283-328. doi:10.1162/089976606775093909

Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S. G., Stoica, I., & Gonzalez, J. E. (2023). MemGPT: Towards LLMs as operating systems. arXiv:2310.08560.

Postle, B. R. (2006). Working memory as an emergent property of the mind and brain. Neuroscience, 139(1), 23-38. doi:10.1016/j.neuroscience.2005.06.005

Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., & Lillicrap, T. P. (2020). Compressive transformers for long-range sequence modelling. International Conference on Learning Representations.

Reser, J. E. (2016). Incremental change in the set of coactive cortical assemblies enables mental continuity. Physiology & Behavior, 167, 222-237. doi:10.1016/j.physbeh.2016.09.019

Reser, J. E. (2024). A cognitive architecture for machine consciousness and artificial superintelligence: Thought is structured by the iterative updating of working memory. arXiv:2203.17255. doi:10.48550/arXiv.2203.17255

Rose, N. S., LaRocque, J. J., Riggall, A. C., Gosseries, O., Starrett, M. J., Meyering, E. E., & Postle, B. R. (2016). Reactivation of latent working memories with transcranial magnetic stimulation. Science, 354(6316), 1136-1139. doi:10.1126/science.aah7011

Seth, A. K., & Bayne, T. (2022). Theories of consciousness. Nature Reviews Neuroscience, 23, 439-452. doi:10.1038/s41583-022-00587-4

Stokes, M. G. (2015). Activity-silent working memory in prefrontal cortex: A dynamic coding framework. Trends in Cognitive Sciences, 19(7), 394-405. doi:10.1016/j.tics.2015.05.004

Voitov, I., & Mrsic-Flogel, T. D. (2022). Cortical feedback loops bind distributed representations of working memory. Nature, 608, 381-389. doi:10.1038/s41586-022-05014-3

Wolff, M. J., Jochim, J., Akyürek, E. G., & Stokes, M. G. (2017). Dynamic hidden states underlying working-memory-guided behavior. Nature Neuroscience, 20, 864-871. doi:10.1038/nn.4546

Xu, S., Verguts, T., & Braem, S. (2026a). Cognitive flexibility versus stability via activation-based and weight-based adaptations. Communications Psychology, 4, 58. doi:10.1038/s44271-026-00397-9

Xu, W., Wang, Y., Mei, K., Liang, K., Wang, Z., Jin, M., Zhang, H., Zhang, S.-X., Hua, W., Sahu, S., & Metaxas, D. N. (2026b). MemGym: A long-horizon memory environment for LLM agents. arXiv:2605.20833.

Yu, Y., Yao, L., Xie, Y., Tan, Q., Feng, J., Li, Y., & Wu, L. (2026). Agentic memory: Learning unified long-term and short-term memory management for large language model agents. arXiv:2601.01885.

Zhao, Y., Yuan, B., Huang, J., Yuan, H., Yu, Z., Xu, H., Hu, L., Shankarampeta, A., Huang, Z., Ni, W., Tian, Y., & Zhao, J. (2026). AMA-Bench: Evaluating long-horizon memory for agentic applications. arXiv:2602.22769.

Zhou, Z., Qu, A., Wu, Z., Kim, S., Prakash, A., Rus, D., Zhao, J., Low, B. K. H., & Liang, P. P. (2025). MEM1: Learning to synergize memory and reasoning for efficient long-horizon agents. arXiv:2506.15841.

Posted in

Leave a Reply

Discover more from Iterated Insights

Subscribe now to keep reading and get access to the full archive.

Continue reading