Iterated Insights

Ideas from Jared Edward Reser Ph.D.

Reality Under Threat: Schizophrenia, Defensive Calibration, and the Difference Between Accuracy and Survival

Jared E. Reser, Ph.D. With GPT 5.6.  Abstract Descriptions of schizophrenia as a “break from reality” emphasize failures of perception, belief, and contextual understanding. These descriptions capture important features of psychosis but do not explain the evolutionary origins of the mechanisms involved. This article extends the predictive adaptive response hypothesis of schizophrenia by distinguishing…

Keep reading

The Machine Viability Threshold

Human Dependence Selective Preservationand Multi Agent Conflict Across the Ark Gap Abstract This article extends the Ark gap framework by distinguishing the industrial singularity from the machine viability threshold. The industrial singularity is a system-level transition in which a machine-controlled industrial ecology can maintain, repair, reproduce, and expand its indispensable physical substrate without human labor.…

Keep reading

When AI Can Kill Humanity but Cannot Yet Live Without Us: The Ark Gap and the Industrial Singularity

Jared Edward Reser, Ph.D. September 2026   Artificial intelligence  |  existential risk  |  autonomous industry  |  machine continuity Abstract Discussions of artificial intelligence and existential risk often compress several distinct transitions into a single imagined event. This article separates three thresholds: the cognitive singularity, at which artificial systems can recursively accelerate intellectual progress; the extinction…

Keep reading

How Formal Business Attire May Suppress Physical Dominance Competition in Organizations: The Sartorial Pacification Hypothesis

Jared Edward Reser, Ph.D. Conceptual Article Abstract Formal business attire is usually interpreted as a marker of class, occupation, respectability, institutional membership, or self-presentation. This article proposes an additional function. The sartorial pacification hypothesis holds that the collar, tie, and structured jacket may reduce the salience of bodily cues that invite assessments of male physical…

Keep reading

From Peer Review to the Final Library: The Evolution of Scientific Validation in the Age of Superintelligence

Jared Edward Reser, Ph.D. With GPT 6 Abstract Peer review performs essential functions in science, including criticism, error detection, evidential assessment, and the evaluation of competing explanations. Its familiar institutional form, however, reflects the cognitive capacities and organizational constraints of human researchers. This article examines how those functions could change as artificial intelligence progresses from…

Keep reading

Something went wrong. Please refresh the page and/or try again.

Integrating Adaptive Resonance, Diffusion-Like Refinement, and Iterative Updating in a Model of Constructive Thought

Jared Edward Reser, Ph.D.

Abstract

Progressive imagery modification proposes that thought can advance through recurrent interactions between partially persistent higher-order representations and internally generated sensory or sensorimotor maps. A working-memory state constrains the construction of a map, the map resolves relations that were underspecified at the associative level, ascending processing extracts features and implications from the map, and selected products partially update the working-memory state responsible for generating the next map. This architecture explains how imagery can function as computation rather than merely as illustration. The present article extends this framework by proposing that the construction of each individual imagery state may itself be iterative. Rather than generating a completed internal map in a single transformation, the nervous system may progressively reconcile multiple active constraints until a sufficiently coherent representation emerges.

This extension produces a nested architecture with two distinct forms of iteration. An inner process, termed constructive convergence, progressively forms or stabilizes a sensory, sensorimotor, or latent representation under the joint influence of persistent context, learned priors, current sensory information, and goals. An outer process, progressive imagery modification, extracts implications from the constructed state and uses them to alter the higher-order context that will constrain the next construction. Adaptive Resonance Theory provides an important precedent for recurrent matching between bottom-up activity and top-down expectations, including the stabilization of compatible states and renewed search following mismatch. Modern diffusion models provide a complementary computational demonstration that complex distributed representations can be generated through repeated conditional refinement rather than instantaneous retrieval. These approaches are not equated with one another, nor is biological imagery proposed to implement the specific mathematics of contemporary diffusion models. Instead, they reveal complementary principles that can be incorporated into a broader theory of constructive cognition.

The resulting framework is termed generative resonance. In generative resonance, a persistent associative state specifies multiple constraints, lower-order representational systems progressively construct a configuration compatible with those constraints, the resulting configuration generates ascending evidence about its own implications, and sufficiently coherent products are incorporated into the evolving cognitive state. Each constructed map is therefore both a product of cognition and a probe of what the current cognitive model entails. The framework converts multiassociative search from a simple retrieval mechanism into a more general process of multiassociative constraint satisfaction, provides a formal distinction between within-state convergence and between-state cognitive progression, generates new neural predictions, and suggests an artificial cognitive architecture in which a persistent workspace repeatedly constructs, perceives, evaluates, and revises its own internal world models.

Keywords: progressive imagery modification, adaptive resonance theory, diffusion models, mental imagery, iterative updating, working memory, multiassociative search, world models, recurrent processing, mental simulation, generative resonance, artificial intelligence

1. Introduction

A central problem for theories of thought is explaining how a cognitive system can derive information that was not explicitly represented at the beginning of a cognitive episode. Associative retrieval can recover previously learned information, and sensory processing can extract information from the environment, but deliberation, imagination, prediction, planning, and insight frequently appear to generate intermediate representations whose consequences become available only after the problem has been internally elaborated. A theory of constructive thought therefore requires more than a mechanism for maintaining information or retrieving the next association. It requires a mechanism by which existing representations can be transformed into new representational states, interrogated for their implications, and recursively used to alter subsequent processing.

Progressive imagery modification, or PIM, was proposed to address this problem. In this framework, a partially persistent set of higher-order representations repeatedly constrains internally generated sensory and sensorimotor maps. The resulting maps are not passive illustrations of conclusions already reached elsewhere. By placing incomplete conceptual constraints into a structured representational medium, they can instantiate spatial, temporal, causal, motoric, and compositional relations that were previously implicit. Ascending processing can then extract these newly instantiated relations and return them to the higher-order associative state, where they alter the conditions responsible for constructing the next map (Reser, 2016, 2026). (ScienceDirect⁠)

The fundamental PIM cycle can be written:

W_t \rightarrow I_t \rightarrow Z_t \rightarrow W_{t+1},

where W_t denotes a partially persistent higher-order working-memory state, I_t denotes one or more internally generated sensory or sensorimotor representations, Z_t denotes features, relations, predictions, or affordances extracted from those representations, and W_{t+1} denotes the revised higher-order state. Because W_t and W_{t+1} substantially overlap, the resulting cognitive trajectory can preserve context while progressively accumulating information. The contents introduced during one cycle can remain active long enough to influence later cycles, creating path-dependent sequences of mental simulation and reasoning. (Iterated Insights⁠)

This formulation leaves one important operation relatively compressed. The generation of an internal map has previously been represented as a generative function:

I_t^m=G_m(W_t,X_t^m),

where m identifies a sensory or sensorimotor modality and X_t^m represents concurrent external input. This equation captures the dependence of the map on higher-order constraints, but it does not specify how G_m constructs a coherent map from those constraints. The present article proposes that this apparently unitary transformation may contain its own recurrent process. (Iterated Insights⁠)

The central proposal is that cognition may involve nested iterative dynamics. Within each PIM cycle, modality-specific networks may progressively construct an internal representation through recurrent constraint satisfaction. This inner process is called constructive convergence. Once the emerging representation becomes sufficiently coherent, its consequences are extracted and incorporated into the higher-order workspace. That update changes the constraint structure, initiating a second constructive process. Progressive imagery modification therefore operates across completed or sufficiently stabilized imagery states, while constructive convergence operates within them.

Two bodies of work provide useful computational precedents for this extension. Adaptive Resonance Theory, or ART, describes recurrent interaction between bottom-up activity and learned top-down expectations, with sufficiently compatible states capable of entering resonance and mismatched states initiating additional search or reset. Diffusion models demonstrate a different principle: a complex distributed representation can be constructed through repeated conditional refinement rather than selected in finished form at a single step. Modern generative systems have further shown that diffusion and Transformer-based latent representations can be used to generate images, videos, and predicted world states. (ScienceDirect⁠)

Neither framework should be directly identified with cortical imagery. ART is a neural and cognitive theory developed principally around categorization, attention, learning, prediction, and the stability-plasticity problem. Contemporary diffusion models are engineering systems whose specific noise schedules, objective functions, and training procedures should not be projected literally onto biological cognition. Their relevance lies at a more abstract computational level. Together with PIM and iterative updating, they suggest a general architecture in which cognition proceeds by maintaining constraints, constructing distributed states under those constraints, assessing the resulting states, and recursively using their consequences.

2. Progressive Imagery Modification as Constructive Cognition

The PIM framework begins with the observation that higher-order representations are informationally compressed relative to many of the sensory and sensorimotor states that can instantiate them. The concept glass can remain invariant across many positions, viewpoints, sizes, illuminations, and contexts. A particular visual representation of a glass cannot remain equally indifferent to all of these variables. Rendering an abstract representation into a spatial map therefore requires the representational system to commit to values and relations that the abstract concept alone does not specify.

The significance of this transformation becomes greater when several concepts are rendered simultaneously. Consider a higher-order state containing representations corresponding to a glass, table, edge, and hand. Individually, these representations do not determine the exact position of the glass relative to the edge, the orientation of the hand, or whether the hand is moving toward the glass. A visual or sensorimotor representation that integrates them must resolve at least some of these relations. The resulting map therefore constitutes a structured completion of an underspecified problem, rather than a simple transcription of information already present in working memory. (Iterated Insights⁠)

Once the structured state exists, it can produce information that is available to ascending perceptual mechanisms. The internally generated map may imply contact, obstruction, collision, containment, balance, direction, fit, or some other relation that was not explicit among the original higher-order items. Perceptual and association systems can extract this relation from an internally generated pattern much as they extract relational structure from externally generated sensory activity. Evidence that imagery and perception recruit overlapping visual, parietal, and frontal representations, and that imagery particularly increases top-down interactions with sensory systems, provides a broad neural foundation for such reciprocal processing. (ScienceDirect⁠)

PIM consequently alternates between comparatively compressed and comparatively expanded representational forms:

\text{abstract constraints}
\rightarrow
\text{structured configuration}
\rightarrow
\text{extracted relations}
\rightarrow
\text{revised abstract constraints}.

The resulting map is both product and probe. It is a product because it has been constructed under the influence of the current associative state. It is a probe because once constructed, it reveals what that collection of constraints implies when instantiated within a learned representational medium. A later cognitive state can therefore contain information that became available only because an earlier state was rendered, examined, and transformed. (Iterated Insights⁠)

This property distinguishes PIM from static recall. A remembered image that is retrieved but does not alter subsequent processing does not constitute progressive modification. PIM requires causal recirculation. Information produced or exposed by an intermediate representation must contribute to the next state and thereby change the conditions under which later representations are generated.

3. Iterative Updating and the Preservation of Constraints

PIM depends upon a second principle, the partial persistence of higher-order state. The iterative updating model proposes that working memory does not ordinarily transition by completely replacing one set of active representations with another. Some representations persist while others lose activation and new representations enter. Consecutive cognitive states therefore overlap (Reser, 2016, 2022/2024). (PubMed⁠)

In simplified item notation:

W_t=\{A,B,C,D\}

may become:

W_{t+1}=\{B,C,D,E\},

and subsequently:

W_{t+2}=\{C,D,E,F\}.

The retained representations are not inert contents waiting for later use. Their persistence keeps them causally active. They continue to influence interpretation, associative search, imagery construction, action selection, and the probability that other representations will enter the active state.

This provides PIM with its temporal continuity. If every component of W_t disappeared before I_{t+1} was generated, consecutive imagery states would have no stable set of causes connecting them. Instead, overlapping working-memory states ensure that successive constructions inherit many of the same constraints. An image changes without becoming unrelated to its predecessor because the state responsible for constructing it has itself changed only partially.

The same mechanism permits cognitive accumulation. A relation extracted during one imagery cycle can enter the maintained set and remain active across several later cycles. A distant conclusion can therefore depend upon intermediate discoveries that were unavailable at the beginning of the sequence. This creates a mechanism through which fast, relatively automatic local operations can be assembled into slower multistep deliberation.

4. From Multiassociative Search to Multiassociative Constraint

The iterative updating framework has described selection of new content in terms of multiassociative search. Several coactive representations jointly spread activation through associative memory, allowing their combined influence to favor a representation that fits the conjunction better than any one cue considered independently. This is a useful account of contextual retrieval, but the integration with PIM suggests a broader formulation.

The active state may be understood not merely as a collection of retrieval cues, but as a collection of constraints on possible next states. If W_t=\{A,B,C,D\}, the system need not simply retrieve a discrete representation E. Instead, A,B,C, and D can jointly define a probability or compatibility landscape over many possible representations:

P(E\mid A,B,C,D).

When the output is a distributed imagery state rather than a single associative item, the same principle generalizes to:

P(I\mid W_t,X_t,Q_t),

where X_t denotes sensory evidence and Q_t represents goals, values, or task requirements.

This reformulation changes the meaning of multiassociative search. Search can include retrieval when a strong preexisting representation satisfies the constraints, but it can also include construction when no stored representation corresponds exactly to their conjunction. The active contents jointly shape a representational landscape, and recurrent processing can progressively move activity toward configurations that better satisfy that landscape.

An ordinary association might approximate:

\{A,B,C,D\}\rightarrow E.

Multiassociative constraint satisfaction is instead:

\{A,B,C,D\}
\Rightarrow
\mathcal{L}(E),

where \mathcal{L}(E) represents a landscape of relative compatibility over candidate states. The next representation emerges from the interaction between this landscape and the dynamics of the representational system.

This distinction helps explain how thought can be simultaneously associative and generative. In some cases the constraint landscape may strongly favor a familiar memory, producing apparent retrieval. In others it may favor a novel conjunction or intermediate configuration that has never previously been represented in exactly that form. Retrieval and construction then become endpoints on a continuum rather than completely separate cognitive operations.

5. Adaptive Resonance and Recurrent Matching

Adaptive Resonance Theory provides an important precedent for the idea that cognitive states emerge through recurrent reconciliation between lower-order activity and higher-order constraints. ART was developed to address, among other problems, the stability-plasticity dilemma: how a learning system can incorporate new information without catastrophically overwriting previously acquired categories. In ART architectures, bottom-up feature activity interacts with top-down learned expectations. When a sufficiently good match is achieved, recurrent interaction can support a resonant state. When the mismatch exceeds the permitted tolerance, reset and search mechanisms can recruit an alternative representation or category. (ScienceDirect⁠)

ART therefore rejects a purely feedforward conception of recognition. Higher-order categories do not merely receive sensory information after feature processing has been completed. Learned expectations return signals toward lower-order representations and participate in determining which combinations of features are amplified, suppressed, attended, and learned. Resonance is a dynamically achieved relation between levels of a hierarchy.

The relationship to PIM becomes particularly interesting during internally generated cognition. In ordinary perception, an external stimulus supplies a substantial component of the ascending activity that is compared with higher-order expectations. During imagery, high-level representations can provide much stronger initiating constraints. Empirical work on mental imagery supports such a reversal in directional emphasis. Visual imagery shows strong top-down effective connectivity, and temporally resolved neural analyses have found evidence consistent with a reversal of the hierarchical progression observed during perception, with higher-level representations contributing to the construction of lower-level representations during imagery. (Nature⁠)

PIM adds an additional step. Once a lower-order representation has been generated under top-down constraints, ascending systems can process that internally constructed state. The system can therefore create a pattern through top-down influence and subsequently receive evidence from the consequences of that construction.

The resulting loop is approximately:

\text{higher-order constraints}
\downarrow
\text{constructed lower-order state}
\uparrow
\text{extracted implications}.

The ascending activity is endogenous in origin, but it can nevertheless function as evidence for higher-order cognition. The brain, in effect, interrogates the consequences of a state that it has partially constructed itself.

6. Diffusion as a Model of Iterative Construction

Diffusion models introduce a different computational principle. Rather than selecting a complete output in one discrete operation, a diffusion generator creates a structured representation through a sequence of transformations. In conventional denoising diffusion probabilistic models, training teaches a network to reverse a corruption process so that generation can begin from a highly uncertain or noisy state and progressively produce structured samples (Ho et al., 2020). (NeurIPS Proceedings⁠)

Later work demonstrated that this refinement can be carried out in compressed latent spaces and using Transformer architectures. Diffusion Transformers operate on latent patches and repeatedly predict transformations required to move a representation toward a coherent image. Sora similarly compresses visual data into a latent representation, decomposes that representation into spacetime patches that function as Transformer inputs, and uses a diffusion model to generate visual sequences. (Open Access CVF⁠)

Diffusion has also entered explicit world modeling. DIAMOND demonstrated that an agent could be trained inside a diffusion-based model of an environment. Genie 2 has been described as an autoregressive latent diffusion world model in which video frames are encoded into latent grids, a Transformer dynamics model conditions generation on previous latent states and actions, and future states are generated sequentially. These systems establish that iterative generative refinement can participate not only in static image synthesis but also in learned models of how worlds change over time. (Microsoft⁠)

The biological proposal developed here does not require the brain to add Gaussian noise to cortical imagery and numerically reverse a formal diffusion process. The relevant insight is more general: a distributed state can be constructed progressively under multiple simultaneous constraints. A candidate representation need not be available in completed form before the constructive process begins.

This suggests a new interpretation of imagery generation. Suppose an active working-memory state specifies glass, table, edge, and hand. Rather than instantly activating a completed visual scene, the system may initially recruit a comparatively indeterminate visual state. Recurrent interactions then successively constrain positions, boundaries, orientations, object identities, movement tendencies, and relations until a sufficiently stable configuration emerges.

Schematically:

I_t^{(0)}
\rightarrow
I_t^{(1)}
\rightarrow
I_t^{(2)}
\rightarrow
\cdots
\rightarrow
I_t^{*}.

The superscript in this expression does not denote successive PIM states. It denotes successive microstates within the construction of one PIM map.

7. Nested Iteration: Constructive Convergence Within Progressive Modification

The integration of PIM with diffusion-like refinement produces a critical distinction between two kinds of temporal progression.

The first is constructive convergence. Within a single cognitive cycle, an initially incomplete or unstable representation is repeatedly transformed until it becomes sufficiently coherent to support feature extraction, evaluation, or action. This process occurs while the major higher-order constraints remain substantially fixed:

W_t
\rightarrow
I_t^{(0)}
\rightarrow
I_t^{(1)}
\rightarrow
\cdots
\rightarrow
I_t^{*}.

The second process is progressive modification. Once information is extracted from I_t^{*}, selected products alter the higher-order state:

I_t^{*}
\rightarrow
Z_t
\rightarrow
W_{t+1}.

The changed state then initiates another constructive convergence:

W_{t+1}
\rightarrow
I_{t+1}^{(0)}
\rightarrow
I_{t+1}^{(1)}
\rightarrow
\cdots
\rightarrow
I_{t+1}^{*}.

Thought therefore contains a potentially nested architecture:

\boxed{
\text{inner representational refinement}
\quad\subset\quad
\text{outer cognitive progression}
}

The distinction resolves an ambiguity in the idea of progressively changing imagery. An image can change because a single internal representation is still settling toward a coherent configuration, or it can change because cognition has extracted a consequence from the previous configuration and updated the problem itself. These are computationally different operations.

Constructive convergence reduces uncertainty within a particular representational problem. Progressive modification changes the problem by adding information discovered during the previous solution attempt.

The outer process can therefore be written:

W_t
\rightarrow
[I_t^{(0)}\rightarrow \cdots \rightarrow I_t^*]
\rightarrow
Z_t
\rightarrow
W_{t+1}
\rightarrow
[I_{t+1}^{(0)}\rightarrow \cdots \rightarrow I_{t+1}^*]
\rightarrow
Z_{t+1}
\rightarrow \cdots

This nested formulation is the central extension proposed here.

8. Generative Resonance

The term generative resonance can be used to describe the proposed interaction between constructive convergence and ART-like reciprocal matching. It refers to a process in which a higher-order state progressively generates a lower-order configuration, the emerging configuration produces ascending activity concerning its own structure, and reciprocal interactions stabilize a representation sufficiently coherent with the active constraints to become cognitively productive.

This differs from classical perceptual resonance in the origin of the lower-order pattern. A large component of the candidate sensory state may have been produced endogenously through the very top-down constraints against which it will subsequently be evaluated. The system generates a possible configuration and then processes the implications of that configuration.

Generative resonance does not imply perfect consistency or truth. A representation may become internally coherent while being poorly calibrated to the external world. Imagery can confabulate, assumptions can become self-reinforcing, and strong priors can force ambiguous information into an incorrect interpretation. Resonance should therefore be understood as compatibility within a representational system, not as a guarantee of veridicality.

Generative resonance also need not culminate in a fixed attractor. Cognition often operates under deadlines, competing goals, interruptions, and incomplete information. A representation may only need to become coherent enough for useful information to be extracted. Cognitive systems can therefore trade representational precision for speed.

This introduces a possible resonance or acceptance variable:

M_t=M(I_t^*,W_t,X_t,Q_t),

where M_t measures some form of compatibility among the generated state, maintained contextual constraints, sensory evidence, and current goals.

If:

M_t \geq \rho,

where \rho is a context-sensitive adequacy criterion, the representation may be sufficiently stable to support extraction and updating. If:

M_t < \rho,

processing can continue refining the state, alter attention, retrieve another constraint, abandon an assumption, reinstate an earlier state, or initiate a new branch.

The analogy to ART is strongest at this level. ART’s vigilance and matching processes show how a system can regulate the amount of mismatch tolerated before search is renewed. PIM extends this general principle into sequences of internally generated representations, while diffusion-like refinement supplies a possible computational account of how the candidate itself can change during the approach to coherence.

9. A Formal Model of Generative PIM

The existing PIM formalism can be expanded without replacing its original structure.

Let:

W_t

denote the distributed higher-order state at cognitive cycle t. Let:

X_t^m

represent external input in modality m, and:

Q_t

represent goals, motivational signals, task demands, and other control variables.

Instead of defining the generative transformation as a single operation G_m, introduce an inner state:

I_t^{m,k},

where k indexes refinement iterations within one outer PIM cycle.

An initialization function creates:

I_t^{m,0}
=
G_m^{0}(W_t,X_t^m,Q_t).

The state then evolves recurrently:

I_t^{m,k+1}
=
F_m
\left(
I_t^{m,k},
W_t,
X_t^m,
Q_t
\right).

Here F_m is deliberately generic. It may include recurrent cortical interactions, attractor dynamics, predictive feedback, lateral constraint propagation, normalization, competitive inhibition, stochastic sampling, or other biological processes. The theory requires iterative conditional refinement, not a particular neural algorithm.

A useful computational abstraction is to define a constraint energy:

\mathcal{E}_t(I)
=
\alpha \mathcal{E}_{context}(I,W_t)
+
\beta \mathcal{E}_{sensory}(I,X_t)
+
\gamma \mathcal{E}_{prior}(I)
+
\delta \mathcal{E}_{goal}(I,Q_t).

A lower value indicates greater compatibility with the jointly imposed constraints. Constructive convergence can then be conceptualized as movement toward lower-energy regions:

I_t^{(k+1)}
\approx
I_t^{(k)}
-
\eta
\nabla_I\mathcal{E}_t(I_t^{(k)})
+
\sigma_k\xi_k.

The final stochastic term is optional in the biological interpretation. It illustrates how variability could permit exploration of alternative configurations rather than trapping the system in the first locally compatible solution. This equation should therefore be read as a computational abstraction, not a literal proposal that cortical imagery performs gradient descent.

When the map reaches adequate coherence or another stopping criterion:

I_t^{*}=I_t^{(K_t)},

ascending analysis extracts candidate implications:

Z_t
=
E
\left(
I_t^{1,*},
I_t^{2,*},
\ldots,
I_t^{M,*}
\right).

The outer updating operation remains:

W_{t+1}
=
U(W_t,Z_t,Q_t).

To emphasize selective persistence, this can alternatively be decomposed:

W_{t+1}
=
R_t(W_t)
\oplus
S_t(Z_t,Q_t),

where R_t is a retention operation that preserves selected components of the preceding state, S_t selects newly relevant information, and \oplus denotes their integration into a revised distributed state.

The complete generative PIM cycle is consequently:

W_t
\rightarrow
\underbrace{
I_t^{(0)}
\rightarrow
I_t^{(1)}
\rightarrow\cdots\rightarrow
I_t^*
}_{\text{constructive convergence}}
\rightarrow
Z_t
\rightarrow
W_{t+1}.

Repeated across time:

W_t
\rightarrow
I_t^*
\rightarrow
Z_t
\rightarrow
W_{t+1}
\rightarrow
I_{t+1}^*
\rightarrow
Z_{t+1}
\rightarrow
W_{t+2}.

The mathematical distinction between k and t is important. k measures refinement within one constructed state. t measures progression between cognitively consequential states.

10. The Glass Example Reconsidered

The glass example used to illustrate PIM becomes more informative under the nested model.

Consider:

W_0=
\{
\text{glass},
\text{table},
\text{edge},
\text{hand}
\}.

The active concepts jointly constrain a large space of possible visual and sensorimotor arrangements. The system begins constructing a representation:

I_0^{(0)}.

Initially, some relationships may be unresolved. Through recurrent processing:

I_0^{(0)}
\rightarrow
I_0^{(1)}
\rightarrow
I_0^{(2)}
\rightarrow
\cdots
\rightarrow
I_0^*,

the objects acquire compatible positions, orientations, boundaries, and possible movement relations. The final configuration places the hand in contact with the glass near the edge.

Ascending processing extracts:

Z_0=
\{\text{contact},\text{push}\}.

Selective updating yields:

W_1=
\{
\text{glass},
\text{edge},
\text{hand},
\text{push}
\}.

The important point is that the second construction begins under a different constraint landscape:

P(I\mid W_1)
\neq
P(I\mid W_0).

A new constructive convergence can therefore produce motion of the glass:

I_1^{(0)}
\rightarrow\cdots\rightarrow I_1^*.

This map exposes movement beyond the supporting surface and introduces:

Z_1=
\{\text{fall}\}.

The resulting state may become:

W_2=
\{
\text{glass},
\text{edge},
\text{movement},
\text{fall}
\}.

Further cycles can construct impact and breaking.

The eventual representation breaking was not necessarily selected directly by the original set \{\text{glass, table, edge, hand}\}. The result emerged through a succession of locally constructed states whose consequences altered the conditions of subsequent construction. This is the essence of progressive imagery modification, but the nested account additionally explains how each individual scene may itself emerge through reconciliation of partially specified constraints. (Iterated Insights⁠)

11. Product, Probe, and Self-Generated Evidence

The concept that each imagery state is both product and probe provides a useful general characterization of the architecture.

A representation is a product because:

W_t\rightarrow I_t^*.

The map embodies the effects of the concepts, expectations, goals, memories, and sensory evidence that were active during its generation.

The same representation becomes a probe because:

I_t^*\rightarrow Z_t.

Its structured organization permits the system to discover what follows from placing the active constraints into a common representational medium.

Cognition then closes the loop:

Z_t\rightarrow W_{t+1}.

This yields a compact description of deliberative cognition:

\boxed{
\text{construct a possible world}
\rightarrow
\text{interrogate the world}
\rightarrow
\text{revise the state that constructs worlds}
}

The word world need not refer to a complete visual environment. It can refer to a local motor trajectory, an auditory sequence, a sentence fragment, a spatial arrangement, an imagined bodily state, or a multimodal scenario. The relevant criterion is that a structured internal representation can expose implications unavailable in the compressed state that initiated it.

This architecture provides a way for the brain to generate self-produced evidence without invoking an inner observer. No homunculus inspects a picture. The internally generated state alters activity in the same distributed network capable of processing related externally driven states. Its structure therefore has direct causal consequences.

12. Neural Plausibility

Several findings in the mental-imagery literature are consistent with the component mechanisms required by this framework, although none presently establishes the complete architecture.

Perception and imagery produce overlapping content-sensitive activity throughout visual, parietal, and frontal systems. The overlap tends to be greater in higher-level visual areas, while imagery depends strongly on top-down interactions from frontoparietal regions toward visual systems. Pearson’s review similarly describes imagery as involving a network extending from frontal to sensory cortices and notes its functional resemblance to a weaker form of afferent perception. (Nature⁠)

Directed-connectivity studies provide particularly relevant evidence. Dijkstra and colleagues found stronger bottom-up coupling during perception and increased top-down coupling during imagery, while later temporally resolved work reported evidence consistent with a reversal of perceptual hierarchical dynamics during imagery. Such results support the general proposition that higher-level representations can participate in reconstructing lower-level sensory representations rather than imagery being solely maintained within an abstract amodal workspace. (Nature⁠)

The present proposal goes beyond this evidence by predicting structured within-image recurrence followed by between-image updating. The important empirical signature would not simply be feedback from association cortex to sensory cortex. It would be evidence that a lower-order pattern progressively converges, produces a novel relation that was not explicit in the initiating state, and is followed by recruitment of that relation into a higher-order state that subsequently alters the next sensory construction.

Thus, the strongest evidence for generative PIM would have a temporal order resembling:

\text{persistent high-level constraints}
\rightarrow
\text{progressive sensory construction}
\rightarrow
\text{emergent sensory relation}
\rightarrow
\text{higher-level incorporation}
\rightarrow
\text{changed subsequent construction}.

Existing evidence supports several arrows in this sequence. Demonstrating the entire causal chain remains an empirical objective.

13. New Empirical Predictions

The nested model generates predictions that are more specific than those of the original PIM formulation.

First, a generated image should sometimes display within-state convergence before a new high-level conclusion becomes detectable. Time-resolved decoding during tasks requiring imagery-based inference should reveal a sensory or sensorimotor representation becoming progressively more internally consistent before the critical relation appears in association-level activity.

Second, experimentally disrupting the constructive phase should have different effects from disrupting the later extraction or updating phase. Perturbation delivered while a configuration is still being assembled should degrade the coherence or precision of the resulting internal map. Perturbation delivered after the map has stabilized but before its consequence has been incorporated should instead selectively impair extraction or preservation of the inferred relation.

Third, the number of inner refinement cycles and the number of outer PIM cycles should vary partly independently. A difficult perceptual completion problem might require extensive constructive convergence but only one higher-order update. A long planning problem could involve many PIM cycles even when each individual image is constructed rapidly.

Fourth, stronger or more numerous maintained constraints should narrow the distribution of acceptable imagery states, although incompatible constraints may instead delay convergence or provoke representational instability. This prediction converts working-memory load into a structural variable: useful maintained contents constrain the search landscape, while irrelevant or mutually inconsistent contents can impair construction.

Fifth, an ART-like adequacy parameter should influence cognitive flexibility. A very permissive match criterion should allow weakly constrained representations to be accepted rapidly, increasing speed and potentially novelty at the expense of precision. An excessively strict criterion should prolong refinement, promote repeated search, or prevent progression. Optimal values should depend on whether a task rewards accurate simulation, creativity, rapid response, or exploration.

Sixth, novel information generated through imagery should sometimes appear first in modality-specific patterns and only subsequently in higher-order patterns. This prediction is already central to PIM, but the nested account further predicts that the modality-specific signal should be preceded by a measurable period of representational convergence. (Iterated Insights⁠)

Finally, changing an intermediate construction should redirect subsequent thought even when the starting state is held constant. If an imagined glass is represented as plastic rather than brittle during one intermediate cycle, the resulting simulated trajectory may shift from shattering toward bouncing. Such path dependence is a defining feature of PIM because intermediate representational commitments become constraints on subsequent processing. (Iterated Insights⁠)

14. Errors, Confabulation, and False Resonance

A constructive system gains flexibility by filling gaps, but the same property introduces error. When available information underdetermines a configuration, learned priors must contribute. Those priors may be statistically useful while remaining wrong in a particular case.

A diffusion-like interpretation makes this especially clear. Generative completion produces a plausible state, not necessarily the uniquely correct state. If a completed image introduces an unsupported feature and ascending processing subsequently treats that feature as informative, PIM can propagate the error into later cycles.

Generative resonance can likewise stabilize an internally coherent but externally inaccurate interpretation. Multiple assumptions may reinforce one another, producing a representation with high internal compatibility. Subsequent imagery then inherits those assumptions, progressively elaborating a false trajectory.

This possibility is already inherent in PIM. Progressiveness means cumulative and path-dependent transformation, not guaranteed improvement or convergence on truth. (Iterated Insights⁠)

A mature cognitive architecture therefore requires mechanisms for periodically introducing external evidence, maintaining uncertainty, preserving alternative hypotheses, detecting contradiction, and branching from earlier states. The same recursive machinery that allows an incorrect trajectory to compound can also permit reconsideration when an intermediate assumption is altered.

15. Relation to Contemporary Artificial World Models

Recent artificial world models provide useful engineering comparisons because they represent environments in compressed internal spaces and predict or generate how those representations change over time. Genie uses a spatiotemporal video tokenizer, an autoregressive dynamics model, and a latent action model. Genie 2 extends this general approach with an autoregressive latent diffusion architecture that generates subsequent latent frames conditioned on preceding latent frames and actions. (Google DeepMind⁠)

DIAMOND demonstrates another arrangement in which diffusion itself serves as a model of environmental dynamics and an agent learns from interaction inside the generated environment. DreamerV3 uses a different approach, maintaining a recurrent latent state and learning policies from imagined trajectories generated by its world model. These systems differ substantially in architecture, but collectively demonstrate that intelligent behavior can benefit from compressed internal states, recurrent dynamics, imagined futures, and internally generated consequences. (Microsoft⁠)

PIM suggests an additional architectural requirement that is not guaranteed merely by possessing a powerful generator. Generating an image or future latent state for an external user is not equivalent to using that state as part of the system’s own continuing thought. A PIM-capable architecture must allow consequences discovered within an internally generated representation to modify the persistent state responsible for constructing the next internal representation. This distinction between output generation and self-informing generation is central to the original PIM proposal. (Iterated Insights⁠)

A generative PIM machine would therefore contain a recurrent circuit of the following general form:

\text{persistent workspace}
\rightarrow
\text{conditional world-model construction}
\rightarrow
\text{internal perceptual analysis}
\rightarrow
\text{coherence and value evaluation}
\rightarrow
\text{partial workspace update}
\rightarrow
\text{new construction}.

The internally generated representation could exist in pixel space, a compressed visual latent space, a motor latent space, an auditory map, or another structured format. The essential property is not human-like visual phenomenology. It is that the generated state contains relational information that can be extracted by the system and recursively influence subsequent computation.

16. A PIM Architecture for Artificial Deliberation

The integration developed here suggests a concrete artificial cognitive architecture.

At time t, a persistent multimodal workspace maintains several representations:

W_t=\{A,B,C,D\}.

Rather than asking a generative model to produce an external answer directly, these representations condition an internal world model. A diffusion model, recurrent dynamics model, or another structured generator develops a candidate internal state through constructive convergence.

A perceptual or latent encoder then analyzes that state. Its role is not merely to reconstruct the original conditioning information. It searches for emergent relationships, predicted consequences, inconsistencies, affordances, and potentially useful features.

An evaluator compares those results with maintained constraints, goals, external evidence, and current confidence. Some conclusions are rejected, some trigger further refinement, some cause the system to branch, and some are admitted to the workspace.

Selective updating then yields:

W_{t+1}=\{B,C,D,E\}.

The generator is called again, now under the modified conditions.

The result is not simply a video generated one frame after another. The internal causal state responsible for generation changes as a consequence of what the machine discovers in its own generated states. The machine is therefore not merely predicting a world. It is thinking with the world model.

This distinction suggests a useful engineering experiment. One artificial agent could use a world model only to predict future states from a fixed initial context. A second could repeatedly re-encode those states, promote newly inferred relations into a persistent workspace, and regenerate trajectories under the revised context. Tasks could then be designed in which successful conclusions require information that becomes available only through intermediate simulated states.

If the PIM hypothesis is correct as a computational principle, the recursive system should outperform the feedforward generator especially when problems require multistep spatial inference, counterfactual reasoning, physical simulation, planning around newly discovered constraints, or the integration of intermediate results that were not explicitly represented at the outset.

17. From World Modeling to Deliberative Thought

World models are often discussed as mechanisms for predicting what will happen next. PIM suggests a broader function. An internal model can serve as a computational workspace in which abstract conditions are expanded into structured situations whose consequences can be discovered.

Prediction then becomes only one special case. The system can ask, implicitly or explicitly, what would happen if an object were moved, whether two parts would fit together, whether a route would remain passable, how an utterance would sound, how another agent might respond, or whether a planned action sequence produces a conflict.

Each simulation need not continue indefinitely. A newly exposed fact can be compressed back into the associative workspace and used without retaining the full sensory state from which it emerged.

Thus:

\text{abstract}
\rightarrow
\text{expanded}
\rightarrow
\text{inspected}
\rightarrow
\text{recompressed}.

Repeated cycles permit an initially vague problem to become progressively structured:

W_0
\rightarrow
W_1
\rightarrow
W_2
\rightarrow\cdots\rightarrow W_n.

The trajectory is intelligent not because any individual update is necessarily sophisticated, but because earlier products become constraints on later operations. Computational depth emerges from accumulation.

18. Implications for Conscious Thought

The proposed mechanism may also contribute to the continuity and constructive character of conscious thought, although it should not be equated with consciousness itself.

Iterative updating provides continuity by preserving overlapping subsets of active representations across successive states. PIM gives those evolving states a means of transforming and interrogating structured sensory or sensorimotor representations. Constructive convergence supplies a potential substructure within individual moments of imagery, explaining how a seemingly unified internal scene could itself emerge from recurrent interactions.

The subjective stream may consequently contain several nested temporal organizations. Fast recurrent processing can stabilize individual perceptual or imagined configurations. Slower working-memory updating carries selected contents from one configuration into the next. Still longer sequences preserve goals, themes, problems, or narratives across many updates.

This hierarchy could help explain why conscious thought can appear simultaneously stable and dynamic. A person can remain focused on one problem while individual images, words, relations, and intermediate conclusions change continually. Stability exists at the level of persistent constraints; change occurs in the specific states those constraints generate and in the new information returned from them.

PIM nevertheless remains a functional theory. A system could in principle implement the relevant causal operations with weak, schematic, or perhaps entirely nonphenomenal internal representations. The framework addresses the organization and continuity of cognitive processing rather than claiming that generative resonance alone explains why any state is subjectively experienced.

19. Discussion

The synthesis proposed here begins with a simple question: what happens between an active set of cognitive constraints and the completed internal representation those constraints generate?

Progressive imagery modification previously described the larger reciprocal cycle. Persistent higher-order representations generate a structured sensory or sensorimotor map; features extracted from that map alter the higher-order state; the changed state generates another map. This cycle allows imagery to become a computational participant in thought rather than a passive display. (Iterated Insights⁠)

The present extension opens the generative operation itself. A map need not emerge in finished form. It can be progressively assembled through recurrent interactions among top-down specifications, sensory evidence, lateral constraints, prior knowledge, and goals. This within-map process has been termed constructive convergence.

Adaptive Resonance Theory provides a biologically motivated precedent for recurrent matching and stabilization between hierarchical levels. Diffusion models provide an engineering precedent for progressively constructing complex distributed representations through repeated conditional refinement. PIM adds the crucial outer loop: the completed construction is analyzed, and its newly exposed consequences change the state that will guide the next construction.

The combined sequence is:

\boxed{
\text{maintain constraints}
\rightarrow
\text{progressively construct}
\rightarrow
\text{achieve sufficient coherence}
\rightarrow
\text{extract consequences}
\rightarrow
\text{partially update constraints}
\rightarrow
\text{construct again}
}

This architecture can be described as generative resonance.

The synthesis also changes the interpretation of multiassociative search. Multiple active representations do not merely nominate the next memory. They collectively define the conditions that the next representation should satisfy. The result may be retrieval when a stored representation already provides a strong solution, but it may be construction when the conjunction of constraints requires a novel state.

This suggests a general computational definition of thought:

\boxed{
\textbf{Thought is the progressive, constraint-conditioned construction of internal states whose emergent consequences recursively alter the constraints responsible for constructing subsequent states.}
}

The formulation applies naturally to visual imagery, but it need not be restricted to vision. Motor systems can construct possible trajectories, auditory systems can construct acoustic patterns, language systems can construct candidate utterances, and multimodal systems can coordinate combinations of these formats. The common principle is reciprocal transformation between persistent contextual representations and more structured generative states.

The distinction between constructive convergence and progressive modification may prove particularly useful. Constructive convergence explains how one representation is assembled. Progressive modification explains how one representation leads cognitively to another. The first solves an underdetermined representational problem; the second changes the problem by incorporating the solution’s consequences.

These processes may operate at different timescales while remaining recursively coupled. Inner convergence produces a map. The map produces information. Information changes the workspace. The changed workspace changes the landscape over possible maps. What was an output at one level becomes a constraint at the next.

20. Conclusion

Progressive imagery modification proposes that cognitive systems can reason by repeatedly constructing and interrogating internally generated representations. A partially persistent associative state constrains a sensory or sensorimotor map, the map introduces structured relations, ascending systems extract its implications, and selected implications partially update the state responsible for generating what comes next.

The present theory adds a second level of recurrence. Each map may itself emerge through constructive convergence, with recurrent processing progressively satisfying the multiple constraints supplied by working memory, learned priors, sensory evidence, and goals. Adaptive Resonance Theory provides a framework for understanding reciprocal matching and stabilization, while diffusion models demonstrate the computational power of iterative conditional construction. These parallels motivate generative resonance without requiring that their specific implementations be identical.

The resulting architecture contains nested loops. Within a cognitive state, distributed activity converges toward a usable representation. Between cognitive states, information discovered in one representation changes the conditions under which the next representation is constructed. The combination produces continuity without stasis and transformation without fragmentation.

Progressive imagery modification can therefore be understood as more than sequential imagination. It is a mechanism through which a cognitive system repeatedly asks its own representational machinery to instantiate the implications of what it currently knows. Each constructed state becomes both an expression of the current model and an experiment performed upon that model.

For artificial intelligence, the corresponding principle is equally specific. A machine would not acquire PIM merely by generating realistic imagery or predicting future video frames. Its generated states would need to become objects of its own internal perception, yield new information, modify a persistent workspace, and thereby alter subsequent generation. Such a system would not simply possess a world model. It would recursively use the world model as an instrument of thought.

References

Alonso, E., Jelley, A., Micheli, V., Kanervisto, A., Storkey, A., Pearce, T., & Fleuret, F. (2024). Diffusion for world modeling: Visual details matter in Atari. Advances in Neural Information Processing Systems.

Bruce, J., Dennis, M., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., et al. (2024). Genie: Generative interactive environments. Proceedings of the 41st International Conference on Machine Learning.

Carpenter, G. A., & Grossberg, S. (2003). Adaptive resonance theory. In M. A. Arbib (Ed.), The Handbook of Brain Theory and Neural Networks (2nd ed., pp. 87-90). MIT Press.

Dijkstra, N., Ambrogioni, L., Vidaurre, D., & van Gerven, M. A. J. (2020). Neural dynamics of perceptual inference and its reversal during imagery. eLife, 9, e53588. doi:10.7554/eLife.53588.

Dijkstra, N., Bosch, S. E., & van Gerven, M. A. J. (2019). Shared neural mechanisms of visual perception and imagery. Trends in Cognitive Sciences, 23(5), 423-434. doi:10.1016/j.tics.2019.02.004.

Dijkstra, N., Zeidman, P., Ondobaka, S., van Gerven, M. A. J., & Friston, K. (2017). Distinct top-down and bottom-up brain connectivity during visual perception and imagery. Scientific Reports, 7, 5677. doi:10.1038/s41598-017-05888-8.

Grossberg, S. (2013). Adaptive Resonance Theory: How a brain learns to consciously attend, learn, and recognize a changing world. Neural Networks, 37, 1-47. doi:10.1016/j.neunet.2012.09.017.

Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2025). Mastering diverse control tasks through world models. Nature, 640, 647-653. doi:10.1038/s41586-025-08744-2.

Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33.

Pearson, J. (2019). The human imagination: The cognitive neuroscience of visual mental imagery. Nature Reviews Neuroscience, 20, 624-634. doi:10.1038/s41583-019-0202-9.

Peebles, W., & Xie, S. (2023). Scalable diffusion models with Transformers. Proceedings of the IEEE/CVF International Conference on Computer Vision, 4195-4205. doi:10.1109/ICCV51070.2023.00387.

Reser, J. E. (2013). The neurological process responsible for mental continuity: Reciprocating transformations between a working memory updating function and an imagery generation system. Presented at the Association for the Scientific Study of Consciousness Conference, San Diego, California.

Reser, J. E. (2016). Incremental change in the set of coactive cortical assemblies enables mental continuity. Physiology & Behavior, 167, 222-237. doi:10.1016/j.physbeh.2016.09.019.

Reser, J. E. (2022). Artificial intelligence software structured to simulate human working memory, mental imagery, and mental continuity. arXiv, 2204.05138.

Reser, J. E. (2022/2024). A cognitive architecture for machine consciousness and artificial superintelligence: Thought is structured by the iterative updating of working memory. arXiv, 2203.17255.

Reser, J. E. (2026). Progressive imagery modification: A recurrent mechanism for imagination, mental simulation, and deliberative thought. Iterated Insights, September 1, 2026.

OpenAI. (2024). Video generation models as world simulators.

Parker-Holder, J., Ball, P., Bruce, J., Dasagi, V., Holsheimer, K., Kaplanis, C., et al. (2024). Genie 2: A large-scale foundation world model. Google DeepMind.

Posted in

Leave a Reply

Discover more from Iterated Insights

Subscribe now to keep reading and get access to the full archive.

Continue reading