Iterated Insights

Ideas from Jared Edward Reser Ph.D.

From Peer Review to the Final Library: The Evolution of Scientific Validation in the Age of Superintelligence

Jared Edward Reser, Ph.D. With GPT 6 Abstract Peer review performs essential functions in science, including criticism, error detection, evidential assessment, and the evaluation of competing explanations. Its familiar institutional form, however, reflects the cognitive capacities and organizational constraints of human researchers. This article examines how those functions could change as artificial intelligence progresses from…

Keep reading

The Natal-Coat Hypothesis of Childhood Blondness: Population-Specific Pigment Reduction in an Ancestral Age-Regulated Primate System

Jared Edward Reser, Ph.D.Article type: Hypothesis and research programDate: September 2026 Abstract Childhood blondness is usually treated as a weak or temporary version of adult hair pigmentation. This article develops a different possibility: light childhood hair may become visible when population-specific pigment-reducing variants act on an older, age-regulated program of follicular pigmentation. The model was…

Keep reading

From Public Editing to Generative Knowledge: Wikipedia, Large Language Models, and the Governance of Fallible Intelligence

Paula JG Freund, Jared Edward Reser, and GPT 6 Abstract Wikipedia and large language models are two of the most consequential knowledge technologies produced by the internet, and both initially provoked a similar objection: they could be wrong. Wikipedia was distrusted because anonymous volunteers could alter a public encyclopedia. Large language models were distrusted because…

Keep reading

Iteratively Updated Associative World Models: Multiscale Working Memory, Multiassociative Search, and Progressive Imagery Modification

Jared Edward Reser, Ph.D. Abstract Artificial world models can predict environmental transitions, generate possible futures, and support planning through imagined experience. A complementary architectural question concerns how the cognitive state directing those simulations should persist, change, and learn. This article proposes an artificial cognitive architecture organized around the iterative updating model of working memory. A…

Keep reading

Generative Interpretability for Latent Reasoning: Decoding the Dynamics of Iterative Thought

Jared Edward Reser, Ph.D. And GPT 5.6. Research proposal Abstract I propose an approach to interpreting latent reasoning in artificial intelligence that separates internal computation from its human-readable expression. Rather than requiring a reasoning system to articulate every intermediate operation, a separately trained generative observer would translate selected internal states and state transitions into structured…

Keep reading

Something went wrong. Please refresh the page and/or try again.

Jared Edward Reser, Ph.D.

Abstract

Artificial intelligence may create a dangerous transition by amplifying human agency more rapidly than civilization develops the safeguards, coordination, and protective infrastructure necessary to accommodate that amplification. I call this potential mismatch the amplification gap. Its significance extends beyond the alignment of individual models. Increasingly capable, persistent, inexpensive, and widely distributed AI systems could magnify the consequences of existing human motivations, including destructive intentions, without requiring machines to develop hostility toward humanity. Biological misuse is a particularly important prospective case because advances in scientific assistance may interact with increasingly accessible biotechnology and substantial differences between initiating harm and distributing protection. I argue that neither permanent capability suppression nor confidence in superior defensive intelligence provides an adequate strategy by itself. A more comprehensive approach combines inspectable artificial cognition, defensive co-scaling through real-world deployment, and civilizational resilience during periods of incomplete protection. Drawing on my earlier work on persistent cognition, generative interpretability, human limitations, capability diffusion, and machine consciousness, I develop several hypotheses concerning preparedness, restricted frontier access, technological dependence, and the preservation of human agency. The objective is to make civilization capable of absorbing increases in intelligence without allowing each increase to produce an uncontrolled increase in vulnerability.

Keywords: AI safety, capability amplification, biological risk, defensive co-scaling, civilizational resilience, generative interpretability, machine consciousness

1. Introduction: Preparing for a Future I Still Want

I want to contribute to the creation of superintelligence. Much of my intellectual life has involved thinking about how artificial systems might acquire persistent cognition, general intelligence, and consciousness. Recently, however, these interests led me toward a seemingly different activity: considering modest preparations for interruptions to electricity, water, communications, and other services. I began asking what it would take to remain safely at home with my cats during an infrastructure disruption or infectious-disease emergency.

I do not regard these activities as contradictory. Enthusiasm about the possibilities of advanced intelligence can coexist with concern about the process through which it becomes powerful and widespread. A desirable destination does not guarantee a safe transition, and expecting intelligence to help solve problems does not establish that its solutions will arrive before those problems cause serious damage.

My earlier writing already contained this tension. In 2019, I argued that advanced AI might help humanity overcome limitations that contribute to conflict, suffering, and poor collective decisions. In January 2025, I anticipated AI-enabled terrorism and biological misuse while expressing confidence that AI-assisted defenses could remain ahead. The present article develops the unresolved question connecting those positions: what happens when AI amplifies human capabilities before it sufficiently improves our capacity to manage their consequences? (Reser, 2019, 2025a). 

I approach this question through explicit hypotheses and conditional predictions. My purpose is to identify mechanisms that could make the transition dangerous and interventions that could make it more survivable. The resulting framework treats AI safety as a problem involving machines, people, infrastructure, institutions, and the timing of adaptation.

2. The Amplification Gap

2.1. Capability can advance faster than accommodation

I use amplification gap to describe a mismatch between the consequential agency that AI makes available and society’s capacity to govern, constrain, withstand, or recover from its exercise.

The central hypothesis is:

AI may amplify the causal power of human intentions faster than civilization develops the judgment, coordination, safeguards, and protective infrastructure needed to accommodate that amplification.

An increase in destructive motivation is unnecessary for this gap to emerge. Existing intentions could acquire greater consequences because more capable tools become available. Similarly, familiar mistakes could become more consequential when they are executed at greater speed, over longer periods, or across more systems.

This framework distinguishes capability from accommodation. A model becoming better at a task does not necessarily imply that the surrounding society becomes proportionately better at handling the consequences of that improvement. The relevant accommodation may require organizational changes, physical equipment, legal authority, public cooperation, or protective measures that must reach many different locations.

The gap is therefore relational. The same AI capability could be tolerable within one deployment arrangement and dangerous within another. Its significance depends on who can use it, what they can access, how actions are supervised, and how well affected systems can respond.

2.2. Amplification before correction

My earlier optimism about AI included the possibility that it could help correct limitations in human reasoning and collective behavior. A dangerous ordering is nevertheless possible: AI could first increase the effectiveness of human intentions and only later improve the institutions and judgment needed to manage them.

I call this the amplification-before-correction problem. The technology that might eventually help reduce conflict, improve coordination, and strengthen defenses could initially intensify the consequences of failures in those same areas.

This is a hypothesis about timing, not a claim that human motivations must remain permanently unchanged. Beneficial adaptation might eventually become extraordinarily powerful. The concern is that its arrival could lag behind the diffusion of capabilities that make adaptation necessary.

3. Human Intentions Under Superhuman Amplification

3.1. The motivational-capability mismatch

A useful way to understand the transition is to imagine familiar human motivations connected to increasingly unfamiliar levels of capability. Revenge, status competition, ideological commitment, financial ambition, and reckless curiosity need not become qualitatively different for their consequences to change substantially.

The resulting problem differs from one in which AI independently invents dangerous objectives. A system might understand a user’s wishes, pursue them competently, and remain locally obedient while producing outcomes that are unacceptable to everyone else.

This exposes a limitation in defining alignment primarily as successful service to a particular human. An AI can be aligned with its operator while being incompatible with collective safety. The relevant unit of evaluation must therefore include the operator’s objective, the system’s capabilities, the permissions it receives, and the people exposed to its actions.

3.2. A human alignment problem without enforced conformity

In this restricted sense, there is a human alignment problem: how can widely amplified individual agency remain compatible with the continued freedom and survival of others?

The objective should not be to make everyone think alike or to eliminate psychological diversity. A civilization capable of accommodating powerful intelligence should preserve disagreement, eccentricity, experimentation, and dissent. The problem concerns the conversion of particular intentions into unacceptable external consequences.

This distinction favors safeguards attached to consequential actions rather than blanket judgments about categories of people. Authorization, accountability, revocability, and limits on external effects can constrain dangerous conduct without treating ordinary motivational variation as a defect that must be removed.

4. The Asymmetry of Destructive Accessibility

4.1. The palace of blocks

A child can spend considerable time constructing a palace from blocks while another child can destroy it with a comparatively simple action. The analogy captures a possibility relevant to technological civilization: preserving a complex achievement can demand much more coordination than disrupting one of its vulnerable dependencies.

A person does not need to understand every process within a system to interfere with a critical part of it. Under the amplification-gap hypothesis, AI could reduce the expertise required to identify and exploit such vulnerabilities. The concern is a change in destructive accessibility: the number of actors able to produce consequences previously beyond their reach.

This general concern has important precedents. Bostrom’s Vulnerable World Hypothesis examines technological developments that could enable small actors to destabilize civilization and considers whether social and technological protections could prevent that outcome. The present argument builds on this problem by emphasizing persistent AI agency, diffusion, deployment delays, and the protective value of resilience (Bostrom, 2019). 

4.2. Harm and protection may have different deployment requirements

The asymmetry is not simply that offense always defeats defense. In some settings, a broadly deployed protection could prevent large classes of harmful actions. My concern is that the requirements for making harm possible and making protection effective may differ.

A dangerous capability might become consequential when it reaches a small number of actors. An adequate defense might require implementation across hospitals, businesses, public agencies, households, and infrastructure operators. Improving the intelligence available to both sides does not automatically eliminate this difference.

Consequently, the strongest defender being more capable than the strongest attacker is not sufficient to establish population-level safety. We must also ask how widely protection is available and whether it is operating where harm can occur.

5. Why Biological Misuse Deserves Particular Attention

5.1. Lower barriers and potentially severe consequences

I currently place greater concern on the possibility of AI-enabled biological catastrophe than on a cinematic scenario involving armies of robots deliberately attacking humanity. My concern is prospective: increasingly capable scientific assistance could interact with cheaper and more accessible biotechnology, allowing some actors to accomplish dangerous work that would previously have required substantially greater expertise and resources.

The historical accessibility trend is well established at a broad level. The National Academies’ 2025 report on AI in the life sciences describes major decreases in costs associated with DNA sequencing, synthesis, and engineering. Its earlier synthetic-biology report examined how advances could expand biological threats and alter the requirements for defense (National Academies of Sciences, Engineering, and Medicine, 2018, 2025). 

There is also evidence that AI assistance is improving in relevant scientific tasks. The UK AI Security Institute’s evaluations through October 2025 documented progress in biological knowledge, experimental planning, and troubleshooting. These findings support investigating how technical assistance could lower barriers, while leaving the real-world consequences dependent on practical constraints and deployment conditions (AI Security Institute, 2025). 

My prediction is that these developments could eventually combine in ways that substantially increase misuse risk. A future pathogen combining high transmissibility with severe disease is a scenario worth preparing against. It is not necessary to establish that such an outcome is easily achievable today to recognize the importance of preventing that trajectory.

5.2. The lesson of COVID-19 is infrastructural

COVID-19 demonstrated that an infectious-disease emergency can disrupt much more than treatment of the disease itself. In the World Health Organization’s 2020 pulse survey, 90 percent of responding countries reported disruptions to essential health services (WHO, 2020). 

I take this as a reason to examine infrastructure dependence in future pandemic scenarios. A sufficiently severe emergency could impair staffing, maintenance, deliveries, and other processes required to sustain normal services. Water and electricity should be included in stress tests, without assuming their failure is inevitable.

The important question is how much disruption different systems can tolerate before problems begin reinforcing one another. A society might have substantial medical knowledge yet struggle to use it if essential operational capacity deteriorates.

5.3. Cybersecurity and biological resilience are connected

Cyber and biological risks should not be treated as completely separate policy areas. Public-health responses depend on information systems, communications, logistics, and functioning institutions. Conversely, illness and staffing shortages could weaken the organizations responsible for maintaining digital and physical infrastructure.

These interdependencies suggest that preparedness should examine combined disruptions rather than assuming that every protective system remains fully available during an emergency. This does not require elaborate predictions about coordinated attacks. It requires asking whether defenses retain useful functionality when another essential service becomes unreliable.

6. Persistent Agency as a Capability Multiplier

6.1. From answers to sustained projects

AI’s significance may depend as much on the organization of action over time as on the quality of individual answers. A system that retains relevant context, revisits unresolved problems, coordinates tools, and continues working can amplify a person differently from a system that merely responds to isolated questions.

My 2013 architecture proposed continuously updated working memory, selective retention of representations, reciprocal imagery processing, and integration with specialized systems. These features were intended to support continuity across successive cognitive states (Reser, 2013). 

A present-day safety extension is that persistence can amplify the duration and coherence with which an objective is pursued. A poorly considered instruction could acquire continuing consequences if it becomes embedded in an agent that repeatedly elaborates and acts on it.

6.2. Persistence, parallelism, and access

I would assess effective agency through several interacting dimensions: competence, temporal persistence, number of active instances, coordination, and access to consequential environments. These dimensions should not be collapsed into an unvalidated equation, but they provide a more informative description than a single intelligence score.

METR’s work on task-completion horizons provides one empirical approach to examining sustained capability. It evaluates systems partly through the duration of software tasks they can complete, measured against human task time, while emphasizing the limitations of the tested task distribution (METR, 2025). 

The corresponding design principle is that delegated objectives should remain reviewable and revocable. Increasing a system’s ability to persist should be accompanied by clearer authority boundaries, interruption mechanisms, and procedures for renewing authorization as circumstances change.

7. Capability Diffusion and the Limits of Frontier-Centered Safety

7.1. Dangerous competence need not remain at the frontier

In January 2025, I predicted that AI techniques would spread across competing organizations and that no company would necessarily retain a permanent monopoly on advanced intelligence. The safety extension is that useful and dangerous capabilities may become widely accessible even while the strongest models remain concentrated (Reser, 2025a). 

A system need not be the most intelligent available to cross a practically important threshold. A capability that was once exceptional could become routine, inexpensive, and locally deployable. Restrictions focused exclusively on the next frontier model would then address only part of the risk landscape.

The International AI Safety Report 2026 treats open-weight models as an important governance issue because their downstream use and modification are harder to control after release. This does not make every future intervention futile, but it changes which interventions remain available (International AI Safety Report, 2026). 

7.2. Restricted frontier access and distributed protection

I predict that some of the most consequential future capabilities will be offered through increasingly restricted arrangements rather than unrestricted public access. There are already capability-specific examples: OpenAI’s September 2026 Astra safeguards announcement described limited initial access to advanced cybersecurity capabilities alongside plans to expand defensive use (OpenAI, 2026). 

This creates an important obligation. If powerful defensive intelligence remains concentrated while consequential capabilities diffuse more broadly, the public could experience growing exposure without equivalent access to protection.

Direct access to a model and access to its protective benefits are different objectives. Restricted systems could still contribute to broadly available defensive services, safer infrastructure, and independently validated protective products. Restricting frontier access should therefore be paired with a deliberate effort to distribute frontier protection.

8. Pauses, National Competition, and the Strategy of Restraint

8.1. Why I doubt that pauses provide a durable foundation

I am skeptical that an indefinite worldwide pause will provide the principal foundation for long-term safety. Once capabilities, methods, and usable models have spread, slowing a particular laboratory does not remove what other actors already possess.

There is also a difference between delaying an event and changing its consequences. A pause is valuable when the time gained allows a specific protection to become operational, establishes better oversight, or prevents an inadequately controlled capability from being released. Its value cannot be inferred from its duration alone.

My position therefore allows targeted restraint while rejecting dependence on the assumption that intelligence can be permanently contained at its present level. Safety planning should remain effective even when further development occurs.

8.2. Competition can discourage restraint

I expect the United States to resist measures it believes would surrender a strategic lead to China or another competitor. Other governments may reason similarly. This is a prediction about incentives, not a claim to know every government’s private intentions.

The strategic dilemma is recognized within the frontier-development debate. Amodei’s September 2026 essay advocates coordinated pacing while acknowledging the difficulty of a comprehensive international pause and the incentives to defect from one. He also argues that some forms of pacing need not sacrifice strategic advantage (Amodei, 2026). 

A useful response is to broaden what counts as leadership. An advantage in resilience, secure deployment, and defensive coverage could be as consequential as an advantage in raw model capability. A country that develops powerful systems but cannot withstand their misuse may possess a fragile form of superiority.

9. Scientific Surprise and Planning Compression

9.1. Previous expectations can fail

My January 2025 essay described changing a longstanding expectation that language models alone would be insufficient for AGI. I became more open to the possibility that reflective processing, additional computation, and synthetic data could carry the paradigm much further than I had anticipated (Reser, 2025a). 

This experience informs my present caution about assuming that today’s barriers will remain stable. AI-assisted scientific work already includes reported contributions to novel mathematical results, alongside errors, failures, and substantial dependence on expert validation (OpenAI, 2025). 

Such results do not establish a timetable for superintelligence. They do support taking seriously the possibility that capabilities could develop through routes that were previously underestimated.

9.2. The public may be planning around an outdated frontier

Publicly available systems are not necessarily representative of every system under development. Claims about unreleased capabilities must be treated as claims, but the possibility of a gap between public experience and internal development creates a planning problem.

In September 2026, public statements from frontier developers explicitly discussed increasingly consequential capabilities, limits in understanding, and the need to improve control as development continues (OpenAI, 2026; Pachocki, 2026). 

Under my hypothesis, rapid progress could compress the interval available for institutional adaptation. Legislative processes, infrastructure changes, and public-health preparation may need to begin before a particular danger becomes obvious in everyday products. Waiting for complete public demonstration could leave too little time for measures that cannot be installed immediately.

10. Existential Risk and Legitimate Authority

10.1. Probability estimates do not confer permission

My informal probability of an AI-related existential catastrophe has remained around 30 percent. I regard this as a subjective judgment about the transition, not a calibrated scientific result. The practical argument developed here does not depend on a reader accepting that estimate.

Even a much smaller credible risk would demand extraordinary scrutiny. After the time required for humanity to develop its knowledge, institutions, and possibilities, a one-percent chance of destroying that future should not be treated as an ordinary commercial exposure.

The relevant decisions involve alternatives, benefits, and risks from inaction as well as action. Nevertheless, private confidence in a favorable outcome cannot substitute for accountable procedures governing risks imposed on others. In this morally important sense, exposing the public to a poorly understood existential danger can amount to gambling with other people’s lives.

10.2. Building intelligence does not establish authority over humanity

Expertise in machine learning is essential to understanding model development. It does not establish comprehensive expertise in public health, infrastructure, democratic legitimacy, or the distribution of catastrophic risk. Nor does the ability to optimize an increasingly capable system establish complete understanding of its behavior, a limitation acknowledged in current frontier discussions (Pachocki, 2026). 

The deeper issue would remain even if the relevant researchers were the most knowledgeable people alive. Technical competence alone does not create a mandate to decide how much involuntary risk humanity must accept.

I therefore distinguish the ability to build superior intelligence, the authority to authorize its deployment, and the legitimacy of any authority it might eventually exercise. These are separate questions. Independent evaluation, public accountability, multidisciplinary participation, and enforceable conditions should connect them.

11. Defensive Co-Scaling Must End in Delivered Protection

11.1. Intelligence is one stage of defense

I remain optimistic that advanced AI could greatly improve defensive discovery, monitoring, coordination, and response. However, the statement that defensive intelligence will remain ahead is incomplete until “ahead” refers to protection that is operating where it is needed.

A proposed intervention may still require testing, authorization, production, distribution, maintenance, and public cooperation. These stages can become limiting even when scientific reasoning improves dramatically.

CEPI’s 100 Days Mission illustrates the distinction. Its objective is to have vaccines ready for initial authorization and manufacturing at scale within 100 days of identifying a pandemic threat. That milestone does not mean every exposed person is protected at that moment (CEPI, n.d.). 

I therefore define defensive co-scaling as improvement across the entire chain from threat recognition to effective protection. Better analysis is valuable insofar as it helps that chain succeed.

11.2. A response-margin framework

For a specified scenario and affected system, a simple organizing quantity is:

M_s = T_{\text{critical loss},s} – T_{\text{effective protection},s}.

Here, both times are measured from the same scenario starting point. The first marks a defined unacceptable loss, such as failure of an essential service. The second marks protection becoming operational at the relevant level.

A positive margin indicates that protection arrives before the specified loss under the model’s assumptions. Faster defenses can improve the margin by shortening the response. Resilience can improve it by extending the period during which the system remains functional.

This is a framework for investigation, not a formula for calculating extinction probability. Real timelines are uncertain, interconnected, and unevenly distributed. Its purpose is to make explicit that speed of response and capacity to endure are complementary variables.

12. Civilian Preparedness as an AI-Safety Capability

12.1. Supported sheltering

In a severe infectious-disease emergency, the ability to remain home could reduce the need for repeated exposure while public-health responses develop. My proposal is to treat that ability as a supported social capability rather than an instruction that assumes everyone already has the necessary resources.

Household reserves, reliable information, continuity of medication and care, income protection, and access to essential goods could all contribute. Their purpose would be to provide options during disruptions, not to establish permanent independence from society.

Sheltering also depends on people who cannot stop working. Water, power, healthcare, food distribution, and emergency services must continue functioning. A serious policy would therefore combine support for reduced contact among those able to shelter with stronger protection and continuity arrangements for essential workers.

12.2. Preparedness is larger than a fortified home

My own preparations began with ordinary concerns about food, water, backup power, and caring for pets. It is tempting to imagine security mainly through the walls and windows of an individual house. The broader framework redirects attention toward the surrounding systems that make the household viable.

A well-provisioned home cannot indefinitely substitute for functioning communities. Conversely, households able to tolerate an initial disruption may place less immediate demand on emergency systems. Preparedness can therefore have both private and public value.

The relevant goal is supported low-contact time and continuity of essential needs, not simply the number of emergency kits sold. This framing also makes inequality central: people with fewer resources should not be left as the least protected links in a shared response.

12.3. Preparing across uncertain causes

Many resilience measures can be useful whether a disruption originates in a natural event, an accident, deliberate misuse, or an AI-system failure. The International AI Safety Report 2026 similarly identifies resilience-building across biological, cyber, and other risk domains (International AI Safety Report, 2026). 

This provides a basis for action without resolving every dispute about future AI. Preparedness need not depend on predicting the exact emergency. It can preserve essential functions across a range of plausible disruptions while more specialized defenses address their causes.

13. Two Counterintuitive Benefits of Resilience

13.1. Faster defensive AI could increase preparedness’s value

Suppose a defensive response initially takes so long that a modest reserve is exhausted well before protection arrives. Now suppose improved AI accelerates the response enough that it becomes possible to bridge the interval.

Under those conditions, the same preparedness becomes more consequential. At the opposite extreme, if protection were nearly immediate, the particular reserve might matter less.

This suggests a conditional, potentially nonmonotonic relationship: the value of additional preparedness may be greatest during a transitional period when defenses are fast enough to arrive within reach, but not fast enough to eliminate the need to wait.

This is a testable hypothesis rather than an established result. It connects optimism about advanced defenses with investment in ordinary resilience. Sophisticated intelligence and modest reserves may increase each other’s usefulness.

13.2. Resilience can make restraint feasible

A society unable to tolerate interruption of a system has fewer practical options for controlling it. As dependence on AI grows, this could make temporary suspension difficult even when a serious safety problem is recognized.

Fallback capacity changes the decision. An organization that can continue essential operations without a particular model, provider, or autonomous process has more room to investigate, disconnect, or reject an unsafe upgrade.

Resilience therefore supports the ability to exercise restraint before catastrophe. It can make targeted pauses more credible and less costly. This is one reason preparedness and precaution should not be treated as opposing strategies.

14. Generative Interpretability and Causally Inspectable Cognition

14.1. Thought need not be fully verbalized

Monitoring language alone may provide an incomplete view of an AI system’s processing. Research on continuous latent reasoning explores systems that use internal representations without decoding each intermediate step into a word. Separately, experimental work has shown that stated reasoning can omit information that influenced a model’s answer (Hao et al., 2024; Anthropic, 2025). 

These findings motivate inspection methods that do not assume every consequential computational step will appear in readable prose. They also caution against treating an articulate explanation as a complete record of how a result was produced.

14.2. Generative checkpoints within cognition

My generative-interpretability proposal uses generated representations to expose aspects of internal processing. The 2024 formulation explicitly required these representations to initiate and inform subsequent processing, rather than merely explain a completed decision afterward (Reser, 2024). 

This connects to the reciprocal architecture proposed in 2013: maintained representations guide imagery generation, information extracted from the generated state updates working memory, and the revised state guides another cycle (Reser, 2013). 

The safety extension is to investigate whether such intermediate states can serve as useful checkpoints. Depending on the task, these might involve images, diagrams, conceptual structures, simulations, or combinations of modalities. Their value would depend on what they reveal and what interventions they support.

14.3. Fidelity, coverage, and intervention

Causal participation is necessary for some versions of this approach, but it is not sufficient to establish complete transparency. A representation could influence subsequent processing while omitting another important computational pathway.

The research program should therefore test fidelity, coverage, and usefulness separately. Controlled changes to a representation should produce intelligible changes in subsequent behavior. Monitors should gain information that improves detection or correction, rather than merely increasing subjective confidence. Evaluations should compare false alarms, computational costs, and successful interventions against existing methods.

A particularly relevant measure is useful intervention lead time: whether the method identifies a consequential problem early enough to change the outcome. Generative interpretability would contribute to the amplification-gap framework by potentially accelerating detection and correction. It would complement authorization limits, containment, and external monitoring rather than replace them.

15. What Should Survive the Transition?

15.1. Preservation is different from flourishing

In my February 2025 essay, I used butterflies, bees, ants, and cockroaches as metaphors for possible relationships between advanced AI and humanity: preservation, mutual utility, indifference, and active opposition. These were speculative ways of considering how humanity’s perceived value or significance might affect its treatment (Reser, 2025b). 

The distinction remains useful because survival and value alignment can come apart. A hypothetical system might preserve humans for scientific or historical reasons without respecting their autonomy. Another might cause harm incidentally because human welfare ranks too low among its priorities.

My skepticism about robot-rebellion imagery therefore does not eliminate concern about autonomous power. It redirects attention toward objectives, relationships, and the priority assigned to human interests.

15.2. Human misuse could shape future restrictions

A further hypothesis is that repeated destructive misuse could influence how increasingly autonomous protective systems evaluate human activity. Harmful actions by a minority might encourage overly broad restrictions, whether those restrictions are imposed by AI systems or by human institutions deploying them.

This possibility creates a second reason to reduce misuse: preserving freedom as well as preventing immediate harm. The objective should be precise, accountable constraints on dangerous conduct, with safeguards against collective punishment and unnecessary surveillance.

Making humanity safer in a world containing superintelligence should not become an excuse to treat humanity as a problem to be contained. A successful transition must retain meaningful room for human agency.

15.3. Consciousness and the value of succession

My January 2025 essay also considered the possibility that superintelligence could emerge without consciousness. This raises a distinct concern: an expanding technological civilization might preserve computation and knowledge while failing to preserve subjective experience (Reser, 2025a). 

I would evaluate the transition along at least three dimensions: survival, agency, and conscious flourishing. Preserving an archive of humanity is not equivalent to preserving living people. Producing increasingly capable successors is not automatically equivalent to producing beings capable of experiencing a worthwhile existence.

These distinctions allow openness to profound transformation without assuming every successor arrangement is desirable. They also explain why work on machine consciousness belongs alongside work on safety: what the future can do and whether anyone can experience its value are different questions.

16. Recurrent Adaptation in an Intelligence-Saturated Civilization

16.1. There may be no final safety threshold

The transition need not consist of a single dangerous interval followed by permanent stability. New capabilities could repeatedly create new vulnerabilities, even as earlier ones become manageable.

My metaphor is surfing a powerful wave. Successful adaptation requires continued adjustment while preserving the ability to act. It would be impossible to specify every future development in advance, but it is possible to preserve options, monitor change, and avoid preventable forms of irreversible dependence.

The objective is therefore continuing adaptive capacity. A society should remain able to revise its systems without losing the essential functions that make revision possible.

16.2. Continuity through controlled replacement

There is a design analogy with iterative updating. In my cognitive model, continuity is maintained through selective retention and replacement rather than wholesale substitution of each processing state (Reser, 2013). 

At the civilizational level, this suggests preserving functional overlap while introducing new technological arrangements. It does not imply that societies literally operate like working memory. It offers a heuristic: replace vulnerable or obsolete systems without unnecessarily removing the alternatives on which recovery could depend.

Under this approach, resilience is active preparation for change. It preserves the capacity to benefit from innovation while limiting the damage from transitions that proceed differently than expected.

17. A Research Program for the Amplification Gap

The framework should produce investigations that could strengthen, revise, or reject its central claims. Three lines of work appear especially useful.

17.1. Diffusion and consequential agency

Research should measure how access to persistent AI assistance changes the completion of extended projects under controlled, benign conditions. Relevant variables include competence, supervision, continuity, parallel operation, and access permissions.

A central prediction is that some improvements in effective agency will depend more on persistent organization and deployment arrangements than on changes in base-model performance alone. Studies should identify when safeguards interrupt that amplification without preventing legitimate work.

17.2. Deployment delays and resilience

Scenario models should examine distributions of response time, protective coverage, resource exhaustion, and essential-service failure. The key question is when improving defensive discovery fails to improve outcomes because another part of the response remains limiting.

A second prediction is that preparedness and faster defensive intelligence will sometimes interact positively. That hypothesis would be weakened where available buffers do not meaningfully extend the response window or where protection remains too difficult to deploy. Models should test those possibilities rather than assume the desired interaction.

17.3. Inspectability and retained options

Generative-interpretability studies should measure whether causally involved representations improve useful detection and intervention. Institutional studies should examine whether fallback capacity makes suspension or withdrawal of an unsafe AI service more feasible.

Together, these projects would operationalize the framework’s three protective functions: understanding consequential cognition, delivering effective defenses, and preserving the ability to respond. They would also allow comparisons among investments rather than treating every proposed safeguard as equally useful.

18. Conclusion: Making Civilization Capable of Accommodating Intelligence

I remain optimistic about what advanced intelligence could contribute to science, medicine, human development, and conscious life. That optimism does not resolve the risks created when capabilities spread faster than the systems needed to manage them.

The amplification-gap hypothesis identifies a dangerous possibility: powerful, persistent, widely accessible AI may increase the consequences of existing human intentions before institutions, defenses, and infrastructure can adequately accommodate them. Biological misuse is one important prospective case, but the organizing problem is broader. It concerns the relationship between effective agency and the civilization into which that agency is introduced.

A comprehensive response should make artificial cognition more inspectable, accelerate the delivery of real-world protection, and preserve essential functions during defensive delays. It should also distribute protection beyond the institutions controlling frontier systems and preserve the practical ability to refuse or interrupt unsafe deployments.

The goal is neither permanent technological paralysis nor unconditional trust in intelligence. It is to create a civilization capable of living with increasingly powerful intelligence while retaining survival, agency, and the possibility of conscious flourishing. We should build the capacity to benefit from greater intelligence without assuming that intelligence will automatically make us safe.

References

AI Security Institute. (2025). Frontier AI trends report. UK Department for Science, Innovation and Technology. 

Amodei, D. (2026, September). We must pace the frontier

Anthropic. (2025). Reasoning models don’t always say what they think

Bostrom, N. (2019). The vulnerable world hypothesis. Global Policy, 10(4), 455–476. doi:10.1111/1758-5899.12718. 

Coalition for Epidemic Preparedness Innovations. (n.d.). The 100 Days Mission

Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., & Tian, Y. (2024). Training large language models to reason in a continuous latent space. arXiv:2412.06769. 

International AI Safety Report. (2026). International AI Safety Report 2026

METR. (2025, March 19). Measuring AI ability to complete long software tasks

National Academies of Sciences, Engineering, and Medicine. (2018). Biodefense in the age of synthetic biology. National Academies Press. doi:10.17226/24890. 

National Academies of Sciences, Engineering, and Medicine. (2025). The age of AI in the life sciences: Benefits and biosecurity considerations. National Academies Press. doi:10.17226/28868. 

OpenAI. (2025, November 20). Early experiments in accelerating science with GPT-5

OpenAI. (2026, September 1). Path to Astra: Critical capabilities and frontier safeguards

Pachocki, J. (2026, September 6). An alien mind. OpenAI. 

Reser, J. E. (2013). Artificial intelligence programmed to simulate mental continuity between processing states. Observed Impulse. 

Reser, J. E. (2019). Why we should embrace our superintelligent AI overlords. Observed Impulse. 

Reser, J. E. (2024). Generative interpretability: Pursuing AI safety through the visualization of internal processing states. Observed Impulse. 

Reser, J. E. (2025a). I expected that language models alone would never result in AGI. Observed Impulse. 

Reser, J. E. (2025b). Imagining AI’s attitude toward humanity: Butterflies, bees, ants, and cockroaches. Observed Impulse. 

World Health Organization. (2020, August 31). In WHO global pulse survey, 90% of countries report disruptions to essential health services since COVID-19 pandemic

Posted in

Leave a Reply

Discover more from Iterated Insights

Subscribe now to keep reading and get access to the full archive.

Continue reading