Paula JG Freund, Jared Edward Reser, and GPT 6
Abstract
Wikipedia and large language models are two of the most consequential knowledge technologies produced by the internet, and both initially provoked a similar objection: they could be wrong. Wikipedia was distrusted because anonymous volunteers could alter a public encyclopedia. Large language models were distrusted because a fluent generator could produce false claims, fabricated references, biased summaries, and confident explanations unsupported by evidence. Neither technology was simply rejected, however. Both were adopted rapidly by the public while schools, professions, and other institutions struggled to define acceptable use. This article argues that their histories are best understood as cases of behavioral acceptance preceding epistemic acceptance. People used them before they agreed on how much to trust them.
The comparison also reveals a decisive difference. A Wikipedia error is usually an error in a persistent public document. It has a location, revision history, discussion page, and potential correction that benefits later readers. A language-model hallucination is an error produced by a generative process. It may appear once, disappear, and never exist in precisely the same form again. Correcting the answer does not necessarily correct the generator. I call this difference error addressability. Wikipedia became institutionally useful because its fallibility was made inspectable and governable. LLMs will require stronger mechanisms for provenance, calibrated uncertainty, retrieval, persistent correction, and public audit.
The two systems also govern bias differently. Wikipedia externalizes disagreement through policies, citations, talk pages, and edit histories. LLMs internalize much of it within training data, model weights, human-feedback procedures, system instructions, and product decisions. The future of digital knowledge may depend on joining Wikipedia’s persistent, versioned, publicly contestable memory with the synthetic and conversational capacities of language models. Such a hybrid points toward the Final Library: an evidence-linked, continuously revised, machine-readable system that preserves not only conclusions, but also their sources, objections, failures, and conceptual genealogies.
Keywords: Wikipedia, large language models, hallucination, bias, crowdsourcing, epistemic governance, provenance, error correction, artificial intelligence, Final Library
1. Introduction: Two Technologies We Were Told Not to Use
An entire generation of students was told not to cite Wikipedia. A later generation was told not to use ChatGPT. Both instructions responded to real problems, and both were immediately undermined by utility. Wikipedia could provide an accessible orientation to almost any established topic within seconds. A large language model could explain, compare, translate, summarize, brainstorm, and draft in response to a natural-language request. People began using each technology before educational and professional institutions had decided what legitimate use should look like.

The parallel is tempting. Wikipedia was rejected because it was crowdsourced. Generative AI was rejected because it hallucinated. Wikipedia could contain a false statement, and an LLM could generate one. Both were digital, internet-dependent knowledge systems. Both challenged traditional ideas about authorship, expertise, authority, and intellectual labor. Both quickly became difficult to exclude from ordinary knowledge work.
Yet the parallel becomes useful only when its limits are taken seriously. Wikipedia is primarily a shared, persistent, collaboratively edited reference object. A language model is primarily a learned generative process that constructs a response for a particular context. Wikipedia publishes a page. An LLM performs an answer. The former invites users into a common artifact with a visible history; the latter commonly gives each user a private, transient, and differently phrased result. Their errors may look similar at the level of a sentence, but they occupy different technical and social architectures.
This article compares the development and acceptance of Wikipedia with the development and acceptance of large language models. Its central claim is that societies do not require a knowledge technology to be perfectly accurate before using it. They require ways to calibrate trust, identify appropriate roles, inspect provenance, contest bias, and correct errors. Wikipedia’s history shows that fallibility can be socially tolerated when it is made governable. The remaining challenge for generative AI is not simply to produce fewer errors, although that remains essential. It is to make errors more addressable.
The comparison also extends earlier work on iterative updating, artificial cognitive architecture, distributed insight synthesis, and the Final Library. My account of mental continuity described thought as a succession of partially overlapping representational states, with continuity preserved through incremental rather than total replacement (Reser, 2016, 2022a, 2022b). Wikipedia exhibits a public, documentary analogue of this principle: the page persists while portions of its content are repeatedly replaced. More recent work described AI-assisted writing as distributed insight synthesis and proposed a Final Library in which knowledge, hypotheses, evidence, failures, and conceptual lineages are persistently organized and revised (Reser, 2025, 2026b, 2026c, 2026d). Wikipedia and LLMs can be interpreted as complementary precursors to that architecture.
2. Wikipedia’s Improbable Proposition
Wikipedia launched on January 15, 2001, as an open companion to Nupedia, an expert-written encyclopedia with a slow review process. The wiki model inverted familiar assumptions about quality control. Instead of asking credentialed authors to pass through editorial gates before publication, it allowed publication first and correction afterward. Readers could become editors, and articles could change at any time. To people raised on signed encyclopedia entries, stable editions, and professional editorial boards, this sounded less like a reference system than an invitation to vandalism.
The early objections were not foolish. Anonymous and pseudonymous contributors could make mistakes, promote ideologies, embellish biographies, insert jokes, or fight over politically charged language. Articles varied greatly in quality. The identity and expertise of a contributor were often unclear. A printed encyclopedia could also be wrong, but its errors arrived wearing a suit. Wikipedia’s errors sometimes arrived under a username created five minutes earlier.
The underlying dispute concerned the location of authority. Traditional encyclopedias placed authority upstream, in the selection of authors and editors. Wikipedia placed much of it downstream, in open revision, public discussion, source requirements, and the possibility that many people would inspect the same page. Its wager was not that every contributor would be reliable. Its wager was that a sufficiently active community, operating under workable rules, could make a shared artifact more reliable over time (Reagle, 2010; Jemielniak, 2014).
Those rules mattered. Wikipedia developed three core content policies: neutral point of view, verifiability, and no original research. Significant viewpoints should be represented fairly and in proportion to their prominence; challenged claims should be attributable to reliable published sources; and editors should not use the encyclopedia to advance novel theories or syntheses of their own (Wikipedia contributors, 2026a). These principles did not remove conflict. They gave conflict a procedural vocabulary. Editors no longer had to settle the philosophical question of what was finally true before acting. They could ask whether a claim had an adequate source, whether a view received undue weight, and whether a proposed synthesis had appeared in the published literature.
This was a major epistemic innovation disguised as a website rulebook. Wikipedia shifted many disputes from personal authority to inspectable procedure. A professor and a teenager could disagree, but both were expected to point to sources. An editor could still misuse sources or apply policy selectively, yet the dispute occurred on a visible page and left a public record. The authority of the encyclopedia came to reside less in the perfection of individual contributors than in the revisability of the collective process.
3. Acceptance in Practice Before Acceptance in Principle
Wikipedia was neither universally rejected nor suddenly accepted. Its incorporation into public life was gradual, uneven, and role-specific. Readers discovered that it was extraordinarily useful for learning basic terminology, locating references, identifying names and dates, surveying unfamiliar debates, and deciding what to investigate next. Many teachers continued to prohibit it as a cited authority, but students and scholars still consulted it. Wikipedia was often accepted behaviorally before it was accepted epistemically.
The distinction matters. Behavioral acceptance occurs when a technology becomes part of ordinary practice. Epistemic acceptance occurs when institutions develop stable norms about what its outputs mean, when they can be trusted, and what verification they require. A person may use Wikipedia every day while refusing to cite it in a journal article. This is not necessarily hypocrisy. It is a form of role differentiation. Wikipedia can be appropriate as an orientation layer and inappropriate as the final authority for a contested claim.
Evidence about quality helped legitimate that limited role. In a widely discussed 2005 comparison, Nature asked experts to review matched scientific entries from Wikipedia and Encyclopaedia Britannica. Across the usable pairs, reviewers identified four serious errors in each source and more total inaccuracies, omissions, or misleading statements in Wikipedia than in Britannica, although the difference was far smaller than many observers expected (Giles, 2005). Britannica disputed the study’s design and interpretation (Encyclopaedia Britannica, 2006). Even with that dispute, the episode changed the public question. The issue was no longer whether an open encyclopedia could contain errors. All encyclopedias could. The more interesting question was how often, how seriously, and how correctably they erred.
Academic resistance remained visible. In 2007, the history department at Middlebury College prohibited students from citing Wikipedia in academic work after faculty encountered inaccurate claims in papers (Jaschik, 2007). The incident is often remembered as a blanket rejection of Wikipedia, but it more precisely concerned citation and scholarly responsibility. A later survey at two Spanish universities found four faculty profiles ranging from averse and reluctant to open and proactive. The open and proactive groups together outnumbered the strictly skeptical groups, while perceptions of quality, usefulness, visibility, and academic culture predicted adoption more strongly than simple demographic stereotypes (Minguillón et al., 2018).
Wikipedia’s eventual acceptance therefore did not amount to a declaration that the site was always correct. It became a piece of epistemic infrastructure. Search engines surfaced it. Journalists consulted it. Teachers designed editing assignments. Experts quietly repaired pages in their fields. Readers learned an informal protocol: begin there, inspect the citations, check the history when controversy matters, and verify consequential claims elsewhere. The technology matured partly because society developed a more granular concept of trust.
Large language models are following a compressed version of this trajectory. ChatGPT’s public release in late 2022 was followed by rapid adoption and equally rapid institutional anxiety. Schools worried about cheating and the erosion of writing practice. Researchers encountered fabricated citations. Lawyers submitted nonexistent cases. Users discovered political, cultural, and demographic biases. Organizations raised concerns about confidentiality, copyright, deskilling, employment, and automation. UNESCO’s guidance on generative AI in education and research emphasized human agency, validation, privacy, and institutionally defined uses rather than uncritical adoption (UNESCO, 2023).
At the same time, people used the systems because they were helpful. LLMs became tutors, coding partners, translators, drafting assistants, search intermediaries, and conversational interfaces to complex material. As with Wikipedia, public use did not wait for a final epistemology. The systems were too useful to remain outside practice and too unreliable to enter practice without qualification. AI, like Wikipedia, was accepted through verbs before it was accepted through doctrines. People asked it, revised it, checked it, copied it, argued with it, and built workflows around it while institutions were still composing policies.
4. Two Kinds of Fallibility
The most important contrast between Wikipedia and LLMs concerns the location of error. A Wikipedia article is a persistent knowledge object. If it contains an incorrect date, the mistake can be located in a particular sentence and revision. Editors can identify when it appeared, who added it, which source was offered, how others responded, and whether the statement was later removed. The error may be socially harmful and may persist for years, but it has an address.
An LLM hallucination is usually different. It is an output event produced by a probabilistic generator in response to a particular prompt, conversational history, system configuration, retrieval context, and sampling path. The false sentence may never have existed in the training data. It may be an improvised combination of true fragments, an incorrect inference, or a plausible citation assembled from familiar bibliographic patterns. Another user can ask the same question and receive a different mistake, a correct answer, or a refusal.
This difference can be described as error addressability, the degree to which a false or misleading claim can be durably located, inspected, attributed, contested, and corrected for future users. Wikipedia generally offers high error addressability. Its pages have stable identifiers, revisions, diffs, citations, talk pages, edit summaries, watchlists, and rollback tools. A conventional ungrounded chatbot response generally offers low error addressability. It may be logged, but its claim does not necessarily belong to a shared public object, and correcting it in one conversation rarely changes the model for everyone else.
The practical consequence is simple. Wikipedia corrects the knowledge object. AI developers must correct the generator, its retrieval environment, its instructions, or the downstream record in which its answer is stored. A page can be repaired with a scalpel. A generator may require retraining, fine-tuning, a system-level rule, a safety intervention, a retrieval correction, or a change in evaluation incentives. Because model behavior is distributed across many parameters and contexts, a local fix can have distant effects, and the same error can reappear in a new form.
Hallucination is therefore not merely Wikipedia-style inaccuracy at higher speed. It is a process-level failure. Language models are trained to predict and generate linguistically appropriate continuations, not to retrieve a verified proposition from a canonical ledger every time they speak. Human-feedback training can make them more useful and better aligned with instructions (Ouyang et al., 2022), but fluency and truth remain separable. OpenAI’s analysis of hallucination emphasizes that common evaluations can reward guessing over calibrated uncertainty, thereby encouraging models to answer when abstention would be more reliable (OpenAI, 2025). NIST accordingly treats confabulation as a central generative-AI risk requiring measurement and management (Autio et al., 2024).
Wikipedia has a memorable badge for incompleteness: [citation needed]. A language model can reproduce the phrase, but it does not automatically possess the epistemic discipline the phrase represents. The deepest requirement for trustworthy generative AI may be the functional equivalent of that tag: explicit links between claims and evidence, visible confidence states, records of disagreement, and the ability to say that a question remains unresolved. The lesson from Wikipedia is not that error does not matter. The lesson is that error needs an address.
5. Crowdsourcing Did Not Disappear. It Became Compressed
The usual contrast describes Wikipedia as crowdsourced and LLMs as machine-generated. This is true at the interface and misleading underneath it. Wikipedia visibly aggregates the labor of volunteer editors. LLMs compress contributions from a much larger and less visible crowd: authors of books and websites, Wikipedia editors, programmers, forum participants, translators, annotators, evaluators, red teams, product designers, and users whose feedback shapes later systems. The crowd did not disappear. It became sedimented in data, weights, policies, and feedback procedures.
Wikipedia’s contributors act directly on a shared representation. An editor changes a sentence, and the public page changes. In an LLM, human influence is statistically mediated. Training transforms vast collections of human-produced symbols into distributed parameters. Reinforcement learning from human feedback further shapes which responses are preferred, while system instructions and safety policies constrain behavior at deployment (Ouyang et al., 2022). The resulting answer has no simple one-to-one author. It is a new composition produced from a network of learned regularities and current instructions.
This makes an LLM a form of compressed crowdsourcing. The phrase does not imply that every training source is represented faithfully or that model weights are a searchable archive. It identifies the social origin of the patterns the model has learned. An LLM can sound like an individual mind because it serializes many distributed influences through one conversational voice. That unity is useful, but it also obscures disagreement. Wikipedia may display an edit war. An LLM often gives the linguistic appearance that the war has already been settled.
The visibility of labor differs as well. Wikipedia exposes usernames, edit histories, discussion archives, and community roles. LLM development typically exposes far less about individual training examples, annotation decisions, filtering criteria, and alignment interventions. Model cards and system documentation can improve transparency (Mitchell et al., 2019), but they do not reproduce the proposition-level genealogy available in a mature wiki. The issue is not only whether a crowd contributed. It is whether users can inspect how the crowd’s conflicts were transformed into the answer they received.
6. Bias as a Governance Problem
Both Wikipedia and LLMs inherit bias from human culture, but they organize it differently. Wikipedia’s bias is relatively externalized. It appears in who edits, which topics receive attention, which sources count as reliable, how viewpoints are weighted, and how persistent participants influence consensus. Its neutral-point-of-view policy does not promise a view from nowhere. It requires significant published positions to be represented fairly and proportionately (Wikipedia contributors, 2026a).
This procedural neutrality has strengths. Editors can dispute wording in public, compare sources, attach warning templates, request additional viewpoints, and revisit earlier decisions. Studies of political articles suggest that Wikipedia’s bias can change as contributions accumulate, although collective editing does not guarantee neutrality and may perform differently across topics (Greenstein & Zhu, 2012, 2018). The revision record makes at least part of the struggle observable.
Wikipedia also reproduces structural inequalities in participation and source availability. The Wikimedia Foundation reports that its contributor population is disproportionately male and geographically uneven, with only a small fraction of editors based in Africa (Wikimedia Foundation, n.d.). Underrepresentation affects which biographies are written, which languages flourish, which histories receive detail, and which gaps are noticed. Projects such as Women in Red respond by deliberately expanding neglected coverage, illustrating that bias mitigation is not a single neutralizing operation. It is sustained institutional work.
LLM bias is more internalized and layered. It can enter through the distribution of training text, decisions about what data to include, frequency patterns within that data, annotation guidelines, reward models, safety policies, system prompts, retrieval sources, and the wording of the user’s request. A model can reproduce stereotypes, privilege highly represented languages and cultures, flatten minority positions, or present contested judgments as settled facts (Bender et al., 2021; Bommasani et al., 2021). Alignment procedures can reduce some harmful behaviors while introducing new tradeoffs about refusal, deference, political framing, and whose preferences define acceptable output.
The two systems share a deeper limitation: neither can fully escape the ecology of recorded knowledge. Wikipedia’s verifiability policy means that a poorly documented community may remain poorly represented even when its members know that the article is incomplete. LLMs likewise learn most easily from what has been digitized, published, repeated, and made accessible. Knowledge that was never recorded cannot simply be recovered from statistical patterns. Missing archives become missing representation, and repeated errors can acquire the appearance of consensus.
This is why bias cannot be handled only as a property of output sentences. It is a governance problem involving participation, documentation, source selection, dispute procedures, measurement, and correction. Wikipedia’s great contribution was not the elimination of bias. It was the construction of public machinery for arguing about bias. LLM systems need comparable machinery at the level of claims, datasets, evaluations, retrieval pipelines, and product policy.
7. How the Two Systems Are Updated
Wikipedia and LLMs are both digital and internet-connected, but they inhabit different temporal regimes. Wikipedia supports continuous, proposition-level updating. An editor can change one date, add one source, revert one act of vandalism, or restructure an entire page. The updated page becomes visible immediately, while the old version remains accessible. Change is local, public, and versioned.
Language-model systems have at least four update layers, and they should not be confused. First, conversation context can update the model’s current behavior temporarily. A user can correct a name or define a term, and the system may follow that correction for the rest of the exchange. Second, retrieval systems can supply recent documents, web pages, databases, or organizational records at response time. Retrieval-augmented generation separates some factual updating from the slower modification of model parameters (Lewis et al., 2020). Third, developers can change system instructions, tools, filters, and product policies relatively quickly. Fourth, the underlying model weights are updated through continued training, fine-tuning, or a new model release, usually in batches rather than one public claim at a time.
These layers produce very different meanings of “the AI knows.” A model may have an obsolete parametric association but retrieve a current source. It may know a correction within one conversation and forget it in the next. A product may block a failure through instructions without changing the underlying model. A new checkpoint may improve one class of answers while degrading another. There is no single edit history equivalent to Wikipedia’s page history because behavior emerges from the interaction of model, prompt, context, retrieval, tools, and sampling.
The contrast can be summarized as follows:
|Dimension |Wikipedia |Large language model system |
|—————-|——————————————|————————————————————————–|
|Primary object |Shared public page |Generated response |
|Main update unit|Claim, sentence, section, or page |Context, retrieved evidence, instructions, fine-tuning, or weights |
|Update timing |Continuous and often immediate |Temporary, real-time through retrieval, or periodic through model releases|
|History |Public revision log and diffs |Often private logs, release notes, and incomplete behavioral documentation|
|Correction scope|Usually benefits later readers of the page|Often limited to one session unless incorporated upstream |
|Rollback |Direct restoration of an earlier revision |Difficult because behavior is distributed and context-dependent |
|Provenance |Page-level and often claim-level citations|Variable; may range from explicit citations to no visible source trail |
This table also explains why “just correct the AI” is an underspecified instruction. A correction can target the transient conversation, the retrieved source, the orchestration layer, the post-training procedure, or the base model. Each target has a different cost and different side effects. Wikipedia taught users to think in revisions. Generative AI requires them to think in layers.
8. A Shared Page and a Private Interlocutor
Wikipedia presents different readers with substantially the same public page at a given moment. Its common object supports collective scrutiny. If an article gives undue weight to a political claim, readers and editors can refer to the same paragraph, compare revisions, and debate a proposed change. Public knowledge remains disputable, but the object of dispute is shared.
An LLM normally produces individualized outputs. Two users may receive different facts, levels of caution, analogies, examples, or conclusions. Personalization can improve teaching and accessibility, yet it can also fragment epistemic experience. The error shown to one user may never be shown to another. The bias introduced by a particular prompt may remain invisible outside the conversation. There may be no common paragraph around which a correction community can form.
The interface amplifies this difference. Wikipedia looks like a document. A chatbot looks like an interlocutor. Conversational language encourages people to attribute understanding, confidence, intention, and social awareness to a system whose internal operation is unlike ordinary human authorship. A polished answer can feel more authoritative than a cluttered Wikipedia page precisely because the negotiation, uncertainty, and citation disputes have been hidden. Ease of use can therefore increase both utility and epistemic risk.
The solution is not to make every answer resemble a talk-page argument. It is to preserve access to the layers beneath the smooth response. Users should be able to inspect sources, identify which statements are inferred rather than retrieved, see meaningful uncertainty, compare alternative interpretations, and contribute corrections that can reach a persistent knowledge layer. The conversational surface should function as an interface to governed knowledge, not as a substitute for it.
9. Wikipedia’s Productive Limitation: No Original Research
Wikipedia became more reliable partly by refusing a task that LLMs are increasingly asked to perform. Its no-original-research policy prohibits editors from using the encyclopedia to publish new theories or syntheses (Wikipedia contributors, 2026a). Wikipedia is designed to summarize established, attributable knowledge. It can describe a scientific frontier, but it should not move that frontier by itself.
Generative AI is routinely invited to cross this boundary. Users ask models to infer, theorize, propose mechanisms, combine disciplines, design experiments, coin terminology, and search for hypotheses that have not been published. This synthetic capacity is one of the technology’s greatest promises. It is also one reason Wikipedia’s governance model cannot simply be copied. A rule requiring every generative claim to have appeared previously in a source would eliminate much of the value of generative reasoning.
The distinction is central to the problem of AI slop. Fluent recombination can resemble scientific thinking without providing independent evidence, novelty checks, discriminating predictions, or serious criticism. Earlier work argued that the road from generative text to a Final Library must pass through provenance, conceptual testing, and explicit differentiation among established findings, plausible hypotheses, and unsupported speculation (Reser, 2026a, 2026b). A generative knowledge system must be allowed to create candidates while being prevented from silently promoting those candidates into facts.
Recursive conceptual prospecting offers one possible architecture. An AI system can generate semirandom conceptual combinations, investigate promising relationships, search prior art, construct mechanisms, invite adversarial criticism, propose predictions, and preserve both successful and failed branches in a synthetic frontier corpus (Reser, 2026d). In such a system, creativity occurs upstream of validation. A hypothesis can be valuable without being true, provided the system labels it correctly and preserves the route by which it might be tested.
Wikipedia’s policy boundary therefore suggests a useful separation of epistemic spaces. One layer represents established, source-grounded knowledge. Another represents unresolved claims, speculative syntheses, and research proposals. A third records tests, criticisms, failures, and changes in confidence. LLMs can move among these layers conversationally, but they should not collapse them. The user should know whether the system is reporting, inferring, or inventing.
10. The Recursive Loop Between Wikipedia and AI
Wikipedia and LLMs are not merely parallel technologies. They now participate in the same recursive information ecosystem. Wikipedia has been an important source of relatively structured, multilingual, publicly licensed text for model training (Bommasani et al., 2021). Language models then generate explanations, articles, summaries, and synthetic prose that enter the wider web. Some of that material may be copied into future reference works, datasets, or even Wikipedia itself. Later models can train on a corpus already altered by earlier models.
The irony is unusually clean. Wikipedia became one of AI’s teachers, and now Wikipedia is trying to stop its student from writing the textbook. As of 2026, English Wikipedia’s guideline generally prohibits using LLMs to generate or rewrite article content, with narrow allowances for reviewed copyediting and translation. The rationale is that LLM-generated text often violates core policies and can be produced faster than other editors can feasibly review it (Wikipedia contributors, 2026b).
This policy addresses a scale asymmetry. Human editors can generate mistakes, but automated systems can generate plausible mistakes at industrial speed. Open collaboration depends on the reviewing capacity of the community. If production becomes much cheaper than verification, the correction system can be overwhelmed. The problem is familiar in cybersecurity, spam, and scientific publishing: a filter that works at one volume may fail when the cost of producing candidate material approaches zero.
The recursive loop also creates a risk of epistemic recycling. A model generates a false but plausible claim. The claim appears on websites, in reports, or in machine-generated references. Later retrieval systems encounter multiple copies and interpret repetition as corroboration. Future models absorb the pattern during training. The claim can then return with greater fluency and apparent cultural support. This process might be called synthetic cultural inbreeding: a knowledge ecosystem repeatedly trains on its own unverified descendants.
Avoiding this outcome requires provenance that survives transformation. Systems should distinguish human-authored evidence from machine-generated summaries, primary findings from derivative repetition, and independent corroboration from copied text. They should preserve negative results and retractions rather than allowing the web’s most repeated formulations to dominate. Without such measures, generative AI could increase the quantity of accessible prose while decreasing the genetic diversity of the evidence beneath it.
11. From Wikipedia and LLMs to the Final Library
Wikipedia can be understood as a collective external memory. It stores a public representation of what sources and editors currently support, together with a history of how that representation changed. LLMs can be understood as engines of linguistic and conceptual synthesis. They can reorganize knowledge for a user, connect distant domains, simulate objections, generate hypotheses, and serialize complex material into coherent explanations. Each system possesses what the other lacks.
Wikipedia has persistence, public addressability, revision history, and citation norms, but limited permission to generate original theory. LLMs have flexibility, synthesis, personalization, and generative reach, but weak default provenance and no universal public correction object. The obvious next step is not the victory of one system over the other. It is their architectural integration.
The Final Library was proposed as a cumulative repository in which machine-generated insights are expanded, tested, cross-referenced, and organized into structures too extensive for direct unaided human navigation (Reser, 2025). Later work developed two relevant components. Distributed insight synthesis describes how transient ideas can be preserved as cognitive checkpoints, elaborated through dialogue, and serialized into larger theoretical structures (Reser, 2026c). Recursive conceptual prospecting describes how AI systems could systematically search underexplored conceptual space while retaining hypotheses, evidence maps, rejected branches, terminology, confidence states, and unresolved questions (Reser, 2026d).
Wikipedia supplies a partial institutional ancestor for this proposal. The Final Library should keep receipts as carefully as Wikipedia does, but at finer granularity and greater scale. Each claim should have a stable identifier, provenance record, version history, epistemic status, supporting and opposing evidence, known replications, conceptual dependencies, and unresolved challenges. Generated syntheses should link back to the claims from which they were constructed. Corrections should propagate through dependent summaries without erasing earlier states.
The library should also retain the forms of uncertainty that conventional encyclopedias exclude. A rejected hypothesis can prevent repeated travel down the same dead end. A null result can correct publication bias. A minority interpretation can be preserved without receiving the same weight as a well-replicated finding. A conceptual genealogy can show that two apparently independent theories share an origin. A model could then answer differently when asked what is known, what is disputed, what has failed, and what remains worth testing.
Such a system would combine three modes of updating. Wikipedia-style editing would support local, public, proposition-level correction. Retrieval would provide immediate access to new evidence without waiting for a new base model. Periodic model training would improve the generator’s broader capacities. The persistent library would remain the public epistemic substrate, while the model would act as interface, analyst, critic, and hypothesis generator.
This architecture also fits the iterative-updating model of cognition. Mental continuity does not require a static representational state; it requires sufficient overlap across successive states for prior structure to constrain what comes next (Reser, 2016). A mature digital knowledge system can operate similarly. It need not freeze conclusions, and it should not regenerate its worldview from nothing for every query. It should preserve a continuously revised state in which new evidence modifies part of the structure while provenance and conceptual continuity remain intact.
12. A Governance Model for Fallible Intelligence
The history of Wikipedia suggests that acceptance depends less on a claim of perfection than on the construction of calibrated trust. People learned what Wikipedia was good for, where it was vulnerable, how to inspect it, and when another source was necessary. The same development is underway for LLMs, but conversational generation requires additional safeguards.
A governable generative knowledge system should include at least six properties. First, it should provide claim-level provenance when factual accuracy matters. Second, it should distinguish retrieved evidence from model inference and speculative generation. Third, it should express uncertainty in ways calibrated to empirical performance rather than rhetorical style. Fourth, it should maintain persistent correction channels so that identified failures can benefit future users. Fifth, it should expose meaningful records of policy and model change. Sixth, it should separate exploratory creativity from the validated knowledge layer.
These requirements do not imply that every casual conversation needs a scholarly apparatus. Wikipedia itself uses stricter protections for biographies of living people and other high-risk domains. Generative systems can also vary their epistemic friction by context. A request for fictional names can prioritize creativity. A medical, legal, scientific, or financial claim should trigger stronger sourcing, uncertainty, and verification. Acceptance becomes possible when the system’s confidence, interface, and governance match the stakes.
The comparison also argues against a simplistic human-versus-machine frame. Wikipedia succeeded through interaction among humans, software, bots, policies, interfaces, and source institutions. LLMs are likewise sociotechnical systems. Their reliability depends on model architecture, training data, evaluators, retrieval tools, organizational incentives, user behavior, and the surrounding information environment. Blaming or praising “the AI” can conceal the many decisions through which its behavior is produced.
The most durable norm may be correctability before certainty. Perfect knowledge is unavailable to human authors, encyclopedias, scientific institutions, and artificial systems. The practical question is whether a claim can be traced, challenged, revised, and prevented from silently reproducing itself. A fallible system with excellent correction architecture may deserve more trust than a highly accurate system whose occasional errors are invisible, untraceable, and difficult to remove.
13. Conclusion
Wikipedia and large language models disrupted knowledge culture for related reasons. Both weakened the traditional connection between an answer and a clearly credentialed individual author. Both offered extraordinary utility at internet scale. Both could be wrong, and both were adopted before institutions knew how to describe their proper use. Their early histories therefore reveal the same general pattern: behavioral acceptance came before epistemic acceptance.
Wikipedia’s path to legitimacy did not require the crowd to become infallible. It required the crowd’s activity to become structured. Neutrality, verifiability, source citation, revision history, discussion, and role-specific norms transformed an improbable experiment into global infrastructure. The system earned trust by exposing enough of its own fallibility for users and editors to manage it.
LLMs face a harder version of the problem. Their mistakes are generated dynamically, their provenance is often unclear, their corrections may remain local, and their smooth conversational form can hide disagreement. A Wikipedia error sits on a page. An LLM error may be an ephemeral performance of the generator. This is why lower hallucination rates, although necessary, are not sufficient. Generative AI needs error addressability.
The likely destination is a hybrid. Wikipedia contributes the logic of persistent public memory, versioned claims, and visible correction. LLMs contribute synthesis, dialogue, personalization, translation, criticism, and hypothesis generation. Joined through strong provenance and explicit epistemic labeling, these capacities could support the Final Library: a continuously updated knowledge architecture that remembers not only what is believed, but why it is believed, what opposes it, how it changed, and what remains unknown.
Wikipedia taught the internet that a reference work could remain unfinished and still become indispensable. Large language models may teach the next lesson: an intelligence can remain fallible and still become trustworthy, but only if its fallibility is designed for inspection, correction, and cumulative learning.
References
Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., & Roberts, K. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.600-1
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). Association for Computing Machinery. https://doi.org/10.1145/3442188.3445922
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., … Liang, P. (2021). On the opportunities and risks of foundation models. arXiv. https://doi.org/10.48550/arXiv.2108.07258
Encyclopaedia Britannica. (2006). Fatally flawed: Refuting the recent study on encyclopedic accuracy by the journal Nature. https://corporate.britannica.com/britannica_nature_response.pdf
Giles, J. (2005). Internet encyclopaedias go head to head. Nature, 438, 900–901. https://doi.org/10.1038/438900a
Greenstein, S., & Zhu, F. (2012). Is Wikipedia biased? American Economic Review, 102(3), 343–348. https://doi.org/10.1257/aer.102.3.343
Greenstein, S., & Zhu, F. (2018). Do experts or crowd-based models produce more bias? Evidence from Encyclopædia Britannica and Wikipedia. MIS Quarterly, 42(3), 945–959. https://doi.org/10.25300/MISQ/2018/14084
Jaschik, S. (2007, January 26). A stand against Wikipedia. Inside Higher Ed. https://www.insidehighered.com/news/2007/01/26/stand-against-wikipedia
Jemielniak, D. (2014). Common knowledge? An ethnography of Wikipedia. Stanford University Press.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems (Vol. 33, pp. 9459–9474). https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html
Minguillón, J., Aibar, E., Lerga, M., Lladós, J., & Meseguer-Artola, A. (2018). Wikipedia in academia as a teaching tool: From averse to proactive faculty profiles. arXiv. https://doi.org/10.48550/arXiv.1801.07138
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220–229). Association for Computing Machinery. https://doi.org/10.1145/3287560.3287596
OpenAI. (2025, September 5). Why language models hallucinate. https://openai.com/index/why-language-models-hallucinate/
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (Vol. 35, pp. 27730–27744). https://arxiv.org/abs/2203.02155
Reagle, J. M., Jr. (2010). Good faith collaboration: The culture of Wikipedia. MIT Press.
Reser, J. E. (2016). Incremental change in the set of coactive cortical assemblies enables mental continuity. Physiology & Behavior, 167, 222–237. https://doi.org/10.1016/j.physbeh.2016.09.019
Reser, J. E. (2022a). A cognitive architecture for machine consciousness and artificial superintelligence: Thought is structured by the iterative updating of working memory. arXiv. https://doi.org/10.48550/arXiv.2203.17255
Reser, J. E. (2022b). Artificial intelligence software structured to simulate human working memory, mental imagery, and mental continuity. arXiv. https://doi.org/10.48550/arXiv.2204.05138
Reser, J. E. (2025, December 2). The Final Library and the last years of human-original ideas. Iterated Insights. https://iteratedinsights.com/2025/12/02/the-final-library-and-the-last-years-of-human-original-thought/
Reser, J. E. (2026a, March 26). AI slop, scientific thinking, and the road to the Final Library. Observed Impulse. https://www.observedimpulse.com/2026/03/ai-slop-scientific-thinking-and-road-to.html
Reser, J. E. (2026b). From peer review to the Final Library: The evolution of scientific validation in the age of superintelligence [Manuscript in preparation].
Reser, J. E. (2026c, September 10). My writing process: Notes, cognitive checkpoints, distributed insight synthesis, and the serialization of thought. Iterated Insights. https://iteratedinsights.com/2026/09/10/my-writing-process-notes-cognitive-checkpoints-distributed-insight-synthesis-and-the-serialization-of-thought/
Reser, J. E. (2026d, August 11). Recursive conceptual prospecting: Mining the latent scientific frontier with artificial intelligence. Iterated Insights. https://iteratedinsights.com/2026/08/11/recursive-conceptual-prospecting-mining-the-latent-scientific-frontier-with-artificial-intelligence/
UNESCO. (2023). Guidance for generative AI in education and research. https://unesdoc.unesco.org/ark:/48223/pf0000386693
Wikimedia Foundation. (n.d.). Change the stats. Retrieved September 14, 2026, from https://wikimediafoundation.org/what-we-do/open-the-knowledge/otk-change-the-stats/
Wikipedia contributors. (2026a). Wikipedia: Core content policies. In Wikipedia. Retrieved September 14, 2026, from https://en.wikipedia.org/wiki/Wikipedia:Core_content_policies
Wikipedia contributors. (2026b). Wikipedia: Writing articles with large language models. In Wikipedia. Retrieved September 14, 2026, from https://en.wikipedia.org/wiki/Wikipedia:Writing_articles_with_large_language_models

Leave a Reply