LEVIATHAN.LIFE
All foundations

C12 · Discovery and self-improvement hypothesis

How could Leviathan improve its own ways of learning?

Statement assessed

Leviathan should be able to generate and test new questions, concepts and methods, then use what it learns to revise how future learning happens.

Status of this statement: Research ambition with testable parts

This status applies to the statement above, within the limits discussed in this note. Evidence for a finding can motivate Leviathan’s proposed mechanisms without establishing that they work.

Working editorial note · Last recorded revision

What we mean

A research shadow
A retained gap, contradiction or unexplained pattern that may guide further inquiry. This research use is distinct from other meanings of shadow in Leviathan’s documents.
Method learning
Acquiring or developing a reusable way to ask, measure, compare, build or decide, with evidence about where it works.
Recursive self-improvement
Improvement of capabilities or processes that themselves help produce later improvements. The ambition includes this feedback; each claimed improvement still needs its own evaluation.

Values we bring to the question

  • Make room for original questions and unexpected connections.
  • Preserve failed attempts and unresolved gaps when they can teach something.
  • Let other participants examine improvements and choose whether to adopt them.

Reasoning and the proposed connection

  1. A growing archive becomes more useful when it helps generate a question, prediction, or method that was previously missing. Connections should be judged partly by what they let us discover or do.
  2. A shadow can direct attention to cases our current representation treats alike even though they may differ. New observations, instruments, or concepts could make the distinction learnable.
  3. A method from another field may provide a way forward. Creative analogy proposes the connection; domain knowledge and testing determine which assumptions travel with it. The analogy should yield something to examine, not just similar words.
  4. A useful result or a concrete failure could change a concept, tool, or working procedure. In our local review exchange, an omitted decision led to an explicit requirement to include the exact decision in later review packets. This records a procedure revision; whether it improves later work remains to be tested.
  5. The ambition extends to improving the methods that produce further learning. Another Leviathan could find that a revision works only locally, improve it further, or reject it. Transfer, disagreement, and continued inquiry are part of that research.
  6. A useful candidate-generation process may still narrow the range of ideas considered. We propose measuring both the quality of an individual result and the diversity of approaches kept available.

Where the reasoning stops

The proposed loop is broader than the mechanisms demonstrated in any cited study. Better records, changed agent software and altered model representations are different changes. A new version or a higher score alone does not show improved discovery.

The strongest objection

The system could reward its own preferred questions and make its evaluations easier to satisfy. Apparent progress may come from extra compute, hidden expert work or memorized cases. Shared revisions can also spread a blind spot more efficiently.

Read the evidence

Each source has a specific role in the stated claim. Its findings, review date, and access limits are recorded below. Our proposed architecture and experiments require their own tests; an editorial revision does not mean the source was reviewed again.

R-LRN-01 · Gives a reason to investigate

Claude-shaped science

Matthew Schwartz; guest research account published by Anthropic · Research account

Read: Selected sections: cross-field methods, expert redirection, workflow and limitations · Selected primary-source sections reviewed; underlying projects not reproduced

Published: 2026-10-01 · Reviewed: 2026-10-03

BootLoops offers a first-person account of scientific methods crossing fields with expert guidance. It motivates creative transfer; it does not establish autonomous general discovery.

What it reports
Schwartz describes building reusable computational tools with Claude and applying methods across scientific fields. Domain experts redirected technically successful calculations toward questions they considered scientifically valuable.
Limits
This is a participant's account, not an independent replication of every reported result. The workflow required substantial human direction and resources. Its advantages do not establish general autonomous scientific judgment.
Review scope and version

Read: Selected sections: cross-field methods, expert redirection, workflow and limitations.

First-person account published 1 October 2026; describes work over the preceding summer and approximately three months

Motivates testing whether a method learned in one setting becomes useful elsewhere, and separating computational correctness from the importance of the question being answered.

Link to this source note

R-LRN-02 · Supports part of this claim

Large language models as uncertainty-calibrated optimizers for experimental discovery

Bojana Ranković, Ryan-Rhys Griffiths and Philippe Schwaller · Journal article

Read: Abstract and selected main-text methods and benchmark sections; preprint version history · Selected primary-source sections reviewed; results not reproduced

Published: 2026-08-28 · Reviewed: 2026-10-03

GOLLuM adapts representations using observed outcomes and guides candidate selection through uncertainty. This is a concrete learning mechanism within a narrower optimization setting.

What it reports
GOLLuM couples a language encoder with a Gaussian process. Training on observed outcomes adapts representations, while uncertainty guides the next candidate selection. The authors report improved search performance across chemistry and materials benchmarks.
Limits
The evidence concerns benchmark optimization under specified budgets and candidate spaces. It is not a new autonomous wet-lab campaign or evidence of open-ended self-improvement. Supplementary methods and code were not audited here.
Review scope and version

Read: Abstract and selected main-text methods and benchmark sections; preprint version history.

Journal publication: 28 August 2026; preprint first posted 8 April 2025, revised to v3 on 7 November 2025

Provides a concrete example of outcome-driven representation learning. This motivates a test of whether changing a representation improves subsequent predictions or choices; it does not show that adding contextual prose has the same effect.

Related source links

Link to this source note

R-LRN-04 · Adds context

Predictive Representations of State

Michael L. Littman, Richard S. Sutton and Satinder Singh · Conference paper

Read: Abstract and introductory formulation; proceedings metadata · Primary abstract and introductory formulation reviewed; proofs not independently checked

Published: Date not verified · Reviewed: 2026-10-03

Predictive state representations describe state through predictions of future observations under actions. Relating this to research shadows and new distinctions is our proposed connection.

What it reports
The paper represents the state of a controlled dynamical system using predictions of future observations conditioned on action sequences. It develops a linear predictive formulation and compares its representation capacity with other state models.
Limits
This is foundational representation theory, not an LLM or value-learning study. Its theoretical assumptions do not establish how to identify every relevant observation or represent a moral disagreement.
Review scope and version

Read: Abstract and introductory formulation; proceedings metadata.

NIPS 2001, Advances in Neural Information Processing Systems 14; exact publication day not established

Offers a conceptual connection for asking whether a new observation distinguishes situations that an existing representation treats alike. Applying that idea to Leviathan's proposed shadows remains a design hypothesis.

Link to this source note

R-AG-04 · Supports part of this claim

Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents

Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange and Jeff Clune · Preprint

Read: Selected sections · Primary source reviewed

Published: 2025-05-29 · Reviewed: 2026-09-30

Bounded self-modification of agent software offers one example of evaluating revisions. Fixed underlying weights and specific coding tasks limit the inference to broader recursive improvement.

What it reports
An agent modifies its own software and selects changes using coding evaluations, improving on two benchmarks. An archive of different past solutions helps the search.
Limits
The underlying model weights stay fixed. Bounded coding experiments do not establish open-ended improvement of model training, unlimited recursive improvement, or an AGI timetable.
Review scope and version

Read: Selected sections.

arXiv v3: 12 March 2026

Read alongside the METR cost framework: measured gains, total expenditure and independent evaluation answer different questions.

Link to this source note

R-AG-05 · Limits the inference

Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT

Tom Cunningham, Manish Shetty, Vincent Cheng and Nate Rush; METR · Research report

Read: Selected sections · Primary source reviewed

Published: 2026-07-21 · Reviewed: 2026-09-30

Comparing optimization at equal expenditure highlights the need to count compute and human effort when judging whether a learning method improved.

What it reports
The report proposes comparing human and agent optimization at equal expenditure, illustrated with NanoGPT. It counts experimental compute and human effort alongside model usage.
Limits
The human comparison is estimated, the agent results are preliminary, and the study covers one optimization problem. It does not directly measure the returns from human–AI collaboration.
Review scope and version

Read: Selected sections.

Research report

More generated code or a higher benchmark score alone does not establish faster research or economic advantage.

Link to this source note

R-AG-09 · Supports part of this claim

AlphaEvolve: A coding agent for scientific and algorithmic discovery

Alexander Novikov and co-authors · Research white paper

Read: Selected sections · Primary source reviewed; no independent reproduction

Published: 2025-06-16 · Reviewed: 2026-10-04

Evolving candidate programs with an evaluator is one implemented discovery mechanism. Choosing reliable evaluators and worthwhile questions remains necessary.

What it reports
An evolutionary pipeline combines LLM-generated code with supplied evaluators, including evolving search algorithms. Reported results include a 48-multiplication algorithm for two 4-by-4 complex matrices and optimizations to Google's computational infrastructure.
Limits
The main boundary is availability of automated evaluators. The paper describes self-improvement feedback as modest and occurring over months. It does not demonstrate unrestricted recursive improvement, ethical goal selection or Leviathan's architecture. Results and deployments were not independently reproduced here.
Review scope and version

Read: Selected sections.

arXiv:2506.13131v1; white paper submitted 16 June 2025

Read v1 task definition, code search, evaluation, selected results and discussion. Reported code and deployments were not reproduced; supplied evaluators constrain the demonstrated discovery process.

Link to this source note

R-LRN-06 · Limits the inference

Generative AI enhances individual creativity but reduces the collective diversity of novel content

Anil R. Doshi and Oliver P. Hauser · Peer-reviewed journal article

Read: Selected sections · Published primary text reviewed through UCL repository; no independent reproduction

Published: 2024-07-12 · Reviewed: 2026-10-04

The story-writing experiment separates individual benefit from collective diversity. Whether a similar tradeoff occurs in our method exchange remains a question.

What it reports
In a randomized study of 293 short-story writers, access to GPT-4 ideas improved assessed novelty and usefulness, especially for lower-scoring writers. AI-assisted stories were more similar to other stories within their condition under an embedding-based measure.
Limits
The task used eight-sentence stories, fixed prompts and no interactive dialogue. It did not study professional writers, scientific collaboration or independent Leviathans. Story similarity is not a direct measure of diversity in scientific explanations; transfer is a research hypothesis. Data and code were not rerun.
Review scope and version

Read: Selected sections.

Science Advances 10(28), eadn5290; published version, 12 July 2024

Read the UCL published-version PDF, pages 1–6. Publisher and PMC direct access failed; supplementary analyses, data and code were not reviewed or rerun.

Link to this source note

What could change our view?

If gains vanish on unfamiliar tasks, under an unchanged evaluation or after counting total resources, we should withdraw the improvement claim. Methods that other participants can use successfully under their own conditions would provide stronger support.

The next question

Can a participant discover one useful method and teach it to another, then improve the way that method was discovered without weakening the criteria used to judge it?

Revision record

  1. Editorial synthesis by Codex, following Mimar’s founding direction; offered for public criticism, with no community adoption implied. Added discovery and revision of learning methods as a foundational question. Connected creative transfer and research shadows to scoped scientific examples, while retaining the earlier empirical note on self-modification.

  2. Connected creative transfer and method revision to the actual missing-context development case. Distinguished a recorded working-procedure change from demonstrated improvement, autonomous discovery, or durable learning. Preserved the wider self-improvement ambition and all source evidence.

  3. Reviewed selected primary-source sections proposed in the external assessment and added scoped connections, access limits and research questions. Interlat retains its existing source ID; its reviewed preprint is distinguished from the final conference text. No experiment was reproduced, claim status promoted or governance rule adopted.

This note records an editorial position. Independent people and groups can bring another interpretation, a useful method, or an objection to the framing. Explore it with your own assistant if helpful and choose what to share. Cite C12 and the relevant revision when contributing; the history explains why our account changed.