LEVIATHAN.LIFE
All foundations

C01 · Empirical question

What does human language carry into AI?

Statement assessed

Some learned internal representations associated with human concepts can causally shape model behavior.

Status of this statement: Supported within a defined scope

This status applies to the statement above, within the limits discussed in this note. Evidence for a finding can motivate Leviathan’s proposed mechanisms without establishing that they work.

Working editorial note · Last recorded revision

What we mean

Internal representation
A pattern of model activity associated with information the model uses.
Causal intervention
Changing an internal pattern and testing whether the model's behavior changes.

Values we bring to the question

  • Understand the systems that increasingly shape our lives.
  • Make the influence of inherited human concepts open to scrutiny.

Reasoning and the proposed connection

  1. If a learned concept changes a model's decisions, understanding its role matters for both capability and safety.
  2. This connects the study of language and culture to mechanisms we can investigate. A human name for a representation remains a hypothesis about its function.
  3. Leviathan proposes local meaning kernels that connect concepts, principles, rules, and their versions. Those explicit records are a design layer; they are distinct from learned internal model representations.
  4. A proposed test would ask whether that context changes interpretation and action on unfamiliar cases, including when a community revises a concept or disagrees with another community. The cited findings motivate this question; they do not answer it.

Where the reasoning stops

The findings concern particular concepts, models and experimental settings. They do not establish that all culture is contained in language or that every model operation can be reduced to a human concept.

The strongest objection

A representation extracted using human labels may partly reflect the researcher's choice of categories. An intervention can change behavior without proving that a model uses the concept as a person does. Similarly, supplying a detailed value context could produce fluent agreement without reliable changes in action.

Read the evidence

Each source has a specific role in the stated claim. Its findings, review date, and access limits are recorded below. Our proposed architecture and experiments require their own tests; an editorial revision does not mean the source was reviewed again.

R-AI-01 · Supports this claim

Emotion Concepts and their Function in a Large Language Model

Nicholas Sofroniew and 15 co-authors; Anthropic · Preprint

Read: Author summary · Primary source reviewed

Published: 2026-04-02 · Reviewed: 2026-09-30

Interventions on emotion-related representations change measured behavior.

What it reports
Vectors associated with 171 preselected emotion concepts were extracted from Claude Sonnet 4.5. Interventions changed preferences and some alignment-related behaviors, supporting a functional role for these learned concepts.
Limits
The 171 concepts are a starting list, not 171 discovered feeling centers. The representations mainly track local context. The blackmail experiment used an early, unreleased snapshot; subjective feeling was not established.
Review scope and version

Read: Author summary.

Research announcement: 2 April 2026; arXiv v1: 9 April 2026

The geometry has a partial replication in another model in R-AI-06. That replication does not repeat the original behavioral interventions.

Link to this source note

R-AI-06 · Supports part of this claim

Replicating the Geometry of Emotion Representations in a Base Open-Weights Model

Adam Hollowell; University of North Carolina at Chapel Hill · Preprint

Read: Selected sections · Primary source reviewed

Published: 2026-09-01 · Reviewed: 2026-09-30

A different base model partly reproduces the representation geometry; causal behavioral effects were not tested.

What it reports
In base Gemma-2-27B, much of the valence geometry and clustering of 171 emotion vectors is recovered. Some structure is already present in token embeddings. The arousal axis does not meet all stability criteria.
Limits
This supports the representation finding, not a causal effect on behavior. The study uses Claude-generated fiction, one different model, reconstructed methods and a linear representation assumption.
Review scope and version

Read: Selected sections.

arXiv:2609.22208v1; submission date listed as 1 September 2026

This is a partial geometric replication. The author does not count arousal as fully replicated and identifies structural-token confounds in some measurements.

Related source links

Link to this source note

R-AI-09 · Supports this claim

How’s it going? Reinforcement learning in language models recruits a functional welfare axis

Andy Q Han, David J. Chalmers and Pavel Izmailov · Preprint

Read: Abstract · Primary source reviewed

Published: 2026-05-28 · Reviewed: 2026-09-30

Minimal reward training recruits pre-existing representations that affect behavior.

What it reports
Reward and punishment vectors extracted after maze training influence behavior in other tasks. The authors' controls support the view that training recruits pre-existing representations rather than creating them from scratch.
Limits
Functional welfare here means an estimate of doing well or badly relative to goals. It is not a claim about experienced pleasure or pain. This review covers the abstract.
Review scope and version

Read: Abstract.

arXiv v1

The study adds evidence about functional internal representations. Its publication date is May 2026, even when later coverage draws attention to it.

Link to this source note

R-AI-03 · Adds context

Emergent Introspective Awareness in Large Language Models

Jack Lindsey; Anthropic · Preprint

Read: Selected sections · Primary source reviewed

Published: 2025-10-29 · Reviewed: 2026-09-30

Some models have limited access to their internal representations; this does not validate every self-report.

What it reports
Some Claude models can identify injected concepts in certain settings and distinguish earlier internal representations from input text. Controls support a limited form of functional introspection.
Limits
Failures are common, and performance depends on context and post-training. The experiments do not establish human-like introspection, the reliability of every self-report, or subjective experience.
Review scope and version

Read: Selected sections.

First published: 29 October 2025; web revision: 1 January 2026; arXiv v1: 5 January 2026

The January web revision adds a control experiment and prompt corrections. The later arXiv submission is a separate publication event, not an independent replication.

Link to this source note

What could change our view?

Independent interventions that fail to reproduce the reported effects, or controls showing that unrelated directions explain them equally well, would weaken the claim. Replication across models, languages and contexts would strengthen it.

The next question

Which effects survive different concept lists, model families, languages, and post-training methods? For our proposed local kernels, compare ordinary instructions with explicit concept–principle relationships at a comparable total budget. Test behavior when meanings change, values conflict, and unfamiliar cases appear; record failures as well as agreement.

Revision record

  1. Added the functional-representation claim with a partial replication and its limits. Clarified that 171 is the study's preselected concept list.

  2. Connected the functional-representation findings to a proposed test of local meaning and value contexts. Distinguished explicit kernels from learned model representations and verbal agreement from behavior. Empirical statement, support status, and source reviews are unchanged.

This note records an editorial position. Independent people and groups can bring another interpretation, a useful method, or an objection to the framing. Explore it with your own assistant if helpful and choose what to share. Cite C01 and the relevant revision when contributing; the history explains why our account changed.