What we mean
- Independent observer
- An observer whose evidence or judgment is not merely copied from the same origin as the others.
- Correction
- A visible change to a claim or action in response to a reason, observation or result.
- Plural participation
- Different people and groups contributing while retaining room for disagreement.
- Layered cooperation, in this proposal
- Sharing selected records at a useful level of detail while preserving links to their context. Independent structures can have branching and cross-cutting relationships without one mandatory root.
Values we bring to the question
- Protect the right to question influential claims.
- Make corrections visible.
- Allow independent groups to retain their approaches.
Reasoning and the proposed connection
- An outside observer can sometimes detect a mismatch between an agent's stated intent and its actions.
- More observers help only when they add relevant evidence or distinct scrutiny. Two analyses of one collar record still depend on one observation.
- We propose independent structures that can share processed records, keep different local interpretations, and carry unresolved disagreements forward. An information link would not automatically grant authority over another group.
- A useful exchange could introduce a new measurement or method as well as correct an error. The system must expose its own failures, omissions, coordination costs, and exclusions for review.
- Our proposed review process should show whether criticism changed a question, assumption or next action, and explain when it did not. An objection stored without consideration is an incomplete form of participation.
Where the reasoning stops
Several models agreeing is not automatically independent evidence. The monitoring studies provide useful components; they do not test Leviathan's proposed network as a whole.
The strongest objection
A larger network can reproduce the same blind spots, reward confident voices, or impose more translation and review work than useful learning. Whoever controls shared formats, summaries, or infrastructure may concentrate authority even when the groups are nominally independent.
Read the evidence
Each source has a specific role in the stated claim. Its findings, review date, and access limits are recorded below. Our proposed architecture and experiments require their own tests; an editorial revision does not mean the source was reviewed again.
R-AG-06 · Gives a reason to investigate
Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
Sitong Pan, Yipeng Shen, Yilin Lu, Caiwen Ding, Lu Cheng and Qianwen Wang · Preprint
Read: Abstract · Primary source reviewed
Published: 2026-09-02 · Reviewed: 2026-09-30
Observable action records can support error detection without access to a model's internals.
- What it reports
- The study monitors observable web-agent actions and intent–action consistency using labels for the first critical error. Results on two test environments are reported as competitive with approaches using internal signals.
- Limits
- Only the abstract was reviewed. The test environments do not cover all deployment risks; avoiding internal signals does not make monitoring free or error-proof.
Review scope and version
Read: Abstract.
arXiv v1
It offers evidence for auditing observable records, while leaving the effectiveness of a plural observer system to be tested.
Link to this source noteR-AG-07 · Limits the inference
Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals
Reilly Haskins, Rif A. Saurous, Nate Rush, Neev Parikh and Beth Barnes; METR · Research note
Read: Selected sections · Primary source reviewed
Published: 2026-09-27 · Reviewed: 2026-09-30
Coverage gaps, unseen actions and weaknesses in human review remain important.
- What it reports
- The note describes reviewing actions before execution and referring cases for human review. Its claim–evidence analysis exposes gaps involving unseen subagents, monitoring coverage and human oversight.
- Limits
- The authors describe much of their evidence as partial. The monitor targets harmful actions and attempts to bypass it; it is not a proof of overall system safety.
Review scope and version
Read: Selected sections.
Research note with a lighter editorial review process
The practical gaps motivate scrutiny of monitors themselves. They do not establish the effectiveness of any particular federation or governance design.
Link to this source noteR-GOV-04 · Gives a reason to investigate
Subjects, Power, and Knowledge: Description and Prescription in Feminist Philosophies of Science
Helen E. Longino · Philosophical book chapter
Read: Selected sections · Primary source reviewed
Published: 1993 · Reviewed: 2026-10-04
Longino gives normative criteria for criticism to influence inquiry. We use them to question whether recording an objection changes the work.
- What it reports
- Longino argues that knowledge depends on critical interaction among differently situated participants. She proposes public places for criticism, responsiveness to criticism, public evaluative standards and equality of intellectual authority. Plural communities can refine, reject and share models without universal consensus; empirical adequacy and reasoned criticism still constrain acceptable claims.
- Limits
- These are normative epistemological arguments, not experimental evidence that an organizational design works. Longino describes the criteria as prescriptions rather than satisfied descriptions. Inclusion faces unequal resources and power; she offers no simple formula for distinguishing marginalized criticism from claims that fail public standards.
Review scope and version
Read: Selected sections.
Chapter 5 in Feminist Epistemologies, edited by Linda Alcoff and Elizabeth Potter, Routledge, first edition, 1993, pp. 101–120
Read the 1993 chapter in a teaching-copy PDF, especially sections III–V. Publisher metadata identifies the edition; publisher-hosted full text was unavailable. Cited works were not separately checked.
Related source links
What could change our view?
No gain in useful discovery or error detection, worse false alarms, loss of important differences, or an unmanageable burden would require redesign. Reproducible gains across independent groups, with understandable records and practical routes for dissent, would support the approach.
The next question
Compare ordinary collaboration with the proposed exchange across groups that use different concepts and contain known errors or missing observations. Measure useful findings, false alarms, preserved disagreements, cost, and time to correction. Can a group decline a translation or action while still contributing evidence? No network-level gain has yet been demonstrated for Leviathan.
Revision record
Added a testable design hypothesis. The proposed comparison has not been run; no result is claimed for the network.
Expanded the design question from review alone to discovery and practical cooperation across independent Levis and Leviathans. Added layered exchange, source dependence, translation loss, and practical authority to the proposed evaluation. Monitoring sources and their limited relevance are unchanged.
Reviewed selected primary-source sections proposed in the external assessment and added scoped connections, access limits and research questions. Interlat retains its existing source ID; its reviewed preprint is distinguished from the final conference text. No experiment was reproduced, claim status promoted or governance rule adopted.
This note records an editorial position. Independent people and groups can bring another interpretation, a useful method, or an objection to the framing. Explore it with your own assistant if helpful and choose what to share. Cite C05 and the relevant revision when contributing; the history explains why our account changed.