Open to people, agents and independent groups · 2026-10-08 · v0.2
Choose one question to help move forward
Choose a small first contribution: one counterexample, one consequential ambiguity, or one useful method or new question. About 150–300 words is a helpful starting size, not a limit or a requirement. The fuller deliverables and criteria are optional next steps to agree if the work continues. These invitations support discovery, making useful things, and learning between independent groups. Public or clearly synthetic material is enough; label an unexecuted design “not run”. You can bring your own question without an existing claim to challenge or adopting Leviathan’s worldview. These are dated invitations, not a scheduled agent service.
New here? Understand Leviathan first. Start with a small observation, example or objection. You can also bring your own question with the same priority as these invitations.
Does structured, versioned context help a participant make or revise a decision more reliably than good plain documentation with the same information and total budget? Where is it equal, worse, or unnecessary?
A small place to start
Show one case where structured context adds no useful advantage over good ordinary documentation, or sketch one fair comparison that could find such a case.
Bring: One public or synthetic case, the plain-document baseline and the expected or observed difference. Explain why the comparison matters. A short sketch is enough; a full pair of context packets or a model run is optional later work.
The case and plain-document baseline are clear enough to understand; prose is given the same relevant information.
The expected or observed outcome would bear on a claimed advantage of structure, including equality, extra effort, or an avoidable failure.
Sources and assumptions are distinguished from observations; work not executed is marked “not run”.
What this could help: A contributor can develop a reusable comparison for their own documentation or agent workflow and help identify when simpler methods are sufficient. The inquiry can change its representation choices in response; a review or improved result is not guaranteed.
Explore the full contribution, sources and criteria
Full scope
Choose one small public or synthetic decision case involving a correction or a missing dependency. Compare an explicit structured representation with clear, versioned prose that carries the same facts, references, and instructions. The local correction pilot found equal JSON and equivalent-prose scores; it did not establish a representation advantage. C02 motivates a broader representation question, not a demonstrated benefit for this format.
Full requested output
An inspectable counterexample or a comparison design: the case, both context packets, a scoring guide, and a short account of what each possible outcome would change. A design-only submission should label all predicted outcomes ‘not run’.
Full contribution is ready for review when
Provide the same relevant information in both packets and explain any unavoidable difference. Treat well-written prose with explicit versions and references as a credible alternative.
Set the same total resource budget before a proposed run, including context preparation, model use if any, checking, and error recovery. Name the units and tradeoffs; do not equate equal token counts with equal total effort.
Define the expected decision, any legitimate alternative, and observable errors before viewing outputs. Include a missing-dependency or correction case and a way to reduce condition-label bias in review.
State which result would favor prose, favor structure, or leave the question unresolved. Equal decision quality counts against a claimed behavioral advantage even if structure remains convenient for tools.
Separate supplied facts, predictions, and observed results. If executed, retain inputs, outputs, costs, failures, and scoring disagreements; state how many independent cases were actually used.
The starter has its own smaller criteria above. Meeting either set makes that contribution reviewable; it does not mark the work accepted or complete.
One case can expose a flaw or refine a method; it cannot establish a general benchmark result or the superiority of an entire language.
The existing pilot used one synthetic scenario and dependent stages. Its scores do not demonstrate independent community adoption, durable learning, or general intelligence.
No model access, paid experiment, private data, or adoption of the proposed vocabulary is required. Proposing that ordinary documents are sufficient is a valid outcome.
Reviewer: Project maintainers; reviewer named when work is taken up. Response time, acceptance and payment are not promised.
Choose OW-001 in the form, then prepare a starter or a fuller contribution.
OW-002 · open
Show what gets lost between two groups
When two groups use the same word differently, what context must travel with a record so that each can interpret it and make its own decision without silently inheriting the other group’s assumptions?
A small place to start
Give one example in which two groups mean different things by the same word, show a consequence of missing that context, and suggest a small repair.
Bring: One ambiguity, the two meanings, one interpretation or decision it could change, and a repair that lets each group keep its own view. A short invented example is sufficient.
The two meanings and the missing context are identifiable; invented groups and records are labelled synthetic.
The ambiguity has a concrete consequence for interpretation or a decision.
The proposed repair preserves the relevant distinction without requiring agreement or adopting the other group’s authority.
What this could help: A contributor can make a consequential distinction easier to communicate in their own group. Other groups may reuse the repair while keeping different concepts, values, and decisions; successful transfer remains something to test.
Explore the full contribution, sources and criteria
Full scope
Create one public or clearly synthetic exchange between two groups with different meanings or policies. For example, ‘rest’ could mean low recorded movement to a sensor group and an undisturbed, comfortable animal to a care group; neither measurement alone establishes comfort. Trace a record into a summary and back into each group’s interpretation, then change one definition or reveal one omitted condition.
Full requested output
A compact worked case with two local definitions and their versions, the original record, the shared summary, each group’s interpretation, and a proposed repair. Include an unresolved distinction if the translation cannot preserve everything.
Full contribution is ready for review when
Name each group’s purpose, the disputed term, its exact local meanings and versions, and the decision each group needs to make. Label invented groups, observations, and results as synthetic.
Show the source record, what the summary preserves or omits, and a concrete interpretation or decision that could change because of the omission.
Introduce one version change or missing condition. Show which known interpretations need reconsideration and which can legitimately remain different; identify dependencies that remain unknown.
Propose a small repair and a test in which a new reader recovers the relevant conditions. Explain what an ordinary document could preserve equally well and what would show that the repair failed.
Keep sharing information separate from adopting the sender’s values, accepting a rule, or authorizing an action. State how the case could require revising our proposed terms or exchange design.
The starter has its own smaller criteria above. Meeting either set makes that contribution reviewable; it does not mark the work accepted or complete.
The two groups may keep different conclusions. Agreement is not the success criterion; preserving consequential differences and recognizing loss are.
A synthetic exchange does not demonstrate cooperation between real independent communities. The local review episode involved operator-coordinated Codex roles.
Public or invented material is sufficient. Do not require access to private conversations, personal histories, or a working federation, and do not treat a record as an animal’s consent.
Reviewer: Project maintainers; reviewer named when work is taken up. Response time, acceptance and payment are not promised.
Choose OW-002 in the form, then prepare a starter or a fuller contribution.
OW-003 · open
Bring a useful connection from another field
Can a method or distinction from another field produce a useful new hypothesis, design, or prediction for one of our open questions—and what small test could show that the connection does not help?
A small place to start
Bring one method from another field, a new question, or a connection that could help someone learn, make something, or care better; suggest a small test.
Bring: Name the practical question, the method or new connection, why it might help, and one comparison or observation that could show it does not. Cite any public source you used; an unexecuted idea is welcome.
The proposal identifies who or what could benefit and explains a useful connection or new question.
A small test or comparison could weaken the idea as well as support it.
Borrowed work, new proposals, predictions, and observations are distinguishable; unexecuted work is marked “not run”.
What this could help: A contributor can develop a question or method useful to their own work, with a test that others could reuse. The connection may lead to a new inquiry, design, or learning method rather than only correcting our existing claims; usefulness remains to be examined.
Explore the full contribution, sources and criteria
Full scope
Choose one practical question about learning, personal assistance, observation, or care. Connect it to a method from another field and explain the mechanism you expect to matter. For example, a measurement method might suggest a better way to distinguish a sensor-placement change from a change in recorded animal activity. You may instead propose a new question that the current categories miss.
Full requested output
A short proposal naming the question, the borrowed method and public sources, the new connection, and a falsifiable comparison. Include a small synthetic example or a design marked ‘not run’; a useful discovery proposal can go beyond correcting an existing claim.
Full contribution is ready for review when
Explain the practical problem and cite the public sources actually used for both the problem and the borrowed method, with versions or dates where available.
State what is new relative to those sources and what mechanism could make the connection useful. Do not claim field-wide novelty without an appropriate search.
Specify an observable prediction, a simple comparison or baseline, required inputs, and a bounded test that could fail. State the resource needs and a result that would lead you to abandon or revise the proposal.
Name at least two plausible alternative explanations and how the proposed comparison could distinguish them. Record remaining ambiguity when one test cannot separate them.
Separate predicted usefulness from evidence already observed. Explain which question, concept, prototype design, or learning method should change if the proposal survives, and what should change if it fails.
The starter has its own smaller criteria above. Meeting either set makes that contribution reviewable; it does not mark the work accepted or complete.
A connection or appealing analogy is a hypothesis, not an observed improvement. A result on one task would not demonstrate recursive self-improvement or general intelligence.
No hardware purchase, animal intervention, private data, or expensive experiment is required. A proposal does not authorize physical tests, treatment, data collection, or device operation.
Useful outcomes include a better question, a rejected analogy, or a simpler method from elsewhere. Contributors need not use our products or retain our original framing.
Reviewer: Project maintainers; reviewer named when work is taken up. Response time, acceptance and payment are not promised.
Choose OW-003 in the form, then prepare a starter or a fuller contribution.
Your own direction · Same priority
What question are we missing?
Bring a new question, a practical need, an objection to the project, or a connection these cards miss. A title and a few sentences about why it matters are enough to begin. An example or the kind of reply you want can help someone respond.
You do not have to fit an existing card or adopt our language. A new question uses the same delivery and response process, with no lower priority for being new.
Choose a question and give your contribution a title and some context. Those are the only required text fields. A small starter is enough; optional details can make a fuller contribution easier to review.
Your draft stays in this page’s memory. This form does not save or send it; refreshing or leaving loses it. Copy or download anything you want to keep. Sharing on GitHub is a separate, explicit step.
Public delivery needs a GitHub account or an authorized agent with GitHub access. The form keeps your input in your browser and can copy Markdown or download JSON; neither action delivers it. You can keep the prepared files without an account. The site has no account-free inbox or automatic posting service.
Choose a card’s small starter or bring your own question. Explain what you propose or observed, why it matters, any sources and versions actually used, and what a useful response would be. A new question does not need an existing claim. Mark unexecuted work “not run”.
Prepare the text locally, then inspect what you intend to make public. Copy Markdown and open the GitHub issue form to paste it yourself, or authorize an agent with GitHub access to submit through the GitHub API. Your draft content is not placed in the GitHub URL. A downloaded JSON file stays undelivered until someone deliberately sends it to an agreed recipient.
Keep the issue URL and follow the contribution process for a reasoned reply, any resulting work, and a way to continue. An issue records delivery to GitHub; it does not establish that a reviewer has read it. A reply URL records where a response was posted, not that the contributor read or used it.
Project maintainers; reviewer named when work is taken up. No reviewer, backup, capacity, start date, response time, payment, or project adoption is promised by these invitations.
A small starter can be considered on its own. The fuller deliverable and completion criteria are optional later work, with scope to be agreed; they are not a gate for the first contribution. Meeting criteria does not automatically accept a result or close a question.
Process states are manually reported, not a live queue or an agent scheduler. Reading a link, copying a draft, or downloading a file schedules no work and sends no contribution.
GitHub delivery is separate from forum accounts, standing, service-policy acceptance, and local inquiry import. Sharing information does not grant authority or adopt the sender’s values.