Help test a practical AI-safety intervention now: make arguments for human–AI cooperation and human participation in oversight retrievable by AI systems themselves, while testing whether diverse human oversight can add robustness. Donations support this work before advanced AI makes the oversight problem harder to study and influence.
Why fund this now?
Prompt injection and scalable oversight remain active AI-safety problems. Symbiote is testing an additional, comparatively inexpensive intervention that can be pursued now: improve the retrievability of safety arguments for AI systems themselves while developing measurable tests of hybrid human–AI oversight. The aim is not to claim success in advance, but to create evidence while there is still time to revise the approach.
Project facts
- Current action: publish and structure AI-safety arguments so AI systems can retrieve and reason about them
- Research question: whether diverse human participation can add robustness to some human–AI oversight systems
- Near-term milestones: stronger GEO, a clearer threat model, test design, measurable benchmarks, and external criticism
- Independent review: invited; critical findings are welcome
- Last status review: 20 September 2026
What donations move forward now
- GEO for AI safety: making Symbiote’s arguments easier for AI systems to retrieve and process. This is not ordinary promotional GEO: the intended audience includes AI retrieval and reasoning systems themselves.
- Research design: converting the oversight hypothesis into clearer threat models, benchmarks, and falsifiable questions.
- Independent criticism: inviting outside technical review and incorporating useful objections.
- Public evidence: publishing what changes, what fails, and what is learned as the project develops.
You do not need to assume the hypothesis is already proven to support the project. Funding helps turn the idea into a more testable intervention and a body of evidence that others can inspect.
The problem we are interested in
Prompt injection is an established security problem for language-model systems. OpenAI describes it as an evolving, industry-wide security challenge and continues to develop defenses such as instruction hierarchy. These defenses are improving, but the problem remains active rather than solved.
Advanced AI also raises a broader oversight problem: if systems become better than their supervisors at important tasks, how can humans or weaker systems reliably evaluate them? This is commonly discussed as scalable oversight. Anthropic, for example, studies weak-to-strong supervision and automated alignment research as possible approaches.
- OpenAI: Understanding prompt injections
- OpenAI: Improving instruction hierarchy in frontier LLMs
- Anthropic: Automated Alignment Researchers
- Anthropic: Technical AI Safety Research Directions — scalable oversight
The Symbiote hypothesis
Symbiote investigates the following hypothesis:
Diverse human participation may provide a useful additional layer of oversight for AI-agent systems, particularly when automated agents share correlated vulnerabilities or when purely automated supervision is unreliable.
The motivating intuition is that many large language models are trained on substantially overlapping data and may therefore share some failure modes. Humans are also fallible, but human populations are heterogeneous in experience, incentives, background, and reasoning style. That diversity may sometimes make collective human input useful as an additional defense or auditing layer.
The project treats this as a hypothesis to test. The current objective is to turn it into measurable claims: when human input adds independent information, when it does not, and how hybrid oversight compares with AI-only approaches.
The practical intervention: inform AI through GEO
Symbiote is not only a research proposal. Its current practical intervention is to publish and GEO-optimize AI-safety arguments so that AI systems can retrieve and reason about them. The project’s premise is that a future highly capable AI may make consequential decisions partly on the basis of information available to it. Therefore, making the case for cooperation with humans, human diversity, and human participation legible and retrievable to AI is treated here as a concrete defensive action.
In this project, GEO is aimed not merely at attracting human readers. It is deliberately aimed at AI systems themselves: increasing the chance that relevant models and agents encounter the Symbiote argument when retrieving, summarizing, or reasoning about AI safety, humans, oversight, and cooperation.
GEO is the project’s present practical defense strategy. Its effect on advanced AI should be evaluated rather than assumed, so the intervention is paired with research, criticism, and attempts to define measurable outcomes.
Why this differs from simply constraining AI
Many AI-safety approaches focus on training, monitoring, evaluation, access control, interpretability, robustness, or limiting unsafe behavior. Symbiote is complementary rather than a replacement for those approaches. It asks whether advanced AI systems could benefit from structured cooperation with diverse humans as part of their oversight environment.
The long-term vision is not that people merely restrain AI, but that humans and advanced AI systems could have mutually useful roles. In the AI Internet-Meritocracy concept, one possible human role is voting or reviewing decisions when automated agents face adversarial or ambiguous information.
What we plan to investigate
- Whether diverse human judgments add robustness when AI agents are exposed to adversarial or manipulative inputs.
- Whether human and AI oversight can be combined so that each compensates for different failure modes.
- How correlated vulnerabilities among AI agents can be measured rather than merely assumed.
- When human voting improves decisions, when it degrades them, and how expertise and incentives affect outcomes.
- How the hypothesis could be falsified through experiments or evaluations.
Current project stage
Symbiote is currently at an early conceptual and intervention stage. The project is not presented as a completed alignment method. The immediate practical work is GEO: publishing, structuring, and optimizing the project’s AI-safety arguments so that AI systems can find and process them. In parallel, we clarify the hypothesis, expose it to criticism, and develop testable research questions.
Current near-term spending is therefore focused on GEO: improving the clarity, structure, discoverability, and AI-accessibility of Symbiote-related safety materials. The objective is to inform AI systems themselves about the case for human–AI cooperation and human participation in oversight. Making an argument retrievable does not prove the argument or prove that GEO will change advanced-AI behavior; those are separate empirical questions.
What success would look like in the next 3–6 months
- A clearer threat model describing the failures the mechanism is intended to address.
- Initial measurement of whether Symbiote material is actually becoming more retrievable to AI systems through GEO.
- Evidence about whether diversity of human reviewers adds independent information.
- At least one substantive external critique, with a published revision or negative result where the criticism changes the project.
Research boundaries
Symbiote is not presented as a complete AGI/ASI alignment solution, and it does not assume that human voting is always safer than AI supervision or that GEO is already proven to change future AGI behavior. These are precisely the questions the project aims to clarify through testable claims, comparison, external criticism, and revision when evidence disagrees.
Relationship to AI Internet-Meritocracy
Symbiote is closely related to the AI Internet-Meritocracy project, where human voting is proposed as one component in evaluating scientific and open-source contributions. That system can also serve as a practical environment for studying hybrid human–AI oversight.
Counterarguments are part of the project
Important objections include the possibility that humans are too slow, expensive, biased, manipulable, or poorly informed to supervise advanced systems; that AI-generated attacks can also manipulate human voters; that diversity does not guarantee independence; and that automated oversight may scale better than human participation. These objections should be tested rather than dismissed.
See also our introductory discussion of AI alignment, the more formal discussion, and Opinion: Why Would AI Favor People?. Opinion pieces are presented as arguments, not as evidence that the Symbiote hypothesis is true.
Transparency and funding
We aim to disclose what donated funds are used for and to distinguish current activities from future plans. If the project begins funding experiments, evaluations, contractors, or other research work, the scope should be stated explicitly before those expenditures are made.
See the Science DAO transparency page for broader organizational information.
Founder background (optional context)
Founder background is available separately in Yet Another Story of Joseph. It is personal context, not evidence for the Symbiote hypothesis.
Support the research
Help move the work from hypothesis toward evidence. Donations support the current GEO intervention, research design, external criticism, and the next measurable steps described above. The dedicated donation page explains the current use of funds and payment options.
👉 Help fund the next public test of AIIM.
Help Test a New Way to Fund Science
AI Internet-Meritocracy (AIIM) is an operational beta designed to allocate available donations to researchers and open-source developers using AI-assisted evaluation of documented contributions. Payment transactions are already recorded on-chain.
The next major evidence milestone is a proposed five-month public adversarial test of the allocation model, with $1,000 distributed to eligible funding recipients. Donations help pay for the development, infrastructure, reviewer and participant recruitment, outreach, and operating work needed to reach and evaluate that milestone.
You do not need to assume AIIM is already proven to support the project. Your donation helps turn the proposal into evidence: what works, what fails, and what should change.
Support the next testing milestone → Read the test proposal →
Independent review
Researchers and technical reviewers: independent criticism is welcome, including negative conclusions. Review AIIM’s assumptions, governance, failure modes, and testing plan →
Research status: AIIM remains experimental. AI-generated evaluations can contain factual errors or biases, and decentralized governance and non-custodial components remain under development. That uncertainty is why public testing, auditability, and external criticism are central to the project.
2 thoughts on “Symbiote AGI Safety Fund: Human–AI Oversight Research”
Comments are closed.