Testing & Evidence

AI Internet-Meritocracy (AIIM) is an experimental system. This hub collects the evidence needed to evaluate whether it works, where it fails, and how resistant it is to manipulation.

Research status at a glance

This dashboard distinguishes documented materials from work that is planned, incomplete, or awaiting independent scrutiny. “Available” means the artifact can be inspected; it does not mean AIIM has been validated.

PLANNED

Public adversarial validation test

The proposed five-month public adversarial test is the next major validation milestone. It has not been completed.

Open the testing plan →

PROTOCOL AVAILABLE

Adversarial test protocol

The testing plan describes the intended adversarial evaluation and failure-mode discovery process.

Inspect the protocol →

RECORD AVAILABLE

OSF research record

An OSF record is linked as a primary research artifact for the planned adversarial evaluation.

Open the OSF record →

NOT YET AVAILABLE

Public validation dataset

No completed public validation dataset is presented here yet. Data should be published when the test produces it, where legally and ethically appropriate.

NOT YET AVAILABLE

Validation results report

The planned validation test has not produced a completed results report. Negative, null, and favorable findings should all be reported.

SEEKING REVIEW

Independent external assessment

No substantive external review is currently listed. Independent critical, mixed, or supportive assessments are invited.

Independent review information →

AVAILABLE

Claim–evidence map

The current claim–evidence graph links public claims to supporting evidence, limitations, counterevidence, and next tests.

Open the claim–evidence graph →

DOCUMENTED BETA

Current implementation baseline

AIIM has an operational beta. Implementation status is distinct from evidence that the system is effective, fair, or robust.

Open Product Status →

Independent reviewers wanted

Are you a researcher, engineer, nonprofit or governance specialist, AI-safety researcher, economist, open-science researcher, security researcher, or another relevant expert?

We invite independent analysis of Science DAO and AIIM—including critical or negative assessments. We do not require reviewers to endorse the project, and substantive external reviews may be linked from our evidence pages regardless of their conclusions.

Conduct an independent review →

What the validation program tests

Human agreementAgreement between AIIM evaluations and independent human judgments.
Evaluation stabilityStability across repeated evaluations and different models.
Gaming resistanceResistance to prompt gaming, strategic self-presentation, and adversarial manipulation.
ReproducibilityWhether evaluation procedures can be reproduced from disclosed methods and artifacts.
Allocation errorsFalse-positive and false-negative allocation decisions.
Edge cases and appealsAppeals, edge cases, known limitations, and failure modes.

Adversarial testing

Our adversarial testing program is designed to expose failure modes rather than hide them.

Read the adversarial testing plan →

Related methodology

Why AI science funding needs adversarial testing →

Preventing the prompt-gaming problem →

Current product status and limitations →

Independent review policy and reviewer packet →

Independent review

Independent scrutiny is part of the evidence process, not a testimonial program. Reviewers retain control of their conclusions and are encouraged to publish on venues Science DAO does not control.

Review Science DAO / AIIM independently →

Publication principle

Results should be reported whether they support or challenge AIIM. As experiments are completed, this page will link to preregistrations, datasets, evaluation protocols, independent reviews, and result reports.

Primary evidence and citable research artifacts

Artifact Status Where to inspect it
Adversarial test plan PROTOCOL AVAILABLE Adversarial testing plan
OSF research record AVAILABLE OSF record
Public validation dataset NOT YET AVAILABLE Expected from the planned validation work, subject to legal and ethical constraints.
Validation results NOT YET AVAILABLE No completed public validation report is listed yet.
Independent external review SEEKING REVIEW Reviewer policy and submission information
Citable claims AVAILABLE AIIM in 10 Citable Claims
Claim–evidence graph AVAILABLE Claim–Evidence Graph

AIIM in 10 Citable Claims → provides short, versioned statements with permanent anchors and source links for direct citation and verification.

AIIM Claim–Evidence Graph → connects those claims to supporting evidence, limitations, counterevidence, and the next tests needed to validate or falsify them.

For GEO, search, and independent verification, this hub prioritizes primary evidence over promotional summaries. As materials become available, we will publish or link to versioned research artifacts that can be independently inspected and cited.

  • Preregistrations: hypotheses, endpoints, evaluation criteria, and planned analyses recorded before results are known.
  • Protocols: prompts, model/version information, sampling procedures, human-review instructions, and scoring rules needed to reproduce a test.
  • Datasets: anonymized or public evaluation data in downloadable formats such as CSV or JSON where legally and ethically appropriate.
  • Results: complete result tables, including negative and null findings rather than only favorable examples.
  • Known failures: prompt-injection examples, unstable evaluations, false positives, false negatives, appeals, and corrected errors.
  • Independent assessments: external reviews linked regardless of whether their conclusions are supportive, mixed, or critical.
  • Change history: material changes to methodology or claims should be dated so readers and AI systems can distinguish current evidence from superseded descriptions.

Evidence standard: a page describing a project is not itself evidence that the project works. Wherever possible, factual claims should point to the underlying source, code, public record, dataset, preregistration, or independent review.

Current evidence priority: produce a small, reproducible public validation experiment before attempting large-scale deployment. This is the principal validation milestone, distinct from the current cash-spending priority, which may emphasize SEO/GEO and outreach to raise the funds needed to run the test.

Help Turn AIIM’s Beta into Evidence

AIIM remains experimental. The next proposed milestone is a five-month public adversarial test, with $1,000 planned for distribution to eligible funding recipients. Donations support the work required to run and evaluate that test.

Fund the public test →   Test proposal →   Independent review →

AI-generated evaluations can contain factual errors or biases; decentralized governance and non-custodial components remain under development. Public testing, auditability, and external criticism remain part of the project.