The Science Funding Goodhart’s Law Problem: Can AIIM Preserve Fairness?

Getting your Trinity Audio player ready...

Science funding faces a persistent problem: when rewards depend on an indicator of scientific value, researchers gain an incentive to improve that indicator—even when doing so adds little scientific value.

Our hypothesis is that adaptive AI funding, as proposed through AI Internet-Meritocracy (AIIM), could preserve fairness more effectively than fixed scoring systems by recognizing manipulation, reassessing evidence, and correcting its evaluation methods. This is a testable possibility, not an established achievement.

The crucial question is whether evaluation can adapt fast enough to keep genuine contributions more rewarding than gaming.

What is Goodhart’s law in science funding?

Goodhart’s law describes how optimizing a proxy can weaken its relationship with the underlying goal. In science funding, the goal might be valuable knowledge, while the proxy is publication count, citations, journal prestige, or an evaluator’s score.

The problem extends beyond deliberate cheating. An imperfect indicator can produce distorted allocations even when everyone follows the rules. Manheim and Garrabrant distinguish several mechanisms through which optimization undermines proxies in their theoretical paper, Categorizing Variants of Goodhart’s Law.

Consider these illustrative incentives:

Funding targetBehavior it could encourageWhat the indicator may miss
Number of papersDividing one contribution into many publicationsDepth and originality
Citation countReciprocal citation arrangementsIndependent scientific usefulness
Journal prestigePrioritizing venue appealValuable work published elsewhere
Software activityProducing numerous minor commitsReliability and essential maintenance
AI assessment scoreTailoring descriptions to the evaluatorThe contribution behind the description

Improving an indicator is not inherently bad. Better documentation can help others use research. The failure occurs when a funding system rewards an improved appearance without a corresponding improvement in contribution.

This concern already informs research assessment reform. The San Francisco Declaration on Research Assessment (DORA) recommends evaluating research on its own merits, avoiding journal metrics as substitutes for individual quality, and recognizing outputs such as datasets and software.

Why AI funding might adapt more effectively

AI Internet-Meritocracy is Science DAO’s experimental approach to allocating donated funds through AI-assisted evaluation of documented scientific and open-source contributions. Its current beta does not establish that adaptive fairness has been achieved.

The proposed advantage is that an AI evaluator could examine relationships among evidence instead of relying on a permanent formula.

A fixed rule might treat every additional citation as equally valuable. An adaptive evaluator could investigate whether citations represent independent reuse, criticism, routine background references, or coordinated promotion. Those distinctions could change how much evidential weight the citations receive.

Three capabilities make the hypothesis plausible.

1. Assessing the substance behind an indicator

An AI system could assist with comparing papers, examining software changes, tracing documented reuse, and identifying discrepancies between claimed and demonstrated contributions.

For example, ten papers that repeat substantially the same result should not automatically receive ten times the recognition of one complete presentation. Detecting overlap could reduce the advantage of inflating publication counts.

These assessments remain fallible. Reading a paper does not establish its correctness, and tracing citations does not prove causal importance.

2. Updating evaluations when manipulation becomes visible

A funding system could revise how it interprets a signal after discovering that the signal is being exploited.

Suppose coordinated citations initially inflate several researchers’ assessments. After investigation, the evaluator could discount that pattern and reassess affected records. The correction should also be tested against legitimate, closely connected research communities to avoid penalizing normal collaboration.

An important distinction follows: updating software is easy compared with demonstrating that the update makes funding fairer.

3. Applying validated corrections broadly

Once a correction has been tested, software can apply it consistently across many assessments. AI could also help identify older decisions affected by the same weakness.

Human funding institutions can adapt too. AIIM’s distinctive hypothesis concerns the speed, cost, breadth, and consistency of adaptation—not an exclusive ability to learn.

Fairness needs stable principles

Adaptation can itself become unfair if researchers cannot understand why their funding changes.

For this hypothesis, fairness should mean that comparable, well-supported contributions receive comparable consideration; irrelevant prestige and superficial manipulation do not systematically improve awards; and mistaken decisions can be challenged.

These principles should remain stable even as methods of interpreting evidence evolve.

Several design requirements follow:

  • Evidence-based explanations: Identify which contributions and supporting records affected an assessment.
  • Context-sensitive evaluation: Account for differences among disciplines, career stages, and types of work.
  • Uncertainty handling: Distinguish insufficient evidence from evidence of low value.
  • Versioned changes: Record which evaluation method produced a decision and why that method changed.
  • Meaningful appeals: Allow factual corrections and independent review of disputed judgments.

These are proposed requirements for preserving fairness, not a claim that every feature is already implemented in AIIM.

An AI score can become another Goodhart target

Replacing citation counts with an AI-generated merit score does not remove the underlying problem. Applicants can optimize for whatever raises that score.

They might discover persuasive wording, exploit prestige biases, fabricate evidence, or insert instructions into documents the evaluator reads. Multiple AI evaluators can share the same weaknesses, so agreement alone does not establish reliability.

Moreover, AI can help attackers generate and test manipulations. Manheim and Garrabrant explicitly identify stronger optimization as a reason Goodhart effects matter for artificial intelligence. Their analysis supports caution about AI’s power; it does not demonstrate that AIIM succeeds or fails.

Adaptive AI funding works only if corrections reduce distortion without creating equally serious new errors.

Making the evaluator unpredictable would not be sufficient. Randomly changing criteria could frustrate manipulators while also harming honest researchers. The aim should be defensible adaptation under publicly understandable principles.

Human oversight also needs scrutiny. Reviewers and voters can have conflicts of interest, favor familiar institutions, or mistake unconventional work for low-quality work.

How to test the AIIM hypothesis

Science DAO’s proposed adversarial testing of AIIM provides a relevant starting point. Testing adaptive fairness would require an explicit comparison over repeated rounds.

A useful experiment would compare:

  1. A fixed bibliometric scoring system.
  2. A frozen AI evaluator.
  3. An adaptive AI evaluator with documented updates and an appeal process.
  4. A human or hybrid review baseline, where resources permit.

All systems should receive comparable evidence and funding budgets. Participants should be allowed repeated attempts to improve their allocations through both legitimate documentation and controlled manipulation.

The evaluation should track:

OutcomeWhat it reveals
Funding gained through changes with no added scientific valueSusceptibility to gaming
Honest contributors wrongly penalizedCost of defensive measures
Time and resources needed to correct an exploitPractical adaptability
Performance against previously unseen attacksWhether corrections generalize
Unexplained allocation changesDecision instability
Evaluation and appeal costsOperational feasibility

Assessments of contribution should use independent evidence and reviewers who do not know which system produced each allocation. Reviewer disagreement should be reported rather than treated as nonexistent.

Even this benchmark is an imperfect proxy for fairness. Longer-term replication, reuse, and expert reassessment can provide additional checks, although their delays make immediate validation difficult.

The hypothesis would gain support if adaptive AIIM repeatedly reduced manipulation gains while protecting legitimate contributors at a sustainable cost. It would be weakened if updates merely shifted the preferred manipulation strategy, increased arbitrary decisions, or performed no better than simpler alternatives.

A funding system that can correct itself

Goodhart’s law gives science funders a reason to design for correction. A useful indicator should remain open to challenge when incentives change what it measures.

AIIM could contribute by making evaluation revisable, connecting decisions to evidence, and applying validated improvements across many funding decisions. Whether it preserves fairness depends on the quality of those corrections and the governance surrounding them.

The strongest promise of adaptive AI funding is the possibility of keeping rewards aligned with genuine scientific contributions even as people learn how the system works. Demonstrating that ability would be a meaningful result for science funding—and an appropriate goal for AIIM’s experiments.

👉 Donate for science.

Support Independent Science

Our flagship product, AI Internet-Meritocracy, is an experimental app designed to allocate donated funds to researchers and open-source developers using AI-assisted evaluation of documented contributions. Payments depend on available funds and eligibility requirements.

Help fund the proposed five-month public test of AIIM’s allocation model. Support the next testing milestone.

Supporting independent science is not only a matter of fairness to researchers whose expertise and work are often underfunded. It is also essential for addressing systemic failures in scientific publishing that delay discoveries and leave important results unnoticed. In science and software, even one missing component can prevent an entire system from working.

Help valuable research and open-source infrastructure move forward. Please make a donation to support independent scientists and free software developers.

Disclaimer

Experimental-system notice: AI Internet-Meritocracy is an experimental funding system. Its AI-generated evaluations are heuristic judgments based on available public or connected-account evidence; they are not validated measurements of a person’s causal economic or scientific impact. Payment transactions are already recorded on-chain and can be verified on the blockchain. The current beta initiates payments off-chain through Node.js and uses custodial and administrative components. Decentralized governance and non-custodial wallets remain under development; on-chain payment records are already available. Evaluations may contain factual errors or biases and should be interpreted together with audit logs, appeals, human oversight, and published test results.

Ads:

Description Action
A Brief History of Time
by Stephen Hawking

A landmark volume in science writing exploring cosmology, black holes, and the nature of the universe in accessible language.

Check Price
Astrophysics for People in a Hurry
by Neil deGrasse Tyson

Tyson brings the universe down to Earth clearly, with wit and charm, in chapters you can read anytime, anywhere.

Check Price
Raspberry Pi Starter Kits
Supports Computer Science Education

Inexpensive computers designed to promote basic computer science education. Buying kits supports this ecosystem.

View Options
Free as in Freedom: Richard Stallman's Crusade
by Sam Williams

A detailed history of the free software movement, essential reading for understanding the philosophy behind open source.

Check Price

As an Amazon Associate I earn from qualifying purchases resulting from links on this page.

Leave a Reply

Your email address will not be published. Required fields are marked *