|
Getting your Trinity Audio player ready...
|
The strongest arguments against AI Internet-Meritocracy (AIIM) concern whether it can measure scientific value reliably, resist manipulation, and turn fair recognition into additional research. These objections challenge the connection between AIIM’s proposed allocation method and its intended benefits.
AIIM is Science DAO’s experimental funding system, using AI assessments of documented research and open-source contributions to guide payments from donated funds. Its appeal is understandable: useful work should be eligible for support without requiring a degree, institutional affiliation, or conventional grant proposal.
However, removing those barriers does not establish that the replacement allocates money well. The following arguments identify risks and unresolved questions; they are not findings that AIIM has already failed.
1. Scientific value may not support the numerical precision AIIM requires
AIIM describes its impact score as a heuristic framed as a share of world GDP, while explicitly acknowledging that it is not a validated measurement of causal impact. That qualification matters. The project’s explanation of its scoring model distinguishes an AI-generated estimate from an established economic measurement.
Consider a theorem that helps improve an algorithm, which becomes part of software used in manufacturing. Assigning economic credit requires assumptions about alternative discoveries, substitute software, implementation work, and what would have happened without the theorem.
Even complete knowledge of the dependency chain would not uniquely determine how credit should be divided. Different defensible allocation rules could produce different payments.
A precise score can express an uncertain judgment without making that judgment more accurate.
Using scores only proportionally does not remove this problem. Multiplying every score by the same number leaves allocations unchanged, but incorrectly rating one contribution ten times above another changes who receives the money.
This objection requires evidence about comparative evaluation quality and sensitivity to assumptions. More decimal places or more confident explanations cannot answer it.
2. AIIM could reproduce the recognition barriers it aims to overcome
An evaluator must distinguish important unfamiliar research from impressive-looking work that contains errors. That is particularly demanding when a contribution uses new definitions, challenges established approaches, or has received little expert attention.
Research by Zheng and colleagues found that language-model judges can approximate human preferences in chatbot evaluation, while also exhibiting position, verbosity, and self-enhancement biases. This is evidence about those evaluation settings, not a direct test of AIIM or a demonstration that models can assess original scientific importance. Read the study on LLM judges.
The concern for AIIM is an inference: polished presentation, familiar terminology, and visible recognition might influence scores more than underlying merit. An independent researcher could face a new gatekeeper trained on the judgments of the old ones.
Testing should therefore include unfamiliar but verified work, plausible-looking incorrect work, and equivalent submissions with different names, affiliations, languages, and writing styles. Agreement with established reviewers is useful evidence, but cannot by itself prove that neglected discoveries receive fair treatment.
3. Financial rewards could turn evaluation into an optimization contest
Once scores influence payments, participants have an incentive to discover what raises those scores. Some improvements will be legitimate, such as documenting contributions clearly. Others could include exaggerating dependencies, producing redundant outputs, or coordinating endorsements.
There is also a technical attack surface. Greshake and colleagues demonstrated indirect prompt injection: instructions embedded in external material could manipulate the behavior of LLM-integrated applications. Their work establishes an attack class, not the current vulnerability of any particular AIIM deployment. Read the indirect prompt-injection research.
The difficult case is manipulation that looks like ordinary scholarly communication. A fabricated claim of importance can mislead an evaluator without containing an obvious instruction to ignore its rules.
AIIM must remain useful when applicants optimize for its evaluator, rather than merely when applicants describe their work honestly.
The proposed AIIM adversarial testing program addresses this concern. However, a small test can reveal vulnerabilities without establishing resistance to better-funded attackers or sustained collusion at a larger scale.
4. Rewarding past merit is different from financing future progress
Two funding questions must be separated:
- Who deserves recognition and compensation for previous contributions?
- Where would additional money enable the most valuable future work?
The answers may differ. A highly accomplished researcher might already have enough support to continue. A newcomer might need a modest payment to finish a first important contribution but have little documented work to evaluate.
This is a structural concern for funding based on past contributions. It can recognize researchers who already crossed the initial funding barrier while leaving others unable to begin.
Retrospective rewards might still enable future research, encourage open publication, or provide valuable income security. Those are plausible mechanisms, but each requires evidence.
AIIM should state whether its primary goal is deserved compensation, additional scientific output, or a combination. A system can succeed at one while performing poorly at another.
5. Human voting can relocate the governance problem
Human oversight can correct AI errors. It also introduces questions about who participates, what evidence they see, and how conflicts of interest are handled.
AIIM’s adversarial-testing proposal includes a governance or ban-voting process. That makes the quality of human decisions part of the system’s security, rather than an external guarantee. See the proposed testing and governance process.
Possible failure modes include coordinated groups protecting one another, participants voting against competitors, and legitimate researchers struggling to appeal a mistaken ban. These are risks to investigate, not allegations about existing participants.
The hardest design tension is that open participation can expose governance to manipulation, while restrictive participation can restore the gatekeeping AIIM seeks to reduce. More automation does not resolve who should have authority over disputed judgments.
6. Blockchain records cannot establish that an allocation was deserved
AIIM states that payment transactions are recorded on-chain, while the current beta initiates payments off-chain through Node.js and retains custodial and administrative components. Decentralized governance and non-custodial wallets remain under development. See AIIM’s current implementation description.
These are distinct properties. A verifiable payment record can show that money moved. It cannot establish that the recipient’s scientific claims were correct or that the allocation rule was fair.
Even a fully on-chain implementation would retain this distinction. Immutable records preserve good and bad decisions alike. Evaluator quality, evidence quality, and correction procedures therefore need separate assessment from transaction transparency.
7. Low administrative overhead may conceal missing services
Direct payments could reduce application work and institutional intermediation. However, comparing payment-processing costs with the entire cost of a research institution would compare different services.
Research can require laboratories, equipment maintenance, data stewardship, technical staff, and long-term coordination. Someone still has to provide and finance these functions when researchers receive money directly.
AIIM also has potential operating costs: evidence verification, identity checks, model usage, disputes, appeals, security, and software maintenance. Unpaid work by applicants and volunteers should count when evaluating efficiency.
This suggests a plausible boundary: direct individual funding may fit some mathematics and software work more readily than research requiring expensive shared facilities. AIIM could complement institutional infrastructure, but cheaper transfers alone would not demonstrate cheaper science.
8. A functioning pilot would not establish superior research outcomes
Successfully calculating scores and sending payments would demonstrate operational feasibility. Favorable participant surveys would provide evidence about perceived fairness. Neither would establish that AIIM produces more valuable discoveries per dollar.
That claim needs a relevant comparison, such as simple expert allocation, equal payments among eligible contributors, or a lottery within a qualified pool. All approaches should be assessed using comparable budgets and explicit objectives.
A useful evaluation would distinguish:
- Reliability: Do equivalent submissions receive similar assessments?
- Fairness: Who is systematically overlooked or incorrectly penalized?
- Security: How much can manipulation change actual payments?
- Cost: What resources do administration and participation consume?
- Additionality: What useful work becomes possible because funding was provided?
Scientific outcomes can take years to emerge. Early evidence should therefore support narrower claims, with uncertainty and unsuccessful results published alongside successes.
What would justify confidence in AIIM?
The strongest case against AIIM is that its central chain of reasoning remains unproven: documented contributions must lead to defensible assessments, those assessments must support appropriate payments, and those payments must advance the stated funding goal.
Failures in conventional funding provide a reason to experiment. They do not establish which alternative works better.
Confidence would grow through independent evaluation, published limitations, meaningful comparisons, effective appeals, and evidence that results survive manipulation and changes of scale. Some findings might support AIIM broadly; others might justify a narrower role or a different scoring method.
AIIM’s credibility should depend on its willingness to change when evidence challenges its design. A serious experimental funding system needs a clear account of what would count against it, as well as what would count in its favor.
Support Independent Science
Our flagship product, AI Internet-Meritocracy, is an experimental app designed to allocate donated funds to researchers and open-source developers using AI-assisted evaluation of documented contributions. Payments depend on available funds and eligibility requirements.
Help fund the proposed five-month public test of AIIM’s allocation model. Support the next testing milestone.
Supporting independent science is not only a matter of fairness to researchers whose expertise and work are often underfunded. It is also essential for addressing systemic failures in scientific publishing that delay discoveries and leave important results unnoticed. In science and software, even one missing component can prevent an entire system from working.
Help valuable research and open-source infrastructure move forward. Please make a donation to support independent scientists and free software developers.
Disclaimer
Experimental-system notice: AI Internet-Meritocracy is an experimental funding system. Its AI-generated evaluations are heuristic judgments based on available public or connected-account evidence; they are not validated measurements of a person’s causal economic or scientific impact. Payment transactions are already recorded on-chain and can be verified on the blockchain. The current beta initiates payments off-chain through Node.js and uses custodial and administrative components. Decentralized governance and non-custodial wallets remain under development; on-chain payment records are already available. Evaluations may contain factual errors or biases and should be interpreted together with audit logs, appeals, human oversight, and published test results.
Ads:
| Description | Action |
|---|---|
|
A Brief History of Time
A landmark volume in science writing exploring cosmology, black holes, and the nature of the universe in accessible language. |
Check Price |
|
Astrophysics for People in a Hurry
Tyson brings the universe down to Earth clearly, with wit and charm, in chapters you can read anytime, anywhere. |
Check Price |
|
Raspberry Pi Starter Kits
Inexpensive computers designed to promote basic computer science education. Buying kits supports this ecosystem. |
View Options |
|
Free as in Freedom: Richard Stallman's Crusade
A detailed history of the free software movement, essential reading for understanding the philosophy behind open source. |
Check Price |
As an Amazon Associate I earn from qualifying purchases resulting from links on this page.