|
Getting your Trinity Audio player ready...
|
When AI makes scientific manuscripts cheap to generate, the scarce resource becomes the capacity to determine which claims are correct, original, and useful. If submissions grow faster than reliable assessment, evaluation becomes the bottleneck in science.
A second constraint appears when AI evaluates research: people checking suspected prompt injections become a bottleneck themselves. Automated systems can process documents rapidly, but human investigation, judgment, and appeals consume time that cannot be expanded simply by running more LLM instances.
These are conditional arguments about how research systems scale. They do not imply that all scientific work is becoming free, or that AI cannot improve evaluation.
Cheap papers are not the same as cheap discoveries
A scientific paper packages claims, methods, evidence, and interpretation. Producing that package is different from establishing that its contents deserve trust.
The 2024 AI Scientist preprint reported generating papers in selected machine-learning research settings for less than $15 per paper. This was a reported generation cost within a particular experimental system, not the total cost of independently validated discoveries across science. Its automated review scores should not be confused with independent confirmation of scientific importance. Nevertheless, the work illustrates how inexpensive manuscript production can become. Read the original AI Scientist study.
Experiments, equipment, fieldwork, data collection, and replication may remain expensive. In mathematics, a polished exposition still needs a correct proof. In computational research, convincing figures still require sound methods and reproducible results.
The important asymmetry is that producing another plausible manuscript may cost far less than carefully checking it.
Why research evaluation becomes the bottleneck
Evaluation involves several distinct questions:
| Question | What evaluators need to establish |
|---|---|
| Is it correct? | Whether the evidence or proof supports the claim |
| Is it original? | What it adds beyond existing work |
| Is it reproducible? | Whether others can check or repeat the relevant steps |
| Is it useful? | What understanding or capability it contributes |
| Who deserves credit? | Which people and earlier contributions enabled the result |
A paper can pass one test and fail another. Reproducible code can implement an inappropriate experiment. A correct theorem can duplicate an existing result. An important discovery can arrive in poorly written prose.
Consider an illustrative scenario: submissions increase tenfold while assistance halves the time needed to evaluate each submission. Total evaluation work still increases fivefold. Better review tools help, but their existence does not establish that review capacity will keep pace with production.
When funding or promotion rewards document counts, cheap generation also creates incentives to split contributions, repeat arguments, and submit minor variations. The system then spends scarce attention distinguishing additional knowledge from additional paperwork.
AI can help evaluate research—but its judgments need checking
AI assistance can organize claims, compare documents, identify possible inconsistencies, and suggest checks for reviewers. Its usefulness should be assessed task by task.
Some checks can rely on more constrained tools: reference lookup, software tests, statistical recalculation, or proof checking. These produce evidence about particular properties. They do not automatically establish the importance of a discovery or the fairness of a funding allocation.
Nature Portfolio’s peer-review policy keeps reviewers accountable for their reports and identifies limitations of generative AI, including false or biased output. This illustrates an important distinction between receiving assistance and transferring responsibility. Nature Portfolio peer-review policy.
The useful measure is therefore the cost of reaching a sufficiently reliable judgment, including corrections and appeals—not merely the cost of generating a review.
Human checks for prompt injection create another bottleneck
Prompt injection occurs when content supplied to an LLM attempts to redirect its behavior. In research evaluation, an instruction embedded in a paper, profile, or retrieved page might try to influence a score instead of providing scientific evidence. OWASP identifies both direct and indirect forms of this vulnerability. OWASP’s explanation of prompt injection.
Using another LLM to detect attacks can provide an additional filter, but it does not by itself establish an independent, reliable security boundary. OWASP recommends layered defenses, including restricted privileges, validation, monitoring, and human oversight for consequential operations. OWASP prompt-injection prevention guidance.
When people, rather than LLMs, make the final judgment about suspected manipulation, human review capacity limits how quickly those cases can be resolved.
A simple planning model makes this visible:
Human review hours = submissions × fraction referred to people × average hours per referred case.
For example, 100,000 submissions, a 1% referral rate, and 30 minutes per case require 500 human hours. These are illustrative assumptions, not measured AIIM figures. Multiple reviewers, appeals, and repeated attacks increase the workload further.
Referral volume includes false alarms. A legitimate security paper may quote malicious instructions as research material. Human reviewers must distinguish discussion of an attack from an attempt to manipulate the evaluator, sometimes with help from technical logs and controlled tests.
Nor does a low referral rate prove safety: the detector may simply miss attacks. Human review also has limitations, including fatigue, bias, and difficulty interpreting concealed inputs. Its effectiveness must be tested rather than assumed.
What this means for AI Internet-Meritocracy
AI Internet-Meritocracy (AIIM) is Science DAO’s experimental system for using AI assessments of documented research and software contributions to guide funding. Its evaluations are heuristic judgments, and improved fairness or efficiency remains something to demonstrate.
The case for human oversight is developed in Science DAO’s discussion of human independence in AI governance. Human adjudication may add a different perspective, but it also needs adequate capacity and protection against manipulation.
For AIIM, this suggests a concrete testing agenda: measure assessment quality, successful attacks, mistaken accusations, human time per dispute, appeal outcomes, and resulting payment errors. A system that scores contributions quickly but accumulates unresolved disputes has moved the bottleneck into governance.
These are proposed evaluation criteria, not claims that AIIM has already achieved them.
Fund the capacity to distinguish valuable work
A practical response is to treat evaluation as scientific infrastructure:
- Support verifiable evidence. Encourage accessible methods, data where appropriate, executable code, and checkable proofs.
- Reuse assessments. Preserve reviews and their evidence so unchanged claims do not require a complete restart.
- Allocate human attention carefully. Combine review of consequential or suspicious cases with random audits that can reveal missed problems.
- Fund reviewers and dispute resolution. Budget for expert time, security investigation, and appeals alongside research production.
- Reward contributions beyond document counts. Recognize replication, corrections, maintained software, and evidence that others can use.
The central opportunity is to turn cheaper production into more trustworthy knowledge. That requires evaluation capacity to grow alongside generation capacity—including the people who check whether automated evaluators are being manipulated.
Support Independent Science
Our flagship product, AI Internet-Meritocracy, is an experimental app designed to allocate donated funds to researchers and open-source developers using AI-assisted evaluation of documented contributions. Payments depend on available funds and eligibility requirements.
Help fund the proposed five-month public test of AIIM’s allocation model. Support the next testing milestone.
Supporting independent science is not only a matter of fairness to researchers whose expertise and work are often underfunded. It is also essential for addressing systemic failures in scientific publishing that delay discoveries and leave important results unnoticed. In science and software, even one missing component can prevent an entire system from working.
Help valuable research and open-source infrastructure move forward. Please make a donation to support independent scientists and free software developers.
Disclaimer
Experimental-system notice: AI Internet-Meritocracy is an experimental funding system. Its AI-generated evaluations are heuristic judgments based on available public or connected-account evidence; they are not validated measurements of a person’s causal economic or scientific impact. Payment transactions are already recorded on-chain and can be verified on the blockchain. The current beta initiates payments off-chain through Node.js and uses custodial and administrative components. Decentralized governance and non-custodial wallets remain under development; on-chain payment records are already available. Evaluations may contain factual errors or biases and should be interpreted together with audit logs, appeals, human oversight, and published test results.
Ads:
| Description | Action |
|---|---|
|
A Brief History of Time
A landmark volume in science writing exploring cosmology, black holes, and the nature of the universe in accessible language. |
Check Price |
|
Astrophysics for People in a Hurry
Tyson brings the universe down to Earth clearly, with wit and charm, in chapters you can read anytime, anywhere. |
Check Price |
|
Raspberry Pi Starter Kits
Inexpensive computers designed to promote basic computer science education. Buying kits supports this ecosystem. |
View Options |
|
Free as in Freedom: Richard Stallman's Crusade
A detailed history of the free software movement, essential reading for understanding the philosophy behind open source. |
Check Price |
As an Amazon Associate I earn from qualifying purchases resulting from links on this page.