{"id":24294,"date":"2026-07-17T11:03:40","date_gmt":"2026-07-17T11:03:40","guid":{"rendered":"https:\/\/science-dao.org\/?p=24294"},"modified":"2026-07-17T11:03:48","modified_gmt":"2026-07-17T11:03:48","slug":"researchers-or-work","status":"publish","type":"post","link":"https:\/\/science-dao.org\/hi\/researchers-or-work\/","title":{"rendered":"Should AI Evaluate Researchers, Research Outputs, or Both?"},"content":{"rendered":"<div id=\"scien-1347456376\" class=\"scien-before-content scien-entity-placement\"><style>\r\n.amazon-support-link {\r\n    display: inline-flex;\r\n    align-items: baseline;\r\n    gap: 0.3rem;\r\n    padding: 0.35rem 0.55rem;\r\n    color: inherit;\r\n    font-size: 0.88rem;\r\n    line-height: 1.2;\r\n    text-decoration: none;\r\n    opacity: 0.72;\r\n    white-space: nowrap;\r\n    transition: opacity 0.2s ease;\r\n}\r\n\r\n.amazon-support-link:hover,\r\n.amazon-support-link:focus-visible {\r\n    opacity: 1;\r\n    text-decoration: underline;\r\n}\r\n\r\n.amazon-support-link small {\r\n    font-size: 0.65rem;\r\n    opacity: 0.7;\r\n}\r\n\r\n@media (max-width: 900px) {\r\n    .amazon-support-link small {\r\n        display: none;\r\n    }\r\n}\r\n<\/style>\r\n<a class=\"amazon-support-link\"\r\n   href=\"https:\/\/www.amazon.com\/?tag=vpf04-20\"\r\n   target=\"_blank\"\r\n   rel=\"nofollow sponsored noopener\"\r\n   aria-label=\"Shop on Amazon and support World Science DAO\">\r\n    Shop on Amazon <span aria-hidden=\"true\">\u2197<\/span>\r\n    <small>affiliate link<\/small>\r\n<\/a><\/div>\n<p class=\"wp-block-paragraph\">AI should evaluate <strong>both research outputs and researchers\u2014but not in the same way or with equal weight<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Research outputs should be the primary unit of scientific evaluation. Papers, datasets, proofs, software, experimental results, replications, and other concrete contributions can be examined for quality, originality, rigor, usefulness, and reproducibility.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Researchers should be evaluated more cautiously. Person-level assessment may be necessary for hiring, grants, leadership roles, or access to long-term funding, but it introduces greater risks of prestige bias, historical lock-in, discrimination, and self-reinforcing rankings.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The best design is therefore a <strong>two-layer assessment system<\/strong>:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Evaluate each research output on its own merits.<\/li>\n\n\n\n<li>Construct a limited, contextual researcher profile from the assessed outputs and other documented contributions.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">AI should not begin with the question, \u201cIs this person an excellent scientist?\u201d It should begin with, \u201cWhat has this person produced, and how strong is the evidence that these outputs are valuable?\u201d<\/p>\n\n\n\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_85 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Content\"><span class=\"ez-toc-js-icon-con\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewbox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewbox=\"0 0 24 24\" version=\"1.2\" baseprofile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/span><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#The_Difference_Between_Evaluating_Research_and_Evaluating_Researchers\" >The Difference Between Evaluating Research and Evaluating Researchers<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Research-output_evaluation\" >Research-output evaluation<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Researcher-level_evaluation\" >Researcher-level evaluation<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Why_Research_Outputs_Should_Be_Evaluated_First\" >Why Research Outputs Should Be Evaluated First<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#It_reduces_prestige_bias\" >It reduces prestige bias<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#It_gives_independent_researchers_a_fairer_entry_point\" >It gives independent researchers a fairer entry point<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#It_makes_decisions_easier_to_audit\" >It makes decisions easier to audit<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#It_allows_scores_to_change_as_evidence_develops\" >It allows scores to change as evidence develops<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Why_AI_Cannot_Evaluate_Outputs_Alone\" >Why AI Cannot Evaluate Outputs Alone<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Future_projects_require_capability_assessment\" >Future projects require capability assessment<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Some_valuable_contributions_are_distributed_across_many_outputs\" >Some valuable contributions are distributed across many outputs<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Integrity_and_reliability_matter\" >Integrity and reliability matter<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-13\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#The_Danger_of_Turning_Researchers_Into_Scores\" >The Danger of Turning Researchers Into Scores<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-14\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Reputation_can_overwhelm_current_evidence\" >Reputation can overwhelm current evidence<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-15\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Historical_data_can_encode_historical_discrimination\" >Historical data can encode historical discrimination<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-16\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#A_universal_researcher_score_collapses_different_abilities\" >A universal researcher score collapses different abilities<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-17\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#What_Responsible_Research-Assessment_Frameworks_Suggest\" >What Responsible Research-Assessment Frameworks Suggest<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-18\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#A_Better_Hybrid_Architecture\" >A Better Hybrid Architecture<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-19\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Layer_One_Assess_Individual_Research_Outputs\" >Layer One: Assess Individual Research Outputs<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-20\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Layer_Two_Build_Contextual_Researcher_Profiles\" >Layer Two: Build Contextual Researcher Profiles<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-21\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Layer_Three_Match_Researchers_to_Specific_Opportunities\" >Layer Three: Match Researchers to Specific Opportunities<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-22\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Layer_Four_Update_Assessments_Over_Time\" >Layer Four: Update Assessments Over Time<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-23\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Layer_Five_Preserve_Human_Appeals_and_Adversarial_Review\" >Layer Five: Preserve Human Appeals and Adversarial Review<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-24\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Should_AI_Use_Citation_Counts\" >Should AI Use Citation Counts?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-25\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Should_Researcher_Identity_Be_Hidden\" >Should Researcher Identity Be Hidden?<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-26\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Identity_should_usually_be_hidden_during_initial_output_assessment\" >Identity should usually be hidden during initial output assessment<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-27\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Identity_may_be_introduced_for_verification\" >Identity may be introduced for verification<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-28\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Sensitive_attributes_should_not_become_ranking_variables\" >Sensitive attributes should not become ranking variables<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-29\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Preventing_AI_From_Rewarding_Quantity_Over_Quality\" >Preventing AI From Rewarding Quantity Over Quality<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-30\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Funding_New_Researchers_Without_a_Track_Record\" >Funding New Researchers Without a Track Record<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-31\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Evaluating_Collaborative_Research\" >Evaluating Collaborative Research<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-32\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Evaluating_Negative_Results_and_Replications\" >Evaluating Negative Results and Replications<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-33\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#The_Appropriate_Role_of_AIIM\" >The Appropriate Role of AIIM<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-34\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Recommended_Decision_Rule\" >Recommended Decision Rule<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-35\" href=\"https:\/\/science-dao.org\/hi\/researchers-or-work\/#Conclusion_Evaluate_Work_Directly_and_People_Contextually\" >Conclusion: Evaluate Work Directly and People Contextually<\/a><\/li><\/ul><\/nav><\/div>\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"The_Difference_Between_Evaluating_Research_and_Evaluating_Researchers\"><\/span>The Difference Between Evaluating Research and Evaluating Researchers<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The distinction may appear minor, but it changes the entire architecture of a funding or assessment system.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Research-output_evaluation\"><\/span>Research-output evaluation<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Output-level evaluation examines identifiable scientific objects, such as:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>journal articles and preprints;<\/li>\n\n\n\n<li>mathematical proofs;<\/li>\n\n\n\n<li>datasets;<\/li>\n\n\n\n<li>laboratory protocols;<\/li>\n\n\n\n<li>source code and scientific software;<\/li>\n\n\n\n<li>replications and negative results;<\/li>\n\n\n\n<li>peer reviews;<\/li>\n\n\n\n<li>theoretical frameworks;<\/li>\n\n\n\n<li>patents or practical applications;<\/li>\n\n\n\n<li>educational and research infrastructure.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The relevant question is:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>How scientifically valuable is this particular contribution?<\/strong><\/p>\n<\/blockquote>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Researcher-level_evaluation\"><\/span>Researcher-level evaluation<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Researcher-level evaluation attempts to estimate a person\u2019s broader capability or expected future contribution.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It may examine:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>the quality and consistency of previous work;<\/li>\n\n\n\n<li>technical expertise;<\/li>\n\n\n\n<li>successful completion of earlier projects;<\/li>\n\n\n\n<li>research integrity;<\/li>\n\n\n\n<li>collaboration and mentorship;<\/li>\n\n\n\n<li>openness and data-sharing practices;<\/li>\n\n\n\n<li>ability to identify important problems;<\/li>\n\n\n\n<li>contributions that do not appear in conventional publications.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The relevant question is:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>How likely is this researcher to produce valuable work under the proposed conditions?<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">These questions overlap, but they are not identical. An excellent researcher can produce a weak paper. An unknown researcher can produce an exceptional result.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Why_Research_Outputs_Should_Be_Evaluated_First\"><\/span>Why Research Outputs Should Be Evaluated First<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The strongest argument for output-first evaluation is epistemic: <strong>science advances through contributions, not reputations<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A theorem does not become true because its author works at a prestigious university. An experiment does not become reproducible because its principal investigator has a high h-index. A software package does not become reliable because it was developed by a famous laboratory.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The San Francisco Declaration on Research Assessment, commonly known as DORA, recommends evaluating research according to its scientific content rather than using the journal in which it appeared as a proxy for quality. DORA also emphasizes that research assessment should recognize outputs beyond conventional journal articles.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An output-first system provides several benefits.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"It_reduces_prestige_bias\"><\/span>It reduces prestige bias<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Traditional evaluation frequently relies on indirect signals:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>university affiliation;<\/li>\n\n\n\n<li>journal reputation;<\/li>\n\n\n\n<li>citation counts;<\/li>\n\n\n\n<li>academic rank;<\/li>\n\n\n\n<li>awards;<\/li>\n\n\n\n<li>previous grants;<\/li>\n\n\n\n<li>recommendations from influential scientists.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">These signals may contain information, but they can also reproduce existing hierarchies. Once a researcher receives early recognition, obtaining later recognition becomes easier. Conversely, researchers outside major institutions may remain invisible even when their work deserves examination.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">AI can reduce this problem by evaluating the content before revealing identity-related information. For example, an initial model could analyze a manuscript, proof, dataset, or repository without receiving the author\u2019s name, institution, country, academic title, or prior funding history.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This does not eliminate bias, because writing style, research topic, citations, and access to equipment may still reveal social information. Nevertheless, it creates a stronger barrier against direct prestige-based reasoning.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"It_gives_independent_researchers_a_fairer_entry_point\"><\/span>It gives independent researchers a fairer entry point<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Researcher-level evaluation often disadvantages people who lack conventional credentials. An independent mathematician, open-source developer, citizen scientist, or early-career researcher may have little institutional history to evaluate.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Output-level assessment allows such contributors to present something concrete:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>a proof that can be checked;<\/li>\n\n\n\n<li>software that can be tested;<\/li>\n\n\n\n<li>data that can be inspected;<\/li>\n\n\n\n<li>an experiment that can be replicated;<\/li>\n\n\n\n<li>a useful synthesis that can be compared with existing literature.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This is especially important for systems such as <a href=\"https:\/\/science-dao.org\/hi\/meritocracy\/\">AI Internet Meritocracy<\/a>, where funding is intended to follow demonstrated scientific value rather than institutional position alone.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"It_makes_decisions_easier_to_audit\"><\/span>It makes decisions easier to audit<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">A researcher score can be vague. Why was one person rated 83 and another 61?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An output assessment can be decomposed into narrower claims:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Was the methodology appropriate?<\/li>\n\n\n\n<li>Are the conclusions supported by the evidence?<\/li>\n\n\n\n<li>Is the work genuinely novel?<\/li>\n\n\n\n<li>Can the data be inspected?<\/li>\n\n\n\n<li>Does the code reproduce the reported results?<\/li>\n\n\n\n<li>Are mathematical steps valid?<\/li>\n\n\n\n<li>Does the output solve an important problem?<\/li>\n\n\n\n<li>Has the contribution been used by other work?<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Each claim can be connected to evidence, uncertainty, and reviewer disagreement. This makes automated evaluation more explainable and easier to contest.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"It_allows_scores_to_change_as_evidence_develops\"><\/span>It allows scores to change as evidence develops<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The value of a research output is not always known at publication.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A paper may initially appear important but later fail replication. A little-noticed software library may become essential infrastructure. A mathematical construction may acquire major applications years after publication.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Output-level evaluation can therefore be dynamic. Scores may change as new evidence appears:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>successful or failed replications;<\/li>\n\n\n\n<li>formal verification;<\/li>\n\n\n\n<li>corrections or retractions;<\/li>\n\n\n\n<li>downstream dependencies;<\/li>\n\n\n\n<li>citations with substantive context;<\/li>\n\n\n\n<li>use in clinical, industrial, or public-policy settings;<\/li>\n\n\n\n<li>incorporation into later proofs or theories.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The researcher\u2019s profile can then be updated from these revised output assessments rather than remaining tied to an early reputation.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Why_AI_Cannot_Evaluate_Outputs_Alone\"><\/span>Why AI Cannot Evaluate Outputs Alone<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Despite its advantages, a pure output-only model is incomplete.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Many funding decisions concern work that does not yet exist. A grant committee must decide whether a proposed project is credible before the final paper, dataset, or discovery is available.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Even retroactive funding systems must sometimes evaluate the people behind an output. Questions of authorship, integrity, responsibility, and capacity cannot always be answered by inspecting the artifact alone.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Future_projects_require_capability_assessment\"><\/span>Future projects require capability assessment<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Suppose two teams submit similar proposals for a technically difficult experiment. One team already operates the required equipment and has completed related work. The other has no demonstrated access to the necessary facilities.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Ignoring this difference would not make the process fairer. It would make the prediction less accurate.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Current NIH peer-review rules illustrate this distinction. The NIH framework considers both the importance and rigor of the proposed research and whether the investigators and research environment are sufficient to conduct it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Researcher-level evidence can therefore be relevant when it is connected to the requirements of a specific project.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Some_valuable_contributions_are_distributed_across_many_outputs\"><\/span>Some valuable contributions are distributed across many outputs<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">A researcher may contribute through:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>maintaining scientific software;<\/li>\n\n\n\n<li>curating datasets;<\/li>\n\n\n\n<li>reviewing others\u2019 work;<\/li>\n\n\n\n<li>mentoring junior researchers;<\/li>\n\n\n\n<li>organizing collaborations;<\/li>\n\n\n\n<li>designing standards;<\/li>\n\n\n\n<li>documenting failed approaches;<\/li>\n\n\n\n<li>developing research tools;<\/li>\n\n\n\n<li>identifying errors in influential work.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Evaluating each contribution separately is possible, but a person-level view may reveal sustained service or cumulative expertise that no single output captures.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Integrity_and_reliability_matter\"><\/span>Integrity and reliability matter<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Research evaluation is not only a ranking of ideas. Funding systems must also consider whether participants fulfill obligations.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Relevant evidence may include:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>whether promised data were released;<\/li>\n\n\n\n<li>whether previous funding was used as agreed;<\/li>\n\n\n\n<li>whether conflicts of interest were disclosed;<\/li>\n\n\n\n<li>whether errors were corrected;<\/li>\n\n\n\n<li>whether collaborators received appropriate credit;<\/li>\n\n\n\n<li>whether ethical and safety requirements were followed.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">These factors should not become an unrestricted personality score. They should be documented, appealable, and limited to behavior relevant to research responsibilities.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"The_Danger_of_Turning_Researchers_Into_Scores\"><\/span>The Danger of Turning Researchers Into Scores<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Evaluating researchers is much more dangerous than evaluating outputs because a person-level score can become a persistent label.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A low score may prevent someone from receiving the resources needed to produce better work. The absence of future work then appears to confirm the original score. This creates a feedback loop:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>The researcher receives a low rating.<\/li>\n\n\n\n<li>The researcher receives less funding and visibility.<\/li>\n\n\n\n<li>Fewer outputs are produced.<\/li>\n\n\n\n<li>The lack of outputs is interpreted as evidence of low ability.<\/li>\n\n\n\n<li>The rating falls further.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">This is particularly harmful for early-career researchers, people changing fields, researchers affected by illness or political instability, and contributors working outside universities.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Reputation_can_overwhelm_current_evidence\"><\/span>Reputation can overwhelm current evidence<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Once a global researcher score exists, evaluators may stop examining individual outputs carefully.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A weak article from a highly rated scientist may be accepted too easily. A strong article from a low-rated scientist may receive excessive scrutiny or be ignored.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That would reproduce the same failure found in journal-based assessment: replacing evaluation of scientific content with an easier proxy.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Historical_data_can_encode_historical_discrimination\"><\/span>Historical data can encode historical discrimination<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">AI models trained on past academic decisions may learn that successful researchers tend to come from particular institutions, countries, demographic groups, or professional networks.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Even when protected characteristics are removed, proxies may remain:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>postal location;<\/li>\n\n\n\n<li>institution names;<\/li>\n\n\n\n<li>writing conventions;<\/li>\n\n\n\n<li>career gaps;<\/li>\n\n\n\n<li>collaboration networks;<\/li>\n\n\n\n<li>publication venues;<\/li>\n\n\n\n<li>access to expensive equipment.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">An AI system may therefore reproduce prestige bias while presenting its decision as mathematically neutral.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"A_universal_researcher_score_collapses_different_abilities\"><\/span>A universal researcher score collapses different abilities<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Scientific ability is multidimensional.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A person may be outstanding at:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>proving theorems;<\/li>\n\n\n\n<li>designing experiments;<\/li>\n\n\n\n<li>writing reliable software;<\/li>\n\n\n\n<li>generating hypotheses;<\/li>\n\n\n\n<li>detecting errors;<\/li>\n\n\n\n<li>collecting difficult data;<\/li>\n\n\n\n<li>explaining complex results;<\/li>\n\n\n\n<li>coordinating interdisciplinary work.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Compressing all these abilities into one number discards information. A researcher who is unsuitable for one project may be ideal for another.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The system should therefore estimate <strong>task-specific capability<\/strong>, not claim to measure a person\u2019s universal scientific worth.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"What_Responsible_Research-Assessment_Frameworks_Suggest\"><\/span>What Responsible Research-Assessment Frameworks Suggest<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Major research-assessment initiatives generally reject simplistic rankings.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">DORA argues that scientific content should be assessed directly and that journal-level metrics should not substitute for evaluating individual work.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The Leiden Manifesto proposes that quantitative evaluation should support qualitative expert assessment rather than replace it. It also emphasizes field differences, transparency, and the need to examine indicators regularly for systemic effects.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The Coalition for Advancing Research Assessment, or CoARA, applies these principles to research, researchers, and research organizations. Its agreement calls for recognition of diverse outputs, activities, and practices, using qualitative judgment supported by responsible quantitative indicators.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">These frameworks were designed primarily around human assessment, but their principles apply equally to AI:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>AI-generated metrics should support reasoned scientific judgment, not replace scientific judgment with an opaque ranking.<\/strong><\/p>\n<\/blockquote>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"A_Better_Hybrid_Architecture\"><\/span>A Better Hybrid Architecture<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">An AI research-funding platform should maintain separate but connected layers of evaluation.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Layer_One_Assess_Individual_Research_Outputs\"><\/span>Layer One: Assess Individual Research Outputs<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Each output should receive a structured assessment rather than a single unexplained score.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Possible dimensions include:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Dimension<\/th><th>Central question<\/th><\/tr><\/thead><tbody><tr><td>Validity<\/td><td>Are the claims supported?<\/td><\/tr><tr><td>Rigor<\/td><td>Were appropriate methods used?<\/td><\/tr><tr><td>Novelty<\/td><td>What is genuinely new?<\/td><\/tr><tr><td>Reproducibility<\/td><td>Can others verify or repeat it?<\/td><\/tr><tr><td>Utility<\/td><td>Does it enable other research or applications?<\/td><\/tr><tr><td>Importance<\/td><td>How significant is the problem addressed?<\/td><\/tr><tr><td>Transparency<\/td><td>Are data, code, assumptions, and limitations available?<\/td><\/tr><tr><td>Robustness<\/td><td>Does the result survive alternative analyses or criticism?<\/td><\/tr><tr><td>Influence<\/td><td>Has it produced meaningful downstream use?<\/td><\/tr><tr><td>Uncertainty<\/td><td>How confident should evaluators be?<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Scores should be accompanied by evidence and explanations. When different AI agents or human reviewers disagree, the disagreement should be preserved rather than hidden through premature averaging.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Layer_Two_Build_Contextual_Researcher_Profiles\"><\/span>Layer Two: Build Contextual Researcher Profiles<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A researcher profile should summarize evidence from evaluated contributions without reducing the person to an immutable rank.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A profile might describe:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>demonstrated areas of expertise;<\/li>\n\n\n\n<li>types of outputs produced;<\/li>\n\n\n\n<li>reliability in completing previous projects;<\/li>\n\n\n\n<li>strengths relevant to a proposed task;<\/li>\n\n\n\n<li>open-science practices;<\/li>\n\n\n\n<li>collaboration history;<\/li>\n\n\n\n<li>unresolved concerns;<\/li>\n\n\n\n<li>uncertainty caused by limited data.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The profile should answer practical questions, such as:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Has this person demonstrated the skills required for the project?<\/li>\n\n\n\n<li>Has the researcher successfully completed similar work?<\/li>\n\n\n\n<li>Are additional collaborators or safeguards needed?<\/li>\n\n\n\n<li>Does the researcher have access to the required infrastructure?<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">It should not make unsupported declarations such as \u201cResearcher A is objectively better than Researcher B.\u201d<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Layer_Three_Match_Researchers_to_Specific_Opportunities\"><\/span>Layer Three: Match Researchers to Specific Opportunities<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The system should evaluate the relationship among three entities:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>the researcher or team;<\/li>\n\n\n\n<li>the proposed work;<\/li>\n\n\n\n<li>the funding opportunity.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">A scientist\u2019s suitability is conditional. A theoretical mathematician and a laboratory chemist cannot be ranked meaningfully on one universal scale, but each can be assessed for a relevant project.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This matching layer should consider:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>required expertise;<\/li>\n\n\n\n<li>available equipment;<\/li>\n\n\n\n<li>project scale;<\/li>\n\n\n\n<li>collaboration requirements;<\/li>\n\n\n\n<li>time horizon;<\/li>\n\n\n\n<li>uncertainty tolerance;<\/li>\n\n\n\n<li>previous evidence;<\/li>\n\n\n\n<li>possible conflicts of interest.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This approach is closer to scientific resource allocation than to social ranking.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Layer_Four_Update_Assessments_Over_Time\"><\/span>Layer Four: Update Assessments Over Time<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Assessments should be revisable.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">New evidence may include:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>a correction;<\/li>\n\n\n\n<li>an independent replication;<\/li>\n\n\n\n<li>an identified flaw;<\/li>\n\n\n\n<li>new software adoption;<\/li>\n\n\n\n<li>practical application;<\/li>\n\n\n\n<li>a successful follow-up project;<\/li>\n\n\n\n<li>discovery that an output was derivative;<\/li>\n\n\n\n<li>clarification of an authorship dispute.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Both output assessments and researcher profiles should show version histories. Users should be able to see what changed, when it changed, and what evidence caused the revision.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Layer_Five_Preserve_Human_Appeals_and_Adversarial_Review\"><\/span>Layer Five: Preserve Human Appeals and Adversarial Review<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Automated systems will make mistakes. Researchers must therefore be able to:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>inspect the evidence used;<\/li>\n\n\n\n<li>challenge factual errors;<\/li>\n\n\n\n<li>identify missing outputs;<\/li>\n\n\n\n<li>contest inappropriate comparisons;<\/li>\n\n\n\n<li>disclose unusual circumstances;<\/li>\n\n\n\n<li>request independent reassessment;<\/li>\n\n\n\n<li>submit counterevidence.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">High-impact systems should also undergo systematic <a href=\"https:\/\/science-dao.org\/hi\/adversarial\/\">adversarial testing<\/a> to identify manipulation, hidden bias, fabricated evidence, and unstable scoring behavior.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The objective is not to prove that an AI evaluator is infallible. It is to make its errors discoverable and correctable.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Should_AI_Use_Citation_Counts\"><\/span>Should AI Use Citation Counts?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Citation counts can provide evidence, but they should not define either output quality or researcher quality.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Citations may indicate attention or influence, but they can be distorted by:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>field size;<\/li>\n\n\n\n<li>publication age;<\/li>\n\n\n\n<li>self-citation;<\/li>\n\n\n\n<li>citation cartels;<\/li>\n\n\n\n<li>fashionable topics;<\/li>\n\n\n\n<li>negative citations;<\/li>\n\n\n\n<li>review articles accumulating more citations than original discoveries;<\/li>\n\n\n\n<li>database coverage;<\/li>\n\n\n\n<li>language and regional inequalities.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">A better system should analyze <strong>citation context<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It should distinguish among citations that:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>use a result;<\/li>\n\n\n\n<li>confirm it;<\/li>\n\n\n\n<li>extend it;<\/li>\n\n\n\n<li>criticize it;<\/li>\n\n\n\n<li>mention it only as background;<\/li>\n\n\n\n<li>copy citations from another source;<\/li>\n\n\n\n<li>depend on associated data or software.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">A small number of substantive downstream uses may be more important than hundreds of superficial mentions.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Should_Researcher_Identity_Be_Hidden\"><\/span>Should Researcher Identity Be Hidden?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Identity should sometimes be hidden, but not always.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Identity_should_usually_be_hidden_during_initial_output_assessment\"><\/span>Identity should usually be hidden during initial output assessment<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Removing names and affiliations can reduce direct prestige bias. The first-pass evaluator can focus on methods, evidence, reasoning, and contribution.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Identity_may_be_introduced_for_verification\"><\/span>Identity may be introduced for verification<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Later stages may require identity information to determine:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>authorship;<\/li>\n\n\n\n<li>conflicts of interest;<\/li>\n\n\n\n<li>access to facilities;<\/li>\n\n\n\n<li>previous fulfillment of grants;<\/li>\n\n\n\n<li>ethical approvals;<\/li>\n\n\n\n<li>relevant technical experience.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Sensitive_attributes_should_not_become_ranking_variables\"><\/span>Sensitive attributes should not become ranking variables<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Nationality, religion, ethnicity, gender, disability, and political affiliation should not be treated as evidence of scientific quality.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">However, simply deleting these fields is not enough. Systems must also test whether proxy variables recreate the same discrimination.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Preventing_AI_From_Rewarding_Quantity_Over_Quality\"><\/span>Preventing AI From Rewarding Quantity Over Quality<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">An AI evaluator may be easy to manipulate if it rewards the number of outputs.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Researchers could split one result into many publications, generate low-value papers, create superficial datasets, or produce large volumes of AI-written material.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The solution is not merely to cap publication counts. The system should recognize relationships among outputs:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>duplicate or overlapping work;<\/li>\n\n\n\n<li>incremental extensions;<\/li>\n\n\n\n<li>dependency chains;<\/li>\n\n\n\n<li>consolidated research programs;<\/li>\n\n\n\n<li>shared datasets;<\/li>\n\n\n\n<li>reused code;<\/li>\n\n\n\n<li>genuine independent contributions.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Marginal value matters. The tenth nearly identical paper should not receive the same reward as the foundational contribution that enabled the entire series.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For related safeguards, see <a href=\"https:\/\/science-dao.org\/hi\/ai-alignment\/\">Preventing the Prompt-Gaming Problem<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Funding_New_Researchers_Without_a_Track_Record\"><\/span>Funding New Researchers Without a Track Record<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A hybrid model must avoid making previous success a prerequisite for all future success.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">New researchers could be evaluated through:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>small initial grants;<\/li>\n\n\n\n<li>staged funding;<\/li>\n\n\n\n<li>technical work samples;<\/li>\n\n\n\n<li>preliminary results;<\/li>\n\n\n\n<li>open research plans;<\/li>\n\n\n\n<li>prediction or replication tasks;<\/li>\n\n\n\n<li>collaboration with established infrastructure;<\/li>\n\n\n\n<li>milestone-based release of funds.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This approach gives unknown researchers an entry path without requiring funders to ignore uncertainty.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The system might allocate a limited exploratory budget first, then increase funding when verifiable outputs appear. This is more informative than rejecting applicants simply because they lack conventional credentials.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Evaluating_Collaborative_Research\"><\/span>Evaluating Collaborative Research<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Research outputs often have many authors, and contribution is rarely equal.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">AI should not divide credit mechanically by author count. Nor should it assume that the first or final author always made the most important contribution.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Where evidence is available, assessment may consider:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>contributor-role statements;<\/li>\n\n\n\n<li>repository commits;<\/li>\n\n\n\n<li>experimental records;<\/li>\n\n\n\n<li>dataset provenance;<\/li>\n\n\n\n<li>protocol authorship;<\/li>\n\n\n\n<li>mathematical sections or proofs;<\/li>\n\n\n\n<li>project-management responsibilities;<\/li>\n\n\n\n<li>statements confirmed by collaborators.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Even then, contribution estimates should include uncertainty. Private intellectual work and informal collaboration are difficult to reconstruct reliably.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Output evaluation and contribution attribution should therefore remain separate:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>The system should first estimate the value of the output, then estimate who contributed what.<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Combining these questions too early can distort both.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Evaluating_Negative_Results_and_Replications\"><\/span>Evaluating Negative Results and Replications<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Person-centered academic systems often reward novelty and visibility. This can undervalue failed experiments, null results, replications, corrections, and error detection.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An output-centered AI system could recognize their actual utility.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A well-conducted experiment showing that a promising hypothesis is false may prevent many laboratories from wasting resources. A replication may establish whether an influential result is dependable. A correction may protect an entire research field from building on an error.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The assessment should ask what information the output adds\u2014not whether it produces an exciting headline.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"The_Appropriate_Role_of_AIIM\"><\/span>The Appropriate Role of AIIM<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">AI Internet Meritocracy should not operate as a machine that permanently sorts scientists from \u201cbest\u201d to \u201cworst.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Its stronger role is to organize evidence about scientific contributions and allocate resources according to transparent, contestable criteria.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A practical AIIM design could:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>evaluate research outputs through multiple specialized AI agents;<\/li>\n\n\n\n<li>estimate novelty, rigor, reproducibility, utility, and uncertainty separately;<\/li>\n\n\n\n<li>collect evidence of downstream scientific use;<\/li>\n\n\n\n<li>attribute contributions cautiously;<\/li>\n\n\n\n<li>construct task-specific researcher profiles;<\/li>\n\n\n\n<li>match researchers and teams to suitable funding opportunities;<\/li>\n\n\n\n<li>permit human challenges and community review;<\/li>\n\n\n\n<li>revise decisions when new evidence appears.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">This model represents <strong>evidence-based merit allocation<\/strong>, not algorithmic social hierarchy.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Recommended_Decision_Rule\"><\/span>Recommended Decision Rule<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The appropriate balance depends on the type of decision.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Decision<\/th><th>Primary emphasis<\/th><th>Secondary emphasis<\/th><\/tr><\/thead><tbody><tr><td>Retroactive reward<\/td><td>Research output<\/td><td>Contribution attribution<\/td><\/tr><tr><td>Publication review<\/td><td>Research output<\/td><td>Relevant conflicts or integrity<\/td><\/tr><tr><td>Replication funding<\/td><td>Proposed method<\/td><td>Team capability<\/td><\/tr><tr><td>Early-stage grant<\/td><td>Proposal and preliminary work<\/td><td>Task-specific researcher evidence<\/td><\/tr><tr><td>Large infrastructure grant<\/td><td>Project design<\/td><td>Team, governance, and execution record<\/td><\/tr><tr><td>Hiring<\/td><td>Relevant outputs and skills<\/td><td>Broader professional contributions<\/td><\/tr><tr><td>Scientific prize<\/td><td>Specific contribution<\/td><td>Attribution and historical context<\/td><\/tr><tr><td>Long-term institutional funding<\/td><td>Portfolio of outputs<\/td><td>Reliability, leadership, and infrastructure<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">There is no defensible reason to use the same weighting for every decision.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Conclusion_Evaluate_Work_Directly_and_People_Contextually\"><\/span>Conclusion: Evaluate Work Directly and People Contextually<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">AI should evaluate both researchers and research outputs, but <strong>research outputs must remain the evidential foundation<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Output-level assessment is better suited to testing validity, novelty, reproducibility, and scientific utility. It reduces reliance on prestige and creates opportunities for independent and early-career researchers.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Researcher-level assessment remains useful when decisions concern future execution, integrity, coordination, or accumulated expertise. However, it should be contextual, multidimensional, revisable, and subordinate to concrete evidence.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The governing principle should be:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>Evaluate scientific work directly. Evaluate researchers only for a defined purpose, using evidence from their work and behavior relevant to that purpose.<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">AI should help science move from reputation-based judgment toward transparent evaluation. It should not replace one academic hierarchy with a more opaque algorithmic one.<\/p>\n<div id=\"scien-879273614\" class=\"scien-after-content scien-entity-placement\"><section>\r\n\r\n<h2>Support Independent Science<\/h2>\r\n\r\n<p>Supporting independent science is not only a matter of fairness to researchers whose expertise and work are often underfunded. It is also essential for addressing <a href=\"https:\/\/science-dao.org\/hi\/who-are-science-marketers\/\">systemic failures in scientific publishing<\/a> that delay discoveries and leave important results unnoticed. In science and software, even one missing component can prevent an entire system from working.<\/p>\r\n\r\n<p><strong>Help valuable research and open-source infrastructure move forward.<\/strong> Please <strong><a href=\"https:\/\/science-dao.org\/hi\/donation\/\">make a donation<\/a><\/strong> to support <a href=\"https:\/\/science-dao.org\/hi\/amateur-scientists\/\">independent scientists<\/a> and <a href=\"https:\/\/science-dao.org\/hi\/free-software\/\">free software developers<\/a>.<\/p>\r\n\r\n<p>\r\n\t\tOur flagship product is <a href=\"https:\/\/science-dao.org\/hi\/meritocracy\/\">AI Internet-Meritocracy<\/a> - an app, that unlike universities distributes money directly to researchers and open source developers, without bureaucracy.\r\n<\/p>\r\n\r\n<\/section><\/div><div id=\"scien-2045048886\" class=\"scien-after-content-2 scien-entity-placement\"><div data-nosnippet style=\"max-width: 800px\">\r\n<p style=\"margin-bottom: 0\">Ads:<\/p>\r\n<style>\r\n    \/* Compact Table Styling *\/\r\n    .amazon-ad-table {\r\n        width: 100%;\r\n        max-width: 800px;\r\n        border-collapse: collapse;\r\n        margin: 10px auto;\r\n        font-family: Arial, sans-serif;\r\n        border: 1px solid #e0e0e0;\r\n    }\r\n    .amazon-ad-table th {\r\n        background-color: #f3f3f3;\r\n        padding: 8px;\r\n        text-align: left;\r\n        border-bottom: 2px solid #ddd;\r\n        font-size: 0.9em;\r\n    }\r\n    .amazon-ad-table td {\r\n        padding: 8px 5px;\r\n        border-bottom: 1px solid #e0e0e0;\r\n        vertical-align: middle;\r\n    }\r\n    \/* Column sizing *\/\r\n    .col-image { width: 15%; text-align: center; vertical-align: top; padding-top: 10px; }\r\n    .col-desc { width: 65%; }\r\n    .col-action { width: 20%; text-align: center; }\r\n\r\n    \/* Image Placeholder Styling *\/\r\n    \/* YOU WILL REPLACE THIS ENTIRE BLOCK WHEN YOU INSERT REAL AMAZON CODE *\/\r\n    .img-placeholder {\r\n        width: 80px;\r\n        height: 110px;\r\n        background-color: #eee;\r\n        border: 1px solid #ddd;\r\n        color: #666;\r\n        font-size: 0.7em;\r\n        display: flex;\r\n        justify-content: center;\r\n        align-items: center;\r\n        margin: 0 auto;\r\n        text-align: center;\r\n    }\r\n    \r\n    \/* Typography - Smaller fonts and tighter line heights *\/\r\n    .product-title {\r\n        font-size: 1em;\r\n        font-weight: bold;\r\n        color: #007185;\r\n        text-decoration: none;\r\n        display: block;\r\n        margin-bottom: 2px;\r\n    }\r\n    .product-title:hover { color: #C7511F; text-decoration: underline; }\r\n    .product-author { \r\n        color: #565959; \r\n        font-size: 0.8em; \r\n        margin-bottom: 4px; \r\n    }\r\n    .product-blurb { \r\n        font-size: 0.85em; \r\n        line-height: 1.25;\r\n        color: #333; \r\n        margin: 0;\r\n    }\r\n\r\n    \/* Compact Button Styling *\/\r\n    .amazon-button {\r\n        display: inline-block;\r\n        background-color: #FFD814;\r\n        border: 1px solid #FCD200;\r\n        border-radius: 20px;\r\n        color: #0F1111;\r\n        padding: 6px 10px;\r\n        text-align: center;\r\n        text-decoration: none;\r\n        font-size: 0.8em;\r\n        font-weight: bold;\r\n        box-shadow: 0 2px 5px rgba(0,0,0,0.1);\r\n        transition: background-color 0.2s;\r\n        white-space: nowrap;\r\n    }\r\n    .amazon-button:hover { background-color: #F7CA00; border-color: #F2C200; cursor: pointer;}\r\n\r\n    \/* Disclosure *\/\r\n    .affiliate-disclosure {\r\n        font-size: 0.75em;\r\n        color: #565959;\r\n        text-align: center;\r\n        margin-top: 5px;\r\n    }\r\n\r\n    \/* Responsive *\/\r\n    @media (max-width: 600px) {\r\n        .amazon-ad-table thead { display: none; }\r\n        .amazon-ad-table tr { display: flex; flex-direction: column; border-bottom: 2px solid #ddd; padding: 10px; }\r\n        .amazon-ad-table td { width: 100%; border: none; padding: 5px 0; text-align: center; }\r\n        .col-desc { text-align: center; }\r\n    }\r\n<\/style>\r\n<table class=\"amazon-ad-table\">\r\n    <thead>\r\n        <tr>\r\n            <th>Description<\/th>\r\n            <th>Action<\/th>\r\n        <\/tr>\r\n    <\/thead>\r\n    <tbody>\r\n        <tr>\r\n            <td class=\"col-desc\">\r\n                <a href=\"https:\/\/www.amazon.com\/dp\/0553380168?tag=vpf04-20\" class=\"product-title\" target=\"_blank\" rel=\"nofollow noopener\">A Brief History of Time<\/a>\r\n                <div class=\"product-author\">by Stephen Hawking<\/div>\r\n                <p class=\"product-blurb\">A landmark volume in science writing exploring cosmology, black holes, and the nature of the universe in accessible language.<\/p>\r\n            <\/td>\r\n            <td class=\"col-action\">\r\n                <a href=\"https:\/\/www.amazon.com\/dp\/0553380168?tag=vpf04-20\" class=\"amazon-button\" target=\"_blank\" rel=\"nofollow noopener\">Check Price<\/a>\r\n            <\/td>\r\n        <\/tr>\r\n\r\n        <tr>\r\n            <td class=\"col-desc\">\r\n                <a href=\"https:\/\/www.amazon.com\/dp\/0393609391?tag=vpf04-20\" class=\"product-title\" target=\"_blank\" rel=\"nofollow noopener\">Astrophysics for People in a Hurry<\/a>\r\n                <div class=\"product-author\">by Neil deGrasse Tyson<\/div>\r\n                <p class=\"product-blurb\">Tyson brings the universe down to Earth clearly, with wit and charm, in chapters you can read anytime, anywhere.<\/p>\r\n            <\/td>\r\n            <td class=\"col-action\">\r\n                <a href=\"https:\/\/www.amazon.com\/dp\/0393609391?tag=vpf04-20\" class=\"amazon-button\" target=\"_blank\" rel=\"nofollow noopener\">Check Price<\/a>\r\n            <\/td>\r\n        <\/tr>\r\n\r\n         <tr>\r\n            <td class=\"col-desc\">\r\n                <a href=\"https:\/\/www.amazon.com\/s?k=raspberry+pi+4+starter+kit&tag=vpf04-20\" class=\"product-title\" target=\"_blank\" rel=\"nofollow noopener\">Raspberry Pi Starter Kits<\/a>\r\n                <div class=\"product-author\">Supports Computer Science Education<\/div>\r\n                <p class=\"product-blurb\">Inexpensive computers designed to promote basic computer science education. Buying kits supports this ecosystem.<\/p>\r\n            <\/td>\r\n            <td class=\"col-action\">\r\n                <a href=\"https:\/\/www.amazon.com\/s?k=raspberry+pi+4+starter+kit&tag=vpf04-20\" class=\"amazon-button\" target=\"_blank\" rel=\"nofollow noopener\">View Options<\/a>\r\n            <\/td>\r\n        <\/tr>\r\n\r\n        <tr>\r\n            <td class=\"col-desc\">\r\n                <a href=\"https:\/\/www.amazon.com\/dp\/0596002874?tag=vpf04-20\" class=\"product-title\" target=\"_blank\" rel=\"nofollow noopener\">Free as in Freedom: Richard Stallman's Crusade<\/a>\r\n                <div class=\"product-author\">by Sam Williams<\/div>\r\n                <p class=\"product-blurb\">A detailed history of the free software movement, essential reading for understanding the philosophy behind open source.<\/p>\r\n            <\/td>\r\n            <td class=\"col-action\">\r\n                <a href=\"https:\/\/www.amazon.com\/dp\/0596002874?tag=vpf04-20\" class=\"amazon-button\" target=\"_blank\" rel=\"nofollow noopener\">Check Price<\/a>\r\n            <\/td>\r\n        <\/tr>\r\n    <\/tbody>\r\n<\/table>\r\n<div class=\"affiliate-disclosure\">\r\n    <p>As an Amazon Associate I earn from qualifying purchases resulting from links on this page.<\/p>\r\n<\/div>\r\n<\/div><\/div>","protected":false},"excerpt":{"rendered":"<p>AI should evaluate both research outputs and researchers\u2014but not in the same way or with equal weight. Research outputs should be the primary unit of scientific evaluation. Papers, datasets, proofs, [&hellip;]<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-24294","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/posts\/24294","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/comments?post=24294"}],"version-history":[{"count":1,"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/posts\/24294\/revisions"}],"predecessor-version":[{"id":24295,"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/posts\/24294\/revisions\/24295"}],"wp:attachment":[{"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/media?parent=24294"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/categories?post=24294"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/science-dao.org\/hi\/wp-json\/wp\/v2\/tags?post=24294"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}