When Incentives Become the Moral LanguagePart I — The Need for Translation
Chapter 3 — The Paper That Must Count
The Paper That Must Count
The researcher has two folders open on her desk.
One contains a question she has been carrying for years: whether a particular intervention changes outcomes in a way existing studies claim but have never tested under the conditions where it would matter most. The design would be slow. The sample would be difficult. The result might be negative, inconclusive, or politically inconvenient within her field. If it succeeds, it could change how practitioners think. If it fails, it could still clarify what the field has been assuming without proof.
The other folder contains a project she knows how to finish. The methods are established. The journals that might take it are identifiable. The citations will accumulate in a recognizable pattern. The work is not trivial, but it is narrower than the question in the first folder, and narrowness has become a kind of professional shelter.
She is not choosing between truth and falsehood. She is choosing between two kinds of risk. The first risk is intellectual: that she spends years on work that does not resolve cleanly. The second is institutional: that she spends years on work her department cannot easily recognize when promotion committees, grant panels, and deans ask what she has produced.
A calendar reminder appears for the department's annual review. Beside it is a spreadsheet template asking for publications, grants, citations, and graduate students placed. The template is neutral, almost kindly in its formatting. It does not ask which question she believes the field most needs answered. It asks what can be counted.
On the shelf behind her, a row of journals faces outward, spines aligned by impact tier. She has learned, without anyone writing it down, which venues count as serious, which count as adequate, and which count as placeholders until something better appears. The shelf looks like a library. It functions like a map.
The Room Where Inquiry Used to Travel
Research institutions exist to produce reliable knowledge. That is a moral claim as much as a technical one. It says some questions are worth pursuing even when answers are slow, uncertain, or unpopular. For a long time, peer review, replication, and publication worked as a shared language for that claim. The system was never pure. It was biased by prestige, access, and politics. But it remained legible enough that scholars could argue about standards in common terms: evidence quality, method fit, explanatory power.
A journal article or dissertation was not only a career token. It was supposed to move understanding forward. Disagreement was expected. Delayed correction was often part of rigor, not proof that the field had failed. Reviewers could reject work for weak reasoning. Editors could ask for clarification. Replication could expose error. The process was slow, human, and often frustrating. It was also, at its best, a way of making judgment travel across people who did not share a lab.
As research scaled across countries, disciplines, and funding regimes, that shared language became harder to defend in institutions that no longer agreed on what counted as serious work—and that increasingly needed decisions they could compare across fields no individual administrator could master. A dean cannot read every paper in every department. A grant agency cannot evaluate every proposal on its own terms. A government minister cannot distinguish methodological nuance in thirty disciplines before allocating budgets. The institution needs something thinner than judgment and thicker than hope.
Judgment did not fracture because scholars stopped caring about truth. It fractured because truth-seeking became expensive to defend in systems that needed fast, comparable, auditable outputs. Disciplines specialized. Journals multiplied. Methods became less mutually intelligible across fields. University budgets tightened and competition intensified. Departments were compared by rankings, grant intake, and publication profile. Administrators needed metrics that traveled.
To say, "this line of inquiry is foundational but slow" asks a committee to trust judgment it cannot independently verify. To say, "this work has strong publication velocity and citation performance" gives the committee numbers it can defend publicly. So the system adapted toward the language that survived scrutiny.
The Journal That Ranks the Field
The journal did not set out to become a gatekeeper. It set out to curate. Someone had to decide which submissions merited the scarce pages of print, which arguments were ready for wider scrutiny, and which claims required more work before they entered the shared record. Peer review was supposed to be that decision made visible.
As publishing grew, the journal acquired a second function. It became a signal of quality to people who could not read the work itself. Hiring committees used venue prestige as shorthand. Funding panels treated certain titles as evidence of seriousness. Rankings aggregated journal reputations into department scores. Eugene Garfield's journal impact factor, invented to help librarians decide which subscriptions to keep, became one of the most influential proxies in modern science—a number attached to a venue rather than to an argument.1
The impact factor was never intended to rank individual researchers. It was absorbed into that role anyway because institutions needed a portable measure of influence. Tools like the h-index later compressed individual careers into comparable integers.2 The San Francisco Declaration on Research Assessment urged institutions to stop misusing journal metrics for hiring and promotion.3 Many universities signed it. Many committees continued to look at the journal shelf anyway, because the shelf offered what judgment could not: a defensible comparison across candidates who had pursued different questions in different fields with different timelines.
None of this required fraud. It required scale. Once enough decisions depend on a proxy, the proxy begins teaching the field what decisions are possible.
What the File Can Defend
The researcher closes the ambitious folder—not forever, she tells herself, but for this cycle—and opens the narrower one. The decision feels small at her desk. At scale, it is one of thousands of similar decisions that gradually teach a field what kinds of questions are worth attempting.
Publication volume became one of the first portable proxies. Hiring, promotion, and tenure processes increasingly treated count and venue pattern as primary evidence of scholarly contribution. A strong file meant a recognizable publication trajectory. Citation metrics followed, compressing complex influence into numbers easy to compare and difficult to interpret without context, but legible in institutional speech. Grant competitiveness added a third pressure. Funding criteria privilege feasibility, measurable outputs, and strategic fit.4 None of these tools claims to replace truth. Each claims to measure activity, impact, or productivity. Together they become a practical definition of merit because they are what institutions can repeat.
A "successful researcher" becomes, in operational terms, someone with steady output, strong citation performance, and competitive funding. That profile protects departments in rankings and audits. It is not identical to contribution, but it is easier to defend. The language performs the same move seen in hospital corridors and feeds: it shifts contested judgment into auditable criteria.
Once output metrics become the default language, certain kinds of work become harder to sustain. Negative results are harder to place. Replications are less rewarded than novelty in many fields. Long-horizon foundational projects can look unproductive inside annual review cycles. Questions that are methodologically important but less fashionable become harder to fund and publish. This does not mean the system is fake or that nothing is true. It means institutional incentives and epistemic priorities diverge in predictable ways.
The Committee Room
Months later, she sits in a conference room while a promotion committee discusses a colleague's file. The candidate's work is excellent by any measure she respects. The discussion, however, does not stay inside the work for long. It moves quickly to counts: articles, grants, citations, doctoral students completed, service rendered in forms the university recognizes.
Someone raises a concern that one strand of the research took years to mature and produced fewer papers than comparable candidates. Someone else replies that the venues were strong. A third person notes that the citation curve is rising. No one says the field is better because this colleague asked the questions she asked. They say the record is strong. The record is strong. The record is also a translation of a career into columns the committee can compare.
The researcher thinks of her two folders. She understands, with sudden clarity, that the committee is not failing. It is doing what large institutions must do when judgment cannot coordinate across specialties. It is trying to be fair. Fairness, here, means applying the same countable standards to people whose work is not commensurable. The fairness protects the institution. It also narrows what the institution can see.
What Replication Revealed
Replication debates in psychology, medicine, and adjacent fields made the narrowing visible. High-profile findings failed to reproduce at expected rates, exposing how novelty and positive-result bias can outrun reliability under publication pressure.5 John Ioannidis argued years earlier that the structure of modern research rewards exciting results over careful ones, producing a literature in which many published findings may not survive scrutiny.6 These were not revelations that science had failed. They were revelations that the translation layer between inquiry and recognition had developed incentives of its own.
Reforms followed, and many were genuine improvements. Preregistration, open data, stronger reporting standards, and replication initiatives asked researchers to make their methods visible earlier and their claims easier to test. Institutions added new badges, checklists, and compliance fields. The reforms were often absorbed into the same incentive environment they were meant to correct: another criterion to satisfy, another line on the annual report, another metric that could be optimized without changing which questions felt worth asking.
The system did not need to oppose truth to adopt a language that made some truths easier to pursue than others.
The Graduate Student's Inheritance
A graduate student appears in her office doorway with a draft proposal. The idea is promising but unwieldy. It would require a method the student is still learning, a collaboration that may take months to arrange, and a timeline that does not fit the department's expectation for progress toward candidacy. The researcher hears herself explaining how to make the project more fundable. She hears the student agree. She also hears what is not said: that the original question may survive only as a footnote to a safer version.
This is how fields change without anyone deciding to change them. Each adviser teaches the next generation what can be defended. Each generation learns which ambitions fit the gate. The student is not corrupted. The student is educated—educated into an institution that must speak in throughput if it is to allocate degrees, jobs, and grants across thousands of people who cannot all be known personally.
The danger is not only evaluation after the fact. Metrics shape what gets attempted in the first place. A lab learns which projects keep the lights on. A discipline learns which methods generate papers reviewers can process quickly. The institution does not need to forbid slow inquiry. It need only make slow inquiry difficult to defend.
What the Shelf Cannot Show
A citation count can tell us how often a paper entered other papers. It cannot tell us which question a researcher abandoned because the answer would take too long. A publication record can show what was completed. It cannot show the investigation that would have mattered but never became fundable. A strong h-index can indicate influence within a citation network. It cannot say whether that influence moved understanding in the direction the field most needed.
These absences do not invalidate the measures. They become dangerous only when absence is mistaken for emptiness—when the institution begins to treat what it can count as the boundary of what knowledge production was for. The missing work does not cease to exist. It survives in conversations after committee meetings, in folders kept open despite the spreadsheet, in drafts shared quietly with colleagues who understand why the official version had to be narrower, in the private unease of people who know the official account of merit is accurate and incomplete.
Institutions often describe this remainder as individual eccentricity, lack of focus, or failure to adapt. Sometimes it is those things. Sometimes it is the part of inquiry the instrument was not built to recognize.
The Hidden Draft
The researcher returns to the ambitious folder late at night, when the spreadsheet on her screen no longer demands an immediate answer. She is not naive about what the institution can reward this year. She also refuses to let the institution be the only place her curiosity lives.
She works on the narrower project because that project can pass through the gates: grants, committees, reviewers trained to expect a certain shape of result. She keeps the larger question alive in the margin—in pilot data, in an unfunded collaboration, in a seminar where she can still speak as if understanding mattered more than velocity. The department's metrics look healthy partly because people like her continue doing work the metrics cannot see until it succeeds enough to be counted.
When the hidden work fails, the failure is private. When it succeeds, the institution may claim it as evidence of excellence. Either way, the subsidy remains unnamed. This is one reason researchers can feel exhausted inside fields they still love. They are not only producing knowledge. They are maintaining two accounts of what their work is for—one the institution can repeat, and one they still believe the field needs.
The annual review ends. The spreadsheet is submitted. The narrower project moves forward on schedule. The ambitious folder remains open.
She does not experience this as heroism or betrayal. She experiences it as the ordinary condition of working inside a system that solved real coordination problems and, in solving them, taught her which questions could be asked in public.
Knowledge keeps moving. Meaning becomes harder to defend at scale.
The two folders will remain on her desk longer than she expected. One will produce the papers the shelf can recognize. The other will produce the question she is not willing to close. She has learned to live between them—the way the nurse in the hospital lives between criteria and readiness, the way the user on the platform lives between connection and capture. The institution sees one folder. She carries both.
Core Principle
What Institutions Can Recognize Eventually Shapes What People Learn to Pursue
Publication metrics began as ways to compare influence and activity. They became, for many institutions, the practical definition of scholarly merit. The translation solves real coordination problems across fields no committee can master. It also teaches researchers, slowly and without any single announcement, which questions are worth asking in public and which must be carried in the margin.
Footnotes
-
Eugene Garfield, "The History and Meaning of the Impact Factor," JAMA 295, no. 1 (2006): 90–93. Garfield argued that the journal impact factor was a tool for evaluating journals, not individual scientists. ↩
-
Jorge E. Hirsch, "An Index to Quantify an Individual's Scientific Research Output," Proceedings of the National Academy of Sciences 102, no. 46 (2005): 16569–16572. ↩
-
San Francisco Declaration on Research Assessment (DORA), 2012, https://sfdora.org/. ↩
-
National Science Foundation, Proposal & Award Policies & Procedures Guide (PAPPG), "Merit Review Principles and Criteria." ↩
-
Open Science Collaboration, "Estimating the Reproducibility of Psychological Science," Science 349, no. 6251 (2015). ↩
-
John P. A. Ioannidis, "Why Most Published Research Findings Are False," PLoS Medicine 2, no. 8 (2005): e124. ↩
