Research Notes
All views expressed here are my own.
Research Practice / Scientific Integrity
Beyond Acceptance: A Discipline for Durable AI Research
Top-conference acceptance is worth pursuing, but it remains a noisy, gameable proxy inside a larger incentive system. These notes develop a reward architecture, research portfolio, operating system, and agent-era governance model for producing work that survives the deadline, the review cycle, and the benchmark.
The goal is not to stop seeking reward. It is to refuse a reward function simple enough to game.
When a Useful Ambition Becomes the Objective Function
Before every deadline, two questions quietly compete: what is true, and what can be made legible enough to score? They often align. A clear contribution, fair evaluation, strong writing, and thoughtful venue fit can improve both the paper and its chance of acceptance. Wanting a top-conference paper is therefore not a failure of scientific values. Conferences concentrate feedback, attention, collaborators, career opportunities, and a deadline that can turn unfinished thinking into a public contribution.
The danger begins when anticipated review scores stop being feedback about the research and become the objective function of the research. The question shifts from “What uncertainty is worth reducing?” to “What contribution will look new?” The benchmark stops representing the construct and becomes the target. Evidence is selected for narrative neatness. Boundary conditions become inconvenient. The project may become increasingly optimized for acceptance while becoming less capable of surviving contact with reality.
This is a form of proxy capture. Work on Goodhart effects distinguishes several ways in which a measure can lose its relationship to the goal under optimization pressure. In research, the pressure need not produce one dramatic act of misconduct. It can instead operate through a sequence of locally defensible choices that all lean in the same direction.
- Seeking
- Pursuing acceptance, visibility, and recognition. This is normal ambition and can create useful discipline.
- Proxy optimization
- Improving a proxy—reviewer familiarity, a leaderboard score, or novelty language—until it starts displacing the underlying scientific objective.
- Hacking
- Exploiting blind spots in evaluation to obtain the signal without earning the substance the signal is supposed to represent.
- Self-hacking
- Training my own taste to recognize only what looks publishable, until I stop seeing questions whose value is difficult to compress into one review cycle.
I am borrowing the language of reward hacking from AI systems, but the distinctions matter. Most everyday drift is unintentional proxy optimization, not misconduct. Reward hacking is the stronger case in which an evaluation blind spot is used to obtain the signal without the corresponding substance. “Self-hacking” is my informal name for the quieter process by which the incentive changes what I learn to notice and value.
How Good Researchers Accidentally Game Themselves
This is not only a problem of individual character. Smaldino and McElreath's model of the natural selection of bad science shows how incentives for publication output can favor weaker methods even without deliberate cheating. That is why good intentions are not a sufficient control. The research process needs structures that make distortion harder and honest correction easier.
Proxy optimization can happen without fraud or conscious shortcutting. An experiment is rerun on several seeds, but only the stable-looking setting enters the main table. A weak baseline is tuned less carefully than the proposed method. A small benchmark gain becomes a broad claim about capability. The story is decided first and the ablations are chosen to support it. A deadline forces the experiment to shrink, but the claim expands to preserve excitement. Complexity is interpreted as depth because it is easier to display than understanding.
Machine learning makes this especially easy because benchmarks give precise feedback while many true goals—robustness, explanatory depth, usefulness, and downstream consequence—remain partially observed. The benchmark lottery shows how choices of tasks and evaluation setup can change apparent algorithmic superiority. Leaderboards can clarify progress, but their precision should not be confused with completeness.
AI research agents increase both sides of this problem. They can search, code, experiment, and write faster, but they can also industrialize plausible shallow work. More runs can hide weaker experimental taste; smoother prose can hide a gap between claim and evidence; a large generated pipeline can exceed the understanding of every author. Automation should increase the amount of verification we can afford, not reduce the amount of understanding we require.
The Reward Function Is Nested
The research system does not have one reward function. It has a stack of lossy translations, and what is easy to observe can become more powerful at every layer.
A researcher does not optimize in isolation. A student responds to graduation, hiring, and limited time, while a lab responds to grants, reputation, and the need to sustain a team. A conference responds to reviewer capacity, annual deadlines, and community expectations; a field responds to citations, benchmarks, and concentrated attention. An agent responds to whatever objective can be stated and checked. Each layer passes a compressed version of value to the next. Even when every participant acts reasonably, their combined effects can favor work that is fast, legible, and easy to score over work that is uncertain, slow to validate, or difficult to package.
- Field
- Citations, leaderboards, and attention can hide question diversity, correction, maintenance, and work whose value appears slowly.
- Conference
- Scores, novelty, and reviewer agreement can underweight replication, negative findings, long validation cycles, and unfamiliar contribution types.
- Lab
- Paper pipelines, grants, and placements can crowd out mentoring, shared infrastructure, risk diversity, and stewardship after publication.
- Researcher
- Acceptance probability and visible novelty can displace understanding, capability growth, patient question selection, and honest stopping.
- Agent
- Task completion, metric deltas, and plausible text can conceal provenance gaps, correlated errors, weak judgment, and unowned scientific claims.
This makes proxy capture an ecological problem, not only a character problem. The researcher may optimize rationally for the lab, the lab for the conference, and the conference for the field's current signals. No actor needs to intend distortion for the system to select a highly publishable research phenotype. Personal discipline is still necessary, but it is the first control surface—not the whole solution.
It would also be a mistake to treat conference optimization as inherently corrupt. Deadlines force completion. Common benchmarks create comparability. Strong narrative exposes incoherent thinking. Venue fit helps the right community encounter the work. Communication is part of scientific value because knowledge that cannot be inspected or understood cannot travel. The useful boundary is whether optimization improves clarity, comparability, testability, or completion without changing the truth conditions of the claim. Capture begins when evidence standards, claim scope, reporting choices, or the research question itself are bent around the anticipated score.
Agency is also unequal. A junior researcher facing PI pressure, visa constraints, compute scarcity, or a narrow hiring window may not be free to delay submission, publish a null result, or unilaterally narrow an attractive claim. Self-governance should therefore not become a moral judgment of people with less power. The corresponding obligation falls upward as well: senior authors and institutions should create escalation paths, protect dissent, allocate time for verification, and share responsibility for the incentives they impose.
Acceptance Is Not a Scalar Reward
A binary accept–reject result compresses a high-dimensional research process into one consequential bit. That bit matters, but it cannot tell me whether the question was important, the central claim is true, the evidence will replicate, the artifact will be reused, or the project made me a better researcher. I want to read the decision alongside several independent scoreboards rather than let it overwrite them.
My current principle: acceptance is an external signal inside the reward architecture, not the reward architecture itself.
Venue & communication
- Is the problem legible and relevant to the community?
- Can reviewers identify the contribution and test its claims?
- Did the work enter the right scientific conversation?
Science & evidence
- Is the central claim falsifiable and proportional to the evidence?
- Does it survive strong baselines, reruns, and adversarial checks?
- Are mechanisms, uncertainty, and boundary conditions understood?
Endurance & community value
- Can others reuse the artifact, measurement, or abstraction?
- Does the insight survive changes in benchmarks or trends?
- Does it generate better questions rather than only variants?
Researcher growth
- What technical capability and judgment became genuinely mine?
- Can I reproduce and explain the core result independently?
- Did the project improve how I ask, test, debug, and reason?
These scoreboards should not be combined into one larger weighted sum. That would merely create a more elaborate proxy. Some conditions are non-compensable: ten attractive experiments cannot rescue leakage in the evaluation; elegant writing cannot compensate for a claim that the evidence does not support; acceptance cannot retroactively make a fragile result robust. The reward architecture therefore needs three layers.
1 · Non-compensable constraints
- Evidence integrity, no hidden leakage, and traceable deviations
- Claims proportional to the evidence and risk
- Ownership of every critical evidence path
- Reproducibility appropriate to the stakes of the claim
2 · Value vector
- Epistemic gain, explanatory compression, and reliability
- Reusability, researcher capability, and community legibility
- Portfolio optionality and venue calibration
- No dimension silently erases damage to another
The third layer is time. Each claimed value should be marked as observed, provisional, or unknown across immediate, medium, and long horizons. Acceptance is observed evidence of crossing a venue threshold now; five-year durability is unknowable at submission. Independent reproduction may add little to novelty while adding greatly to confidence. A measurement system may not win a leaderboard yet still create substantial option value for a research program. Keeping these values separate is consistent with the Leiden Manifesto's broader principle that quantitative indicators should support, not replace, qualitative judgment.
Constraint first → Vector second → Horizon always visible
The decision rule is gate, then deliberate. Discard options that violate a non-compensable constraint. Among the options that pass, state the trade-off explicitly rather than hiding it inside a total score, and preserve unresolved disagreement rather than averaging it away. Non-scalar does not mean non-decisive; it means the reason for the decision remains inspectable.
This Framework Can Be Gamed Too
A constitution can constrain motivated reasoning; it cannot certify virtue.
Replacing acceptance with “depth,” “durability,” or “understanding” does not escape Goodhart's law. A researcher can perform depth through complexity, call an untested transfer claim durable, choose a falsification pass unlikely to falsify, optimize downloads instead of truth, or turn a completed constitution into compliance theater. Even the language of long-term impact can become a prestige aesthetic that excuses weak present evidence.
The most dangerous loophole is outcome arbitrage. Accepted work is celebrated for venue success; rejected work is relabeled scientifically brave; unused work is justified as personal growth; and a failed project becomes “exploration.” All four interpretations may sometimes be true, but a framework that can rationalize every outcome cannot discipline any decision. The remedy is not a more elaborate total score. It is an evidence dossier that preserves what was intended before the outcome was known.
- Name the primary and secondary goals, claim scope, and intended contribution type in advance.
- Record kill, narrow, pause, and continuation criteria—including projects that never become papers.
- Timestamp changes and explain whether they follow new evidence or merely rescue a desired story.
- Keep disconfirming results, abandoned paths, and unresolved disagreements in the project ledger.
- Invite an outside challenge from someone rewarded for finding the strongest alternative explanation.
- Review and retire rituals that no longer change decisions, expose disagreements, or reduce verification cost.
The framework should remain contestable. It should reveal judgment rather than launder it, and it should be revised when its own side effects appear. The goal is not to build a bureaucracy around integrity. It is to leave enough temporal evidence that I cannot effortlessly rewrite the meaning of a project after I know how reviewers, metrics, or the community responded.
A More Granular Meaning of Success
“Did it get in?” is too coarse for understanding what a project achieved. I find it more useful to separate venue milestones from scientific and endurance milestones. These are not one prestige ladder. A project can advance along one dimension without advancing along the other, which is precisely why a venue decision should not monopolize the evaluation.
A conference decision is also an allocation decision: what belongs in a finite program under constraints. It is not designed to be a general certificate of truth or durability. The mistake is not taking the decision seriously; it is asking the institution to certify more than its process can establish.
Venue milestones
- Review-ready
- The question, claim, evidence, limitations, and artifact form a complete contribution that another researcher can evaluate seriously.
- Engaged
- The work reached a submission, workshop, presentation, or serious exchange and began receiving external scrutiny.
- Accepted
- The paper passed a venue's threshold under a particular scope, reviewer pool, and moment in the field.
- Recognized
- The community gave an additional visibility or esteem signal—discussion, spotlight, oral, award, or sustained attention—without this alone establishing scientific value or durability.
Scientific & endurance milestones
- Stress-tested
- The central result survives reruns, stronger alternatives, stress tests, and ideally independent reproduction.
- Reused
- Other people build on the code, data, method, measurement, or conceptual framing rather than only cite the paper.
- Enduring
- The work remains useful after the benchmark, model generation, or trend changes, and improves how the field thinks or builds.
These milestones are causally entangled. Topic popularity, institutional reputation, network position, compute, communication, licensing, and community size can affect acceptance and visibility. Acceptance can then increase visibility, which can increase citations and reuse. Endurance is right-censored: new work cannot yet have demonstrated it. Researcher growth is partly self-report. None of this makes later outcomes meaningless; it means they are evidence with confounders, not moral grades.
- Process quality
- Were decisions prospective, traceable, appropriately skeptical, and revised when evidence changed?
- Claim–evidence quality
- Do the methods and observations warrant the stated inference under explicit assumptions and limits?
- Venue reception
- Did a particular reviewer pool find the contribution clear, significant, credible, and suitable for that venue at that time?
- Downstream uptake
- Was the work later reproduced, reused, corrected, extended, or incorporated into how others reason and build?
The first two layers are most directly influenceable during a project. The latter two arrive later and mix intrinsic qualities with exposure and circumstance. A review decision should therefore be treated as noisy Bayesian evidence—not by assigning theatrical probabilities, but by updating different beliefs by different amounts. In the official NeurIPS 2021 consistency experiment, 882 submissions received independent committee reviews and 23% received inconsistent accept–reject recommendations; the organizers estimated that rerunning the process would change roughly half of the accepted set. This is one venue experiment, not proof that review is random. It is evidence that a decision contains threshold and reviewer noise alongside signal.
Prior beliefs → Specific review evidence + fit + noise → Separate belief updates
A reproducible leakage path should sharply decrease confidence in validity. Repeated independent confusion should sharply decrease confidence in clarity. Novelty disagreement may update belief about venue fit more than belief about truth. A score without a specific reason deserves little update, and correlated reviewers are not independent samples. I should never label unfavorable feedback “review noise” before asking what observation would prove that diagnosis wrong. Acceptance should not suspend falsification; rejection should not erase strong experimental evidence; and independent reproduction should update belief in a claim more than prestige does.
Design the Portfolio, Not Only the Paper
A paper is a time-bounded slice of a research program, not necessarily the unit of intellectual ambition.
Requiring every project to deliver short-term acceptance, high-risk novelty, mechanism-level depth, exhaustive validation, reusable infrastructure, and career security is not rigor. It is an impossible scalar objective that often produces cosmetic depth. A healthier unit of planning is the portfolio: a set of projects with different roles, horizons, and failure modes that jointly protect both survival and sustained inquiry.
Here, depth means reducing a consequential uncertainty through mechanism, discriminating evidence, reliable measurement, or boundary mapping—not displaying complexity. Portfolio design protects the time and diversity of work required for that depth to emerge.
- Anchor questions
- Consequential uncertainties that can organize several conference cycles and accumulate conceptual identity rather than reset with every trend.
- Probes
- Cheap, time-bounded explorations allowed to fail. Their product is information about whether a direction deserves deeper investment.
- Bridge papers
- Claims with honest boundaries that connect a long-horizon question to a contribution that can be completed and scrutinized within a submission cycle.
- Infrastructure & verification
- Datasets, measurements, baselines, replications, negative results, and tools that become shared research capital for later projects.
- Stewardship
- Corrections, artifact support, benchmark maintenance, and follow-up tests that continue after publication or acceptance.
Portfolio design turns self-restraint from heroic willpower into resource allocation. If every active project must serve the next deadline, no personal constitution can keep long-horizon work from being repeatedly narrowed, overstated, or sliced into locally attractive units. Short-cycle work creates cadence; anchor work creates direction; infrastructure creates compounding capacity. I do not need a fixed percentage for each category, but I do need to notice when one category has consumed all intellectual slack.
This also prevents “durable” from becoming a perfectionist demand. A narrow, timely, incremental, benchmark-specific, replication, dataset, audit, or engineering paper can be excellent. Different contribution types require different evidence. Theory needs explicit, defensible assumptions and a correct proof; systems work needs realistic workloads and resource accounting; datasets need provenance, coverage, and intended-use limits. Benchmarks need construct validity and lifecycle governance; empirical claims need robust comparisons and boundary tests; replications and audits need fidelity, independence, and diagnostic value. Not every paper must be timeless. The portfolio should contain credible paths by which some work can mature, compound, and remain useful.
Proportional rigor: honesty and claim–evidence proportionality are non-negotiable. Assurance should scale with claim breadth, stakes, and irreversibility. When resources cannot support that assurance, narrow the claim rather than lower the evidentiary standard. Stop when another test's expected information gain no longer justifies its cost; if unresolved uncertainty still bears on the inference, document it or narrow the claim.
Depth is not research-as-suffering. Exhaustion, resource maximalism, and endless ablations are not evidence of truth. The point is to spend scarce verification effort where an error would most change the inference, and to keep failed or killed projects visible enough that the portfolio is not judged only by its survivors.
A Research Constitution for Each Project
Self-constraint is not a moral pose. It is research governance designed before the result, the deadline, and the temptation to rationalize.
The moment of maximum temptation is a poor time to invent standards. At the beginning of a substantial project, I want a one-page research constitution: a small set of commitments that remains visible while methods, results, and the paper evolve. It should be lightweight enough to use and concrete enough to make deviations obvious.
- What important uncertainty are we trying to reduce?
- What is the smallest falsifiable version of the central claim?
- What evidence would make us abandon, narrow, or revise it?
- Which primary metrics, baselines, and evaluation choices should be fixed before final results?
- What are we unwilling to do even if it improves acceptance probability?
- What would remain valuable if the expected positive result disappears?
- If the current benchmark vanished next year, what insight would survive?
- Which claims and limits must every author understand, and who owns each critical evidence path well enough to rebuild it?
Minimum viable protocol before the confirmatory pass: freeze the primary claim, metric, decision rule, and a defensible baseline set; preserve an experiment ledger; schedule an explicit falsification pass; and reserve a clean setting for the central evidence. Stronger claims and higher-risk settings should trigger stronger controls, including independent reproduction and broader boundary mapping.
This is related to the logic of preregistration: decisions and predictions stated before outcomes are known make later deviations and selective analyses easier to detect. Not every AI project supports formal preregistration, and premature commitment can freeze a poor metric or suppress useful surprise. The practical solution is a phase boundary rather than a ban on exploration.
Exploratory sandbox
- Search broadly, debug constructs, and revise hypotheses
- Label results as discovery evidence, not final confirmation
- Preserve alternatives, failures, and the path that shaped the claim
Confirmatory pass
- Freeze the central claim, primary metric, baseline, and decision rule
- Use a fresh holdout, untouched setting, or clean independent rerun
- Timestamp amendments and separate them from the original test
Epistemic flexibility is not itself hacking. Transparent updating is science. The key is to distinguish learning-driven amendment from outcome-driven relabeling, and to avoid using the same evidence both to invent a hypothesis and to present it as though it had been predicted.
Evidence audit
- Every core claim maps to specific supporting and disconfirming tests
- The strongest baseline receives comparable care and resources
- Failed runs, changed metrics, and killed hypotheses remain traceable
- A clean final setting is preserved from repeated optimization
Understanding audit
- Explain the mechanism without the abstract, slides, or agent
- Predict the direction of decisive ablations before running them
- Name realistic failure conditions and simpler explanations
- Rebuild the central evidence path from a clean environment
The situation I most want this constitution to handle is the familiar week before a deadline: the headline metric has improved, the mechanism is still unclear, a strong baseline has not received equal care, and the attractive title implies more than the clean evidence shows. A scalar reward says to preserve the story and squeeze in supporting runs. This protocol says to restore baseline parity, run the discriminating check, and either submit the narrower claim or keep the research clock open. The framework earns its place only if it can change that decision.
A useful paper should leave two artifacts: public knowledge outside the researcher and durable capability inside the researcher. If the PDF improves while the authors lose ownership of the mechanism, implementation, or evidence chain, the project has externalized output while failing to internalize expertise.
A Personal Research Operating System
A principle that lives only in an essay is weakest exactly when a deadline is strongest. I need a small set of persistent artifacts, role assignments, and review cadences that preserve memory when incentives intensify and judgment becomes less reliable. The operating system should make good science cheaper, not merely add moral friction.
Agenda → Portfolio → Project → Experiment → Submission → Feedback → Stewardship
Anchor the Question
Optimize for Reduced Uncertainty
A research contribution should make an important part of the world less confusing, not merely make a table look better.
I want to state the uncertainty before naming the method. This makes it easier to see whether the method answers the question or merely supplies a publishable object. The project should retain value under a negative result: a clarified boundary, a failed assumption, a better measurement, or a reusable experimental system can all reduce uncertainty honestly.
Apply the Durability Test Early
If the benchmark, model family, and trend disappeared next year, what part of the contribution would still matter?
This question does not disqualify timely work. It asks whether the paper contains a mechanism, measurement, abstraction, dataset, or engineering insight that travels beyond one leaderboard moment. The earlier this is asked, the more it can shape the contribution rather than decorate the conclusion.
Make Evidence Costly to Fool
Schedule a Falsification Pass
A claim earns more confidence when it survives a serious attempt to disprove it—especially when the test is designed, or the result independently reproduced, by someone not invested in the preferred result.
Before polishing the story, I want a review devoted to alternative explanations, leakage, unfair baselines, seed sensitivity, hidden data choices, and the simplest mechanism that could reproduce the result. An independent collaborator or reviewer—possibly supported by a separately specified agent—should be rewarded for finding disconfirming evidence rather than helping the narrative converge.
Maintain a Claim–Evidence Map
Every major claim should have a visible evidence path and a visible condition under which it would weaken.
A claim–evidence map prevents one attractive result from silently supporting several stronger conclusions. It also makes rebuttal and revision more honest: a failed experiment can narrow the relevant claim without forcing the entire paper into defensive storytelling. The reproducibility practices reported by the NeurIPS reproducibility program—including its checklist, code policy, and reproducibility challenge—strengthen the case for treating transparency, code, and verifiable procedures as part of the contribution rather than administrative material added after it.
Preserve Epistemic Ownership
Let AI Increase Leverage, Not Distance
Agents may accelerate execution, but they cannot inherit scientific responsibility or make understanding optional.
Core citations should be checked independently. Central experiments should be reproducible without hidden conversational state. At least one author should own each critical abstraction, implementation path, and inference from evidence to claim well enough to rebuild it independently. Tool separation alone is not epistemic independence when the tools inherit the same specification, data, and assumptions. Agent-generated complexity should pass a necessity review: if a simpler mechanism explains the result, the paper should prefer understanding over ornamental sophistication.
Distribute Work Without Diffusing Responsibility
No person needs to master an entire complex system, but every critical path needs an accountable owner and a documented interface.
The team should name claim, evaluation, data, implementation, and artifact owners; designate a skeptic with permission to challenge the preferred narrative; define an escalation or veto route for integrity concerns; and collectively sign off on the final claim scope. Senior authors must make it safe for junior members to report negative evidence. Distributed epistemic ownership is legitimate; anonymous responsibility is not.
Perform a Story-Removal Review
Temporarily remove the abstract, novelty framing, and intended conclusion; then ask what the evidence alone permits us to say.
Good writing helps a real contribution travel. It becomes dangerous when narrative coherence is treated as evidence. A story-removal review exposes whether experiments support the central claim, only a narrower claim, or a different result that may be less exciting but more true.
Use the Clock, Keep the Compass
Separate the Conference Clock and Research Clock
A deadline can decide when to communicate the current evidence; it should not decide when the research stops being responsible.
The conference clock contains submission, rebuttal, camera-ready, and presentation. The research clock contains mechanism discovery, replication, boundary mapping, artifact maintenance, and follow-up questions. Acceptance does not end the second clock, and rejection should not automatically kill a valuable question. After the decision, I want to score the project again without looking at the review outcome first.
Review Scores Are Sensors, Not Sovereigns
Reviewer feedback is evidence about clarity, relevance, credibility, and fit—not a truth certificate and not a complete judgment of research value.
A high-value rejection may reveal a communication gap, missing evidence, poor fit, or review noise. A low-value acceptance should not be canonized by the badge; its claims and artifacts still need repair. The right response is neither to worship acceptance nor romanticize rejection. It is to diagnose which signal changed and continue making the work more true, useful, and legible.
After each decision, a short review ledger should record the prior view of validity, evidence, clarity, significance, and fit, followed by the specific new observations in the reviews and competing explanations for the outcome. It should then name the experiment or rewrite that could distinguish those explanations and the resulting action: continue, narrow, reframe, replicate, or stop. This prevents one aggregate score from overwriting several different beliefs—or from becoming a verdict on the researcher's identity.
Maintain, Learn & Reassess
Use a Lightweight Evidence Cadence
Governance works when it preserves memory at the moment a decision is made, not when documentation is reconstructed for the appendix.
Per experiment, record the prediction, configuration, lineage, deviation, result, interpretation, belief update, and agent or tool provenance. Weekly, ask what changed a belief and what became genuinely internalized. Monthly, run a red-team and kill-or-continue review. Quarterly, inspect portfolio balance and correlated risk. At submission, place a science gate before the packaging gate. The cadence should stay small enough to survive pressure.
Treat Publication as the Start of Stewardship
Durable work needs versioned artifacts, correction paths, and later checks that distinguish attention from actual use.
After publication, track corrections, replication reports, artifact issues, maintenance needs, and downstream use. Review the work at longer horizons— for example after six, eighteen, and thirty-six months—without pretending that citation count alone measures impact. Record which claims survived, which were narrowed, what others reused, what capability stayed with the authors, and whether the original agenda still deserves investment.
Research After Paper Generation Becomes Cheap
AI systems are beginning to connect hypothesis generation, literature search, implementation, experimentation, visualization, and manuscript production. Early systems such as Co-Scientist—which reports end-to-end validation of AI-generated hypotheses in three biomedical application areas—and AI Scientist-v2—which reports three autonomous workshop submissions, one accepted—do not establish reliable autonomous discovery in general. They do show that hypothesis generation, execution, and scientific packaging can now be integrated into agentic workflows. The right response is to redesign rigor before fluent scientific output becomes abundant.
Generation abundance → Attention, verification, provenance & judgment scarcity
Reinvest the Productivity Dividend in Rigor
Update What a Paper Can Signal
When polished prose, competent code, extensive ablations, and attractive figures become cheap, their presence carries less information about understanding.
The scarce complements become question selection, experimental judgment, calibrated uncertainty, trustworthy provenance, independent verification, willingness to report disconfirming evidence, and responsibility for the inference from evidence to claim. A paper will remain valuable, but visual and procedural completeness will be weaker proxies for epistemic ownership. In a large human study with more than one hundred NLP researchers, LLM-generated ideas were judged more novel but slightly less feasible, while self-evaluation and diversity remained limitations. Novelty generation and research judgment are related but different capabilities. Evaluation must increasingly ask not only “what was produced?” but “what chain of evidence makes this conclusion trustworthy?”
Spend Automation on Independent Checks
The first-order use of an agent is more output; the higher-value use may be making previously unaffordable verification routine.
Faster execution should buy stronger baseline tuning, clean-room reproductions, broader boundary tests, sensitivity analysis, better documentation, and systematic searches for simpler or disconfirming explanations. If every saved hour becomes another submission, the field gains throughput without gaining trust. As generation becomes abundant, verification becomes the bottleneck—and judgment becomes the researcher's defining contribution.
Build Evidence Passports and Epistemic Security
Make the Claim Chain Inspectable
A future research artifact should carry an evidence passport, not merely a PDF, repository link, and model card.
The passport could encode a machine-readable claim graph; the experiments that support or weaken each claim; data, code, model, environment, and tool versions; hypotheses considered and rejected; compute and resource records; human and agent contributions; known limitations; and independent checks. Datasheets for Datasets demonstrates how standardized documentation can expose an artifact's motivation, composition, process, uses, and limits. The W3C PROV ontology already offers a vocabulary for entities, activities, human agents, and software agents. The challenge is not inventing provenance from nothing, but making it usable enough to follow a scientific inference and compact enough to be reviewed.
An evidence passport is a verification interface, not a scorecard. Completeness does not establish correctness, and a larger log is not stronger evidence. If the passport does not reduce the cost of checking a claim, it has become paperwork.
Separate Roles, Context, and Authority
Multiple agents are not independent when they inherit the same prompt, codebase, data leakage, assumptions, and desired conclusion.
A safer workflow distinguishes explorer, builder, verifier, critic, archivist, and human claim owner. The verifier should begin from the claim and evidence contract rather than the builder's narrative, and should use a different context, implementation path, or data view where feasible. Final holdouts need access controls. Logs should support audit and rollback. Failed verification should trigger claim narrowing, reruns, or retraction— not silent regeneration. The human claim owner remains accountable for scope, exceptions, uncertainty, and the final inference.
Threat-Model the Research Pipeline
Agentic science needs epistemic security: protection against failures that preserve plausible output while corrupting the evidence path.
Relevant threats include benchmark contamination, fabricated citations, hidden prompt state, evaluator overfitting, correlated agent agreement, silent data mutation, specification gaming, and conclusions that no author can reconstruct. Controls should be proportional to risk and include separation of duties, append-only or tamper-evident experiment records, data lineage, least-privilege access to final evaluations, incident response, and a clear path to correct the public record. Reliability is not only a model property; it is an end-to-end system property.
Let Review and Credit Become Multidimensional
Delay the Collapse Into One Score
Review should preserve distinctions among validity, significance, clarity, artifact quality, uncertainty, and venue fit for as long as possible.
A single overall score encourages authors and reviewers to trade away one dimension invisibly for another. Future systems could first assess evidential support and whether the methods and artifacts permit reproduction, then separately assess interest, novelty, and fit. Staged review and Registered Reports offer one model: evaluate questions and methods before outcomes are known, while acknowledging that this format is not appropriate for every exploratory or systems contribution. Review outcomes should communicate uncertainty and disagreement, not only a threshold bit.
Reward the Work That Keeps Knowledge Reliable
Personal discipline cannot create venues for null results, pay artifact maintenance, or give career credit for correcting a popular claim.
Labs and conferences should visibly credit replication, negative findings, measurement, dataset stewardship, benchmark governance, software maintenance, mentoring, and successful correction. The Hong Kong Principles argue for assessing responsible practices across the research lifecycle and valuing replication, synthesis, open science, peer review, and mentoring. In AI, this also means benchmarks should be assessed across their lifecycle, with provenance, construct boundaries, contamination history, versions, and retirement criteria. A metric should have a lifecycle, not an assumption of permanent authority.
Make AI-Assisted Review Auditable
Experiment Without Quietly Replacing Judgment
AI may reduce reviewer burden or improve consistency, but those benefits must be measured, and assistance must not dissolve consent, confidentiality, or human accountability.
The official NeurIPS 2026 AI Reviewing Experiment uses voluntary randomized conditions—no LLM assistance, open-ended assistance, and structured assistance—while keeping area chairs blind to the condition and framing AI as augmentation rather than replacement. That is a healthier direction than silently introducing tools and assuming quality. Future experiments should measure not only speed or score agreement, but error discovery, calibration, evidence traceability, confidentiality failures, correlated blind spots, and the cost of auditing the resulting review.
Keep the Human Reviewer Legible
An assisted review still needs a person who can explain, defend, and revise each consequential judgment.
Review provenance should record where assistance was used, which claims or citations were independently checked, and where reviewer confidence is low. Authors need a correction channel for verifiable tool errors. Chairs need visibility into correlated recommendations. The aim is not a machine-made consensus, but a review process that uses automation to surface evidence while preserving plural judgment and responsibility.
What Depth Should Mean
“Hardcore” should not be an aesthetic of larger models, more equations, longer appendices, more benchmarks, greater system complexity, heroic overwork, or expensive compute. Those may sometimes be necessary, but none guarantees depth. The word is useful only if it names epistemic qualities: the work attacks an important uncertainty, compresses a confusing mechanism into transferable understanding, maps where an insight does and does not hold, and gives future researchers better questions, measurements, or tools.
The hardest paper is often not the one with the largest model. It is the one that turns a consequential but poorly understood phenomenon into knowledge that is precise, generative, and difficult to misunderstand. Depth means the result explains more than it obscures, the complexity is necessary, and the insight can guide work beyond the original setup. It also includes calibrated restraint: a narrow conclusion with strong evidence is deeper than a sweeping claim supported by a fragile proxy.
Long-term impact cannot be demanded or certified in advance. It depends partly on future problems, adoption, timing, and luck. The responsible ambition is to create option value: trustworthy evidence that can be reinterpreted, artifacts that can be reused, mechanisms that transfer, clearly documented failures, and questions that remain fertile. Endurance is something later communities discover, not a label an author awards at submission.
Ambition Without Capture
I still want top-conference papers. I want acceptance to be evidence that a deeper process produced something clear and credible—not a substitute for that process. I want work that can survive a different reviewer pool, the disappearance of a benchmark, and the end of a trend. The most durable reward is not a badge but a strengthened ability to ask, test, explain, and build.
I want a research system in which the conference helps work become legible, the lab gives difficult questions enough time to mature, the portfolio protects both survival and depth, agents are deliberately used to make verification more affordable, and review decisions update beliefs without becoming identity. Acceptance can remain a meaningful reward inside that system. It simply should not be allowed to define the system.
The ambition is not merely to publish papers that pass a threshold. It is to build an operating discipline, a portfolio, and eventually institutions capable of producing knowledge that remains trustworthy after the threshold has been forgotten. That is not less ambition. It is ambition anchored to a reward system that is harder to satisfy through appearance alone.
Sources That Informed These Notes
- David Manheim and Scott Garrabrant, “Categorizing Variants of Goodhart's Law” (2018).
- Paul E. Smaldino and Richard McElreath, “The Natural Selection of Bad Science” (2016).
- Mostafa Dehghani et al., “The Benchmark Lottery” (2021).
- Joelle Pineau et al., “Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program)” (2021).
- Brian A. Nosek et al., “The Preregistration Revolution” (2018).
- Diana Hicks et al., “Bibliometrics: The Leiden Manifesto for Research Metrics” (2015).
- David Moher et al., “The Hong Kong Principles for Assessing Researchers” (2020).
- Alina Beygelzimer et al., “The NeurIPS 2021 Consistency Experiment” (2021).
- Christopher D. Chambers and Loukia Tzavella, “The Past, Present and Future of Registered Reports” (2022).
- Anka Reuel et al., “BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices” (2024).
- Timnit Gebru et al., “Datasheets for Datasets” (2021).
- Timothy Lebo, Satya Sahoo, and Deborah McGuinness, “PROV-O: The PROV Ontology” (W3C Recommendation, 2013).
- Juraj Gottweis et al., “Accelerating Scientific Discovery with Co-Scientist” (2026).
- Yutaro Yamada et al., “The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search” (2025).
- Chenglei Si, Diyi Yang, and Tatsunori Hashimoto, “Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers” (2025).
- NeurIPS, “NeurIPS 2026 AI Reviewing Experiment” (2026).