As AI takes on more demanding intellectual tasks, universities face a practical challenge: how to use these systems while preserving the learning, accountability, and collaboration that make research meaningful. Sasha Rakhlin, director of the MIT Statistics and Data Science Center, argues that academic institutions should rethink both how they train researchers and what they recognize as valuable work.
Fast verification helps AI progress
Mathematics offers a striking view of how quickly AI capabilities can change. Last year, a model reached gold-medal level at the International Mathematical Olympiad. Only a year later, models were producing new research results, including a recently proposed solution to one of the Millennium Prize Problems.
Rakhlin points to verification as one reason progress can accelerate in some fields. Formalized proofs can be checked automatically, and programs can be run and tested. When results can be evaluated quickly and reliably, AI systems can generate possibilities, learn from the outcomes, and improve.
The same cycle can apply to AI research itself. Better models can help improve models, training procedures, and supporting tools, which may contribute to further advances. That feedback loop creates opportunities for experienced researchers to tackle questions that once seemed too technically demanding.
But producing a result and understanding it are different tasks. Researchers still need to ask why a solution works, whether it applies elsewhere, and what question should follow. Rakhlin cautions universities against assuming that judgment, abstraction, or problem formulation will remain outside AI's reach. Institutions should prepare for AI to surpass people in many aspects of intellectual work.
Academic credit needs a clearer basis
A polished paper may become a less reliable sign of an individual's expertise as AI takes on more parts of research. Departments therefore need to reconsider what they reward, rather than defining valuable work only by what AI cannot yet do.
Rakhlin suggests that asking good questions, replicating results, connecting ideas, reporting informative negative results, and sharing datasets may deserve greater recognition. Evaluation should make clear what a researcher contributed and takes responsibility for, including when AI performed substantial parts of the work.
Those expectations matter beyond publication. They can shape hiring, promotion, and funding, and should be made explicit to current and incoming PhD students. Clearer standards could help departments assess a researcher's role even when the final product does not show who did which parts of the work.
Graduate training must preserve learning by doing
AI can help students attempt more ambitious projects, but handing routine work to a system may also remove experiences that build expertise. Calculations, coding, unsuccessful approaches, and small discoveries can help students develop intuition and judgment.
The challenge is to distinguish wasted effort from productive practice. Students need opportunities to build the abilities that let them use AI critically: formulating problems, checking model outputs, reproducing results, and explaining their decisions. A strong grasp of fundamentals may become more valuable as students work with these tools.
This changes the question from whether students should use AI to how their training can combine its capabilities with the intellectual work through which researchers learn. Universities will need to help students take advantage of AI without treating every difficult or repetitive step as disposable.
Shared infrastructure could connect laboratories
Rakhlin sees a role for universities in pursuing questions over long periods, sharing results openly, and evaluating claims independently. Industry partnerships will remain essential, but commercial priorities may not cover the full breadth of science or stay aligned with it. Some technological independence will therefore matter.
Universities also hold knowledge that is often missing from published papers. Failed experiments, abandoned directions, and the reasons an approach did not work may never appear in the formal record. Scientists and engineers gain practical knowledge about what is likely to fail and why. Rakhlin suggests that this missing context may help explain why models can struggle, especially in empirical sciences, to anticipate consequences that experts recognize. Recording negative results and researchers' interpretations could make AI more useful for scientific exploration.
He imagines laboratories linked by shared AI research infrastructure. For example, if a neuroscience laboratory needed better methods for segmenting neurons in microscopy images, an AI agent might identify relevant work from a computer vision group, connect the teams, suggest benchmarks, and support further iteration.
Building toward that kind of collaboration would require workflows that capture hypotheses, interventions, outcomes, failures, and interpretations. Shared systems could then connect those records and let agents use information and tools across laboratories with appropriate permissions. Rakhlin says making this possible will require substantial public and institutional investment in compute, secure data systems, and expertise in adapting and post-training AI models.
Keeping a record of how research develops could also make the lineage of ideas and contributions, including graduate students' work, easier to recognize. With agreed rules for consent and credit, researchers may be more willing to collaborate openly. For Rakhlin, the opportunity is to use AI's ability to connect information and ideas to help researchers work together on difficult scientific and engineering problems—and for universities to invest now in the systems that make that possible.