Paper deep dive
Personalized Auto-Research: Towards a True AI Co-Scientist
Bo Ni, Franck Dernoncourt, Hongjie Chen, Yu Wang, Nesreen K. Ahmed, Zhengzhong Tu, Tyler Derr, Ryan A. Rossi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/18/2026, 4:50:29 AM
Summary
This paper introduces 'Personalized Auto-Research,' a framework for AI co-scientists that conditions every stage of the research process (retrieval, hypothesis generation, experimentation, writing, review) on a graph-grounded representation of the individual researcher. It argues that current systems are researcher-agnostic, leading to a 'one-size-fits-all' failure mode that erases tacit knowledge and scientific heterogeneity. The proposed framework uses a heterogeneous research graph to derive researcher contexts, enabling personalized evidence retrieval, hypothesis scoring, and output synthesis to ensure feasibility, alignment, and novelty relative to the specific scientist.
Entities (12)
Relation Signals (7)
Bo Ni → affiliatedwith → Vanderbilt University
confidence 99% · Bo Ni email: bo.ni@vanderbilt.edu Affiliation: Vanderbilt University
Ryan A. Rossi → affiliatedwith → Adobe Research
confidence 99% · Ryan A. Rossi email: ryarossi@gmail.com Affiliation: Adobe Research
Personalized Auto-Research → proposessolutionfor → Researcher-Agnostic AI Systems
confidence 95% · In this work, we introduce the problem of personalized auto-research... state-of-the-art systems remain researcher-agnostic
Personalized Auto-Research → usesmethod → Graph-Grounded Researcher Representation
confidence 92% · The framework consists of three fundamental components: (i) graph-grounded researcher representations
Researcher-Agnostic Systems → causes → One-Size-Fits-All Failure Mode
confidence 90% · Notably, we highlight a one-size-fits-all failure mode where distinct researchers issuing the same goal receive essentially the same research
Personalized Auto-Research → implementsas → Algorithm 1
confidence 90% · Algorithm 1 gives the proposed personalized auto-research procedure.
AI Scientist → precedes → Personalized Auto-Research
confidence 85% · Since the 2024 AI Scientist prototype, the field has moved rapidly... In this work, we introduce the problem of personalized auto-research
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to change how research is carried out. Despite this rapid progress, state-of-the-art systems remain researcher-agnostic: given a research goal, they optimize novelty, validity, or reviewer score while ignoring the individual scientist who will use the output. This overlooks a fundamental fact about research, namely, that what counts as novel, valuable, or feasible depends on the researcher, including their prior work, methodological repertoire, and the collaborators and communities in which they are embedded. In this work, we introduce the problem of personalized auto-research, which conditions every stage of the research process on a representation of the individual researcher. We argue that personalization is not a convenience layer, but rather the fundamental property that allows an AI system to serve as a genuine co-scientist rather than a generic instrument. To address this problem, we propose a general and flexible framework that threads a graph-grounded researcher context through retrieval, hypothesis search, experimentation, writing, and review. The framework consists of three fundamental components: (i) graph-grounded researcher representations, (ii) personalization across the full research pipeline, and (iii) evaluation grounded in the individual. Notably, we highlight a one-size-fits-all failure mode where distinct researchers issuing the same goal receive essentially the same research, erasing the tacit knowledge through which novel ideas arise. Finally, we discuss fundamental open problems and challenges.
Tags
Links
- Source: https://arxiv.org/abs/2608.14881v1
- Canonical: https://arxiv.org/abs/2608.14881v1
Trouble viewing inline? Open PDF directly →
Full Text
41,459 characters extracted from source content.
Expand or collapse full text
Personalized Auto-Research: Towards a True AI Co-ScientistConference: ACM AI Leadership Summit; 2026; ACM AI Leadership Summit 2026CCS: Computing methodologies Artificial intelligenceCCS: Computing methodologies Machine learningCCS: Information systems Data mining Bo Ni email: bo.ni@vanderbilt.edu Affiliation: Vanderbilt University , Nashville , Tennessee , USA , Franck Dernoncourt email: dernonco@adobe.com Affiliation: Adobe Research , San Jose , California , USA , Hongjie Chen email: hojiechen@gmail.com Affiliation: Dolby Laboratories , San Francisco , California , USA , Yu Wang email: yu.wang6@uga.edu Affiliation: University of Georgia , Athens , Georgia , USA , Nesreen K. Ahmed email: n.kamel@gmail.com Affiliation: Cisco AI Research , San Jose , California , USA , Zhengzhong Tu email: tzz@tamu.edu Affiliation: Texas A&M University , College Station , Texas , USA , Tyler Derr email: tyler.derr@vanderbilt.edu Affiliation: Vanderbilt University , Nashville , Tennessee , USA and Ryan A. Rossi email: ryarossi@gmail.com Affiliation: Adobe Research , San Jose , California , USA 2026© , 2026; Abstract. AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to change how research is carried out. Despite this rapid progress, state-of-the-art systems remain researcher-agnostic: given a research goal, they optimize novelty, validity, or reviewer score while ignoring the individual scientist who will use the output. This overlooks a fundamental fact about research, namely, that what counts as novel, valuable, or feasible depends on the researcher, including their prior work, methodological repertoire, and the collaborators and communities in which they are embedded. In this work, we introduce the problem of personalized auto-research, which conditions every stage of the research process on a representation of the individual researcher. We argue that personalization is not a convenience layer, but rather the fundamental property that allows an AI system to serve as a genuine co-scientist rather than a generic instrument. To address this problem, we propose a general and flexible framework that threads a graph-grounded researcher context through retrieval, hypothesis search, experimentation, writing, and review. The framework consists of three fundamental components: (i) graph-grounded researcher representations, (i) personalization across the full research pipeline, and (i) evaluation grounded in the individual. Notably, we highlight a one-size-fits-all failure mode where distinct researchers issuing the same goal receive essentially the same research, erasing the tacit knowledge through which novel ideas arise. Finally, we discuss fundamental open problems and challenges. Keywords: Auto-research, AI co-scientists, personalization goal ggresearcher-agnosticco-scientist fi(g,o<i)f_i(g,o_<i)QQQsame forevery user u(a) Researcher-agnosticu1u_1fi(g,o<i∣cu1)f_i(g,o_<i c_u_1)u1Q_u_1u1→cu1z_u_1\!→\!c_u_1(b) Personalizedu2u_2fi(g,o<i∣cu2)f_i(g,o_<i c_u_2)u2Q_u_2u2→cu2z_u_2\!→\!c_u_2same goal g≠ foreach user u Figure 1. (a) A researcher-agnostic co-scientist maps the same abstract goal g to an identical research package Q for every researcher. (b) Personalized auto-research conditions every stage on a graph-derived context cuc_u. For example, researcher u1u_1 lies in a dense network cluster, while u2u_2 bridges a structural hole. Consequently, the same goal naturally yields distinct evidence, search trajectories, and output packages (u1≠u2Q_u_1 _u_2). Each package is thereby optimized for feasibility, alignment, and novelty relative to the individual researcher’s domain. 1. Introduction Scientific publication has grown exponentially for decades (27), and no individual researcher can fully navigate the literature of their own field, let alone integrate insights from adjacent disciplines. This fundamental problem motivates the recent line of work on AI co-scientists and auto-research systems, that is, language-model agents that generate hypotheses, retrieve and synthesize related work, design and execute experiments, and draft manuscripts (12; 4; 33; 23; 25). Since the 2024 AI Scientist prototype, the field has moved rapidly to progressive agentic tree search, experiment-manager control, shared preprint-style archives, cross-run memory, human-in-the-loop co-research, and explicit provenance and safety mechanisms (22; 36; 26; 11; 28; 6). Despite their importance and rapid progress, these systems share a fundamental structural limitation: they are researcher-agnostic. Given the same goal, they produce the same distribution of outputs whether the requester is a first-year doctoral student or a senior professor, a graph-mining researcher or a computational biologist (Figure 1). The output may be competent, but it is also interchangeable, and this interchangeability contradicts how scientific ideas are often discovered. Intuitively, novel directions frequently arise from a scientist’s idiosyncratic combination of prior failures, methodological taste, and hard-won experimental intuition, and a homogeneous system erases precisely the heterogeneity that makes research creative. The stakes span every level: (i) epistemic, since most scientific capability is tacit and absent from the literature, and is thus unreachable by any literature-conditioned system; (i) systemic, since many researchers querying one generic engine drives the field toward a scientific monoculture, racing the same ideas while counterfactually valuable directions go unexplored; and (i) practical, since researchers can only verify, and will only adopt, directions matched to their expertise and resources. In this work, we introduce personalized auto-research, the problem of conditioning every stage of the research process on a representation of the individual researcher. Notably, this problem is fundamentally different from both existing AI co-scientist systems and prior work on personalized language models (21; 18): existing co-scientists automate research but largely ignore the researcher, whereas existing personalization work models user-specific outputs but not the sequence of scientific decisions that constitute research. In personalized auto-research, the object being personalized is an end-to-end research trajectory, not merely a response. Summary of contributions. The key contributions of this work are as follows: (1) We formalize the problem of personalized auto-research (§3). (2) We propose a general and flexible framework (Algorithm 1) for this problem that personalizes every step such as retrieval, hypothesis search, experimentation, writing, citation, and review, etc (§4). (3) We present a vision for personalized auto-research organized around three fundamental components: researcher representation, personalization across the full research pipeline, and evaluation grounded in the individual (§5). (4) Finally, we discuss open problems and challenges. (§6). 2. Background AI Co-Scientists and Auto-Research. Long-horizon reasoning and tool use (29; 35) have enabled agents that carry out research end-to-end. The AI Scientist established the autonomous loop of proposing, implementing, writing up, and reviewing experiments (12; 13), and AI Scientist-v2 replaced its template-dependent loop with progressive agentic tree search, an experiment manager, parallel execution, and vision-language figure feedback (33). A growing set of systems extends this across multi-agent hypothesis generation, staged human-feedback workflows, automated data science, and long-horizon discovery (4; 5; 23; 25; 34; 31; 17; 19). A parallel line builds the surrounding infrastructure, namely shared archives, cross-run memory, and population-level search over code states (22; 36; 28; 11; 6; 26), and a third studies evaluation and risk through discovery benchmarks, manuscript verification, safety, and critiques of implementation and evaluation bottlenecks (2; 14; 24; 39; 40; 15); recent surveys organize the space by research stage and autonomy level (16; 3; 37; 20; 30; 38). Auto-research is thus no longer speculative, spanning agentic search, automated implementation, full-paper generation, and safety controls, with domain systems such as AlphaFold showing how AI complements expert judgment in high-stakes science (7). Yet the dominant objective is task-, benchmark-, or field-conditioned: quality is optimized with respect to the goal, literature, evidence, or reviewer model, never with respect to the individual researcher who will adopt the output. Even scientist-in-the-loop systems (4; 23) condition only on manual, session-level guidance, so interaction is not personalization, which requires a learned, persistent representation of what the scientist cannot articulate in a prompt. Auto-research and co-scientist are therefore not synonyms: the former names a capability (the automation of research) while the latter names a relationship with a specific researcher. Table 1 makes this explicit: personalization is orthogonal to the autonomy axis along which the field has advanced, and the bottom-right quadrant, where interaction updates the personalization itself, is the true AI co-scientist we pose and the gap this work addresses. Table 1. Personalization is orthogonal to autonomy. Existing systems advance along the autonomy axis while remaining researcher-agnostic, whereas our vision of Personalized Auto-Research (Alg. 1) and our True Personalized AI Co-Scientist (Alg. 1 with human-in-the-loop). Researcher-agnostic Personalized Fully autonomous AI Scientist (12; 33), DeepScientist (31), … Personalized Auto-Research (Alg. 1) Human-in-the-loop Co-Scientist (4; 5), AutoResearchClaw (11), … True AI Co-Scientist (Alg. 1 w/ human-in-the-loop) Personalization. Recommender systems learn latent user representations from interactions (8), while preference alignment (18), retrieval-augmented generation (10), and the LongLaMP benchmark (9) establish that user-conditioned language modeling is tractable and beneficial. However, personalizing a recommendation or a single output is far narrower than personalizing science. In personalized auto-research, the object is a sequence of decisions: what to retrieve, what to hypothesize, which experiments to run, how to frame the write-up, how to revise the resulting artifact, etc. 3. Problem Formulation More formally, let U denote the population of researchers and let =(,ℰ,τV,τE)G=(V,E_G, _V, _E) denote a heterogeneous graph over the research landscape, where V contains researcher, paper, venue, institution, method, dataset, and topic nodes; ℰE_G contains co-authorship, citation, publication, affiliation, usage, and topic-assignment edges; and τV,τE _V, _E assign node and edge types (ℰE_G is distinct from the literature index ℐI used below). For a researcher u∈u , let uS_u denote observed signals (papers, citations, code, review history, venue preferences, explicit constraints), let u∈ℝdz_u ^d be a graph-derived representation of u, and let cuc_u be the operational context given to the auto-research system: (1) u=Enc(u,),cu=Φ(u,u).z_u=Enc_G(u;G), c_u= (S_u,z_u). Definition 3.1 (Personalized Auto-Research). Let g be a research goal and =(p1,…,pK)P=(p_1,…,p_K) the stages of the research process (literature retrieval, hypothesis generation, experiment design, code execution, writing, citation, refinement, review). Whereas a researcher-agnostic co-scientist parameterizes each stage by the goal alone, oi=fi(g,o<i)o_i=f_i(g,o_<i), personalized auto-research learns stage models conditioned on the researcher, (2) oi=fi(g,o<i∣cu),o_i=f_i(g,o_<i c_u), where o<io_<i denotes outputs of earlier stages. The desiderata are threefold: the output should be (i) feasible, respecting the researcher’s capabilities, resources, and constraints; (i) aligned, compatible with the researcher’s scientific identity, community, and style; and (i) novel, new to the field and distinct from the researcher’s prior work. Feasibility and alignment are personalized properties, whereas novelty is partly field-level and partly user-relative. This distinction is fundamentally important, since without it, personalization collapses into a recommender that predicts more of the same. 4. Personalized Auto-Research Algorithm 1 gives the proposed personalized auto-research procedure. The key idea is to formulate personalization as a modification of the SOTA agentic auto-research loop, that is, agentic tree search over partial research states with an experiment manager, code execution, figure refinement, manuscript writing, automated review, and provenance logging (33; 11; 28; 6), rather than the linear 2024 template loop. Personalization adds a user u, a context cuc_u, a personalized evidence set ℛu(g)R_u(g), and a user-conditioned utility, where retrieval scores documents d in corpus D jointly by goal and researcher and Topk _k returns the k highest-scoring items: (3) su(d∣g) s_u(d g) =⟨η(d),ρ(g,cu)⟩,ℛu(g)=Topk(;su(⋅∣g)), = η(d),\,ρ(g,c_u) , _u(g)= _k\! (D;s_u(· g) ), (4) U(h∣g,u) U(h g,u) =αNov(h,ℐ,cu)+βRel(h,g,cu)+γFeas(h,0,cu). =α (h,I,c_u)+β (h,g,c_u)+γ (h,W_0,c_u). Notably, Eq. (4) is one key distinction: hypotheses are ranked by personalized novelty, relevance, and feasibility, rather than filtered by a binary, global novelty test. Intuitively, the same idea can be infeasible for one researcher, obvious to another, and transformative for a third whose graph position makes a new bridge credible. Algorithm 1 Personalized Auto-Research 1: language-model agents π, experiment manager μ, vision-language reviewer ω, research goal g, initial workspace 0W_0, seed archive J, literature index ℐI over corpus D, research graph G, user u with signals uS_u, context encoder Φ , tree-search budget N, branch factor b, package count m, exp. budget B 2: personalized reprod. research packages uA_u 3: // Personalized Context and Literature Grounding 4: u←Enc(u,)z_u _G(u;G) 5: cu←Φ(u,u)c_u← (S_u,z_u) 6: ℛu(g)←Topk(;su(⋅∣g))R_u(g)← _k(D;s_u(· g)) 7: ℬu←(0,∅,∅,∅,0)B_u←\(W_0, , , ,0)\ 8: // Personalized Hypothesis Search and Experimentation 9: for r=1r=1 to N do 10: x←μ(select,ℬu,g,ℛu(g),cu)x←μ( select,B_u,g,R_u(g),c_u) 11: ←π(⋅∣expand,x,g,,ℛu(g),cu,b)X←π(· expand,x,g,J,R_u(g),c_u,b) 12: for each child state x′∈x do 13: h∼π(⋅∣hypothesize,x′,g,,ℛu(g),cu)h π(· hypothesize,x ,g,J,R_u(g),c_u) 14: U(h∣g,u)←αNov(h,ℐ,cu)+βRel(h,g,cu)+γFeas(h,0,cu)U(h g,u)←α (h,I,c_u)+β (h,g,c_u)+γ (h,W_0,c_u) 15: ∼π(⋅∣implement,x′,h,0,cu)C π(· implement,x ,h,W_0,c_u) 16: (,ℱ,Λu)←Execute(,h,B,cu)(O,F, _u)← (C,h,B,c_u) 17: ℱ∼ω(⋅∣figure-review,ℱ,,g,cu)F ω(· figure-review,F,O,g,c_u) 18: qu←μ(score,U(h∣g,u),h,,,ℱ,g,cu)q_u←μ( score,U(h g,u),h,C,O,F,g,c_u) 19: ℬu←ℬu∪(h,,,ℱ,qu,Λu)B_u _u∪\(h,C,O,F,q_u, _u)\ 20: end for 21: optionally: er←u(feedback,Top1(ℬu;qu))e_r← u ( feedback, _1(B_u;q_u) ), u←u∪erS_u _u∪\e_r\, cu←Φ(u,u)c_u← (S_u,z_u) 22: end for 23: // Personalized Research Package Synthesis 24: u←∅A_u← 25: for each state (h,,,ℱ,qu,Λu)∈Topm(ℬu;qu)(h,C,O,F,q_u, _u)∈ _m(B_u;q_u) do 26: y∼π(⋅∣write,g,h,,,ℱ,ℛu(g),cu)y π(· write,g,h,C,O,F,R_u(g),c_u) 27: y∼π(⋅∣cite-refine,y,ℛu(g),,ℱ,cu)y π(· cite-refine,y,R_u(g),O,F,c_u) 28: v∼π(⋅∣review,y,g,,ℱ,cu)v π(· review,y,g,O,F,c_u) 29: u←(h,U(h∣g,u),,,ℱ,y,v,Λu,cu,ℛu(g))Q_u←(h,U(h g,u),C,O,F,y,v, _u,c_u,R_u(g)) 30: u←u∪uA_u _u∪\Q_u\ 31: end for 32: return uA_u Personalized context and evidence Lines 4–6 construct the user-specific state: the graph encoder produces uz_u, the context encoder produces cuc_u, and retrieval produces ℛu(g)R_u(g). Notably, this evidence set is not simply the topically closest literature to g. It is the one that is relevant to the goal and useful for this researcher, given their prior work, collaborators, methods, resources, and position in G. Personalized hypothesis search Lines 9–19 replace a universal agentic tree with a user-conditioned tree ℬuB_u: the manager selects partial states, the language model expands them, and every hypothesis is scored by U(h∣g,u)U(h g,u), all under cuc_u. Two researchers with the same goal therefore induce different search trees, since their contexts change the selection policy, the expansion distribution, and the utility. Furthermore, candidates are implemented and executed under the user’s constraints (Lines 15–16), and even vision-language figure feedback may depend on cuc_u (Line 17): a graph-mining researcher and a biomedical collaborator need different visual encodings and terminology for the same quantitative result. The researcher also remains in the loop: at any iteration, u may optionally critique the current best state (Line 21). Notably, this feedback does more than steer the session, as in existing human-in-the-loop systems (4; 23; 11), where guidance evaporates when the session ends. Here it is folded into uS_u and re-encoded into cuc_u, and thus the representation of the relationship itself is updated, both within a run and across runs. Personalized package synthesis Lines 25–32 write, cite, refine, and review the top-m states under the same context. Each terminal artifact is a reproducible research package u=(h,U(h∣g,u),,,ℱ,y,v,Λu,cu,ℛu(g))Q_u=(h,U(h g,u),C,O,F,y,v, _u,c_u,R_u(g)) consisting of the hypothesis, its personalized score, code, outputs, figures, paper, automated review, provenance log, and the context and evidence that conditioned the run, that is, the metadata needed to audit why this package was produced for this researcher. 5. Personalized Auto-Research Vision The distinction between an instrument and a collaborator is not one of capability but of relationship: a capable instrument returns high-quality outputs, while a collaborator returns outputs tailored to the person it works with. Personalization is therefore not a marginal improvement to AI co-scientists, but rather what makes the metaphor accurate. We organize the resulting research agenda around three fundamental components. Researcher Representation. A researcher’s publications alone are a thin description of their scientific identity. Intuitively, two researchers with similar publication lists can have fundamentally different research identities if one lies in a dense cluster of theorists while the other bridges that cluster to computational biology. This leads us to propose learning uz_u from the researcher’s position in G, aggregating multi-hop neighborhoods that capture who their collaborators work with, where their extended network publishes, and which topics are adjacent to their community. Notably, this grounds personalization in structural properties of science: the most valuable direction for a researcher is often not an extension of their work but a structural hole, that is, a bridge between a region they inhabit and a nearby region not yet connected (1). Pipeline Personalization. Personalizing only hypothesis generation while leaving retrieval, experiment design, writing, citation, and review researcher-agnostic is internally inconsistent: a hypothesis can be aligned yet require resources the researcher lacks and cite a literature that omits their community. The design principle is to separate exploitation, which extends the researcher’s trajectory, from exploration, which uses the profile to judge feasibility and framing while keeping the candidate space broad. The goal is to personalize the how without narrowing the what. Personalization also need not stop at the stages: the components Algorithm 1 treats as fixed can themselves be learned per researcher. The weights in Eq. (4) and the NovNov, RelRel, and FeasFeas functions can be fit from research traces such as revisions, reviews, and abandoned versus published projects; feasibility can be grounded in the researcher’s actual repositories, compute, and datasets rather than a self-reported profile, and review calibrated to what they can rigorously verify, keeping autonomous science inside their expertise and auditable. Evaluation Grounded in the Individual. Predicting a researcher’s next papers is a flawed proxy: held-out papers record what they happened to work on under path-dependent incentives, not what they should have, and the proxy penalizes a co-scientist that recommends something better than anything they pursued. Evaluation should instead combine (i) feasibility alignment, whether directions are executable given documented capabilities; (i) expert-assessed quality, whether blinded experts judge ideas novel, significant, and appropriate for the profile; and (i) longitudinal impact, whether co-scientist use yields more influential work over multi-year horizons. Held-out papers nonetheless serve as a scalable necessary condition: holding out a paper at time t and building cuc_u from the record prior to t, one conditions the system on the general concept and measures fidelity, whether it recovers a hypothesis and experimental path close to what the researcher pursued, and contrast, whether the researcher-agnostic system stays generic while different contexts diverge on the same concept. High fidelity with high contrast shows cuc_u carries signal; failure falsifies the representation. Fully automatable from bibliographic data, this protocol validates the personalization signal rather than the quality ceiling, and building benchmark infrastructure for all of these criteria is itself a major missing contribution. 6. Open Challenges We now discuss key open problems and challenges, which are deep tensions in the problem itself rather than engineering obstacles. Creativity Collapse. A researcher-agnostic system applies one map from goal to output distribution, so the epistemic loss compounds at population scale: when many researchers query the same engine, the field’s portfolio of explored hypotheses collapses toward a monoculture, and globally “best” ideas are raced redundantly while directions of high counterfactual value, those only a particular researcher is positioned to pursue, go unexplored. Personalization is therefore a decorrelation mechanism for collective discovery, not merely a convenience for individuals. The challenge is to exploit a researcher’s experience without collapsing into biographical mimicry, treating it as a signal for experiments a generic system would not propose yet this researcher can develop. This calls for rewarding counterfactual complementarity: ideas unlikely under both a universal model and the user’s past work alone, yet plausible and valuable given their accumulated experience. One realization trains a mimicry model of the researcher and optimizes for high utility U(h∣g,u)U(h g,u) at low likelihood under both the mimicry and universal models. Lifecycle Dependence and Cold Start. For a senior researcher with a dense graph, value lies in adjacent unexplored territory and bridging structural holes, whereas for an early-career researcher with sparse history, the system must help establish an identity rather than extend one. Notably, this exceeds the classical cold-start problem, since the objective function itself changes with career stage. Team-Personalized Auto-Research. Most impactful research is carried out by teams rather than individuals (32), yet the formulation above personalizes for a single researcher. The framework extends naturally by replacing u with a team T⊆T and defining cT=Ψ((u,u)u∈T)c_T= (\(S_u,z_u)\_u∈ T), where a team is simply a subgraph of G. Notably, the desiderata aggregate asymmetrically: feasibility is a union, since the team can execute what any member can execute; novelty is relative to the union of prior work; and alignment is closer to an intersection, since framing must be evaluable by every community the team spans. This asymmetry is precisely why scientists collaborate, and it enables fundamentally new capabilities, including routing experiments to the member best positioned to execute them, writing and citing for the union of the team’s communities, and recommending the collaborator whose addition closes the structural hole that makes a hypothesis credible (1). Notably, Algorithm 1 supports this setting with only local modifications: Lines 4–5 encode each member and aggregate via Ψ to obtain cTc_T; retrieval (Line 6) scores documents against the team context; the utility (Line 14) becomes U(h∣g,T)U(h g,T) with Feas(h,0,cT)=maxu∈TFeas(h,0,cu) (h,W_0,c_T)= _u∈ T (h,W_0,c_u) and alignment taken as a minimum over members; implementation and execution (Lines 15–16) assign each candidate to argmaxu∈TFeas(h,0,cu) _u∈ T (h,W_0,c_u); and the feedback step (Line 21) collects critiques from multiple members, updating each uS_u and re-aggregating cTc_T. However, aggregating conflicting member preferences into a single cTc_T is a social-choice problem, the max-based feasibility assumes frictionless handoffs between members, and multi-party privacy becomes harder when the user is itself a group. Evaluation Without Ground Truth. The fundamental difficulty is that the quantity we need to evaluate, namely the value of a recommended direction for a specific researcher, is never observed. History records only the single trajectory each researcher actually followed, and that trajectory was shaped by funding, advisors, reviewing, and chance rather than by an oracle over alternatives. Scoring a system by similarity to this record therefore rewards mimicry and penalizes better recommendations (§5). The held-out protocol of §5 confirms only a necessary condition, namely that cuc_u carries researcher-specific signal; it cannot certify that a recommended direction is good, since the value of an unpursued alternative is never recorded. Resolving that distinction requires human judgment (expert panels), time (longitudinal studies), or interventions (counterfactual designs), each of which necessitates further research. 7. Conclusion We introduced personalized auto-research, the problem of conditioning the full research process on a representation of the individual researcher. The core claim is simple: a system cannot be a true co-scientist if it does not know whom it is collaborating with. The goal is complementarity rather than similarity, just as scientists choose collaborators for the expertise they lack, and so a stronger universal engine does not resolve the problem; it still returns the same high-scoring package to everyone, which makes personalization orthogonal to, and required on top of, the current SOTA. We therefore pose personalized auto-research as a grand challenge for the AI ecosystem, spanning graph learning for representation, agentic systems for execution, HCI for consent and control, and community benchmarks for individual-grounded evaluation. No single group can deliver it alone, and the tensions it raises around novelty, equity, privacy, lifecycle, and evaluation define the field as much as its algorithms do. References Burt (2004) R. S. Burt Structural holes and good ideas. American Journal of Sociology 110 (2), p. 349–399. External Links: Document Cited by: §5, §6. Chen et al. (2025) T. Chen, S. Anumasa, B. Lin, V. Shah, A. Goyal, and D. Liu Auto-Bench: an automated benchmark for scientific discovery in LLMs. External Links: 2502.15224, Document, Link Cited by: §2. Eger et al. (2025) S. Eger, Y. Cao, J. D’Souza, A. Geiger, C. Greisinger, S. Gross, Y. Hou, B. Krenn, A. Lauscher, Y. Li, C. Lin, N. S. Moosavi, W. Zhao, and T. Miller Transforming science with large language models: a survey on AI-assisted scientific discovery, experimentation, content generation, and evaluation. External Links: 2502.05151, Document, Link Cited by: §2. Gottweis et al. (2025) J. Gottweis, W. Weng, A. Daryin, T. Tu, A. Palepu, P. Sirkovic, A. Myaskovsky, F. Weissenberger, K. Rong, R. Tanno, K. Saab, D. Popovici, J. Blum, F. Zhang, K. Chou, A. Hassidim, B. Gokturk, A. Vahdat, P. Kohli, Y. Matias, A. Carroll, K. Kulkarni, N. Tomasev, Y. Guan, V. Dhillon, E. D. Vaishnav, B. Lee, T. R. D. Costa, J. R. Penadés, G. Peltz, Y. Xu, A. Pawlosky, A. Karthikesalingam, and V. Natarajan Towards an AI co-scientist. External Links: 2502.18864, Document, Link Cited by: §1, Table 1, §2, §2, §4. Gottweis et al. (2026) J. Gottweis, W. Weng, A. Daryin, T. Tu, P. Sirkovic, A. Myaskovsky, G. Glowaty, F. Weissenberger, A. Orlandi, D. Popovici, A. Palepu, K. Rong, R. Tanno, K. Saab, F. Zhang, J. Blum, A. Carroll, K. Kulkarni, N. Tomašev, D. Zverinski, I. Rendulic, E. Vedadi, F. Hasler, L. Rimanic, M. Boia, I. Budiselic, B. Feinstein, M. Bellaiche, T. Sheffer, J. Freyberg, J. Ratcliff, O. Bertolli, K. Chou, A. Hassidim, B. Gokturk, A. Vahdat, Y. Guan, V. Dhillon, E. D. Vaishnav, B. Lee, T. R. D. Costa, J. R. Penadés, G. Peltz, Y. Matias, J. Manyika, D. Hassabis, Y. Xu, P. Kohli, A. Pawlosky, A. Karthikesalingam, and V. Natarajan Accelerating scientific discovery with co-scientist. Nature. External Links: Document, Link Cited by: Table 1, §2. Jeddi et al. (2026) A. Jeddi, M. N. Le, H. C. Karaimer, K. G. Derpanis, and B. Taati GEAR: genetic AutoResearch for agentic code evolution. External Links: 2605.13874, Document, Link Cited by: §1, §2, §4. Jumper et al. (2021) J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P. Kohli, and D. Hassabis Highly accurate protein structure prediction with AlphaFold. Nature 596 (7873), p. 583–589. External Links: Document Cited by: §2. Koren et al. (2009) Y. Koren, R. Bell, and C. Volinsky Matrix factorization techniques for recommender systems. Computer 42 (8), p. 30–37. External Links: Document Cited by: §2. Kumar et al. (2024) I. Kumar, S. Viswanathan, S. Yerra, A. Salemi, R. A. Rossi, F. Dernoncourt, H. Deilamsalehy, X. Chen, R. Zhang, S. Agarwal, et al. Longlamp: a benchmark for personalized long-form text generation. arXiv preprint arXiv:2407.11016. Cited by: §2. Lewis et al. (2020) P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, Vol. 33, p. 9459–9474. External Links: Link Cited by: §2. Liu et al. (2026) J. Liu, S. Qiu, M. Li, B. Li, H. Ji, S. Han, X. Ye, P. Xia, Z. Dong, M. Chen, C. Zhang, L. Zhang, G. Chen, H. Tu, X. Yang, L. Feng, X. Zhao, H. Chen, J. Zhou, X. Wang, W. Zhang, H. Zhu, Y. Li, J. Mei, H. Fei, J. Zhang, L. Li, L. Zhang, Y. Zhou, S. Wang, C. Xiong, J. Zou, Z. Zheng, C. Xie, M. Ding, and H. Yao AutoResearchClaw: self-reinforcing autonomous research with human-AI collaboration. External Links: 2605.20025, Document, Link Cited by: §1, Table 1, §2, §4, §4. Lu et al. (2024) C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha The AI scientist: towards fully automated open-ended scientific discovery. External Links: 2408.06292, Document, Link Cited by: §1, Table 1, §2. Lu et al. (2026) C. Lu, C. Lu, R. T. Lange, Y. Yamada, S. Hu, J. Foerster, D. Ha, and J. Clune Towards end-to-end automation of AI research. Nature 651, p. 914–919. External Links: Document, Link Cited by: §2. Luo et al. (2025a) E. Luo, J. Jia, Y. Xiong, X. Li, X. Guo, B. Yu, M. Hao, L. Wei, and X. Zhang Benchmarking AI scientists for omics data driven biological discovery. External Links: 2505.08341, Document, Link Cited by: §2. Luo et al. (2025b) Z. Luo, A. Kasirzadeh, and N. B. Shah The more you automate, the less you see: hidden pitfalls of AI scientist systems. External Links: 2509.08713, Document, Link Cited by: §2. Luo et al. (2025c) Z. Luo, Z. Yang, Z. Xu, W. Yang, and X. Du LLM4SR: a survey on large language models for scientific research. External Links: 2501.04306, Document, Link Cited by: §2. Mitchener et al. (2025) L. Mitchener, A. Yiu, B. Chang, M. Bourdenx, T. Nadolski, A. Sulovari, E. C. Landsness, D. L. Barabasi, S. Narayanan, N. Evans, S. Reddy, M. Foiani, A. Kamal, L. P. Shriver, F. Cao, A. T. Wassie, J. M. Laurent, E. Melville-Green, M. Caldas, A. Bou, K. F. Roberts, S. Zagorac, T. C. Orr, M. E. Orr, K. J. Zwezdaryk, A. E. Ghareeb, L. McCoy, B. Gomes, E. A. Ashley, K. E. Duff, T. Buonassisi, T. Rainforth, R. J. Bateman, M. Skarlinski, S. G. Rodriques, M. M. Hinks, and A. D. White Kosmos: an AI scientist for autonomous discovery. External Links: 2511.02824, Document, Link Cited by: §2. Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, Vol. 35, p. 27730–27744. External Links: Link Cited by: §1, §2. Pu et al. (2025) Y. Pu, T. Lin, and H. Chen PiFlow: principle-aware scientific discovery with multi-agent collaboration. External Links: 2505.15047, Document, Link Cited by: §2. Ren et al. (2025) S. Ren, C. Xie, P. Jian, Z. Ren, C. Leng, and J. Zhang Towards scientific intelligence: a survey of LLM-based scientific agents. External Links: 2503.24047, Document, Link Cited by: §2. Salemi et al. (2024) A. Salemi, S. Mysore, M. Bendersky, and H. Zamani LaMP: when large language models meet personalization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, p. 7370–7392. External Links: Link Cited by: §1. Schmidgall and Moor (2025) S. Schmidgall and M. Moor AgentRxiv: towards collaborative autonomous research. External Links: 2503.18102, Document, Link Cited by: §1, §2. Schmidgall et al. (2025) S. Schmidgall, Y. Su, Z. Wang, X. Sun, J. Wu, X. Yu, J. Liu, M. Moor, Z. Liu, and E. Barsoum Agent laboratory: using LLM agents as research assistants. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, p. 5977–6043. External Links: Document, Link Cited by: §1, §2, §2, §4. Son et al. (2025) G. Son, J. Hong, H. Fan, H. Nam, H. Ko, S. Lim, J. Song, J. Choi, G. Paulo, Y. Yu, and S. Biderman When AI co-scientists fail: SPOT-a benchmark for automated verification of scientific research. External Links: 2505.11855, Document, Link Cited by: §2. Tang et al. (2025) J. Tang, L. Xia, Z. Li, and C. Huang AI-researcher: autonomous scientific innovation. External Links: 2505.18705, Document, Link Cited by: §1, §2. Tie et al. (2026) G. Tie, J. Shi, D. Song, Y. Huang, Z. Sheng, X. Zhou, D. Liu, P. Zhou, Y. Chen, R. Xu, L. He, Q. Wen, M. Li, C. Lu, S. Li, P. Xie, Y. Yuan, R. Meng, L. Xing, L. Sun, C. Xiong, P. S. Yu, and J. Gao AutoResearch AI: towards AI-powered research automation for scientific discovery. External Links: 2605.23204, Document, Link Cited by: §1, §2. Wang and Barabási (2021) D. Wang and A. Barabási The science of science. Cambridge University Press. External Links: Document Cited by: §1. Wang and Luan (2026) Y. Wang and Z. Luan PARNESS: a paper harness for end-to-end automated scientific research with dynamic workflows, full-text indexing, and cross-run knowledge accumulation. External Links: 2605.05258, Document, Link Cited by: §1, §2, §4. Wei et al. (2022) J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, Vol. 35, p. 24824–24837. External Links: Link Cited by: §2. Wei et al. (2025) J. Wei, Y. Yang, X. Zhang, Y. Chen, X. Zhuang, Z. Gao, D. Zhou, G. Wang, Z. Gao, J. Cao, Z. Qiu, M. Hu, C. Ma, S. Tang, J. He, C. Song, X. He, Q. Zhang, C. You, S. Zheng, N. Ding, W. Ouyang, N. Dong, Y. Cheng, S. Sun, L. Bai, and B. Zhou From AI for science to agentic science: a survey on autonomous scientific discovery. External Links: 2508.14111, Document, Link Cited by: §2. Weng et al. (2025) Y. Weng, M. Zhu, Q. Xie, Q. Sun, Z. Lin, S. Liu, and Y. Zhang DeepScientist: advancing frontier-pushing scientific findings progressively. External Links: 2509.26603, Document, Link Cited by: Table 1, §2. Wuchty et al. (2007) S. Wuchty, B. F. Jones, and B. Uzzi The increasing dominance of teams in production of knowledge. Science 316 (5827), p. 1036–1039. External Links: Document Cited by: §6. Yamada et al. (2025) Y. Yamada, R. T. Lange, C. Lu, S. Hu, C. Lu, J. Foerster, J. Clune, and D. Ha The AI scientist-v2: workshop-level automated scientific discovery via agentic tree search. External Links: 2504.08066, Document, Link Cited by: §1, Table 1, §2, §4. Yang et al. (2025) X. Yang, X. Yang, S. Fang, Y. Zhang, J. Wang, B. Xian, Q. Li, J. Li, M. Xu, Y. Li, H. Pan, Y. Zhang, W. Liu, Y. Shen, W. Chen, and J. Bian R&D-Agent: an LLM-agent framework towards autonomous data science. External Links: 2505.14738, Document, Link Cited by: §2. Yao et al. (2023) S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations, External Links: Link Cited by: §2. Zhang et al. (2025) P. Zhang, X. Hu, G. Huang, Y. Qi, H. Zhang, X. Li, J. Song, J. Luo, Y. Li, S. Yin, C. Dai, E. H. Jiang, X. Zhou, Z. Yin, B. Yuan, J. Dong, G. Su, G. Qiao, H. Tang, A. Du, L. Pan, Z. Lan, and X. Liu aiXiv: a next-generation open access ecosystem for scientific discovery generated by AI scientists. External Links: 2508.15126, Document, Link Cited by: §1, §2. Zheng et al. (2025) T. Zheng, Z. Deng, H. T. Tsang, W. Wang, J. Bai, Z. Wang, and Y. Song From automation to autonomy: a survey on large language models in scientific discovery. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, p. 17733–17750. External Links: Document, Link Cited by: §2. Zhou et al. (2025) L. Zhou, H. Ling, C. Fu, Y. Huang, M. Sun, W. Yu, X. Wang, X. Li, X. Su, J. Zhang, X. Chen, C. Liang, X. Qian, H. Ji, W. Wang, M. Zitnik, and S. Ji Autonomous agents for scientific discovery: orchestrating scientists, language, code, and physics. External Links: 2510.09901, Document, Link Cited by: §2. Zhu et al. (2025a) K. Zhu, J. Zhang, Z. Qi, N. Shang, Z. Liu, P. Han, Y. Su, H. Yu, and J. You SafeScientist: toward risk-aware scientific discoveries by LLM agents. External Links: 2505.23559, Document, Link Cited by: §2. Zhu et al. (2025b) M. Zhu, Q. Xie, Y. Weng, J. Wu, Z. Lin, L. Yang, and Y. Zhang AI scientists fail without strong implementation capability. External Links: 2506.01372, Document, Link Cited by: §2.