Paper deep dive
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
Judah Goldfeder, Philippe Wyder, Yann LeCun, Ravid Shwartz Ziv
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/20/2026, 7:59:20 AM
Summary
The paper critiques the concept of Artificial General Intelligence (AGI) as a flawed and overly broad goal for AI development, arguing that human intelligence is not truly general but specialized. It introduces Superhuman Adaptable Intelligence (SAI) as a more precise and useful framework, defined as AI that can learn to exceed human performance in any important task and fill human skill gaps. The authors argue that focusing on SAI, which emphasizes adaptability and specialization via self-supervised learning and world models, leads to clearer goals and more rapid progress than the ambiguous pursuit of AGI.
Entities (12)
Relation Signals (9)
Judah Goldfeder â authored â AI Must Embrace Specialization via Superhuman Adaptable Intelligence
confidence 99% · AI Must Embrace Specialization via Superhuman Adaptable Intelligence Judah Goldfeder * 1
Ravid Shwartz-Ziv â authored â AI Must Embrace Specialization via Superhuman Adaptable Intelligence
confidence 99% · Yann LeCun 3 Ravid Shwartz-Ziv 3
Yann LeCun â authored â AI Must Embrace Specialization via Superhuman Adaptable Intelligence
confidence 99% · Philippe Wyder * 2 Yann LeCun 3
Philippe Wyder â authored â AI Must Embrace Specialization via Superhuman Adaptable Intelligence
confidence 99% · Judah Goldfeder * 1 Philippe Wyder * 2
Superhuman Adaptable Intelligence â definedas â intelligence that can learn to exceed humans at anything important that we can do
confidence 98% · SAI is defined as intelligence that can learn to exceed humans at anything important that we can do, and that can fill in the skill gaps where humans are incapable.
Artificial General Intelligence â criticizedby â Superhuman Adaptable Intelligence
confidence 95% · We argue that AI must embrace specialization, rather than strive for generality... and introduce Superhuman Adaptable Intelligence (SAI).
Human Intelligence â isnot â General
confidence 95% · We argue that this notion is fundamentally misguided... In truth, we are only good at the specific subset of tasks that are important to our existence
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Everyone from AI executives and researchers to doomsayers, politicians, and activists is talking about Artificial General Intelligence (AGI). Yet, they often don't seem to agree on its exact definition. One common definition of AGI is an AI that can do everything a human can do, but are humans truly general? In this paper, we address what's wrong with our conception of AGI, and why, even in its most coherent formulation, it is a flawed concept to describe the future of AI. We explore whether the most widely accepted definitions are plausible, useful, and truly general. We argue that AI must embrace specialization, rather than strive for generality, and in its specialization strive for superhuman performance, and introduce Superhuman Adaptable Intelligence (SAI). SAI is defined as intelligence that can learn to exceed humans at anything important that we can do, and that can fill in the skill gaps where humans are incapable. We then lay out how SAI can help hone a discussion around AI that was blurred by an overloaded definition of AGI, and extrapolate the implications of using it as a guide for the future.
Tags
Links
- Source: https://arxiv.org/abs/2602.23643v1
- Canonical: https://arxiv.org/abs/2602.23643v1
Trouble viewing inline? Open PDF directly â
Full Text
50,514 characters extracted from source content.
Expand or collapse full text
AI Must Embrace Specialization via Superhuman Adaptable Intelligence Judah Goldfeder * 1 Philippe Wyder * 2 Yann LeCun 3 Ravid Shwartz-Ziv 3 Abstract Everyone from AI executives and researchers to doomsayers, politicians, and activists is talking about Artificial General Intelligence (AGI). Yet, they often donât seem to agree on its exact defi- nition. One common definition of AGI is an AI that can do everything a human can do, but are humans truly general? In this paper, we address whatâs wrong with our conception of AGI, and why, even in its most coherent formulation, it is a flawed concept to describe the future of AI. We explore whether the most widely accepted defini- tions are plausible, useful, and truly general. We argue that AI must embrace specialization, rather than strive for generality, and in its specialization strive for superhuman performance, and introduce Superhuman Adaptable Intelligence (SAI). SAI is defined as intelligence that can learn to exceed humans at anything important that we can do, and that can fill in the skill gaps where humans are in- capable. We then lay out how SAI can help hone a discussion around AI that was blurred by an overloaded definition of AGI, and extrapolate the implications of using it as a guide for the future. 1. Introduction The AI community has become increasingly fractured over where the field is headed. On one side are âdoomers,â who argue we are headed towards a gruesome societal endgameâ mass unemployment, loss of human agency, and a future in which humanity becomes subordinate to artificial overlords. On the other side are those who expect advanced artificial in- telligence to bring something close to utopia, ending hunger, suffering, and scarcity. A third camp frames AI as a ânor- mal technology,â forecasting major impacts but rejecting extreme narratives (Narayanan & Kapoor, 2025). Central to all of these views is the concept of Artificial General Intelligence or AGI. Yet, as is often the case in 1 Columbia University, New York, NY, USA 2 Distyl, New York, NY, USA 3 New York University, New York, NY, USA. Correspon- dence to: Judah Goldfeder<jag2396@columbia.edu>. Preprint. March 2, 2026. widely public debates, much of the disagreement stems less from evidence than from terminology: AGI is invoked constantly, but rarely defined precisely, and the resulting ambiguity has made the debate far more confusing and far more polarizedâthan it needs to be. Much of the discourse uses human intelligence as a paradigm of generality, but we argue that this notion is fun- damentally misguided. As humans, we struggle to perceive our own blind spots; this leads to the illusion of generality. In truth, we are only good at the specific subset of tasks that are important to our existence, but are completely incapable of performing tasks outside this narrow range. Awareness of human limitation gives rise to a critical realization: humans may be specialized creatures, but are nonetheless capable of accomplishing or quickly learning a wide range of in- credible things. We argue that the current focus on AGI and generality as the North Star of the field, should be replaced with an emphasis on adaptability, including the time it takes to learn a new task, and the range of tasks capable of be- ing learned. We refer to this as Superhuman Adaptable Intelligence (SAI). A natural corollary of an emphasis on adaptability is the need for a model with strong assumptions about the world. This suggests self-supervised learning (SSL) as a promising way to acquire generic knowledge, and world models as a useful mechanism for planning and zero-shot task transfer. We believe that recentering the dis- course around SAI will lead towards better communication, clearer goals, and more rapid progress. [Pos. #1] Human intelligence is not general in any meaningful way [Pos. #2] Generality is not a requirement for an intelli- gence to be extremely useful [Pos. #3] There is no consensus on the meaning of the term AGI in industry or academia [Pos. #4] Existing definitions are insufficient [Pos. #5] We should instead focus on Superhuman Adaptable Intelligence, which points toward SSL and world models 1 arXiv:2602.23643v1 [cs.AI] 27 Feb 2026 AI Must Embrace Specialization via Superhuman Adaptable Intelligence 2. Human Intelligence is Specialized While the idea of human intelligence as the paradigm of generality is ubiquitous in the literature, two related but distinct notions of generality are often conflated: 1.The average, educated human is capable of a wide range of tasks that are very âgeneralâ in nature, and enable a wide range of objectives to be accomplished. This includes things like complex planning and loco- motion, fine motor skills, abstract thinking, self simu- lation, spatial reasoning, and visual understanding. 2. Human intelligence as a whole is âgeneralâ because it can be specialized to âanyâ given task, whether it be medicine, advanced mathematics, plumbing, or playing chess. Both of these claims make the same error: circularly defin- ing generality in human terms, and then asserting hu- manity as its paradigm. Evolution has honed humanity over time to be highly spe- cialized in the domain of skills necessary for survival in the physical world. The things most innate to us are not always the most simple, but the most critical for our sur- vival. This observation has given rise to Moravecâs Paradox, where the tasks we find easiest, like locomotion, are difficult for computers, but tasks that we find difficult, which are not essential to our survival, like playing chess, turn out to be much easier for computers. This clearly illustrates the illusion of our generality. While the average abilities of an educated human are truly remarkable, one only need ask them to play chess like a grandmaster, or compose a musi- cal symphony like Beethoven, to truly realize the hubris of calling such an intelligence general. The above argument serves to dispel the first definition of generality. The second notion of generality is more subtle in its error. By identifying specialization/adaptation as the core component of generality, it is closer to the definition of SAI that we are arguing in favor of. However, our point of contention is that we object to calling human adaptation general. While we are excellent at adapting to the tasks that were of high evolutionary importance, we are simply incapable of adapting to many tasks outside of this range at a high level. Take chess as an example. Magnus Carlsen is widely regarded as the greatest chess player of all time, and as such represents the pinnacle of human adaptation when it comes to playing chess. But this begs the question: Is Magnus actually any good at chess? When compared with the best computers, the answer is clearly no. Even more damning is that with modern day computers, creating a program that plays chess at a much higher level than Magnus is not particularly difficult. Our perception of his ability is colored by the limitations of humanity. Humans in general are bad at chess; Magnus is much better than most humans. The conclusion, that Magnus is good at chess, is a perfect illustration of our own human centric biases. Magnus Carlsen is not objectively good at chess, he is good at chess with respect to human performance levels. By any objective metric, playing chess at a much higher level is not difficult from a computational perspective, but it is something that humans are incapable of. Relatedly, many animals can perform tasks that humans cannot do at a high level, such as echolocation. So what are humans then, if not a paradigm of general intelligence? The evidence points to specialized adaptation. We have an incredible ability to adapt and specialize, within the range of tasks that we evolved to address (Russell & Norvig, 2010). This is the main contention of [Pos. #1]. Alternative Views Several objections have been raised to our assertion that hu- mans are not general. Elon Musk and Demis Hassabis have claimed that our argument conflates General Intelligence with Universal Intelligence. They further argue that the hu- man brain is indeed general in the Turing Machine sense, capable of learning anything computable given enough time, memory, and data. They therefore claim that âbrains are the most exquisite and complex phenomena we know of in the universe (so far), and they are in fact extremely gen- eralâ (Hassabis, 2025; Musk, 2025). In response, it is indeed important to clarify terms. Univer- sal Intelligence refers to the ability to act intelligently over all computable environments (Legg & Hutter, 2007). Gen- eral Intelligence, as Demis is using it, seemingly refers to adaptation to any computable task given time and resources. Far from conflating the terms, we are arguing that humans are not capable of either of these things. While the issue of whether approximate Turing- completeness under idealized conditions matters for defining intelligence is a legitimate question, it is missing the point. Even if we grant the fact that human brains are approximately Turing-complete (far from an obvious fact), under real constraints such as finite memory, finite time, and finite attention, we handle only a tiny sliver of possible problems. The space of possible functions is unimaginably vast, and we can represent an infinitesimal fraction. We feel general because we canât perceive our blind spots, not because we lack them. 3.Implications for AI North Star Terminology 3.1. A Survey of Definitions Measuring machine intelligence is non-trivial. Language- based tests, such as the Turing Test, where a machine has 2 AI Must Embrace Specialization via Superhuman Adaptable Intelligence LEARN (Adaptability) DO (Performance) AnythingAnything ImportantAnything Humans Can Do Anything Important Humans Can Do Legg & Hutter (2007) Universal Intelligence Xu (2024) Open Environments Chollet (2019) Skill-Acquisition Efficiency Hendrycks (2025) Cognitive Capabilities Morris et al. (2023) Levels of AGI Wozniak (2010) Embodiment âCoffeeâ Test Nilsson (2005) Employment Test OpenAI (2018) Economically Valuable Legend Adaptive Generalists (Focus: Learning) Cognitive Mirrors (Focus: Human Tasks) Economic Engines (Focus: Utility/Jobs) SAI Chollet (2019) Skill-Acquisition Efficiency Adaptable Intelligence inside and outside human domain Figure 1. A two-dimensional semantic map organizing prominent definitions for AGI and other North Star measures of artificial intelligence, along two axes. The vertical axis represents the source of intelligence, ranging from performance-based capabilities (DO, bottom) to learning and adaptability (LEARN, top). The horizontal axis represents the scope of tasks, from universal/open-ended domains (left) to human-centric and economically-focused domains (right). Definitions cluster into three categories: Adaptive Generalists (teal) emphasize learning efficiency and generalization in open environments; Cognitive Mirrors (violet) focus on replicating human-level cognitive capabilities across broad task domains; Economic Engines (orange) prioritize practical utility and economic value in human- relevant tasks. Superhuman Adaptable Intelligence (SAI) falls into the realm of adaptable AI that can do anything that is important both inside and outside the human realm. to pretend to be a human well enough to fool a human to believe the machine is human (Turing, 1950) and the Wino- grad schema challenge that tests common-sense reasoning and natural language understanding (Levesque et al., 2012) are helpful to measure aspects of intelligence, but not a true measure of whether AGI is achieved. Steve Wozniakâs Coffee Testâwhether a machine could make a cup of cof- fee if sent to a random kitchenâdraws attention to the fact that, despite claims to the contrary(Chen et al., 2026), lan- guage alone is not sufficient to be considered intelligent and that human intelligence adapts well to unseen environ- ments (Wozniak, 2010). AGI definitions commonly fall into categories along two axes, the first one defining what capabilities we are referring to, and the second defining the required scope of those capabilities: 1.Axis 1 (capability): (A) AI that can learn to do tasks vs (B) AI that can do tasks out of the box 2.Axis 2 (scope): (I) Anything, (I) Anything important, (I) Anything humans can do, (IV) Anything humans can do that is important We visualized popular definitions of AGI in accordance with this two-dimensional framework in Fig. 1. There is a reasonable argument to be made for a third axis that spans the space from observable capability to subjective understanding (Searle, 1980), thereby including the dimen- sion of the internalist view. According to this view, an AI could meet any performance benchmark for AGI, yet if it lacks subjective experience (qualia), it remains merely a simulation of intelligence rather than the genuine article. While the exploration of this dimension is profound, we consider it outside the scope of this work. Our focus is on operational definitionsâmetrics that can be observed and measuredâwhereas the internalist objection currently resides in the realm of metaphysics and philosophy of mind, offering no falsifiable test for engineering progress. Regardless, one thing is clear: AGI as a term is overloaded with varying definitions from high-impact sources. This confusion has even led to claims that AGI has already ar- rived(Ag Ì uera y Arcas & Norvig, 2023; Chen et al., 2026). The varying definitions plotted in Fig. 1 and the imprecise nature of the public discussions being had by high-profile 3 AI Must Embrace Specialization via Superhuman Adaptable Intelligence individuals around AGI, as shown in the previous section, clearly demonstrate [Pos. #3]. 3.2. Why Existing Definitions are Insufficient [Pos. #1] has the following implications for these defini- tions: 1. Humanity is still quite useful, so AI does not need to be general to still be groundbreaking and powerful ([Pos. #2]). 2. Any definition focused exclusively on humanity as a goal cannot claim to be general. 3.Focusing exclusively on humans is also not ideal, since there are many tasks we cannot do that are still high utility and important. In addition, for a definition to be useful, it must meet the following criteria: 1.It must be feasible. If a goal is not possible to be re- alized from a theoretical perspective, it provides ques- tionable value. 2.It must be internally consistent. If a definition claims to be general, it must actually be general in a meaningful way. 3. It must be assessable. The goal as presented should lead to clear subgoals and strategies, and there must be a clear metric with which to measure progress. Having established these criteria, we can now demonstrate [Pos. #4], namely that existing definitions come up short. First, definitions of AGI that claim true generality fall prey to the âNo Free Lunchâ theoremâno single, general-purpose machine learning algorithm or optimization strategy works best for every problem (Wolpert & Macready, 1997). Or to frame it differently: given finite energy, an approach that directs available energy towards learning a finite set of tasks will reasonably outperform an approach that distributed the finite energy over an infinite amount of tasks. At the limit, the amount of energy dedicated to each of the infinite tasks approaches zero. Thus, any definition that defines the scope as literally anything computable fails our criteria by not being feasible. Second, any definition of AGI that focuses on a subset of tasks, or that emphasizes specialization and adaptation as key metrics, can not truly be said to be general. Sim- ilarly, AGI measured by the âgeneralâ nature of humans is not truly general. Chollet acknowledges this problem and states that human intelligence âis only âgeneralâ in a limited senseâ (Chollet, 2019), but we contend that this is simply an inherent contradiction of terminology. For this reason, Shane Legg and Marcus Hutter speak of âUniversal Intelligenceâ rather than AGI because human intelligence is âfar too limitedâ (Legg & Hutter, 2007), since a definition of AGI that is human-centric excludes the infinite space of non-human intelligence (Wang, 2019). Despite the above objections, AGI defined specifically as the ability to match human cognitive breadth is quite popular: Hendrycks et al. and Morris et al. argue that human generality is the only general example for the concept of intelligence (Hendrycks et al., 2025; Morris et al., 2024). While the above empha- sized the ability to do anything humans can do, others argue that AGI must be able to learn to do or do anything impor- tant that humans can do. Their arguments acknowledge that the domain of human intelligence is finite and that it is desirable for AI to be able to perform or learn to perform a subset of important tasks: tasks that generate economic value. In the words of Nilsson, âSystems with true human- level intelligence should be able to perform the tasks for which humans get paidâ (Nilsson, 2005). 1 , an idea further echoed in the OpenAI Charter (OpenAI). We suspect that the cause for such a widespread conflation of human intelligence with generality stems from the urge for self-flattery, and the difficulty of truly conceiving of our own limitations. Regardless, all the definitions in this category fail our criteria by not being internally consistent. One might raise the objection that our contention here is merely one of semantics, and that these definitions can still be valuable North Stars for the field, even if they misuse the term âgeneralâ. In response, we argue that when defining the end goal of an entire field, semantics are extremely important. A misapprehension of generality is dangerous for several reasons. It obscures how such an intelligence can actually be realized, which violates our feasibility criteria. Further, it can lead to unnecessarily narrow conceptions of what the end goal should be. For example, the belief that humans are general has led to several definitions of AGI as mimicking humans, which is certainly far too limited a goal for what AI is and can become. Third, definitions of AGI that cannot be assessed or evalu- ated are not practical or useful. Legg and Hutter acknowl- edge this issue in their paper: âvarious practical challenges will need to be addressed before universal intelligence can be used to construct an effective intelligence testâ (Legg & Hutter, 2007). The ability to measure progress is critical for many reasons. An enormous body of evidence suggests that the precise ability to measure progress is one of the 1 Nilsson doesnât use the term AGI, he speaks of âstrong AIâ or âhuman-level artificial intelligenceâ (the term was popularized later by Shane Legg and Ben Goertzel), but still pushes the idea of humans being âmore-or-lessâ general purpose 4 AI Must Embrace Specialization via Superhuman Adaptable Intelligence AGI DefinitionFailure modeExplanation âMatch or exceed the cognitive versatility and proficiency of a well-educated adult.â (Hendrycks et al., 2025) Not ConsistentHuman cognition is not general in any meaningful way. This definition is also unnecessarily narrow âHighly autonomous systems that outperform humans at most economically valuable work.â (Morris et al., 2024) Not ConsistentThe focus here is explicitly on a subset of tasks that are of economic worth. Clearly not General âA system that should be able to do pretty much any cognitive task that humans can do.â â Demis Hassabis (DeepMind CEO) (Mitchell, 2024) Not ConsistentThis definition is not actually general both in its focus on humans, and also in its focus on âcognitive tasksâ, which seems to be to the exclusion of physical tasks like locomotion âWe need precise, quantitative definitions and measures of intelligence â in particular human-like general intelligence.... ...The intelligence of a system is a measure of its skill-acquisition efficiency over a scope of tasks, with respect to priors, experience, and generalization difficultyâ (Chollet, 2019) Not ConsistentChollet himself admits as much, calling human cognition âonly âgeneralâ in a limited senseâ, a contradiction of terms âWe define AGI as a system that demonstrates broad generality (performing a wide range of tasks) and high performance (matching or exceeding human levels).â (Morris et al., 2024) Not FeasibleWhile they acknowledge generality requires exceeding human levels, with a focus on direct performance over adaptation, such a system is not realizable with finite resources. âIntelligence measures an agentâs ability to achieve goals in a wide range of environments.â (Legg & Hutter, 2007) Not FeasibleLegg and Hutter define the domain of environments as all that are computable. They further emphasize ability over adaptability. Strong ability on such a vast set of tasks is not realizable with finite resources âHighly autonomous systems that outperform humans at most economically valuable workâ (OpenAI). Not Assessable The focus on performance means that any evaluation would have to benchmark against an ever growing set of tasks Table 1. The failure of most AGI definitions. Note: some definitions fail for multiple reasons, but we only highlight one. strongest catalysts of progress itself (Wyder et al., 2025). Relatedly, clear metrics usually give an idea of what sorts of subgoals and strategies are useful. Even more fundamen- tally, a definition that is not measurable is not really much of a definition at all, and is often indicative of a lack of precision, or a hand-wavy nature. This criterion highlights a key difference between the two categories of definitions on our first axis (capability). Any definition that focuses on learning or adapting implicitly has a clear metric with which to evaluate intelligence: speed of adaptation to new tasks. Conversely, definitions focused on doing and performing often lack any obvious way to measure this, other than benchmarking the AIâs ability to do everything, an ever-expanding and ill-defined set of bench- marks. Table 1 elaborates on our issues with many popular AGI definitions ([Pos. #4]). 4. Why Specialization Wins To motivate [Pos. #5], it behooves us to explore the impor- tance of specialization. Specialization is not an accident of biology; it is a predictable consequence of limited resources, competing objectives, and environments that reward per- formance on a small subset of evolutionarily relevant chal- lenges. Forister et al. state that a generalist organism carries genetic traits suited to various environments, but never the ideal combination for thriving in any one of them (Forister et al., 2012). Organisms face persistent trade-offs: improv- ing performance on one niche often reduces performance elsewhere, and selection therefore tends to favor designs that are sharply tuned to the local payoff landscape rather than uniformly competent across all possible conditions (Fu- tuyma & Moreno, 1988). In markets and organizations, the same logic appears under a different name: entities that fail to meet the performance threshold disappear, so competi- tion acts as a selection mechanism that amplifies effective strategies and eliminates ineffective ones (Hannan & Free- man, 1977; Loasby, 1983). AI systems are not exempt from this pressure: models that are too costly, too unreliable, or insufficiently accurate in the domains that matter will be ne- glected in favor of systems that are better matched to those domains. In machine learning, the core mathematical point is that performance gains require assumptions about the problem class i.e. the target distribution. Again, âNo Free Lunchâ. An algorithm wins by being a good fit for the target problem. As AI improves, specialized systems can improve too: if it is possible to attain a higher performance on a task, a system that concentrates that capability on a narrower task can typically realize larger gains than a system that must spend capacity and compute covering additional unrelated tasks. 5 AI Must Embrace Specialization via Superhuman Adaptable Intelligence Practically, this means that generality is intractable. Al- though multi-task learning can benefit performance when tasks share an underlying structure, it can lead to ânegative transferâ when tasks compete for representational capac- ity or impose conflicting gradients, and thereby harm task performance (Ruder, 2017). Models that route queries to specialized subsets of model parameters depending on the task are a technological acknowledgment of this limitationâ these systems attain breadth and scale through repeated, modular specialization rather than uniform shared param- eters for all inputs(Fedus et al., 2022). Although seem- ingly âgeneralâ, these models achieve their best performance through internal specialization. Universal generality is a theoretical concept, but in practical terms it is a myth. A large fraction of what we intuitively mean by âdo anythingâ reduces to planning and decision- making under uncertainty. Classical planning problems quickly become intractable in worst case (e.g., propositional STRIPS variants) (Bylander, 1994), and probabilistic plan- ning inherits similarly severe complexity barriers (Littman et al., 1998). This does not mean planning is impossible in practice; it means that broad generality across arbitrary envi- ronments has no reason to be computationally cheap. A spe- cialized agent that restricts the space of environments, goals, and action models it must handle can leverage structure and avoid worst-case blowups. This is similar for humans, as our biases, genetic makeup, and environment naturally drive us towards âhuman things,â a mere sliver of universal generality. Empirically, specialized AI systems repeatedly demonstrate the advantage of concentrating model design, data curation, and evaluation around a single domain objective. Protein structure prediction is an archetypal example: AlphaFold achieved dramatic gains by targeting a specific scientific task with task-specific training and architectural choices, and it set a new bar for accuracy and usefulness in that domain (Jumper et al., 2021). It is therefore plausibleâ indeed, expected under both the No Free Lunch framing and negative-transfer dynamicsâthat an AI system asked to âfold proteins and fold laundryâ will not match a protein- folding specialist on protein-folding performance unless it internally recovers specialization (e.g., via routing, modular- ity, or dedicated submodels) (Wolpert & Macready, 1997; Ruder, 2017; Fedus et al., 2022; Jumper et al., 2021). Specialization also clarifies why AI can be uniquely valu- able: it can target precisely the domains where human cogni- tion is systematically miscalibrated. Humans exhibit stable biases and heuristics that are often sensible under ances- tral constraints but error-prone in modern settings (Tversky & Kahneman, 1974). More broadly, the evolutionary mis- match hypothesis argues that many psychological mecha- nisms were tuned for past selection regimes and can there- fore produce maladaptive outputs in contemporary envi- ronments (Li et al., 2018). This creates an opportunity: specialized AI systems can be designed to excel exactly where humans are weak but where correctness now matters (e.g., high-dimensional statistical inference, optimization under constraints, complex mechanistic modeling) (Tversky & Kahneman, 1974; Domingos, 2012). Finally, none of this implies that generality is âbadâ. It im- plies a narrower, more operational claim: we must embrace specialization rather than fight it. Even in domains that feel like demonstrations of âgeneral intelligence,â the history of AI milestones frequently reflects intense domain targeting rather than broad competence, while newer âgeneralâ meth- ods still succeed by exploiting strong structure in the task family (Silver et al., 2018). For high-stakes applications (e.g., scientific discovery, medicine), the correct aspiration is not to preserve the romance of a single generalist mind, but to build the strongest available specialistsâand, where needed, compose them into systems whose coordination is engineered rather than assumed. We should also note that this claim does not dispute the bitter lesson (Sutton, 2019). The bitter lesson is the obser- vation that approaches that scale with computational power tend to outperform ones based on domain knowledge, a claim that we agree with. The diminishing usefulness of domain knowledge is distinct from the usefulness of do- main specialization. As scaling progresses, we will need to know less about proteins to build a system that does protein folding; however, such a system still benefits from focusing specifically on proteins. Figure 2. Illustration of the task space overlap between the human domain and the AI domain within the universal task space. Awareness of the ânarrownessâ of humans and the benefit of specialization allows us to exploit the complementary nature of AI as it is filling in the gaps in the human domain 6 AI Must Embrace Specialization via Superhuman Adaptable Intelligence where it matches or eclipses human performance, while also being able to perform tasks outside the human domain (see figure 2). 5. Towards Superhuman Adaptable Intelligence Given the utility of specialization, we propose Superhuman Adaptable Intelligence (SAI) as the id Ì e fixe of AI research ([Pos. #5]). Unlike the earlier AI North Star terminologies that we challenged, our definition of SAI sidesteps issues of feasibility by focusing on adaptation to tasks with human utility, as opposed to the performance of simply doing the task. We embrace the necessity of specialization, and avoid the pitfall of claiming generality. Further, we broaden the task domain beyond the human task domain, while not re- quiring that the AI is master of the human task domain as a whole. Finally, adaptation speedâthe speed with which an agent can acquire new skills and learn new tasks, can be measured, and thus our approach is practical. Definition Superhuman Adaptable Intelligence (SAI) is capa- ble of adapting to exceed humans at any task humans can do, while also being able to adapt to tasks outside the human domain that have utility. Our definition is most similar to Cholletâs (Chollet, 2019), except that we object to calling such a definition general, and also reject his view that we âshould benchmark progress specifically against human intelligenceâ. While human per- formance can be a useful reference point during early devel- opment, we argue that anchoring benchmarks to human baselines is ultimately orthogonal to the route to super- human capability. AI models and systems that optimize well-defined objectives and improve through self-play, evo- lutionary search, or large-scale exploration in simulation can surpass human performance without imitation (Zhao et al., 2025). We believe that over-indexing on âhuman- levelâ metrics risks misspecifying the target and limiting evaluation to anthropocentric tasks and constraints. More broadly, any evaluation scheme that treats intelligence as a checklist of fixed competenciesâwhether anchored to human baselines or to an ever-growing catalog of tasksâ misses the point of SAI. Instead, the focus should be on minimizing adaptation time. The space of possible skills is effectively unbounded, so individually testing skills be- comes a Sisyphean endeavor. Figure 3. Illustration of autoregressive model divergence Metric SAI is measured by the speed with which it takes an agent to acquire new skills and learn new tasks. Our vision towards SAI as a North Star is potentially real- izable via self-supervised learning (SSL). We believe that learning in the embedding space as opposed to in the token space may drive performance gains. We also believe that world models may help us advance towards SAI. Simultane- ously, we reject the concept of a single model or architecture as the âone paradigm to rule them all,â as it would suggest that the evolution of artificial intelligence will come to a halt once that architecture has been discovered. It is also important to note that our definition emphasized tasks outside the human domain that have utility. The pur- pose of this clause was to exclude a potential infinitely set of useless tasks from our definition, but we have not as of yet precisely defined utility, or how we determine task im- portance. Many definitions have been proposed, such as economic value or societal agreement. The exact definition one prefers is largely orthogonal to our arguments here, and we leave debating which one is most appropriate to other work. 5.1. Why Self Supervised Learning By shifting the focus from performance to adaptation, SAI points to SSL as a potential pathway. Specializing to a wide 7 AI Must Embrace Specialization via Superhuman Adaptable Intelligence range of tasks requires the ability to learn generic knowl- edge. In many real-world settings, supervised learning is not feasible in practice because it presupposes access to large, reliably labeled datasets (LeCun et al., 2015)âan assump- tion that often fails outside carefully curated benchmarks. In contrast, SSL can be applied to any data that contain ex- ploitable internal structure (Balestriero et al., 2023). Further, perhaps even more powerful, SSL has actually been shown to be on par with and even exceed SL even when supervision is abundant (He et al., 2020; Grill et al., 2020; Chen et al., 2020; He et al., 2022). SSL fueled the rise of GPTs, and has reached SOTA performance in most domains. 5.2. The substrate for fast adaptation Adaptation and specialization can be produced by many architectures and paradigms, yet which architecture is most performant remains an open research question. Designing maximally adaptable algorithms remains a central pursuit of meta learning (Finn et al., 2017). The brain is not a monolith, but a system of systems. This suggests that no single system will be able to adapt in the way that humans do. Thus, we believe that adaptation requires hierarchy and diversity of models and modalities. Specifically, we believe that adaptation is benefitted by a world model and by moving from token level prediction to latent prediction architectures such as Dreamer 4, Genie 2, or Joint Embedding Prediction Architecture (JEPA) (Van As- sel et al., 2025; Assran et al., 2023; Hafner et al., 2025; Bruce et al., 2024). Pixels are not state. The physical world is too rich and too stochastic for pixel-level prediction to be a meaningful objective; what matters is learning and fore- casting a compact representation that captures the systemâs dynamics. It has long been posited that humans and animals make heavy use of world models in their cognition (Craik, 1967). A world model allows for simulation, and therefore planning (Schrittwieser et al., 2020). As such, it is the hall- mark of zero shot and few shot adaptation (LeCun, 2022). Although, we find this argument towards a particular group of architectures persuasive, SAI doesnât dictate a specific architecture. 5.3. On the importance of diversity Homogeneity kills research. Autoregressive LLMs and LMMs have become the dominant architecture in the state- of-the-art âgeneralâ AI space (Huang et al., 2024; Su, 2025). This concentration is understandableâshared tooling and benchmarks create momentumâbut it also narrows the search space. Progress is most rapid when a greater di- versity of solutions are explored. In addition to slowing progress, these homogeneous solu- tions are often only local optima. GPTs and similar autore- gressive models are no exception, they have many flaws (Lin et al., 2021); their errors diverge exponentially with predic- tion length (LeCun, 2024), as shown in figure 3. In practice, compounding prediction error makes long-horizon interac- tion brittle. SAI counters homogenization and drives diver- sity in AI development. It provides a more coherent and reasonable target that fosters a diversity of specialization profiles. Embracing specialization counteracts incentives that lead to fast convergence towards the mean. 6. Discussion The AGI discourse is often framed as a single destination, benchmarked against an ill-defined notion of âhuman-levelâ generality. We argue that this framing is both scientifically unhelpful and operationally misleading. Human intelligence is not a universal competence engine; itâs a collection of specialized capabilities shaped by constraints and selective pressures. There is no reason to expect the most capable artificial systems to mirror the human task distribution, nor to treat human performance as the natural reference point for progress. We propose Superhuman Adaptable Intelligence (SAI) as a more concrete and productive North Star: the ability to rapidly adapt to important tasks inside and outside the hu- man domain. The central quantity is not a checklist of skills, but the speed and efficiency with which new skills are acquired under realistic resource constraints. This re- frames evaluation away from human-centric benchmarks and toward measurable adaptation dynamics. Key Insight The AI that folds our proteins should not be the AI that folds our laundry! Finally, SAIâs specialization focus fosters an environment that promotes diverse engineering approaches. Progress wonât come from a single architecture optimized for next- token prediction. We believe instead that systems that learn general latent structure from unlabeled data, build world models that support planning, and compose specialized mod- ules are better suited to fast adaptation. Put another way: it is highly unlikely that an AI tasked to fold both proteins and laundry will exceed a protein-folding specialist at protein folding or a laundry-folding specialist at laundry folding. Given limited resources, capability should be allocated to the tasks that carry utility rather than to an anthropocen- tric notion of universal competence. One promising path forward is therefore to emphasize self-supervised learning approaches, predictive world models, and modularityâand to judge advances by how quickly and reliably they produce new competence, rather than by how closely they imitate human behavior. 8 AI Must Embrace Specialization via Superhuman Adaptable Intelligence References Ag Ì uera y Arcas, B. and Norvig, P.Artificial general intelligence is already here. Noema Magazine, Octo- ber 2023.URLhttps://w.noemamag.com/ artificial-general-intelligence-is-already-here/. Accessed: Y-M-D. Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., LeCun, Y., and Ballas, N. Self-supervised learning from images with a joint-embedding predictive architecture. In 2023 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), p. 15619â 15629, 2023. doi: 10.1109/CVPR52729.2023.01499. Balestriero, R., Ibrahim, M., Sobal, V., Morcos, A., Shekhar, S., Goldstein, T., Bordes, F., Bardes, A., Mialon, G., Tian, Y., Schwarzschild, A., Wilson, A. G., Geiping, J., Garrido, Q., Fernandez, P., Bar, A., Pirsiavash, H., LeCun, Y., and Goldblum, M. A cookbook of self-supervised learning, 2023. URLhttps://arxiv.org/abs/ 2304.12210. Bruce, J., Dennis, M., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steiger- wald, R., Apps, C., Aytar, Y., Bechtle, S., Behbahani, F., Chan, S., Heess, N., Gonzalez, L., Osindero, S., Ozair, S., Reed, S., Zhang, J., Zolna, K., Clune, J., de Freitas, N., Singh, S., and Rockt Ì aschel, T.Ge- nie: Generative interactive environments, 2024. URL https://arxiv.org/abs/2402.15391. Bylander, T. The computational complexity of propo- sitional strips planning.Artificial Intelligence, 69(1):165â204, 1994.ISSN 0004-3702.doi: https://doi.org/10.1016/0004-3702(94)90081-7. URLhttps://w.sciencedirect.com/ science/article/pii/0004370294900817. Chen, E. K., Belkin, M., Bergen, L., and Danks, D. Does ai already have human-level intelligence? the evidence is clear.Nature, 650:36â40, 2026.doi: 10.1038/ d41586-026-00285-6. URL https://w.nature. com/articles/d41586-026-00285-6.Ac- cessed: Y-M-D. Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual rep- resentations. In International conference on machine learning, p. 1597â1607. PmLR, 2020. Chollet, F. On the measure of intelligence, 2019. URL https://arxiv.org/abs/1911.01547. Craik, K. J. W. The nature of explanation, volume 445. CUP Archive, 1967. Domingos, P. A few useful things to know about ma- chine learning. Commun. ACM, 55(10):78â87, Octo- ber 2012. ISSN 0001-0782. doi: 10.1145/2347736. 2347755.URLhttps://doi.org/10.1145/ 2347736.2347755. Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: scaling to trillion parameter models with simple and effi- cient sparsity. J. Mach. Learn. Res., 23(1), January 2022. ISSN 1532-4435. Finn, C., Abbeel, P., and Levine, S. Model-agnostic meta- learning for fast adaptation of deep networks. In Interna- tional conference on machine learning, p. 1126â1135. PMLR, 2017. Forister, M. L., Dyer, L. A., Singer, M. S., Stire- man I, J. O., and Lill, J. T.Revisiting the evolution of ecological specialization, with empha- sis on insectâplant interactions. Ecology, 93(5):981â 991, 2012.doi: https://doi.org/10.1890/11-0650.1. URLhttps://esajournals.onlinelibrary. wiley.com/doi/abs/10.1890/11-0650.1. Futuyma, D. J. and Moreno, G.The evolution of ecological specialization.Annual Review of Ecol- ogy, Evolution, and Systematics, 19(Volume 19, 1988):207â233, 1988.ISSN 1545-2069.doi: https://doi.org/10.1146/annurev.es.19.110188.001231. URLhttps://w.annualreviews.org/ content/journals/10.1146/annurev.es. 19.110188.001231. Grill, J.-B., Strub, F., Altch Ì e, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems, 33:21271â21284, 2020. Hafner, D., Yan, W., and Lillicrap, T. Training agents inside of scalable world models, 2025. URLhttps: //arxiv.org/abs/2509.24527. Hannan, M. T. and Freeman, J. The population ecology of organizations. American Journal of Sociology, 82(5): 929â964, 1977. doi: 10.1086/226424. URLhttps: //doi.org/10.1086/226424. Hassabis, D.Yann is just plain incorrect here, heâs confusing general intelligence with universal in- telligence.X (formerly Twitter) post, December 2025. URLhttps://x.com/demishassabis/ status/2003097405026193809 .Posted by @demishassabis. 9 AI Must Embrace Specialization via Superhuman Adaptable Intelligence He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Mo- mentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9729â9738, 2020. He, K., Chen, X., Xie, S., Li, Y., Doll Ì ar, P., and Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 16000â16009, 2022. Hendrycks, D., Song, D., Szegedy, C., Lee, H., Gal, Y., Brynjolfsson, E., Li, S., Zou, A., Levine, L., Han, B., Fu, J., Liu, Z., Shin, J., Lee, K., Mazeika, M., Phan, L., Ingebretsen, G., Khoja, A., Xie, C., Salaudeen, O., Hein, M., Zhao, K., Pan, A., Duvenaud, D., Li, B., Omohundro, S., Alfour, G., Tegmark, M., McGrew, K., Marcus, G., Tallinn, J., Schmidt, E., and Bengio, Y. A definition of agi, 2025. URLhttps://arxiv.org/abs/2510. 18212. Huang, D., Yan, C., Li, Q., and Peng, X. From large language models to large multimodal models: A liter- ature review. Applied Sciences, 14(12), 2024. ISSN 2076-3417. doi: 10.3390/app14125068. URLhttps: //w.mdpi.com/2076-3417/14/12/5068. Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Ë Z Ì Ä±dek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P., and Hassabis, D. Highly accurate pro- tein structure prediction with alphafold. Nature, 596 (7873):583â589, Aug 2021. ISSN 1476-4687. doi: 10.1038/s41586-021-03819-2. URLhttps://doi. org/10.1038/s41586-021-03819-2. LeCun, Y. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. Open Review, 62(1):1â62, 2022. LeCun, Y. Objective-driven ai: Towards ai systems that can learn, remember, reason, and plan, 2024. URL https://cmsa.fas.harvard.edu/media/ lecun-20240328-harvard_reduced.pdf . Harvard CMSA Ding Shum Lecture; slide includes P(correct) = (1âe) n and âdiverges exponentiallyâ. LeCun, Y., Bengio, Y., and Hinton, G. Deep learning. Na- ture, 521(7553):436â444, May 2015. ISSN 1476-4687. doi: 10.1038/nature14539. URLhttps://doi.org/ 10.1038/nature14539. Legg, S. and Hutter, M. Universal intelligence: A definition of machine intelligence. Minds and machines, 17(4): 391â444, 2007. Levesque, H. J., Davis, E., and Morgenstern, L. The wino- grad schema challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning, KRâ12, p. 552â561. AAAI Press, 2012. ISBN 9781577355601. Li, N. P., van Vugt, M., and Colarelli, S. M. The evo- lutionary mismatch hypothesis: Implications for psy- chological science.Current Directions in Psycho- logical Science, 27(1):38â44, 2018.doi: 10.1177/ 0963721417731378. URLhttps://doi.org/10. 1177/0963721417731378. Lin, C.-C., Jaech, A., Li, X., Gormley, M. R., and Eisner, J. Limitations of autoregressive models and their alter- natives. In Proceedings of the 2021 conference of the North American chapter of the association for compu- tational linguistics: Human language technologies, p. 5147â5173, 2021. Littman, M. L., Goldsmith, J., and Mundhenk, M. The Computational Complexity of Probabilistic Plan- ning. Journal of Artificial Intelligence Research, 9:1â 36, August 1998. ISSN 1076-9757. doi: 10.1613/ jair.505. URLhttps://jair.org/index.php/ jair/article/view/10208. Loasby, B. J. An evolutionary theory of economic change. The Economic Journal, 93(371):652â654, 09 1983. ISSN 0013-0133. doi: 10.2307/2232409. URLhttps:// doi.org/10.2307/2232409. Mitchell, M.Debates on the nature of artificial gen- eral intelligence.Science, 383(6689):eado7069, 2024.doi:10.1126/science.ado7069.URL https://w.science.org/doi/abs/10. 1126/science.ado7069. Morris, M. R., Sohl-Dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., and Legg, S. Position: levels of agi for operationalizing progress on the path to agi. In Proceedings of the 41st International Conference on Machine Learning, ICMLâ24. JMLR.org, 2024. Musk, E. Demis is right. X (formerly Twitter) post, De- cember 2025. URLhttps://x.com/elonmusk/ status/2003165966243598738 .Posted by @elonmusk. Narayanan, A. and Kapoor, S. Ai as normal technology. Knight First Amendment Institute, 2025. 10 AI Must Embrace Specialization via Superhuman Adaptable Intelligence Nilsson, N. J. Human-level artificial intelligence? be serious!AI Magazine, 26(4):68â75, 2005.doi: https://doi.org/10.1609/aimag.v26i4.1850.URL https://onlinelibrary.wiley.com/doi/ abs/10.1609/aimag.v26i4.1850. OpenAI. OpenAI Charter. URLhttps://openai. com/charter/. Ruder, S. An overview of multi-task learning in deep neural networks, 2017. URLhttps://arxiv.org/abs/ 1706.05098. Russell, S. J. and Norvig, P. Artificial Intelligence: A Mod- ern Approach, volume 3. Prentice Hall, Upper Saddle River, NJ, 2010. ISBN 978-0136042594. Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D.Mastering atari, go, chess and shogi by plan- ning with a learned model. Nature, 588(7839):604â 609, Dec 2020.ISSN 1476-4687.doi: 10.1038/ s41586-020-03051-4. URLhttps://doi.org/10. 1038/s41586-020-03051-4. Searle, J. R. Minds, brains, and programs. Behavioral and Brain Sciences, 3(3):417â424, 1980. doi: 10.1017/ S0140525X00005756. Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Has- sabis, D. A general reinforcement learning algorithm that masters chess, shogi, and go through self-play. Sci- ence, 362(6419):1140â1144, 2018. doi: 10.1126/science. aar6404. URLhttps://w.science.org/doi/ abs/10.1126/science.aar6404. Su, W. Do large language models (really) need statisti- cal foundations?, 2025. URLhttps://arxiv.org/ abs/2505.19145. Sutton, R. The bitter lesson. Incomplete Ideas (blog), 13(1): 38, 2019. Turing, A. M. I.âcomputing machinery and intelligence. Mind, LIX(236):433â460, 10 1950. ISSN 0026-4423. doi: 10.1093/mind/LIX.236.433. URLhttps://doi. org/10.1093/mind/LIX.236.433. Tversky, A. and Kahneman, D.Judgment under un- certainty: Heuristics and biases. Science, 185(4157): 1124â1131, 1974.doi: 10.1126/science.185.4157. 1124.URLhttps://w.science.org/doi/ abs/10.1126/science.185.4157.1124. Van Assel, H., Ibrahim, M., Biancalani, T., Regev, A., and Balestriero, R. Joint embedding vs reconstruction: Prov- able benefits of latent space prediction for self supervised learning. arXiv preprint arXiv:2505.12477, 2025. Wang, P. On defining artificial intelligence. Journal of Artificial General Intelligence, 10:1â37, 08 2019. doi: 10.2478/jagi-2019-0002. Wolpert, D. and Macready, W. No free lunch theorems for optimization. IEEE Transactions on Evolutionary Com- putation, 1(1):67â82, 1997. doi: 10.1109/4235.585893. Wozniak, S. Could a computer make a cup of coffee? Fast Company Live, March 2010. URLhttps://w. fastcompany.com . Interview proposing a physical benchmark for AGI. Wyder, P. M., Goldfeder, J., Yermakov, A., Zhao, Y., Riva, S., Williams, J. P., Zoro, D., Rude, A. S., Tomasetto, M., Germany, J., et al. Common task framework for a criti- cal evaluation of scientific machine learning algorithms. arXiv preprint arXiv:2510.23166, 2025. Zhao, A., Wu, Y., Yue, Y., Wu, T., Xu, Q., Yue, Y., Lin, M., Wang, S., Wu, Q., Zheng, Z., and Huang, G. Absolute zero: Reinforced self-play reasoning with zero data, 2025. URL https://arxiv.org/abs/2505.03335. 11