Paper deep dive
The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era
Rudra Jadhav, Janhavi Danve
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/10/2026, 3:45:39 AM
Summary
The paper introduces the Skill Automation Feasibility Index (SAFI) to benchmark four frontier LLMs (LLaMA 3.3 70B, Mistral Large, Qwen 2.5 72B, and Gemini 2.5 Flash) across 35 O*NET skills. It identifies a 'capability-demand inversion' where skills most demanded in AI-exposed jobs are those LLMs perform least well at, suggesting that current AI adoption is primarily augmentation-driven rather than automation-driven. The authors propose an AI Impact Matrix to categorize skills by displacement risk and augmentation potential.
Entities (7)
Relation Signals (3)
SAFI â benchmarks â LLM
confidence 100% ¡ We present the Skill Automation Feasibility Index (SAFI), benchmarking four frontier LLMs
AI Impact Matrix â uses â SAFI
confidence 95% ¡ we propose an AI Impact Matrixâan interpretive framework that positions skills along four quadrants
LLM â performson â O*NET
confidence 90% ¡ benchmarking four frontier LLMs... across 35 skills in the U.S. Department of Labor's O*NET taxonomy
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As Large Language Models reshape the global labor market, policymakers and workers need empirical data on which occupational skills may be most susceptible to automation. We present the Skill Automation Feasibility Index (SAFI), benchmarking four frontier LLMs -- LLaMA 3.3 70B, Mistral Large, Qwen 2.5 72B, and Gemini 2.5 Flash -- across 263 text-based tasks spanning all 35 skills in the U.S. Department of Labor's O*NET taxonomy (1,052 total model calls, 0% failure rate). Cross-referencing with real-world AI adoption data from the Anthropic Economic Index (756 occupations, 17,998 tasks), we propose an AI Impact Matrix -- an interpretive framework that positions skills along four quadrants: High Displacement Risk, Upskilling Required, AI-Augmented, and Lower Displacement Risk. Key findings: (1) Mathematics (SAFI: 73.2) and Programming (71.8) receive the highest automation feasibility scores; Active Listening (42.2) and Reading Comprehension (45.5) receive the lowest; (2) a "capability-demand inversion" where skills most demanded in AI-exposed jobs are those LLMs perform least well at in our benchmark; (3) 78.7% of observed AI interactions are augmentation, not automation; (4) all four models converge to similar skill profiles (3.6-point spread), suggesting that text-based automation feasibility may be more skill-dependent than model-dependent. SAFI measures LLM performance on text-based representations of skills, not full occupational execution. All data, code, and model responses are open-sourced.
Tags
Links
- Source: https://arxiv.org/abs/2604.06906v1
- Canonical: https://arxiv.org/abs/2604.06906v1
Trouble viewing inline? Open PDF directly â
Full Text
41,934 characters extracted from source content.
Expand or collapse full text
The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era Rudra Jadhav Department of Computer Science Savitribai Phule Pune University Pune, India roodra.jadhav@gmail.com Janhavi Danve Department of Computer Science Savitribai Phule Pune University Pune, India janhavi.danve@gmail.com April 2026 | Preprint â Under Review Abstract. As Large Language Models reshape the global labor market, policymakers and workers need empirical data on which occupational skills may be most suscep- tible to automation. We present the Skill Automa- tion Feasibility Index (SAFI), benchmarking four frontier LLMsâLLaMA 3.3 70B, Mistral Large, Qwen 2.5 72B, and Gemini 2.5 Flashâacross 263 text-based tasks spanning all 35 skills in the U.S. Department of Laborâs O*NET taxonomy (1,052 total model calls, 0% failure rate). Cross-referencing with real-world AI adoption data from the Anthropic Economic Index (756 occupations, 17,998 tasks), we propose an AI Impact Matrixâan interpretive framework that positions skills along four quadrants: High Displacement Risk, Upskilling Required, AI-Augmented, and Lower Displacement Risk. Key findings: (1) Mathematics (SAFI: 73.2) and Pro- gramming (71.8) receive the highest automation feasi- bility scores; Active Listening (42.2) and Reading Com- prehension (45.5) receive the lowest; (2) a âcapability- demand inversionâ where skills most demanded in AI- exposed jobs are those LLMs perform least well at in our benchmark; (3) 78.7% of observed AI interactions are augmentation, not automation; (4) all four mod- els converge to similar skill profiles (3.6-point spread), suggesting that text-based automation feasibility may be more skill-dependent than model-dependent. SAFI mea- sures LLM performance on text-based representations of skills, not full occupational execution. All data, code, and model responses are open-sourced. Keywords: AI labor markets ¡ skill automation ¡ LLM benchmarking ¡ O*NET ¡ workforce transition ¡ SAFI 1 Introduction The rapid advancement of Large Language Models has shifted the conversation about AIâs economic impact from speculative forecasting to observable reality. In Febru- ary 2026, JPMorgan Chase CEO Jamie Dimon confirmed that his bank has already experienced AI-driven work- force displacement, stating: âWe have displaced people from AI, and we offer them other jobsâ [6]. Speaking at the Hill & Valley Forum in March 2026, he warned that AI-driven disruption âmay be quickerâ than past techno- logical transitions and called for coordinated government- business efforts to âretrain, reskill, and redeployâ affected workers [7]. Goldman Sachs CEO David Solomon of- fered a contrasting perspective, declaring âIâm not in the job apocalypse campâ while acknowledging that the pace of AI adoption means âthe short-term disruption might be a little bit higherâ than prior technology shifts [8]. Anthropic CEO Dario Amodei issued perhaps the stark- est warning, predicting AI could eliminate up to 50% of entry-level white-collar jobs within five years [10]. These industry perspectives are supported by empir- ical data. Anthropicâs Economic Index, based on anal- ysis of over four million Claude conversations, found that AI usage primarily concentrates in software devel- opment and writing tasks, with approximately 36% of occupations using AI for at least a quarter of their as- sociated tasks [12]. The âAgents of Chaosâ study [16] demonstrated through live deployment that autonomous AI agents exhibit both significant capabilities and criti- cal failure modesâunderscoring the urgency of system- atic skill-level evaluation. Yet a critical gap persists: while we know AI is be- ing adopted, relatively few studies offer empirical data on which specific skills LLMs can perform in text-based settings and how this capability relates to the labor mar- ketâs skill structure. Existing studies rely on expert opin- ion [11], employer surveys [17], or theoretical exposure indices [15]âfew directly benchmark LLM capabilities against the standardized skill taxonomy that underpins arXiv:2604.06906v1 [cs.CL] 8 Apr 2026 workforce planning. This paper addresses that gap with three contributions: 1. The Skill Automation Feasibility Index (SAFI): An empirically-derived score (0â100) for each of the 35 O*NET skills, based on benchmarking four frontier LLMs across 263 purpose-designed tasks with 1,052 total responses. 2. The AI Impact Matrix: An interpretive framework that cross-references SAFI scores with real-world AI adoption data from the Anthropic Economic Index, positioning skills along four quadrants to inform work- force planning discussions. 3. The Capability-Demand Inversion: Evidence from our benchmark that skills most demanded in AI- exposed occupations are those LLMs score lowest on in text-based evaluationâconsistent with the inter- pretation that the current AI wave is augmentation- dominant rather than automation-dominant. 2 Related Work 2.1 AI and Labor Market Impact The study of technologyâs impact on employment has a rich history [4, 1]. [11] estimated that approximately 80% of the U.S. workforce could have at least 10% of their tasks affected by GPTs, based on expert annotations of O*NET tasks. However, their methodology relied on sub- jective human ratings rather than empirical capability testing. The Anthropic Economic Index [12, 2] repre- sents the most comprehensive empirical study of actual AI usage patterns, analyzing millions of conversations mapped to O*NET occupations and tasks. Their findings established that augmentation dominates automation by roughly 4:1 in current usage patterns. More recently, An- thropicâs economic primitives framework [3] introduced measures of task complexity, skill level, and AI auton- omy, finding that more complex tasks are actually sped up more by AIâa finding our results complement at the skill level. 2.2 LLM Evaluation and Benchmarking Standard LLM benchmarks (MMLU, HumanEval, MATH) measure academic performance but do not map onto workforce skills [13, 5]. The âAgents of Chaosâ study [16] deployed six autonomous agents with persistent mem- ory, email, and shell access, documenting eleven failure cases alongside six genuine safety behaviorsâreinforcing the need for skill-level evaluation in real-world contexts. Additionally, recent work on LLM evaluation biases has shown that model outputs are influenced by surface-level features such as writing style and linguistic register [14], a finding that motivated our choice of heuristic scoring over LLM-as-judge approaches in this study. 2.3 Workforce Transition Frameworks The World Economic Forum [17] projects skill demand based on employer surveys across 55 economies. Pew Re- search [15] classified occupations by AI exposure using O*NET work activities. Our work differs from both by: (a) operating at the skill level rather than occupation or activity level; (b) using empirical LLM benchmarks rather than surveys; and (c) cross-referencing with actual AI adoption data from production systems. 3 Data and Methods 3.1 Data Sources O*NET Database (v30.2). The U.S. Department of Laborâs Occupational Information Network provides im- portance (IM) and level (LV) ratings for 35 skills across 1,016 occupations. Skills are organized into seven cate- gories: Content (6 skills), Process (4), Social (6), Com- plex Problem Solving (1), Technical (11), Systems (3), and Resource Management (4). Anthropic Economic Index. We utilize data from five releases (February 2025 through March 2026): job exposure scores for 756 occupations (range: 0.0 to 0.745), task penetration rates for 17,998 O*NET tasks, and au- tomation versus augmentation interaction patterns for 3,364 tasks with five interaction modes (directive, feed- back loop, task iteration, validation, learning). LLM Benchmark Data (ours). 263 tasks across all 35 O*NET skills administered to four frontier LLMs, yielding 1,052 responses with a 100% completion rate. 3.2 Skill Taxonomy Figure 1 shows the average importance of each skill cate- gory across all O*NET occupations. Process Skills (3.18) and Content Skills (3.05) rank highest, reflecting their universal importance across the labor market. Technical Skills (1.90) rank lowest because the category includes highly specialized skills (Installation: 1.24, Equipment Maintenance: 1.71) that apply to a narrow set of occupa- tions. 3.3 AI Exposure Landscape The distribution of AI exposure across 756 occupations is heavily right-skewed (Figure 2), with a mean exposure of 0.077 and median of 0.000. Only 67 occupations (8.9%) have exposure scores above 0.30. The top five most AI- exposed occupations are Computer Programmers (0.745), Customer Service Representatives (0.701), Data Entry Keyers (0.671), Medical Records Specialists (0.667), and Market Research Analysts (0.648). 3.4 LLM Benchmark Design 3.4.1 Task Battery We designed 263 tasks covering all 35 O*NET skills at three difficulty levels (easy, medium, hard). Tasks are 2 Figure 1: Average skill importance by category across all 1,016 O*NET occupations. Process and Content skills are universally important; Technical skills are more specialized. Figure 2: Distribution of AI exposure across 756 occupations from the Anthropic Economic Index. Most occu- pations have minimal AI exposure; a long tail of 67 occupations exceeds 0.30. text-based prompts designed to elicit the cognitive and communicative dimensions of each skill as defined by O*NET; they do not capture physical, embodied, or real- time interactive aspects of skill execution. For example, Reading Comprehension tasks require identifying contra- dictions between passages and evaluating methodological limitations; Negotiation tasks present multi-party scenar- ios requiring strategic thinking; Programming tasks re- quire writing functional code and debugging errors. Each task includes a standardized grading rubric. Figure 3: The 25 most AI-exposed occupations. Computing, data, and customer-facing roles dominate. 3.4.2 Model Selection Four LLMs were selected for geographic, architectural, and licensing diversity: ⢠LLaMA 3.3 70B (Meta, US, open-source) via Groq ⢠Mistral Large (Mistral AI, France, closed-source) ⢠Qwen 2.5 72B (Alibaba, China, open-source) via HuggingFace ⢠Gemini 2.5 Flash (Google, US, closed-source) All models received identical prompts with temperature 0.3. The pipeline achieved 1,052 successful responses with zero failures, executed over approximately 172 minutes with built-in resume capability and rate limiting. 3.4.3 Scoring Methodology To avoid LLM-as-judge bias (where a model evaluates its own output favorably), responses were scored using a multi-signal heuristic engine across four dimensions: Re- sponse Completeness (0â3), Response Depth (0â3), Rea- soning Quality (0â2), and Difficulty-Adjusted Bonus (0â 2). Skill-specific adjustments add credit for mathematical calculations in Mathematics tasks and functional code in Programming tasks. Total scores range 0â10. 3.4.4 SAFI Computation The Skill Automation Feasibility Index is computed as the average normalized score across all models and tasks for a given skill: SAFI(s) = 100 |M| X mâM 1 |T s | X tâT s score(t,m) 10 (1) where s is a skill, M is the set of 4 models, and T s is the set of tasks for skill s. SAFI ranges from 0 (no measurable 3 text-based performance) to 100 (perfect performance on all text-based tasks). Importantly, SAFI reflects LLM performance on textual representations of skillsâit does not directly measure the feasibility of automating the full occupational contexts in which these skills are applied. 4 Results 4.1 SAFI Scores Across 35 Skills Figure 4 presents the complete SAFI ranking. Mathemat- ics (73.2) and Programming (71.8) receive substantially higher scores than all other skills, with a clear gap be- tween the top two and the remaining 33 skills (range: 42.2â61.7). Figure 4: SAFI ranking across all 35 O*NET skills. Mathe- matics and Programming receive the highest text- based automation feasibility scores. Active Lis- tening and Reading Comprehension score lowest. Color indicates skill category; zone annotations in- dicate relative automation feasibility. The bottom five skillsâActive Listening (42.2), Read- ing Comprehension (45.5), Speaking (48.5), Writing (51.0), and Social Perceptiveness (51.5)âare all Con- tent or Social skills. These are precisely the skills rated most important across the broadest range of occupations in O*NET, creating a pattern we term the âcapability- demand inversion.â Notably, Content Skills has the highest within-category variance (Ď = 1.93), driven by the extreme range be- Table 1: SAFI scores by skill category. Technical Skills score highest; Content Skills score lowest in text-based evaluation. CategorySAFI Ď Skills Technical Skills62.1 0.9911 Systems Skills60.0 0.673 Complex Problem Solving 59.2 0.531 Resource Management59.2 0.544 Process Skills57.9 1.054 Social Skills57.3 0.956 Content Skills53.1 1.936 tween Mathematics (73.2) and Active Listening (42.2)â both classified as Content Skills in O*NET. This sug- gests that the Content category contains fundamentally different skill types: structured quantitative reasoning (which LLMs excel at) and nuanced human communi- cation (which LLMs struggle with). 4.2 SAFI Heatmap: SkillĂ Model Figure 5 presents the complete SAFI matrix across all 35 skills and 4 models. Several patterns emerge: (1) Math- ematics shows the highest cross-model variance, with LLaMA scoring 82 and Gemini scoring 58; (2) the mid- dle band of skills (SAFI 58â62) shows remarkable unifor- mity across models; (3) Active Listening shows the widest performance gap among low-scoring skills (LLaMA: 45, Gemini: 35). 4.3 Model Comparison All four models exhibit remarkably similar performance profiles (Figure 6). Mistral Large achieved the highest overall SAFI (60.0), followed by LLaMA 3.3 70B (58.2), Qwen 2.5 72B (56.7), and Gemini 2.5 Flash (56.4). The narrow 3.6-point spread suggests that text-based au- tomation feasibility, as measured here, may be more skill-dependent than model-dependent. The most notable inter-model variation occurs in Con- tent Skills, where Mistral Large (57) outperforms Gemini 2.5 Flash (48) by 9 pointsâthe largest gap in any cate- gory. This suggests model architecture and training data influence performance on communication-intensive tasks more than on structured reasoning. 4.4 Difficulty Scaling All models show increasing SAFI from easy to hard tasks (Figure 7). Averaged across models, easy tasks score 53.2, medium tasks 59.0, and hard tasks 61.0. At the per-model level, the pattern is consistent: Mistral Large scores 56.5 / 61.2 / 61.9 across easy / medium / hard; LLaMA 3.3 70B scores 54.5 / 58.5 / 61.4; Qwen 2.5 72B scores 50.9 / 58.4 / 60.4; and Gemini 2.5 Flash scores 50.8 / 57.8 / 60.2. The gap between easy and hard is largest for Qwen (9.5 points) and smallest for Mistral (5.4 points), suggesting Mistral maintains more consistent performance regardless 4 Figure 5: Complete SAFI heatmap: 35 skills Ă 4 models. Color strip on right indicates skill category. Note the gradient from green (lower SAFI, top) to red (higher SAFI, bottom). Figure 6: SAFI scores by model across skill categories. All four models show similar profiles, with Technical Skills consistently highest and Content Skills low- est. of task complexity. This counterintuitive resultâbetter scores on harder tasksâis a known artifact of our length- sensitive scoring methodology: harder tasks elicit longer, more structured responses that earn higher completeness and reasoning scores. We report this transparently as a limitation of the heuristic scoring approach rather than a Figure 7: SAFI scores by difficulty level and model. All models show higher scores on harder tasksâa scoring artifact of longer, more structured re- sponses rather than genuinely better problem- solving. Mistral Large leads at every difficulty level. finding about model capability. 4.5 Skill-Exposure Correlations Cross-referencing O*NET skill importance ratings with Anthropic Economic Index exposure data across 744 matched occupations yields Pearson correlations for each skill (Figure 8). Figure 8: Correlation between skill importance and real- world AI exposure across 744 occupations. Pro- gramming (+0.455) is most concentrated in AI- exposed jobs; Operation and Control (â0.456) is most concentrated in occupations with low AI ex- posure. The split between cognitive and physical skills is notable. Programming shows the strongest positive correlation (+0.455), meaning it is most important in occupations 5 with high AI exposure. Content skills (Reading Com- prehension: +0.372, Writing: +0.369, Active Listening: +0.338) are also concentrated in AI-exposed occupations. In contrast, physical technical skills (Operation and Con- trol: â0.456, Operations Monitoring: â0.404, Trou- bleshooting: â0.397) concentrate in occupations with low AI exposure. 4.6 SAFI vs. Real-World Exposure Figure 9 plots SAFI scores against real-world AI exposure correlations, revealing the capability-demand inversion. The Pearson correlation between SAFI and exposure cor- relation is r = â0.196 (p = 0.26, n = 35); the Spearman rank correlation is Ď =â0.300 (p = 0.08). While neither reaches conventional statistical significanceâlikely due to the small sample of 35 skillsâthe consistent negative di- rection across both tests supports the interpretation that skills more important in AI-exposed occupations tend to receive lower SAFI scores. Figure 9: SAFI score vs. real-world AI exposure correla- tion (r = â0.196, p = 0.26, n = 35). The neg- ative trend, while not statistically significant at Îą = 0.05, is consistent with a capability-demand inversion: skills most important in AI-exposed jobs tend to receive lower SAFI scores. This pattern is consistent with the interpretation that occupations currently adopting AI most heavily may not be the ones most susceptible to skill-level automation. Rather, AI appears to be used predominantly as a collab- orative tool in roles requiring communication skills where LLMs show lower benchmark performance. 4.7 Automation vs. Augmentation Analysis of 3,364 task-level interaction patterns from the Anthropic Economic Index confirms this interpretation (Figure 10). Of all observed AI-task interactions, 78.7% represent augmentation (collaborative interaction) versus 21.3% automation (directive task completion). The dom- inant augmentation modes are feedback loops (32.5%), where humans iteratively refine AI outputs, and learn- ing interactions (29.9%), where AI serves an educational role. Pure directive automationâwhere the AI completes a task end-to-end without human iterationâaccounts for only one in five interactions. Figure 10: Automation vs. augmentation patterns from 3,364 task interactions. Nearly four in five AI interactions are collaborative, not replacement- oriented. This breakdown is important because it challenges the common narrative of AI as a âjob killer.â The data sug- gest that, at present, AI is overwhelmingly used to en- hance human work rather than replace itâa finding con- sistent with both Anthropicâs original analysis [12] and our SAFI results showing that the skills most involved in AI-exposed occupations are those LLMs score lowest on. 4.8 Model Skill Profiles Figure 11 presents individual SAFI profiles for each model across seven skill categories. Mistral Large shows the most balanced profile, with consistently above-average performance across all categories. LLaMA 3.3 70B shows particular strength in Technical Skills but scores rela- tively lower on Social Skills. Qwen 2.5 72B and Gem- ini 2.5 Flash show similar overall shapes despite different origins (China vs. US) and architectures, reinforcing the finding of cross-model convergence. The convergence of radar shapes across models from three different countries and two licensing regimes (open- source and closed-source) is notable. It suggests that at the 70B+ parameter scale, frontier LLMs develop a shared capability profile for workforce-relevant text tasksâregardless of the specific training pipeline. For workforce planners, this means that skill-level assess- ments of AI capability are likely to remain relatively sta- ble across the model landscape, at least within the current generation of frontier systems. 5 The AI Impact Matrix Synthesizing SAFI benchmarks with real-world AI adop- tion data, we propose the AI Impact Matrix (Fig- ure 12)âan interpretive framework that positions each skill along two dimensions: text-based automation feasi- bility (SAFI) and real-world AI exposure correlation. The matrix is intended as a heuristic for structuring workforce 6 Figure 11: SAFI radar profiles for each model across seven skill categories. Mistral Large shows the most balanced capability; all four models converge on similar shapes despite diverse origins. planning discussions, not as a direct forecast of displace- ment outcomes, which depend on many factors beyond LLM text performance. Figure 12: The AI Impact Matrix: SAFI score Ă real-world AI exposure correlation for all 35 O*NET skills. Each dot represents one skill; color indicates cat- egory. Programming sits in the High Displace- ment Risk quadrant; Content skills cluster in the AI-Augmented quadrant; physical Technical skills fall in the Upskilling Window. Quadrant I â Higher Displacement Risk (High SAFI + Positive AI Exposure). Programming sits here: LLMs score well on its text-based tasks (SAFI: 71.8) and it concentrates in heavily AI-exposed occupations (+0.455). Workers relying primarily on structured pro- gramming tasks may face elevated near-term displace- ment risk, though current evidence points toward AI- augmented development rather than full replacement. In practical terms, this means junior developers and entry-level software engineersâwhose work often involves implementing well-specified features, writing boilerplate code, and fixing routine bugsâare more exposed than se- nior architects who design systems and make judgment calls about tradeoffs. Universities and coding bootcamps may need to shift curricula from syntax fluency toward system design, AI-assisted development workflows, and the ability to evaluate and debug AI-generated code. In our benchmark, LLMs scored highest on self-contained coding tasks (writing functions, implementing algorithms, generating SQL queries) and lowest on tasks requiring multi-file architectural reasoning or ambiguous specifica- tion interpretationâreinforcing the distinction between automatable routine coding and resilient system-level de- sign work. Quadrant I â AI-Augmented (Low SAFI + Pos- itive AI Exposure). Content skills (Reading, Writing, Speaking, Active Listening) and some Social skills (Per- suasion, Negotiation) cluster here. AI is used heavily in occupations requiring these skills, but as a collaborative toolâLLMs cannot replicate the nuanced human com- munication these skills demand. The implications are concrete: customer support representatives are using AI to draft responses faster, but the empathy, de-escalation, and contextual judgment that define excellent service re- main human. Financial analysts use LLMs to summarize reports and generate first drafts, but the critical read- ing, skeptical evaluation, and client-facing communica- tion that drive investment decisions stay with the ana- lyst. In education, teachers may use AI to generate les- son materials, but active listening to a struggling student and adapting instruction in real time are irreplaceable. For entry-level white-collar workersâadministrative as- sistants, junior analysts, content coordinatorsâthe shift is not from employment to unemployment, but from rou- tine execution to AI-augmented productivity, where the workers who thrive will be those who learn to direct, eval- uate, and refine AI outputs effectively. Quadrant I â Upskilling Window (Moder- ate/High SAFI + Negative AI Exposure). Physical Tech- nical skills (Equipment Maintenance, Troubleshooting, Operation and Control) appear here. LLMs can discuss these skills abstractly (moderate SAFI) but the occu- pations requiring them are not AI-exposed because the skills demand physical presence. As AI-assisted diagnos- tics, predictive maintenance platforms, and remote mon- itoring advance, this quadrant represents a window for proactive upskilling. HVAC technicians, manufacturing operators, and field maintenance workers are currently insulated from AI disruptionâbut as sensor data and AI diagnostic tools enter their workflows, those who can com- bine hands-on expertise with data literacy will be best po- sitioned. Trade schools and vocational programs have an opportunity to integrate AI-assisted troubleshooting into their curricula now, before the transition accelerates. 7 Quadrant IV â Lower Displacement Risk (Low SAFI + Negative/Neutral AI Exposure). Skills requir- ing embodied human judgment in occupations with low current AI exposure fall here, suggesting relatively lower near-term disruption from text-based AI systems. 6 Discussion 6.1 The Capability-Demand Inversion Our most notable finding is that the skills most con- centrated in AI-exposed occupations are not the skills LLMs score highest on in our text-based benchmark. This âcapability-demand inversionâ has important implications for workforce planning. It is consistent with the inter- pretation that the current AI adoption wave is driven by augmentation demandâworkers in communication-heavy roles using AI as a productivity toolârather than by au- tomation capabilityâAI replacing the core skills of those roles. If this pattern holds, the most immediate work- force challenge may not be mass displacement of communication-intensive roles, but rather a gradual shift in what those roles require: from pure execution to AI- augmented execution, demanding new meta-skills like prompt engineering, AI output evaluation, and human-AI workflow design. Consider the paralegal who now uses AI to draft legal summaries but must still catch hallucinated case citations, or the marketing analyst who generates campaign copy with AI but must ensure it resonates with the target audienceâs cultural context. The skill itself is not automated; the workflow around the skill is restruc- tured. 6.2 Model Convergence The narrow 3.6-point SAFI spread across four diverse models (US, France, China; open-source and closed- source) suggests that current frontier LLMs may have converged to a similar capability profile for text-based workforce skill tasks. If confirmed by broader evaluations, this would mean workforce planning need not track indi- vidual model releasesâthe skill-level performance profile appears relatively stable across the current frontier. 6.3 The Content Skills Paradox Content Skills simultaneously contain the skill with the highest SAFI score (Mathematics: 73.2) and the lowest (Active Listening: 42.2). This 31-point within-category spreadâthe largest of any categoryâhighlights a distinc- tion between structured content processing (mathemati- cal reasoning, where LLMs score well) and unstructured content understanding (active listening, which in occupa- tional practice involves emotional intelligence, contextual inference, and nonverbal cue interpretationâdimensions that text-based models are not designed to capture). 6.4 Industry Perspectives and Policy Implications The debate among industry leaders illustrates pre- cisely why empirical skill-level data matters. Dimonâs positionâthat displacement is happening now and soci- ety needs phased responses including âretraining, reloca- tion, and income assistanceâ [6]âimplies a need for gran- ular knowledge of which skills are affected. Solomonâs counterpositionâthat the economy is âincredibly broad and nimbleâ enough to absorb displaced workers [8]âstill requires understanding where the absorption will occur. Our SAFI index and AI Impact Matrix may help struc- ture both perspectives. Solomon offered a vivid illustration of AIâs produc- tivity impact in financial services: at Goldman Sachs, AI can now draft 95% of an IPO prospectus (S-1 filing) in minutesâa task that previously required a six-person team working for two weeks [9]. Yet as Solomon noted, âthe last 5% now matters because the rest is now a com- modity.â This precisely mirrors our Quadrant I find- ing: the structured, text-amenable components of finan- cial analysis (high SAFI) are being automated, while the judgment, client communication, and regulatory interpre- tation (low SAFI) become more valuable. While our framework cannot directly predict displace- ment outcomes, it may help differentiate policy responses: skills in Quadrant I (Higher Displacement Risk) war- rant attention for transition planning; skills in Quadrant I (AI-Augmented) suggest training in AI collaboration tools may be beneficial; skills in Quadrant I (Upskilling Window) may offer time for proactive preparation. For governments, this means workforce retraining programs should not treat âAI exposureâ as monolithicâa data en- try clerk (Quadrant I) needs a fundamentally different in- tervention than a customer service representative (Quad- rant I) or an equipment technician (Quadrant I). For corporations, internal training budgets may be better al- located toward AI collaboration skills for communication- heavy roles than toward wholesale role elimination. And for individual workers, the message is nuanced: the ques- tion is less âwill AI take my job?â and more âhow will AI change what my job requires?â 6.5 Recommendations Based on our findings, we offer the following actionable recommendations, organized by stakeholder. For policymakers and governments: Workforce re- training programs should be differentiated by skill quad- rant, not treated as one-size-fits-all âAI readinessâ initia- tives. Workers in Quadrant I occupations (e.g., data en- try clerks, junior programmers) need funded transition pathways to adjacent rolesâour SAFI scores, combined with O*NET skill-adjacency data, could help identify the shortest reskilling routes. Workers in Quadrant I roles (e.g., analysts, customer support, educators) need sub- sidized training in AI collaboration toolsâprompt engi- neering, output verification, and human-AI workflow de- 8 sign. Workers in Quadrant I occupations (e.g., HVAC technicians, manufacturing operators) have a narrow but real window for proactive upskilling before AI diagnostics reshape their fields. Community colleges and vocational programs should integrate AI-assisted troubleshooting and data literacy modules now, while the window remains open. For corporations: Rather than framing AI adoption as headcount reduction, our data supports a redeployment model consistent with JPMorganâs approach [6]. Internal training investments should prioritize teaching existing employees to work with AIâevaluating outputs, catching errors, maintaining qualityârather than replacing them. The 78.7% augmentation rate suggests that most cur- rent AI usage already follows this pattern; formalizing it through structured training programs is the logical next step. For educational institutions:University curriculaâparticularly in computer science, busi- ness, and communicationâshould evolve to emphasize the skills that sit at the intersection of high labor market demand and low AI capability: critical evaluation of AI-generated content, complex interpersonal communica- tion, ethical judgment under uncertainty, and the ability to design workflows that combine human strengths with AI efficiency. The era of teaching skills in isolation from their AI context is ending. For individual workers: Identify where your pri- mary skills sit on the AI Impact Matrix. If you rely heav- ily on Quadrant I skills (structured programming, data processing), invest in system-level thinking, architecture, and AI-augmented development practices. If your work centers on Quadrant I skills (writing, analysis, communi- cation), learn to use AI as a force multiplierâthe workers who thrive will not be those who resist AI, but those who integrate it most effectively into their craft. 6.6 Limitations Several important limitations should be noted. First, and most fundamentally, SAFI measures LLM per- formance on text-based task representations, not the full occupational execution of skills. Many O*NET skillsâparticularly Social and Technical skillsâ involve physical, embodied, real-time, or interpersonal di- mensions that cannot be captured through text prompts. A high SAFI score for a skill like âOperation and Controlâ reflects the modelâs ability to discuss the skillâs cognitive components, not its ability to physically operate machin- ery. Readers should interpret SAFI as an upper-bound proxy for the text-amenable component of each skill, not as a direct measure of occupational automation feasibil- ity. Second, our heuristic scoring methodology, while avoid- ing LLM-as-judge bias, relies on surface-level response features (length, structural markers, reasoning keywords) that may not fully capture response qualityâparticularly for nuanced Social Skills tasks where expert human eval- uation would be more appropriate. The scoring does not verify factual correctness of responses. Third, 263 tasks across 35 skills, while covering the full O*NET taxonomy, represent a limited sample of the vast space of possible skill applications. Skills with fewer tasks (3 tasks for some Technical skills vs. 10 for Content skills) have less statistical power. Fourth, Anthropic Economic Index data reflects Claude usage specifically and may not generalize to all AI plat- forms or to non-English-speaking labor markets. Fifth, our study is cross-sectional; SAFI scores will change as models improve, and the AI Impact Matrix reflects a snapshot of current capabilities rather than a forecast. Sixth, we do not incorporate BLS employment data, which would allow weighting by workforce size and esti- mating the number of workers affected; this is planned for future work. Finally, the AI Impact Matrix is an interpretive frame- work for organizing findings, not a predictive model. Ac- tual displacement outcomes depend on many factors be- yond LLM text performance, including organizational adoption decisions, regulatory environments, economic conditions, and the pace of complementary technology development. 7 Future Work We identify four priority directions for extending this re- search: (1) Reskilling Transition Pathway Maps. Us- ing SAFI scores combined with O*NETâs skill-adjacency data (which tracks how similar skills are across occupa- tions), we plan to construct shortest-path transition maps that recommend specific career moves for workers in high- displacement-risk occupations. For example: a data en- try clerk (Quadrant I, high displacement risk) shares skill overlap with administrative coordinators and project as- sistants (Quadrant I, AI-augmented)âquantifying these transitions would make retraining programs more tar- geted and efficient. (2) Longitudinal SAFI Tracking. As new models are released (GPT-5, Claude 4, Gemini 3, LLaMA 4), re-running our benchmark would measure how fast the automation frontier is advancing per skill. If Active Lis- teningâs SAFI jumps from 42 to 65 in 18 months, that changes the policy calculus significantly. We plan to es- tablish a semi-annual benchmarking cadence. (3) AI-Emergent Skill Taxonomy. O*NETâs 35 skills were designed before LLMs existed. New skills have emergedâprompt engineering, AI output evaluation, multi-agent orchestration, human-AI workflow designâ that are not captured in the current taxonomy. Cat- aloging and benchmarking these emergent skills would complete the picture of the evolving labor market. (4) BLS Employment Integration. Incorporating Bureau of Labor Statistics employment data would al- 9 low us to weight SAFI scores by the number of workers affected, transforming skill-level insights into workforce- level impact estimates (e.g., âProgramming automation feasibility affects approximately 1.8 million U.S. work- ersâ). 8 Conclusion The prevailing narrative about AI and work is binary: either AI will take your job, or it wonât. Our data suggests a third possibility that is both more nuanced and more urgent. The capability-demand inversion reveals that AI is not advancing uniformly across the skill landscape. It is ex- ceptionally strong at structured, rule-bound reasoningâ mathematics, programming, systems analysisâand mea- surably weak at the unstructured, deeply human skills that the labor market values most: listening, reading, speaking, writing, social perception. These are not pe- ripheral skills. They are the connective tissue of the mod- ern economy, the skills that make organizations function, clients trust, and teams collaborate. This asymmetry means that the most significant eco- nomic impact of AI in the near term is not displacement but restructuring. The job stays. The workflow changes. The human becomes responsible not for the first draft but for the last mileâthe judgment, the nuance, the contex- tual awareness that a language model trained on internet text cannot access. This is augmentation in a precise, measurable sense: 78.7% of real-world AI interactions al- ready follow this pattern. But this finding is not cause for complacency. It is cause for urgency. The capability-demand inversion holds today. It holds for this generation of models, bench- marked at the 70B-parameter frontier in early 2026. There is no guarantee it will hold in 2028. If future models close the gap on communication and social skillsâas mul- timodal architectures, real-time audio models, and em- bodied AI agents suggest they mightâthe augmentation- dominant equilibrium could shift rapidly toward automa- tion. The window for proactive preparation is open. It will not stay open indefinitely. We release our complete task battery, all 1,052 model responses, SAFI scores, and analysis code at https: //github.com/rudrajadhav/ai-skills-shiftânot as a finished answer, but as a foundation. The question of which skills AI can and cannot perform is not static. It re- quires continuous, empirical measurement. We hope this work contributes to a culture of evidence over speculation in a debate where the stakes are measured in livelihoods. Data Availability All datasets used in this study are publicly avail- able: O*NET Database (v30.2) from onetcenter. org, Anthropic Economic Index from huggingface.co/ datasets/Anthropic/EconomicIndex. Our benchmark task battery, model responses, and SAFI scores are re- leased under MIT license. Acknowledgments We thank the developers of the O*NET database, An- thropic for open-sourcing the Economic Index data, and the teams behind LLaMA (Meta), Mistral (Mistral AI), Qwen (Alibaba), and Gemini (Google) for providing API access. Special thanks to the Groq and HuggingFace teams for free inference infrastructure that made this re- search possible without institutional funding. References [1] Acemoglu, D. and Restrepo, P. (2020). Robots and jobs: Evidence from US labor markets. Journal of Political Economy, 128(6):2188â2244. [2] Appel, R., McCrory, P., and Tamkin, A. (2025). An- thropic Economic Index report: Uneven geographic and enterprise AI adoption. arXiv:2511.15080. [3] Appel, R., Massenkoff, M., McCrory, P., et al. (2026). Anthropic Economic Index report: Economic primitives. Anthropic Research. [4] Autor, D. H. (2015). Why are there still so many jobs? The history and future of workplace automa- tion. Journal of Economic Perspectives, 29(3):3â30. [5] Chen, M., Tworek, J., Jun, H., et al. (2021). Evaluating large language models trained on code. arXiv:2107.03374. [6] Dimon, J. (2026). Remarks at JPMorgan Chase Investor Day, February 24â25, 2026. Reported by CNBC (https://w.cnbc.com/2026/02/24/ jpm-ceo-jamie-dimon-ai-reshaping-workforce-redeployment. html)andFortune(https: //fortune.com/2026/02/25/ jamie-dimon-society-prepare-ai-job-displacement/). [7] Dimon, J. (2026). Remarks at the Hill & Valley Forum, Washington, D.C., March 24, 2026. Re- ported by CNBC (https://w.cnbc.com/2026/ 03/24/jamie-dimon-ai-job-loss.html). [8] Solomon, D. (2026). Remarks on Goldman Sachs Exchanges Podcast, January 20, 2026. Reported by Fortune (https://fortune.com/2026/01/23/ no-job-apocalypse-goldman-sachs-ceo-david-solomon-ai-hiring-nightmare/). [9] Solomon, D. (2025). Remarks at Cisco AI Sum- mit, Palo Alto, January 15, 2025. Reported by Fortune (https://fortune.com/2025/01/17/ goldman-sachs-ceo-david-solomon-ai-tasks-ipo-prospectus-s1-filing-sec/). 10 [10] Amodei, D. (2026). On the unpredictabil- ity of AI: Reflections on economic impact. Published January 27, 2026. Reported by CNBC(https://w.cnbc.com/2026/01/27/ dario-amodei-warns-ai-cause-unusually-painful-disruption-jobs. html). [11] Eloundou, T., Manning, S., Mishkin, P., and Rock, D. (2023). GPTs are GPTs: An early look at the la- bor market impact potential of large language mod- els. arXiv:2303.10130. [12] Handa, K., Tamkin, A., et al. (2025). Which eco- nomic tasks are performed with AI? Evidence from millions of Claude conversations. arXiv:2503.04761. [13] Hendrycks, D., Burns, C., Basart, S., et al. (2021). Measuring massive multitask language understand- ing. arXiv:2009.03300. [14] Jadhav, R., Danve, J., and Shaw, S. (2026). Implicit grading bias in large language models. arXiv:2603.18765. [15] Pew Research Center (2023). Which U.S. workers are more exposed to AI on their jobs? https://w. pewresearch.org/social-trends/2023/07/26/ which-u-s-workers-are-more-exposed-to-ai-on-their-jobs/. [16] Shapira, N., Wendler, C., Yen, A., et al. (2026). Agents of Chaos. arXiv:2602.20021. [17] WorldEconomicForum(2025).The Future of Jobs Report 2025. https: //w.weforum.org/publications/ the-future-of-jobs-report-2025/. 11