Paper deep dive
Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation?
Wensu Li, Atin Aboutorabi, Harry Lyu, Kaizhi Qian, Martin Fleming, Brian C. Goehring, Neil Thompson
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/1/2026, 1:30:31 AM
Summary
This paper presents a unified microeconomic framework for evaluating the optimal degree of task automation, moving beyond binary 'automate-or-not' models. By modeling automation as a continuous choice, the authors demonstrate that due to the convex cost structure of AI (driven by scaling laws), partial automationโwhere AI handles part of a task and humans handle the residualโis frequently the cost-minimizing equilibrium rather than a mere transitional phase. The framework is calibrated using O*NET data, expert surveys, and GPT-4o task decompositions, specifically applied to computer vision, finding that approximately 11% of computer-vision-exposed labor compensation is economically attractive to automate.
Entities (5)
Relation Signals (3)
Wensu Li โ authored โ Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation?
confidence 100% ยท Paper authorship list
Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation? โ utilizes โ O*NET
confidence 100% ยท We calibrate the framework with O*NET task data
Scaling Laws โ causes โ Partial Automation
confidence 95% ยท the convexity of scaling-law cost functions makes the jump from partial to full automation disproportionately expensive
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper develops a unified framework for evaluating the optimal degree of task automation. Moving beyond binary automate-or-not assessments, we model automation intensity as a continuous choice in which firms minimize costs by selecting an AI accuracy level, from no automation through partial human-AI collaboration to full automation. On the supply side, we estimate an AI production function via scaling-law experiments linking performance to data, compute, and model size. Because AI systems exhibit predictable but diminishing returns to these inputs, the cost of higher accuracy is convex: good performance may be inexpensive, but near-perfect accuracy is disproportionately costly. Full automation is therefore often not cost-minimizing; partial automation, where firms retain human workers for residual tasks, frequently emerges as the equilibrium. On the demand side, we introduce an entropy-based measure of task complexity that maps model accuracy into a labor substitution ratio, quantifying human labor displacement at each accuracy level. We calibrate the framework with O*NET task data, a survey of 3,778 domain experts, and GPT-4o-derived task decompositions, implementing it in computer vision. Task complexity shapes substitution: low-complexity tasks see high substitution, while high-complexity tasks favor limited partial automation. Scale of deployment is a key determinant: AI-as-a-Service and AI agents spread fixed costs across users, sharply expanding economically viable tasks. At the firm level, cost-effective automation captures approximately 11% of computer-vision-exposed labor compensation; under economy-wide deployment, this share rises sharply. Since other AI systems exhibit similar scaling-law economics, our mechanisms extend beyond computer vision, reinforcing that partial automation is often the economically rational long-run outcome, not merely a transitional phase.
Tags
Links
- Source: https://arxiv.org/abs/2603.29121v1
- Canonical: https://arxiv.org/abs/2603.29121v1
Trouble viewing inline? Open PDF directly โ
Full Text
210,612 characters extracted from source content.
Expand or collapse full text
Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation? Wensu Li ํ Atin Aboutorabi ํ Harry Lyu ํ Kaizhi Qian ํ Martin Fleming ํ Brian C. Goehring ํ Neil Thompson ํโ ํ Massachusetts Institute of Technology ํ ฬ Ecole Polytechnique F ฬ ed ฬ erale de Lausanne ํ IBM Research ํ IBMโs Institute for Business Value wensu, hlyu, marti264, neil t@mit.edu atin.aboutorabi@epfl.ch kqian@ibm.com goehring@us.ibm.com Abstract This paper develops a unified framework for evaluating the optimal degree of task automation. Moving beyond binary automate-or-not assessments, we model automation intensity as a continuous choice in which firms minimize costs by selecting an AI accuracy level, with outcomes ranging from labor-only production (no automation) to partial automation (humanโAI collaboration) to full automation. Our framework has two components. On the supply side, we estimate an AI production function through fine-tuning scaling-law experiments that link model performance to data, training steps, and model size. Because large language models, computer vision systems, and foundation- model-based AI more generally exhibit predictable but diminishing returns to these inputs, the cost of achieving higher accuracy is convex: reaching good performance may be relatively inexpensive, but pushing toward near-perfect accuracy becomes disproportionately costly. Full automation is therefore often not cost-minimizing; instead, partial automation โ an interior solution in which firms choose an intermediate level of AI deployment and retain human workers for the residual workload โ frequently emerges as the cost-minimizing equilibrium. On the demand side, we take an information-theoretic approach, introducing an entropy-based measure of task complexity that maps model accuracy into a labor substitution ratio, quantifying how much human work time AI displaces at each accuracy level and formalizing the division of work between AI and human collaborators. We calibrate the framework with O*NET task data, a survey of 3,778 domain experts, and GPT-4o- derived task decompositions, and implement it empirically in computer vision โ a domain where abundant scaling-law data exists. We find that task complexity shapes labor substitution: tasks with few subtasks and low complexity see high substitution rates, while tasks with many subtasks and high complexity favor limited partial automation. Scale of deployment is a fundamental determinant of the automation frontier: AI-as-a-Service and AI agents spread fixed development costs across many users, sharply expanding the set of economically viable tasks. At the firm level, cost-effective automation captures approximately 11% of computer-vision-exposed labor compensation; under economy-wide deployment, the economically viable share rises sharply. Since large language models and other AI systems exhibit similar scaling-law economics, the mech- anisms we identify extend beyond computer vision. Our findings reinforce that partial automation โ where AI assists rather than replaces human judgment โ is often the economically rational long-run outcome, not merely a transitional phase on the path to full automation. Keywords: Artificial Intelligence, Computer Vision, Labor Substitution, Scaling Laws, Partial Automation, HumanโAI Collaboration. โ Corresponding author: neilt@mit.edu arXiv:2603.29121v1 [econ.GN] 31 Mar 2026 Human-AI Collaboration: Partial vs. Full Automation 1 Introduction The central strategic question facing firms in the AI era is no longer whether tasks can be automated, but whether they should be automated, and, critically, to what extent. Emerging AI systems have fueled expectations of large productivity gains while simultaneously raising concerns about job loss and the pace of labor market adjustment. Much of the early economics of automation framed this problem as a binary choice: a task is either automated or not, and exposure is measured by whether a technology is technically capable of performing the required task. Yet for most real-world tasks, automation is better understood as a continuum. 2 Firms can adopt AI systems that fully replace human effort, or they can deploy AI as a tool that handles part of the work while humans perform the remainder. Understanding when partial automation is the optimal solution โ rather than full automation or no automation โ is central to predicting how AI will reshape production and employment. Why does partial automation deserve this central role? The answer lies in the cost structure of modern AI systems. Large language models, computer vision systems, and foundation-model-based AI more generally exhibit scaling laws: model performance improves predictably as training data, model size, and compute are scaled up, but with sharply diminishing returns at higher accuracy levels. As a result, the marginal cost of improving AI performance rises steeply. Achieving โgoodโ accuracy on a task may be relatively inexpensive, but pushing from good to near-perfect accuracy โ the level often required for full automation โ can be orders of magnitude more costly. When the marginal cost of further accuracy improvement exceeds the marginal labor saving it enables, firms optimally stop short of full automation. The AI system handles the portion of the task it can perform cost-effectively, and human workers resolve the remaining uncertainty. Partial automation is therefore not merely a transitional state on the path to full automation; it is frequently the long-run cost-minimizing equilibrium. This cost structure implies three possible outcomes for any given task: no automation, partial automation, and full automation. Which outcome is optimal depends on whether the labor savings from improved AI performance justify the sharply rising cost of achieving that performance. Our framework formalizes these three cases and shows that partial automation is often the most prevalent outcome, precisely because the jump from partial to full automation is disproportionately expensive for many tasks. Figure 1 illustrates this logic schematically. The horizontal axis traces the automation rate from zero to full automation. When the convex AI cost function lies below the labor-saving benefit throughout the relevant performance range, full automation is both feasible and optimal. When the steeply rising marginal cost curve intersects the marginal benefit curve before the required accuracy level is reached, the firm optimally stops at an interior solution โ partial automation โ and human workers complete the residual workload. When fixed development costs alone exceed the potential labor savings, no automation is warranted. The key insight is that the middle region โ partial automation โ occupies the largest share of the task space, precisely because the convexity of scaling-law cost functions makes the jump from partial to full automation disproportionately expensive for most tasks. Firms increasingly begin with foundation models or off-the-shelf AI systems that already deliver some baseline level of performance. For some tasks, that baseline is sufficient. For many others, achieving the accuracy required for reliable deployment requires additional investment in task-specific data, training compute, or larger models. The economic problem is therefore not simply whether AI can perform a task, but whether it is worth paying to move far enough up the performance frontier. In contemporary task models (Autor et al. 2003, Acemoglu and Autor 2011, Acemoglu and Restrepo 2018b,a), automation results in the reallocation of tasks from labor to capital. Over the past decade, with the advent of large language models and deep learning models, the literature has moved beyond labor-capital substitution. Automation research has focused on exposure to tasks potentially automated through these technologies or robotics (Frey and Osborne 2017, Brynjolfsson and Mitchell 2017, Eloundou et al. 2024). A growing literature now examines the conditions under which technically feasible automation is economically viable. One strand emphasizes the composition 2 In contrast to the economics literature, the computer science literature has long recognized automation as a continuum of human-machine interaction. See Sheridan and Verplank (1978); Parasuraman et al. (2000), who propose a multi-dimensional continuum; and Endsley and Kaber (1999), who propose a taxonomy of 10 levels of automation for dynamic control tasks. 2 Human-AI Collaboration: Partial vs. Full Automation Automation rate: 0% FullautomationisoptimalPartialautomationisoptimalZeroautomationisoptimal Fullautomationisfeasible Partialautomationisfeasible 0% 100% Figure 1: Breakdown of occupation compensation by task-level automation type: different cases of partial equilibrium, ranked from most to least benefit from AI automation. of tasks and the expertise required for execution: whether AI substitutes for low-expertise or high-expertise tasks shapes both wage and employment outcomes (Autor and Thompson 2025). Another strand studies humanโAI collaboration, showing that many tasks are most productively executed when AI and workers share responsibility (Agarwal et al. 2025, Brynjolfsson et al. 2025, Shao et al. 2025). A third line of work investigates AI scaling laws, documenting how performance improves โ and costs escalate โ with sharply diminishing returns at high accuracy levels (Kaplan et al. 2020, Rosenfeld 2021, Hoffmann et al. 2022, Thompson et al. 2022). Macroeconomic analyses link these micro-level choices to aggregate employment outcomes, emphasizing that the net impact of AI depends on the balance between automation of existing tasks and the creation of new, complementary activities (Acemoglu 2025, Hampole et al. 2025). These strands offer important insights but are typically studied in isolation. What remains absent is a unified framework that connects (i) the technical feasibility of task automation, (i) the costs of achieving different levels of AI performance, (i) the required accuracy for economically acceptable task performance, and (iv) the scale at which AI systems are deployed. In particular, we lack a microeconomic model that treats full automation, partial automation, and no automation as competing options within a single optimization problem, and the data to quantify how much labor is optimally automated with current technologies. Without such a framework, it is difficult to answer basic questions: When is it optimal for firms to invest in automation? When they do, is it optimal to fully replace workers on a task, or to retain them as collaborators? And how does the scale of deployment โ within a firm, an industry, or the entire economy โ shift the boundary between these choices? In this paper, we develop such a framework. We reconceptualize automation for the era of large language models and deep learning. Prior automation models were shaped by robotics, where capital directly replaces labor. Modern AI systems introduce a fundamentally different economic structure: the binding constraint is accuracy rather than physical capability, and achieving the accuracy required for reliable deployment means navigating the steep cost curves imposed by scaling laws. We model each task as a bundle of discrete classification decisions with a required accuracy threshold and decompose it into automatable and non-automatable components. On the supply side, we estimate an AI production function from fine-tuning scaling-law experiments that link performance to data, training steps, and model size. On the demand side, we introduce an entropy-based mapping from model accuracy to labor substitution, which quantifies how much human work time AI displaces at each level of performance. Together, these components yield 3 Human-AI Collaboration: Partial vs. Full Automation a task-level optimization problem in which firms choose among labor-only production, partial automation, and full automation. The supply side generates the convex cost structure that drives the partial-automation result: because each input to the AI production function exhibits diminishing returns, the cost of achieving progressively higher accuracy rises steeply. The demand side provides a quantitative account of how work is divided between AI and human collaborators along the automation continuum. Our information-theoretic approach leverages the well-established relationship between entropy and human processing time: higher model accuracy reduces the residual uncertainty that humans must resolve, and this reduction in uncertainty translates into labor savings. This is what allows us to model partial automation as a precise, quantifiable outcome rather than a vague intermediate category. We implement this framework empirically in computer vision โ one of the most developed areas of AI and one where abundant scaling-law data exists (Svanberg et al. 2024). We bring the framework to data by combining several new empirical inputs with existing economic statistics. First, we use O*NET to identify 420 computer-vision-exposed tasks across 263 occupations and measure how much worker time in each occupation is allocated to each task. 3 Second, we use a large survey of workers and domain experts to elicit task-specific required accuracy, which we translate into entropy-based complexity and performance targets. Third, we use GPT-4o to extract for each task the number of vision subtasks, the number of classes per subtask, and the share of the task that is inherently visual, with manual validation by human coders. Finally, we integrate wage, employment, and firm-size distributions from U.S. statistical agencies to scale task-level decisions up to the occupation, firm, industry, and economy-wide levels. Calibrated to these data, our model delivers four main findings: First, at the task level, the convexity of the AI cost function and the entropy-based mapping from accuracy to labor savings imply that marginal costs of improving model performance eventually exceed marginal labor savings for many tasks. As a result, partial automation is often cost-minimizing, even when full automation is technically feasible. This occurs both when AI falls short of the accuracy required for full task replacement and when only a subset of subtasks is economical to automate. Second, task complexity shapes labor substitution. For tasks with fewer subtasks and low complexity, high labor substitution is often optimal. Holding complexity fixed, moving from small-scale deployment to medium-scale deployment increases labor substitution. For more complex tasks, labor substitution is reduced at both small- and medium-scale deployment. Third, at the firm level, we find that approximately 11% of the labor compensation allocated to computer-vision-exposed tasks is economically attractive to automate. Of this, most labor saving comes from partial rather than full automation, with AI systems handling a share of computer-vision-related work while humans retain the residual workload. These estimates pertain to a single AI modality; the economically viable share of automation would be considerably larger when the framework is extended to large language models and other foundation models that cover additional task types. Fourth, the scale of deployment fundamentally shapes automation incentives. When a single AI-as-a-Service offering or AI agent can be shared across all firms performing the same task, fixed development costs are distributed over a much larger user base. This substantially expands the range of tasks for which adoption is economically viable and shifts the optimum toward higher-quality models and higher automation rates. Conditional on adoption, firms typically deploy systems that automate the large majority of the targeted task, leaving only a small residual share to human workers. Under economy-wide deployment, the economically viable share of computer-vision-exposed labor compensation rises substantially. Because large language models and other foundation models exhibit similar scaling-law cost structures, the analytical framework and the qualitative mechanisms we develop are designed to extend to AI-based automation more broadly. 3 An O*NET task refers to the standardized task descriptors in the U.S. Department of Laborโs Occupational Information Network (O*NET), which provides a detailed taxonomy of job tasks for each occupation. We later discuss how these task definitions are incorporated into our data construction in Section 5. Across the entire O*NET database, there are 1,016 occupations and 18,796 tasks, for an average of 18.50 tasks per occupation. There are also 73,308 direct work activities โ or subtasks โ for an average of 3.81 subtasks per task. 4 Human-AI Collaboration: Partial vs. Full Automation Our contributions are fivefold. First, we provide a unified microeconomic framework that endogenizes the choice among full automation, partial automation, and no automation at the task level, explicitly linking these choices to task complexity, required accuracy, and AI scaling laws. Second, we introduce an entropy-based mapping from model performance to labor saving that provides a quantitative framework for modeling humanโAI collaboration and defining feasibility and optimality along the automation continuum. Third, we show that scaling-law diminishing returns create a fundamental asymmetry: the cost of moving from partial to full automation can be orders of magnitude larger than the cost of partial automation itself, so that full automation is often not cost-minimizing, while partial automation frequently emerges as the interior optimum. Fourth, we combine new experimental and survey evidence with large administrative datasets to calibrate this framework to real-world computer vision tasks across the U.S. economy, yielding quantitative estimates of how much labor is optimally automated under current technologies and prices. Fifth, by comparing firm-level deployment, AI-as-a-Service, and AI agents, we show how economies of scale in AI development reshape the automation frontier and help explain why early adoption is concentrated in large firms and standardized tasks, rather than in small firms and complex, high-variability activities. More broadly, automation potential depends jointly on technical feasibility and economic feasibility. Occupations with standardized, low-complexity tasks are more likely to satisfy both conditions, while occupations involving high variability, tacit knowledge, and complex classification demands remain less attractive targets for automation. Although we calibrate the framework to computer vision โ where abundant scaling-law data permits precise empirical implementation โ the underlying cost structures and optimization logic apply to any AI system exhibiting diminishing returns to scale, including large language models and multi-modal foundation models. Our work highlights that partial automation is not merely a transitional state on the path to full automation, but often the long-run cost-minimizing outcome. As costs fall and delivery models such as AI-as-a-Service expand access, the frontier of economically viable automation will continue to widen โ but the fundamental asymmetry between partial and full automation costs implies that humanโAI collaboration will remain a central and durable mode of adoption across a broad range of tasks. The remainder of the paper proceeds as follows. Section 2 reviews related work, covering the transition from technical feasibility to economic viability, the reconceptualization of automation as a continuum, scaling-law regularities, and the connections between micro-level decisions and macro-level employment outcomes. Section 3 presents the theory of cost optimization, including the supply-side AI cost function, the demand-side entropy-based mapping from accuracy to labor substitution, and the firmโs optimization problem across the three automation regimes. Section 4 estimates the AI production function from fine-tuning scaling-law experiments and characterizes its properties, including performance elasticities and input substitutability. Section 5 quantifies the automation frontier by bringing the framework to data โcalibrating it to computer vision tasks using O*NET task data, a domain expert survey, and GPT-4o-derived task decompositions โ and presents results at the task, firm, industry, and economy-wide levels. Section 6 concludes. 2 Related Work The fundamental question facing firms in the AI era is not whether tasks can be automated, but whether they should be automated, and if so, to what extent. This question has driven the economic analysis of AI automation to evolve from early exposure-based assessments toward frameworks that evaluate the feasibility and optimality of humanโAI collaboration. This evolution rests on three key insights that directly motivate our microeconomic approach: (1) technical possibility does not guarantee economic viability; (2) automation exists on a spectrum rather than as a binary choice; and (3) the economics of AI deployment depend fundamentally on scaling laws and cost structures that vary across tasks and required accuracy levels. Building on these insights, the related literature can be grouped into four strands that collectively inform our framework. Section 2.1 reviews the transition from technical feasibility to economic viability. Section 2.2 surveys evidence on humanโAI collaboration and the reconceptualization of automation as a continuum. Section 2.3 links scaling-law regularities to the cost structure of AI systems. Section 2.4 connects these micro foundations to macro-level employment outcomes. Section 2.5 positions our contribution within this broader literature. 5 Human-AI Collaboration: Partial vs. Full Automation 2.1 From Technical Feasibility to Economic Viability Early task-based automation theories trace back to Autor et al. (2003), who built on Zeira (1998). Acemoglu and Autor (2011) combine elements of prior work and consider a continuum of tasks, with technologies augmenting factors of production through a range of possible outcomes: increasing worker productivity, increasing capital productivity, automating work, or creating new tasks (Acemoglu and Restrepo 2024). With the combination of automation and new tasks, this literature established that automation induces the contraction of the range of tasks performed by labor. In recent years, initial AI automation research focused primarily on exposureโthe technical possibility of tasks being automated through machine learning or robotics (Frey and Osborne 2017, Brynjolfsson and Mitchell 2017, Eloundou et al. 2024). However, recent scholarship has pivoted toward the more economically relevant question: when does technical feasibility translate into profitable automation decisions? This shift toward economic viability is exemplified by Autor and Thompson (2025)โs expertise framework, which demonstrates that the wage-and-employment impact of automation depends critically on whether AI eliminates low- expertise or high-expertise subtasks. When automation removes inexpert work, it raises average task expertise, bidding up wages while reducing employment among workers whose skills the automated tasks replaced; eliminating expert work produces the opposite effect. This expertise perspective strengthens the argument that task compositionโnot just exposureโdetermines economic viability and directly affects the cost-benefit computation that firms face. This is precisely the margin our cost-minimization model addresses by incorporating both required accuracy levels and the complexity structure of tasks. Complementing this task-composition insight, Hampole et al. (2025) provide firm-level evidence that labor-demand effects depend on the distribution of AI exposure across tasks within firmsโthat is, within-firm heterogeneity. They find that while mean exposure reduces labor demand, concentrated exposure in specific tasks can increase demand through labor reallocation toward remaining work. This finding validates our frameworkโs emphasis on partial automation: when AI handles concentrated, high-exposure subtasks while workers focus on complementary activities, overall productivity can increase without proportional job displacement. Beyond technical feasibility, adoption frequently requires substantial supporting investments in managerial know- how, organizational restructuring, and process redesign. Brynjolfsson et al. (2021) characterize this pattern through a โproductivity J-curve,โ in which productivity gains materialize only after firms make costly intangible and organizational investments. These studies highlight that adoption costs extend beyond model development to include the broader investments needed to integrate AI systems into production. 2.2 Reconceptualizing Automation The recognition that automation operates along a spectrumโfrom full automation to partial automationโhas motivated a growing body of research on how workers and AI jointly contribute to task execution. However, much of the economic literature has assumed that automation yields full task replacement of workers with machines. With the advent of generative AI and deep learning, the wholesale replacement of labor with capital is no longer the only choice. Partial automation has become a possibility: AI models are able to perform tasks and subtasks while augmenting the skills that workers bring, making the choice more than binary. At the occupation level, some roles can be automated while others continue to be performed by co-workers. At the task level, automated tasks can free up workers for tasks requiring human skills. At the subtask level, automation requires workers to augment subtasks with new or existing capabilities. The choice of automation versus augmentation can thus be conceptualized as partial versus full automation. Human involvement for task completion and quality is the focus of Shao et al. (2025), who introduce the Human Agency Scale (HAS), a five-level audit framework (H1โH5). The scale centers on the ability of workers to act independently while adopting AI agents and provides a shared language to capture the spectrum between automation and augmentation. HAS is based on assessments from 1,500 workers across 104 occupations and annotations from 52 AI experts, covering 844 occupational tasks. Their WORKBank dataset shows that although roughly 46% of tasks fall into an โAutomation Green-Lightโ category, a comparably large share lies in โRed-Lightโ or โR&D Opportunityโ zones 6 Human-AI Collaboration: Partial vs. Full Automation where some degree of human involvement is preferred. This evidence highlights that many tasks are not suitable for full automation and instead benefit from configurations in which AI systems provide assistance while workers continue to guide or refine outputs. Field evidence also demonstrates the potential value of partial automation. Brynjolfsson et al. (2025) and Handa et al. (2025) show that partially automated workflows can substantially increase productivity without proportional reductions in labor. In a large-scale field experiment with radiologists, Agarwal et al. (2023) find that simply providing AI predictions does not always improve diagnostic accuracy; gains arise only when experts are able to integrate AI outputs with contextual information. This supports the broader insight that human and machine inputs are often complementary, particularly when tasks require nuanced judgment or contextual understanding. On the theoretical side, Agarwal et al. (2025) develop a sufficient-statistic approach for allocating decisions between people and predictive systems in classification settings. They show that optimal policies direct cases with high- confidence AI predictions to automated decision-making and assign uncertain cases to humans. Importantly, they find that once such allocation is in place, providing AI predictions to humans yields limited additional benefit: humans tend to under-respond to model output and reduce effort when shown confident AI predictions. These results highlight that the primary efficiency gains come from assigning cases to the agentโhuman or AIโbest suited to handle them, rather than from simultaneous joint decision-making. However, the most suitable agent could change with learning. Our paper takes a novel information-theoretic approach to quantifying humanโAI collaboration. We leverage entropyโ a measure of uncertainty and information contentโto map AI model accuracy into effective labor savings. Higher model accuracy reduces the residual uncertainty that humans must resolve, and this reduction translates proportionally into less required labor time. This connection between entropy and processing time is well supported empirically: Hick (1952) and Hyman (1953) establish that reaction time in choice tasks increases linearly with task entropy, and subsequent work extends this relationship to reading time and binary decision tasks (Lowder et al. 2018, Hu et al. 2022). By embedding these information-theoretic relationships into an economic framework, we provide a rigorous quantitative bridge between AI model performance and labor market outcomes. 2.3 Scaling Laws and the Economics of AI Quality The economic feasibility of both full and partial automation depends not only on technical feasibility but also on the cost structure required to achieve a given level of performance. Recent advances in understanding how AI performance scales with data, compute, and model size provide an empirical basis for assessing these costs. These relationshipsโ commonly referred to as AI scaling lawsโdescribe how increases in training data, computational resources, and model parameters translate into performance improvements (Kaplan et al. 2020, Rosenfeld 2021, Hoffmann et al. 2022, Thompson et al. 2022). The relationship typically follows a power law with diminishing returns as systems approach fundamental limits, thereby quantifying the cost of improving model quality. For computer vision specifically, Svanberg et al. (2024) demonstrate that transfer-learning curves exhibit consistent diminishing-return exponents across classification tasks, implying sharply rising compute costs at very high accuracy levels. Economically, moving from โgoodโ to โnear-perfectโ accuracy can be vastly more expensive, yielding limited incremental benefit. This convex cost structure explains why firms often select partial automation: when the marginal cost of accuracy exceeds its marginal benefit, retaining human oversight becomes the cost-minimizing choice. Our framework embeds these empirical scaling relationships into a formal cost function linking model quality, required accuracy, and task complexity, thereby quantifying when automation remains economically feasible and when human collaboration dominates. These firm-level cost boundaries naturally aggregate into broader labor-market outcomes, providing a bridge to macroeconomic analyses of AI adoption. 2.4 Integrating Micro-Foundations with Macro-Employment Effects The broader employment implications of AI adoption depend fundamentally on the micro-level decisions that determine how tasks are allocated between workers and automated systems. Recent macro-level research emphasizes that 7 Human-AI Collaboration: Partial vs. Full Automation aggregate effects arise from the interplay of task-level substitution, within-firm reallocation, and economy-wide productivity gains. At the aggregate level, employment outcomes emerge from how task-level automation decisions propagate through firms and industries. Hampole et al. (2025) use firm-level variation in AI adoption to show that task- level substitution is largely offset by productivity gains and the reallocation of labor toward complementary activities, yielding modest net employment effects. This finding aligns with Acemoglu (2025), who argues that employment outcomes depend on the relative pace of new-task creation versus automation-driven substitution. In this view, labor- market effects depend not only on technological capability, but also on how firms reorganize production and reallocate tasks in response to it. 2.5 Positioning Our Contribution Existing research has advanced understanding along several dimensions of the automation processโranging from the expertise composition of tasks and the role of human agency to the scaling-law cost structures that govern model performance and the macro employment effects of technological adoption. Yet these strands remain largely separate. What is missing is a unified perspective that links these insights to the firmโs choice among full automation, partial automation, and no automationโthe margin through which technical progress ultimately shapes economic outcomes. Our framework addresses this gap by connecting the technical foundations of AI performance with the economic determinants of adoption. We draw on the expertise-based view of task composition to capture heterogeneity in automation potential; on evidence of humanโAI complementarity to motivate partial automation; on scaling-law research to characterize the cost of achieving required accuracy levels; and on macroeconomic findings to interpret how such micro decisions aggregate into economy-wide outcomes. When full automation is costly or produces unreliable performanceโconditions that both micro and macro evidence shows to be commonโfirms rationally choose partial automation and reallocate labor toward complementary activities rather than eliminating it. This mechanism provides a micro-founded rationale for why AI tends to augment, rather than displace, labor in practice. 3 Theory The starting point for our framework is that firms today can access foundation models โ pre-trained AI systems that deliver some baseline level of performance across a broad class of tasks without task-specific investment. For some tasks, this baseline performance is sufficient for deployment. For many others, however, the accuracy required for reliable task execution exceeds what the foundation model provides out of the box. In such cases, the firm must invest in task-specific fine-tuning: additional training data, more training steps, and potentially larger model architectures. These investments follow the scaling laws documented in the AI literature: each additional unit of data, compute, or model capacity yields diminishing returns in accuracy, so that the cost of improving AI performance is convex and steeply increasing at high accuracy levels. This cost structure gives rise to the central economic trade-off of our model. The labor savings from higher AI accuracy are approximately proportional to the reduction in uncertainty that the AI system achieves (as formalized in Section 3.4), while the costs of achieving that improvement are convex (as formalized in Section 3.3). When the benefit exceeds the cost throughout the relevant performance range, full automation is optimal. When the convex cost curve overtakes the benefit before the required accuracy is reached, partial automation โ deploying AI at an interior accuracy level and retaining human workers for the residual โ is cost-minimizing. When even modest AI investment is uneconomical, no automation is warranted. The formal model below makes these three cases precise. 3.1 Model Set Up To fix ideas, it is helpful to begin with a concrete example, though our framework is not specific to this setting. Consider radiology: each diagnostic case (e.g., reading an X-ray or CT scan) involves a sequence of discrete classification decisions such as determining whether an image is normal or abnormal, identifying the likely condition, and, when applicable, assessing severity. This example simply illustrates that many tasks can be decomposed into a number of discrete decisions and that different decisions may require different accuracy levels. Motivated by this general structure, 8 Human-AI Collaboration: Partial vs. Full Automation we model each task ํ as requiring ํ ํ discrete decisions and define ํ ํ as the accuracy level at which those decisions must be performed. These primitives apply broadlyโto any setting in which AI can potentially automate part of a multi-step decision processโand they will allow us to formalize the economics of full and partial automation in the sections that follow. Consider a firm that must complete task ํ as an integral component of its production process. Previously, this task has been entirely performed by human labor. The firm is now assessing the potential adoption of AI models to either partially or fully automate task ํ. In the second stage, conditional on adoption (ํ ํ = 1), the firm chooses the accuracy level ํ ํ . Throughout, we interpret ํ ํ as the quality of the AI system, expressed in terms of its operating accuracy, the accuracy level at which the firm chooses to operate its AI system. Higher ํ ํ corresponds to deploying a more capable and better-performing model. Assume that the total output of the task is given exogenously as ํ ํ (the number of discrete decisions the firm must make for task ํ over the period). In our context, โoutputโ refers to the total number of completed decisions required for task ํ, which must achieve the target accuracy ํ ํ . Denote the total cost of performing the task ํ as ํถ ํ (ํ ํ ,ํ ํ ,ํ ํ ). If the firm chooses labor-based production (ํ ํ = 0), it will not face any further decision regarding the quality of AI, and the cost does not depend on ํ ํ . In this case, the total cost with a human-only platform is given by: ํถ ํ (0,ํ ํ ,ํ ํ ) = ํถ ํ (0,ํ ํ ) = ํค ํ ํ ํ ํ ํ .(1) ํค ํ represents the exogenous labor wage, and ํ ํ represents the labor time it takes to produce one unit of the output. Here, we assume that each task consists of two proportions: a vision component and a non-vision component. 4 When the firm completes the task using labor only, workers spend a ํฟ ํ fraction of their time on the vision component which a human relies on vision but could be replaced by computer vision AI to some extent and (1โ ํฟ ํ ) on the non-vision component which requires other cognitive or physical abilities and cannot be substituted by computer vision AI. 5 We interpret ํฟ ํ โ [0, 1] as the fraction of task i that is technically automatable by computer vision AI. When ํฟ ํ = 0, the task cannot be automated at all; when ํฟ ํ = 1, the entire taskโs input is technically automatable. If the firm chooses to introduce computer vision AI, i.e. ํ โ ํ = 1, then AI and human labor collaborate to accomplish the task, the total cost of producingํ ํ is ํถ ํ (1,ํ ํ ,ํ ํ ) = ํฟ ํ (1โ ํ ํ (ํ ํ ))+(1โ ํฟ ํ ) ํค ํ ํ ํ ํ ํ + ํ ํ (ํ ํ ,ํ ํ ).(2) Equation (2) consists of two terms. The first term corresponds to the labor cost, and the second term is the cost associated with training and deploying the AI system. For the labor cost, recall that workers originally spent (1โ ํฟ ํ ) of their time on the non-vision portion of the task, which is not automatable by the AI system and must be completed by human labor. For the computer vision portion, we assume that the AI system saves a proportion ํ ํ of the required labor, so human workers still need to perform the remaining(1โํ ํ ) proportion. ํ ํ = 1 corresponds to the case of full automation, where all the vision part is performed by the AI system. 0 < ํ ํ (ํ ํ ) < 1 corresponds to the case of partial automation, where human labor still needs to participate in a portion of the production process; ํ ํ = 0 corresponds to no automation. The degree of automation, ํ ํ , depends on the quality of the computer vision system chosen by the firm, described as accuracy ํ ํ . If the firm chooses a high-accuracy AI system, more human labor will be saved. Thus ํ ํ (ํ ํ ) is a monotonically increasing function. Section 3.4.5 describes our approach to obtain ํ ํ (ํ ํ ). 4 In this paper, we focus on computer vision AI. For other technologies, the โvision/non-visionโ split can be replaced with โAI/non-AIโ or a technology-appropriate decomposition. 5 For example, consider one of gambling managersโ tasks that requires them to circulate among gaming tables to ensure that operations are conducted properly, that dealers follow house rules, and that players are not cheating. The aspects that can be replaced by surveillance camera and computer vision include analyzing the state of the game at each table, detecting abnormal betting patterns such as consistently high winnings, and identifying unusual hand movements of customers. Computer vision cannot detect audio-based cheating, such as whispered communication or coded language. 9 Human-AI Collaboration: Partial vs. Full Automation In the radiology example, partial automation corresponds to natural coarse-to-fine workflows: an AI system may triage cases (normal vs. abnormal), or narrow a large label set down to a short list of plausible findings, while the radiologists perform the remaining fine-grained decisions among the narrowed options. We formalize this division of work in Section 3.4.5. The second term ํ ํ (ํ ํ ,ํ ํ ) represents the cost associated with the building and use of the AI system to produce the output of the task. It is increasing monotonically with respect to system quality ํ ํ and task output quantity ํ ํ . Section 3.3.2 describes our approach to derive ํ ํ (ํ ํ ,ํ ํ ). Equation (2) highlights the trade-off the firm faces in choosing different quality levels of the AI-system. The higher the accuracy, the more human labor will be saved, but the AI-related costs will also become higher. Essentially, ํ ํ (ํ ํ ,ํ ํ ) describes the supply of AI technology in terms of quality; ํ ํ (ํ ํ ) describes the demand of AI technology in terms of quality. Section 3.3 and Section 3.4 will detail our modeling of the supply and demand aspects of AI technology, respectively. Appendix D describes the backward induction approach to solve the firmโs optimization problem and summarizes different scenarios. For each task, the firmโs optimization problem is summarized as follows min ํ ํ ,ํ ํ ํถ ํ (ํ ํ ,ํ ํ ,ํ ํ ).(3) Assume that the total output of the task is given exogenously as ํ ํ (the number of discrete decisions the firm must make for task ํ over the period). In our context, โoutputโ refers to the total number of completed decisions required for task ํ, which must achieve the target accuracy ํ ํ . Denote the total cost of performing the task ํ as ํถ ํ = ํถ ํ (ํ ํ ,ํ ํ ,ํ ํ ). 3.2 Assumptions and Scope Our analysis is conducted within a partial equilibrium framework. Specifically, we take the wage of labor ํค ํ , the cost of computing and data used by AI systems, and the output level ํ ํ of each task as given. These variables are not endogenously determined within the model. Additionally, we assume that the set of tasks required in the economy is fixed and exogenously specified. That is, we do not consider the possibility of new tasks emerging due to advances in AI capabilities. The task composition within each occupation is also assumed to remain constant. Moreover, when evaluating the benefits of AI adoption, we focus exclusively on the labor-saving aspect - i.e., the cost reduction achieved through substitution of human labor with AI in the vision component of the task. We do not account for potential additional gains from AI systems that could arise from improved performance or accuracy levels beyond human levels, which might enhance the quality or output of a task. For more capable systems, the question is not whether systems will be created that have better capabilities than human workers. Rather, when and if more capable systems are economically attractive to build, so are systems with capabilities equal to human workers. Consequently, our modeling approach will correctly identify the extent and timing of automation. The challenge to our approach would occur if building a more capable system becomes economically attractive before the equal-capabilities system. We argue that this is unlikely to be a common occurrence because improving the capability of AI systems results in an enormously rapid increase in the cost of these systems, as shown by Thompson et al. (2022) and as is consistent with foundational computer science work in this area (Kaplan et al. (2020), Henighan et al. (2020), Mikami et al. (2021)). Since less capable systems are unlikely to be able to substitute for human workers, and more capable ones are likely to become economically attractive only later, the modeling that will best predict the automation of human labor is the model with equivalent capabilities. And, because such a system provides similar benefits to the human doing that task, one can compare the economic attractiveness of these systems by comparing their costs. Similarly, any economic benefits stemming from the creation of entirely new products or services enabled by AI technologies are not considered. 10 Human-AI Collaboration: Partial vs. Full Automation 3.3 Supply of AI: In Terms of Both Quality and Quantity In this section, we discuss how to derive the function ํ ํ (ํ ํ ,ํ ํ ). Recall that ํ ํ (ํ ํ ,ํ ํ ) characterizes the cost of deploying an AI system with a given quality requirement ํ ํ and usage levelํ ํ โ that is, the task output produced using the system. Specifically, ํ ํ (ํ ํ ,ํ ํ ) represents the cost associated with implementing an AI system that delivers quality level ํ ํ and generates task outputํ ํ . One of the foundational discoveries in the scaling law literature is that AI models exhibit predictable improvements in performance, ํ ํ , when the inputs to training are scaled up in a systematic manner. 6 We identify three key input factors that affect the quality of an AI system: data, model size, and training steps. Increasing any of these inputs leads to higher compute requirements, and both data and compute contribute directly to the overall cost of developing the AI system. Once a model with a given quality level has been successfully trained, we can consider the fixed cost associated with achieving that level of quality to be determined. However, increasing the modelโs usageโthat is, increasing the number of inference runsโraises the required amount of compute during deployment. As a result, the cost associated with AI system also increases with the usage levelํ ํ . 3.3.1 Scaling Law: The Production Function for AI Quality To illustrate our model, return to the radiology example. One core task in this occupation that can be augmented or substituted by a computer vision system is the diagnosis of diseases and abnormalities from medical images such as X-rays or CT scans. In this context, the accuracy of the computer vision system for this diagnostic task, ํ ํ = ํ(ํท ํ ,ํ ํ , ํ ํ ;ํ ํ ,ํ ํ ),(4) can be measured by its diagnostic performance and is modeled as a production function of three critical inputs: data ํท ํ , training steps ํ ํ , and model size ํ ํ , together with task-specific parameters(ํ ํ ,ํ ํ ). Data (ํท ํ ) refers to the amount of data used for training the computer vision system. Specifically, in our example, it consists of paired X-ray images and their corresponding diagnoses. Training steps (ํ ํ ) represents the number of iterations during the training process where the modelโs parameters are updated. Multiple training steps constitute an epoch, where the entire dataset is passed through the model once. We useํ ํ = ํท ํ ยทํํํํโ ํ as a proxy for training steps. Model Size (ํ ํ ) denotes the size of the computer vision system, characterized by the number of parameters within the AI model. 7 In equation (4), ํ ํ , the number of subtasks, and ํ ํ , the number of classes per subtask, serve as key parameters that characterize the complexity of a given O*NET task. In general, as the number of distinct diseases that a diagnostic system is able to detect increases, the value of ํ ํ will increase accordingly. Moreover, if diagnosis becomes more granular, for example, distinguishing between early-stage and advanced-stage illness, or between mild and severe forms, then each subtask will require more output classes, and ํ ํ will exceed 2. In both cases, increases in ํ ํ and ํ ํ reflect a rise in task complexity. Achieving higher diagnostic accuracy in such complex settings will require a greater investment in resources in terms of input data, computational power, or model sophistication. 8 3.3.2 Cost Minimization for AI Systems The firm needs to address the cost minimization problem: given a required level of model performance ํ ํ and usage levelํ ํ , how to determine the optimal input bundles in developing and adopting the AI system? 6 In this context, training specifically refers to fine-tuning a foundation computer vision (CV) model on a task tailored to a specific occupational setting. 7 A commonly included input in traditional production functionsโlaborโis notably absent from our production function. This omission reflects the idea that labor is not a key driver of improvements in the performance (i.e., quality) of an AI system. While the costs associated with employing AI experts, engineers, and support staff must be accounted for, we incorporate these labor inputs into the fixed cost component (introduced in the following section). 8 In Appendix F, we explain how the number of tasks ํ ํ influences the interpretation of required accuracy. Simply put: when two O*NET tasks share the same required accuracy, the one with a higher ํ ํ value will impose stricter requirements on both individual vision task accuracy and overall, AI model performance. 11 Human-AI Collaboration: Partial vs. Full Automation The objective function in this cost minimization problem is: ํพ(ํท ํ ,ํ ํ , ํ ํ , ํผ ํ ) = ํ ํน + ํ ํท ํท ํ + ํ ํ ํ ํ ํ ํ + ํ ํผ ํผ ํ .(5) In our earlier discussion of AI quality, we introduced three types of inputs that affect system performance. We now extend the framework by introducing a fourth input, denoted by ํผ ํ , which does not affect model quality but captures the computational resources expended for model inference. 9 Specifically, ํผ ํ reflects the GPU hours invested to perform inference and accomplish the task over the period. Unlike the other inputs, ํผ ํ does not affect the quality of the model but is instead directly linked to the quantity of output, that is, the level of usage. We assume that ํผ ํ = ํ GPU ํ ํ ํ ํ is the amount of GPU time needed to produce one unit of task output. ํผ ํ is proportional to ํ ํ because larger models have more parameters and more computations, so they require more GPU time to process inputs during inference. ํ GPU is the number of GPU hours required for a model size ํ ํ usedํ ํ times. The term ํ ํผ ํผ ํ captures the variable cost incurred during deployment, which increases with system usage. We use ํ ํน to denote the fixed cost component, which does not vary with the modelโs quality or usage level. In the context of this paper, ํ ํน primarily reflects the cost of hiring an engineering team to develop, train and maintain the AI system. ํ ํท ํท ํ +ํ ํ ํ ํ ํ ํ represent the variable costs required to train and maintain a model to a given quality level. 10 Here, ํ ํท is the cost of increasing data, and ํ ํ is the cost of increasing training computation. Larger model size and more training steps both increase the total amount of training compute. This cost structure arises from the inherent characteristics of training and deploying computer vision models. For instance, increasing model size typically requires more compute per training epoch and greater computational resources per inference. A more detailed definition of each cost term in the formal cost function is provided in Appendix B, and a broader discussion of the economic interpretation of these costs appears in Section 3.3.3. We can now formally define the corresponding cost minimization problem as follows: min ํท ํ ,ํ ํ ,ํ ํ ,ํผ ํ ํ ํ (ํท ํ ,ํ ํ , ํ ํ , ํผ ํ ) s.t. ํ(ํท ํ ,ํ ํ , ํ ํ ;ํ ํ ,ํ ํ ) โฅ ํ ํ ํผ ํ ํ GPU ํ ํ โฅ ํ ํ (6) The objective is to minimize cost while delivering the required accuracy (ํ ํ ) and supporting the volume of decisions (ํ ํ ) necessary. See Appendix C for the first order conditions. 3.3.3 Cost Components in AI System Development and Deployment To interpret the solution to the cost minimization problem, it is helpful to describe how the different terms in the objective function reflect the economic costs of developing and deploying an AI system. The reduced-form cost function ํ ํ (ํ ํ ,ํ ํ ) represents the minimum expenditure needed to achieve a system with accuracy level ํ ํ that can support a usage level ofํ ํ . Conceptually, these costs fall into two broad categories: fixed costs associated with building and training the model, and variable costs associated with using it at scale. The fixed component corresponds to the one-time investment necessary to create a model capable of meeting the target accuracy level. This includes engineering and development labor involved in designing the system, setting up 9 Inference refers to the process of using a trained model to generate outputs. 10 Model maintenance includes monitoring performance, integrating updated data, retraining or fine-tuning as needed, addressing model drift, and ensuring the long-term reliability and safety of the deployed system. 12 Human-AI Collaboration: Partial vs. Full Automation the training pipeline, and maintaining the model throughout its lifecycle. 11 It also includes the cost of acquiring and preparing training data, as reflected in the term proportional to ํท ํ , since higher accuracy generally requires a larger and more carefully curated dataset. In addition, training a more accurate model requires greater computational resources, which is captured by the component proportional to ํ ํ ํ ํ : larger models require more compute per update, and more training steps increase the total computational workload. Together, these fixed elements determine the minimum development cost necessary to produce a model of quality ํ ํ . A second part of the cost arises from deployment and scales with the amount of output the model produces. Once the model is trained, each inference run requires computational resources that depend on the model size. Because larger models involve more parameters and higher per-call compute requirements, the term proportional to ํผ ํ captures the expenditure associated with running the system to produce the required ํ ํ task outputs. This component therefore reflects the variable cost of using the model in practice, and increases with both the model size chosen to achieve accuracy ํ ํ and the usage levelํ ํ that the firm must support. These two componentsโdevelopment costs that depend on achieving accuracy ํ ํ , and deployment costs that depend on supporting usage ํ ํ โtogether constitute the overall cost structure summarized in ํ ํ (ํ ํ ,ํ ํ ). The structure implies two useful properties. First, the cost is increasing in ํ ํ , because achieving higher accuracy requires more data, more computation, or larger models. Second, the cost is increasing inํ ํ , since each additional model call requires inference compute that grows with the model size. These properties clarify how AI-system costs enter the firmโs optimization problem in Section 3.5, with fixed costs governing whether AI adoption is economically viable at all, and variable costs determining the marginal trade-off between AI usage and human labor for task ํ. 3.4 Demand for AI Quality In this section, we explain our modeling of ํ ํ (ํ ํ ), which is the proportion of labor time (within the AI-automatable subtask) that the AI system could save given the accuracy of the AI system is ํ ํ . ํ ํ (ํ ํ ) serves as a core equation characterizing the demand for AI quality. While the introduction of AI can generate economic gains from different channels, e.g. increase in output, better products, we do not attempt to incorporate all benefits into the analysis. To address this, we use the reduction in labor compensation resulting from labor-saving substitution as a proxy for the benefit of AI adoption. We focus specifically on the demand for AI quality. As we construct a one-to-one correspondence between AI quality and the extent of labor substitution by ํ ํ (ํ ํ ), once quality is determined, both the proportion of labor saving and the nature of human-AI collaboration are determined. On the other hand, based on the assumption that the output at task levelํ ํ is given exogenously in our partial equilibrium framework, we do not require a separate function to explicitly determine the demand for AI quantity or usage. To bridge AI quality and labor saving, we draw on two concepts from information theory: entropy and cross-entropy loss. The entropy concept, well suited for computer vision models, is a measure of the amount of missing information before reception. Entropy captures the clarity of vision or the fidelity of sound, quantifying the average level of uncertainty or information associated with the variableโs potential states or possible outcomes. Cross-entropy loss is the standard measure for classification problems, such as image recognition. It directly compares the predicted probabilities to the true labels. It provides more informative gradients than functions like mean squared error, which can be slow to converge when predictions are confidently wrong. It also provides a measure of the โdistanceโ between the predicted probability distribution and the true distribution, with the goal of making them as close as possible. Minimizing cross-entropy loss is equivalent to maximizing the log-likelihood of the data. Lower cross-entropy loss increases confidence in correct predictions and generally delivers higher accuracy. Cross-entropy loss penalizes confident incorrect predictions more than uncertain predictions. As a result, minimizing cross-entropy loss during training tends to maximize accuracy, but not with perfect correlation. 11 For example, maintenance includes managing updates, integrating new training data, addressing model drift, and ensuring the long-run reliability of the deployed system. 13 Human-AI Collaboration: Partial vs. Full Automation A body of literature in psychology provides empirical support for a relationship between entropy and work time. Intuitively, higher model accuracy reduces uncertainty in outcomes, and a reduction in uncertainty corresponds to fewer effective decisions that a human needs to makeโthereby translating into less required labor time. The cross- entropy loss commonly used in training AI modelsโboth in the mathematical sense and in the context of our settingโis highly correlated with model accuracy. Together, these relationships allow us to establish a mapping from accuracy to labor saving, linking the technical performance of the AI system to the economic outcome of interest. 3.4.1 Entropy and Task Complexity In information theory, classification is essentially a process of narrowing down possibilities. The task entropy, ํป ํกํํ ํ , measures the number of probabilistically equivalent decisions to make in order to rule out all the possibilities. Formally, denote ํ(ํฟ = ํ) as the prior probability of each class ํ โ L. Then, the task entropy is defined as ํป ํกํํ ํ =โ โ๏ธ ํโL ํ(ํฟ = ํ) ln ํ(ํฟ = ํ).(7) We introduced the concept of task complexity in Section 3.3.1, characterizing it using two parameters, ํ ํ (number of subtasks) and ํ ํ (number of classes). Task complexity can be formally captured by its associated task entropy ํป ํกํํ ํ , with higher values of ํ ํ and ํ ํ generally corresponding to higher levels of entropy. 3.4.2 Cross-Entropy Loss and Classifier Performance Cross-entropy loss is a standard metric for evaluating the performance of probabilistic classifiers, encompassing both human and AI decision-makers. It measures the divergence between the true label distribution and the predicted probability distribution over possible outcomes. It characterizes the expected number of additional equivalent decisions required to identify the true class, conditional on the classifierโs probabilistic output. A lower cross-entropy loss indicates a closer alignment between predicted beliefs and actual outcomes, signifying a more accurate classifier. Specifically, denote ํ as an input image, and ํ(ํฟ = ํ|ํ) as true class probabilities conditional on the image. In addition, denote ํ(ํฟ = ํ|ํ) as the predicted probability of the AI system for each class. Formally, the cross-entropy loss of a classifier, ฬ ํป, is defined as ฬ ํป =โE ํ โ๏ธ ํโL ํ(ํฟ = ํ|ํ) lnํ(ํฟ = ํ|ํ) .(8) Cross-entropy loss ฬ ํป and accuracy ํ are both commonly used metrics for evaluating classifier performance. When the task complexity dimensions, ํ and ํ are held fixed, ฬ ํป and ํ approximately follow a monotonic relationship: ฬ ํป = ํน(ํ;ํ,ํ)(9) In several parts of our analysis, we employ this function to transform the accuracy levels into the corresponding cross- entropy loss values. A detailed description of the estimation of this mapping function is provided in Appendix E. As expected, higher accuracy tends to correspond to lower entropy levels, reflecting improved certainty in classification. When a computer vision classification task is performed by a human, workers rarely execute classification tasks with complete precision; a certain margin of error is typically accepted as part of routine performance. If we take the typical accuracy achieved by humans on a given classification task as the required accuracy level for any technology performing that task, and denote it by ํ ํํํ , then based on Equation (9), this yields an associated cross-entropy loss of the completed task, denoted by ฬ ํป ํํํ = ํน(ํ ํํํ ;ํ,ํ). The cross-entropy loss of the AI system, ฬ ํป ํดํผ , measures how many more probabilistically equivalent decisions to make given the output of AI systems. It is governed by the following inequality: ฬ ํป ํํํ โค ฬ ํป ํดํผ โค ฬ ํป ํกํํ ํ .(10) 14 Human-AI Collaboration: Partial vs. Full Automation Figure 2: Quantifying Human-AI Work Allocation via Entropy and Cross-Entropy Loss 3.4.3 Quantifying AI-Human Collaboration As established above, completing a vision classification task involves bringing the cross-entropy loss to ฬ ํป ํํํ . With the introduction of an AI tool, the level of complexity that must be handled by the human worker is reduced. This implies that the AI system effectively completes a portion of the task corresponding to the reduction to ฬ ํป ํดํผ . The remaining complexity - from ฬ ํป ํดํผ to ฬ ํป ํํํ - is resolved by human effort. Figure 2 illustrates this division of the amount of work between the AI system and the human worker. In this framework, we further assume that human work time is proportional to the number of equivalent decisions that must be made by the human worker. Without AI assistance, human effort corresponds to the reduction in cross-entropy loss from the random guess accuracy to the required accuracy level: Human Work Timeโ ฬ ํป ํํํํ โ ฬ ํป ํํํ = ํป ํกํํ ํ โ ฬ ํป ํํํ .(11) With AI assistance, the humanโs contribution is reduced, corresponding only to the remaining gap between the AI systemโs performance and the required accuracy: Human Work Timeโ ฬ ํป ํดํผ โ ฬ ํป ํํํ .(12) As can be seen, our modeling approach implicitly assumes that AI and human labor are substitutes in the production of task-level output. Given the difficulty of directly measuring output at the task level, we restrict our analysis for the purpose of this paper to the benchmark case of perfect substitutability between AI and human inputs. 3.4.4 Empirical Evidence on the Relationship Between Entropy and Human Work Time There is abundant empirical evidence supporting the above assumption that the amount of human labor needed to perform a classification task should be linearly proportional to the number of decisions required, and hence proportional to the reduction in cross-entropy loss/entropy. In particular, Hick (1952) and Hyman (1953) empirically derived the well-known Hick-Hyman law, which states that reaction time in a choice task increases linearly with task entropy. Furthermore, additional studies extend these findings across different domains and time scales. For example, Lowder et al. (2018) demonstrate a positive relationship between surprisal and entropy. Since surprisal is defined as the negative log probability of a word given its preceding context, higher surprisal values are associated with longer reading times. Also, Hu et al. (2022) show that reaction times in binary decision tasks escalate with increasing uncertainty. Together, these works and other similar works robustly support the notion that processing time is directly linked to the entropy of the task. Figure 3 highlights this line of literature, illustrating how the impact of entropy on processing time spans from rapid perceptual decisions to more extended cognitive tasks. 15 Human-AI Collaboration: Partial vs. Full Automation Stimulus information as a determinant of reaction time R. Hyman (1953) Lexical predictability during natural reading: effects of surprisal and entropy reduction M. W. Lowder et al. (2018) Human Decision Time in Uncertain Binary Choice L. Hu et al. (2022) interaction terms were removed from the models because the models would not converge otherwise. Statistical significance was computed using the lmerTest package (Kuznetsova, Brockhoff, & Christensen, 2013) in R. 3.Results We observed a moderate, positive correlation between surprisal and entropy reduction (r=.29,p<.001). This relationship is depicted in Fig. 1. Results of the reading-time analyses are presented in Table 1. Consistent with previous findings, we observed robust effects of word frequency and word length on all reading-time measures, such that increases in word frequency were associated with decreased reading times, whereas increases in word length were associated with increased reading times. Beyond the word- level effects of frequency and length, we also observed a significant main effect of text difficulty in all reading-time measures, such that increases in text difficulty were associ- ated with increased reading times. Crucially, we also observed significant effects of surprisal and entropy reduction. 2 The effect of surprisal was significant across the eye-movement record, such that increases in surprisal were associated with increased reading times in all reading-time measures. In contrast, the effect of entropy reduction was only significant in the early measures of first 0123456 0 1 2 3 4 5 Surprisal Entropy Reduction Fig. 1. Relationship between surprisal and entropy reduction. 1174M. W. Lowder et al. / Cognitive Science 42 (2018) 15516709, 2018, S4, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cogs.12597 by Epfl Library Bibliothรจque, Wiley Online Library on [19/03/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 192RAY HYMAN cases the nonlinear variance is signifi- cant at the 1% level. The first five conditions in Table 2 produced reaction times which were significantly lower than the reaction times for the latter three conditions after the means were adjusted for the linear trend. Practically all of the nonlinear variance for each of t he four Ss was produced by t he one degree of freedom used to compare these two groups of conditions; the nonlinear Q O LJ 600 400 200 0 o a a " G. - EX P. EXP. EX P. C. I n m f 1 f Yยซ r = * 212 + 985 1 3 I53X 1 F.K. - I S A Yยซ 1 I 0 165 t 1 a I27X 0 1 LU 2 800 o U 600 o: 400 200 F.P. Y* IBOt2l5X r โข .955 L.S. Yยป 160+ I99X r ยซ .938 BITS PER STIMULUS PRESENTATION FIG. 1. Reaction time as a function of stimulus information (expressed in bits) when amount of information is varied in three different ways. Data are for Ss G. C, F. K., F. P., and L. S., respectively. Long-Time Scale Symmetry2022,14, 20111 of 17 (a) Figure 4.Results of Experiment 1: (a) Observed data and fitted regression model; (b) Residuals of data; (c) Normal probability density distributions of residuals regarding different cases; (d) Residuals and upper and lower bounds given a significance level of 0.05. According to Equation (16), we used the linear equation y=bยทx+q(25) to fit the data and eventually determine the fitted regression model, i.e., the best fitting line: y (1) =0.1854ยทx (1) +0.6366,(26) wherex (1) refers to information entropy of stimuli andy (1) refers to choice reaction time. It should be noted that the value of parameterqin Equation (25) was set to be equal to the meanโ0.6366โof choice reaction times of the case in which the number of stimuli is one. The reason for this is that, in the HHL, the constant is regarded as the sum of those processing latencies that are unrelated to the reduction of uncertainty [1], e.g., the time spent on encoding the stimulus and executing the response, and this should be the mean choice reaction time in the case in which only one stimulus occurs. According to Equation (18), we calculated the residuals of data(# (1) 1 ,# (1) 2 ,ยท,# (1) 300 ), and plotted them in Figure4b. It can be clearly seen from the figure that there is obvious heteroscedasticity existing in the residuals plot. To make it clearer, we plotted the normal Short-Time Scale *surprisal: negative log probability of a word, given its preceding context; higher surprisal values have been shown to be associated with longer reading times * Figure 3: Highlighted empirical evidence linking task entropy to processing time across different time scales. 3.4.5 Calculating the Labor Substitution Ratio ํ ํ (ํ ํ ) is the proportion of labor time within the AI-automatable subtask that the AI system could save given the accuracy ํ ํ of the AI system. ํ ํ (ํ ํ ) serves as a core equation characterizing the demand for AI quality. We distinguish between two scenarios for the adoption of AI: Scenario 1: Full Automation If the AI system is capable of independently meeting the task requirement, i.e., ฬ ํป req = ฬ ํป AI , then the automation rate is ํ ํ (ํ ํ ) = 1. In this case, the task is fully executed by the AI, with no human involvement. We refer to this as full automation. Scenario 2: Partial Automation When the required cross-entropy loss is lower than what the AI system can achieve alone, i.e., ฬ ํป req < ฬ ํป AI , human-AI collaboration is necessary. In this case, the AI is first deployed to reduce the cross-entropy loss from ํป task to ฬ ํป AI . Subsequently, human workers contribute to further reduce it to the required level ฬ ํป req . We term this arrangement partial automation, and define the corresponding labor substitution ratio for a given task ํ as: ํ ํ (ํ ํ ) = ํป task,ํ โ ฬ ํป AI ํป task,ํ โ ฬ ํป req,ํ = ํป task,ํ โ ํน(ํ ํ ;ํ ํ ,ํ ํ ) ํป task,ํ โ ฬ ํป req,ํ . These two cases can be unified under the following expression: ํ ํ (ํ ํ ) = min ํป ํกํํ ํ,ํ โ ฬ ํป ํดํผ ํป ํกํํ ํ,ํ โ ฬ ํป ํํํ,ํ , 1 = min ํป ํกํํ ํ,ํ โ ํน(ํ ํ ;ํ ํ ,ํ ํ ) ํป ํกํํ ํ,ํ โ ฬ ํป ํํํ,ํ , 1 .(13) It can be easily verified that when the accuracy of the AI model meets the required level, (i.e. ํ ํ = ํ ํํํ,ํ , and thus, ฬ ํป ํดํผ,ํ = ฬ ํป ํํํ,ํ ), ํ ํ (ํ ํ ) = 1, which corresponds to the case of full automation. Otherwise, if the performance of the AI system falls short of the required threshold, (i.e., ํ ํ < ํ ํํํ,ํ , and ฬ ํป ํดํผ,ํ > ฬ ํป ํํํ,ํ ), ํ ํ (ํ ํ ) < 1, which corresponds to the case of partial automation. 16 Human-AI Collaboration: Partial vs. Full Automation With ํ ํ (ํ ํ ), a measure of demand, and ํ ํ (ํ ํ ,ํ ํ ), a measure of supply, we can solve the firmโs cost minimization problem in Equation (3). Recall that the firmโs problem is a two-stage decision. In the first stage, the firm chooses to adopt an AI system. If the firm adopts AI, in the second stage, it selects the quality level ํ ํ of the AI system. In Appendix D, we solve this problem using backward induction. 3.5 Automating Work Through AI Agents In the preceding analysis, we have assumed that each firm operates as an independent decision-making entity, choosing whether to adopt AI for each computer vision task on a firm-specific basis. Under this assumption, any AI system developed is exclusively deployed within the firm that produced it and cannot be shared across organizational boundaries. However, in many occupations, the same task may be performed in a highly similar manner across different firms. When the production of task-level output is sufficiently standardized and tradable across firm boundaries, a decentralized, firm-by-firm approach to AI deployment may lead to inefficiencies due to the duplication of fixed development costs. In such cases, there is scope for the emergence of intermediaries that supply AI agents or AI-as-a-service, developing and offering AI systems for a given task within an industry to multiple downstream firms. Such third-party providers already exist in the LLM space, where firms such as OpenAI, Google, and Anthropic provide AI model services for a known price per use. However, at the industry, subsector, and sector level, domain-specific AI agents and AI-as-a-service providers are generally not yet available. Whether legacy firms or startups innovate such offers, the economics are similar. The adoption principle remains unchanged: buyers will undertake automation whenever expected total benefits exceed total costs. What differs under AI agents or AI-as-a-Service is the level at which these quantities are evaluated. At the firm level, the benefit of automation corresponds to the labor compensation saved within an individual organization. Under AI agents or AI-as-a-Service, the same logic applies, but both benefits and costs are assessed jointly across all participating firms. Fixed development and training costs that would otherwise be incurred repeatedly at the firm level are shared instead, while inference costs continue to scale with total usage. Because the benefit pool now reflects the combined labor compensation of all users, the potential gains from automation expand substantially. Conceptually, the optimization problem determining the optimal level of automation remains conceptually identical to the firm-level case; only the scale of the costโbenefit components change, now aggregated over the group of firms rather than defined for a single firm. In Section 5.5, we formalize this aggregation by applying the same task-level primitives from the firm analysis while replacing firm-specific quantities with their industry-level counterparts, allowing a direct comparison between firm-level automation and AI agent or AI-as-a-Service within a unified analytical framework. 4 Scaling Law Experiments: Estimating the Production Function of AI Quality In Section 3, we outlined our approach to analyzing the quality of each computer vision task, with the aim of identifying the optimal level of quality provision. Central to this process is the necessity of defining the production function for computer vision quality. For the majority of industrial applications, we posit that an effective strategy involves leveraging freely available, pre-trained computer vision models and fine-tuning the models for task-specific applications. In other words, the production of computer vision quality, as discussed in this paper, refers to the process of fine-tuning a pre-trained computer vision foundation model. The AI scaling law literature examines empirical relationships that demonstrate how certain performance metrics of artificial intelligence models, such as model accuracy or loss, improves as a function of increased computational resources, training data, and the number of parameters in AI models. Hestness et al. (2017) and Rosenfeld (2021), enable prediction of improvements in model error, as data and compute are scaled. Subsequent research, exemplified by Kaplan et al. (2020) and Henighan et al. (2020), elucidates optimal input arrangements under constraints on total compute. However, the direct application of findings from existing research is not feasible for three primary reasons. First, the majority of scaling law studies have concentrated on AI models in the context of large language models, while research 17 Human-AI Collaboration: Partial vs. Full Automation specific to scaling laws in computer vision is relatively limited and does not fully meet the requirements of this study. Second, we have identified three critical factors that simultaneously affect both model performance and cost: data, model size, and training steps. For a scaling law to be appropriately utilized as an AI production function, it must account for all three of these factors. However, to our knowledge, the vast majority of existing scaling law studies incorporate only two of these factors. Most commonly, these studies examine the optimal combination of data and model size while holding compute resources constant. Third, existing scaling law literature has largely overlooked the heterogeneity in task difficulty, i.e., the number of classes. The prevailing computer vision scaling law studies often treat the number of classes as a fixed parameter, commonly using the 1,000 classes found in the ImageNet dataset. However, based on a detailed manual analysis conducted by our AI experts across more than 461 selected computer vision tasks, the number of classes encountered in practical economic activities ranges from as few as two to several thousand, with the majority of tasks involving fewer than ten classes. This variation in the number of classes is markedly different from the assumptions made in existing scaling law studies. To bridge this gap, we expand the canonical form of scaling law into the following form: ln(ํป ํดํผ ) = ln ํผ ํท ํ + ํฝ ํ ํ + ํ ํ ํ + ํบ + ํ(14) In the equation above, ํป ํดํผ denotes the cross-entropy loss of the AI model. The logarithmic function on the right side of Equation (14) includes three fractions, corresponding to the three input factors: data (ํท), training steps (ํ ), and model size (ํ ). The remaining variables in the equation represent the parameters to be estimated in the scaling law analysis. The functional form of this scaling law largely follows the classical multi-input scaling laws proposed in prior work by Rosenfeld (2021), Hoffmann et al. (2022), and others. It clearly deviates from the Constant Elasticity of Substitution (CES) forms commonly used in economic modeling. Instead, it can be viewed as one of the simplest and most general non-CES specifications. The CES production function is a neoclassical production function that displays a constant percentage change in the factor proportions due to a percentage change in the marginal rate of technical substitution. ํ = ํนยท ( ํผยท ํพ ํ +(1โ ํผ)ยท ํฟ ํ ) ํ/ํ (15) Where ํ = quantity of output, ํน = total factor productivity, ํพ ,ํฟ = quantities of capital and labor, ํผ = share parameter, ํ = substitution parameter ํ = ((ํโ 1))/ํ, and ํ = elasticity of substitution. Also, ํ is the degree of homogeneity of the production function, when ํ = 1 there are constant returns to scale, ํ < 1 there are decreasing returns to scale, and ํ > 1 there are increasing returns to scale. CES production technology has a constant percentage change in factor proportions (e.g., capital and labor) due to a percentage change in the marginal rate of technical substitution. In the case of AI, and software more generally, the increase in output from additional data, training steps and/or model size requires little or no increase in labor. 12 Further, because the nature of AI production technology is a reduction in loss (increased accuracy and improved decision making), increased factor resources (data, model size and training steps) results in reduced output (cross-entropy loss) as opposed to the traditional increase in produced units. Consequently, a production function in the form of Equation (14) is used in what follows. The choice between using a CES versus a non-CES functional form reflects a fundamental trade-off. Adopting a CES form requires imposing the restrictive assumption that the ability to substitute one factor input for another remains constant over different production levels, which appears to contradict the empirical motivation and flexibility underlying the scaling law literature. Conversely, using a non-CES specification sacrifices access to the standard economic analysis 12 The two-factor CES production function was introduced by Solow (1956) and expanded by Arrow et al. (1961). The scholarship surrounding the CES production function innovation was well before the economics of software and AI development and production existed. 18 Human-AI Collaboration: Partial vs. Full Automation tools built upon CES assumptions. Moreover, the non-CES functional form (Equation (14)) implies that the inputs are complements. Since entropy consistently decreases as any input increases, the parameters ํ, ํ, and ํ must take positive values to reflect this relationship. As a result, it can be shown that the inputs are pairwise complementary. Interestingly, this also provides a useful observation from an economic perspective: computer scientists appear to implicitly assume that the key inputs in AI model production are fundamentally complementary. To incorporate the task complexity, ํ, into the scaling law, we assume a log-log dependency of the scaling law parameters on ํ. Thus, the scaling law is expanded as ln( ฬ ํป ํดํผ ) = ln ํ ํด 0 +ํด 1 ln(ํ) ( ํท ํ ) ํ 0 +ํ 1 ln(ํ) + ํ ํต 0 +ํต 1 ln(ํ) ํ ํ 0 +ํ 1 ln(ํ) + ํ ํถ 0 +ํถ 1 ln(ํ) ํ ํ 0 +ํ 1 ln(ํ) + ํบ ! + ํพ ln(ํ)(16) A lower value of ฬ ํป ํดํผ corresponds to improved performance of the AI system, whereas increases in the values of the three input factorsโdata, training steps, and model sizeโare expected to enhance the AI systemโs performance. Consequently, the anticipated outcome is that, within a reasonable range of ํ values (from 2 to several thousand), the exponents associated with each of the three inputs will be positive, ensuring that ฬ ํป decreases as the levels of these input factors increase. Equation (16) defines the parametric relationship between the cross-entropy loss ฬ ํป ํดํผ and the input factors, which we abbreviate as ฬ ํป ํดํผ = ฬ ํป ํดํผ (ํท,ํ, ํ ;ํ). Combining Equations (9) and (16), the production functionํ(ํท ํ ,ํ ํ , ํ ํ ;ํ ํ ,ํ ํ ) (Equation (4)) is derived as: ํ(ํท ํ ,ํ ํ , ํ ํ ;ํ ํ ,ํ ํ ) = ํน โ1 ( ฬ ํป ํดํผ (ํท ํ ,ํ ํ , ํ ํ ;ํ ํ );ํ ํ ,ํ ํ )(17) where ํน โ1 ( ฬ ํป;ํ,ํ) represents the inverse function of ํน(ํ;ํ,ํ) with respect to ํ. To calibrate this AI quality production function, we fine-tuned a Swin Transformer model under 80 different settings. 13 The computer vision foundation model was pre-trained on 500 randomly selected classes from the ImageNet dataset. 14 These 80 settings capture variation across four key dimensions: task complexity, data size, model size, and training steps. Cross-entropy loss (ํป) measures how closely the modelโs predicted probability distribution aligns with the true labels. It penalizes cases where the model assigns low probability to the correct class, so lower values indicate better predictive performance. Prediction performance is evaluated as the out-of-sample cross-entropy loss when testing the AI model in an object classification task. The experiment systematically varies four dimensions of the fine-tuning process. First, task complexity (ํ) is varied across four levels, measured by the number of distinct outcome classes in the image classification task: 2, 10, 100, and 500 classes, spanning simple binary classification to highly complex multi-class recognition problems. Second, training data size (ํท) is varied across five levels, specified as 13, 65, 130, 650, and 1,300 samples per class, corresponding to total training dataset sizes that scale proportionally with the number of classes. For example, a task with 100 classes and 130 samples per class yields a total training set of 13,000 images. Third, model size (ํ ) is varied across four configurations: 7.3 thousand, 0.4 million, 28.3 million, and 87.8 million parameters, capturing a wide range of model capacities from very small to large-scale architectures. Fourth, these three dimensions together yield 80 unique experimental configurations, each independently replicated 50 times (ํ ), yielding a total of 4,000 observations. 13 The Swin Transformer is a hierarchical vision transformer that processes images by dividing them into local windows and calculating self-attention within these windows. See Liu et al. (2021). 14 The ImageNet database contains 1,000 classes. For this study, we randomly selected 500 classes for pre-training and used the remaining 500 classes to simulate applications of varying complexity, spanning from 2-class to 500-class classification tasks. 19 Human-AI Collaboration: Partial vs. Full Automation The scaling law function related to computer vision fine-tuning is estimated with nonlinear least squares. Following our parametric scaling law introduced in Equation (16), to obtain robust estimates for its parameters, we implemented a multi-run optimization strategy. In our approach, we split the data into training (80%) and test (20%) sets, and we performed 20 independent runs with different random initializations for the parametric scaling law. Using SciPyโs nonlinear least-squares optimizer (curve fit) with a maximum evaluation limit (500, 000), we fitted the proposed scaling function to the training data by minimizing the residuals between the logarithm of the observed cross entropy loss and the model predictions. For each run, we computed the ํ 2 score on the training data to assess the goodness- of-fit, and the best-performing parameter set (highest training ํ 2 ) was selected and further evaluated on the test set. In addition to the best parameter estimates, we aggregated the results from all runs to calculate the mean and standard deviation for each parameter, providing a measure of the uncertainty in the estimation process. The initial values of the parameters are randomly drawn from a Gaussian distribution with mean 0 and variance 0.1. The resulting parameter estimates and their standard errors are reported in Table 1. Table 1: Estimated Parameters for the Scaling Law ParameterEstimateStd. Error ํด 0 -1.4480.046 ํด 1 0.7520.050 ํ 0 -0.0340.005 ํ 1 0.0770.003 ํต 0 1.4740.080 ํต 1 1.0490.000 ํ 0 0.3830.009 ํ 1 0.0200.005 ํถ 0 4.0541.756 ํถ 1 0.3080.409 ํ 0 0.6140.216 ํ 1 -0.0410.045 ํบ-0.2960.173 ํพ-0.1500.044 ํ 2 test 0.963 Our objective is to formulate an empirical scaling law that accurately describes the regime relevant to our study, rather than a universal structural law over all possible combinations of task complexity and resources. Accordingly, Equation (16) should be interpreted as a fitted relation over the supported regime covered by our experiments and by common computer vision settings. In particular, the monotonicity of the fitted law with respect to task complexity is understood to be verified empirically within this domainโthat is, over the range of class counts, data budgets, optimization budgets, and model sizes represented in our experiments and in practically relevant CV configurations. We therefore use the law as a descriptive and comparative model within this regime, and do not claim unrestricted monotonicity or guaranteed validity under arbitrary extrapolation far outside it. 5 Results: Quantifying the Automation Frontier This section describes our approach to quantifying the optimal AI automation decision in the real-world economy, with estimates of cross entropy loss, ฬ ํป ํดํผ , and labor substitution, ํ ํ (ํ ํ ), grounded in the theoretical model developed in Section 3. To measure the required accuracy of each task, we conducted an extensive survey in late 2023. The two key variables obtained from the survey are: required accuracy and random-guess accuracy for each task, which we then feed into a prediction function (9) that assesses the corresponding cross-entropy loss for the computer vision classification tasks. 5.1 Occupation and Task Characteristics Survey To evaluate whether a task is exposed to computer vision technology, it is essential to differentiate between vision tasks and non-vision tasks within the economy. Our analysis relies on data from the O*NET Database 27.1 U.S. Department 20 Human-AI Collaboration: Partial vs. Full Automation of Labor (2023), which provides standardized information about jobs and workers in the United States. The dataset includes descriptions for 1,016 occupations and a total of 18,796 unique tasks. Several recent studies, including Webb (2019), Eloundou et al. (2024), and Brynjolfsson et al. (2018), have utilized the O*NET task framework to analyze the impact of AI. Svanberg et al. (2024) also works within this framework and identifies 420 tasks as vision related. In our analysis, we use the set of vision-related tasks identified by Svanberg et al. (2024). We collected answers from 3,778 respondents. Each of them chose one occupation that they have familiarity with within the 263 occupations that has suitable vision tasks. At least 5, and on average 10 data points are collected for each of the 461 computer vision tasks, and we use the average values for the variables. 15 More details of the survey could be found in Appendix A. The required accuracy variable comes from the question: โWhat is a typical error rate for employees currently doing this task? (i.e. how likely are workers who perform this task to make mistakes?)โ The random accuracy variable comes from the question: โWhat would be the error rate on this task for a worker that had to guess their answer without any information? (e.g. while blindfolded).โ The survey also examines task frequency and time allocation to capture the prevalence and importance of visual tasks within the respondentsโ occupations. Respondents indicate how often they perform the task, with frequencies ranging from rare (less than once per year) to extremely frequent (over 3,000 times a day). They also report the percentage of their overall work time dedicated to performing the task, offering insight into its significance in their day-to-day responsibilities. This information is critical for understanding the role of these tasks in professional settings and for assessing the feasibility of automating tasks that are both prevalent and time intensive. Appendix E provides the equation for estimating cross-entropy loss as a function of required accuracy and number of classes. 5.2 Quantifying Task Complexity and Visual Intensity with ChatGPT-4o The complexity of applying computer vision systems varies significantly across O*NET job tasks, leading to heteroge- neous costs in developing suitable AI models for automation. To systematically measure these costs, we characterize the complexity of the computer vision task associated with each O*NET task along two key dimensions: the number of computer vision tasks required (henceforth, number of tasks), i.e. ํ ํ as in Equation (4); and the number of possible outcome categories within each task (henceforth, number of classes), i.e. ํ ํ as in Equation (4). The number of tasks captures the extent to which an O*NET task decomposes into multiple distinct computer vision classification sub-tasks. A higher value indicates that completing the O*NET task necessitates a system composed of multiple specialized models. The number of classes, in turn, reflects the granularity of classification required within each sub-task, i.e., the number of distinct categories among which the AI system must differentiate. For instance, the O*NET task โinspect motor vehiclesโ for light truck drivers comprises sub-tasks, such as determining โwhether the gas system is operational,โ โwhether the oil level is adequate,โ โwhether the washer fluid level is sufficient,โ among othersโamounting to approximately 20 distinct sub-tasks. In this case, the number of tasks is 20. Each sub-task requires a binary classificationโe.g., โworkingโ vs. โnot workingโโimplying a number of classes of two per sub-task. By contrast, a computer vision system assisting zoologists and wildlife biologists in analyzing animal characteristics to classify species may involve a single overarching classification task (number of tasks = 1) yet requires distinguishing among 500 species (number of classes = 500). A complementary dimension in characterizing O*NET tasks is the vision proportion โํฟ ํ in Equation (2) โ that is, the fraction of the task that is inherently reliant on visual information rather than other sensory modalities or cognitive processes. Many computer-vision-exposed O*NET tasks involve additional non-visual components. For example, in the light truck driver case, workers may need to listen to engine sounds to detect mechanical issues or manually assess tire pressure through tactile feedback. To avoid overestimating the potential automation of such tasks via computer vision systems, we scale task-level estimates by the proportion of time that workers allocate specifically to vision-dependent activities. 15 Our analysis focused on 420 out of the 461 surveyed tasks, as certain occupations were excluded due to incomplete employment or wage data. 21 Human-AI Collaboration: Partial vs. Full Automation To quantify these measures across over 420 O*NET tasks with computer vision exposure, we employed ChatGPT- 4o to generate initial estimates for number of tasks, number of classes, and vision proportion. These outputs were subsequently validated by a team of five researchers through a structured manual review process to ensure alignment with domain knowledge and practical feasibility. The prompts used to generate these estimates are provided in Appendix H. In addition to these original data sources, we incorporate several external datasets to support the simulation imple- mentation of the model. For task and occupation definitions, we use the Occupational Information Network (O*NET), a comprehensive database maintained by the U.S. Department of Labor that provides standardized descriptors of job tasks across occupations. From O*NET, we also extract the share of time that workers in each occupation allocate to specific tasks and the distribution of occupations in different industries/sectors. Employment and wage data come from the 2024 Occupational Employment and Wage Statistics (OEWS) published by the U.S. Bureau of Labor Statistics (BLS). 16 Since scale effects play a critical role in firmsโ automation decisions, it is essential to obtain estimates of the firm size distribution in the economy. We use data from the 2022 Statistics of U.S. Businesses (SUSB), which reports the number of establishments and firm size distributions by six-digit NAICS industry codes. Regarding the engineers and domain experts required for AI training and implementation, as well as their wage levels, the costs for training data and computational resources are based on the findings of Svanberg et al. (2024). 5.3 Computer Vision Tasks Identification Based on O*NET While it is difficult to directly measure task-level output ํ ํ , we estimated ํ ํ ํ ํ โthe product of labor time allocated to the task and outputโusing the following strategy, which is sufficient for the purposes of our analysis: ํค ํ ํ ํ ํ ํ = ํค ํํ ํ ํํ โ ํ ํ ํ ํ = ํ ํ ํ ํํ (18) where ํค ํ is the annual wage for occupation ํ, which can be obtained from Occupational Employment and Wage Statistics (OEWS). ํ ํ represents the total number of employees in each business entity (firms or industry subgroups). The most granular employment data we obtain come from the 4-digit NAICS level in the 2024 Occupational Employment and Wage Statistics (OEWS). While firm-level employment data are available from the 2022 Statistics of U.S. Businesses (SUSB), they require adjustments to be usable in our analysis. To address this, we follow the imputation strategy outlined in Svanberg et al. (2024), as detailed in Appendix G. ํ ํํ represents the time proportion that a human employee with occupation ํ would spend on task ํ. We calculated ํ ํํ by weighting each task according to its score on the O*NET task importance scale, following the methodology of Brynjolfsson et al. (2018) and Webb (2019). Recall that improved performance of AI systems corresponds to reduced cross entropy loss which in turn is a function of available data, model size, and training steps along with the number of classes per subtask (Equation (14)). Consequently, the optimization problem is to minimize cost subject to the constraints of required accuracy and cost for required GPU time to produce tasks incurred during deployment (See Equation (6)). The solution seeks the optimal combination of data, model size, and training steps relative to the cost of labor to perform the same task or subtasks. We discuss the results of this experiment, optimal combination with respect to the reported required task accuracy, over the next subsections. Labor substitutability must consider not only the required accuracy but also the task or subtask complexity. 5.4 Scaling Law With the parameter estimates in Table 1, the implementation of Equation (16) will vary according to the number of classes (n). The estimation of scaling law described in Equation (14), is summarized in Table 2. The table shows the relationship between task complexity and cross-entropy loss, which in turn is a determinant of accuracy. As the number of classes increase, task complexity increases. As each of the classes increase, ํผ, ํฝ, and ํ increase as well, 16 To convert wages into total employer-side labor costs, we apply the BLS-reported wage-to-compensation ratio for civilian workers, which is 1.4498. 22 Human-AI Collaboration: Partial vs. Full Automation contributing to increased cross-entropy loss. With greater complexity, the models have more difficulty in executing tasks with the required accuracy. In addition, the denominators of the first two terms of Equation (14) have coefficient values for a and b that increase as the number of classes increase, finding that as the volume of data increase and the number of training steps increase, the cross-entropy loss decreases. With greater complexity, the models reduce loss and improve accuracy with more data and training. Conversely, as the number of classes increase, the coefficient c decreases, finding that as model sizes increase, the cross-entropy loss increases โ suggesting there are diminishing marginal returns to model size. Our estimated scaling-law coefficients are broadly consistent in magnitude with prior work on data and model scaling. Rosenfeld et al. (2019) report fitted power-law exponents on data size and model size in the range ํผ โ 0.40โ1.10 and ํฝ โ 0.51โ1.16 across vision datasets, suggesting our estimates are within the range documented in established work. The positive returns to data and compute are directionally consistent with Hoffmann et al. (2022), and the diminishing returns to model size as task complexity grows are conceptually aligned with Sharma and Kaplan (2022), who show theoretically that the parameter-scaling exponent decreases as the intrinsic dimension of the data manifold increases, though their notion of complexity is intrinsic dimension rather than class count. We emphasize that our scaling law is empirical and calibrated to our computer vision setting; it is intended to describe behavior within the observed data regime rather than to claim a universal law. Table 2: Scaling Law Parameters under Varying Number of Classes (n) Number of Classes ํผaํฝbํcGk 20.400.029.030.4071.350.59-0.30-0.10 50.790.0923.620.4294.610.55-0.30-0.24 101.330.1448.870.43117.130.52-0.30-0.35 504.460.27264.340.46192.260.46-0.30-0.59 1007.520.32546.910.48238.010.43-0.30-0.69 50025.240.442958.460.51390.690.36-0.30-0.93 100042.520.506120.920.52483.640.33-0.30-1.04 5.4.1 Performance Elasticity Combining Equation (13) and Equation (14), we can obtain the labor substitution ratio as a function of the fine-tuning inputs: ํ ํ (ํท ํ ,ํ ํ , ํ ํ ) = min ยฉ ยญ ยญ ยซ ํป task,ํ โ ํ ํ ํผ ํท ํ ํ + ํฝ ํ ํ ํ + ํ ํ ํ ํ + ํบ ํป task,ํ โ ฬ ํป req,ํ , 1 ยช ยฎ ยฎ ยฌ .(19) In the context of this paper, the extent to which labor can be substituted depends jointly on task complexity, the required accuracy, and the achieved accuracy of the model. For convenience, we define the performance elasticity of input ํ as ํ ํ = ํํ ํํ ยท ํ ํ ,(20) capturing the percentage change in performance r (defined as the labor substitution ratio) in response to a 1% change in input factor X. We obtain the formulas for performance elasticities for data, training steps, and model size, which are reported in Appendix I.2. However, it is important to keep in mind that this elasticity is itself influenced by the intrinsic difficulty of the task. To facilitate a more intuitive analysis, we examine four representative cases based on the parameter values reported in Table 2. Specifically, we consider two levels of task complexity, proxied by the number of classes equal to 2 and 500, respectively, with corresponding task entropy set to ln 2 and ln 500 under the assumption of a uniform distribution (i.e., 23 Human-AI Collaboration: Partial vs. Full Automation entropy equals ln(ํ)). For each level, we evaluate the performance elasticity under two input bundles: a small-scale configuration with ํท = 25,000,ํ = 200,000, and ํ = 250,000; and a medium-scale configuration with ํท = 100,000, ํ = 1,000,000, and ํ = 5,000,000. 17 Table I.1 reports the results. Table 3: Performance Elasticity ScenarioPerformanceDataTraining StepsModel SizeTotal (ํ )(ํ ํท )(ํ ํ )(ํ ํ )Elasticity 2-Class Classification Task Small Scale (I)0.8040.0100.0460.0460.102 Medium Scale (I)0.9110.0090.0210.0070.037 500-Class Classification Task Small Scale (I)0.3510.0230.5440.2830.849 Medium Scale (I)0.7510.0060.1120.0450.163 Notes: ํ denotes the labor-substitution ratio. Scenario I (Small Scale): ํท = 25,000, ํ = 200,000, ํ = 250,000. Scenario I (Medium Scale): ํท = 100,000, ํ = 1,000,000, ํ = 5,000,000. ํท is data size, ํ training steps, ํ model size. Elasticities are evaluated at the respective input bundles. Table 3 shows the labor substitution ratios and the performance elasticities โ the responsiveness of labor to data, training steps, and model size. When the performance elasticity approaches one (e.g. 0.80), increases in model elements (data, training steps, and model size) are nearly fully reflected in labor substitution increases. Conversely, when the performance elasticity approaches zero (e.g. 0.10), increases in model elements show limited response in labor substitution. However, total elasticities remain below one in every scenario, signaling decreasing returns to scale in AI fine-tuning. First, task complexity clearly shapes both the labor-substitution ratio and the elasticity patterns. Compare the top panel of Table 4 (less complexity) with the bottom of Table 3 (more complexity). For the simple binary-classification problem (n = 2), a small-scale configuration delivers a substitution ratio of 80.4%, while moving to medium scale lifts the substitution ratio to 91.1%. By contrast, the 500-class task attains only 35.1% under small-scale training yet rises sharply to 75.1% at medium scale. Second, elasticities systematically decline as training scale expands. In the 500-class task, training-steps elasticity falls from 0.544 to 0.112, and model-size elasticity drops from 0.283 to 0.045 when moving from small to medium scale. These reductions confirm diminishing marginal gains: each additional unit of input yields progressively smaller performance improvements at larger scales. The ranking of elasticities varies with task complexity and resource constraints. Under small-scale conditions the 2-class task shows ํ ํ > ํ ํ > ํ ํท , whereas the 500-class task reverses to ํ ํ > ํ ํ > ํ ํท . Organizations should therefore adapt their investment strategy to both the complexity of the classification problem and their budget. Simple tasks benefit most from enlarging the model, while complex tasks gain more from extending training steps, at least until diminishing returns set in. 17 In large-scale computer vision projects, the input bundle can comprise up to a billion images, with training often requiring tens of millions of training steps and models containing hundreds of millions of parameters. However, we deliberately abstract away from such scenarios. The primary reason is that our analysis focuses on fine-tuning a model to perform a specific task, typically a classification task. In contrast, large-scale foundation models are often designed to be general-purpose, capable of handling a wide range of tasksโincluding generative onesโrather than being tailored to a single narrowly defined use case. In practical applications, especially in industrial settings, the vast majority of tasks are relatively simple binary classification problemsโfor example, detecting the presence or absence of a defect. As such, the scale of models considered in our framework is aligned with the practical requirements of these tasks. Notably, in our setting, excessively large models would trivially achieve full automation (the corner solution) at reasonable task complexity levels. In these corner solutions, the performance elasticity becomes irrelevant. 24 Human-AI Collaboration: Partial vs. Full Automation 5.4.2 Input Substitutability This section quantifies how the three key inputs to model performanceโtraining data (ํท), training steps (ํ ), and model size (ํ )โsubstitute for one another as task complexity grows. Table 4 reports local elasticities of substitution, each evaluated at the same baseline input bundle ํท,ํ, ํ = 100,000, 1,000,000, 5,000,000 . 18 Figure 4 visualises the corresponding isoquant contours for a simple (2-class) and a complex (500-class) task, each drawn with one input held at its baseline level while the other two vary; the elasticities reported in Table 4 are measured along these constant-performance isoquants. We begin with a few observations that align closely with basic intuition. At low input levels the contours are tightly packed, so a modest increase in data, steps, or parameters quickly boosts performance; farther from the origin the contours spread out, illustrating diminishing marginal product (second derivative of output with respect to the input). None of the panels reach the ํ = 1 isoquant (perfect performance), which lies well outside the plotted range. The ํ = 0 isoquant (performance no better than random guessing) sits noticeably higher in the 500-class task than in the 2-class task, confirming the intuition that more complex problems require a larger โentry feeโ of resources before the model becomes useful at all. NearโCobbโDouglas behaviour at low complexity. For the 2-class task, both ํ ํทํ and ํ ํทํ are essentially unity (0.982), placing the technology at the CobbโDouglas production function boundary. In practical terms, a 1% reduction in training steps can be offset by roughly a 1% increase in either data volume or model parameters. The left and middle isoquants in the top row of Figure 4 confirm this: contours are smooth and almost hyperbolic, indicating that the marginal rate of technical substitution changes proportionally with the input ratio. Table 4: Pairwise elasticities of substitution by task complexity Number of classes Elasticity of substitution ํ ํทํ ํ ํทํ ํ ํํ 20.9820.9820.716 50.9180.9180.706 100.8750.8750.698 500.7900.7900.685 1000.7580.7580.694 5000.6930.6950.732 1 0000.6680.6820.749 Notes: Elasticities are evaluated at the baseline bundle ํท,ํ, ํ = 100,000, 1,000,000, 5,000,000 . Rising complementarity with task complexity. The functional form we estimate mechanically keeps each pairwise elasticity between 0 and 1, so some degree of complementarity is always present. Even within that band, however, clear patterns emerge. As the number of classes grows to 50, 100, and 500, ํ ํทํ and ํ ํทํ fall monotonically to about 0.69. In the isoquant panels this shows up as contours that bend ever more sharply toward the axes, meaning that adding just one type of input delivers rapidly diminishing returns unless the other input is expanded in tandem. Put differently, rising task complexity pushes the technology from โalmost substitutableโ toward โstrongly complementary,โ with the curves becoming progressively more convex toward the origin. Stable but low substitutability between steps and model size. The elasticity between training steps and model parameters, ํ ํํ , remains in a narrow band (0.68โ0.75) across all complexity levels. This reflects hardware and optimisation constraintsโe.g. memory limits, gradient stabilityโthat restrict the extent to which longer training can compensate for a smaller model (or vice versa). 18 We obtain the formulas for pairwise elasticities of substitution, which are reported in Appendix I.2. Because the underlying production function is not CES, elasticities in principle vary with the evaluation point. Sensitivity checks using alternative bundles change the levels only marginally, whereas task complexity drives pronounced variation; hence we omit additional bundles for brevity. 25 Human-AI Collaboration: Partial vs. Full Automation T T T T T T Figure 4: Pairwise Isoquant Contours Illustrating Substitutability among Data, Training Steps, and Model Size across Task Complexity (2-Class vs. 500-Class Classification Tasks) Implications for resource allocation. For low-complexity tasks, practitioners enjoy a high degree of freedom in trading off data collection, compute time, and parameter count, enabling cost-efficient rapid prototyping. For more complex tasks, the diminishing elasticities mandate a balanced expansion of all three resources: doubling data alone or scaling parameters alone is insufficient to maintain performance. 5.5 Labor Compensation Saving from Full and Partial Automation by Computer Vision This section addresses two central research questions: (1) How much of total labor compensation can be fully or partially automated by computer vision (CV) systems? (2) How do automation outcomes differ when AI deployment occurs at the firm level versus AI-as-a-Service? To evaluate these questions, we define an automation rate that quantifies the proportion of labor compensation optimally automated under our model. Specifically, we measure this automation rate from two complementary perspectives: first, by considering all tasks within a given deployment scale (whether at the firm level or under an industry-wide AI-as-a-Service arrangement), and second, by restricting attention to computer vision tasks within that same deployment scale to isolate the contribution of CV automation specifically. Let ํธ denote the number of distinct entities that require separate AI systems under each deployment assumption. So far, we have considered the scope of deployment at the firm level, hence, different entities denote different firms. Now, under an industry-wide AI agent or AI-as-a-Service arrangement, we can broaden our deployment scale, namely, changing entities to industry groups (4-digit NAICS), subsectors (3-digit NAICS), sectors (2-digit NAICS), or the entire economy. As the deployment scale expands from individual firms to any of these deployment scales, ํธ decreases monotonically. We index entities by ํ, and solving the optimization problem for any given deployment scale yields the optimal AI automation decisions(ํ โ ํํ ,ํ โ ํํ ) for task ํ within entity ํ. We then define the normalized automation rate ํ as the proportion of labor compensation that is optimally automated conditional on tasks being technically automatable by computer vision (i.e., tasks whose ํฟ ํ โ 0). That is, ํ captures the optimal automation share within the subset of tasks feasible to be performed by computer vision. Formally, 26 Human-AI Collaboration: Partial vs. Full Automation ํ = ร ํ,ํ ํฟ ํ ํค ํ ํ ํํ ํ ํํ I[ํ โ ํํ = 1] ร ํ,ํ ํฟ ํ ํค ํ ํ ํํ ํ ํํ ํ ํ (ํ โ ํํ ).(21) whereI[ยท] is an indicator function, which equals one if the statement within its brackets is true, and zero otherwise. The numerator represents the total labor compensation of CV-related tasks that are economically automated, that is, tasks with ํ โ ํํ and optimal substitution rate ํ ํ (ํ โ ํํ ) > 0. The denominator represents the total compensation of tasks that are technically automatable by CV (ํฟ ํ โ 0). Hence, ํ measures the optimal automation rate within automatable activities. Tasks with ํ ํ (ํ โ ํํ ) = 1. are fully automated, meaning AI systems completely replace human performance, while tasks with 0 < ํ ํ (ํ โ ํํ ) < 1 are partially automated, indicating humanโAI complementarity. The aggregate automation rate in Eq. (21) can be decomposed into contributions from fully automated and partially automated tasks: ํ = ํ full + ํ partial , ํ full = ร ํ,ํ ํฟ ํ ํค ํ ํ ํํ ํ ํํ I[ํ โ ํํ = 1,ํ ํ (ํ โ ํํ ) = 1] ร ํ,ํ ํฟ ํ ํค ํ ํ ํํ ํ ํํ ํ ํ (ํ โ ํํ ), ํ partial = ร ํ,ํ ํฟ ํ ํค ํ ํ ํํ ํ ํํ I[ํ โ ํํ = 1, 0 < ํ ํ (ํ โ ํํ ) < 1] ร ํ,ํ ํฟ ํ ํค ํ ํ ํํ ํ ํํ ํ ํ (ํ โ ํํ ). (22) This decomposition directly aligns with the entropy-based definitions in Section 3.4.5, where ํ ํ (ํ โ ํํ ) = 1 corresponds to complete substitution (zero residual entropy, compared to required entropy) and 0 < ํ ํ (ํ โ ํํ ) < 1 corresponds to partial substitution (positive residual entropy, compared to required entropy). Equation (22) also makes explicit that the aggregate automation rate ํ is an aggregation of task-level automation intensitiesํ ํ (ํ โ ํํ ) across firms or industries. We distinguish between two modes of AI agent or AI-as-a-Service deployment. First, AI systems are shared within industry groups, for example, across firms operating within the same 2-, 3-, or 4-digit NAICS category the model is trained once and used by firms performing comparable visual tasks. Second, a unified AI agent or AI-as-a- Service platform serves across the entire economy, allowing a single model to be deployed wherever equivalent visual recognition tasks appear, regardless of industry. The within-industry case reflects realistic near-term diffusion, while the economy-wide case represents an upper bound on attainable efficiency gains from shared deployment. In Figure 5, we vary assumptions about the extent to which a single AI system can be shared across users performing the same task. The leftmost bar represents the most restrictive case, where we assume that an AI system developed for a specific computer vision task is used exclusively by workers in a single occupation within a single firm. The second to fourth bars reflect progressively broader assumptions under which AI is provided as an agent or as a service, enabling the same system to be deployed by all workers in the same occupation performing the same task across progressively larger economic aggregatesโnamely, within a NAICS 4-digit, 3-digit, 2-digit industry, and eventually the entire U.S. economy. The red segments of Figure 5 represent the share of CV-automatable labor compensation for which full automation is optimal. At the firm deployment scale, only 2.4% of labor compensation is fully automated. Under economy-wide deployment, however, this share rises to 95.8%. The remaining automated portion shown in blue corresponds to tasks where partial automation and humanโAI complementarity are optimal. At the firm level, total automation accounts for 10.8% of labor compensation: 8.4% through partial automation and 2.4% through full automation. Thus, 89.2% of technically feasible CV tasks remain economically unattractive to automate at firm scale, as fixed development costs outweigh potential labor savings when systems cannot be shared across firms. As deployment scope expands, automation potential increases sharply: 96.7% at the 4-digit NAICS level; 98.3% at the 3-digit level; 99.1% at the 2-digit level; and 99.6% under economy-wide deployment. This monotonic increase reflects a fundamental economic mechanism: broader deployment amortizes AI development costs across larger user bases, making tasks that are unprofitable to automate in isolation increasingly attractive at scale. 27 Human-AI Collaboration: Partial vs. Full Automation To place these results in the context of the overall economy, we compute an unnormalized automation rate ํ โฒ , defined relative to total labor compensation across all activities, including those not feasible for CV automation: ํ โฒ = ร ํ,ํ ํฟ ํ ํค ํ ํ ํํ ํ ํํ I[ํ โ ํํ = 1] ร ํ,ํ ํค ํ ํ ํํ ํ ํ (ํ โ ํํ ).(23) ํ and ํ โฒ provide different perspectives on AI automation. While ํ describes the economic attractiveness within technically feasible computer vision tasks only, ํ โฒ describes the economic attractiveness to AI automation considering all tasks in the economy. Individual Businesses Industry Groups (NAICS 4d) Subsectors (NAICS 3d) Sectors (NAICS 2d) U.S. Economy Deployment Scale 0% 20% 40% 60% 80% 100% Fraction of compensation tasks 2% 77% 82% 88% 96% 8% 20% 16% 11% 4% Full AutomationPartial AutomationNo Automation Figure 5: Composition of Full and Partial Automation of Vision-Task Labor Compensation Across Deployment Scales While Figure 5 considers the percentage of the automation fraction among all AI-exposed vision tasks, Figure 6 presents its proportion across all economic activities. As shown in Figure 5, the percentage of the automation fraction among all AI-exposed vision tasks as automation within automatable CV activities approaches completeness at higher aggregation levels, reflecting near-total substitution potential when shared AI systems are deployed widely. In contrast, Figure 6 indicates that when considered by all economic activities, the overall automation rate remains below 4% of total labor compensation, underscoring that only a limited share of the U.S. economy is currently automatable Individual Businesses Industry Groups (NAICS 4d) Subsectors (NAICS 3d) Sectors (NAICS 2d) U.S. Economy Deployment Scale 0% 1% 2% 3% 4% 5% Fraction of Computer Vision Task Total Compensation 0% 4% 4% 4% 4% Full Automation Rate Partial Automation Rate Figure 6: Fraction of Computer Vision Task Compensation Economically Attractive to Automate. 28 Human-AI Collaboration: Partial vs. Full Automation by existing CV technologies. Taken together, these results demonstrate that broader aggregation magnifies attainable benefits and increases automation potential, yet the overall share of total labor compensation affected remains small when evaluated across the entire economy. This dual perspective, spanning firm-level and AI agent or AI-as-a-Service deployment, as well as CV- specific and economy-wide scopes clarifies both the magnitude and the limits of current automation potential. 5.6 Characteristics of Computer Vision System Adoption Figure 7 presents the distribution of firms by automation level choice. Notably, the proportion of firms opting for automation is significantly lower than the corresponding share observed in Figure 5. For instance, at the firm level, over 99.99% of firms do not adopt any form of automation. This outcome can be attributed to the predominance of small firms in the economy, as automation tends to be implemented first by larger firms. Scale plays a crucial role in AI-driven automation. However, despite only a small subset of large firms initially adopting automation, the total labor compensation affected remains substantial (see Figure 5). As we move from the left to the right side of the figure, assuming that an AI platform can serve an entire industry group, subsector, or even sector, the share of deployment entities capable of adopting AI models increases significantly. Figure 8 displays the distribution of partial-automation intensities at the NAICS 4-digit deployment scale, conditional on adoption. As shown in Figure 7, only 10.6% of industry groups adopt partial automation at this scale, but those that do tend to implement systems with very high automation intensity. The histogram in Figure 8 is heavily right-skewed: nearly all systems automate between 90% and close to 100% of the task, and the mean partial automation rate is 93.4%. This indicates that, conditional on adoption, firms typically deploy computer-vision systems that perform almost the entire visual component of the task, leaving only a small residual share to human workers. The left tail of the distribution reflects heterogeneity in task requirements and the diminishing returns to automating small or idiosyncratic subtasks, but the concentration of mass near full automation underscores that partial automationโwhen chosenโis generally high-intensity rather than marginal. Taken together, the results show that partial automation at the 4-digit industry level operates predominantly as a high-coverage technology: although only a subset of industry groups may adopt it, those that do tend to automate nearly all CV-feasible components of the task. Individual Businesses Industry Groups (NAICS 4d) Subsectors (NAICS 3d) Sectors (NAICS 2d) U.S. Economy Deployment Scale 0% 20% 40% 60% 80% 100% Fraction of U.S. Businesses 5% 11% 26% 82% 11% 18% 40% 18% 100% 84% 71% 34% Full AutomationPartial AutomationNo Automation Figure 7: Fraction of the U.S. businesses (by business counts) over different deployment scales. 29 Human-AI Collaboration: Partial vs. Full Automation 30%40%50%60%70%80%90%100% Partial Automation Rate (%) 0 500 1000 1500 2000 2500 3000 Counts of Deployed Computer Vision Models Distribution of Partial Automation Rates for Industry Groups (NAICS 4d) Deployment Scale Mean = 93.40% Figure 8: Counts of deployed computer vision models at industry group (4d) deployment scale by partial automation rates. 5.7 Automation Across Occupation Figure 9 presents automation rates for the top 20 occupations at the NAICS 4-digit deployment scale; Appendix Figure J.1 provides the complete occupation-level breakdown. The x-axis represents the proportion of an occupationโs total tasksโboth computer vision -exposed and non-exposedโthat can be substituted by computer vision systems. The automation rate is influenced by two key factors: (i) the share of an occupationโs tasks that are computer vision-exposed and (i) the extent to which these exposed tasks are suitable for full or partial automation. In the figure, the red segments indicate tasks that can be fully automated, while the blue segments correspond to tasks that can be partially automated. Some occupations have bars containing both red and blue segments, signifying that multiple tasks within the occupation are automatable, with some being more suited for full automation and others for partial automation. According to our findings, occupations with the highest potential for computer vision system adoption include: Transportation Security Screeners (24% of work time automated), Orthodontists (16%), Lifeguards, Ski Patrol, and Other Recreational Protective Service Workers (14%), Cartographers and Photogrammetrists (15%), and Gambling Surveillance Officers (17%), among others. It is important to note that the figure only displays occupations that contain at least one computer vision-exposed task. Many occupations that lack tasks suitable for computer vision applications are therefore not included. As a result, occupations appearing at the bottom of the figure should not be interpreted as the least affected by AI. However, we also observe that several occupations at the lower end of the figure exhibit an automation rate of 0%. This indicates that while some tasks within these occupations are technically computer vision-exposed, further economic feasibility analysis suggests that none of them are viable candidates for automation through computer vision AI systems. To illustrate these aggregate patterns, Table 5 provides a task-level breakdown for two representative occupations: Transportation Security Screeners, and Zoologists and Wildlife Biologists. These examples highlight how technical feasibility (e.g., number of tasks and classes) and economic feasibility jointly shape realized automation outcomes. Transportation Security Screeners demonstrate both high technical feasibility and an optimal automation rate. Most daily activities involve standardized image inspection or object recognition, tasks that CV systems can replicate with high consistency and minimal contextual variation. The overall automation rate of 24% closely mirrors their task time proportion (0.24), indicating that nearly all visually intensive activities are represented and economically meaningful. 30 Human-AI Collaboration: Partial vs. Full Automation 0%5%10%15%20%25% Automation Rate (%) Bakers Agricultural Inspectors Farmworkers and Laborers, Crop, Nursery, and Greenhouse Paper Goods Machine Setters, Operators, and Tenders Crossing Guards and Flaggers Laundry and Dry-Cleaning Workers Prepress Technicians and Workers Umpires, Referees, and Other Sports Officials Diagnostic Medical Sonographers Transportation Vehicle, Equipment and Systems Inspectors, Except Aviation Chefs and Head Cooks First-Line Supervisors of Farming, Fishing, and Forestry Workers Graders and Sorters, Agricultural Products Forest Fire Inspectors and Prevention Specialists Lifeguards, Ski Patrol, and Other Recreational Protective Service Workers Weighers, Measurers, Checkers, and Samplers, Recordkeeping Cartographers and Photogrammetrists Orthodontists Gambling Surveillance Officers and Gambling Investigators Transportation Security Screeners Occupation 9.62% 9.66% 9.87% 10.01% 10.03% 10.21% 10.31% 10.42% 10.48% 10.68% 13.05% 13.09% 13.13% 13.46% 13.73% 14.25% 14.74% 15.67% 17.27% 24.07% 0%5%10%15%20%25% Automation Rate (%) Top 20 Automation Rates Across Occupations for Industry Groups (NAICS 4d) Deployment Scale Full Automation Rate Partial Automation Rate Automatable Not Automated Figure 9: Top 20 Automation Rates across Occupations at Industry Group (4d) Deployment Scale. See Appendix Figure J.1 for all occupations. This reflects not only the dominance of visual inspection tasks but also their relatively low complexity, averaging fewer than ten visual classes per task. The small number of classes and subtasks enables high substitution rates (full automation exceeding 90% of automatable portions), consistent with economies of scale achieved through AI agents or AI-as-a-Service deployment. In this domain, automation effectively reduces labor time while maintaining reliability standards. By contrast, zoologists illustrate the opposite pattern, high task complexity and low optimal automation rate. Despite only two automatable tasks, each accounts for a substantial time proportion (0.13 overall), nearly double the aggregate automation rate (7%). These tasks, such as identifying or classifying animal species, involve extremely large clas- sification spaces (over 250 classes per task) and require reasoning under uncertainty. The resulting automation rate, entirely partial, indicates that while technical feasibility exists, full automation is economically infeasible due to high model development costs and the necessity of human interpretive oversight. These results show how the combination of high complexity and large time shares amplifies both economic and technical barriers to automation. Across occupations, two factors jointly determine automation potential. The first is technical feasibilityโthe share of total work time and subtasks that can be automated, influenced by task standardization and complexity (e.g., number of visual classes). The second is economic feasibilityโthe profitability of implementing automation given the costs of model development, integration, and required human oversight. The interaction between these two is mediated by time proportion: tasks occupying a larger share of total work time have greater economic leverage, yet only when technical feasibility is sufficiently high. Occupations with low complexity and high standardization, such as Transportation Security Screeners, exhibit both strong technical feasibility and high time alignment, making automation economically viable. By contrast, occupations such as Zoologists demonstrate that even when tasks are visually intensive or consume substantial time, high variability, tacit knowledge, and extensive classification complexity constrain economic attractiveness. These findings reinforce that partial automation, where computer vision systems 31 Human-AI Collaboration: Partial vs. Full Automation Table 5: Task-Level Automation Profiles Occupation / Task / DWATask AutomationFullPartialTime Proportion #Task #Class Transportation Security Screeners Task: Inspect carry-on items using x-ray equipment to determine whether items contain objects that warrant further investigation. DWA: Inspect cargo to identify potential hazards. 2.80%2.72%0.08%2.81%210 Transportation Security Screeners Task: Inspect checked baggage for signs of tampering. DWA: Inspect cargo to identify potential hazards. 3.24%3.24%0.00%3.24%27 Transportation Security Screeners Task: Locate suspicious bags pictured in printouts sent from remote monitoring areas, and set these bags aside for inspection. DWA: Locate suspicious objects or vehicles. 4.09%4.09%0.00%4.09%23 Transportation Security Screeners Task: Monitor passenger flow through screening checkpoints to ensure order and efficiency. DWA: Monitor access or flow of people to prevent problems. 3.05%3.05%0.00%3.05%23 Transportation Security Screeners Task: Record information about any baggage that sets off alarms in monitoring equipment. DWA: Record information about suspicious objects. 2.79%2.71%0.08%2.79%133 Transportation Security Screeners Task: Send checked baggage through automated screening machines, and set bags aside for searching or rescreening as indicated by equipment. DWA: Inspect cargo to identify potential hazards. 3.08%2.99%0.09%3.09%37 Transportation Security Screeners Task: View images of checked bags and cargo using remote screening equipment and alert baggage screeners or handlers to possible problems. DWA: Inspect cargo to identify potential hazards. 2.10%2.04%0.06%2.11%215 Transportation Security Screeners Task: Watch for potentially dangerous persons whose pictures are posted at checkpoints. DWA: Maintain surveillance of individuals or establishments. 3.21%3.21%0.00%3.21%21 Transportation Security Screeners (Agg Results)24.07%23.77%0.30%24.38%-- Zoologists and Wildlife Biologists Task: Inventory or estimate plant and wildlife populations. DWA: Measure environmental characteristics. 0.75%0.00%0.75%6.08%10251 Zoologists and Wildlife Biologists Task: Analyze characteristics of animals to identify and classify them. DWA: Examine characteristics or behavior of living organisms. 5.13%0.00%5.13%7.13%1260 Zoologists and Wildlife Biologists (Agg Results)5.87%0.00%5.87%13.21%-- assist rather than replace human judgment remains the most prevalent and economically rational trajectory for adoption in practice. 6 Conclusion This paper develops a microeconomic framework to determine not just if a task can be automated, but to what extent automation is optimal. By integrating the technical realities of AI scaling laws with an entropy-based measure of task complexity, we bridge the gap between computer science metrics (such as cross-entropy loss) and economic outcomes (labor substitution). This unified approach moves the analysis of automation beyond binary exposure measures, allowing for a rigorous evaluation of the trade-off between full automation and partial, human-in-the-loop collaboration. Our results highlight a sharp divergence between technical feasibility and economic viability. While computer vision systems are technically capable of performing a wide range of visual tasks, the convex cost structure of achieving human- level accuracyโdriven by diminishing returns to data, model size and training stepsโoften renders full automation economically suboptimal. Instead, we find that partial automation is frequently the cost-minimizing equilibrium. In our calibrated model, the optimal strategy for most viable tasks involves AI systems reducing vision-related labor with human workers retaining the high-entropy share of the workload. This finding provides a micro-foundation rationale for why AI deployment has thus far been characterized more by augmentation than by wholesale displacement. 32 Human-AI Collaboration: Partial vs. Full Automation Furthermore, our analysis demonstrates that the scale of deployment is a fundamental determinant of the automation frontier. Under a firm-level deployment model, high fixed development costs act as a barrier to entry, making automation feasible only for the largest firms or the most standardized tasks. However, when costs are distributed through AI agents or an AI-as-a-Service model, the range of economically viable tasks expands significantly, shifting the optimal choice toward higher-quality models and higher automation rates. Despite this expansion, we estimate that even under optimistic economy-wide deployment assumptions, the total share of U.S. labor compensation attractive for computer-vision automation remains below 4%. This suggests that while AI will be transformative for specific occupationsโparticularly those with standardized, low-complexity visual componentsโits aggregate labor market impact may be more gradual than purely exposure-based estimates suggest. While our empirical application focuses on computer vision, the logic of our framework is modality-agnostic. The interplay among scaling laws, task entropy, and cost components applies equally to Large Language Models (LLMs) and multimodal systems. As these technologies continue to mature, our framework provides a tractable tool for predicting which cognitive tasks will remain in the human domain and which will be ceded to machines. Ultimately, our findings suggest that the future of work will not be defined by a simple race against the machine, but by a complex optimization of human-AI collaboration, governed as much by the economics of model training and its life cycle overall as by the capabilities of the models themselves. The framework opens multiple directions for future work. Extending the model to include quality-adjusted output or error-sensitive payoff functions would capture settings where exceeding human-level accuracy yields additional economic value. Incorporating endogenous task creation would enable richer macroeconomic implications, including how AI reshapes occupations over time. Finally, applying the approach to multimodal or language-based AI systems would broaden the analysis beyond computer vision. This work contributes to a deeper understanding of how advances in AI technology propagate through firms, tasks, and labor markets, and provides a foundation for evaluating the economic consequences of humanโAI collaboration. 33 Human-AI Collaboration: Partial vs. Full Automation References Acemoglu, D. (2025). The simple macroeconomics of AI. Economic Policy, 40(121):13โ58. Acemoglu, D. and Autor, D. (2011). Skills, tasks and technologies: Implications for employment and earnings. In Handbook of labor economics, volume 4, pages 1043โ1171. Elsevier. Acemoglu, D. and Restrepo, P. (2018a). Modeling automation. In AEA Papers and Proceedings, volume 108, pages 48โ53. Acemoglu, D. and Restrepo, P. (2018b). The race between man and machine: Implications of technology for growth, factor shares, and employment. American Economic Review, 108(6):1488โ1542. Acemoglu, D. and Restrepo, P. (2024). A task-based approach to inequality. Oxford Open Economics, 3(Supplement 1):i906โi929. Agarwal, N., Moehring, A., Rajpurkar, P., and Salz, T. (2023). Combining human expertise with artificial intelligence: Experimental evidence from radiology. Working Paper 31422, National Bureau of Economic Research. Agarwal, N., Moehring, A., and Wolitzky, A. (2025). Designing human-AI collaboration: A sufficient-statistic approach. Working Paper 33949, National Bureau of Economic Research. Arrow, K. J., Chenery, H. B., Minhas, B. S., and Solow, R. M. (1961). Capital-labor substitution and economic efficiency. The Review of Economics and Statistics, 43(3):225โ250. Assets, F. (2003). Consumer durable goods in the united states, 1925-99. Washington, DC: US Government Printing Office, September. Autor, D. and Thompson, N. (2025). Expertise. Working Paper 33941, National Bureau of Economic Research. Autor, D. H., Levy, F., and Murnane, R. J. (2003). The skill content of recent technological change: An empirical exploration. The Quarterly Journal of Economics, 118(4):1279โ1333. Brynjolfsson, E., Li, D., and Raymond, L. (2025). Generative ai at work. The Quarterly Journal of Economics. Brynjolfsson, E. and Mitchell, T. (2017). What can machine learning do? workforce implications. Science, 358(6370):1530โ1534. Brynjolfsson, E., Mitchell, T., and Rock, D. (2018). What can machines learn, and what does it mean for occupations and the economy? volume 108, pages 43โ47. AEA papers and proceedings. Brynjolfsson, E., Rock, D., and Syverson, C. (2021). The productivity j-curve: How intangibles complement general purpose technologies. American Economic Journal: Macroeconomics, 13(1):333โ72. Eloundou, T., Manning, S., Mishkin, P., and Rock, D. (2024). GPTs are GPTs: Labor market impact potential of llms. Science, 384(6702):1306โ1308. Endsley, M. R. and Kaber, D. B. (1999). Level of automation effects on performance, situation awareness and workload in a dynamic control task. Ergonomics, 42(3):462โ492. Frey, C. B. and Osborne, M. A. (2017). The future of employment: How susceptible are jobs to computerisfation? Technological forecasting and social change, 114:254โ280. Hampole, M., Papanikolaou, D., Schmidt, L. D., and Seegmiller, B. (2025). Artificial intelligence and the labor market. Working paper, National Bureau of Economic Research. Handa, K., Tamkin, A., McCain, M., Huang, S., Durmus, E., Heck, S., Mueller, J., Hong, J., Ritchie, S., Belonax, T., et al. (2025). Which economic tasks are performed with ai? evidence from millions of claude conversations. arXiv preprint arXiv:2503.04761. Henighan, T., Kaplan, J., Katz, M., Chen, M., et al. (2020). Scaling laws for autoregressive generative modeling. Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y. (2017). Deep learning scaling is predictable, empirically. Hick, W. E. (1952). On the rate of gain of information. Quarterly Journal of Experimental Psychology, 4(1):11โ26. Hobbhahn, M. and Besiroglu, T. (2022). Trends in gpu price-performance. Epoch AI. Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L., Welbl, J., Clark, A., et al. (2022). Training compute-optimal large language models (2022). arXiv preprint arXiv:2203.15556. Hu, L., Pan, X., Ding, S., and Kang, R. (2022). Human decision time in uncertain binary choice. Symmetry, 14(2):201. Hyman, R. (1953). Stimulus information as a determinant of reaction time. Journal of Experimental Psychology, 45(3):188โ196. Kadra, A., Janowski, M., Wistuba, M., and Grabocka, J. (2023). Scaling laws for hyperparameter optimization. In Thirty-seventh Conference on Neural Information Processing Systems. 34 Human-AI Collaboration: Partial vs. Full Automation Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020). Scaling laws for neural language models. arXiv preprint arXiv:2001.08361. Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L., and Doll ฬ ar, P. (2015). Microsoft coco: Common objects in context. Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9992โ10002. Lowder, M. W., Choi, W., Ferreira, F., and Henderson, J. M. (2018). Lexical predictability during natural reading: Effects of surprisal and entropy reduction. Cognitive science, 42:1166โ1183. Mikami, H., Fukumizu, K., Murai, S., Suzuki, S., Kikuchi, Y., Suzuki, T., ichi Maeda, S., and Hayashi, K. (2021). A scaling law for synthetic-to-real transfer: How much is your pre-training effective? NVIDIA (2024). Megatron-lm: Training multi-billion parameter language models using model parallelism. https://raw. githubusercontent.com/NVIDIA/Megatron-LM/main/README.md. Accessed: 2024-07-21. Parasuraman, R., Sheridan, T. B., and Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and CyberneticsโPart A: Systems and Humans, 30(3):286โ297. Rosenfeld, J. S. (2021). Scaling laws for deep learning. arXiv preprint arXiv:2108.07686. Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N. (2019). A constructive prediction of the generalization error across scales. arXiv preprint arXiv:1909.12673. Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. (2015). Imagenet large scale visual recognition challenge. Shao, Y., Zope, H., Jiang, Y., Pei, J., Nguyen, D., Brynjolfsson, E., and Yang, D. (2025). Future of work with AI agents: Auditing automation and augmentation potential across the U.S. workforce. Sharma, U. and Kaplan, J. (2022). Scaling laws from the data manifold dimension. Journal of Machine Learning Research, 23(9):1โ34. Sheridan, T. B. and Verplank, W. L. (1978). Human and computer control of undersea teleoperators. Technical report, MIT Man-Machine Systems Laboratory. Singh, S., Yadav, A., Jain, J., Shi, H., Johnson, J., and Desai, K. (2024). Benchmarking object detectors with coco: A new path forward. Solow, R. M. (1956). A contribution to the theory of economic growth. The Quarterly Journal of Economics, 70(1):65โ94. Sullivan, B. (2023). Average stock market return. Forbes Advisor, 16. Svanberg, M., Li, W., Fleming, M., Goehring, B., and Thompson, N. (2024). Beyond AI exposure: Which tasks are cost-effective to automate with computer vision? SSRN Working Paper No. 4700751. Thompson, N., Borge, N. J., Pande, A., and Fleming, M. (2021). Demand forecasting with a.i.: Building the business case. MIT Sloan Research Brief. Thompson, N., Fleming, M., Tang, B. J., Pastwa, A. M., Borge, N., Goehring, B. C., and Das, S. (2024). A model for estimating the economic costs of computer vision systems that use deep learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 23012โ23018. AAAI Press. Thompson, N. C., Greenewald, K., Lee, K., and Manso, G. F. (2022). The computational limits of deep learning. U.S. Department of Labor (2023). O*NET. https://w.onetcenter.org/overview.html, Accessed: 2023-04-04. Webb, M. (2019). The Impact of Artificial Intelligence on the Labor Market. Available at SSRN 3482150. Wistuba, M., Kadra, A., and Grabocka, J. (2022). Supervising the multi-fidelity race of hyperparameter configurations. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K., editors, Advances in Neural Information Processing Systems. Zeira, J. (1998). Workers, machines, and economic growth. The Quarterly Journal of Economics, 113(4):1091โ1117. Zhang, Y., Zong, R., Shang, L., Zeng, H., Yue, Z., Wei, N., and Wang, D. (2023). On optimizing model generality in ai- based disaster damage assessment: A subjective logic-driven crowd-ai hybrid learning approach. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence (IJCAI), pages 6317โ6325. Zimmer, L., Lindauer, M., and Hutter, F. (2021). Auto-pytorch: Multi-fidelity metalearning for efficient and robust autodl. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(9):3079โ3090. 35 Human-AI Collaboration: Partial vs. Full Automation A Occupation and Task Characteristics Survey A.1 Survey Participants The survey leveraged Prolific, a widely regarded online platform known for its diverse and reliable respondent pool, to recruit participants and direct them to the survey. The target population comprised U.S. adults aged 18 and older, with at least two years of work experience and familiarity with occupations involving the visual tasks under investigation. Screening and survey implementation were conducted via Qualtrics, a platform that enabled dynamic question flows, tailored question paths, and streamlined participant navigation. The final sample exhibited substantial demographic diversity, encompassing respondents aged 18 to 85, with a median age of 35 years. Gender representation within the sample was approximately 44.34% female and 41.21% male, with the remainder identifying as non-binary, other genders, or preferring not to disclose their gender. The average annual household income of respondents was approximately $94,000, with 8.44% reporting incomes below $25,000 and 9.26% reporting incomes above $150,000. A.2 Survey Questions The survey questions fall into following categories: Task Navigation Questions. These questions direct participants to identify their industry, occupation, and familiar computer vision tasks. Once an occupation is selected, the first computer vision task associated with that occupation is presented. Participants are then asked whether they are familiar with the task. If they indicate familiarity, they proceed to answer a series of detailed questions related to the task; otherwise, the survey bypasses the task and moves on to the next. Tasks within the same occupation are presented in random order. Among the 263 occupations included in the survey, 153 are associated with two computer vision tasks, with an overall average of 1.75 tasks per occupation. After completing the questions for a given task, participants are asked whether they are willing to continue with an additional task. By allowing participants to skip tasks or exit the survey at their discretion, we mitigate the risk of response quality deterioration due to cognitive limitation or participant fatigue. Questions on Task Characteristics. These questions aim to collect detailed information about the nature of visual tasks, focusing on their outcomes, error tolerance, and frequency, as well as their significance within the respondentโs occupation. Participants are guided to categorize the vision task into classification, visual generation (e.g., drawing or films), or other types of vision tasks according to the outcome of the task. For classification tasks, which is the focus of this paper, respondents further determine the specific type: Binary classification, e.g., whether thereโs defect in the product. Single-choice multi-class classification, involving the selection of one option from a predefined list. E.g., determining whether it is sunny, cloudy, or rainy based on weather images. Multi-choice classification, requiring the identification of multiple items from a set. E.g., determining which products at a grocery store need to be re-stocked. Respondents are also asked to list the number of classes and potential classes relevant to the task. The survey includes a series of error tolerance questions to evaluate performance standards for tasks. Respondents are asked to estimate the typical error rate under normal conditions for competent workers, providing a benchmark for the expected level of accuracy. Additionally, they are asked to identify the worst acceptable error rate, beyond which workers would no longer be considered qualified to perform the task. To contextualize the difficulty of the task, respondents also report the error rate that would occur under random guessing. These measures allow for a nuanced assessment of the precision requirements of each task and the potential challenges of automating them while maintaining acceptable performance levels. The survey also examines task frequency and time allocation to capture the prevalence and importance of visual tasks within the respondentsโ occupations. Respondents indicate how often they perform the task, with frequencies ranging from rare (less than once per year) to extremely frequent (over 3,000 times a day). They also report the percentage of their overall work time dedicated to performing the task, offering insight into its significance in their day-to-day 36 Human-AI Collaboration: Partial vs. Full Automation responsibilities. This information is critical for understanding the role of these tasks in professional settings and for assessing the feasibility of automating tasks that are both prevalent and time-intensive. Questions on Data Availability and Data Collection Costs. The survey examines the feasibility of collecting and recording data, a critical component for training AI systems and automating vision tasks. Participants were asked to determine whether the necessary task inputsโsuch as images, videos, and task outcomesโare already being recorded in their workflows. For tasks lacking sufficient data, the survey investigated the estimated costs of data collection, ranging from โvery cheapโ to โvery expensive.โ Additionally, respondents identified any adjustments required to facilitate data capture, including the need for specialized equipment, technical expertise, or changes in existing procedures. A.3 Survey Data Quality Control To ensure the reliability and accuracy of the collected data, the survey implemented rigorous quality control measures throughout its design and administration. Key methods included: Attention Checks. Specific criteria were embedded within the survey to identify inconsistent or illogical responses. For example, respondents were required to logically align their answers to questions on error rates. Workers were expected to make fewer errors than randomly guessing and less than the worst acceptable error rate. Responses failing these criteria were flagged and excluded from the dataset. Screening for Familiarity. At the start of the survey and before each task, participants were asked to confirm their familiarity with the chosen occupation and the listed tasks. This ensured that only qualified respondents contributed data, enhancing the relevance and credibility of the results. Survey Duration Monitoring. The average completion time for each wave was recorded to identify outliers. We excluded respondents who completed the survey too quickly, potentially indicating rushed or careless responses. Iterative Survey Refinement.To ensure high-quality survey results, we conducted a trial run (Wave 1) in July 2023, which collected 451 valid responses, and followed it with a formal survey (Wave 2) in November 2023, yielding 3,327 valid responses. The trial run provided an initial dataset for analyzing survey data quality and identifying areas for improvement. Based on insights from Wave 1, we refined the survey design by rephrasing ambiguous questions, introducing additional response options to capture more nuanced perspectives, and enhancing the instructions on both the Prolific and Qualtrics platforms. Furthermore, the payment plan was revised to better align with respondent expectations and boost engagement. To preserve data integrity, we ensured that no respondents participated in both waves. After thorough post-survey processing, we confirmed that the results from the two waves could be seamlessly combined following appropriate adjustments. This iterative approach allowed us to address potential issues early and significantly enhance data clarity and accuracy in the formal survey. B Cost Function Specifications To interpret the solution to the cost minimization problem, it is helpful to describe how the different terms in the objective function reflect the economic costs of developing and deploying an AI system. The reduced-form cost function ํ ํ (ํ ํ ,ํ ํ ) represents the minimum expenditure needed to achieve a system with accuracy level ํ ํ that can support a usage level ofํ ํ . Conceptually, these costs fall into two broad categories: fixed costs associated with building and training the model, and variable costs associated with using it at scale. The fixed component corresponds to the one-time investment necessary to create a model capable of meeting the target accuracy level. This includes engineering and development labor involved in designing the system, setting up the training pipeline, and maintaining the model throughout its lifecycle. It also includes the cost of acquiring and preparing training data, as reflected in the term proportional to ํท ํ , since higher accuracy generally requires a larger and more carefully curated datasets. In addition, training a more accurate model requires greater computational resources, which is captured by the component proportional to ํ ํ and ํ ํ : larger models require more compute per update, and 37 Human-AI Collaboration: Partial vs. Full Automation more training steps increase the total computational workload. Together, these fixed elements determine the minimum development cost necessary to produce a model of quality ํ ํ . A second part of the cost arises from deployment and scales with the amount of output the model produces. Once the model is trained, each inference run requires computational resources that depend on the model size. Because larger models involve more parameters and higher per-call compute requirements, the term proportional to ํผ ํ captures the expenditure associated with running the system to produce the required ํ ํ task outputs. This component therefore reflects the variable cost of using the model in practice and increases with both the model size chosen to achieve accuracy ํ ํ and the usage levelํ ํ that the firm must support. These two componentsโdevelopment costs that depend on achieving accuracy ํ ํ , and deployment costs that depend on supporting usage ํ ํ โtogether constitute the overall cost structure summarized in ํ ํ (ํ ํ ,ํ ํ ). The structure implies two useful properties. First, the cost is increasing in ํ ํ , because achieving higher accuracy requires more data, more computation, or larger models. Second, the cost is increasing inํ ํ , since each additional model call requires inference compute that grows with the model size. These properties clarify how AI-system costs enter the firmโs optimization problem in Section 3.3.2, with fixed costs governing whether AI adoption is economically viable at all, and variable costs determining the marginal trade-off between AI usage and human labor for task ํ. Our approach to deriving the cost function ํ (ํท,ํ, ํ, ํผ;ํ) is primarily based on the cost framework established by Svanberg et al. (2024). Specifically, ํ (ํท,ํ, ํ, ํผ;ํ) is decomposed as ํ (ํท,ํ, ํ, ํผ;ํ) = ํ eng + ํ(ํ data (ํท)+ ํ train (ํ, ํ)+ ํ inf (ํ, ํผ)).(B.1) Here, ํ ํํํ refers to the cost of hiring a group of engineering team and subject matter experts to write the training and inference code. We assume that the coding effort holds constant regardless of the change in other input factors (which mostly just involves changing the hyperparameters in the code). ํ ํํํกํ refers to the cost of collecting and manually labeling the data, which scales linearly with ํท. ํ ํกํํํํ refers to the GPU cost of training the AI system, which scales linearly withํ and ํ . ํ ํํํ refers to the GPU cost of deploying the AI system to produce the task output, which scales linearly with ํ and ํผ. In the case where there are multiple classification tasks, i.e., ํ > 1, we assume that a distinct CV system is developed for each classification task, each undergoing respective data collection, training, and inference processes. Hence, the corresponding cost terms in Equation (B.1) is proportional to ํ. On the other hand, we assume that the engineering effort remains constant regardless of ํ. Since an AI system, once developed, can be used over a finite operational lifespan, let ํฟ denotes the number of years the system remains in use. 19 Each of the cost terms above should thus be interpreted as the present value of the discounted stream of costs incurred over the ํฟ-year period with a discount rate of ํ.In the following subsections, we detail the specifications for each term. B.1 Data Collection Cost We assume that data collection is a continuous effort. Specifically, to collect and maintain a training data set of size ํท, we assume there is an initial round of collecting ํท images, and ํ recur rounds of data renewal per year for ํฟ years. Each data renewal round collects a new set of ํท data. Therefore, the data collection cost is specified as follows ํ data (ํท) = ํ data ํท+ ํฟ โ๏ธ ํก=1 ํ recur ํ data ํท (1+ ํ) ํก ,(B.2) where ํ data represents the cost of collecting and labeling one datum. ํ represents the discount rate. We assume that the cost of collecting and labeling one datum for a visual-judgment task equals the opportunity cost of the worker performing that judgment within their regular occupation. In other words, this cost is proportional to the workerโs wage and the share of time spent on the taskโs vision-related component, and is normalized by the frequency 19 U.S. Bureau of Economic Analysis specifies a 5-year service life for custom software in the context of fixed assets. We use this number as the operational lifespan of a deployed AI system. 38 Human-AI Collaboration: Partial vs. Full Automation with which such judgments are performed within the occupation. Formally, for each task ํ, we define the per-datum data-collection cost as ํ data = ํค ํ ํ ํํ ํฟ ํ ํน ํ ฮฆ ํ ,(B.3) where ํค ํ , ํ ํํ , and ํฟ ํ follow the definitions introduced in Equation (2). ํน ํ denotes the annualized frequency of visual-judgment decisions for task ํ, derived directly from survey questions detailed in Appendix A.2. ฮฆ ํ represents the multi-annotator reliability factor, indicating the number of independent judgments required per datum to achieve consensus-quality labels. We adopt ฮฆ ํ = 8 consistent with established standards for multi-annotator labeling practices in canonical computer vision datasets (Russakovsky et al. 2015, Lin et al. 2015, Singh et al. 2024) to ensure inter-rater reliability in visual annotation across diverse task settings. To prevent unrealistically low cost estimates when judgment frequency is exceptionally high or wage data are missing, we apply a conservative minimum threshold of $0.05 per datum based on Zhang et al. (2023), ensuring that our cost model remains both stable and realistic. Thus, ํ data,ํ captures the marginal opportunity cost of labeling one datum for task ํ, expressed in wage-time per decision. This measure forms the micro-foundation for the aggregate data-collection cost component employed in the automation-rate analysis in Section 5. B.2 Training Cost Similar to data collection, training a CV system is also modelled as a continuous effort, consisting of ํ init rounds of initial training (for algorithmic development) and ํ recur rounds of re-training per year. Each training step involves approximately ํน GPU ํยทํ input GPU FLOPs, where ํน GPU denotes the number of GPU FLOPs per floating point operation, ํ input denotes the number of pixels of the input image. Hence each training round, withํ training steps in total, requires ํน GPU ํ input ํํ GPU FLOPs. The training cost involves the GPU cost, which is paid by GPU hours. Denote ํ ํบํํ as the price for one GPU hour, and we assume that the price will drop by 1+ ํ GPU each year. Denote ํ FLOP as the number of FLOPs the GPU can perform per hour, andํ GPU as GPU utilization, which is the percentage of time when the GPU is being utilized. Then the total training cost, including the initial and recurring training rounds, amounts to ํ train (ํ, ํ) = ํ init + ํฟ โ๏ธ ํก=1 ํ recur (1+ ํ GPU ) ํก (1+ ํ) ํก ยท ํ GPU ํ FLOP ํ GPU ยท ํน GPU ํ input ํํ.(B.4) B.3 Inference Cost The inference cost involves the GPU cost to perform inference. The number of GPU hours needed for inference per year is ํผ. In the main text we assume that adopting AI does not change the total quantity of outputโthat is, the number of times the task is completed remains the same as it would under human labor. However, it is very difficult to estimate how many times a human worker performs a given task over the course of a year. Given that the inference cost term is relatively small compared to other cost components in computer vision applications, we adopt a conservative principle to avoid underestimating it. Specifically, we assume an extremely efficient human worker who can make one classification decision per second (recall that the share of time a worker spends on a given task varies across tasks). This assumption gives us an upper bound on the number of task completions, which in turn provides a reasonable approximation forํ ํ . Using this, and considering the inference speed of typical neural networks such as VGG-Net, we estimate ํ ํ GPU = 2.01138ร 10 13 , which, combined with Equation (D.6), allows us to eliminate the unknown ํ ํ in the cost minimization process as in Equation (3). 39 Human-AI Collaboration: Partial vs. Full Automation Taking into account the annual price drop of the GPU and the discount rate, the inference cost of ํฟ years can be computed as ํ inf (ํ,ํ) = ํฟ โ๏ธ ํก=1 1 (1+ ํ GPU ) ํก (1+ ํ) ํก ยท ํ GPU ํ ํบํํ ํผ.(B.5) B.4 Engineering Cost The engineering cost consists of two terms, one is the implementation cost,ํถ impl , incurred during the initial CV system development stage, and the other is the maintenance cost, ํถ maint , incurred annually. So the total engineering cost amounts to ํ eng = ํถ impl + ํฟ โ๏ธ ํก=1 ํถ maint (1+ ํ) ํก .(B.6) We used the original implementation and maintenance costs reported in Thompson et al. (2021). The original case study, based on a real-world IBM deep learning deployment project for time series forecasting, reports an upfront implementation cost of $1,765,000 and an annual maintenance cost of $242,840, primarily reflecting the compensation of IBM engineers, client engineers, and subject matter experts. To account for changes in labor costs over time, we adjust the original costs reported to 2024 dollars. Given that these expenditures are driven by high-skilled technical labor, we apply an inflation adjustment using the U.S. Employment Cost Index (ECI) for Private Industry Workers in Professional and Technical Services. According to the BLS ECI data, the index value was approximately 130.2 in Q1 2018 and rose to approximately 158.3 by Q4 2024, reflecting a cumulative increase of about 21.5% over this period. 20 Applying this adjustment, the inflation-adjusted 2024 values are approximately $2.14 million for the implementation cost and $295,000 for the annual maintenance cost. B.5 Cost Function Parameters The values of the cost function parameters are listed below. Table B.1: Cost Function Parameters SymbolValueDescriptionSource ํ init 1000Number of initial fine-tuning rounds for algorithmic developmentKadra et al. (2023), Wistuba et al. (2022), Zimmer et al. (2021) ํ recur 6Retraining frequency per yearThompson et al. (2024) ํ GPU 0.22GPU-specific annual price discount rateHobbhahn and Besiroglu (2022) ํ0.05Discount rate applied across the economySullivan (2023) ํฟ5System operational lifespan in yearsAssets (2003) ํ data dynamicCost of collecting and labeling one datumThompson et al. (2024), Singh et al. (2024), Russakovsky et al. (2015), Lin et al. (2015) ํน GPU 6FLOPs multiplier for training per token per parameter (accounts for forward + backward passes) Kaplan et al. (2020) ํ FLOP 4ร 10 12 FLOPs per GPU-hourSvanberg et al. (2024) ํ GPU 0.34GPU hour cost for 4 FP-32 TFLOPS GPU on cloudSvanberg et al. (2024) ํ GPU 0.4Average GPU utilization rate during trainingNVIDIA (2024) ํ input 256 2 Input image size in pixelsImageNet image size and standard CV model input size ํถ impl 2,144,475Initial implementation cost for system developmentThompson et al. (2021) ํถ maint 295,123Annual system maintenance costThompson et al. (2021) C Cost Minimization First-Order Conditions To find the first-order conditions for the AI cost minimization problem in Equation (6), the Lagrangian is: L = ํ ํ (ํท ํ ,ํ ํ , ํ ํ , ํผ ํ )+ ํ 1 [ ํ ํ โ ํ(ํท ํ ,ํ ํ , ํ ํ ;ํ ํ ,ํ ํ ) ] + ํ 2 ํ ํ โ ํผ ํ ํ GPU ํ ํ 20 U.S. Bureau of Labor Statistics, Employment Cost Index (ECI) for Private Industry Workers in Professional and Technical Services, available at: https://w.bls.gov/news.release/eci.t01.htm 40 Human-AI Collaboration: Partial vs. Full Automation Table B.2: Input prices under default parameter values (2024 USD). Coeff.ExpressionValue (USD) ํ ํน ํถ impl + 1โ ( 1/(1+ ํ) ) ํฟ 1โ 1/(1+ ํ) ํถ maint 3,486,090 ํ ํท ํ 1+ ํ recur 1โ ( 1/(1+ ํ) ) ํฟ 1โ 1/(1+ ํ) ํ data 6.19 ํ ํ 6ํ ๏ฃฎ ๏ฃฏ ๏ฃฏ ๏ฃฏ ๏ฃฏ ๏ฃฏ ๏ฃฐ ํ init + ํ recur 1โ 1 (1+ํ)(1+ํ GPU ) ํฟ 1โ 1 (1+ํ)(1+ํ GPU ) ๏ฃน ๏ฃบ ๏ฃบ ๏ฃบ ๏ฃบ ๏ฃบ ๏ฃป ํ GPU ํ input ํ FLOP ํ GPU 3.83ร 10 โ6 ํ ํผ ํ ๏ฃฎ ๏ฃฏ ๏ฃฏ ๏ฃฏ ๏ฃฏ ๏ฃฏ ๏ฃฐ 1โ 1 (1+ํ)(1+ํ GPU ) ํฟ 1โ 1 (1+ํ)(1+ํ GPU ) ๏ฃน ๏ฃบ ๏ฃบ ๏ฃบ ๏ฃบ ๏ฃบ ๏ฃป ํ GPU ํ GPU 40ร 50ํฟ ํ ํ ํ ํ/ํ GPU 1.29ร 10 โ8 whereํ 1 โฅ 0 is the Lagrange multiplier on the quality constraint andํ 2 โฅ 0 is the Lagrange multiplier on the inference capacity constraint. C.1 Partial Derivatives of the Lagrangian The first-order condition with respect to ํท ํ (Data) is: ํL ํํท ํ = ํํ ํ ํํท ํ โ ํ 1 ํํ ํํท ํ = 0 and rearranging: ํํ ํ ํํท ํ = ํ 1 ํํ ํํท ํ The marginal cost of data is the shadow price of quality times the marginal product of data. The first-order condition with respect to ํ ํ (Training Steps) is: ํL ํํ ํ = ํํ ํ ํํ ํ โ ํ 1 ํํ ํํ ํ = 0 and rearranging: ํํ ํ ํํ ํ = ํ 1 ํํ ํํ ํ The marginal cost of training steps is the shadow price of quality times the marginal product of training steps. The first-order condition with respect to ํ ํ (Model Size) is: ํL ํํ ํ = ํํ ํ ํํ ํ โ ํ 1 ํํ ํํ ํ + ํ 2 ํผ ํ ํ GPU ํ 2 ํ = 0 and rearranging: ํํ ํ ํํ ํ = ํ 1 ํํ ํํ ํ โ ํ 2 ํผ ํ ํ GPU ํ 2 ํ The marginal cost of model size is the shadow price of quality times the marginal product of model size minus the shadow price of inference capacity times the effect of model size on inference capacity. Larger models reduce inference capacity per GPU hour, hence the negative term. The first-order condition with respect to ํผ ํ (Inference Compute) is: ํL ํํผ ํ = ํํ ํ ํํผ ํ โ ํ 2 ํ GPU ํ ํ = 0 41 Human-AI Collaboration: Partial vs. Full Automation and rearranging: ํํ ํ ํํผ ํ = ํ 2 ํ GPU ํ ํ The marginal cost of inference compute is the shadow value of inference capacity times the marginal contribution of inference to capacity. C.2 Partial Derivatives of the Cost Function From Equation (B.1) in Appendix B, compute partial derivatives of ํ : ํ (ํท,ํ, ํ, ํผ;ํ) = ํ eng + ํ ํ data (ํท)+ ํ train (ํ, ํ)+ ํ inf (ํ, ํผ) From Equation (B.2): ํ data = ํ data ํท+ ํฟ โ๏ธ ํก=1 ํ recur ํ data ํท (1+ ํ) ํก ํํ ํ ํํท ํ = ํ ํ data " 1+ ํฟ โ๏ธ ํก=1 ํ recur (1+ ํ) ํก # From Equation (B.4): ํ train (ํ, ํ) = ํ int + ํฟ โ๏ธ ํก=1 ํ recur (1+ ํ GPU ) ํก (1+ ํ) ํก ! ยท ํ GPU ํ FLOP ํ GPU ยท ํน GPU ํ input ํํ ํํ ํ ํํ = ํ ํ int + ํฟ โ๏ธ ํก=1 ํ recur (1+ ํ GPU ) ํก (1+ ํ) ํก ! ยท ํ GPU ํ FLOP ํ GPU ยท ํน GPU ํ input ํ ํํ ํ ํํ = ํยท ํ ํ ยท ํ where ํ ํ is the effective training cost coefficient. Model size appears in both training and inference costs: ํํ ํ ํํ = ํ ํํ train ํํ + ํํ inf ํํ From the training component: ํํ train ํํ = ํ ํ ยทํ From the inference component (Equation (B.5)): ํ inf (ํ,ํ) = ํฟ โ๏ธ ํก=1 1 (1+ ํ GPU ) ํก (1+ ํ) ํก ยท ํ GPU ํ GPU ยท ํผ But ํผ = ํ GPU ํํ (from the binding constraint), so: ํํ inf ํํ = ํฟ โ๏ธ ํก=1 1 (1+ ํ GPU ) ํก (1+ ํ) ํก ยท ํ GPU ํ GPU ยท ํ GPU ํํ = ํ ํผ ยทํ where ํ ํผ is the inference cost coefficient. Therefore: ํํ ํ ํํ = ํ [ ํ ํ ยทํ + ํ ํผ ยทํ ] 42 Human-AI Collaboration: Partial vs. Full Automation From Equation (B.5): ํ inf (ํ,ํ) = ํฟ โ๏ธ ํก=1 1 (1+ ํ GPU ) ํก (1+ ํ) ํก ยท ํ GPU ํ GPU ยท ํผ ํํ inf ํํผ = ํฟ โ๏ธ ํก=1 1 (1+ ํ GPU ) ํก (1+ ํ) ํก ยท ํ GPU ํ GPU ํํ inf ํํผ = ํยท ํ ํผ C.3 Partial Derivatives of ํ From the scaling law (Equation (14)): ln( ฬ ํป AI ) = ln ํผ ํท ํ + ํฝ ํ ํ + ํ ํ ํ + ํบ + ํ and ํ is related to accuracy through Equation (17): ํ(ํท ํ ,ํ ํ , ํ ํ ,ํ ํ ,ํ ํ ) = ํน โ1 ฬ ํป AI (ํท,ํ, ํ,ํ); ํ ํ ,ํ ํ Using the chain rule: ํํ ํํท = ํํน โ1 ํ ฬ ํป AI ยท ํ ฬ ํป AI ํํท From the scaling law: ํ ฬ ํป AI ํํท = ฬ ํป AI ยท โํผํ/ํท ํ+1 ํผ/ํท ํ + ํฝ/ํ ํ + ํ/ํ ํ Since ํน is decreasing in accuracy (higher accuracy results from lower entropy): ํํน โ1 ํ ฬ ํป AI < 0 Therefore ํํ/ํํท > 0 Increased data results in higher accuracy. By similar logic, ํํ/ํํ > 0 More training steps also results in higher accuracy. And ํํ/ํํ > 0 larger models also provide higher accuracy. C.4 Solve for Shadow Prices Assuming both constraints bind, the system of first-order conditions gives: For data: ํ ํ data " 1+ ํฟ โ๏ธ ํก=1 ํ recur (1+ ํ) ํก # = ํ 1 ํํ ํํท For training steps: ํ ํ ํ ํ = ํ 2 ํํ ํํ 43 Human-AI Collaboration: Partial vs. Full Automation For model size: ํ [ ํ ํ ยทํ + ํ ํผ ยทํ ] = ํ 1 ํํ ํํ ํ โ ํ 2 ํผ ํ ํ GPU ํ 2 ํ With ํผ = ํ GPU ํํ : ํ [ ํ ํ ยทํ + ํ ํผ ยทํ ] = ํ 1 ํํ ํํ ํ โ ํ 2 ํ ํ For inference: ํยท ํ ํผ = ํ 2 ํ GPU ํ From the inference first-order condition: ํ 2 = ํยท ํ ํผ ยท ํ GPU ํ The shadow price of requiring higher inference capacity, ํ 2 = ํํ /ํํ , is the cost of producing one more unit of task output. Substituting into the model size first-order condition and solving for ํ 1 : ํ 1 = ํ ํํ/ํํ ํ ( [ ํ ํ ยทํ + ํ ํผ ยทํ ] + ํ ํผ ยท ํ GPU ํ ยท ํผ ํ ํ GPU ํ 2 ํ ) The shadow price of achieving higher accuracy, ํ 1 = ํํ /ํํ ํ , is the increase in total cost per unit increase in required accuracy โ the โmarginal costโ that appears in the Stage 2 optimization (Equation (D.6)). The optimality conditions for input ratios combining first-order conditions yields optimal input ratios. From the first-order conditions for data and training steps: ํํ/ํํท ํํ/ํํ = ํ data ํท 1+ ร ํฟ ํก=1 ํ recur (1+ ํ) ํก ํ ํ ํ Marginal rate of technical substitution between data and training steps equals their cost ratio. From the first order conditions for training steps and model size. More complex due to inference constraint interaction, but principle is same: equalize marginal product per dollar across inputs. C.5 Complete System of First-Order Conditions The complete system of first-order conditions is: ํํ ํํท = ํ 1 ํํ ํํท ํํ ํํ = ํ 1 ํํ ํํ ํํ ํํ = ํ 1 ํํ ํํ โ ํ 2 ํ ํ ํํ ํํผ = ํ 2 ํ GPU ํ ํ(ํท,ํ, ํ,ํ,ํ) = ํ ํ ํผ = ํ GPU ยท ํ ยทํ ํ 1 , ํ 2 โฅ 0 These seven equations (plus non-negativity) determine the six unknownsํท โ ,ํ โ , ํ โ , ํผ โ ,ํ 1 ,ํ 2 . The solution char- acterizes the cost-minimizing way to build an AI system that achieves accuracy ํ ํ and can processํ ํ task outputs. 44 Human-AI Collaboration: Partial vs. Full Automation D Task Level Optimization in AI Adoption Before proceeding, we introduce four terms that will be used throughout the analysis. Total Benefit (TB) refers to the labor cost savings generated by automation. Total Cost (TC) denotes the cost of deploying an AI system with accuracy ํ ํ . Marginal Benefit (MB) refers to the incremental labor-saving gain from a marginal improvement in AI performance, and Marginal Cost (MC) denotes the incremental increase in AI-system cost associated with such an improvement. These definitions will allow us to characterize the firmโs optimization problem using standard cost-benefit analysis. D.1 Stage Two Optimization We begin by analyzing the firmโs choice of the optimal AI quality ํ ํ conditional on adopting AI (ํ ํ = 1) in stage one. To elucidate the structure of the firmโs cost minimization problem, we rewrite Equation (2) as: ํถ ํ (1,ํ ํ ,ํ ํ ) = ํค ํ ํ ํ ํ ํ โ TB ํ (ํ ํ ,ํ ํ )+ TC ํ (ํ ํ ,ํ ํ )(D.1) where total cost consists of the baseline labor cost ํค ํ ํ ํ ํ ํ , net of automation benefits and plus automation costs. The term TB ํ (ํ ํ ,ํ ํ ) = ํฟ ํ ํ ํ (ํ ํ ) ํค ํ ํ ํ ํ ํ (D.2) captures the labor cost savings enabled by the AI systemโinterpreted as the automation benefitโwhere ํฟ ํ denotes the proportion of the task technically feasible to be automated and ํ ํ (ํ ํ ) the labor substitution ratio at AI quality level ํ ํ . The cost of implementing AI, denoted by TC ํ (ํ ํ ,ํ ํ ) = ํ ํ (ํ ํ ,ํ ํ )(D.3) represents the total expenditure associated with deploying an AI system of quality ํ ํ to produce outputํ ํ . The firmโs optimization problem in Stage two is therefore: max [ TB ํ (ํ ํ ,ํ ํ )โ TC ํ (ํ ํ ,ํ ํ ) ] (D.4) Equation (9) establishes a monotonic transformation between accuracy ํ and cross-entropy loss ฬ ํป. Since cross-entropy loss aligns more linearly with labor time, we recast the optimization problem using cross-entropy loss as the choice variable. The reformulated problem is: max ํ ํ TB ํ ํ ํ ( ฬ ํป AI,ํ ),ํ ํ โ TC ํ ํ ํ ( ฬ ํป AI,ํ ),ํ ํ (D.5) The first-order condition over ฬ ํป AI,ํ is: MB โ ฬ ํป = MC โ ฬ ํป ,where MB โ ฬ ํป = ํ TB ํ ํ(โ ฬ ํป AI,ํ ) ,MC โ ฬ ํป = ํ TC ํ ํ(โ ฬ ํป AI,ํ ) (D.6) For ease of interpretation, we introduce a negative sign to ฬ ํป AI,ํ such that increases in the AI systemโs quality correspond to decreases in the AI modelโs cross-entropy loss ฬ ํป AI,ํ . The sufficient conditions for an interior solution are: ํ TB ํ ํ(โ ฬ ํป AI,ํ ) = ํ TC ํ ํ(โ ฬ ํป AI,ํ ) and ํ 2 TB ํ ํ(โ ฬ ํป AI,ํ ) 2 < ํ 2 TC ํ ํ(โ ฬ ํป AI,ํ ) 2 45 Human-AI Collaboration: Partial vs. Full Automation Because TB is piecewise linear in entropy and MB is piecewise constant while TC is convex due to AI scaling laws and MC is increasing, an interior solution arises when the MB curve intersects the rising MC curve before reaching ํ req . See Figure 1, panel 2 (partial automation). Conversely, the sufficient conditions for a corner solution at ํ = ํ req are MBโฅ MC throughout the interval. Equiva- lently: ํ TB ํ ํ(โ ฬ ํป AI,ํ ) โฅ ํ TC ํ ํ(โ ฬ ํป AI,ํ ) โ ํ ํ โค ํ req When the marginal benefit of improving accuracy exceeds the marginal cost throughout, it is optimal to fully automate to the required accuracy level. See Figure 1, panel 1 (full automation). There is also a corner solution for no automation when at ํ = ํ req , MBโค MC. Even at the lowest feasible accuracy, the cost of minimal AI exceeds benefits. See Figure 1, panel 3 (no automation). D.2 Stage One Optimization The stage one optimization problem is straightforward. The firm compares two options: (i) adopting the AI system and incurring the minimized cost from the stage two optimization; or (i) not adopting AI and bearing the baseline labor cost. The firm chooses the option that yields the lower total cost. Denote ฬ ํป โ ํ and ํ โ ํ as the optimal cross-entropy loss and optimal accuracy of the AI system, respectively. The second step of the backward induction is to solve for the optimal ํ ํ , formally: ํ โ ํ = ( 1 if ํถ ํ (1,ํ ํ ,ํ ํ ) < ํถ ํ (0,ํ ํ ) 0 otherwise (D.7) To support the subsequent analysis, we formally define a set of terms related to the feasibility and optimality of task automation. These definitions clarify under what conditions automation may be considered possible or desirable from a cost-minimization perspective. โข Full automation is feasible if the cost of deploying an AI system that delivers performance at the required level is less than or equal to the baseline cost of performing the task entirely with human labor: ํถ ํ (1,ํ req,ํ ,ํ ํ ) โค ํถ ํ (0,ํ ํ ) โข Partial automation is feasible if there exists an AI system accuracy level ํ ํ โค ํ req,ํ such that the total cost of performing the task through collaboration between human workers and an AI system ํ ํ is less than or equal to the baseline cost of performing the task entirely with human labor: โ ํ ํ โค ํ req,ํ such that ํถ ํ (1,ํ ํ ,ํ ํ ) โค ํถ ํ (0,ํ ํ ) โข Full automation is optimal if full automation is feasible and the cost-minimizing accuracy ํ โ ํ satisfies: ํ โ ํ โฅ ํ req,ํ โข Partial automation is optimal if partial automation is feasible and the cost-minimizing accuracy ํ โ ํ satisfies: ํ โ ํ < ํ req,ํ Figure 1 presents illustrative marginal benefitโmarginal cost (MBโMC) and total costโtotal benefit (TCโTB) curves under three distinct automation scenarios. The horizontal axis represents improvements in AI system performance, corresponding to reductions in cross-entropy loss. Moving to the right on the axis thus indicates better AI model accuracy and performance. 46 Human-AI Collaboration: Partial vs. Full Automation To the left of the required performance threshold, according to Equation (13) and Equation (D.2), TB ํ is piecewise linear with respect to ํป AI,ํ , and hence the MB curves are piecewise constant. Panel 1 illustrates the full automation case. In this scenario, the marginal cost (MC) of improving AI performance lies below the MB curve throughout the entire interval between random-guess cross-entropy loss and the required threshold. The two curves may intersect only beyond the required level of performance. Hence, it is optimal to adopt full automation, and the AI system achieves performance equal to or exceeding that of human workers. Panel 2 represents the case of partial automation. In this case, the MB and MC curves intersect before reaching the required performance threshold. The total task workload is captured by the horizontal distance between the random- guess and required cross-entropy loss levels (denoted as segment ํ). The portion of the task completed by the AI system corresponds to the distance between the random-guess level and the AI systemโs optimal cross-entropy loss (segment ํ). The remaining portion, from the AIโs optimal performance to the required threshold, is completed by humans. The labor substitution ratio is therefore given by the ratio ํ/ํ. Panel 3 shows the no automation scenario. In this case, the total cost curve remains above the total benefit curve across the entire performance spectrum, rendering AI adoption infeasible at any level of model performance. This outcome reflects the presence of substantial costs in developing AI systems. Even if MB were to exceed MC marginally at certain levels, the fixed cost may still make adoption suboptimal. The fixed cost primarily arises from assembling and maintaining an engineering team to develop domain-specific AI models. Additional details on the cost parameters used in our analysisโincluding fixed costs of developing and deploying AI systemsโcan be found in Appendix B. For related modeling assumptions and empirical estimates of such costs, see also Svanberg et al. (2024). D.3 Aggregating Occupation-Level Benefits from Labor Saving We define the benefits of labor savings at the occupation level by aggregating benefits at the task level in specific automation scenarios. Suppose that occupation ํ comprises a set of computer vision tasks indexed by ํ. For each task ํ, we calculate the task-level benefit ํํต ํ under its optimal automation choice. Let ํ denote a particular automation condition (e.g., full automation is optimal). We identify the subset of tasks within occupation ํ that satisfy condition ํ . The aggregate benefit under condition ํ is defined as: TB ํ ํ = โ๏ธ ํโํ TB ํ This formulation allows us to calculate, for each occupation, the total benefit from labor saving across all computer visionโrelated tasks under various automation conditions. To facilitate comparative analysis, we report the following three occupation-level labor-saving scenarios in the results section: โข Full automation is both feasible and optimal; โข Both full and partial automation are feasible, but partial automation is optimal; โข Full automation is not feasible, but partial automation is both feasible and optimal. Figure 1 also illustrates how labor compensation associated with all computer visionโrelated tasks in a given occupation is distributed across different automation outcomes. The 100% reference point represents the total labor compensation that would be incurred if no AI-based automation were applied. The red segment corresponds to the share of compensation attributable to tasks that should be fully automated. The green segment reflects the labor cost savings achieved through partial automationโi.e., tasks for which AI systems substitute for part, but not all, of the human input. The grey segment represents the portion of compensation that is not substituted by AI. This residual can be further decomposed into three components: (i) the proportion of a task that is technically non-automatable (e.g., the (1โ ํฟ ํ ) 47 Human-AI Collaboration: Partial vs. Full Automation portion); (i) tasks for which the optimal decision is not to automate, despite being technically feasible; and (i) tasks that are optimally partially automated but still require human input (e.g., the residual (1โ ํ ํ (ํ โ ํ )) portion). Taken together, components (i) and (i) account for the portion of the automatable share that is not replaced by AI, which can be summarized compactly as (1โ ํ ํ (ํ โ ํ )) ํฟ ํ ; that is, within the automatable part of a task, the residual share corresponds to 1โ ํ ํ (ํ โ ํ ). E Functional Mapping from Accuracy to Cross-Entropy Loss We utilize the data from the scaling law experiment described in Section 4 to estimate the relationship between cross-entropy loss and accuracy (Equation (9)). The function form is specified as ฬ ํป ํดํผ = ํฝ 0 + ํฝ 1 ํ+ ํฝ 2 ํ 2 + ํฝ 3 ํ 3 + ํพ 1 ํ log(ํ)+ ํพ 2 (1โ ํ) log(1โ ํ) + ํฟ 1 ํ log(ํ)+ ํฟ 2 log(ํ)+ ํฟ 3 1 ํ (E.1) which is then estimated via OLS. Note that, similar to the case in the scaling law function, we allow this relationship to be adjusted based on task complexity ํ. Table E.1 presents the fitted parameters, highlighting the quantitative relationship between the two metrics in our setting. Table E.1: Fitted Coefficients for Mapping Accuracy to Cross-Entropy Loss. ํฝ 0 ํฝ 1 ํฝ 2 ํฝ 3 ํพ 1 ํพ 2 ํฟ 1 ํฟ 2 ํฟ 3 ํ 2 fitted Coefficient2.8614.22-27.1010.4811.99-1.82-0.700.61-0.620.98 Std. Error0.060.821.740.960.520.290.010.010.05 The parameters were estimated using ordinary least squares linear regression on a feature matrix constructed from the accuracy and the number of classes, with additional polynomial and logarithmic transformations to capture nonlinear relationships. After fitting the model, the covariance matrix of the estimated coefficients was computed as Cov( ห ํฝ) = ํ 2 (ํ โค ํ) โ1 , where ํ 2 is the error variance, i.e., estimated by the residual sum of squares divided by the degrees of freedom. The standard errors of the coefficients were then derived as the square roots of the diagonal elements of this covariance matrix, providing a quantification of the uncertainty in each parameter estimate. The model achieved an ํ 2 score of 0.98, indicating an excellent fit between the predicted and observed cross-entropy loss values. F Required Accuracy in Multiple Classification Tasks When ํ > 1, i.e., an O*NET task involves multiple vision classification tasks, the required accuracy for each classification task, ํ ํํํ,ํ should be discounted from the O*NET task-level required accuracy, denoted as ํ โฒ ํํํ,ํ , as follows 1โ ํ ํํํ,ํ = 1โ ํ โฒ ํํํ,ํ ํ ํ .(F.1) This is because, assuming the probability of making an error for each classification task is independent, if each classification task has an accuracy of ํ, then the O*NET task would have an accuracy of ํ ํ ํ . If ํ is close to one, then we can apply the first-order Taylor approximation: ํ ํ ํ = (1โ(1โ ํ)) ํ ํ โ 1โ ํ ํ (1โ ํ). In other words, the aggregated error rate of all the classification tasks of an O*NET task is ํ ํ times that of each classification task. 48 Human-AI Collaboration: Partial vs. Full Automation G Firm Size Estimation The firm-level analysis requires estimating the occupation-specific firm size distributions, i.e., for each occupation, what is the distribution of the number of employees with this distribution across all US firms. Formally, denote ํ ํํ ํ (ํ) as the firm size distribution for occupation ํ: ํ ํํ ํ (3) = 500 means there are 500 firms with 3 employees with occupation ํ. The detailed data for ํ ํํ ํ (ํ) are not available. However, the data of firm size distributions for each NAICS-4-digit sub-sector can be obtained from Business Dynamic Statistics 21 . In what follows, we will describe how we estimated the ํ ํํ ํ (ํ) from sub-sector firm size distributions. G.1 Estimating Occupation-Specific Firm Size Distributions Denote the subsector-level firm size distribution as ํ 4ํ ํ (ํ): ํ 4ํ ํ (100) = 500 means there are 500 firms with 100 employees working in sub-sector ํ. We made an assumption that the same NAICS-4-digit sector shares the same proportion of employees working in each occupation. Hence, the proportion of occupation ํ within sub-sectorํ can be estimated by ํ ํํ /ํ ํ , where ํ ํํ denotes the total number of employees in sub-sector ํ and occupation ํ, and ํ ํ denotes the total number of employees in sub-sector ํ. Accordingly, ํ ํํ ํ (ํ) can be estimated as ํ ํํ ํ (ํ) = โ๏ธ ํ ํ 4ํ ํ ํ ํ ํ ํํ ยท ํ .(G.1) G.2 Firm Size Imputation BDS only provides the firm size distribution at coarse bins: 1-4, 5-9, 10-19, 20-99, 100-499, 500-999, 1,000-2,499, 2,500-4,999, 5,000-9,999, and 10,000+ employees, respectively. To obtain finer-grained firm-size distributions, we adopt the firm size imputation method in (Svanberg et al. 2024). Except for the 10,000+ bin, each bin is divided into 10 sub-bins uniformly on the logarithmic scale. The number of firms within each coarse bin is evenly divided into its 10 sub-bins. For the 10,000+ bin, we perform the Zipf law extrapolation โ where we assume a linear decay, in the log-log scale, in ํ 4ํ ํ (ํ) with a slope ofโ1. The extrapolation is cut off at the largest firm size in the US. H Prompt for Computer Vision Task Characteristics You will be given a job, a specific task of the job, and a direct working activity (DWA) involved in the specific task. The definitions of task, job, and DWA come from the ONET dataset. The given activity involves or partially involves an image classification problem. Your goal is to identify the classification problem, as well as the number of classes and classification tasks involved in each decision in the given problem. ### Requirements: 1. **Show your reasoning process** and provide examples of your identified classes and tasks before giving your final answer. 2. **Output must contain exactly 6 fields**: - "thought_process" - "example_tasks" - "example_classes" - "number_of_classification_tasks" - "number_of_classes_per_task" - "proportion_of_vision_tasks" 21 https://w.census.gov/programs-surveys/bds.html 49 Human-AI Collaboration: Partial vs. Full Automation 3. **In the โthought_process โ field**, define each classification task and derive the number of tasks/classes from the definition. Rate the importance score of these vision tasks, from 1-10. - **If the DWA includes tasks that require other cognitive or motor skills to achieve the same goal, list these as alternative tasks.** - **Ensure all major non-vision-based skills involved in this DWA are considered.** - **Do not stop at one alternative task - consider multiple reasonable ways the DWA is performed without vision.** - **If no alternative tasks are directly relevant to the DWA, do not list any.** 4. **In the โexample_classes โ and โexample_tasks โ fields**, enumerate a full or partial list: - If there are **fewer than 20**, list all. - If there are **more than 20**, list a subset of 20. 5. **For โnumber_of_classification_tasks โ and โnumber_of_classes_per_task โ fields**: - Always output a **list of two integers** representing the minimum and maximum estimates. - If the estimate has no uncertainty , the minimum and maximum can be the same. - If greater than 20, provide an incomplete list of 20 examples. 6. **For โproportion_of_vision_tasks โ**: - Compute the ratio of the importance score of the vision tasks over the sum of the importance scores of all the listed tasks under this DWA in the thought process . - **Alternative tasks must be directly related to the DWA.** Do not include unrelated activities. - **Alternative tasks must primarily involve different skills applied to the same DWA goal.** - **Ensure a complete list of relevant alternative tasks is provided.** - **If no strong alternative tasks exist, vision tasks should take up 100% --- ### Examples: #### Example 1: **Job:** Light truck drivers **Task:** Inspect and maintain vehicle supplies and equipment , such as gas, oil, water , tires, lights, or brakes, to ensure that vehicles are in proper working condition . **DWA:** Inspect motor vehicles. **Expected JSON output:** โjson "thought_process": ["This activity involves checking multiple parts of the vehicle and deciding whether each part is functioning or not. Therefore , it can be cast as multiple binary classification tasks. Each binary decision involves deciding whether a specific part is working or not. The number of tasks is equal to the number of vehicle parts to be inspected. Since the number of vehicle parts ranges from 6 to 20, so does the number of tasks. The importance score for these vision tasks is **8**. Alternative tasks include: (1) Listening to engine sounds to detect issues without a visual inspection (importance score: 1), (2) Feeling tire pressure manually without relying on visual readings (importance score: 1)."], "example_tasks": ["Whether gas is in working condition", "Whether oil is in working condition", "...", "Whether washer fluid level is sufficient"], 50 Human-AI Collaboration: Partial vs. Full Automation "example_classes": ["working", "not working"], "number_of_classification_tasks": [6, 20], "number_of_classes_per_task": [2, 2], "proportion_of_vision_tasks": [8/10] // Vision tasks take 8 out of 10 importance points --- #### Example 2: **Job:** Zoologists and Wildlife Biologists **Task:** Analyze the characteristics of animals to identify and classify them. **DWA:** Examine characteristics or behavior of living organisms. **Expected JSON output:** "thought_process": ["The classification task is to classify animals based on their observed characteristics. This requires visual examination of multiple features, such as species type, coloration , body shape, and behavior. Therefore , it contains only one classification task. The number of classes should equal the number of distinct species in the inventory , which typically ranges between 20 to 500. The importance score for these vision tasks is **10**, as the entire DWA consists of examining and classifying organisms. Since there are no significant alternative tasks that do not involve vision, the proportion of vision tasks is 100% "example_tasks": ["What type of animal the target animal belongs to."], "example_classes": ["Lion", "Tiger", "Elephant", "Giraffe", "Zebra", "Kangaroo", " Panda", "Cheetah", "Leopard", "Rhinoceros", "Hippopotamus", "Crocodile", " Alligator", "Gorilla", "Orangutan", "Koala", "Sloth", "Wolf", "Fox", "Bald Eagle "], "number_of_classification_tasks": [1,1], "number_of_classes_per_task": [20,500], "proportion_of_vision_tasks": [1.0] // 100% --- #### Example 3: **Job:** Teachers , Secondary School **Task:** Evaluate and grade students โ work in class. **DWA:** Assess educational performance of students. **Expected JSON output:** "thought_process": ["The classification task involves assessing studentsโ performance , but only some of these assessments involve vision-based classification. Specifically , grading image-based problems - such as diagrams , geometric drawings , or labeled scientific illustrations - requires visual recognition and classification. Other grading activities , such as evaluating numerical problems or essay responses , rely on reasoning or language comprehension instead of visual classification and should be treated as alternative tasks. The importance score for grading image-based problems is **4**. Alternative grading tasks include: (1) Evaluating numerical responses using reasoning skills (importance score: 5), (2) Assessing essays based on language comprehension (importance score: 3), (3) 51 Human-AI Collaboration: Partial vs. Full Automation Entering grades into a digital system without requiring classification (importance score: 3)."], "example_tasks": ["Classify the accuracy of a studentโs labeled diagram in Biology", "Classify a studentโs hand-drawn geometric proof in Mathematics", "Classify a studentโs artistic composition in Art class"], "example_classes": ["Excellent", "Good", "Satisfactory", "Needs Improvement", "Fail "], "number_of_classification_tasks": [1,3], "number_of_classes_per_task": [5,10], "proportion_of_vision_tasks": [4/15] // Vision tasks take 4 out of 15 importance points --- Below is the actual task you need to label: Job: JOB Task: TASK DWA: DWA Provide the output in strict JSON format as follows: "thought_process": ["Your detailed reasoning here."], "example_tasks": ["Example task 1", "Example task 2", "...", "Example task 20"], "example_classes": ["Example class 1", "Example class 2", "...", "Example class 20"], "number_of_classification_tasks": [minimum number of tasks, maximum number of tasks], "number_of_classes_per_task": [minimum number of classes per task,maximum number of classes per task], "proportion_of_vision_tasks": ["..."] 1. The JSON object must contain exactly these 5 fields: - "thought_process" - "example_tasks" - "example_classes" - "number_of_classification_tasks" - "number_of_classes_per_task" 2. Do not include any text or explanations outside of the JSON object. I Detailed Analysis of the Scaling Law Production Function I.1 Output Elasticity Following Equation (20) We calculate the performance elasticities for data, training steps, and model size, which are given by: ํ ํท = ํํผยท ํ ํ ํท ํ ยท ํป ํกํํ ํ,ํ โ ํ ํ ยท ํผ ํท ํ + ํฝ ํ ํ + ํ ํ ํ + ํบ (I.1) 52 Human-AI Collaboration: Partial vs. Full Automation ํ ํ = ํํฝยท ํ ํ ํ ํ ยท ํป ํกํํ ํ,ํ โ ํ ํ ยท ํผ ํท ํ + ํฝ ํ ํ + ํ ํ ํ + ํบ (I.2) ํ ํ = ํํยท ํ ํ ํ ํ ยท ํป ํกํํ ํ,ํ โ ํ ํ ยท ํผ ํท ํ + ํฝ ํ ํ + ํ ํ ํ + ํบ (I.3) Table I.1: Performance Elasticity Analysis Results ScenarioPerformanceDataTraining StepsModel SizeTotal (ํ )(ํ ํท )(ํ ํ )(ํ ํ )Elasticity 2-Class Classification Task Small Scale (I)0.8040.0100.0460.0460.102 Medium Scale (I)0.9110.0090.0210.0070.037 5-Class Classification Task Small Scale (I)0.8660.0160.0350.0320.083 Medium Scale (I)0.9600.0130.0160.0060.034 10-Class Classification Task Small Scale (I)0.8600.0160.0400.0340.089 Medium Scale (I)0.9610.0120.0180.0060.036 50-Class Classification Task Small Scale (I)0.7710.0150.0800.0560.151 Medium Scale (I)0.9250.0090.0320.0120.052 100-Class Classification Task Small Scale (I)0.6950.0150.1220.0780.215 Medium Scale (I)0.8940.0070.0440.0170.068 500-Class Classification Task Small Scale (I)0.3510.0230.5440.2830.849 Medium Scale (I)0.7510.0060.1120.0450.163 1000-Class Classification Task Small Scale (I)0.0800.0883.4591.6295.176 Medium Scale (I)0.6360.0060.1880.0750.269 Notes: ํ represents the labor-substitution ratio. Scenario I (Small Scale): ํท = 25,000, ํ = 200,000, ํ = 250,000. Scenario I (Medium Scale): ํท = 100,000, ํ = 1,000,000, ํ = 5,000,000. ํท denotes data size, ํ training steps, and ํ model size. Performance elasticities are calculated at the respective input bundles. Table I.1 reveals that, for a given bundle of inputsโdata, compute, and model parametersโthe attainable labor- substitution ratio ํ differs systematically with task complexity, proxied here by the number of output classes (ํ). The direction of the relationship is shaped by two opposing forces. First, greater task complexity raises the technical difficulty of the prediction problem. Holding the input bundle fixed, a more complex task requires a model with higher effective capacity to achieve the same accuracy. Because the bundle 53 Human-AI Collaboration: Partial vs. Full Automation is not scaled up accordingly, model performance deteriorates as the number of classes ํ rises, exerting a downward pressure on ํ . Second, as ํ increases, humans must spend more time or cognitive effort to complete the task unaided. This lengthens the benchmark completion time against which the AIโs output is compared, enlarging the scope for automation. The result is an upward pressure on ํ . Between the binary (2-class) and the modestly complex (10-class) problems, the second effect marginally outweighs the first: the potential time savings for humans rises faster than model accuracy falls, so ํ edges upward. Beyond 10 classes, howeverโespecially in the jump to 1,000 classesโthe accuracy penalty dominates. Without commensurate increases in data, training steps, or parameter count, the labor-substitution ratio declines sharply with further increases in task complexity. Across the small- and medium-sized computer-vision bundles we analyse, the overwhelming majority of configurations already lie in the region of diminishing returns to scale: the sum of input elasticities is less than one. Because more complex tasks require larger quantities of data, compute, and parameters to yield comparable performance, the slide into diminishing returns occurs more gradually as task complexity risesโthe total elasticity falls toward unity at a slower pace for harder problems. The single exception is the 1 000-class task at the small-scale bundle, where total elasticity exceeds unity, indicating increasing returns. Once that inputs are upgraded to the medium-scale bundle, however, total elasticity falls below one, and the system quickly enters the diminishing-returns region. I.2 Elasticity of Substitution The pairwise (AllenโUzawa) elasticities of substitution evaluated at any point(ํท,ํ, ํ) holding the performance level fixed are ํ ํทํ = (ํ+ 1) ํํฝํ ํ + (ํ+ 1) ํํผ ํท ํ ํํฝํ ํ + ํํผ ํท ํ ,(I.4) ํ ํํ = (ํ+ 1) ํํ ํ ํ + (ํ+ 1) ํํฝํ ํ ํํ ํ ํ + ํํฝํ ํ ,(I.5) ํ ํทํ = (ํ+ 1) ํํ ํ ํ + (ํ+ 1) ํํผ ํท ํ ํํ ํ ํ + ํํผ ํท ํ .(I.6) Substituting the baseline bundle into Eqs. (I.4)โ(I.6) yields the values listed in Table 4. I.3 Equation (19) Implementation Step 1: Data Loading (main()) Three data sources are merged: โข BeyondAIExposure(Averageindustry)rev.xlsx โ survey data with task characteristics (error rates, judgment frequency, etc.) โข promptengineeringoutput.csv โ task complexity information (ํ class , numtasks) โข oesmnaicsgranularity.csv โ BLS employment and wage data at each industry level Step 2: Merging (map soctoocc()) Creates one row per (SOC Code ร Task ID ร DWA ร NAICS) combination by joining the three sources. Also computes: โข ํ class , numtasks โ from prompt engineering output 54 Human-AI Collaboration: Partial vs. Full Automation โข ํ ํ = dwa/occร importancescore โ fraction of work time spent on this visual task Step 3: Row-level Extraction (extractscenariodata()) For each row, extracts: โข accuracy = 1โ Q5 error rateโ required human-level accuracy โข ฬํ rand = Q7 random guess error rate โข ํค ํ , numemplyee, ํ ํ , ํ class , numtasks Step 4: Loss Computation Converts accuracy/error to cross-entropy loss space via entropyfn(): ฬ ํป req,ํ = entropy fn(reqerr/numtasks, ํ class ) ํป task,ํ = min ( entropyfn( ฬํ rand , ํ class ), entropyfn(1โ 1/ํ class , ํ class ) ) Step 5: Cost Structure Setup (optim fns.init()) Initializes all cost components: โข data costcoef โ annotation cost over model lifetime with retraining โข computecostcoefval โ GPU training cost โข infercostcoef โ inference cost scaled by employees โข fixed cost โ implementation and maintenance (NPV) โข ํ data โ row-specific data cost from Q8 (judgment frequencyร wage) Step 6: Scaling Law (crossentropyfn()) Maps fine-tuning inputs(ํท ํ ,ํ ํ , ํ ํ ) to predicted cross-entropy loss using a fitted scaling law: ฬ ํป AI = exp ํด ํท ํ + ํต ํ ํ + ํถ ํ ํ + ํ ร ํ ํ class This is the ํ ํ(ยท) term inside Equation (19) . Step 7: Cost Minimization (cost min()) For a given target loss, finds the cheapest(ํท ํ ,ํ ํ , ํ ํ ) combination: min ํท ํ ,ํ ํ , ํ ํ ํ data (ํท ํ )+ ํ train (ํ ํ , ํ ํ )+ ํ inf (ํ ํ ) s.t. ฬ ํป AI (ํท ํ ,ํ ํ , ํ ํ ) โค ฬ ํป req,ํ , ํ ํ โฅ ํท ํ ร complxร 1000 Solved via scipy trust-constr. Step 8: Marginal Cost and Benefit โข MB = TB ํ ํป task,ํ โ ฬ ํป req,ํ ,where TB ํ = ํ empl ร ํค ํ ร ํ ํ ร NPV factorร compratio โข MC = Lagrange multiplier from costmin(), representing ํํ /ํ ฬ ํป 55 Human-AI Collaboration: Partial vs. Full Automation Step 9: Profit Maximization (profitmax()) Three cases arise: 1. If MB(ํป task,ํ ) โค 0โ no automation, ํ ํ = 0 2. If MB( ฬ ํป req,ํ ) โฅ 0โ full automation, ํ ํ = 1 3. Otherwiseโ partial automation; find optimal loss via inversemarginalcost(): โข Solves min ฬ ํป ฬ ํป+ ํ ( ฬ ํป)/MB to find where MC = MB โข ํ ํ = ํป task,ํ โ ฬ ํป โ ํป task,ํ โ ฬ ํป req,ํ Then checks profitability: if TB ํ โ ํ var ํ โ ํ fix ํ โค 0, sets ํ ํ = 0. This implements Equation (19). Step 10: Output and Aggregation Results are saved per row to profit maximizationresultsgranularityindividualinvmc.csv with fields: Replace Ratio, Optimal Accuracy, Optimal Data/Model Size, Optimal Training Steps, Total Benefit, Vari- able Cost, and Fixed Cost. The plotting script reads these CSVs and computes weighted automation rates across occupations and industries at each granularity level: Automation Rate = ร ํ ํ ํ ร ํ empl,ํ ร ํค ํ ร ํ ํ ร ํ ํ empl,ํ ร ํค ํ 56 J Automation Rates across All Occupations at Industry Group (4d) Deployment Scale 0%10%20% Automation Rate (%) Food Science Technicians Soil and Plant Scientists Telephone Operators Office Machine Operators, Except Computer Farmers, Ranchers, and Other Agricultural Managers Nurse Practitioners Tool Grinders, Filers, and Sharpeners Fiberglass Laminators and Fabricators Probation Officers and Correctional Treatment Specialists Park Naturalists Word Processors and Typists Painting, Coating, and Decorating Workers Bill and Account Collectors Acute Care Nurses Executive Secretaries and Executive Administrative Assistants File Clerks Residential Advisors Passenger Attendants Cutting, Punching, and Press Machine Setters, Operators, and Tenders, Metal and Plastic Geological Technicians, Except Hydrologic Technicians Textile Cutting Machine Setters, Operators, and Tenders Semiconductor Processing Technicians Electro-Mechanical and Mechatronics Technologists and Technicians Commercial Pilots Receptionists and Information Clerks Printing Press Operators Food Cooking Machine Operators and Tenders Office Clerks, General Adhesive Bonding Machine Operators and Tenders Maids and Housekeeping Cleaners Audiologists Rail Car Repairers Plating Machine Setters, Operators, and Tenders, Metal and Plastic Parts Salespersons First-Line Supervisors of Police and Detectives Animal Breeders Environmental Engineers Food Service Managers Security Guards Real Estate Sales Agents Cutters and Trimmers, Hand Refuse and Recyclable Material Collectors Camera Operators, Television, Video, and Film Manicurists and Pedicurists Ophthalmologists, Except Pediatric Medical and Health Services Managers Cooks, Institution and Cafeteria Shoe and Leather Workers and Repairers Aircraft Structure, Surfaces, Rigging, and Systems Assemblers First-Line Supervisors of Housekeeping and Janitorial Workers Cutting and Slicing Machine Setters, Operators, and Tenders First-Line Supervisors of Personal Service Workers Bailiffs Environmental Science and Protection Technicians, Including Health Physicians, Pathologists Forensic Science Technicians Machinists Library Technicians Emergency Management Directors Switchboard Operators, Including Answering Service Radiation Therapists Customer Service Representatives Crushing, Grinding, and Polishing Machine Setters, Operators, and Tenders Procurement Clerks Subway and Streetcar Operators Childcare Workers Property, Real Estate, and Community Association Managers Cargo and Freight Agents Mail Clerks and Mail Machine Operators, Except Postal Service Tree Trimmers and Pruners Etchers and Engravers Athletic Trainers Inspectors, Testers, Sorters, Samplers, and Weighers Geographers Textile Bleaching and Dyeing Machine Operators and Tenders Ushers, Lobby Attendants, and Ticket Takers Hotel, Motel, and Resort Desk Clerks Retail Salespersons Airline Pilots, Copilots, and Flight Engineers Transportation, Storage, and Distribution Managers Furniture Finishers Police and Sheriff's Patrol Officers Insurance Sales Agents Family Medicine Physicians Hairdressers, Hairstylists, and Cosmetologists Railroad Brake, Signal, and Switch Operators and Locomotive Firers Nurse Midwives Intelligence Analysts Upholsterers Locker Room, Coatroom, and Dressing Room Attendants First-Line Supervisors of Food Preparation and Serving Workers Recycling and Reclamation Workers Surveying and Mapping Technicians Tire Builders Biological Technicians Dental Hygienists Animal Trainers Coaches and Scouts Conveyor Operators and Tenders Precision Agriculture Technicians Cooks, Fast Food Woodworking Machine Setters, Operators, and Tenders, Except Sawing Nurse Anesthetists Industrial Truck and Tractor Operators Tank Car, Truck, and Ship Loaders Stockers and Order Fillers Cleaners of Vehicles and Equipment Bridge and Lock Tenders Makeup Artists, Theatrical and Performance Construction and Building Inspectors Multiple Machine Tool Setters, Operators, and Tenders, Metal and Plastic Optometrists Fish and Game Wardens Food and Tobacco Roasting, Baking, and Drying Machine Operators and Tenders Aircraft Mechanics and Service Technicians First-Line Supervisors of Gambling Services Workers Lodging Managers Acupuncturists Grinding, Lapping, Polishing, and Buffing Machine Tool Setters, Operators, and Tenders, Metal and Plastic Fabric and Apparel Patternmakers Anthropologists and Archeologists Carpenters Photographic Process Workers and Processing Machine Operators Postal Service Mail Carriers Postal Service Clerks Magnetic Resonance Imaging Technologists Neurologists Cleaning, Washing, and Metal Pickling Equipment Operators and Tenders Mining and Geological Engineers, Including Mining Safety Engineers Couriers and Messengers Machine Feeders and Offbearers Cooks, Restaurant Rolling Machine Setters, Operators, and Tenders, Metal and Plastic Barbers Fire Inspectors and Investigators Cabinetmakers and Bench Carpenters Baggage Porters and Bellhops Food Batchmakers Brickmasons and Blockmasons Athletes and Sports Competitors Automotive Glass Installers and Repairers Locomotive Engineers Postal Service Mail Sorters, Processors, and Processing Machine Operators Landscape Architects Dentists, General Radiologic Technologists and Technicians Gambling Managers Cardiovascular Technologists and Technicians Radiologists Library Assistants, Clerical Microbiologists Ophthalmic Laboratory Technicians Pourers and Casters, Metal Geoscientists, Except Hydrologists and Geographers Heavy and Tractor-Trailer Truck Drivers Correctional Officers and Jailers Animal Caretakers Chemical Equipment Operators and Tenders Compliance Officers Parking Enforcement Workers Textile Winding, Twisting, and Drawing Out Machine Setters, Operators, and Tenders Gambling Change Persons and Booth Cashiers Zoologists and Wildlife Biologists Occupational Health and Safety Specialists Crane and Tower Operators Museum Technicians and Conservators Meat, Poultry, and Fish Cutters and Trimmers Print Binding and Finishing Workers Airfield Operations Specialists Private Detectives and Investigators Pest Control Workers Butchers and Meat Cutters Food Scientists and Technologists Podiatrists Farmworkers, Farm, Ranch, and Aquacultural Animals First-Line Supervisors of Non-Retail Sales Workers Forest and Conservation Workers First-Line Supervisors of Firefighting and Prevention Workers Shipping, Receiving, and Inventory Clerks Sailors and Marine Oilers Pesticide Handlers, Sprayers, and Applicators, Vegetation Hosts and Hostesses, Restaurant, Lounge, and Coffee Shop First-Line Supervisors of Retail Sales Workers Animal Control Workers Light Truck Drivers Firefighters Oral and Maxillofacial Surgeons Textile Knitting and Weaving Machine Setters, Operators, and Tenders Helpers--Production Workers Packaging and Filling Machine Operators and Tenders Metal-Refining Furnace Operators and Tenders Order Clerks First-Line Supervisors of Landscaping, Lawn Service, and Groundskeeping Workers Captains, Mates, and Pilots of Water Vessels Sewing Machine Operators Dermatologists Sawing Machine Setters, Operators, and Tenders, Wood Chiropractors Parking Attendants Skincare Specialists Packers and Packagers, Hand Flight Attendants Counter and Rental Clerks Veterinarians Foresters Bakers Agricultural Inspectors Farmworkers and Laborers, Crop, Nursery, and Greenhouse Paper Goods Machine Setters, Operators, and Tenders Crossing Guards and Flaggers Laundry and Dry-Cleaning Workers Prepress Technicians and Workers Umpires, Referees, and Other Sports Officials Diagnostic Medical Sonographers Transportation Vehicle, Equipment and Systems Inspectors, Except Aviation Chefs and Head Cooks First-Line Supervisors of Farming, Fishing, and Forestry Workers Graders and Sorters, Agricultural Products Forest Fire Inspectors and Prevention Specialists Lifeguards, Ski Patrol, and Other Recreational Protective Service Workers Weighers, Measurers, Checkers, and Samplers, Recordkeeping Cartographers and Photogrammetrists Orthodontists Gambling Surveillance Officers and Gambling Investigators Transportation Security Screeners Occupation 0.34% 0.40% 0.49% 0.62% 0.78% 0.86% 1.00% 1.05% 1.14% 1.14% 1.15% 1.20% 1.20% 1.22% 1.23% 1.26% 1.33% 1.42% 1.50% 1.54% 1.56% 1.65% 1.67% 1.69% 1.72% 1.75% 1.77% 1.79% 1.82% 1.83% 1.84% 1.84% 1.84% 1.85% 1.87% 1.88% 1.93% 1.93% 1.96% 2.01% 2.03% 2.03% 2.07% 2.10% 2.10% 2.14% 2.18% 2.18% 2.22% 2.28% 2.41% 2.45% 2.46% 2.48% 2.49% 2.55% 2.57% 2.61% 2.61% 2.62% 2.62% 2.65% 2.66% 2.66% 2.71% 2.71% 2.72% 2.79% 2.80% 2.82% 2.82% 2.85% 2.86% 2.88% 2.91% 2.92% 2.92% 2.95% 3.01% 3.02% 3.07% 3.07% 3.11% 3.16% 3.17% 3.20% 3.21% 3.21% 3.25% 3.25% 3.27% 3.28% 3.33% 3.35% 3.40% 3.42% 3.45% 3.45% 3.50% 3.51% 3.52% 3.52% 3.53% 3.56% 3.59% 3.59% 3.59% 3.63% 3.68% 3.70% 3.71% 3.77% 3.77% 3.79% 3.85% 3.90% 3.90% 3.90% 3.94% 3.95% 3.98% 4.03% 4.03% 4.06% 4.07% 4.08% 4.13% 4.13% 4.19% 4.21% 4.21% 4.23% 4.30% 4.30% 4.35% 4.36% 4.42% 4.43% 4.43% 4.44% 4.47% 4.48% 4.76% 4.86% 4.89% 4.97% 5.00% 5.11% 5.21% 5.21% 5.24% 5.29% 5.33% 5.34% 5.40% 5.51% 5.56% 5.56% 5.65% 5.66% 5.86% 5.87% 5.87% 5.90% 5.95% 6.09% 6.10% 6.19% 6.30% 6.33% 6.38% 6.39% 6.43% 6.52% 6.60% 6.64% 6.65% 6.66% 6.72% 6.75% 6.84% 6.86% 6.95% 6.99% 7.12% 7.17% 7.21% 7.21% 7.24% 7.39% 7.43% 7.56% 7.79% 7.90% 7.99% 8.04% 8.25% 8.40% 8.75% 9.07% 9.16% 9.43% 9.50% 9.57% 9.61% 9.62% 9.66% 9.87% 10.01% 10.03% 10.21% 10.31% 10.42% 10.48% 10.68% 13.05% 13.09% 13.13% 13.46% 13.73% 14.25% 14.74% 15.67% 17.27% 24.07% 0%10%20% Automation Rate (%) Automation Rates Across Occupations for Industry Groups (NAICS 4d) Deployment Scale Full Automation Rate Partial Automation Rate Automatable Not Automated Figure J.1: Automation rates across occupations at industry group (4d) deployment scale. 57