Paper deep dive
Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous Computing
Joel Klein, Rebecca Pelke, Roberto Laudani, Jan Moritz Joseph, Rainer Leupers
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/7/2026, 5:24:32 PM
Summary
This paper presents an Integer Linear Programming (ILP)-based framework for partitioning ML workloads between host CPUs and Computing-in-Memory (CIM) accelerators. It minimizes end-to-end inference latency by accounting for RRAM constraints, parallelism, and inter-device transfer costs, achieving up to 30.9x speedup on edge CPUs and 7.3x on high-performance CPUs.
Entities (12)
Relation Signals (8)
ILP Framework → minimizes → End-to-end inference latency
confidence 97% · It minimizes end-to-end inference latency under RRAM constraints
Computing-in-Memory (CIM) → executes → Matrix-Vector Multiplication
confidence 96% · Computing-in-Memory (CIM) accelerators execute Matrix-Vector Multiplications in memory
ResNet-18 → evaluatedon → ILP Framework
confidence 95% · We evaluate ResNet-18 and ResNet-50 [8] as convolution-heavy architectures
Integer Linear Programming (ILP) → optimizes → Workload Partitioning
confidence 95% · we propose an Integer Linear Programming (ILP)-based workload partitioning framework
Resistive Random Access Memory (RRAM) → constrains → CIM Accelerators
confidence 94% · existing ML workload partitioning approaches for CIM accelerators do not fully account for Resistive Random Access Memory (RRAM) constraints
Gurobi → solves → ILP Formulation
confidence 94% · All ILP instances are solved using Gurobi [19].
ARM Cortex-A72 → pairedwith → CIM Accelerator
confidence 93% · The platform comprises a single-core host CPU and a PUMA-based CIM accelerator [4] connected through a shared bus.
Design Space Exploration → analyzes → Hardware Parameters
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Computing-in-Memory (CIM) accelerators execute Matrix-Vector Multiplications (MVMs) in memory, making them a compelling solution for Machine Learning (ML) workloads. However, existing ML workload partitioning approaches for CIM accelerators do not fully account for Resistive Random Access Memory (RRAM) constraints such as limited memory, high write latency, and limited endurance. They also neglect parallelism, low-level architectural effects, or the Central Processing Unit (CPU) as a complementary compute resource. To address these limitations, we propose an Integer Linear Programming (ILP)-based workload partitioning framework for heterogeneous CPU-CIM systems. It minimizes end-to-end inference latency under RRAM constraints, captures parallelism, and combines empirical profiling with analytical models. Using our framework, heterogeneous CPU-CIM execution achieves speedups of up to 30.9x over CPU-only execution on an edge CPU and 7.3x over a high-performance CPU. A Design Space Exploration (DSE) yields further design insights for future CIM accelerators.
Tags
Links
- Source: https://arxiv.org/abs/2607.05240v1
- Canonical: https://arxiv.org/abs/2607.05240v1
Trouble viewing inline? Open PDF directly →
Full Text
30,413 characters extracted from source content.
Expand or collapse full text
1T1R One-Transistor One-Resistor ADC Analog-to-Digital Converter BIP Binary Integer Programming CIM Computing-in-Memory CNN Convolutional Neural Network CPU Central Processing Unit DAC Digital-to-Analog Converter DAG Directed Acyclic Graph DSE Design Space Exploration FPGA Field-Programmable Gate Array GeMM General Matrix Multiplication GPEU General-Purpose Execution Unit GPU Graphics Processing Unit ILP Integer Linear Programming ISA Instruction Set Architecture LP Linear Programming MAC Multiply-Accumulate ML Machine Learning MIP Mixed-Integer Programming MVM Matrix-Vector Multiplication MVMU Matrix-Vector Multiplication Unit N Neural Network ONNX Open Neural Network Exchange RAM Random Access Memory ReLU Rectified Linear Unit ResNet Residual Neural Network RRAM Resistive Random Access Memory SoC System-on-a-Chip SRAM Static Random Access Memory YOLO You Only Look Once Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous Computing Joel Klein, Rebecca Pelke, Roberto Laudani, Jan Moritz Joseph, Rainer Leupers This work was funded by the Federal Ministry of Research, Technology and Space, Germany, in the project NeuroSys I (03ZU2106CA). Abstract cim accelerators execute Matrix-Vector Multiplications in memory, making them a compelling solution for Machine Learning (ML) workloads. However, existing ML workload partitioning approaches for Computing-in-Memory (CIM) accelerators do not fully account for Resistive Random Access Memory (RRAM) constraints such as limited memory, high write latency, and limited endurance. They also neglect parallelism, low-level architectural effects, or the Central Processing Unit (CPU) as a complementary compute resource. To address these limitations, we propose an Integer Linear Programming (ILP)-based workload partitioning framework for heterogeneous CPU–CIM systems. It minimizes end-to-end inference latency under RRAM constraints, captures parallelism, and combines empirical profiling with analytical models. Using our framework, heterogeneous CPU–CIM execution achieves speedups of up to 30.9× 30.9× over CPU-only execution on an edge CPU and 7.3× 7.3× over a high-performance CPU. A Design Space Exploration (DSE) yields further design insights for future CIM accelerators. I Introduction Efficient execution of ML workloads requires new hardware architectures. fuses computation and memory to address the von Neumann bottleneck [1], executing MVMs with (1)O (1 ) latency using crossbar arrays [2, 3, 4]. is a promising CIM candidate due to its high device density, low power consumption, non-volatility, and CMOS compatibility [5, 6, 1]. However, CIM accelerators rely on a host CPU to execute unsupported operators, requiring workload partitioning between CPU and CIM accelerator [4]. Furthermore, CIM storage is limited by the available crossbars [7]. Modern Convolutional Neural Networks often exceed this capacity [8, 9, 10, 11, 7]. Moreover, RRAM write latency and endurance limits make dynamic weight replacement impractical [6, 12, 5]. Hence, deployment requires static assignment of selected operators to the CIM accelerator. Existing ML partitioning approaches [13, 14, 15, 16] treat devices as black boxes, neglect low-level architectural characteristics or assume sequential execution. -specific methods either assume full-model fit [17, 18] or rely on weight reprogramming [7, 11]. Furthermore, workload partitioning between the host CPU and the CIM accelerator remains an open problem. Fig. 1: Overview of the proposed partitioning framework We propose a system-level workload partitioning framework for heterogeneous CPU–CIM systems. It covers latency characterization and static operator assignment. The resulting partitioning between CPU and CIM achieves up to 30.9× 30.9× and 7.3× 7.3× speedup over CPU-only execution on ARM edge and x86 CPUs, respectively. Fig.˜1 illustrates the framework’s components. We highlight the following contributions: I A hybrid latency characterization methodology, integrating empirical host CPU profiling with an analytical CIM accelerator and interconnect performance model. I An ILP formulation for static operator assignment that minimizes the end-to-end inference latency accounting for limited CIM memory budget, operator support, and parallelism constraints. I A DSE using the proposed framework to analyze the impact of model architectures and hardware parameters on system performance, providing design insights for future CIM accelerators. Sections˜I and I provide background and review prior work. Section˜IV presents our framework. Section˜V reports results and a DSE. Section˜VI concludes the paper. I Background This section presents background related to RRAM-based CIM accelerators and the ILP concepts used in this work. I-A RRAM-based CIM Accelerators rram is a memristive non-volatile memory technology that stores data as resistance states [6, 5, 1]. By arranging RRAM cells in crossbar arrays, CIM accelerators perform analog MVMs via Ohm’s and Kirchhoff’s laws [2, 3, 4]. Since programming RRAM cells is slow and repeated writes degrade device endurance [6, 12], a weight-stationary dataflow, where weights are programmed once before inference, is particularly well-suited for RRAM-based CIM accelerators [2, 4]. The PUMA architecture [4] is hierarchically organized into CIM cores, connected by an internal bus. Each core contains a control unit, instruction memory, a Matrix-Vector Multiplication Unit (MVMU), a General-Purpose Execution Unit (GPEU), and a data buffer. The MVMU comprises multiple RRAM crossbars with peripheral circuits. The GPEU performs simple arithmetic and a limited set of functions, including activation functions. I-B Integer Linear Programming ilp is a class of optimization problems where decision variables are restricted to integers. The objective is to maximize or minimize a linear function with linear constraints. allows continuous variables. mixes integer and continuous variables. State-of-the-art solvers [19] combine fundamental methods such as branch-and-bound with heuristics to identify optimal or near-optimal solutions. The relative Mixed-Integer Programming (MIP) gap measures solution quality as zinc−zbound|zinc| z_inc-z_bound |z_inc |, where zincz_inc, the incumbent, is the best known feasible integer solution and zboundz_bound is the proven bound obtained from the Linear Programming (LP) relaxation. However, a large gap does not necessarily indicate a suboptimal solution, since the LP bound may be significantly looser than the true ILP optimum. I Related Work Prior research has explored ML workload partitioning for heterogeneous systems and CIM accelerators specifically. Neurosurgeon [13] splits ML execution between a mobile device and a datacenter at a layer granularity, selecting the optimal split point via per-layer prediction models. Viramontes et al. [14] extend this to a multi-device ILP formulation for edge-hub-cloud hierarchies, using weight preloading and multiple device transitions to minimize inference latency. CoEdge [15] and HiDP [16] further generalize to cooperative and hierarchical partitioning across heterogeneous edge nodes. These methods largely treat devices as black boxes, rely on high-level latency estimates, or assume sequential layer execution. Regarding CIM-specific approaches, Li et al. [17] optimize crossbar allocation by determining the per-layer weight replication factor for static mapping, and CIM-MLC [18] provides a multi-level compilation stack for operator scheduling across device, circuit, and architecture tiers. Both assume the complete model fits on the accelerator. Gao et al. [7] address the resource-constrained case by statically scheduling weight programming to minimize reprogramming overhead. COMPASS [11] addresses capacity limitations through weight partitioning, replication, and dynamic weight replacement. For RRAM-based targets, however, this degrades device endurance and introduces write latency, making reliable long-term deployment infeasible. Furthermore, none of the prior works specifically addresses joint CPU–CIM execution of ML workloads. In contrast, we target system-level CPU–CIM partitioning, combining empirical CPU profiling with analytical models. Our ILP formulation accounts for memory constraints, operator support, and parallelism, with a weight-stationary dataflow. IV Partitioning Framework Fig. 2: Overview of the heterogeneous hardware setup The framework targets a hardware as illustrated in Fig.˜2. The platform comprises a single-core host CPU and a PUMA-based CIM accelerator [4] connected through a shared bus. The CIM accelerator comprises multiple CIM cores (see Section˜I-A). IV-A Hybrid Latency Characterization To accurately characterize operator execution latencies on both devices, we adopt a hybrid approach combining empirical CPU profiling with analytical CIM and interconnect modeling. IV-A1 Host CPU Profiling The operator latency on the CPU, Thost,vT_host,v, is measured by profiling each operator using Open Neural Network Exchange (ONNX) Runtime [20, 21] on the target hardware. This empirical approach captures all architectural and optimization effects, which are difficult to model analytically, while supporting flexible replacement of the host CPU. We perform multiple iterations with warmup runs to eliminate startup effects, using the median latency for robustness. For model-level statistics, we enable all available graph and layout optimizations. IV-A2 CIM Accelerator Modeling We model the CIM accelerator latency analytically, which enables evaluation across a broad range of accelerator configurations. crossbars execute one MVM per input vector in (1)O (1 ) time [2, 3, 4]. However, limited crossbar size requires large operators to span multiple crossbars [2, 22, 23, 17]. The number of crossbars Ncrossbar,vN_crossbar,v required for an operator v is defined as: Ncrossbar,v=⌈MvM⌉⋅⌈NvN⌉,N_crossbar,v= M_vM · N_vN , (1) where MvM_v and NvN_v are the dimensions of the weight matrix for operator v, and M and N are the dimensions of a single RRAM crossbar, as illustrated in Fig.˜2. Operators without a 2D weight matrix are converted to a General Matrix Multiplication (GeMM). For convolutions, this is achieved using im2col [24, 22]. The number of MVMs Nmvm,vN_mvm,v required to execute a GeMM equals the number of input vectors applied to the crossbars. For instance, a Conv2D layer (Cin,Hin,Win)→(Cout,Hout,Wout) (C_in,H_in,W_in )→ (C_out,H_out,W_out ) with kernel size (HK,WK) (H_K,W_K ) can be expressed as Nmvm,v=Hout⋅WoutN_mvm,v=H_out· W_out MVMs with dimensions Mv=CoutM_v=C_out and Nv=Cin⋅HK⋅WKN_v=C_in· H_K· W_K. The operator latency on the accelerator, Tacc,vT_acc,v, is modeled as the sum of the load, compute, and store latency [4, 22]: Tacc,v=Tload,v+Tmvm,v+Tstore,v.T_acc,v=T_load,v+T_mvm,v+T_store,v. (2) Here, Tmvm,v=Nmvm,v⋅TmvmT_mvm,v=N_mvm,v· T_mvm, with TmvmT_mvm denoting the latency of a single MVM. The latency is independent of the number of crossbars, as these operate in parallel [4, 22]. Load and store latencies depend on the internal bandwidth BcimB_cim, element size belemb_elem, and element counts Nin,vN_in,v and Nout,vN_out,v: Tload,v=belem⋅Nin,vBcim,Tstore,v=belem⋅Nout,vBcim.T_load,v= b_elem· N_in,vB_cim, T_store,v= b_elem· N_out,vB_cim. (3) We assume the partial-sum accumulation and activation application on the GPEU are pipelined, thus not contributing to the overall latency [4]. IV-A3 Inter-device Transfer Modeling For directly data-dependent operators u and v mapped to different devices, the output of u is transferred over the shared bus. The transfer latency is: Ttransfer,(u,v)=belem⋅Nout,uBshared,T_transfer, (u,v )= b_elem· N_out,uB_shared, (4) where BsharedB_shared is the bus bandwidth between CPU and CIM. IV-B ILP-based Workload Partitioning We formulate the operator partitioning as an ILP on a Directed Acyclic Graph (DAG) G=(V,E)G=(V,E), where V is the operator set and each directed edge (u,v)∈E⊆V×V(u,v)∈ E V× V encodes a direct data dependency from u to v. Each operator v has a binary decision variable xvx_v, where xv=1x_v=1 maps to host CPU and xv=0x_v=0 maps to CIM. Only a subset of operators Vacc⊆V_acc V is supported by the accelerator: xv∈0,1,∀v∈Vacc,xv=1,∀v∈V∖Vacc.x_v∈ \0,1 \,\;\>∀ v∈ V_acc, x_v=1,\;\>∀ v∈ V V_acc. (5) The objective of the ILP is to minimize the makespan T∈ℝ≥0T _≥ 0, constrained by the finish times of all terminal nodes: T≥fv,∀v∈u∈V∣∄w∈V:(u,w)∈E.T≥ f_v, ∀ v∈ \u∈ V \,w∈ V (u,w )∈ E \. (6) The finish time fvf_v depends on the start time sv∈ℝ≥0s_v _≥ 0 and the device-specific computation latency: fv=sv+xv⋅Thost,v+(1−xv)⋅Tacc,v,∀v∈V.f_v=s_v+x_v· T_host,v+(1-x_v)· T_acc,v, ∀ v∈ V. (7) The start time svs_v is bounded by all predecessors pred(v)=u∈V∣(u,v)∈Epred (v )=\u∈ V (u,v )∈ E\ and any required transfer latency: sv≥fu+zu,v⋅Ttransfer,(u,v),∀u∈pred(v),s_v≥ f_u+z_u,v· T_transfer, (u,v ), ∀ u (v ), (8) The binary auxiliary variable zu,vz_u,v is 11 when u and v are assigned to different devices, and 0 otherwise, enforced by: zu,v≥xu−xv,zu,v≥xv−xu,∀(u,v)∈E.z_u,v≥ x_u-x_v, z_u,v≥ x_v-x_u, ∀ (u,v )∈ E. (9) IV-B1 Parallel Execution Operators without data dependencies can execute in parallel on separate crossbars or across the host and the CIM accelerator. We identify such pairs via the transitive closure G+=(V,E+)G^+=(V,E^+), where E+E^+ contains all pairs (u,v)(u,v) connected by a path in G. The parallelizable pairs are: P=u,v⊆V∣u≠v,(u,v)∉E+,(v,u)∉E+.P= \ \u,v \ V u≠ v, (u,v )∉ E^+, (v,u )∉ E^+ \. (10) To serialize pairs when both operators are assigned to the host, we introduce a binary ordering variable ou,vo_u,v for each u,v∈P\u,v\∈ P, where ou,v=1o_u,v=1 if u precedes v: sv s_v ≥fu−M⋅(3−xu−xv−ou,v), ≥ f_u-M· (3-x_u-x_v-o_u,v ), ∀u,v∈P, ∀\u,v\∈ P, (11) su s_u ≥fv−M⋅(2−xu−xv+ou,v), ≥ f_v-M· (2-x_u-x_v+o_u,v ), ∀u,v∈P. ∀\u,v\∈ P. (12) We set M to an upper bound on the inference latency, where Tacc,v≔0T_acc,v 0 for v∉Vaccv∉ V_acc: M=∑v∈Vmax(Thost,v,Tacc,v)+∑(u,v)∈ETtransfer,(u,v),M= _v∈ V (T_host,v,T_acc,v )+ _(u,v)∈ ET_transfer, (u,v ), (13) which deactivates the serializing constraint if either operator is assigned to the CIM accelerator or the order does not apply. IV-B2 Resource Constraint Finally, the available RRAM crossbar count C∈ℕC limits how many operators can be offloaded: ∑v∈Vacc(1−xv)⋅Ncrossbar,v≤C, _v∈ V_acc(1-x_v)· N_crossbar,v≤ C, (14) where Ncrossbar,vN_crossbar,v is the number of RRAM crossbars required to execute operator v on the accelerator, as defined in (1). V Results In this section, we evaluate our framework in terms of speedup, operator allocation, and ILP scalability on representative CNNs. All ILP instances are solved using Gurobi [19]. Experiments are conducted on two single-core hosts: an AMD Ryzen 9 3900X ( 3.8 , 32 DDR4) for high-performance x86 and an ARM Cortex-A72 ( 1.5 , 4 LPDDR4) for edge deployment. The CPU–CIM bus bandwidth is [per-mode=symbol]20.48 on x86 and [per-mode=symbol]10.24 on ARM. We sweep 25 , 50 , 100, and 2002550100200 crossbars and 50;100;500;1000 MVM latency. This DSE covers both hardware parameters and model architecture effects on system performance. Each crossbar stores up to [product-symbol = ×]256 x 256 weights, and the internal CIM bus bandwidth matches the shared CPU–CIM bus bandwidth. We evaluate ResNet-18 and ResNet-50 [8] as convolution-heavy architectures, MobileNetV2 [9] for depthwise separable convolutions, and YOLOv5n [10] for complex branching patterns. All models are quantized to 8 integer precision [25]. V-A CIM Acceleration Speedup 8.78.78.88.830.930.930.930.97.87.87.97.921.121.121.121.14.24.24.34.36.06.06.06.02.72.72.72.73.13.13.13.1Number of CrossbarsSpeedupARM Host Setup4.74.74.84.87.37.37.37.33.23.23.23.24.04.04.04.01.01.01.01.01.01.01.01.01.01.01.01.01.01.01.01.0Number of Crossbarsx86 Host Setup 50 100 500 1000 255010020002020404025501002000224466881010 Fig. 3: Optimal partitioning speedup for ResNet-18 across MVM latencies and crossbar counts Fig.˜3 reports ResNet-18 speedups over CPU-only execution. On ARM, the speedup reaches 30.9× 30.9× at 50 MVM latency with 100100 or more crossbars available. Even at 100 MVM latency, the speedup remains high at 21.1× 21.1×, demonstrating substantial acceleration potential for edge deployments. Beyond 100100 crossbars, speedups saturate because all profitable operators are already offloaded. On x86, the peak speedup is 7.3× 7.3× at 50 MVM latency, reflecting the stronger baseline performance of the desktop processor. For MVM latencies above 500 , offloading on x86 is no longer beneficial, because the layers can be executed faster on the CPU alone. Fig.˜4 extends the analysis to all evaluated architectures at 50 MVM latency. ResNet-18 exhibits the highest speedups, followed by ResNet-50. MobileNetV2 shows limited acceleration potential because its depthwise separable convolutions cannot be efficiently mapped to RRAM crossbars. Each depthwise filter operates on a single input channel, reducing the operation to a vector-vector multiplication that leaves most crossbar cells idle. YOLOv5n achieves peak speedups of 3.1× 3.1× on ARM and 1.9× 1.9× on x86. For most models, scaling beyond 100100 crossbars yields diminishing returns. As MVM latency increases, speedups decrease proportionally across all models. 8.78.78.88.830.930.930.930.92.22.26.46.414.614.614.914.92.12.13.53.53.93.93.93.91.41.41.91.93.13.13.13.1Number of CrossbarsSpeedupARM Host Setup4.74.74.84.87.37.37.37.32.02.02.82.83.13.13.13.11.21.21.21.21.21.21.21.21.41.41.71.71.91.91.91.9Number of Crossbarsx86 Host SetupResNet-18ResNet-50MobileNetV2YOLOv5n255010020002020404025501002000224466881010 Fig. 4: Optimal partitioning speedup at 50 MVM latency across models and crossbar counts V-B Operator Allocation Fig.˜5 shows the latency share and operator counts for YOLOv5n, which is representative of all evaluated models. 13595930165959301614233332237228218219255352238226198199235221192192235221192192Number of CrossbarsLatency ( )ARM Host Setup2316666624524524524511111112240240240239152332282362282192234848235228203203Number of Crossbarsx86 Host Setup 50 100 500 1000 2550100200010010020020030030025501002000202040406060 Fig. 5: Latency distribution between CPU (darker shade) and CIM accelerator (lighter shade) after optimal partitioning of YOLOv5n. Bar values indicate operator counts per device. On ARM with 2525 crossbars at 50 MVM latency, only 1616 of 251251 operators are offloaded, as the crossbar budget cannot accommodate more layers. Scaling to 100100 crossbars nearly quadruples the offloaded count to 5959, more than halving the total latency from 270 to 127 . Increasing to 200200 crossbars does not change the allocation, confirming that crossbar scaling saturates once all compute-intensive layers are covered. A higher MVM latency reduces offloading drastically, as it becomes less profitable. For instance, at 1000 MVM latency and 100100 crossbars, only 3333 operators are offloaded. On x86, the stronger host performance leads to more conservative offloading. At 50 MVM latency and 100100 crossbars, only 4848 operators are offloaded compared to 5959 on ARM. Furthermore, the impact of a high MVM latency is more severe. At 1000 MVM latency, offloading reduces to 66 operators for all crossbar counts. This confirms that the ILP adapts to platform and hardware parameters by balancing computation and communication costs. V-C Scalability of the ILP Approach TABLE I: ILP formulation statistics for workload partitioning on the ARM setup Model Operators Decision Variables Constraints Mean Solving Time ( ) Mean MIP Gap ( ) ResNet-18 3838 4444 204204 <<0.010.01 0.020.02 ResNet-50 7979 9191 412412 0.240.24 0.370.37 MobileNetV2 7171 7171 331331 <<0.010.01 0.070.07 YOLOv5n 251251 3,0413,041 7,0357,035 503503 6.976.97 Table˜I summarizes the ILP sizes and solver metrics on ARM. The ILPs of ResNet-18 and MobileNetV2 are solved in under 0.01 with a mean MIP gap below 0.1. The larger ResNet-50 converges in 0.24 with a mean MIP gap of 0.37. This confirms that larger linear networks remain manageable. YOLOv5n exhibits a substantial increase in problem complexity due to its heavily branched architecture. Despite having only 3.2× 3.2× more operators than ResNet-50, its branching topology leads to 33× 33× more decision variables and 17× 17× more constraints. Both arise from the complex execution orderings across parallel paths. The solver often reaches the 600 timeout, yielding solutions with a mean MIP gap of 6.97. As discussed in Section˜I-B, the LP relaxation bound may be significantly looser than the true ILP optimum, so these solutions may still be optimal or near-optimal. These results show that graph branching, rather than operator count, drives ILP complexity and solving time. VI Conclusion We presented an ILP-based ML workload partitioning framework for heterogeneous systems comprising a host CPU and an RRAM-based CIM accelerator. It combines empirical host CPU profiling with analytical CIM and bus latency models. The ILP formulation minimizes the end-to-end inference latency under RRAM memory budget and operator support constraints, exploiting both intra-accelerator and inter-device parallelism. Our experiments on ResNet-18, ResNet-50, MobileNetV2, and YOLOv5n demonstrate speedups of up to 30.9× 30.9× over an ARM edge CPU and 7.3× 7.3× over a high-performance x86 CPU. The optimal strategy consistently offloads compute-intensive convolutions. The DSE shows that the MVM latency dominates hardware sensitivity, while crossbar scaling saturates once all compute-intensive layers are covered. The ILP formulation scales well for linear and sparsely branched architectures and remains practical for all evaluated edge-relevant models. Future work could explore heuristic or hybrid methods to reduce solving time for complex branching topologies. References [1] H. Amrouch, N. Du, A. Gebregiorgis, S. Hamdioui, and I. Polian, “Towards Reliable In-Memory Computing:From Emerging Devices to Post-von-Neumann Architectures,” in 2021 IFIP/IEEE 29th International Conference on Very Large Scale Integration (VLSI-SoC), Oct. 2021, p. 1–6. [2] A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P. Strachan, M. Hu, R. S. Williams, and V. Srikumar, “ISAAC: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars,” SIGARCH Comput. Archit. News, vol. 44, no. 3, p. 14–26, Jun. 2016. [3] P. Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y. Liu, Y. Wang, and Y. Xie, “PRIME: A Novel Processing-in-Memory Architecture for Neural Network Computation in ReRAM-Based Main Memory,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), Jun. 2016, p. 27–39. [4] A. Ankit, I. E. Hajj, S. R. Chalamalasetti, G. Ndu, M. Foltin, R. S. Williams, P. Faraboschi, W.-m. W. Hwu, J. P. Strachan, K. Roy, and D. S. Milojicic, “PUMA: A Programmable Ultra-efficient Memristor-based Accelerator for Machine Learning Inference,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, ser. ASPLOS ’19. New York, NY, USA: Association for Computing Machinery, Apr. 2019, p. 715–731. [5] W. Wan, R. Kubendran, C. Schaefer, S. B. Eryilmaz, W. Zhang, D. Wu, S. Deiss, P. Raina, H. Qian, B. Gao, S. Joshi, H. Wu, H.-S. P. Wong, and G. Cauwenberghs, “A compute-in-memory chip based on resistive random-access memory,” Nature, vol. 608, no. 7923, p. 504–512, Aug. 2022. [6] F. Zahoor, T. Z. Azni Zulkifli, and F. A. Khanday, “Resistive Random Access Memory (RRAM): An Overview of Materials, Switching Mechanism, Performance, Multilevel Cell (mlc) Storage, Modeling, and Applications,” Nanoscale Research Letters, vol. 15, no. 1, p. 90–115, Apr. 2020. [7] X. Gao, H. Wang, Y. Chen, Y. Zhang, Z. Shen, and L. Ju, “Static Scheduling of Weight Programming for DNN Acceleration with Resource Constrained PIM,” ACM Trans. Embed. Comput. Syst., vol. 23, no. 6, p. 89:1–89:22, Sep. 2024. [8] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, NV, USA: Institute of Electrical and Electronics Engineers (IEEE), Jun. 2016, p. 770–778. [9] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, UT, USA: Institute of Electrical and Electronics Engineers (IEEE), Jun. 2018, p. 4510–4520. [10] G. Jocher, “Ultralytics YOLOv5,” Zenodo, 2020. [11] J. Park, J. Choe, D. Kim, and J.-J. Kim, “COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators,” in 2025 Design, Automation & Test in Europe Conference (DATE), Mar. 2025, p. 1–7. [12] Z. Swaidan, R. Kanj, J. El Hajj, E. Saad, and F. Kurdahi, “RRAM Endurance and Retention: Challenges, Opportunities and Implications on Reliable Design,” in 2019 26th IEEE International Conference on Electronics, Circuits and Systems (ICECS), Nov. 2019, p. 402–405. [13] Y. Kang, J. Hauswald, C. Gao, A. Rovinski, T. Mudge, J. Mars, and L. Tang, “Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge,” in Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, ser. ASPLOS ’17. New York, NY, USA: Association for Computing Machinery, Apr. 2017, p. 615–629. [14] R. Viramontes and A. Davoodi, “Neural Network Partitioning for Fast Distributed Inference,” in 2023 24th International Symposium on Quality Electronic Design (ISQED), Apr. 2023, p. 1–7. [15] L. Zeng, X. Chen, Z. Zhou, L. Yang, and J. Zhang, “CoEdge: Cooperative DNN Inference With Adaptive Workload Partitioning Over Heterogeneous Edge Devices,” IEEE/ACM Transactions on Networking, vol. 29, no. 2, p. 595–608, Apr. 2021. [16] Z. Taufique, A. Vyas, A. Miele, P. Liljeberg, and A. Kanduri, “HiDP:Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms,” in 2025 Design, Automation & Test in Europe Conference (DATE), Mar. 2025, p. 1–7. [17] W. Li, Y. Han, and X. Chen, “Mathematical Framework for Optimizing Crossbar Allocation for ReRAM-based CNN Accelerators,” ACM Trans. Des. Autom. Electron. Syst., vol. 29, no. 1, p. 21:1–21:24, Dec. 2023. [18] S. Qu, S. Zhao, B. Li, Y. He, X. Cai, L. Zhang, and Y. Wang, “CIM-MLC: A Multi-level Compilation Stack for Computing-In-Memory Accelerators,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ser. ASPLOS ’24, vol. 2. New York, NY, USA: Association for Computing Machinery, Apr. 2024, p. 185–200. [19] Gurobi Optimization, Llc, “Gurobi Optimizer Reference Manual,” 2026. [Online]. Available: https://w.gurobi.com [20] Onnx Developers, “ONNX: Open Neural Network Exchange,” 2026. [Online]. Available: https://onnx.ai/ [21] Onnx Runtime Developers, “ONNX Runtime,” 2021. [Online]. Available: https://onnxruntime.ai/ [22] R. Pelke, N. Bosbach, J. Cubero, F. Staudigl, R. Leupers, and J. M. Joseph, “Mapping of CNNs on multi-core RRAM-based CIM architectures,” in 2023 IFIP/IEEE 31st International Conference on Very Large Scale Integration (VLSI-SoC), Oct. 2023, p. 1–6. [23] R. Pelke, J. Klein, J. Cubero-Cascante, N. Bosbach, J. M. Joseph, and R. Leupers, “Mixed-Precision Training and Compilation for RRAM-based Computing-in-Memory Accelerators,” in 2026 Design, Automation & Test in Europe Conference (DATE). Verona, Italy: IEEE, Apr. 2026, p. 1–7. [24] K. E. Jeon, J. Rhe, and J. H. Ko, “Low-Rank Compression for IMC Arrays,” in 2025 Design, Automation & Test in Europe Conference (DATE). Lyon, France: IEEE, Mar. 2025, p. 1–7. [25] B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, p. 2704–2713.