Paper deep dive
Machine Learning-Driven Intelligent Memory System Design: From On-Chip Caches to Storage
Rahul Bera, Rakesh Nadig, Onur Mutlu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 3/22/2026, 5:10:58 AM
Summary
The paper proposes a machine learning-driven approach to memory system design, replacing static, human-designed heuristics with adaptive, data-driven policies. It introduces three specific ML-guided architectures: Pythia (reinforcement learning for on-chip cache prefetching), Hermes (perceptron learning for off-chip load prediction), and Sibyl (reinforcement learning for hybrid storage data placement). These designs demonstrate significant performance and efficiency improvements over traditional methods with modest hardware overhead.
Entities (5)
Relation Signals (3)
Pythia → usesmethod → Reinforcement Learning
confidence 100% · Pythia, a reinforcement learning-based data prefetcher
Hermes → usesmethod → Perceptron Learning
confidence 100% · Hermes, a perceptron learning-based off-chip predictor
Sibyl → usesmethod → Reinforcement Learning
confidence 100% · Sibyl, a reinforcement learning-based data placement policy
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Despite the data-rich environment in which memory systems of modern computing platforms operate, many state-of-the-art architectural policies employed in the memory system rely on static, human-designed heuristics that fail to truly adapt to the workload and system behavior via principled learning methodologies. In this article, we propose a fundamentally different design approach: using lightweight and practical machine learning (ML) methods to enable adaptive, data-driven control throughout the memory hierarchy. We present three ML-guided architectural policies: (1) Pythia, a reinforcement learning-based data prefetcher for on-chip caches, (2) Hermes, a perceptron learning-based off-chip predictor for multi-level cache hierarchies, and (3) Sibyl, a reinforcement learning-based data placement policy for hybrid storage systems. Our evaluation shows that Pythia, Hermes, and Sibyl significantly outperform the best-prior human-designed policies, while incurring modest hardware overheads. Collectively, this article demonstrates that integrating adaptive learning into memory subsystems can lead to intelligent, self-optimizing architectures that unlock performance and efficiency gains beyond what is possible with traditional human-designed approaches.
Tags
Links
- Source: https://arxiv.org/abs/2603.14583v1
- Canonical: https://arxiv.org/abs/2603.14583v1
Trouble viewing inline? Open PDF directly →
Full Text
83,999 characters extracted from source content.
Expand or collapse full text
This is an extended and updated version of an article published in IEEE Micro 2026 https://doi.org/10.1109/M.2026.3667076 Machine Learning-Driven Intelligent Memory System Design: From On-Chip Caches to Storage Rahul Bera Rakesh Nadig Onur Mutlu SAFARI Research Group, ETH Zürich, Switzerland Despite the data-rich environment in which memory systems of modern computing platforms operate, many state-of-the-art architectural policies employed in the memory system rely on static, human-designed heuristics that fail to truly adapt to the workload and system behavior via principled learning methodolo- gies. In this article, we propose a fundamentally different design approach: using lightweight and practical machine learning (ML) methods to enable adaptive, data-driven control throughout the memory hierarchy. We present three ML-guided architectural policies: (1) Pythia, a reinforcement learning-based data prefetcher for on-chip caches, (2) Hermes, a perceptron learning-based off-chip predictor for multi-level cache hierarchies, and (3) Sibyl, a reinforcement learning-based data placement policy for hybrid storage systems. Our evaluation shows that Pythia, Hermes, and Sibyl signifi- cantly outperform the best-prior human-designed policies, while incurring modest hardware overheads. Collectively, this article demonstrates that integrating adaptive learning into memory subsystems can lead to intelligent, self- optimizing architectures that unlock performance and efficiency gains beyond what is possible with traditional human-designed approaches. 1. INTRODUCTION Today’s computation is heavily bottlenecked by data. Many key applications (e.g., machine learning, graph analytics, large- scale recommendation systems, genome sequencing and analy- sis), irrespective of whether they are executed in cloud servers or mobile devices, are all data-intensive [16, 17, 25, 33–37, 69– 71, 73, 76, 131, 140–143, 145, 151, 174, 179]. They operate on large amounts of application data that overwhelm the storage and retrieval capability of high-performance memory systems employed in today’s processors, rendering memory the key performance and energy bottleneck [141, 142]. To alleviate this, architects have progressively optimized the memory system with numerous sophisticated speculative (e.g., prefetching data to on-chip caches before the processor demands it [97]) and control (e.g., placing a data block in an appropriate device in a hybrid memory/storage system [146, 175]) policies to reduce the data access overheads. Even though these architectural policies observe a vast amount of application data (as well as system-generated meta- data) during their online operation, we observe that these policies often make decisions based on simple, rigid heuristics, unable to learn in a principled manner from massive amounts of easily-available data. This is primarily because most exist- ing architectural policies make human-design-driven decisions, which predominantly rely on fixed and often myopic human- crafted heuristics that provide limited adaptability to complex system state and workload demands. For example, architects have proposed numerous data prefetching policies to predict the addresses of future memory requests and fetch their data into traditional on-chip caches before the processor demands it. Yet, most prefetching policies rely on a single piece of program context information (e.g., the program counter value of a load instruction), that a human architect selects at design time, to identify patterns in a program’s memory access stream. As a result, these prefetching policies often fail to adapt to changing workload behavior across a diverse set of workloads [27]. Our goal is to improve upon many such rigid human- designed policies found in modern state-of-the-art memory system designs by making them data-driven and fundamentally adaptive in a principled manner. To this end, we propose harnessing lightweight and practical machine learning (ML) to guide architectural decision-making, enabling intelligent, data-driven control policies throughout the memory hierarchy: from the on-chip caches to the stor- age subsystem. Specifically, we describe three ML-driven ar- chitectural policies: (1) Pythia [27], a reinforcement learning (RL)-based framework for intelligent hardware prefetching in on-chip caches, (2) Hermes [26], a perceptron learning-based mechanism for accelerating long-latency memory requests via accurate off-chip load prediction in a multi-level cache hierar- chy, and (3) Sibyl [175], an RL-based adaptive data placement policy for hybrid storage systems. Our evaluation shows that Pythia, Hermes, and Sibyl not only provide significant performance and energy improve- ments on top of a well-optimized memory hierarchy design, but also consistently outperform prior best human-designed policies across a wide range of workloads and system configu- rations - owing to their adaptive and principled data-driven decision-making. With detailed hardware design and proto- typing, we also show that all three proposed policies not only incur modest area and power overheads, but can be practically employed in today’s computing systems. Collectively, these proposals illustrate the significant po- tential of ML-driven policies for designing adaptive and self- optimizing memory systems, fundamentally transforming ar- chitectural decision-making processes. We make the following key contributions in this article: •We analyze various human-designed state-of-the-art archi- tectural policies proposed for three components in the mem- ory system: data prefetching, off-chip prediction, and data placement. We show that these policies leave a significant 1 arXiv:2603.14583v1 [cs.AR] 15 Mar 2026 performance potential on the table due to their rigid and myopic designs. •We propose three novel policies, Pythia, Hermes, and Sibyl, that employ various forms of machine learning (reinforce- ment and perceptron learning) to make adaptive and princi- pled data-driven decisions online. • We show that Pythia, Hermes, and Sibyl provide significant benefits over the best human-designed policies across a wide range of workloads and system configurations. 2.SHORTCOMINGS OF HUMAN-DESIGNED MICROARCHITECTURAL POLICIES The ever-growing data footprint of many modern workloads (and likely future workloads) greatly overwhelms the storage and retrieval capability of the memory systems of modern ma- chines [16,34,35,39,63,98,192]. As such, data access has become the key performance and energy bottleneck [140, 142, 143, 145]. In response, architects have developed numerous speculative (e.g., data prefetching) and control (e.g., data placement and management) policies to mitigate data access overheads. Al- though these policies observe a vast quantity of application data and system-generated metadata during their online oper- ation, their decision-making process often relies on simple and often myopic human-designed heuristics, failing to dynami- cally adapt to complex system states and evolving workload demands in a principled manner [25, 84, 139, 142]. As a result, they often leave a large potential for performance and energy efficiency improvement on the table [25, 142]. In this section, we motivate the need for data-driven poli- cies by examining the shortcomings of conventional human- design-driven policies through three case studies that span the memory hierarchy: from the on-chip caches to the storage system. 2.1. Case Study 1: Prefetching in On-Chip Caches Prefetching is a well-studied speculation technique that pre- dicts the addresses of long-latency memory requests and fetches their corresponding data from the main memory to on-chip caches before the processor demands it [13–15, 20, 22– 24, 29, 30, 43, 45, 47, 52–54, 60, 64, 67, 77, 80, 85, 86, 95–97, 99, 102, 103, 105, 107, 137, 153–155, 158, 170, 171, 177, 178, 180, 188, 193– 197, 208]. To identify patterns within a program’s memory access stream, a prefetcher typically correlates memory ad- dresses with program context information (also called program feature). Numerous prefetching techniques proposed in the literature consistently exhibit two primary limitations: (1) reliance on a single, static, human-designed program feature to identify memory access patterns, significantly limiting their efficacy across diverse workloads [20, 29, 30, 47, 49, 65–67, 85, 89, 97, 102, 105, 107, 137, 147, 153, 158, 170, 171, 178, 180], and (2) lack of inherent system awareness (e.g., memory bandwidth usage), resulting in performance degradation in resource-constrained systems [29, 44, 58–60, 108–112, 119, 120, 144, 180, 212]. To illustrate this, we show the coverage (i.e., the frac- tion of program memory requests correctly predicted by the prefetcher) and overpredictions (i.e., prefetched memory re- quests that are not demanded by the program) of two recently proposed prefetchers, signature path prefetcher (SPP [102]) and Bingo [23] for six example workloads in Figure 1(a) and their performance improvement in Figure 1(b). We make two key observations from this figure. First, since SPP and Bingo rely on two different yet static program features, neither of them provides the best performance benefit even across this small set of workloads. Second, even if Bingo achieves a similar prefetch coverage inLigra-CCas compared toPARSEC-Cannealwhile generating significantly lower overpredictions, Bingo degrades performance by1.9% inLigra-CCbut improves performance by6.4% inPARSEC-Canneal, over a no-prefetching baseline. This contrasting outcome is due to Bingo’s lack of awareness of memory bandwidth usage of the system. We conclude that, despite extensive research in prefetching, the static and human-designed nature of most conventional prefetching policies significantly limits their efficacy over a wide range of workloads and system configurations. 2.2.Case Study 2: Off-Chip Prediction in Multi- Level Cache Hierarchies To cater to the ever-increasing data footprints of modern work- loads, architects continue to increase the size (and also num- ber of levels) of the on-chip caches in commercial processors, which inadvertently increases on-chip cache access latencies. As a result, a large fraction of the latency of an off-chip memory request is spent accessing the on-chip cache hierarchy to solely determine that it needs to go off-chip [4, 5, 26, 87, 134, 167, 168]. To mitigate this, architects have proposed off-chip prediction, a technique that predicts which memory requests might go off- chip and speculatively fetches its data directly from the main memory [26]. However, predicting which memory requests will miss the entire cache hierarchy is challenging, especially due to the presence of complementary speculative techniques like data prefetching. Our analysis, as depicted in Figure 1(c), shows that on average only5.1% of the total loads generated by a workload go off-chip in presence of a sophisticated data prefetcher. While researchers have proposed techniques to ac- curately predict off-chip memory requests [87,159,167,168,206], these techniques either rely on a fixed set of human-crafted program features to correlate with the off-chip requests [206] or expensive hardware structures to track cacheline tags in the entire cache hierarchy [87, 159, 167, 168]. Yet, they fall short in providing accurate off-chip prediction across a wide range of workloads. As Figure 1(d) shows, the prior-best hit-miss predictor (HMP [206]) and the tag-tracking-based predictor (TTP [26, 87]) provide only47% and16.6% accuracy (i.e. the fraction of predicted off-chip loads that actually go off-chip), respectively. This shows that the current off-chip predictors have a sig- nificant room for improvement due to their reliance on fixed, human-crafted program features. 2 0% 50% 100% 150% 200% 250% SPPBingoSPPBingoSPPBingoSPPBingoSPPBingoSPPBingo 482.sphinx3-417BPARSEC-CannealPARSEC-Facesim459.GemsFDTD-765BLigra-CCLigra-PageRankDelta Fraction of baseline LLC misses CoveredUncoveredOverpredicted -20% -10% 0% 10% 20% 30% 40% 50% IPC improvement over baseline ( % ) SPPBingo (a) (b) 574%368% 0 2 4 6 8 10 12 0% 2% 4% 6% 8% 10% SPEC06SPEC17 PARSEC Ligra CVP AVG LLC misses per kilo instructions Fraction of loads that goes off - chip Off-chip rateLLC MPKI 0% 20% 40% 60% 80% SPEC06SPEC17 PARSEC Ligra CVP AVG Accuracy % HMPTTP Slow-Only CDE RNN-HSS Oracle (c)(d) (e) Performance-oriented HSS(f) Cost-oriented HSS sphinx3 417B PARSEC Canneal PARSEC Facesim GemsFDTD 765B Ligra C Ligra PageRank sphinx3 417B PARSEC Canneal PARSEC Facesim GemsFDTD 765B Ligra C Ligra PageRank 302%529% Figure 1: (a) Coverage, overprediction, and (b) performance comparison of two recently-proposed prefetchers: SPP and Bingo. (c) Percentage of loads that miss the LLC and go off-chip (on the left y-axis) and the LLC MPKI (on the right y-axis) in the baseline system with a state-of-the-art prefetcher. (d) Accuracy of two off-chip predictors, HMP and TTP. Average request latency of CDE and RNN-HSS on (e) performance-oriented and (f) cost-oriented HSS. The average request latency is normalized to Fast-Only policy. Figures adapted from our MICRO 2021 [27], MICRO 2022 [26], and ISCA 2022 [175] papers. 2.3.Case Study 3: Data Placement in Hybrid Storage Systems This case study explores the storage system, another crit- ical component of the memory hierarchy. Modern high- performance systems employ hybrid storage, combining fast- yet-small storage devices (e.g., SSDs) with large-yet-slow de- vices (e.g., HDDs) to deliver high capacity at low latency [18, 19, 21, 32, 38, 40–42, 48, 51, 55, 56, 62, 78, 100, 101, 104, 106, 113–115, 117, 118, 122–124, 126, 127, 136, 148–150, 152, 161, 164, 165, 176, 181, 184, 185, 190, 201, 203–205, 209, 210, 213]. The key challenge in designing a high-performance, cost-effective hybrid storage system (HSS) is to accurately identify the performance-critical application data and place it in the best-fit storage device [148]. Prior data placement techniques [46, 50, 57, 61, 75, 79, 81, 116, 128, 130, 132, 156, 160, 162, 163, 172, 182, 189, 191, 198–200, 202, 207, 211] propose hand-crafted policies that (i) consider only a limited number of workload characteristics (e.g., access frequency or access recency) to identify the storage device for an incoming I/O request [46, 61, 75, 116, 128, 129, 132, 138, 182, 189, 191], and (i) do not consider the variations in read/write latencies of storage devices [46, 57, 61, 75, 116, 128, 132, 162, 182, 189, 191], or the number and types of storage devices in the HSS [46, 61, 75, 116, 128, 132, 132, 133, 162, 182, 191]. As a result, such data placement techniques cannot easily adapt to the dynamic workload demands and diverse real-world hybrid storage system configurations. To demonstrate this lack of adaptivity, we show the average request latencies of two prior data-placement techniques, CDE [132] and RNN-HSS [57], for four representative workloads on a performance-oriented HSS in Figure 1(e) and cost-oriented HSS in Figure 1(f). We add three boundary scenarios to our evaluation: (1) Slow-Only, where all the data is placed in the slow storage device, (2) Fast- Only, where all the data is placed in the fast storage device, and (3) Oracle [135], which performs data placement with complete knowledge of future I/O access patterns. We make two key observations. First, CDE and RNN-HSS show a large average performance loss compared to Oracle: 41.1% (32.6%) and 34.4% (47.6%) in performance-oriented (cost- oriented) HSS, respectively. While both techniques achieve comparable performance to Oracle in select workloads (e.g., usr _ 0for CDE in cost-oriented HSS,hm _ 1for RNN-HSS in performance-oriented HSS), they show sub-optimal perfor- mance in other workloads, showing limited adaptability to workload diversity. Second, both techniques demonstrate in- consistent behavior across the two HSS configurations: in hm _ 1workload, both CDE and RNN-HSS underperform the Slow-Only in performance-oriented HSS, but outperform the Slow-Only in cost-oriented HSS. This shows the inability of prior approaches to adapt to varying storage device character- istics and system heterogeneity. We conclude that prior hand-crafted data placement tech- niques fail to adapt to diverse workload demands or changes in storage device characteristics due to their rigid design choices and limited system awareness. 3. A PRIMER ON REINFORCEMENT AND PER- CEPTRON LEARNING 3.1. What is Reinforcement Learning? Reinforcement learning (RL) [183] is the algorithmic approach to learn how to take an action in a given situation to maximize a numerical reward. A typical RL system consists of two main components: the agent and the environment, as illustrated in Fig. 2. At every timestept, the agent observes the current state of the environmentS t and takes an actionA t , for which the environment provides a rewardR t+1 . The agent’s goal is to find the optimal policy that maximizes the cumulative re- 3 ward collected from the environment over time. The expected cumulative reward by taking an actionAin a given stateS is defined as the Q-value of the state-action pair (denoted asQ(S,A)). The agent iteratively optimizes its policy in two steps: (1) updating Q-value using the reward collected from the environment, and (2) optimizing the policy using the updated Q-value. Agent Environment State(S t )State(S t )Action(A t )Action(A t )Reward(R t+1 )Reward(R t+1 ) Figure 2: Overview of a reinforcement learning system. Updating Q-Values. Prior research has proposed numerous algorithms to update the Q-values with varying computational complexities [183]. SARSA is one such algorithm that strikes a good trade-off between the learning accuracy and the com- putational complexity, which makes it suitable for online ar- chitectural decision making. In SARSA, if at a given timestep t, the agent observes a stateS t , takes an actionA t , while the environment transitions to a new stateS t+1 and emits a re- wardR t+1 and the agent takes actionA t+1 in the new state, the Q-value of the old state-action pairQ(S t ,A t )is iteratively optimized using Eq. (1): Q (S t ,A t )← Q (S t ,A t ) + α [R t+1 + γQ (S t+1 ,A t+1 )− Q (S t ,A t )] (1) Here,αis the learning rate parameter that controls the con- vergence rate of Q-values, andγis the discount factor that controls the “far-sighted" planning capability of the agent. Optimizing Policy. To maximize cumulative rewards over time, a purely greedy agent always exploits the action provid- ing the highest Q-value. However, greedy exploitation may leave the state-action space under-explored. Thus, to balance exploration and exploitation, anε-greedy agent randomly se- lects an action with a small probabilityε(exploration rate); otherwise, it chooses the action with the highest Q-value. Prior works have applied RL for various microarchitectural decision-making processes, including main memory schedul- ing [84, 139], cache replacement [121, 125, 169] and software- hinted hardware prefetching [157]. In this article, we describe Pythia, which is the first RL-based software-transparent hard- ware prefetching mechanism. 3.2. What is Perceptron Learning? Perceptron learning is a simplified learning model to mimic bio- logical neurons. Fig. 3 shows a single-layer perceptron network where each input is connected to the output via an artificial neuron. Each artificial neuron is represented by a numeric value, called weight. The perceptron network as a whole itera- tively learns a binary classification functionf(x)(shown in Eq. 2), a function that maps the inputX(a vector ofnvalues) to a single binary output. 1 x 1 x 2 x n y w 0 w 1 w 2 w n ... Perceptron inputs Perceptron output Artificial neuron Figure 3: Overview of a single-layer perceptron model. f(x) = 1if w 0 + P n i=1 w i x i > 0 0otherwise (2) The perceptron learning algorithm starts by initializing the weight of each neuron and iteratively trains the weights using each input vector from the training dataset in two steps. First, for an input vectorX, the perceptron network computes a binary output using Eq. 2 and the current weight values of its neurons. Second, if the computed output differs from the desired output for that input vector provided by the dataset, the weight of each neuron is updated. This iterative process is repeated until the error between the computed and desired output falls below a user-specified threshold. Prior research has successfully demonstrated perceptron learning for predicting branch direction [90–93, 186], branch target [68], branch confidence [12], cache reuse [94, 187], prefetch usefulness [30, 88]. In this article, we describe Hermes, which is the first work that applies perceptron learning for off-chip prediction. 4. DATA PREFETCHING USING REINFORCE- MENT LEARNING 4.1. Formulating Prefetching as an RL Agent We formulate prefetching as a reinforcement learning (RL) problem, as shown in Figure 4(a). Specifically, Pythia is an RL- agent that learns accurate, timely, and system-aware prefetch decisions by interacting with the processor and memory sub- system. At each timestep, corresponding to a new demand request, Pythia observes system state and takes a prefetch ac- tion. Each action (including not issuing a prefetch) yields a numerical reward based on the accuracy and timeliness of the prefetch request, and various system-level feedback informa- tion. Pythia’s goal is to identify an optimal policy maximizing accurate and timely prefetches, while taking system-level feed- back information into account. While Pythia’s framework can be extended to various types of system-level feedback, in this work we demonstrate Pythia using memory bandwidth usage as a feedback. State. We define the state as ak-dimensional vector of program features, each comprising of two components: (1) control-flow, and (2) data-flow. The control-flow component includes basic information like load-PC or branch-PC, and a history indicat- ing whether it is derived from the current or past requests. The data-flow component includes cacheline address, physi- cal page number, page offset, cacheline delta, and associated 4 Prefetcher Processor & Memory Subsystem Reward Prefetch from address A+offset (O) Features of memory request to address A (e.g., PC) Evaluation Queue(EQ) Demand Request 1 Assign reward to correspondingEQ entry Lookup QVStore State Vector Q-Value Store (QVStore) 2 3 5 Insert prefetch action & State-Action pair in EQ 6 Prefetch Fill A1 A2 A3 Memory Hierarchy Generate prefetch Evict EQ entry and update QVStore 4 Find the Action with max Q-Value 7 Max S1 S2 S3 S4 Set filled bit (a) (b) Figure 4: (a) Formulating prefetcher as an RL agent. (b) Overview of Pythia. Figures adapted from our MICRO 2021 paper [27]. history. Although Pythia can utilize many features, we fix the state-vector size (i.e.,k) at design time given a limited hard- ware storage budget. However, the exact set ofkfeatures is configurable online via configuration registers. Action. We define the action of the RL-agent as selecting a prefetch offset (i.e., a delta between the predicted and the demanded cacheline address) from a set of candidate prefetch offsets. To limit prefetch requests within the physical page of the triggering demand request, the list of prefetch offsets only contains values in the range of [-63,63] for a system with a traditionally-sized4KB page and64B cacheline. A zero offset is also a valid action for Pythia that signifies no prefetch request is generated. Reward. The reward structure defines the prefetcher’s objec- tive. We define five different reward levels: • Accurate and timely (R AT ) reward, assigned to an action whose corresponding prefetch address gets demanded after the prefetch fill. • Accurate but late (R AL ) reward, assigned to an action whose corresponding prefetch address gets demanded before the prefetch fill. •Loss of coverage (R CL ) reward, assigned to an action whose corresponding prefetch address is to a different physical page than the triggering demand request. • Inaccurate (R IN ) reward, assigned to an action whose cor- responding prefetch address does not get demanded in a temporal window. The reward is classified into two sub- levels: inaccurate given low bandwidth usage (R L IN ) and inaccurate given high bandwidth usage (R H IN ). •No-prefetch (R NP ) reward, assigned when Pythia decides not to prefetch. This reward level is also classified into two sub-levels: no-prefetch given low bandwidth usage (R L NP ) and no-prefetch given high bandwidth usage (R H NP ). The reward levels, in unison, provide prefetching objective to Pythia.R AT andR AL guide Pythia toward accurate and timely prefetches.R CL encourages prefetching within the same physical page as the triggering request.R IN andR NP shape Pythia’s strategy based on memory bandwidth usage. 4.2. Pythia: Design Overview Figure 4(b) shows a high-level overview of Pythia, comprising two primary hardware structures: Q-Value Store (QVStore) and Evaluation Queue (EQ). QVStore records Q-values for observed state-action pairs. EQ maintains a FIFO list of recently taken actions, where each entry contains three information: (1) the selected action, (2) corresponding prefetch address, and (3) a filled bit indicating cache fill status. Upon receiving a demand request, Pythia checks the EQ for the requested memory address ( 1 ). If present, it implies a previously issued prefetch was useful, prompting Pythia to as- sign a reward (eitherR AT orR AL ), depending on the filled bit. Next, Pythia constructs a state-vector from the demand request attributes (e.g., PC, address, cacheline delta) ( 2 ) and consults QVStore to identify the action with the highest Q-value (3). Pythia selects this action and issues the corresponding prefetch request to the cache hierarchy. Simultaneously, Pythia inserts the chosen action, its prefetch address, and state-vector into EQ (5). Actions involving no prefetch or addresses outside the current physical page are also inserted, with rewards assigned immediately. Upon eviction of an EQ entry, Pythia updates Q-values using the stored state-action pair and reward (6). When a prefetch request gets filled in cache, Pythia sets the filled bit in its EQ entry ( 7 ), marking it as timely or late. 4.2.1. Storage Overhead. Overall, Pythia requires only 25.5KB of storage per core, in which QVStore and EQ consume 24 KB and 1.5 KB, respectively. 4.3. Evaluation Methodology We use the ChampSim [72] trace-driven simulator to evaluate Pythia. We simulate an Intel Skylake [1]-like multi-core proces- sor that supports up to12cores. For single-core (multi-core) simulations, we warm up the core using100million (50mil- lion) instructions, followed by simulating the next500million (150million) instructions. We evaluate Pythia using a diverse set of150memory-intensive workload traces spanningSPEC CPU2006[10],SPEC CPU2017[11],PARSEC[31],Ligra[173], andCloudsuite[63] benchmark suites, and compare it against five state-of-the-art prior prefetchers: SPP [102], SPP+PPF [30], 5 SPP+DSPatch [29], Bingo [23], and MLOP [170]. All prefetch- ers are trained on L1-cache misses and fill prefetched lines into L2 and last-level cache (LLC). We also prototype Pythia us- ing Chisel HDL [2] to accurately estimate its area, power, and latency overhead, and incorporate them in the performance model. We open-source our artifact-evaluated implementa- tion of Pythia, and all necessary evaluation infrastructure to facilitate future research [7]. 4.4. Key Results 4.4.1. Performance Evaluation with Varying Number of Cores. Figure 5(a) illustrates the average performance im- provement across all traces for prefetchers in single-core to 12-core systems. We make two key observations. First, Pythia consistently outperforms MLOP, Bingo, and SPP across all configurations. Second, Pythia’s performance gain over all evaluated prior prefetchers increases with an increase in the number of processor cores. As core count increases, per-core memory bandwidth becomes increasingly constrained, limit- ing the effectiveness of bandwidth-unaware prefetchers. By explicitly adapting its prefetching policy based on bandwidth usage, Pythia achieves larger performance gains as core count increases. In single-core systems, Pythia outperforms MLOP, Bingo, SPP, and aggressive SPP with perceptron filtering (PPF) by3.4%,3.8%,4.3%, and1.02%, respectively. For four-core (and twelve-core) systems, Pythia exceeds MLOP, Bingo, SPP, and SPP+PPF by5.8% (7.7%),8.2% (9.6%),6.5% (6.9%), and 3.1% (5.2%), respectively. SPPBingoMLOPSPP+DSPatchSPP+PPFPythia 0.8 0.9 1 1.1 1.2 1.3 100200400800 160032006400 12800 Geomean speedup over no prefetching DRAM MTPS (in log scale) 1.1 1.15 1.2 1.25 1.3 1.35 024681012 Geomean speedup over no prefetching Number of cores (a)(b) Figure 5: Average performance improvement of prior prefetch- ers and Pythia in systems with varying (a) number of cores and (b) DRAM million transfers per second (MTPS). Figures adapted from our MICRO 2021 paper [27]. 4.4.2. Performance Evaluation with Varying DRAM Bandwidth. To evaluate Pythia for bandwidth-constrained commercial server-class processors (where each core receives only a fraction of a DRAM channel), we simulate a single- core, single-channel configuration by scaling DRAM band- width (Figure 5(b)). Pythia consistently outperforms all com- peting prefetchers across bandwidth levels ranging from 1 16 × to4×the baseline. MLOP and Bingo suffer substantial per- formance drops under low DRAM bandwidth due to overpre- diction. By trading prefetch coverage for accuracy based on bandwidth usage, Pythia outperforms MLOP, Bingo, SPP, and SPP+PPF by16.9%,20.2%,3.7%, and9.5%, respectively, in the most constrained150-MTPS configuration. Even under ample bandwidth at9600-MTPS, Pythia leads MLOP, Bingo, SPP, and SPP+PPF by 3%, 2.7%, 4.4%, and 0.8%, respectively. These results succinctly demonstrate that Pythia’s RL-based prefetching policy dynamically adapts to a wide range of work- loads and system configurations, providing consistent perfor- mance gains over many prior-best human-designed prefetchers. Additional results can be found in our MICRO 2021 paper [27]. 5.OFF-CHIP PREDICTION VIA PERCEPTRON LEARNING 5.1. Hermes: Design Overview Figure 6(a) provides a high-level overview of Hermes [26]. POPET is the key component of Hermes that accurately pre- dicts which load requests will go off-chip (i.e., access DRAM) by correlating a load request with multiple program features via perceptron learning. For each processor-generated demand load instruction, POPET predicts whether the request will go off-chip (1). If so, Hermes issues a speculative memory re- quest (called a Hermes request) directly to the main memory controller immediately after obtaining the physical address (2). This Hermes request proceeds in parallel with the regular load accessing the on-chip cache. If the prediction is correct, the regular load eventually misses in the LLC and waits for the ongoing Hermes request, effectively hiding on-chip cache access latency ( 3 ). If the Hermes request completes but no regular load targets the same address (e.g., a predicted off-chip load request hitting on-chip cache), Hermes discards it, avoid- ing unnecessary cache fills and preserving coherence. Hermes uses each retired load to train POPET for future predictions (4). 5.2. Design of the POPET Off-Chip Load Predictor POPET is composed of multiple one-dimensional weight tables, each associated with a single program feature. Each table entry holds a5-bit saturating signed integer that records the correlation between a feature value and the actual outcome (i.e., whether the load actually went off-chip). A weight value closer to+15or−16indicates strong positive or negative correlation, respectively, while values near zero indicate weak correlation. During the training (step 4 in Figure 6(a)), these weights are updated based on the true outcome to refine POPET’s predictions. 5.2.1. Making Predictions using POPET. During load queue (LQ) allocation for a core-generated load (step1in Figure 6(a)), POPET performs a binary prediction on whether the load will go off-chip, following a three-stage process shown in Fig- ure 6(b). First, POPET extracts various program features from the current load and prior requests. Second, it hashes each feature value to index into the corresponding weight table and retrieves the weight. Third, it sums all weights to compute the cumulative perceptron weight (W σ ). IfW σ exceeds the activation threshold (τ act ), POPET predicts that the load will go off-chip. The hashed feature values,W σ , and the predic- tion are stored in the LQ entry to train POPET when the load completes (step4in Figure 6(a)). 6 Core L1-D L2 LLC MC Main Memory POPET 1 2 3 4 Predict whether the load will go off-chip (a) Train POPET Regular load request missing the LLC waits for the Hermes request to finish Existing datapath New datapath Feature 1 # Weight Table 1 hash index Feature 2 # Weight Table 2 hash index Feature N # Weight Table N hash index 횺 weight 1 weight 2 weight n Activation Sum weights Predict to go off-chip . . . . . . . . (e.g., PC + offset) Stage 1Stage 2Stage 3 Issue Hermes request for predicted off-chip load (b) Figure 6: (a) Overview of Hermes. (b) Hardware stages to make a prediction by POPET. Figures adapted from our MICRO 2022 paper [26]. 5.2.2. Training POPET. POPET training begins when a de- mand load returns to the core and releases its load queue (LQ) entry (step 4 in Figure 6(a)). Loads that miss the LLC and ac- cess main memory are labeled as true off-chip requests. POPET uses this true outcome and the stored prediction in the LQ entry to update feature weights in two stages. First, the previously computedW σ is retrieved. IfW σ lies between the positive and negative training thresholds,T P andT N , respectively, training is triggered. This check prevents over-saturation of weights and enables quick adaptation to program phase changes. Second, if training proceeds, POPET retrieves each feature’s weight using the hashed indices from the LQ entry. If the true outcome is positive (off-chip), the weights are incre- mented by one; if negative, they are decremented by one. This update mechanism steers each feature weight toward the true outcome, improving prediction accuracy over time. 5.2.3. Storage Overhead. Overall, Hermes requires only4KB of storage per core, in which POPET and LQ metadata consume 3.2 KB and 0.8 KB, respectively [26]. 5.3. Evaluation Methodology We use the ChampSim [72] trace-driven simulator to evalu- ate Hermes. We faithfully model the latest-generation Intel Alder Lake performance-core with its large reorder buffer, large caches with publicly-reported on-chip cache access latencies. For single-core (multi-core) simulations, we warm up the core using100million (50million) instructions, followed by simulat- ing the next500million (100million) instructions. We evaluate Hermes using110memory-intensive workload traces span- ningSPEC CPU2006[10],SPEC CPU2017[11],PARSEC[31], Ligra[173], and commercial workloads from the 2nd data value prediction championship (CVP[8]), and compare Hermes against (1) five state-of-the-art data prefetchers (i.e., Pythia [27], Bingo [23], SPP with PPF [30,102], MLOP [170], and SMS [178]), and (2) two prior off-chip prediction mechanisms: hit/miss predictor (HMP [206]) and cacheline tag-tracking predictor (TTP [26, 87]). We also evaluate Hermes using a range of la- tencies for issuing a Hermes request directly to the memory controller (step2in Figure 6(a)) to account for diverse on-chip interconnect designs. We open-source our artifact-evaluated implementation of Hermes, and all necessary evaluation in- frastructure to facilitate future research [3]. 5.4. Key Results 5.4.1. Performance Evaluation with Different Off-Chip Predictors. Figure 7(a) shows the performance of Hermes with POPET, HMP, TTP, and an ideal off-chip predictor, with and without an underlying baseline prefetcher (in this case, Pythia) and normalized to a no-prefetching system for single-core workloads. We highlight two key observations. First, Hermes with POPET outperforms Hermes-HMP and Hermes-TTP. This is due to POPET’s high accuracy off-chip prediction (77.1% on average) as compared to HMP (47%) and TTP (16.6%). ∗ On average, Hermes-HMP, Hermes-TTP, and Hermes with POPET improve performance over Pythia by0.8%,1.7%, and 5.4%, respectively. Second, Hermes-POPET achieves nearly 90% of the performance gain of the ideal off-chip predictor that assumes perfect prediction. These results show that POPET’s high prediction accuracy and coverage are central to Hermes’s effectiveness, underscoring the importance of a well-designed off-chip predictor. 1.0 1.1 1.2 1.3 1.4 HMPTTPPOPETIdeal Geomean speedup over the no - prefetching system Off-chip predictor types Off-chip predictor-only Pythia + Off-chip predictor 1.0 1.1 1.2 1.3 1.4 PythiaBingoSPPMLOPSMS Geomean speedup over the no - prefetching system Prefetcher types Prefetcher-only Prefetcher + Hermes (a)(b) Figure 7: Performance comparison by varying (a) off-chip pre- dictors and (b) the underlying prefetcher. Figures adapted from our MICRO 2022 paper [26]. 5.4.2. Performance Evaluation with Different Underly- ing Prefetchers. We evaluate Hermes naively-combined (i.e., without any implicit coordination) with five state-of-the-art ∗ Our MICRO 2022 paper [26] further shows that POPET’s high-accuracy off- chip prediction incurs minimal overhead: increasing main memory requests by 5.5% and dynamic power by 3.6% over no prefetching. 7 data prefetchers at last-level cache. Figure 7(b) shows the per- formance normalized to a no-prefetching single-core baseline. Hermes combined with any prefetcher consistently outper- forms the prefetcher alone. Specifically, Hermes improves per- formance over Bingo, SPP, MLOP, and SMS by6.2%,5.1%,7.6%, and7.7%, respectively. Our recent work, Athena [28], shows that an RL-based coordination policy to synergize off-chip pre- dictor with prefetcher can provide even more performance improvement than their naive combination. These results suggest that by employing a learning-based, adaptive off-chip prediction mechanism, Hermes consistently outperforms prior off-chip predictors across diverse workloads and system configurations. More results can be found in our MICRO 2022 paper [26]. 6. DATA PLACEMENT IN HYBRID STORAGE SYSTEMS USING REINFORCEMENT LEARN- ING We propose Sibyl [175], a reinforcement learning (RL)-based data placement technique for hybrid storage systems. While RL is a promising alternative to existing data placement tech- niques, its effectiveness depends on how the data placement problem is cast as an RL task. 6.1.Formulating Data Placement as an RL Problem Figure 8(a) shows our formulation of data placement as an RL problem. We design Sibyl as an RL agent that learns to per- form accurate and system-aware data placement decisions by interacting with the hybrid storage system. For each storage request, Sibyl observes multiple workload and system-level features as a state to make a placement decision. After every action, Sibyl receives a reward based on the I/O request latency, which reflects both the data placement decision and the HSS state. Sibyl aims to learn an optimal data placement policy that improves performance under the current workload and system configuration by minimizing the average request la- tency. This involves maximizing the use of the fast storage device while avoiding the eviction penalty of non-performance- critical pages. State. At each time-stept, the state features for a particular read/write request are collected in an observation vector. We perform feature selection to determine the best state features to include in Sibyl’s observation vector. We use a limited number of features to reduce the implementation overhead of our mechanism. Sibyl’s observation vector is a 6-dimensional tuple: O t = (size t ,type t ,intr t ,cnt t ,cap t ,curr t ).(3) Table 1 lists our six selected features. We quantize the state features into a small number of bins to reduce the storage overhead. The current request size (size t ) helps distinguish se- quential from random access patterns. The request type (type t ) differentiates between read and write requests, and captures the asymmetry in read and write latencies. The access interval (intr t ) and access count (cnt t ) represent the temporal and spa- tial locality, respectively.cap t tracks the available capacity in the fast storage, guiding the agent to manage the fast storage capacity while avoiding eviction penalties effectively.curr t is the current placement of the requested page, enabling Sibyl to perform past-aware decisions. Table 1: State features used by Sibyl FeatureDescription# of binsEncoding (bits) size t Size of the requested page (in pages)88 type t Type of the current request (read/write)24 intr t Access interval of the requested page648 cnt t Access count of the requested page648 cap t Remaining capacity in the fast storage device88 curr t Current placement of the requested page (fast/slow)24 Action. At each time-stept, in a given state, Sibyl selects an action from all possible actions. In a hybrid storage system with two devices, possible actions are: placing data in (1) the fast storage device or (2) the slow storage device. This is easily extensible to a large number of storage devices. Reward. After every data placement decision at time-stept, † Sibyl gets a reward from the environment at time-stept + 1, serving as feedback for the previous action. We craft Sibyl’s reward function R as follows: R = 1 L t if no eviction of a page from the fast storage to the slow storage max(0, 1 L t − R p ) in case of eviction (4) whereL t andR p represent the latency of the last served re- quest and eviction penalty, respectively. If the fast storage is running out of free space, Sibyl evicts non-performance- critical pages to the slow storage. To discourage excessive use of fast storage, we incorporate an eviction penalty (R p ) into the reward to guide Sibyl to place only performance-critical pages in the fast storage. We empirically setR p to 0.001×L e (L e is the time required to evict pages from fast to slow stor- age), which prevents the agent from excessively using the fast storage device. Sibyl’s reward is a function of the request latency because it captures the state of the hybrid storage system, as it signifi- cantly varies depending on the request type, device type, and the internal state and characteristics of the device (e.g., read- /write latencies, garbage collection latency, queuing delays, and error handling latencies). 6.2. Sibyl: Design Overview We implement Sibyl in the storage management layer of the host system’s operating system. As shown in Figure 8(b), Sibyl is composed of two parallel threads: (1) the RL decision thread, which performs data placement4of the current I/O request while collecting experiences7(i.e., actions4and rewards 6 ) in an experience buffer 5 , and (2) the RL training thread, which uses the collected experiences ‡ 8 to update its decision- making policy online9. Sibyl continuously learns from its † In HSS, a time-step is defined as a new storage request. ‡ Experience is a representation of a transition from one time step to another, in terms of ⟨State, Action, Reward, NextState⟩. 8 Hybrid Storage System Sibyl Features of the current request and system Request latency (of last served request) Select storage device to place the current page (a) Storage Request (from OS) Inference Network Max HSS Observation Vector State State Action Reward RL Decision Thread Sibyl Policy Training Network RL Training Thread Batch Training Dataset Periodic Policy Weight Update Experience Buffer Collect Experiences 1 2 3 4 5 6 7 8 9 10 (b) Figure 8: (a) Formulation of data placement as an RL Problem. (b) Design of Sibyl. Figures adapted from our ISCA 2022 paper [175]. past decisions and their impact. Our two-threaded decoupled implementation ensures that policy training does not interfere with latency-critical placement decisions. To enable parallel execution, we use two separate identical neural networks: an inference network 2 for making placement decisions, and a training network9for background learning. Sibyl period- ically copies the training network weights to the inference network 10 to adapt the policy for current workload and system conditions. For every new storage request to the HSS, Sibyl uses the state information1to make a data placement decision4. The inference network predicts the Q-value for each available action given the state information. Sibyl’s policy3selects the action with the maximum Q-value and accordingly performs the data placement (and potentially eviction). 6.3. Evaluation Methodology We evaluate Sibyl on a real system using two dual HSS config- urations: (1) performance-oriented HSS: high-end device (H) [82] and middle-end device (M) [83], and (2) cost-oriented HSS: high- end device (H) [82] and low-end device (L) [166]. The high-end, middle-end, and low-end devices have a sequential read (write) throughput of2.4(2) GB/s,550(510) MB/s, and520(450) MB/s, respectively. The HSS devices appear to the OS as one contigu- ous logical block address space. We implement a lightweight custom block driver to orchestrate I/O requests to the stor- age devices in HSS. We develop Sibyl using the TF-Agents framework [74]. We use fourteen diverse storage traces that have distinct I/O-access patterns (i.e., randomness and hotness properties) from Microsoft Research Cambridge (MSRC) [6], collected on real enterprise servers. We compare Sibyl against two heuristic-based (i.e., cold data eviction (CDE [132]) and history-based page selection (HPS [135])) and two supervised- learning-based (i.e., Archivist [162] and RNN-HSS [57]) HSS data placement techniques. We compare the above policies with three boundary scenarios: (1) Slow-Only, where all data resides in the slow storage, (2) Fast-Only, where all data re- sides in the fast storage, and (3) Oracle [135], which exploits complete knowledge of future I/O-access patterns to perform optimal data placement. We open-source Sibyl to facilitate future research [9]. 6.4. Key Results Figure 9 compares the average request latency of Sibyl against the baseline policies for performance-oriented (Figure 9(a)) and cost-oriented (Figure 9(b)) HSS configurations. All values are normalized to Fast-Only. We observe that Sibyl consistently outperforms all the baselines for all the evaluated workloads in both HSS configurations. In the performance-oriented HSS (Figure 9(a)), where the latency difference between two de- vices is relatively smaller than that of cost-oriented HSS, Sibyl improves average performance by 28.1%, 23.2%, 36.1%, and 21.6% over CDE, HPS, Archivist, and RNN-HSS, respectively. In the cost-oriented HSS (Figure 9(b)), where there is a large difference between the latencies of the two storage devices, Sibyl improves performance by 19.9%, 45.9%, 68.8%, and 34.1% over CDE, HPS, Archivist, and RNN-HSS, respectively. The larger the latency disparity, the greater the benefits of placing only performance-critical pages in the fast storage to avoid evictions. We evaluate Sibyl on Tri-HSS with three heterogeneous storage devices. Compared to the state-of-the-art heuristic- based policy, Sibyl improves performance by up to 48.2% on average, which demonstrates that Sibyl’s adaptive placement policy generalizes to more devices in an HSS. Extending Sibyl to support additional devices requires only minimal changes to the state and action space, highlighting its ease of extensibility. RNN-HSS Sibyl Oracle Slow-Only CDE HPS Archivist (a) Performance-Oriented HSS (b) Cost-Oriented HSS Figure 9: Average request latency under two different hy- brid storage configurations (normalized to Fast-Only). Figures adapted from our ISCA 2022 paper [175]. We conclude that Sibyl’s RL-based data placement technique consistently outperforms prior best data placement techniques across diverse workloads and different HSS configurations. Sibyl’s workload- and system-aware online learning policy 9 enables it to adapt dynamically to different conditions, making it a generalizable solution for hybrid storage systems. More results can be found in our ISCA 2022 paper [175]. While Sibyl manages the data placement in HSS, it does not address the data migration between different storage devices of an HSS. Our recent proposal, Harmonia [146], demonstrates the first multi-agent RL-based framework for HSS to manage both data placement and data migration in a synergistic way. 7.KEY TAKEAWAYS AND FUTURE OPPORTU- NITIES This article presents our recent efforts on machine learning (ML)-driven intelligent memory system design. We present three case studies, Pythia, Hermes, and Sibyl, that employ reinforcement-learning and perceptron-learning to enable adaptive, data-driven decision-making across the multi-level cache hierarchies and hybrid storage systems. Through exten- sive studies we show that these ML-driven policies consistently outperform prior best human-designed approaches with mod- est overheads, highlighting the practicality of online learning in real systems. These case studies also reveal broader opportu- nities, and challenges, in how such mechanisms are integrated and extended across the memory hierarchy. Better Integration and Coordination of ML-Driven Mech- anisms. Although Pythia, Hermes, and Sibyl improve perfor- mance independently and in combination (e.g., Hermes with Pythia), their interactions expose additional opportunities for deeper system-level integration. We present here two concrete examples of how synergizing these mechanisms can unlock further benefits. First, while integrating Hermes with Pythia improves performance on average, we observe that a naive integration often fails to fully realize their combined perfor- mance potential. This motivates designing holistic coordina- tion mechanisms—potentially ML-driven—that jointly coordi- nate prefetching and off-chip prediction mechanisms. Recently, we have proposed Athena [28], that exploits online reinforce- ment learning to coordinate off-chip prediction with multiple prefetchers employed throughout the cache hierarchy. Second, while Sibyl manages data placement within HSS, it does not address the complementary problem of data migration across the heterogeneous storage devices that constitute an HSS. To bridge this gap, our recent proposal, Harmonia [146], intro- duces the first multi-agent RL framework that jointly orches- trates data placement and inter-device data migration, enabling coordinated decision making in an HSS. By allowing multiple learning agents to collaborate and adapt to dynamic workload characteristics, Harmonia manages these two interdependent tasks synergistically, improving the overall performance of an HSS. Extending ML-Driven Design Principles across the Mem- ory Hierarchy. The ML-driven design principles presented in this article can be broadly extended for many data-driven decision-making processes across the memory hierarchy, in- cluding, but not limited to, (1) co-optimization of caching, prefetching, and memory scheduling mechanisms, (2) co- optimization of thread and memory scheduling decisions, (3) data placement and migration in disaggregated memory sys- tems, (4) coordinated data placement and migration in hybrid storage systems. These problems exhibit complex, non-linear interactions between workload behavior and system state, mak- ing them well suited for lightweight, online ML approaches. Collectively, these directions highlight the potential of ML- driven approaches to enable adaptive, self-optimizing memory systems that deliver performance and efficiency gains beyond traditional designs, and we hope this article inspires further research in intelligent ML-driven memory system design. Acknowledgments We thank all SAFARI Research Group members for providing a stimulating, inclusive, intellectual and scientific environment. We acknowledge the generous gifts from our funding partners: Futurewei, Google, Huawei, Intel, Microsoft, and VMware. This work is supported in part by the Semiconductor Research Corporation and the ETH Future Computing Laboratory. References [1] “6th Generation Intel® Processor Family,” https://w.intel.com/content/w/ us/en/processors/core/desktop-6th-gen-core-family-spec-update.html. [2] “Chisel/FIRRTL Hardware Compiler Framework,” https://w.chisel-lang.org. [3] “Hermes GitHub Repository,” https://github.com/CMU-SAFARI/Hermes. [4] “Intel Core i5-12600K DDR4 Alder Lake CPU Review,” https://w.thefpsreview. com/2021/12/08/intel-core-i5-12600k-ddr4-alder-lake-cpu-review/6/. [5]“L3 Cache Latency Comparison at Base Frequency,” https://w.cpuagent.com/ cpu/intel-core-i9-10900k/benchmarks/l3-cache-latency-at-base-frequency/ nvidia-geforce-rtx-2080-ti?res=1&quality=ultra. [6] “MSR Cambridge Traces., http://iotta.snia.org/traces/388.” [7] “Pythia GitHub Repository,” https://github.com/CMU-SAFARI/Pythia. [8] “Second Championship Value Prediction (CVP-2),” https://w.microarch.org/ cvp1/cvp2/rules.html. [9] “Sibyl GitHub Repository,” https://github.com/CMU-SAFARI/Sibyl. [10] “SPEC CPU 2006,” https://w.spec.org/cpu2006/. [11] “SPEC CPU 2017,” https://w.spec.org/cpu2017/. [12] H. Akkary, S. T. Srinivasan, R. Koltur, Y. Patil, and W. Refaai, “Perceptron-Based Branch Confidence Estimation,” in HPCA, 2004. [13]L. M. AlBarakat, P. V. Gratz, and D. A. Jiménez, “MTB-Fetch: Multithreading Aware Hardware Prefetching for Chip Multiprocessors,” IEEE CAL, 2018. [14]L. M. AlBarakat, P. V. Gratz, and D. A. Jiménez, “SB-Fetch: Synchronization Aware Hardware Prefetching for Chip Multiprocessors,” in ICS, E. Ayguadé, W. W. Hwu, R. M. Badia, and H. P. Hofstee, Eds., 2020. [15] L. M. AlBarakat, P. V. Gratz, and D. A. Jiménez, “SLAP-C: Set-Level Adaptive Prefetching for Compressed Caches,” in ICCD, 2022. [16] M. Alser, Z. Bingöl, D. S. Cali, J. Kim, S. Ghose, C. Alkan, and O. Mutlu, “Accelerat- ing Genome Analysis: A Primer on an Ongoing Journey,” IEEE Micro, 2020. [17] M. Alser, J. Lindegger, C. Firtina, N. Almadhoun, H. Mao, G. Singh, J. Gomez-Luna, and O. Mutlu, “From Molecules to Genomic Variations: Accelerating Genome Analysis via Intelligent Algorithms and Architectures,” CSBJ, 2022. [18]R. Appuswamy, D. C. van Moolenbroek, and A. S. Tanenbaum, “Cache, Cache Everywhere, Flushing All Hits Down the Sink: On Exclusivity in Multilevel, Hybrid Caches,” in MSST, 2013. [19]S. H. Baek and K.-W. Park, “A Fully Persistent and Consistent Read/Write Cache Using Flash-Based General SSDs for Desktop Workloads,” in ICEIS, 2016. [20]J.-L. Baer and T.-F. Chen, “An Effective On-chip Preloading Scheme to Reduce Data Access Penalty,” in SC, 1991. [21]K. A. Bailey, P. Hornyack, L. Ceze, S. D. Gribble, and H. M. Levy, “Exploring Storage Class Memory with Key Value Stores,” in SOSP, 2013. [22]M. Bakhshalipour, P. Lotfi-Kamran, and H. Sarbazi-Azad, “Domino Temporal Data Prefetcher,” in HPCA, 2018. [23]M. Bakhshalipour, M. Shakerinava, P. Lotfi-Kamran, and H. Sarbazi-Azad, “Bingo Spatial Data Prefetcher,” in HPCA, 2019. [24]M. Bekerman, S. Jourdan, R. Ronen, G. Kirshenboim, L. Rappoport, A. Yoaz, and U. Weiser, “Correlated Load-Address Predictors,” ISCA, 1999. [25]R. Bera, “Mitigating the Memory Bottleneck with Machine Learning-Driven and Data-Aware Microarchitectural Techniques,” Ph.D. dissertation, ETH Zurich, 2025. [26]R. Bera, K. Kanellopoulos, S. Balachandran, D. Novo, A. Olgun, M. Sadrosadati, and O. Mutlu, “Hermes: Accelerating Long-Latency Load Requests via Perceptron- Based Off-Chip Load Prediction,” in MICRO, 2022. 10 [27]R. Bera, K. Kanellopoulos, A. Nori, T. Shahroodi, S. Subramoney, and O. Mutlu, “Pythia: A Customizable Hardware Prefetching Framework Using Online Rein- forcement Learning,” in MICRO, 2021. [28]R. Bera, Z. Lang, C. Hengartner, K. Kanellopoulos, R. Kumar, M. Sadrosadati, and O. Mutlu, “Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning,” in HPCA, 2026. [29]R. Bera, A. V. Nori, O. Mutlu, and S. Subramoney, “DSPatch: Dual Spatial Pattern Prefetcher,” in MICRO, 2019. [30]E. Bhatia, G. Chacon, S. Pugsley, E. Teran, P. V. Gratz, and D. A. Jiménez, “Perceptron-Based Prefetch Filtering,” in ISCA, 2019. [31]C. Bienia, S. Kumar, J. P. Singh, and K. Li, “The PARSEC Benchmark Suite: Charac- terization and Architectural Implications,” in PACT, 2008. [32]T. Bisson and S. A. Brandt, “Reducing Hybrid Disk Write Latency with Flash- Backed I/O Requests,” in MASCOTS, 2007. [33]A. Boroumand, “Practical Mechanisms for Reducing Processor-Memory Data Movement in Modern Workloads,” Ph.D. dissertation, Carnegie Mellon University, 2020. [34]A. Boroumand, S. Ghose, B. Akin, R. Narayanaswami, G. F. Oliveira, X. Ma, E. Shiu, and O. Mutlu, “Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks,” in PACT, 2021. [35]A. Boroumand, S. Ghose, Y. Kim, R. Ausavarungnirun, E. Shiu, R. Thakur, D. Kim, A. Kuusela, A. Knies, P. Ranganathan, and O. Mutlu, “Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks,” in ASPLOS, 2018. [36]A. Boroumand, S. Ghose, G. F. Oliveira, and O. Mutlu, “Polynesia: Enabling High- Performance and Energy-Efficient Hybrid Transactional/Analytical Databases with Hardware/Software Co-Design,” in ICDE, 2022. [37]A. Boroumand, S. Ghose, M. Patel, H. Hassan, B. Lucia, R. Ausavarungnirun, K. Hsieh, N. Hajinazar, K. T. Malladi, H. Zheng, and O. Mutlu, “CoNDA: Efficient Cache Coherence Support for Near-Data Accelerators,” in ISCA, 2019. [38]K. Bu, M. Wang, H. Nie, W. Huang, and B. Li, “The Optimization of the Hierarchical Storage System Based on the Hybrid SSD Technology,” in ISDEA, 2012. [39]D. S. Cali, G. S. Kalsi, Z. Bingöl, C. Firtina, L. Subramanian, J. S. Kim, R. Ausavarung- nirun, M. Alser, J. Gomez-Luna, A. Boroumand, A. Nori, A. Scibisz, S. Subramoney, C. Alkan, S. Ghose, and O. Mutlu, “GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis,” in MICRO, 2020. [40]M. Canim, G. A. Mihaila, B. Bhattacharjee, K. A. Ross, and C. A. Lang, “SSD Bufferpool Extensions for Database Systems,” in VLDB, 2010. [41] Y. Chai, Z. Du, X. Qin, and D. A. Bader, “WEC: Improving Durability of SSD Cache Drives by Caching Write-Efficient Data,” in TC, 2015. [42] H.-P. Chang, S.-Y. Liao, D.-W. Chang, and G.-W. Chen, “Profit Data Caching and Hybrid Disk-Aware Completely Fair Queuing Scheduling Algorithms For Hybrid Disks,” in SPE, 2015. [43]M. Charney, “Correlation-Based Hardware Prefetching,” Ph.D. dissertation, Cornell University, 1995. [44] M. J. Charney and T. R. Puzak, “Prefetching and Memory System Behavior of the SPEC95 Benchmark Suite,” IBM Journal of Research and Development, 1997. [45] M. J. Charney and A. P. Reeves, “Generalized Correlation-Based Hardware Prefetch- ing,” Cornell Univ., Tech. Rep., 1995. [46] F. Chen, D. A. Koufaty, and X. Zhang, “Hystor: Making the Best Use of Solid State Drives in High Performance Storage Systems,” in SC, 2011. [47] T.-F. Chen and J.-L. Baer, “Effective Hardware-Based Data Prefetching for High- Performance Processors,” in IEEE TC, 1995. [48]X. Chen, W. Chen, Z. Lu, P. Long, S. Yang, and Z. Wang, “A Duplication-Aware SSD-Based Cache Architecture for Primary Storage in Virtualization Environment,” in ISJ, 2015. [49] Z. Chen, C. Wu, Y. Gu, R. Jia, J. Li, and M. Guo, “Gaze into the Pattern: Char- acterizing Spatial Patterns with Internal Temporal Correlations for Hardware Prefetching,” in HPCA, 2025. [50]P. Cheng, Y. Lu, Y. Du, Z. Chen, and Y. Liu, “Optimizing Data Placement on Hierarchical Storage Architecture via Machine Learning,” in NPC, 2019. [51]Y. Cheng, W. Chen, Z. Wang, X. Yu, and Y. Xiang, “AMC: An Adaptive Multi-Level Cache Algorithm in Hybrid Storage Systems,” in CCPE, 2015. [52]T. M. Chilimbi and M. Hirzel, “Dynamic Hot Data Stream Prefetching for General- Purpose Programs,” in PLDI, 2002. [53]Y. Chou, “Low-cost Epoch-based Correlation Prefetching for Commercial Applica- tions,” in MICRO, 2007. [54]R. Cooksey, S. Jourdan, and D. Grunwald, “A Stateless, Content-Directed Data Prefetching Mechanism,” ASPLOS, 2002. [55]N. Dai, Y. Chai, Y. Liang, and C. Wang, “ETD-Cache: An Expiration-Time Driven Cache Scheme to Make SSD-Based Read Cache Endurable and Cost-Efficient,” in CF, 2015. [56]J. Do, D. Zhang, J. M. Patel, D. J. DeWitt, J. F. Naughton, and A. Halverson, “Tur- bocharging DBMS Buffer Pool Using SSDs,” in SIGMOD, 2011. [57]T. D. Doudali, S. Blagodurov, A. Vishnu, S. Gurumurthi, and A. Gavrilovska, “Kleio: A Hybrid Memory Page Scheduler with Machine Intelligence,” in HPDC, 2019. [58]E. Ebrahimi, C. J. Lee, O. Mutlu, and Y. N. Patt, “Prefetch-aware Shared Resource Management for Multi-core Systems,” in ISCA, 2011. [59]E. Ebrahimi, O. Mutlu, C. J. Lee, and Y. N. Patt, “Coordinated Control of Multiple Prefetchers in Multi-Core Systems,” in MICRO, 2009. [60]E. Ebrahimi, O. Mutlu, and Y. N. Patt, “Techniques for Bandwidth-Efficient Prefetch- ing of Linked Data Structures in Hybrid Prefetching Systems,” in HPCA, 2009. [61]A. Elnably, H. Wang, A. Gulati, and P. J. Varman, “Efficient QoS for Multi-Tiered Storage Systems,” in HotStorage, 2012. [62]W. Felter, A. Hylick, and J. Carter, “Reliability-Aware Energy Management for Hybrid Storage Systems,” in MSST, 2011. [63]M. Ferdman, A. Adileh, O. Kocberber, S. Volos, M. Alisafaee, D. Jevdjic, C. Kaynak, A. D. Popescu, A. Ailamaki, and B. Falsafi, “Clearing the Clouds: A Study of Emerging Scale-out Workloads on Modern Hardware,” ASPLOS, 2012. [64]M. Ferdman and B. Falsafi, “Last-touch Correlated Data Streaming,” in ISPASS, 2007. [65]M. Ferdman, S. Somogyi, and B. Falsafi, “Spatial Memory Streaming with Rotated Patterns,” in In 1st JILP Data Prefetching Championship, 2009. [66]J. W. C. Fu and J. H. Patel, “Data Prefetching in Multiprocessor Vector Cache Memories,” in ISCA, 1991. [67]J. W. C. Fu, J. H. Patel, and B. L. Janssens, “Stride Directed Prefetching in Scalar Processors,” in MICRO, 1992. [68]E. Garza, S. Mirbagher-Ajorpaz, T. A. Khan, and D. A. Jimenez, “Bit-level Perceptron Prediction for Indirect Branches,” in ISCA, 2019. [69]N. M. Ghiasi, T. Güloglu, H. Mustafa, C. Firtina, K. Koliogeorgi, K. Kanellopoulos, H. Mao, R. Nadig, M. Sadrosadati, J. Park, and O. Mutlu, “SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence Analysis,” in HPCA, 2026. [70]N. M. Ghiasi, J. Park, H. Mustafa, J. Kim, A. Olgun, A. Gollwitzer, D. S. Cali, C. Firtina, H. Mao, N. A. Alserr, R. Ausavarungnirun, N. Vijaykumar, M. Alser, and O. Mutlu, “GenStore: A High-Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis,” in ASPLOS, 2022. [71]N. M. Ghiasi, M. Sadrosadati, H. Mustafa, A. Gollwitzer, C. Firtina, J. Eudine, H. Mao, J. Lindegger, M. B. Cavlak, M. Alser, J. Park, and O. Mutlu, “MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage Processing,” in ISCA, 2024. [72]N. Gober, G. Chacon, L. Wang, P. V. Gratz, D. A. Jiménez, E. Teran, P. Seth, and J. Kim, “The Championship Simulator: Architectural Simulation for Education and Competition,” in arXiv, 2022. [73]Y. Gu, A. Khadem, S. Umesh, N. Liang, X. Servot, O. Mutlu, R. Iyer, and R. Das, “PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference,” in ASPLOS, 2025. [74]S. Guadarrama, A. Korattikara, O. Ramirez, P. Castro, E. Holly, S. Fishman, K. Wang, E. Gonina, N. Wu, E. Kokiopoulou, L. Sbaiz, J. Smith, G. Bartók, J. Berent, C. Harris, V. Vanhoucke, and E. Brevdo, “TF-Agents: A Library for Reinforcement Learning in TensorFlow, https://github.com/tensorflow/agents,” 2018. [75] J. Guerra, H. Pucha, J. S. Glider, W. Belluomini, and R. Rangaswami, “Cost Effective Storage Using Extent Based Dynamic Tiering,” in FAST, 2011. [76] Y. He, H. Mao, C. Giannoula, M. Sadrosadati, J. Gómez-Luna, H. Li, X. Li, Y. Wang, and O. Mutlu, “PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System,” in ASPLOS, 2025. [77]Z. Hu, M. Martonosi, and S. Kaxiras, “TCP: Tag Correlating Prefetchers,” in HPCA, 2003. [78]S. Huang, Q. Wei, D. Feng, J. Chen, and C. Chen, “Improving Flash-Based Disk Cache with Lazy Adaptive Replacement,” in TOS, 2016. [79]J. Hui, X. Ge, X. Huang, Y. Liu, and Q. Ran, “E-HASH: An Energy-Efficient Hybrid Storage System Composed of One SSD and Multiple HDDs,” in ICSI, 2012. [80]S. Iacobovici, L. Spracklen, S. Kadambi, Y. Chou, and S. G. Abraham, “Effective Stream-Based and Execution-Based Data Prefetching,” in ICS, 2004. [81]I. Iliadis, J. Jelitto, Y. Kim, S. Sarafijanovic, and V. Venkatesan, “ExaPlan: Queueing- Based Data Placement and Provisioning for Large Tiered Storage Systems,” in MASCOTS, 2015. [82]Intel,“IntelOptaneSSDDCP4801XSeries,https: //ark.intel.com/content/w/us/en/ark/products/149365/ intel-optane-ssd-dc-p4801x-series-100gb-2-5in-pcie-x4-3d-xpoint.html.” [83] Intel, “Intel SSD D3-S4510 Series, https://w.intel.com/content/w/us/ en/products/memory-storage/solid-state-drives/data-center-ssds/d3-series/ d3-s4510-series/d3-s4510-1-92tb-2-5inch-3d2.html.” [84]E. Ipek, O. Mutlu, J. F. Martínez, and R. Caruana, “Self-Optimizing Memory Con- trollers: A Reinforcement Learning Approach,” in ISCA, 2008. [85]Y. Ishii, M. Inaba, and K. Hiraki, “Access Map Pattern Matching for Data Cache Prefetch,” in ISC, 2009. [86]A. Jain and C. Lin, “Linearizing Irregular Memory Accesses for Improved Correlated Prefetching,” in MICRO, 2013. [87]M. Jalili and M. Erez, “Reducing Load Latency with Cache Level Prediction,” in HPCA, 2022. [88]A. V. Jamet, G. Vavouliotis, D. A. Jiménez, L. Alvarez, and M. Casas, “A Two Level Neural Approach Combining Off-Chip Prediction with Adaptive Prefetch Filtering,” in HPCA, 2024. [89]S. Jiang, Q. Yang, and Y. Ci, “Merging Similar Patterns for Hardware Prefetching,” in MICRO, 2022. [90] D. A. Jiménez, “Fast Path-Based Neural Branch Prediction,” in MICRO, 2003. [91]D. A. Jiménez, “Multiperspective Perceptron Predictor,” in 5th Championship Branch Prediction (CBP-5), 2016. [92]D. A. Jiménez and C. Lin, “Dynamic Branch Prediction with Perceptrons,” in HPCA, 2001. [93]D. A. Jiménez and C. Lin, “Neural Methods for Dynamic Branch Prediction,” TOCS, 2002. [94]D. A. Jiménez and E. Teran, “Multiperspective Reuse Prediction,” in MICRO, 2017. [95]D. A. Jiménez, E. Teran, and P. V. Gratz, “Last-Level Cache Insertion and Promotion 11 Policy in the Presence of Aggressive Prefetching,” IEEE CAL, 2023. [96]D. Joseph and D. Grunwald, “Prefetching using Markov Predictors,” in ISCA, 1997. [97]N. P. Jouppi, “Improving Direct-mapped Cache Performance by the Addition of a Small Fully-associative Cache and Prefetch Buffers,” in ISCA, 1990. [98] S. Kanev, J. P. Darago, K. Hazelwood, P. Ranganathan, T. Moseley, G.-Y. Wei, and D. Brooks, “Profiling a Warehouse-scale Computer,” in ISCA, 2015. [99]M. Karlsson, F. Dahlgren, and P. Stenstrom, “A Prefetching Technique for Irregular Accesses to Linked Data Structures,” in HPCA, 2000. [100]T. Kgil and T. Mudge, “FlashCache: A NAND Flash Memory File Cache for Low Power Web Servers,” in CASES, 2006. [101]T. Kgil, D. Roberts, and T. Mudge, “Improving NAND Flash Based Disk Caches,” in ISCA, 2008. [102]J. Kim, S. H. Pugsley, P. V. Gratz, A. Reddy, C. Wilkerson, and Z. Chishti, “Path Confidence Based Lookahead Prefetching,” in MICRO, 2016. [103]J. Kim, E. Teran, P. V. Gratz, D. A. Jiménez, S. H. Pugsley, and C. Wilkerson, “Kill the Program Counter: Reconstructing Program Behavior in the Processor Cache Hierarchy,” in ASPLOS, 2017. [104]Y. Klonatos, T. Makatos, M. Marazakis, M. D. Flouris, and A. Bilas, “Azor: Using Two-Level Block Selection to Improve SSD-Based I/O Caches,” in NAS, 2011. [105]S. Kondguli and M. Huang, “Division of Labor: A More Effective Approach to Prefetching,” in ISCA, 2018. [106]K. Krish, B. Wadhwa, M. S. Iqbal, M. M. Rafique, and A. R. Butt, “On Efficient Hierarchical Storage for Big Data Processing,” in CCGrid, 2016. [107]S. Kumar and C. Wilkerson, “Exploiting Spatial Locality in Data Caches using Spatial Footprints,” in ISCA, 1998. [108]C. J. Lee, E. Ebrahimi, V. Narasiman, O. Mutlu, and Y. N. Patt, “DRAM-Aware Last-Level Cache Replacement,” in HPS Technical Report, 2010. [109] C. J. Lee, O. Mutlu, V. Narasiman, and Y. N. Patt, “Prefetch-aware DRAM Con- trollers,” in MICRO, 2008. [110]C. J. Lee, O. Mutlu, V. Narasiman, and Y. N. Patt, “Prefetch-Aware Memory Con- trollers,” TC, 2011. [111]C. J. Lee, V. Narasiman, E. Ebrahimi, O. Mutlu, and Y. N. Patt, “DRAM-Aware Last-Level Cache Writeback: Reducing Write-Caused Interference in Memory Systems,” in HPS Technical Report, 2010. [112]C. J. Lee, V. Narasiman, O. Mutlu, and Y. N. Patt, “Improving Memory Bank-level Parallelism in the Presence of Prefetching,” in MICRO, 2009. [113]D. Lee, C. Min, and Y. I. Eom, “Effective SSD Caching For High-Performance Home Cloud Server,” in ICCE, 2015. [114] S. Lee, Y. Won, and S. Hong, “Mining-Based File Caching in a Hybrid Storage System,” in JISE, 2014. [115] Y. Li, L. Guo, A. Supratak, and Y. Guo, “Enabling Performance as a Service For a Cloud Storage System,” in CLOUD, 2014. [116] Z. Li, “GreenDM: A Versatile Tiering Hybrid Drive for the Trade-Off Evaluation of Performance, Energy, and Endurance,” Ph.D. dissertation, Stony Brook University, NY, 2014. [117] Y. Liang, Y. Chai, N. Bao, H. Chen, and Y. Liu, “Elastic Queue: A Universal SSD Lifetime Extension Plug-in for Cache Replacement Algorithms,” in SYSTOR, 2016. [118]L. Lin, Y. Zhu, J. Yue, Z. Cai, and B. Segee, “Hot Random Off-Loading: A Hybrid Storage System with Dynamic Data Migration,” in MASCOTS, 2011. [119] W.-F. Lin, S. Reinhardt, and D. Burger, “Reducing DRAM Latencies with an Inte- grated Memory Hierarchy Design,” in HPCA, 2001. [120] W.-F. Lin, S. Reinhardt, D. Burger, and T. Puzak, “Filtering Superfluous Prefetches using Density Vectors,” in ICCD, 2001. [121]E. Z. Liu, M. Hashemi, K. Swersky, P. Ranganathan, and J. Ahn, “An Imitation Learning Approach for Cache Replacement,” 2020. [122]Y. Liu, J. Huang, C. Xie, and Q. Cao, “RAF: A Random Access First Cache Manage- ment to Improve SSD-Based Disk Cache,” in NAS, 2010. [123]Y. Liu, X. Ge, X. Huang, and D. H. Du, “MOLAR: A Cost-Efficient, High- Performance SSD-Based Hybrid Storage Cache,” in CLUSTER, 2013. [124]N. Lu, I.-S. Choi, S.-H. Ko, and S.-D. Kim, “A PRAM Based Block Updating Man- agement for Hybrid Solid State Disk,” in ELEX, 2012. [125]X. Lu, H. Najafi, J. Liu, and X.-H. Sun, “CHROME: Concurrency-Aware Holistic Cache Management Framework with Online Reinforcement Learning,” in HPCA, 2024. [126]Z.-W. Lu and G. Zhou, “Design and Implementation of Hybrid Shingled Recording RAID System,” in PiCom, 2016. [127]D. Luo, J. Wan, Y. Zhu, N. Zhao, F. Li, and C. Xie, “Design and Implementation of a Hybrid Shingled Write Disk System ,” in TPDS, 2015. [128]Y. Lv, X. Chen, G. Sun, and B. Cui, “A Probabilistic Data Replacement Strategy for Flash-Based Hybrid Storage System,” in APWeb, 2013. [129]Y. Lv, B. Cui, X. Chen, and J. Li, “Hotness-Aware Buffer Management For Flash- Based Hybrid Storage Systems,” in CIKM, 2013. [130]S. Ma, H. Chen, Y. Shen, H. Lu, B. Wei, and P. He, “Providing Hybrid Block Storage for Virtual Machines using Object-based Storage,” in ICPADS, 2014. [131]H. Mao, M. Alser, M. Sadrosadati, C. Firtina, A. Baranwal, D. S. Cali, A. Manglik, N. A. Alserr, and O. Mutlu, “GenPIP: In-memory Acceleration of Genome Analysis via Tight Integration of Basecalling and Read Mapping,” in MICRO, 2022. [132]C. Matsui, C. Sun, and K. Takeuchi, “Design of Hybrid SSDs with Storage Class Memory and NAND Flash Memory,” in IEEE, 2017. [133]C. Matsui, T. Yamada, Y. Sugiyama, Y. Yamaga, and K. Takeuchi, “Tri-Hybrid SSD with Storage Class Memory (SCM) and MLC/TLC NAND Flash Memories,” Proc. IEEE, 2017. [134]G. Memik, G. Reinman, and W. H. Mangione-Smith, “Just Say No: Benefits of Early Cache Miss Determination,” in HPCA, 2003. [135]M. R. Meswani, S. Blagodurov, D. Roberts, J. Slice, M. Ignatowski, and G. H. Loh, “Heterogeneous Memory Architectures: A HW/SW Approach for Mixing Die- Stacked and Off-package Memories,” in HPCA, 2015. [136] J. Meza, Y. Luo, S. Khan, J. Zhao, Y. Xie, and O. Mutlu, “A Case for Efficient Hardware/Software Cooperative Management of Storage and Memory,” in WEED, 2013. [137] P. Michaud, “Best-offset Hardware Prefetching,” in HPCA, 2016. [138]D. Montgomery, “Extent Migration For Tiered Storage Architecture,” in USPTO, 2014. [139]J. Mukundan and J. F. Martinez, “MORSE: Multi-objective Reconfigurable Self- Optimizing Memory Scheduler,” in ISCA, 2012. [140] O. Mutlu, “Memory Scaling: A Systems Architecture Perspective,” in IMW, 2013. [141] O. Mutlu, “Intelligent Architectures for Intelligent Machines,” in VLSI-DAT, 2020. [142] O. Mutlu, “Intelligent Architectures for Intelligent Computing Systems,” in DATE, 2021. [143]O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “A Modern Primer on Processing in Memory,” in Emerging computing: from devices to systems: looking beyond Moore and Von Neumann. Springer, 2022. [144]O. Mutlu, H. Kim, D. N. Armstrong, and Y. N. Patt, “Using the First-Level Caches as Filters to Reduce the Pollution Caused by Speculative Memory References,” IJPP, 2005. [145] O. Mutlu, A. Olgun, and İ. E. Yüksel, “Memory-Centric Computing: Solving Com- puting’s Memory Problem,” in IMW, 2025. [146]R. Nadig, V. Arulchelvan, R. Bera, T. Shahroodi, G. Singh, M. Sadrosadati, J. Park, and O. Mutlu, “Harmonia: A Multi-Agent Reinforcement Learning Approach to Data Placement and Migration in Hybrid Storage Systems,” in ICS, 2026. [147] A. Navarro-Torres, B. Panda, J. Alastruey-Benedé, P. Ibáñez, V. Viñals-Yúfera, and A. Ros, “Berti: An Accurate Local-Delta Data Prefetcher,” in MICRO, 2022. [148]J. Niu, J. Xu, and L. Xie, “Hybrid Storage Systems: A Survey of Architectures and Algorithms,” in IEEE Access, 2018. [149]Y. Oh, J. Choi, D. Lee, and S. H. Noh, “Caching Less For Better Performance: Balancing Cache Size and Update Cost of Flash Memory Cache in Hybrid Storage Systems,” in FAST, 2012. [150]Y. Oh, E. Lee, C. Hyun, J. Choi, D. Lee, and S. H. Noh, “Enabling Cost-Effective Flash Based Caching with an Array of Commodity SSDs,” in Middleware, 2015. [151]G. F. Oliveira, J. Gómez-Luna, S. Ghose, A. Boroumand, and O. Mutlu, “Accelerating Neural Network Inference with Processing-in-DRAM: From the Edge to the Cloud,” IEEE Micro, 2022. [152]J. Ou, J. Shu, Y. Lu, L. Yi, and W. Wang, “EDM: An Endurance-Aware Data Migration Scheme for Load Balancing in SSD Storage Clusters,” in IPDPS, 2014. [153]S. Pakalapati and B. Panda, “Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware Prefetching,” in ISCA, 2020. [154]B. Panda and S. Balachandran, “Hardware Prefetchers for Emerging Parallel Ap- plications,” in PACT, 2012. [155] B. Panda and S. Balachandran, “XStream: Cross-Core Spatial Streaming Based MLC Prefetchers for Parallel Applications in CMPs,” in PACT, 2014. [156] D. Park and D. H. Du, “Hot Data Identification for Flash-based Storage Systems Using Multiple Bloom Filters,” in MSST, 2011. [157] L. Peled, S. Mannor, U. Weiser, and Y. Etsion, “Semantic Locality and Context-Based Prefetching using Reinforcement Learning,” in ISCA, 2015. [158] S. H. Pugsley, Z. Chishti, C. Wilkerson, P.-f. Chuang, R. L. Scott, A. Jaleel, S.-L. Lu, K. Chow, and R. Balasubramonian, “Sandbox Prefetching: Safe Run-Time Evaluation of Aggressive Prefetchers,” in HPCA, 2014. [159] M. K. Qureshi and G. H. Loh, “Fundamental Latency Trade-off in Architecting DRAM Caches: Outperforming Impractical SRAM-Tags with a Simple and Practical Design,” in MICRO, 2012. [160]A. Raghavan, A. Chandra, and J. B. Weissman, “Tiera: Towards Flexible Multi- Tiered Cloud Storage Instances,” in Middleware, 2014. [161]D. Reinsel and J. Rydning, “Breaking the 15K-rpm HDD Performance Barrier with Solid State Hybrid Drives,” in IDC, 2013. [162]J. Ren, X. Chen, Y. Tan, D. Liu, M. Duan, L. Liang, and L. Qiao, “Archivist: A Machine Learning Assisted Data Placement Mechanism for Hybrid Storage Systems,” in ICCD, 2019. [163]R. Salkhordeh, H. Asadi, and S. Ebrahimi, “Operating System Level Data Tiering Using Online Workload Characterization,” in JSC, 2015. [164]M. Saxena and M. M. Swift, “Design and Prototype of a Solid-State Cache,” in TOS, 2014. [165]M. Saxena, M. M. Swift, and Y. Zhang, “FlashTier: A Lightweight, Consistent and Durable Storage Cache,” in EuroSys, 2012. [166]Seagate, “Seagate Barracuda Datasheet, https://w.seagate.com/w-content/ datasheets/pdfs/3-5-barracuda-3tbDS1900-10-1710US-en_US.pdf .” [167]A. Sembrant, E. Hagersten, and D. Black-Schaffer, “The Direct-To-Data (D2D) Cache: Navigating the Cache Hierarchy with a Single Lookup,” ISCA, 2014. [168]A. Sembrant, E. Hagersten, and D. Black-Schaffer, “A Split Cache Hierarchy for Enabling Data-Oriented Optimizations,” in HPCA, 2017. [169]S. Sethumurugan, J. Yin, and J. Sartori, “Designing a Cost-Effective Cache Replace- ment Policy using Machine Learning,” in HPCA, 2021. [170]M. Shakerinava, M. Bakhshalipour, P. Lotfi-Kamran, and H. Sarbazi-Azad, “Multi- lookahead Offset Prefetching,” 3rd Data Prefetching Championship, 2019. [171]M. Shevgoor, S. Koladiya, R. Balasubramonian, C. Wilkerson, S. H. Pugsley, and Z. Chishti, “Efficiently Prefetching Complex Address Patterns,” in MICRO, 2015. [172]H. Shi, R. V. Arumugam, C. H. Foh, and K. K. Khaing, “Optimal Disk Storage 12 Allocation for Multitier Storage System,” in TMAG, 2013. [173]J. Shun and G. E. Blelloch, “Ligra: a Lightweight Graph Processing Framework for Shared Memory,” in PPoPP, 2013. [174]G. Singh, M. Alser, D. S. Cali, D. Diamantopoulos, J. Gómez-Luna, H. Corporaal, and O. Mutlu, “FPGA-Based Near-Memory Acceleration of Modern Data-Intensive Applications,” IEEE Micro, 2021. [175]G. Singh, R. Nadig, J. Park, R. Bera, N. Hajinazar, D. Novo, J. Gómez-Luna, S. Stuijk, H. Corporaal, and O. Mutlu, “Sibyl: Adaptive and Extensible Data Placement in Hybrid Storage Systems using Online Reinforcement Learning,” in ISCA, 2022. [176] C. W. Smullen, J. Coffman, and S. Gurumurthi, “Accelerating Enterprise Solid-State Disks With Non-Volatile Merge Caching,” in IGSC, 2010. [177]S. Somogyi, T. F. Wenisch, A. Ailamaki, and B. Falsafi, “Spatio-Temporal Memory Streaming,” in ISCA, 2009. [178] S. Somogyi, T. F. Wenisch, A. Ailamaki, B. Falsafi, and A. Moshovos, “Spatial Memory Streaming,” in ISCA, 2006. [179]M. Soysal, K. Koliogeorgi, C. Firtina, N. M. Ghiasi, R. Nadig, H. Mao, G. F. Oliveira, Y. Liang, K. Zambaku, M. Sadrosadati, and O. Mutlu, “MARS: Processing-In- Memory Acceleration of Raw Signal Genome Analysis Inside the Storage Subsys- tem,” in ICS, 2025. [180]S. Srinath, O. Mutlu, H. Kim, and Y. N. Patt, “Feedback Directed Prefetching: Improving the Performance and Bandwidth-Efficiency of Hardware Prefetchers,” in HPCA, 2007. [181] M. Srinivasan, P. Saab, and V. Tkachenko, “Flashcache,” in Facebook, 2010. [182] C. Sun, K. Miyaji, K. Johguchi, and K. Takeuchi, “A High Performance and Energy- Efficient Cold Data Eviction Algorithm for 3D-TSV Hybrid ReRAM/MLC NAND SSD,” in CAS, 2013. [183]R. S. Sutton and A. G. Barto, “Reinforcement Learning: An Introduction,” in The MIT Press, 2017. [184]J. Tai, B. Sheng, Y. Yao, and N. Mi, “SLA-Aware Data Migration in a Shared Hybrid Storage Cluster,” in C, 2015. [185] M. Tarihi, H. Asadi, A. Haghdoost, M. Arjomand, and H. Sarbazi-Azad, “A Hy- brid Non-Volatile Cache Design for Solid-State Drives Using Comprehensive I/O Characterization,” in TC, 2015. [186]D. Tarjan and K. Skadron, “Merging Path and Gshare Indexing in Perceptron Branch Prediction,” TACO, 2005. [187] E. Teran, Z. Wang, and D. A. Jiménez, “Perceptron Learning for Reuse Prediction,” in MICRO, 2016. [188]D. A. Varkey, B. Panda, and M. Mutyam, “RCTP: Region Correlated Temporal Prefetcher,” in ICCD, 2017. [189]E. Vasilakis, V. Papaefstathiou, P. Trancoso, and I. Sourdis, “Hybrid2: Combining Caching and Migration in Hybrid Memory Systems,” in HPCA, 2020. [190]C. Wang, D. Wang, Y. Chai, C. Wang, and D. Sun, “Larger, Cheaper, but Faster: SSD-SMR Hybrid Storage Boosted by a New SMR-Oriented Cache Framework,” in MSST, 2017. [191]H. Wang and P. Varman, “Balancing Fairness and Efficiency in Tiered Storage Systems with Bottleneck-Aware Allocation,” in FAST, 2014. [192]L. Wang, J. Zhan, C. Luo, Y. Zhu, Q. Yang, Y. He, W. Gao, Z. Jia, Y. Shi, S. Zhang, C. Zheng, G. Lu, K. Zhan, X. Li, and B. Qiu, “BigDataBench: A Big Data Benchmark Suite from Internet Services,” in HPCA, 2014. [193] T. F. Wenisch, M. Ferdman, A. Ailamaki, B. Falsafi, and A. Moshovos, “Practical Off-Chip Meta-Data for Temporal Memory Streaming,” in HPCA, 2009. [194] T. F. Wenisch, M. Ferdman, A. Ailamaki, B. Falsafi, and A. Moshovos, “Making Address-Correlated Prefetching Practical,” IEEE Micro, 2010. [195]T. F. Wenisch, S. Somogyi, N. Hardavellas, J. Kim, A. Ailamaki, and B. Falsafi, “Temporal Streaming of Shared Memory,” in ISCA, 2005. [196]H. Wu, K. Nathella, J. Pusdesris, D. Sunwoo, A. Jain, and C. Lin, “Temporal Prefetch- ing Without the Off-Chip Metadata,” in MICRO, 2019. [197]H. Wu, K. Nathella, D. Sunwoo, A. Jaleel, and C. Lin, “Efficient Metadata Manage- ment for Irregular Data Prefetching,” in ISCA, 2019. [198]X. Wu and A. N. Reddy, “Managing Storage Space in a Flash and Disk Hybrid Storage System,” in MASCOTS, 2009. [199]X. Wu and A. N. Reddy, “Exploiting Concurrency to Improve Latency and through- put in a Hybrid Storage System,” in MASCOTS, 2010. [200]X. Wu and A. N. Reddy, “Data Organization in a Hybrid Storage System,” in ICNC, 2012. [201]W. Xiao, H. Dong, L. Ma, Z. Liu, and Q. Zhang, “HS-BAS: A Hybrid Storage System Based on Band Awareness of Shingled Write Disk,” in ICCD, 2016. [202]J. Xue, F. Yan, A. Riska, and E. Smirni, “Storage Workload Isolation via Tier Warming,” in ICAC, 2014. [203]G. Yadgar, M. Factor, K. Li, and A. Schuster, “Management of Multilevel, Multiclient Cache Hierarchies with Application Hints,” in TOCS, 2011. [204]J. Yang, N. Plasson, G. Gillis, N. Talagala, S. Sundararaman, and R. Wood, “HEC: Improving Endurance of High Performance Flash-Based Cache Devices,” in SYSTOR, 2013. [205]F. Ye, J. Chen, X. Fang, J. Li, and D. Feng, “A Regional Popularity-Aware Cache Replacement Algorithm to Improve the Performance and Lifetime of SSD-Based Disk Cache,” in NAS, 2015. [206]A. Yoaz, M. Erez, R. Ronen, and S. Jourdan, “Speculation Techniques for Improving Load Related Instruction Scheduling,” in ISCA, 1999. [207]G. Zhang, L. Chiu, C. Dickey, L. Liu, P. Muench, and S. Seshadri, “Automated Lookahead Data Migration in SSD-enabled Multi-tiered Storage Systems,” in MSST, 2010. [208]T. Zhang, B. Grot, W. He, Y. Lv, P. Qu, F. Su, W. Wang, G. Zhang, X. Zhang, and Y. Zhang, “Hierarchical Prefetching: A Software-Hardware Instruction Prefetcher for Server Applications,” ASPLOS, 2025. [209]Z. Zhang, Y. Kim, X. Ma, G. Shipman, and Y. Zhou, “Multi-level Hybrid Cache: Impact and Feasibility,” in ORNL Tech. Rep, 2012. [210] D. Zhao, K. Qiao, and I. Raicu, “Towards Cost-Effective and High-Performance Caching Middleware for Distributed Systems,” in IJBDI, 2016. [211]X. Zhao, Z. Li, and L. Zeng, “FDTM: Block Level Data Migration Policy in Tiered Storage System,” in NPC, 2010. [212]X. Zhuang and H.-H. Lee, “A Hardware-based Cache Pollution Filtering Mechanism for Aggressive Prefetches,” in ICPP, 2003. [213]Z. Zong, R. Fares, B. Romoser, and J. Wood, “FastStor: Data-Mining-Based Multi- layer Prefetching for Hybrid Storage Systems,” in C, 2014. 13