Search arXiv⌕ Search

arXiv · 2610.12164

Mole: Tier-Specific Hotness Profiling Driven Memory Tiering for Multi-Tiered Memory Systems

Abstract

Multi-tiered memory systems combine fast, small upper tiers with slow, large lower tiers to improve performance and cost efficiency. Existing designs, such as AutoTiering and MTM, rely on greedy promotion and stepwise demotion, which can intensify contention for limited capacity in faster tiers. We observe that promotions directly improve performance, whereas demotions primarily reclaim space. Based on this asymmetry, we propose Mole, a memory tiering system that separates promotion and demotion destinations. Mole employs staging demotion to move cold pages directly to the lowest tier, bypassing intermediate tiers, and targeted promotion to place hot pages in tiers that match their current hotness. The lowest tier serves as a staging area from which pages can be promoted when they become hot again. This separation preserves intermediate-tier capacity for promotions, reducing tier contention and migration failures. However, staging demotion requires timely identification of pages that become hot after demotion. To meet this requirement with low profiling overhead, Mole uses tier-specific profiling: it profiles the lowest tier at high frequency to detect reactivated hot pages, while profiling upper tiers at lower frequency to identify cold pages. Experimental results show that Mole reduces migration failures and improves performance under dynamic access patterns.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mingyang Liu, Congming Gao, Xufeng Yang, Fang Wu, Youmin Chen, Renhui Chen, Jiwu Shu. 2026-10-08. Mole: Tier-Specific Hotness Profiling Driven Memory Tiering for Multi-Tiered Memory Systems. https://arxiv.org/abs/2610.12164

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

AKTS: Sub-Microsecond Kernel Policy Switching for Language-Model Agents

GPU-backed LLM servers multiplex interactive requests with background batch work on the same CPUs; a fixed kernel policy serves one objective and loses the other, so agentic OS control needs to switch scheduler behavior as the workload changes. The hard part is applying the switch safely and fast enough for the kernel: scheduler events occur every 1-10 $μ$s, and any code running there must satisfy the eBPF verifier. Scalar knobs are fast but limited, while generating eBPF policy code puts compilation, verification and possible rejection on the runtime path. We present AKTS, which verifies a policy library once, at load time, and reduces the agent's runtime action to writing an integer index into an in-kernel array of preverified policies, resolved by a tail call. Because the agent emits an index rather than code, verifier failure is not a runtime outcome. On Linux 6.14, AKTS applies a policy switch in 920 ns (p50); makes an invalid index inert across 60,217 invocations on an attached scheduler; and switches policies in a vLLM workload to capture 97% of a throughput policy's batch work while matching a latency policy's burst response.

cs.OS↗

Behind a Simple Read: Understanding Work and Waiting in the Linux I/O Stack

A simple read interface provides uniform functional semantics, but not equally simple or predictable performance behavior. We use synchronous large-buffer reads in Linux as an observation window and decompose the buffered-read path end to end, distinguishing work volume, processing time, and critical-path exposure. We find that Linux reduces metadata work through large folios and overlaps most cold-read copying with device waiting, but these mechanisms depend on access advice, folio granularity, cache state, and backend execution. Using an experimental kernel prototype, we validate additional opportunities from out-of-order early copying and opportunistic parallelism, while showing that faster request submission does not improve end-to-end performance when SSD supply is already sufficient. To explain device waiting, we abstract first-completion wait $F$ and subsequent completion capacity $B_{CQ}$ from finite request-batch completion timelines. Measurements on a real SSD characterize how they vary with request size and batch size, while MQSim experiments connect them to internal device mechanisms and distinguish finite-batch completion from sustained throughput. We then construct a cross-layer wait chain along the actual folio--bio--request mapping. Independently calibrated device parameters predict waiting at the block layer and at read_pages with errors no greater than approximately 6.9% and 3.2%, respectively. Finally, four use cases apply the model and wait chain to parallel submission, Linux readahead, dependent-read layout, and polling versus sleeping. Their gains, no gains, and gain reversals form a closed loop from observation and modeling to explanation and control, providing a measurable basis for performance decisions in layered I/O stacks.

cs.OS↗

FDP: The Data Placement Promise of Modern NVMe SSDs

NVMe SSDs are now widely deployed as the storage tier in data centers. As SSDs have evolved over the past decade, the commu- nity has continued to debate the interfaces they expose and how operating systems and storage systems should exploit them. The NVMe Flexible Data Placement (FDP) proposal is the latest point in this design space. FDP introduces an interface based on Reclaim Units that enables explicit data placement to reduce device write amplification without the software engineering costs of sequential- write constraints and host garbage collection. FDP-enabled SSDs are emerging in commercial products and early data center de- ployments. Their compatibility with conventional block I/O allows existing applications to run unchanged, allowing a frictionless adop- tion in industry. This paper presents an experimental evaluation of FDP SSDs to characterize their data placement guarantees over the raw device interface. We then revisit two widely deployed and distinct open- source storage systems, MySQL and RocksDB, and examine whether lifetime-based data separation and distinct write patterns built into their architectures can be mapped onto FDP SSDs without inva- sive changes. Our evaluation shows end-to-end WAF reductions at higher device utilization, along with QoS and throughput improve- ments under synthetic and real-world workloads. These results demonstrate that FDP provides a practical and deployable cross- layer mechanism for data placement with open-source ecosystem support on Linux. They also highlight why FDP SSDs are gaining traction in industry.

cs.OS↗