2026
-
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Tsim
-
Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving
-
Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving
-
RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
-
Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention
2025
-
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
-
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
-
Managing Scalable Direct Storage Accesses for GPUs with GoFS
-
ReGate: Enabling Power Gating in Neural Processing Units
-
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
-
SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design
-
FleetIO: Managing Multi-Tenant Cloud Storage with Multi-Agent Reinforcement Learning
-
ByteFS: System Support for (CXL-based) Memory-Semantic Solid-State Drives
-
Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud Platforms
2024
-
Exploring the Efficiency of Renewable Energy-based Modular Data Centers at Scale
-
Scaling Deep Learning Computation over the Inter-Core Connected Intelligence Processor with T10
-
Hardware-Assisted Virtualization of Neural Processing Units for Cloud Platforms
-
HADES: Hardware-Assisted Distributed Transactions in the Age of Fast Networks and SmartNICs
2023
-
Learning to Drive Software-Defined Solid-State Drives
-
G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor Migrations
-
RackBlox: A Software-Defined Rack-Scale Storage System with Network-Storage Co-Design
-
System Virtualization for Neural Processing Units
-
The Security War in File Systems: An Empirical Study from A Vulnerability-Centric Perspective
-
V10: Hardware-Assisted NPU Multi-tenancy for Improved Resource Utilization and Fairness
-
LeaFTL: A Learning-Based Flash Translation Layer for Solid-State Drives
2022
-
Learning to Drive Software-Defined Storage
-
BlockFlex: Enabling Storage Harvesting with Software-Defined Flash in Modern Cloud Platforms
-
RSSD: Defend Against Ransomware with Hardware-Isolated Network-Storage Codesign and Post-Attack Analysis
-
Understanding and Detecting Deep Memory Persistency Bugs in NVM Programs with DeepMC
-
RSSD: Defend Against Ransomware with Hardware-Isolated Network-Storage Codesign and Post-Attack Analysis
-
UniHeap: Managing Persistent Objects Across Managed Runtimes for Non-Volatile Memory
-
IceClave: A Trusted Execution Environment for In-Storage Computing