Reading List
Papers and references for the Advanced Topics lectures (Weeks 9–14).
Topics
- Heterogeneous Computing in Cloud — Lec 16–17
- Modern Resource Scheduling — Lec 18–19
- ML in Cloud — Lec 20–21
- Cloud and IoT — Lec 22–23
- Cloud Security — Lec 24
- Future Cloud Infrastructure — Lec 25
- Cloud in the Agentic-AI Era — Lec 26
Heterogeneous Computing in Cloud
Lec 16–17 · 10/20 & 10/22
Lec 16
- Caulfield, Adrian M., et al. "A cloud-scale acceleration architecture." 2016 49th Annual IEEE/ACM international symposium on microarchitecture (MICRO). IEEE, 2016.
- NVIDIA Multi-Instance GPU User guide.
- Khawaja, Ahmed, et al. "Sharing, protection, and compatibility for reconfigurable fabric with AmorphOS." OSDI 18.
- Jeon, Myeongjae, et al. "Analysis of Large-Scale Multi-Tenant GPU clusters for DNN training workloads." USENIX ATC 19.
Lecture:
Student Presentation:
Lec 17
- Li, Suyi, et al. "Heterogeneity at Hyperscale: Characterization and Scheduling of Large Production AI Clusters at Alibaba (Operational Systems)." OSDI 26.
- Yang, Lingyun, et al. "GPU-Disaggregated serving for deep learning recommendation models at scale." NSDI 25.
- Chen, Jiyang, et al. "μShell: A Microkernel-based FPGA Shell Architecture." OSDI 26.
- He, Congjie, et al. "WaferLLM: Large language model inference at wafer scale." OSDI 25.
- Hei, Chenyang, et al. "HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters." NSDI 26.
Lecture:
Student Presentation:
Optional Reading:
Modern Resource Scheduling
Lec 18–19 · 10/27 & 10/29
Lec 18
- Tang, Chunqiang, et al. "Twine: A unified cluster management system for shared infrastructure." OSDI 20.
- Rzadca, Krzysztof, et al. "Autopilot: workload autoscaling at google." EuroSys 20.
- Cortez, Eli, et al. "Resource central: Understanding and predicting workloads for improved resource management in large cloud platforms." SOSP 17.
- Gog, Ionel, et al. "Firmament: Fast, centralized cluster scheduling at scale." OSDI 16.
Lecture:
Student Presentation:
Optional Reading:
Lec 19
- Chai, Xiaohu, et al. "Fork in the road: Reflections and optimizations for cold start latency in production serverless systems." OSDI 25.
- Zhang, Zhengtong, et al. "DVLA: Dynamic VM Lifetime Aware Scheduling for Drifting Lifetime Distributions and Long-Lived VM Placement Debt (Operational Systems)." OSDI 26.
- Fu, Yuqi, et al. "ALPS: An Adaptive Learning, Priority OS Scheduler for Serverless Functions." ATC 24.
- Liu, Qingyuan, et al. "Harmonizing efficiency and practicability: Optimizing resource utilization in serverless computing with jiagu." ATC 24.
Lecture:
Student Presentation:
Optional Reading:
ML in Cloud
Lec 20–21 · 11/3 & 11/5
Lec 20
- Qiao, Aurick, et al. "Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning." OSDI 21.
- Xiao, Wencong, et al. "Gandiva: Introspective cluster scheduling for deep learning." OSDI 18.
- Xiao, Wencong, et al. "AntMan: Dynamic scaling on GPU clusters for deep learning." OSDI 20.
- Weng, Qizhen, et al. "MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters." NSDI 22.
Lecture:
Student Presentation:
Optional Reading:
Lec 21
- Xie, Zhiqiang, et al. "Strata: Hierarchical context caching for long context language model serving." OSDI 26.
- Kwon, Woosuk, et al. "Efficient memory management for large language model serving with paged attention." SOSP 23.
- Wu, Bingyang, et al. "FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference." NSDI 26.
- Khare, Alind, et al. "SuperServe: Fine-Grained Inference Serving for Unpredictable Workloads." NSDI 25.
Lecture:
Student Presentation:
Optional Reading:
Cloud and IoT
Lec 22–23 · 11/10 & 11/12
Lec 22
- Cuervo, Eduardo, et al. "Maui: making smartphones last longer with code offload." MobiSys 10.
- Gordon, Mark S., et al. "COMET: Code offload by migrating execution transparently." OSDI 12.
- Chun, Byung-Gon, et al. "Clonecloud: elastic execution between mobile device and cloud." EuroSys 11.
- Cuervo, Eduardo, et al. "Kahawai: High-quality mobile gaming using gpu offload." MobiSys 15.
Lecture:
Student Presentation:
Optional Reading:
Lec 23
- Kang, Yiping, et al. 2017. "Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge." ASPLOS 17.
- Shen, Zheyu, et al. "Edgelora: An efficient multi-tenant llm serving system on edge devices." MobiSys 25.
- Prakash, Shvetank, et al. "Lifetime-Aware Design for Item-Level Intelligence at the Extreme Edge." ASPLOS 26.
- Sen, Tanmoy, Haiying Shen, and Anand Padmanabha Iyer. "Flex: Fast, accurate DNN inference on low-cost edges using heterogeneous accelerator execution." EuroSys 25.
Lecture:
Student Presentation:
Optional Reading:
Cloud Security
Lec 24 · 11/17
Lec 24
- Ristenpart, Thomas, et al. "Hey, you, get off of my cloud: exploring information leakage in third-party compute clouds." CCS 09.
- Baumann, Andrew, et al. "Shielding Applications from an Untrusted Cloud with Haven." OSDI 14.
- Shi, Jiacheng, et al. "Serverless functions made confidential and efficient with split containers." USENIX Security 25.
- Zhao, Zirui, et al. "Everywhere all at once: Co-location attacks on public cloud FaaS." ASPLOS 24.
- Wang, Shiwen, et al. "Enjoy the Free Lunch, Someone Paid for Us: Escaping Resource Limits of MicroVM-based Containers." USENIX Security 26.
- Liu, Xunqi, et al. "The Dark Side of Flexibility: Detecting Risky Permission Chaining Attacks in Serverless Applications." NDSS 26.
Lecture:
Student Presentation:
Optional Reading:
Future Cloud Infrastructure
Lec 25 · 11/19
Lec 25
- Kuchler, Tom, et al. "Unlocking true elasticity for the cloud-native era with dandelion." SOSP 25.
- Ruan, Zhenyuan, et al. "Quicksand: Harnessing stranded datacenter resources with granular computing." NSDI 25.
- Li, Quanxi, et al. "Beehive: A scalable disaggregated memory runtime exploiting asynchrony of multithreaded programs." NSDI 25.
- Yi, Shushu, et al. "Espresso: Constructing Cost-Efficient CXL JBOF via Inter-SSD Computing Resource Sharing." OSDI 26.
- Qiu, Ziyue, et al. "Moirai: Optimizing Placement of Data and Compute in Hybrid Clouds." SOSP 25.
- Lou, Chiheng, et al. "HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds." NSDI 26.
Lecture:
Student Presentation:
Optional Reading:
Cloud in the Agentic-AI Era
Lec 26 · 11/24
Lec 26
- Chaudhry, Gohar Irfan, et al. "Murakkab:Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms." OSDI 26.
- Gao, Wei, et al. "RollArt: Disaggregated Multi-Task Agentic RL Training at Scale." OSDI 26.
- Xiang, Dawei, et al. "LLM-as-Scheduler: Agentic Workflow Dynamic Scheduling." ACL 26.
- Liu, Shiyi, et al. "AIMS: Cost-Efficient LLM-Based Agent Deployment in Hybrid Cloud-Edge Environments." EuroSys 26.
- Fang, Taosong, et al. "Flashagents: Accelerating multi-agent llm systems via streaming prefill overlap." MLSys 26.
Lecture:
Student Presentation:
Optional Reading: