๐Ÿ“ ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System - ICML'26 Spotlight

June 18, 2026

๐Ÿ“ DroidSpeak: KV Cache Sharing Across Fine-tuned Model Variants - NSDI'26

June 7, 2026

๐Ÿ“ Empirical Recipes for Efficient and Compact Vision-Language Models

March 24, 2026

๐Ÿ“ Empower Vision Applications with LoRA LMM - Eurosys'25

December 15, 2025

๐Ÿ“ Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding - INFOCOM'26

December 12, 2025

๐Ÿ“ Elastic On-Device LLM Service - MobiCom '25

December 8, 2025

๐Ÿ“ RServe: Overlapping Encoding and Prefill for Efficient LMM Inference

December 6, 2025

๐Ÿ“ ModServe - Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving - SoCC'25

November 27, 2025

๐Ÿ“ Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices - INFOCOM'25

November 26, 2025

๐Ÿ“ Breaking the Wall: Unifying Edge GPUs and NPUs into Pipeline Parallelism for Efficient LLM Fine-Tuning

November 17, 2025

๐Ÿ“ Efficiently Serving Large Multimodal Models Using EPD Disaggregation - ICML'25

November 15, 2025

๐Ÿ“ ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism - NeurIPS'25 Oral

November 13, 2025