๐ Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding - INFOCOM'26
๐ ModServe - Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving - SoCC'25
๐ Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices - INFOCOM'25
๐ Breaking the Wall: Unifying Edge GPUs and NPUs into Pipeline Parallelism for Efficient LLM Fine-Tuning
๐ ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism - NeurIPS'25 Oral