π DOCPRUNE: Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning - CVPR'26
π PRUNE REDUNDANCY, PRESERVE ESSENCE: VISION TOKEN COMPRESSION IN VLMS VIA SYNERGISTIC IMPORTANCEβDIVERSITY - ICLR'26
π ONE PATCH DOESNβT FIT ALL: ADAPTIVE PATCHING FOR NATIVE-RESOLUTION MULTIMODAL LARGE LANGUAGE MODELS - ICLR'26
π Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models - EMNLP'25
π Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding - INFOCOM'26
π ModServe - Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving - SoCC'25
π ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism - NeurIPS'25 Oral
π Stop Looking for βImportant Tokensβ in Multimodal Language Models: Duplication Matters More - EMNLP'25
π Efficient Multi-modal Large Language Models via Progressive Consistency Distillation - NeurIPS'25
π QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models - NeurIPS'25 Spotlight
π SPECVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning - EMNLP'25
π Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs - *NeurIPS'25