← Search

Li Yu

29 accepted papers

2026

Joint Geometric and Trajectory Consistency Learning for One-Step Real-World Super-Resolution

ICML 2026poster

Diffusion-based Real-World Image Super-Resolution (Real-ISR) achieves impressive perceptual quality but suffers from high computational costs due to iterative sampling. While recent distillation approaches leveraging large-scale Text-to-Image (T2I) priors have enabled one-step generation, they are t…

Cited by 0SourceScholar
2026

LENS: Learning to Segment Anything with Unified Reinforced Reasoning

AAAI 2026technical

Text-prompted image segmentation enables fine-grained visual understanding and is critical for applications such as human-computer interaction and robotics. However, existing supervised fine-tuning methods typically ignore explicit chain-of-thought (CoT) reasoning at test time, which limits their ab

Cited by 0SourcePDFScholar
2026

Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment

AAAI 2026technical

Recent advancements in weakly-supervised video anomaly detection have achieved remarkable performance by applying the multiple instance learning paradigm based on multimodal foundation models such as CLIP to highlight anomalous instances and classify categories. However, their objectives may tend to

Cited by 0SourcePDFScholar
2026

SGE-GLoc: Semantic Gaussian Ellipsoid Scene Graphs for Efficient LiDAR Global Localization

RA-L 2026

Global localization, encompassing robust place recognition and precise transformation estimation, is crucial for mobile robot navigation when the global navigation satellite system (GNSS) is unavailable. While LiDAR-based approaches are favored for their accuracy in 3D perception and resilience to i

Cited by 0SourceScholar
2026

TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection

CVPR 2026

Co-salient Object Detection (CoSOD) aims to segment salient objects that consistently appear across a group of related images. Despite the notable progress achieved by recent training-based approaches, they still remain constrained by the closed-set datasets and exhibit limited generalization. Howev

Cited by 0SourcecodeScholar
2025

CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection

AAAI 2025technical

Recent research on universal object detection aims to introduce language in a SoTA closed-set detector and then generalize the open-set concepts by constructing large-scale (text-region) datasets for training. However, these methods face two main challenges: (i) how to efficiently use the prior info…

Cited by 0SourcePDFScholar
2025

Group Modeling and Recommendation Based on Multi-Behavior Interactions in Live Streaming E-Commerce

ICASSP 2025accepted

The live streaming e-commerce paradigm is experiencing a meteoric surge, distinguished by hosts dynamically showcasing products via live videos and fostering interactive engagement with their viewers. Although a streamer may not cater to every individual’s tastes, it is imperative that it resonates…

Cited by 0SourceScholar
2025

Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity

CVPR 2025highlight

How can we enable models to comprehend video anomalies occurring over varying temporal scales and contexts?Traditional Video Anomaly Understanding (VAU) methods focus on frame-level anomaly prediction, often missing the interpretability of complex and diverse real-world anomalies. Recent multimodal…

2025

Structural Pruning via Spatial-aware Information Redundancy for Semantic Segmentation

AAAI 2025technical

In recent years, semantic segmentation has flourished in various applications. However, the high computational cost remains a significant challenge that hinders its further adoption. The filter pruning method for structured network slimming offers a direct and effective solution for the reduction o…

2024

DTMFormer: Dynamic Token Merging for Boosting Transformer-Based Medical Image Segmentation

AAAI 2024technical

Despite the great potential in capturing long-range dependency, one rarely-explored underlying issue of transformer in medical image segmentation is attention collapse, making it often degenerate into a bypass module in CNN-Transformer hybrid architectures. This is due to the high computational comp…

2024

Exploration of Visual Prompt in Grounded Pre-Trained Open-Set Detection

ICASSP 2024accepted

Text prompts are crucial for generalizing pre-trained open-set object detection models to new categories. However, current methods for text prompts are limited as they require manual feedback when generalizing to new categories, which restricts their ability to model complex scenes, often leading to…

Cited by 0SourceScholar
2024

FedA3I: Annotation Quality-Aware Aggregation for Federated Medical Image Segmentation against Heterogeneous Annotation Noise

AAAI 2024technical

Federated learning (FL) has emerged as a promising paradigm for training segmentation models on decentralized medical data, owing to its privacy-preserving property. However, existing research overlooks the prevalent annotation noise encountered in real-world medical datasets, which limits the perfo…

2024

From Optimization to Generalization: Fair Federated Learning against Quality Shift via Inter-Client Sharpness Matching

IJCAI 2024poster

Due to escalating privacy concerns, federated learning has been recognized as a vital approach for training deep neural networks with decentralized medical data. In practice, it is challenging to ensure consistent imaging quality across various institutions, often attributed to equipment malfunction…

2024

Panoramic Image Inpainting with Gated Convolution and Contextual Reconstruction Loss

ICASSP 2024accepted

Deep learning-based methods have demonstrated encouraging results in tackling the task of panoramic image inpainting. However, it is challenging for existing methods to distinguish valid pixels from invalid pixels and find suitable references for corrupted areas, thus leading to artifacts in the inp…

Cited by 0SourceScholar
2024

Pointsoup: High-Performance and Extremely Low-Decoding-Latency Learned Geometry Codec for Large-Scale Point Cloud Scenes

IJCAI 2024poster

Despite considerable progress being achieved in point cloud geometry compression, there still remains a challenge in effectively compressing large-scale scenes with sparse surfaces. Another key challenge lies in reducing decoding latency, a crucial requirement in real-world application. In this pape…

2024

Unsupervised Anomaly Detection via Masked Diffusion Posterior Sampling

IJCAI 2024poster

Reconstruction-based methods have been commonly used for unsupervised anomaly detection, in which a normal image is reconstructed and compared with the given test image to detect and locate anomalies. Recently, diffusion models have shown promising applications for anomaly detection due to their pow…

Cited by 3SourcePDFScholar
2023

FedNoRo: Towards Noise-Robust Federated Learning by Addressing Class Imbalance and Label Noise Heterogeneity

IJCAI 2023poster

Federated noisy label learning (FNLL) is emerging as a promising tool for privacy-preserving multi-source decentralized learning. Existing research, relying on the assumption of class-balanced global data, might be incapable to model complicated label noise, especially in medical scenarios. In this…

2022

Constrained Variable Impedance Control using Quadratic Programming

ICRA 2022poster

This paper proposes a quadratic programming (QP)-based variable impedance control (VIC) algorithm to solve contact-rich trajectory tracking problems with impedance, position and velocity constraints. To the best of our knowledge, the impedance constraints which are significant to ensure the worst co…

Cited by 6SourceScholar
2022

Denoising-Guided Deep Reinforcement Learning For Social Recommendation

ICASSP 2022accepted

Social recommendation (SR) aims to enhance the performance of recommendations by incorporating social information. However, such information is not always reliable, e.g., some of the friends may share similar preferences with the user on a specific item, while others may be irrelevant to this item d…

Cited by 0SourceScholar
2022

Denoising-Oriented Deep Hierarchical Reinforcement Learning for Next-Basket Recommendation⋆

ICASSP 2022accepted

Next basket recommendation aims to provide users a basket of items on the next visit by considering the sequence of their historical baskets. However, since a user’s purchase interests vary over time, historical baskets often contain many irrelevant items to his/her next choices. Therefore, it is ne…

Cited by 0SourceScholar
2022

Imbalance-Aware Uplift Modeling for Observational Data

AAAI 2022technical

Uplift modeling aims to model the incremental impact of a treatment on an individual outcome, which has attracted great interests of researchers and practitioners from different communities. Existing uplift modeling methods rely on either the data collected from randomized controlled trials (RCTs) o…

Cited by 6SourcePDFScholar
2022

MTAF: Shopping Guide Micro-Videos Popularity Prediction Using Multimodal and Temporal Attention Fusion Approach

ICASSP 2022accepted

Predicting the popularity of shopping guide micro-videos incorporating merchandise is crucial for online advertising. What are the significant factors affecting the popularity of the micro-video? How to extract and effectively fuse multiple modalities for the micro-video popularity prediction? This…

Cited by 0SourceScholar
2021

Co-Capsule Networks Based Knowledge Transfer for Cross-Domain Recommendation

ICASSP 2021accepted

Cross-domain recommendation (CDR) technology is proved to be an effective way to tackle the difficulties encountered by traditional recommender technology (e.g. CF), such as data sparsity and cold-start. However, on account of the heterogeneity, it is difficult to enhance the representation of user…

Cited by 0SourceScholar
2020

DA4AD: End-to-End Deep Attention-based Visual Localization for Autonomous Driving

ECCV 2020poster

We present a visual localization framework based on novel deep attention aware features for autonomous driving that achieves centimeter level localization accuracy. Conventional approaches to the visual localization problem rely on handcrafted features or human-made objects on the road. They are kno…

Cited by 57SourcePDFScholar