← Search

Xiaoqiang Li

12 accepted papers

2025

Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception

CVPR 2025poster

Large Vision-Language Models (LVLMs) have achieved impressive results across various cross-modal tasks. However, hallucinations, i.e., the models generating counterfactual responses, remain a challenge. Though recent studies have attempted to alleviate object perception hallucinations, they focus on…

2025

One Arrow, Two Hawks: Sharpness-aware Minimization for Federated Learning via Global Model Trajectory

ICML 2025poster

Federated learning (FL) presents a promising strategy for distributed and privacy-preserving learning, yet struggles with performance issues in the presence of heterogeneous data distributions. Recently, a series of works based on sharpness-aware minimization (SAM) have emerged to improve local lea…

2025

Optimized Gradient Clipping for Noisy Label Learning

AAAI 2025technical

Previous research has shown that constraining the gradient of loss function w.r.t. model-predicted probabilities can enhance the model robustness against noisy labels. These methods typically specify a fixed optimal threshold for gradient clipping through validation data to obtain the desired robust…

2025

ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision Experts

ICLR 2025poster

Vision-language (VL) learning requires extensive visual perception capabilities, such as fine-grained object recognition and spatial perception. Recent works typically rely on training huge models on massive datasets to develop these capabilities. As a more efficient alternative, this paper proposes…

Cited by 0SourcePDFScholar
2025

Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration

ICML 2025poster

Large Vision-Language Models (LVLMs) have manifested strong visual question answering capability. However, they still struggle with aligning the rationale and the generated answer, leading to inconsistent reasoning and incorrect responses. To this end, this paper introduces Self-Rationale Calibratio…

Cited by 0SourcePDFScholar
2024

DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic Segmentation

CVPR 2024poster

Recently One-stage Weakly Supervised Semantic Segmentation (WSSS) with image-level labels has gained increasing interest due to simplification over its cumbersome multi-stage counterpart. Limited by the inherent ambiguity of Class Activation Map (CAM) we observe that one-stage pipelines often encoun…

2024

Generating and Reweighting Dense Contrastive Patterns for Unsupervised Anomaly Detection

AAAI 2024technical

Recent unsupervised anomaly detection methods often rely on feature extractors pretrained with auxiliary datasets or on well-crafted anomaly-simulated samples. However, this might limit their adaptability to an increasing set of anomaly detection tasks due to the priors in the selection of auxiliary…

Cited by 18SourcePDFScholar
2023

Active Negative Loss Functions for Learning with Noisy Labels

NeurIPS 2023poster

Robust loss functions are essential for training deep neural networks in the presence of noisy labels. Some robust loss functions use Mean Absolute Error (MAE) as its necessary component. For example, the recently proposed Active Passive Loss (APL) uses MAE as its passive loss function. However, MAE…

2023

GradPU: Positive-Unlabeled Learning via Gradient Penalty and Positive Upweighting

AAAI 2023technical

Positive-unlabeled learning is an essential problem in many real-world applications with only labeled positive and unlabeled data, especially when the negative samples are difficult to identify. Most existing positive-unlabeled learning methods will inevitably overfit the positive class to some exte…

Cited by 8SourcePDFScholar
2023

Hierarchical Semantic Contrast for Weakly Supervised Semantic Segmentation

IJCAI 2023poster

Weakly supervised semantic segmentation (WSSS) with image-level annotations has achieved great processes through class activation map (CAM). Since vanilla CAMs are hardly served as guidance to bridge the gap between full and weak supervision, recent studies explore semantic representations to make C…

2022

Occlusion-Robust Face Alignment Using a Viewpoint-Invariant Hierarchical Network Architecture

CVPR 2022oral

The occlusion problem heavily degrades the localization performance of face alignment. Most current solutions for this problem focus on annotating new occlusion data, introducing boundary estimation, and stacking deeper models to improve the robustness of neural networks. However, the performance de…

Cited by 18PDFcodeScholar
2021

Improving Robustness of Facial Landmark Detection by Defending Against Adversarial Attacks

ICCV 2021poster

Many recent developments in facial landmark detection have been driven by stacking model parameters or augmenting annotations. However, three subsequent challenges remain, including 1) an increase in computational overhead, 2) the risk of overfitting caused by increasing model parameters, and 3) the…

Cited by 35PDFcodeScholar