← Search

Lingxiao Yang

17 accepted papers

2025

OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model

NeurIPS 2025oral

Understanding and synthesizing realistic 3D hand-object interactions (HOI) is critical for applications ranging from immersive AR/VR to dexterous robotics. Existing methods struggle with generalization, performing well on closed-set objects and predefined tasks but failing to handle unseen objects o…

Cited by 0SourceScholar
2025

Training-Free Class Purification for Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Fine-tuning pre-trained vision-language models has emerged as a powerful approach for enhancing open-vocabulary semantic segmentation (OVSS). However, the substantial computational and resource demands associated with training on large datasets have prompted interest in training-free methods for OVS…

2024

Guidance with Spherical Gaussian Constraint for Conditional Diffusion

ICML 2024poster

Recent advances in diffusion models attempt to handle conditional generative tasks by utilizing a differentiable loss function for guidance without the need for additional training. While these methods achieved certain success, they often compromise on sample quality and require small guidance step…

2024

MMA: Multi-Modal Adapter for Vision-Language Models

CVPR 2024poster

Pre-trained Vision-Language Models (VLMs) have served as excellent foundation models for transfer learning in diverse downstream tasks. However tuning VLMs for few-shot generalization tasks faces a discrimination -- generalization dilemma i.e. general knowledge should be preserved and task-specific…

2024

Spike-Temporal Latent Representation for Energy-Efficient Event-to-Video Reconstruction

ECCV 2024poster

"Event-to-Video (E2V) reconstruction aims to recover grayscale video from neuromorphic event streams, with Spiking Neural Networks (SNNs) being promising energy-efficient models for this task. Event voxels effectively compress event streams for E2V reconstruction, yet their temporal latent represent…

Cited by 2SourcePDFScholar
2023

CuNeRF: Cube-Based Neural Radiance Field for Zero-Shot Medical Image Arbitrary-Scale Super Resolution

ICCV 2023poster

Medical image arbitrary-scale super-resolution (MIASSR) has recently gained widespread attention, aiming to supersample medical volumes at arbitrary scales via a single model. However, existing MIASSR methods face two major limitations: (i) reliance on high-resolution (HR) volumes and (ii) limited g…

Cited by 40PDFcodeScholar
2023

Neural Prediction Errors enable Analogical Visual Reasoning in Human Standard Intelligence Tests

ICML 2023poster

Deep neural networks have long been criticized for lacking the ability to perform analogical visual reasoning. Here, we propose a neural network model to solve Raven's Progressive Matrices (RPM) - one of the standard intelligence tests in human psychology. Specifically, we design a reasoning block b…

2023

RuleMatch: Matching Abstract Rules for Semi-supervised Learning of Human Standard Intelligence Tests

IJCAI 2023poster

Raven's Progressive Matrices (RPM), one of the standard intelligence tests in human psychology, has recently emerged as a powerful tool for studying abstract visual reasoning (AVR) abilities in machines. Although existing computational models for RPM problems achieve good performance, they require a…

2023

Spike Count Maximization for Neuromorphic Vision Recognition

IJCAI 2023poster

Spiking Neural Networks (SNNs) are the promising models of neuromorphic vision recognition. The mean square error (MSE) and cross-entropy (CE) losses are widely applied to supervise the training of SNNs on neuromorphic datasets. However, the relevance between the output spike counts and predictions…

2023

Texture-Guided Saliency Distilling for Unsupervised Salient Object Detection

CVPR 2023poster

Deep Learning-based Unsupervised Salient Object Detection (USOD) mainly relies on the noisy saliency pseudo labels that have been generated from traditional handcraft methods or pre-trained networks. To cope with the noisy labels problem, a class of methods focus on only easy samples with reliable l…

2022

A Weighting-Based Tabu Search Algorithm for the p-Next Center Problem

IJCAI 2022poster

The p-next center problem (pNCP) is an extension of the classical p-center problem. It consists of locating p centers from a set of candidate centers and allocating both a reference and a backup center to each client, to minimize the maximum cost, which is the length of the path from a client to its…

Cited by 0SourcePDFScholar
2022

AcroFOD: An Adaptive Method for Cross-Domain Few-Shot Object Detection

ECCV 2022poster

"Under the domain shift, cross-domain few-shot object detection aims to adapt object detectors in the target domain with a few annotated target data. There exists two significant challenges: (1) Highly insufficient target domain data; (2) Potential over-adaptation and misleading caused by inappropri…

2022

Exploring Dual-Task Correlation for Pose Guided Person Image Generation

CVPR 2022poster

Pose Guided Person Image Generation (PGPIG) is the task of transforming a person image from the source pose to a given target pose. Most of the existing methods only focus on the ill-posed source-to-target task and fail to capture reasonable texture mapping. To address this problem, we propose a nov…

Cited by 102PDFcodeScholar
2022

Self-Supervised Image-Specific Prototype Exploration for Weakly Supervised Semantic Segmentation

CVPR 2022poster

Weakly Supervised Semantic Segmentation (WSSS) based on image-level labels has attracted much attention due to low annotation costs. Existing methods often rely on Class Activation Mapping (CAM) that measures the correlation between image pixels and classifier weight. However, the classifier focuses…

Cited by 192PDFcodeScholar
2021

SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks

ICML 2021spotlight

In this paper, we propose a conceptually simple but very effective attention module for Convolutional Neural Networks (ConvNets). In contrast to existing channel-wise and spatial-wise attention modules, our module instead infers 3-D attention weights for the feature map in a layer without adding par…

2020

Interactive Two-Stream Decoder for Accurate and Fast Saliency Detection

CVPR 2020poster

Recently, contour information largely improves the performance of saliency detection. However, the discussion on the correlation between saliency and contour remains scarce. In this paper, we first analyze such correlation and then propose an interactive two-stream decoder to explore multiple cues,…

Cited by 444PDFcodeScholar
2019

Dynamic Anchor Feature Selection for Single-Shot Object Detection

ICCV 2019poster

The design of anchors is critical to the performance of one-stage detectors. Recently, the anchor refinement module (ARM) has been proposed to adjust the initialization of default anchors, providing the detector a better anchor reference. However, this module brings another problem: all pixels at a…

Cited by 53PDFScholar