← Search

Lei Sun

50 accepted papers

2026

AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving

ICLR 2026poster

Vision–Language–Action (VLA) models in autonomous driving systems have recently demonstrated transformative potential by integrating multimodal perception with decision-making capabilities. However, the interpretability and coherence of the decision process and the plausibility of action sequences r…

Cited by 0SourceScholar
2026

Bridging the Perception Gap in Image Super-Resolution Evaluation

CVPR 2026

As super-resolution (SR) techniques advance, we observe a growing distrust of evaluation metrics in recent SR research. An inconsistency often emerges between certain evaluation criteria and human perceptual preference. Although current SR research employs varying metrics to evaluate SR performance,

Cited by 0SourceScholar
2026

From Scale to Speed: Adaptive Test-Time Scaling for Image Editing

CVPR 2026

Image Chain-of-Thought (Image-CoT) is a test-time scaling paradigm that improves image generation by extending inference time. Most Image-CoT methods focus on text-to-image (T2I) generation. Unlike T2I generation, image editing is goal-directed: the solution space is constrained by the source image

Cited by 0SourceScholar
2026

Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers

CVPR 2026

Region-instructed layout control in text-to-image generation is highly practical, yet existing methods suffer from limitations: (i) training-based approaches inherit data bias and often degrade image quality, and (ii) current techniques struggle with occlusion order, limiting real-world usability. T

Cited by 0SourcecodeScholar
2026

Learning Latent Transmission and Glare Maps for Lens Veiling Glare Removal

CVPR 2026

Beyond the commonly recognized optical aberrations, the imaging performance of simplified optical systems--including single-lens and metalens designs--is often further degraded by veiling glare caused by stray-light scattering from non-ideal optical surfaces and coatings, particularly in complex rea

Cited by 0SourcecodeScholar
2026

PG-Match: A Pose-Guided Generalizable Framework for Semi-Dense Feature Matching

ICRA 2026poster

Feature matching is a fundamental technique in visual perception, essential for tasks such as 3D reconstruction, SLAM, and visual localization. Existing detector-free methods often struggle with generalization due to their reliance on depth data, which is not available in many datasets. We propose P…

Cited by 0Scholar
2026

RaCo-SLAM: A Physics-Informed 4D Radar SLAM with Co-Visibility Consistency Factor

ICRA 2026poster

Robust all-weather localization is a critical capability for autonomous systems. While 4D mmWave radar offers superior resilience to adverse environmental conditions compared to LiDAR and cameras, its application in high-precision Simultaneous Localization and Mapping (SLAM) is hindered by significa…

Cited by 0codeScholar
2026

RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space

ICML 2026poster

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background video, human trajectory, and action patterns. In this paper, …

Cited by 0SourceScholar
2026

SCALAR: Scale-wise Controllable Visual Autoregressive Learning

AAAI 2026technical

Controllable image synthesis, which enables fine-grained control over generated outputs, has emerged as a key focus in visual generative modeling. However, controllable generation remains challenging for Visual Autoregressive (VAR) models due to their hierarchical, next-scale prediction style. Exist

Cited by 0SourcePDFScholar
2026

Semantic Context Matters: Improving Conditioning for Autoregressive Models

CVPR 2026

Recently, autoregressive (AR) models have shown strong potential in image generation, offering better scalability and easier integration with unified multi-modal models compared to diffusion methods.However, extending AR models to controllable image editing remains challenging due to weak and ineffi

Cited by 0SourcecodeScholar
2026

Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysis

CVPR 2026

Prevalent Computational Aberration Correction (CAC) methods are typically tailored to specific optical systems, leading to poor generalization and labor-intensive re-training for new lenses.Developing CAC paradigms capable of generalizing across diverse photographic lenses offers a promising solutio

Cited by 0SourcecodeScholar
2026

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools

ICLR 2026poster

Multimodal large language models (MLLMs) have demonstrated remarkable potential in bridging visual and textual reasoning, yet their reliance on text-centric priors often limits their ability to disentangle semantically similar actions in open-vocabulary scenarios. To address this, we propose Video-S…

Cited by 0SourceScholar
2025

Latent Swap Joint Diffusion for 2D Long-Form Latent Generation

ICCV 2025poster

This paper introduces Swap Forward (SaFa), a modality-agnostic and efficient method to generate seamless and coherent long spectrum and panorama using a latent swap joint diffusion process across multi-views. We first investigate spectrum aliasing problem in spectrum-based audio generation caused by…

2025

Learning-based Keypoints Detection with Topological Order on Deformable Linear Objects from Incomplete Point Clouds

IROS 2025

Detection of deformable linear objects (DLOs) in three-dimensional space is essential for robotic manipulation of DLOs. However, their complex deformations and high degrees of freedom make perception highly susceptible to occlusions, noise, and data missing. To address these challenges, we propose a

Cited by 0SourceScholar
2025

Low-Light Image Enhancement Using Event-Based Illumination Estimation

ICCV 2025poster

Low-light image enhancement (LLIE) aims to improve the visibility of images captured in poorly lit environments. Prevalent event-based solutions primarily utilize events triggered by motion, i.e., "motion events" to strengthen only the edge texture, while leaving the high dynamic range and excellent…

2025

VoxEKF-RIO: A 4D Radar Inertial Odometry Based on Incremental Voxel Map and Iterated Kalman Filter

IROS 2025

4D mmWave radar provides the point cloud with range, azimuth, elevation, Doppler velocity and operates normally in severe weather conditions. However, due to wavelength characteristics, the noisy and sparse point cloud that 4D radar collects poses great challenges for SLAM research. In this paper, w

Cited by 0SourceScholar
2024

A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition

ICASSP 2024accepted

Deep learning (DL)-based speaker diarization methods have proven powerful performance comparing to traditional clustering-based methods for multi-talker speech diarization and recognition in farfield scenes. However, most DL-based approaches cannot utilize the spatial information well due to the poo…

Cited by 0SourceScholar
2024

Bio-Inspired Pupal-Mode Actuator with Ultra-Crossing Capability for Soft Robots

ICRA 2024poster

Robot-assisted Natural Orifice Translu-minal Endoscopic Surgery (NOTES) represents a paradigm shift in surgical practice, significantly mini-mizing patient morbidity. However, the variability of inner diameter and the inter-luminal crossing within the luminal tracts lead to challenge for effective r…

Cited by 0SourceScholar
2024

Implicit Enhancement of Target Speaker in Speaker-Adaptive ASR through Efficient Joint Optimization

ICASSP 2024accepted

In multi-speaker scenarios, automatic speech recognition (ASR) models rely on pre-processed audio after speaker separation. However, when the target speaker is not accurately separated, ASR models face limitations in reaching their peak performance. To address this issue, we propose a speaker-adapti…

Cited by 0SourceScholar
2024

MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models

IJCAI 2024poster

Foundation models have demonstrated significant emergent abilities, holding great promise for enhancing embodied agents' reasoning and planning capacities. However, the absence of a comprehensive benchmark for evaluating embodied agents with multimodal observations in complex environments remains a…

2024

ODA: Observation-Driven Agent for integrating LLMs and Knowledge Graphs

ACL 2024findings

The integration of Large Language Models (LLMs) and knowledge graphs (KGs) has achieved remarkable success in various natural language processing tasks. However, existing methodologies that integrate LLMs and KGs often navigate the task-solving process solely based on the LLM’s analysis of the quest…

2024

Outlier-Robust Geometric Perception: A Novel Thresholding-Based Estimator with Intra-Class Variance Maximization

IROS 2024poster

Geometric perception problems are fundamental tasks in robotics and computer vision. In real-world applications, they often encounter the inevitable issue of outliers, preventing traditional algorithms from making correct estimates. In this paper, we present a novel general-purpose robust estimator…

Cited by 0SourcecodeScholar
2024

Revealing Personality Traits: A New Benchmark Dataset for Explainable Personality Recognition on Dialogues

EMNLP 2024main

Personality recognition aims to identify the personality traits implied in user data such as dialogues and social media posts. Current research predominantly treats personality recognition as a classification task, failing to reveal the supporting evidence for the recognized personality. In this pap…

2023

A Question-Answering Approach to Key Value Pair Extraction from Form-Like Document Images

AAAI 2023technical

In this paper, we present a new question-answering (QA) based key-value pair extraction approach, called KVPFormer, to robustly extracting key-value relationships between entities from form-like document images. Specifically, KVPFormer first identifies key entities from all entities in an image with…

Cited by 13SourcePDFScholar
2023

An Experimental Study on Sound Event Localization and Detection Under Realistic Testing Conditions

ICASSP 2023accepted

We study four data augmentation (DA) techniques and two model architectures on realistic data for sound event localization and detection (SELD). First, based on ResNet-Conformer (RC), we compare the four DA approaches on the realistic DCASE 2022 SELD test set which is often not easy to handle due to…

Cited by 0SourceScholar
2023

Event-Based Frame Interpolation With Ad-Hoc Deblurring

CVPR 2023poster

The performance of video frame interpolation is inherently correlated with the ability to handle motion in the input scene. Even though previous works recognize the utility of asynchronous event information for this task, they ignore the fact that motion may or may not result in blur in the input vi…

2023

Reducing the GAP Between Streaming and Non-Streaming Transducer-Based ASR by Adaptive Two-Stage Knowledge Distillation

ICASSP 2023accepted

Transducer is one of the mainstream frameworks for streaming speech recognition. There is a performance gap between the streaming and non-streaming transducer models due to limited context. To reduce this gap, an effective way is to ensure that their hidden and output distributions are consistent, w…

Cited by 0SourceScholar
2022

Event-Based Fusion for Motion Deblurring with Cross-Modal Attention

ECCV 2022poster

"Traditional frame-based cameras inevitably suffer from motion blur due to long exposure times. As a kind of bio-inspired camera, the event camera records the intensity changes in an asynchronous way with high temporal resolution, providing valid image degradation information within the exposure tim…

2022

Improving Separation-Based Speaker Diarization Via Iterative Model Refinement And Speaker Embedding Based Post-Processing

ICASSP 2022accepted

In this paper, we propose an iterative separation-based speaker diarization (ISSD) approach to cope with the realistic data conditions. In the proposed ISSD, we iteratively generate adaptation data ac-cording to speaker priors and fine-tune the separation model, which leads to a gradual performance…

Cited by 0SourceScholar
2022

On the Connection between Local Attention and Dynamic Depth-wise Convolution

ICLR 2022spotlight

Vision Transformer (ViT) attains state-of-the-art performance in visual recognition, and the variant, Local Vision Transformer, makes further improvements. The major component in Local Vision Transformer, local attention, performs the attention separately over small local windows. We rephrase local…

2022

RANSIC: Fast and Highly Robust Estimation for Rotation Search and Point Cloud Registration Using Invariant Compatibility

RA-L 2022

Correspondence-based rotation search and point cloud registration are two fundamental problems in robotics and computer vision. However, the presence of outliers, sometimes even occupying the great majority of the putative correspondences, can make many existing algorithms either fail or have very h

Cited by 32SourceScholar
2022

TriVoC: Efficient Voting-Based Consensus Maximization for Robust Point Cloud Registration With Extreme Outlier Ratios

RA-L 2022

Correspondence-based point cloud registration is a cornerstone in robotics perception and computer vision, which seeks to estimate the best rigid transformation aligning two point clouds from the putative correspondences. However, due to the limited robustness of 3D keypoint matching approaches, out

Cited by 25SourceScholar
2021

Conditional DETR for Fast Training Convergence

ICCV 2021poster

The recently-developed DETR approach applies the transformer encoder and decoder architecture to object detection and achieves promising performance. In this paper, we handle the critical issue, slow training convergence, and present a conditional cross-attention mechanism for fast DETR training. Ou…

Cited by 829PDFcodeScholar
2020

A Study of Child Speech Extraction Using Joint Speech Enhancement and Separation in Realistic Conditions

ICASSP 2020accepted

In this paper, we design a novel joint framework of speech enhancement and speech separation for child speech extraction in realistic conditions, targeting the problem of extracting child speech from daily conversations in BabyTrain mega corpus. To the best of our knowledge, it is the first discussi…

Cited by 0SourceScholar
2020

P-KDGAN: Progressive Knowledge Distillation with GANs for One-class Novelty Detection

IJCAI 2020poster

One-class novelty detection is to identify anomalous instances that do not conform to the expected normal instances. In this paper, the Generative Adversarial Networks (GANs) based on encoder-decoder-encoder pipeline are used for detection and achieve state-of-the-art performance. However, deep neur…

Cited by 0SourcePDFScholar
2020

Progressive Multi-Target Network Based Speech Enhancement with Snr-Preselection for Robust Speaker Diarization

ICASSP 2020accepted

In this paper, we design a novel front-end processing system for speaker diarization under realistic conditions with challenging background noises. To cope with diversified environments, we first extend our perviously proposed progressive learning based speech enhancement model by adding multi-task…

Cited by 0SourceScholar
2020

Real-Time Fusion Network for RGB-D Semantic Segmentation Incorporating Unexpected Obstacle Detection for Road-Driving Images

RA-L 2020

Semantic segmentation has made striking progress due to the success of deep convolutional neural networks. Considering the demands of autonomous driving, real-time semantic segmentation has become a research hotspot these years. However, few real-time RGB-D fusion semantic segmentation studies are c

Cited by 161SourcecodeScholar
2019

A Two-stage Single-channel Speaker-dependent Speech Separation Approach for Chime-5 Challenge

ICASSP 2019accepted

In this paper, we design a two-stage single-channel speaker-dependent speech separation approach for the CHiME-5 Challenge, targeting the problem of far-field and multi-talker conversational speech recognition in dinner party scenarios involving background noises, reverberations and overlapping spee…

Cited by 0SourceScholar
2019

Channel Adversarial Training for Cross-channel Text-independent Speaker Recognition

ICASSP 2019accepted

The conventional speaker recognition frameworks (e.g., the i-vector and CNN-based approach) have been successfully applied to various tasks when the channel of the enrolment dataset is similar to that of the test dataset. However, in real-world applications, mismatch always exists between these two…

Cited by 0SourceScholar
2018

A Novel LSTM-Based Speech Preprocessor for Speaker Diarization in Realistic Mismatch Conditions

ICASSP 2018accepted

In this study, we investigate on the effects of deep learning based speech enhancement as a preprocessor to speaker diarization in quite challenging realistic environments involving the background noises, reverberations and overlapping speech. To improve the generalization capability, the advanced l…

Cited by 0SourceScholar
2018

Enhancement and Analysis of Conversational Speech: JSALT 2017

ICASSP 2018accepted

Automatic speech recognition is more and more widely and effectively used. Nevertheless, in some automatic speech analysis tasks the state of the art is surprisingly poor. One of these is “diarization”, the task of determining who spoke when. Diarization is key to processing meeting audio and clinic…

Cited by 0SourceScholar
2018

Kinematic-Free Orientation Control for a Deformable Manipulator Based on the Geodesic in Rotation Group SO(3)

RA-L 2018

Orientation control is an important topic with many applications including robotic manipulation, spacecraft control, etc. The orientation control is particularly challenging when the system model is unknown due to the fact that the rotation group is a nonlinear manifold. In this letter, we propose a

Cited by 19SourceScholar
2017

Steady-state mean square performance of a sparsified kernel least mean square algorithm

ICASSP 2017accepted

In this paper, we investigate the convergence performance of a sparsified kernel least mean square (KLMS) algorithm in which the input is added into the dictionary only when the prediction error in amplitude is larger than a preset threshold. Under certain conditions, we derive an approximate value…

Cited by 0SourceScholar
2016

A parameter-free Cauchy-Schwartz information measure for independent component analysis

ICASSP 2016accepted

Independent component analysis (ICA) by an information measure has seen wide applications in engineering. Different from traditional probability density function based information measures, a probability survival distribution based Cauchy-Schwartz information measure for multiple variables is propos…

Cited by 0SourceScholar
2016

Nonlinear disturbance observer based torque control for series elastic actuator

IROS 2016poster

This paper presents a practical control approach for series elastic actuators(SEAs) to generate the desired torque. Specifically, the controller is applicable to both linear and nonlinear SEAs and it works well even in the presence of unknown payload parameters and external disturbances. Via the ana…

Cited by 11SourceScholar