← Search

Ming Yang

105 accepted papers

2026

ACTIVE-o3 : Empowering MLLMs with Active Perception via Pure Reinforcement Learning

ICML 2026poster

Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of efficient perception and decision-making in humans and advanced embodied agents. With the rise of Multimodal Large Language M…

Cited by 0SourceScholar
2026

DSAP: Enhancing Generalization in Goal-Conditioned Reinforcement Learning

AAAI 2026technical

Goal-conditioned Reinforcement Learning (RL) is a promising direction for training agents capable of tackling a variety of tasks. However, generalizing to new goals in different environments remains a central challenge for goal-conditioned RL agents. Existing methods often rely on state abstraction,

Cited by 0SourcePDFScholar
2026

Distributional Priors Guided Diffusion for Generating 3D Molecules in Low Data Regimes

AAAI 2026technical

Can we train a 3D molecule generator using data from dense regions to generate samples in sparse regions? This challenge can be framed as an out-of-distribution (OOD) generation problem. While prior research on OOD generation predominantly targets property shifts, structural shifts, such as differen

Cited by 0SourcePDFScholar
2026

LineageFlow: Flow Matching for High-Fidelity Family-Aware Protein Sequence Generation

ICML 2026poster

Protein sequence generation for engineering requires samples that are biophysically plausible and, when targeting a family/domain, remain recognizable members while exploring within-family diversity. Current discrete generative models typically start from uniform or masked-token noise, which discard…

Cited by 0SourceScholar
2026

Multi-Objective Protein Design via Memory-Aware Test-Time Scaling in Diffusion Models

ICML 2026poster

Multi-objective protein design is essential for meeting the complex demands of synthetic biology. To adapt to shifting multi-functional targets without the prohibitive cost of retraining, test-time scaling has emerged as a flexible, training-free alternative. However, current test-time diffusion met…

Cited by 0SourceScholar
2026

On the Learnability of Test-Time Adaptation: A Recovery Complexity Perspective

ICML 2026poster

Test-time adaptation (TTA) aims to adapt models to maintain reliable performance on non-stationary test streams without requiring labeled data. Despite its empirical success, the learnability of TTA under distributional non-stationarity remains unexplored. A key challenge is lacking of a principled …

Cited by 0SourceScholar
2026

ROAD: Responsibility-Oriented Reward Design for Reinforcement Learning in Autonomous Driving

RA-L 2026

Reinforcement learning (RL) in autonomous driving employs a trial-and-error mechanism, enhancing robustness in unpredictable environments. However, crafting effective reward functions remains challenging, as conventional approaches rely heavily on manual design and demonstrate limited efficacy in co

Cited by 0SourceScholar
2026

SCAN: Self-Calibrated AutoregressioN for High-Quality Visual Generation

AAAI 2026technical

Human artists can continuously refine their coarse sketches during artistic creation. This is quite different from existing autoregressive generation, where a token is determined once sampled. Aiming to flexibly refine the generated contents, this paper presents a Self-Calibrated AutoregressioN (SCA

Cited by 0SourcePDFScholar
2026

Tensorized Label Learning via Balanced Tensor Regression

AAAI 2026technical

The multi-view clustering methods based on tensor regression can make full use of the potential structural information between views and achieve data-level fusion. However, existing tensor regression-based approaches for anchor graph often overlook the probabilistic nature of anchor graph, focusing

Cited by 0SourcePDFScholar
2026

Unified View Extraction with Low-Rankness and Smoothness Fusion for Multi-View Subspace Clustering

AAAI 2026technical

Tensor-based multi-view subspace clustering (MVSC) has achieved significant success by capturing high-order inter-view correlations. However, existing approaches face two principal limitations. First, most methods either exclusively emphasize the inter-view low‑rankness (R) prior while neglecting th

Cited by 0SourcePDFScholar
2025

AVP Scene Graph: Hierarchical Visual Language Mapping and Navigation for Autonomous Valet Parking

IROS 2025

Autonomous valet parking (AVP) aims to help the human drivers navigate to the desired location in the parking lot. Currently, the AVP task is not flexible enough to perform the open-vocabulary navigation tasks such as "navigate to the exit" or "park near the elevator". The widely used map formats fo

Cited by 0SourceScholar
2025

Animate-X: Universal Character Image Animation with Enhanced Motion Representation

ICLR 2025poster

Character image animation, which generates high-quality videos from a reference image and target pose sequence, has seen significant progress in recent years. However, most existing methods only apply to human figures, which usually do not generalize well on anthropomorphic characters commonly used…

Cited by 14SourcePDFScholar
2025

AutoSelecter: Efficient Synthetic Nighttime Images Make Object Detector Stronger

RA-L 2025

Object detection has achieved significant advancements despite the challenges posed by adverse conditions like low-light nighttime environments, where annotated data is not only scarce but also challenging to accurately label. Instead of designing special network, we focus on the creation and effici

Cited by 0SourceScholar
2025

BMIP: Bi-directional Modality Interaction Prompt Learning for VLM

IJCAI 2025

Vision-language models (VLMs) have exhibited remarkable generalization capabilities, and prompt learning for VLMs has attracted great attention for the ability to adapt pre-trained VLMs to specific downstream tasks. However, existing studies mainly focus on single-modal prompts or uni-directional mo

Cited by 0SourcePDFScholar
2025

Building Hybrid Omnidirectional Visual-Lidar Map for Visual-Only Localization

IROS 2025

Recently, there has been growing interest in using low-cost sensor combinations, such as cameras and IMUs, to achieve accurate localization within pre-built pointcloud maps. In this paper, we propose a novel hybrid visual-Lidar mapping and visual-only re-localization framework, specifically designed

Cited by 0SourceScholar
2025

CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance

ICCV 2025poster

Semi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we p…

2025

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding

CVPR 2025poster

The challenge in LLM-based video understanding lies in preserving visual and semantic information in long videos while maintaining a memory-affordable token count. However, redundancy and correspondence in videos have hindered the performance potential of existing methods. Through statistical learni…

Cited by 2SourcePDFScholar
2025

Embodied Escaping: End-to-End Reinforcement Learning for Robot Navigation in Narrow Environment

IROS 2025

Autonomous navigation is a fundamental task for robot vacuum cleaners in indoor environments. Since their core function is to clean entire areas, robots inevitably encounter dead zones in cluttered and narrow scenarios. Existing planning methods often fail to escape due to complex environmental cons

Cited by 2SourceScholar
2025

Engage for All: Making Ordinary Image Descriptions Appealing Again!

ICCV 2025poster

In recent years, multi-modal large language models (MLLMs) have been successfully adopted to generate humorous and engaging descriptions for internet memes. While, it is challenging for the same approaches to apply to ordinary images which lack of inherent funny or exaggerated contents. Thus, crafti…

2025

FedSaaS: Class-Consistency Federated Semantic Segmentation via Global Prototype Supervision and Local Adversarial Harmonization

IJCAI 2025

Federated semantic segmentation enables pixel-level classification in images through collaborative learning while maintaining data privacy. However, existing research commonly overlooks the fine-grained class relationships within the semantic space when addressing heterogeneous problems, particularl

Cited by 0SourcePDFScholar
2025

Flow-Aware Navigation of Magnetic Micro-Robots in Complex Fluids via PINN-Based Prediction

IROS 2025

While magnetic micro-robots have demonstrated significant potential across various applications, including drug delivery and microsurgery, the open issue of precise navigation and control in complex fluid environments is crucial for in vivo implementation. This paper introduces a novel flow-aware na

Cited by 1SourceScholar
2025

From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots

NeurIPS 2025spotlight

Achieving general agile whole-body control on humanoid robots remains a major challenge due to diverse motion demands and data conflicts. While existing frameworks excel in training single motion-specific policies, they struggle to generalize across highly varied behaviors due to conflicting control…

Cited by 0SourceScholar
2025

GAP: a Global Adaptive Pruning Method for Large Language Models

EMNLP 2025

The deployment of Large Language Models (LLMs) faces significant challenges due to high computational costs,driving the demand for effective pruning techniques. Existing structured pruning methods employ uniform compression rates across network layers, neglecting the varying importance of different

2025

HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography Estimation

AAAI 2025technical

Feature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved substantial performance enhancements and established a widely-accepted coarse-to-fine paradigm. However, the majority…

Cited by 0SourcePDFScholar
2025

Human-Like Walking Motion Generation for Self-Balancing Lower Limb Rehabilitation Exoskeletons

ICRA 2025

Self-balancing lower limb rehabilitation exoskeletons (SLLREs) allow individuals with lower limb dysfunction to walk without the use of crutches. Stable and human-like walking motions are crucial for SLLREs because achieving a close imitation of healthy human walking is a key goal in rehabilitation

Cited by 0SourceScholar
2025

Hybrid Feature Collaborative Reconstruction Network for Few-Shot Fine-Grained Image Classification

ICASSP 2025accepted

Our research focuses on few-shot fine-grained image classification (FS-FGIC), which faces two main challenges: the similarity of fine-grained objects and a limited number of samples. Traditional feature reconstruction networks enhance key features through spatial reconstruction and error minimizatio…

Cited by 0SourceScholar
2025

Learning to Follow Infrared Prior Repersentation for Image Dehazing

ICASSP 2025accepted

The infrared image can distinguish the targets from the background based on radiation differences, providing more significant target visibility under dense haze. Fusion of visible haze images with infrared prior representations can generate high-quality fused images for high-level tasks. Consequentl…

Cited by 0SourceScholar
2025

Mimir: Improving Video Diffusion Models for Precise Text Understanding

CVPR 2025poster

Text serves as the key control signal in video generation due to its narrative nature. To render text descriptions into video clips, current video diffusion models borrow features from text encoders yet struggle with limited text comprehension. The recent success of large language models (LLMs) show…

Cited by 4SourcePDFScholar
2025

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation

CVPR 2025poster

The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns, yet there lacks a reliable motion estimator for training such…

Cited by 4SourcePDFScholar
2025

RL-OGM-Parking: Lidar OGM-Based Hybrid Reinforcement Learning Planner for Autonomous Parking

ICRA 2025

Autonomous parking has become a critical application in automatic driving research and development. Parking operations often suffer from limited space and complex environments, requiring accurate perception and precise maneuvering. Traditional rule-based parking algorithms struggle to adapt to diver

Cited by 6SourceScholar
2025

Reversing Flow for Image Restoration

CVPR 2025poster

Image restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, wh…

Cited by 0SourcePDFScholar
2025

Robotic Sim-to-Real Transfer for Long-Horizon Pick-and-Place Tasks in the Robotic Sim2Real Competition

ICRA 2025

This paper presents a fully autonomous robotic system that performs sim-to-real transfer in complex longhorizon tasks involving navigation, recognition, grasping, and stacking in an environment with multiple obstacles. The key feature of the system is the ability to overcome typical sensing and actu

Cited by 0SourcecodeScholar
2025

Robust Seizure Prediction Based on Riemannian Manifold Enhanced Denoising Adversarial Autoencoder

ICASSP 2025accepted

The seizure early warning devices based on multichannel EEG signals is one of the most used assisted-living strategies for drug-resistant epileptic patients. One of the challenges in the development of these devices is that existing algorithms cannot avoid the effects of electrode loosening. To alle…

Cited by 0SourceScholar
2025

SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories

CVPR 2025poster

While MLLMs have demonstrated adequate image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications. Current evaluation tasks like VQA and visual grounding remain too coarse to assess fine-grained pixel comprehension accurately. Though s…

2025

SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling

CVPR 2025poster

Open-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two c…

2025

Social Debiasing for Fair Multi-modal LLMs

ICCV 2025poster

Multi-modal Large Language Models (MLLMs) have dramatically advanced the research field and delivered powerful vision-language understanding capabilities. However, these models often inherit deep-rooted social biases from their training data, leading to uncomfortable responses with respect to attrib…

Cited by 0SourcePDFScholar
2025

Stochastic Momentum Methods for Non-smooth Non-Convex Finite-Sum Coupled Compositional Optimization

NeurIPS 2025poster

Finite-sum Coupled Compositional Optimization (FCCO), characterized by its coupled compositional objective structure, emerges as an important optimization paradigm for addressing a wide range of machine learning problems. In this paper, we focus on a challenging class of non-convex non-smooth FCC…

Cited by 0SourceScholar
2025

Unified Video Generation via Next-Set Prediction in Continuous Domain

ICCV 2025poster

Existing video generation strategies can be categorized into two categories, i.e., the diffusion and autoregressive (AR) methods. While AR methods achieves high efficiency by predicting the next token based on known visual cues, they generally fall short of diffusion models in terms of video quality…

Cited by 0SourcePDFScholar
2025

VCSearch: Bridging the Gap Between Well-Defined and Ill-Defined Problems in Mathematical Reasoning

EMNLP 2025

Large language models (LLMs) have demonstrated impressive performance on reasoning tasks, including mathematical reasoning. However, the current evaluation mostly focuses on carefully constructed benchmarks and neglects the consideration of real-world reasoning problems that present missing or contr

Cited by 0SourcePDFScholar
2025

VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions

ACL 2025long

Complex video question-answering (VQA) requires in-depth understanding of video contents including object and action recognition as well as video classification and summarization, which exhibits great potential in emerging applications in education and entertainment, etc. Multimodal large language m…

2024

A Closed-loop Control for Lower Limb Exoskeleton Considering Overall Deformations: A Simple and Direct Application Method

IROS 2024

In this paper, considering overall deformations of the exoskeleton, we couple deformations relationship network (DRN) with fractional order viscoelastic (FOV) controller, proposing a novel DRN-FOV closed-loop control method, endowing exoskeleton with stable dynamic walking ability. Simply by utilizi

Cited by 0SourceScholar
2024

A Numerical Approximation Approach for Deriving Computational Efficient Inverse Dynamics of 6-DOF Parallel Robots Based on Principle of Virtual Work

RA-L 2024

The inverse dynamics of the six degree-of-freedom (6-DOF) parallel robot (PR) presents an inherent complexity due to the closed-loop kinematic chains. To derive computational efficient inverse dynamics for real-time control, this study presents a numerical approximation (NA) approach based on the pr

Cited by 4SourceScholar
2024

Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight

NeurIPS 2024poster

This paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs). Our approach employs a sequence of visual resamplers that capture visual details at various spacial scales. This architecture not only leverages globa…

Cited by 3SourcePDFScholar
2024

Active Vehicle Re-localization Based on Non-repetitive LiDAR with Gimbal Motion Strategy

IROS 2024poster

The installation of a multi-layer 3D LiDAR atop the vehicle is a widely adopted hardware configuration for map-matching-based localization in intelligent driving. By offering a comprehensive 360° horizontal Field of View (FoV), this setup aims to achieve precise matching outcomes through the imposit…

Cited by 0SourceScholar
2024

An Online Automatic Calibration Method for Infrastructure-Based LiDAR-Camera via Cross-modal Object Matching

IROS 2024poster

In indoor environments where the Global Navigation Satellite System (GNSS) isn’t available, the infrastructure-based LiDAR-camera joint array can provide high-precision localization for mobile robots, such as Autonomous Valet Parking (AVP). The primary challenge in employing the infrastructure-based…

Cited by 0SourceScholar
2024

CSPG: Crossing Sparse Proximity Graphs for Approximate Nearest Neighbor Search

NeurIPS 2024poster

The state-of-the-art approximate nearest neighbor search (ANNS) algorithm builds a large proximity graph on the dataset and performs a greedy beam search, which may bring many unnecessary explorations. We develop a novel framework, namely *corssing sparse proximity graph (CSPG)*, based on random par…

Cited by 1SourcePDFScholar
2024

Cross-Modal Registration Using Adaptive Modeling in Infrastructure-based Vehicle Localization*

ICRA 2024poster

Infrastructure-based vehicle localization, in comparison to single-agent approaches, offers several advantages including reduced system cost, extended perception range, enhanced data fusion capabilities, and energy savings. Many conventional approaches impose limitations on the types of objects due…

Cited by 1SourceScholar
2024

Cross-Modal Visual Relocalization in Prior LiDAR Maps Utilizing Intensity Textures

IROS 2024poster

Cross-modal localization has drawn increasing attention in recent years, while the visual relocalization in prior LiDAR maps is less studied. Related methods usually suffer from inconsistency between the 2D texture and 3D geometry, neglecting the intensity features in the LiDAR point cloud. In this…

Cited by 0SourceScholar
2024

DeCoOp: Robust Prompt Tuning with Out-of-Distribution Detection

ICML 2024poster

Vision-language models (VLMs), such as CLIP, have demonstrated impressive zero-shot capabilities for various downstream tasks. Their performance can be further enhanced through few-shot prompt tuning methods. However, current studies evaluate the performance of learned prompts separately on base and…

2024

EVE: Efficient Zero-Shot Text-Based Video Editing With Depth Map Guidance and Temporal Consistency Constraints

IJCAI 2024poster

Motivated by the superior performance of image diffusion models, more and more researchers strive to extend these models to the text-based video editing task. Nevertheless, current video editing tasks mainly suffer from the dilemma between the high fine-tuning cost and the limited generation capacit…

2024

EcoMatcher: Efficient Clustering Oriented Matcher for Detector-free Image Matching

ECCV 2024poster

"Detector-free local feature matching methods have demonstrated significant performance improvements since leveraging the power of Transformer architecture. The global receptive field allows for simultaneous interaction among all elements, proving particularly beneficial in regions with low texture…

Cited by 1SourcePDFScholar
2024

Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis

CVPR 2024poster

Recent works in implicit representations such as Neural Radiance Fields (NeRF) have advanced the generation of realistic and animatable head avatars from video sequences. These implicit methods are still confronted by visual artifacts and jitters since the lack of explicit geometric constraints pose…

2024

MOSFormer: A Transformer-based Multi-Modal Fusion Network for Moving Object Segmentation

IROS 2024poster

3D moving object segmentation (MOS) is vital for autonomous systems, providing essential information for downstream tasks like mapping and localization. However, current MOS methods face challenges due to the limitation of existing datasets, which are sparse in moving objects and limited in scene di…

Cited by 0SourceScholar
2024

MapLocNet: Coarse-to-Fine Feature Registration for Visual Re-Localization in Navigation Maps

IROS 2024

Robust localization is the cornerstone of autonomous driving, especially in challenging urban environments where GPS signals suffer from multipath errors. Traditional localization approaches rely on high-definition (HD) maps, which consist of precisely annotated landmarks. However, building HD map i

Cited by 29SourceScholar
2024

Monocular Localization with Semantics Map for Autonomous Vehicles

ICRA 2024poster

Accurate and robust localization remains a significant challenge for autonomous vehicles. The cost of sensors and limitations in local computational efficiency make it difficult to scale to large commercial applications. Traditional vision-based approaches focus on texture features that are suscepti…

Cited by 0SourceScholar
2024

ParkingE2E: Camera-based End-to-end Parking Network, from Images to Planning

IROS 2024

Autonomous parking is a crucial task in the intelligent driving field. Traditional parking algorithms are usually implemented using rule-based schemes. However, these methods are less effective in complex parking scenarios due to the intricate design of the algorithms. In contrast, neural-network-ba

Cited by 18SourcecodeScholar
2024

Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs

CVPR 2024poster

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities in various multi-modal tasks. Nevertheless their performance in fine-grained image understanding tasks is still limited. To address this issue this paper proposes a new framework to enhance the fine-grained image understand…

2024

Referencing Where to Focus: Improving Visual Grounding with Referential Query

NeurIPS 2024poster

Visual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted considerable attention, as they directly predict the coordinates of the target object without relying on additional effort…

Cited by 1SourcePDFScholar
2024

SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery

CVPR 2024poster

Prior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless these works primarily focus on a single modality without temporal and geo-context modeling hampering their capabilities for diverse tasks. In this study we pre…

Cited by 140SourcePDFScholar
2024

Stability and Generalization of Stochastic Compositional Gradient Descent Algorithms

ICML 2024poster

Many machine learning tasks can be formulated as a stochastic compositional optimization (SCO) problem such as reinforcement learning, AUC maximization and meta-learning, where the objective function involves a nested composition associated with an expectation. Although many studies have been devote…

Cited by 2SourcePDFScholar
2024

StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models

ECCV 2024poster

"Despite the burst of innovative methods for controlling the diffusion process, effectively controlling image styles in text-to-image generation remains a challenging task. Many adapter-based methods impose image representation conditions on the denoising process to accomplish image control. However…

2024

SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment

ICML 2024poster

Multimodal alignment between language and vision is the fundamental topic in current vision-language model research. Contrastive Captioners (CoCa), as a representative method, integrates Contrastive Language-Image Pretraining (CLIP) and Image Caption (IC) into a unified framework, resulting in impre…

Cited by 4SourcePDFScholar
2024

Towards Better Vision-Inspired Vision-Language Models

CVPR 2024poster

Vision-language (VL) models have achieved unprecedented success recently in which the connection module is the key to bridge the modality gap. Nevertheless the abundant visual clues are not sufficiently exploited in most existing methods. On the vision side most existing approaches only use the last…

Cited by 2SourcePDFScholar
2023

Centerless Multi-View K-means Based on the Adjacency Matrix

AAAI 2023technical

Although K-Means clustering has been widely studied due to its simplicity, these methods still have the following fatal drawbacks. Firstly, they need to initialize the cluster centers, which causes unstable clustering performance. Secondly, they have poor performance on non-Gaussian datasets. Inspir…

2023

Cross-Modal Monocular Localization in Prior LiDAR Maps Utilizing Semantic Consistency

ICRA 2023poster

Visual localization for mobile robots and intelligent vehicles in prior LiDAR maps can achieve high accuracy and low cost. However, algorithms for finding the cross-modal correspondences between images and LiDAR map points are not yet stable. In this paper, we propose a monocular visual localization…

Cited by 14SourceScholar
2023

Efficient Potential-based Exploration in Reinforcement Learning using Inverse Dynamic Bisimulation Metric

NeurIPS 2023poster

Reward shaping is an effective technique for integrating domain knowledge into reinforcement learning (RL). However, traditional approaches like potential-based reward shaping totally rely on manually designing shaping reward functions, which significantly restricts exploration efficiency and introd…

Cited by 9SourcePDFScholar
2023

Fast Robust Principle Component Analysis Using Gauss-Newton Iterations

ICASSP 2023accepted

Robust Principal Component Analysis (RPCA) is an optimization problem that decomposes a data matrix into a low-rank and a sparse matrix. However, solving this problem using alternating procedures requires sequentially computing singular value decompositions (SVDs) of large matrices, which is computa…

Cited by 0SourceScholar
2023

High-Level Semantic Feature Matters Few-Shot Unsupervised Domain Adaptation

AAAI 2023technical

In few-shot unsupervised domain adaptation (FS-UDA), most existing methods followed the few-shot learning (FSL) methods to leverage the low-level local features (learned from conventional convolutional models, e.g., ResNet) for classification. However, the goal of FS-UDA and FSL are relevant yet dis…

Cited by 2SourcePDFScholar
2023

Orthogonal Non-negative Tensor Factorization based Multi-view Clustering

NeurIPS 2023poster

Multi-view clustering (MVC) based on non-negative matrix factorization (NMF) and its variants have attracted much attention due to their advantages in clustering interpretability. However, existing NMF-based multi-view clustering methods perform NMF on each view respectively and ignore the impact of…

Cited by 32SourcePDFScholar
2023

TTC4MCP: Monocular Collision Prediction Based on Self-Supervised TTC Estimation

IROS 2023poster

Vision-based collision prediction for autonomous driving is a challenging task due to the dynamic movement of vehicles and diverse types of obstacles. Most existing methods rely on object detection algorithms, which only predict predefined collision targets, such as vehicles and pedestrians, and can…

Cited by 2SourceScholar
2022

ATF-3D: Semi-Supervised 3D Object Detection With Adaptive Thresholds Filtering Based on Confidence and Distance

RA-L 2022

Performance of current point cloud-based outdoor 3D object detection relies heavily on large-scale high-quality 3D annotations. However, such annotations are usually expensive to collect and outdoor scenes easily accumulate massive unlabeled data containing rich scenes. Semi-supervised learning is a

Cited by 12SourceScholar
2022

BAANet: Learning Bi-directional Adaptive Attention Gates for Multispectral Pedestrian Detection

ICRA 2022poster

Thermal infrared (TIR) image has proven effectiveness in providing temperature cues to the RGB features for multispectral pedestrian detection. Most existing methods directly inject the TIR modality into the RGB-based framework or simply ensemble the results of two modalities. This, however, could l…

Cited by 61SourceScholar
2022

G3DOA: Generalizable 3D Descriptor With Overlap Attention for Point Cloud Registration

RA-L 2022

Point cloud registration (PCR) is a key problem for robotics, autonomous driving, and other applications. Constructing generalizable 3D descriptors and determining whether a 3D descriptor is in the overlapping area are challenging tasks in PCR. Despite the fast evolution of learning-based 3D descrip

Cited by 11SourceScholar
2022

Group-based Interleaved Pipeline Parallelism for Large-scale DNN Training

ICLR 2022poster

The recent trend of using large-scale deep neural networks (DNN) to boost performance has propelled the development of the parallel pipelining technique for efficient DNN training, which has resulted in the development of several prominent pipelines such as GPipe, PipeDream, and PipeDream-2BW. Howev…

2022

HR-Planner: A Hierarchical Highway Tactical Planner based on Residual Reinforcement Learning

ICRA 2022poster

Tactical planning is crucial for safe and efficient driving on the highway. However, the problem is complicated by the uncertain intention of surrounding vehicles, as well as observation noise caused by measurement noise and perception errors. Rule-based tactical planning methods are ineffective in…

Cited by 3SourceScholar
2022

Inertial Navigation System Based Vehicle Temporal Relative Localization With Split Covariance Intersection Filter

RA-L 2022

In autonomous vehicle navigation, it usually involves a basic process of estimating the vehicle pose at one instant relative to the vehicle pose at a previous instant, and we refer to this process as temporal relative localization (TRL). Accurate and reliable TRL is valuable for vehicle localization

Cited by 10SourceScholar
2022

Joint Global-Local Alignment for Domain Adaptive Semantic Segmentation

ICASSP 2022accepted

Unsupervised domain adaptation has shown promising results in leveraging synthetic (source) images for semantic segmentation of real (target) images. One key issue is how to align data distributions between the source and target domains. Adversarial learning has been applied to align these distribut…

Cited by 0SourceScholar
2022

Momentum Accelerates the Convergence of Stochastic AUPRC Maximization

AISTATS 2022poster

In this paper, we study stochastic optimization of areas under precision-recall curves (AUPRC), which is widely used for combating imbalanced classification tasks. Although a few methods have been proposed for maximizing AUPRC, stochastic optimization of AUPRC with convergence guarantee remains an u…

Cited by 27SourcePDFScholar
2022

Towards Accurate Facial Motion Retargeting with Identity-Consistent and Expression-Exclusive Constraints

AAAI 2022technical

We address the problem of facial motion retargeting that aims to transfer facial motion from a 2D face image to 3D characters. Existing methods often formulate this problem as a 3D face reconstruction problem, which estimates the face attributes such as face identity and expression from face images.…

2021

Back-Tracing Representative Points for Voting-Based 3D Object Detection in Point Clouds

CVPR 2021poster

3D object detection in point clouds is a challenging vision task that benefits various applications for understanding the 3D visual world. Lots of recent research focuses on how to exploit end-to-end trainable Hough voting for generating object proposals. However, the current voting strategy can onl…

Cited by 126PDFcodeScholar
2021

CentroidReg: A Global-to-Local Framework for Partial Point Cloud Registration

RA-L 2021

Point cloud registration is a key problem for robotics, computer vision, and other applications. Previous global registration algorithms are sensitive to noises or partial occlusion, while local registration algorithms are highly dependent on initial angles. To solve these problems, we propose Centr

Cited by 13SourceScholar
2021

Recall and Learn: A Memory-augmented Solver for Math Word Problems

EMNLP 2021finding

In this article, we tackle the math word problem, namely, automatically answering a mathematical problem according to its textual description. Although recent methods have demonstrated their promising results, most of these methods are based on template-based generation scheme which results in limit…

2021

Robust Knowledge Transfer via Hybrid Forward on the Teacher-Student Model

AAAI 2021technical

When adopting deep neural networks for a new vision task, a common practice is to start with fine-tuning some off-the-shelf well-trained network models from the community. Since a new task may require training a different network architecture with new domain data, taking advantage of off-the-shelf m…

Cited by 13SourcePDFScholar
2021

Stacked Homography Transformations for Multi-View Pedestrian Detection

ICCV 2021poster

Multi-view pedestrian detection aims to predict a bird's eye view (BEV) occupancy map from multiple camera views. This task is confronted with two challenges: how to establish the 3D correspondences from views to the BEV map and how to assemble occupancy information across views. In this paper, we p…

Cited by 56PDFScholar
2021

Track To Detect and Segment: An Online Multi-Object Tracker

CVPR 2021poster

Most online multi-object trackers perform object detection stand-alone in a neural net without any input from tracking. In this paper, we present a new online joint detection and tracking model, TraDeS (TRAck to DEtect and Segment), exploiting tracking clues to assist detection end-to-end. TraDeS in…

Cited by 452PDFcodeScholar
2020

3D Instance Embedding Learning With a Structure-Aware Loss Function for Point Cloud Segmentation

RA-L 2020

This letter presents a framework for 3D instance segmentation on point clouds. A 3D convolutional neural network is used as the backbone to generate semantic predictions and instance embeddings simultaneously. In addition to the embedding information, point clouds also provide 3D geometric informati

Cited by 33SourceScholar
2020

ROI-cloud: A Key Region Extraction Method for LiDAR Odometry and Localization

ICRA 2020poster

We present a novel key region extraction method of point cloud, ROI-cloud, for LiDAR odometry and localization with autonomous robots. Traditional methods process massive point cloud data in every region within the field of view. In dense urban environments, however, processing redundant and dynamic…

Cited by 18SourceScholar
2020

Temporal-Context Enhanced Detection of Heavily Occluded Pedestrians

CVPR 2020poster

State-of-the-art pedestrian detectors have performed promisingly on non-occluded pedestrians, yet they are still confronted by heavy occlusions. Although many previous works have attempted to alleviate the pedestrian occlusion issue, most of them rest on still images. In this paper, we exploit the l…

Cited by 69PDFScholar
2019

Bi-Directional Cascade Network for Perceptual Edge Detection

CVPR 2019poster

Exploiting multi-scale representations is critical to improve edge detection for objects at different scales. To extract edges at dramatically different scales, we propose a Bi-Directional Cascade Network (BDCN) structure, where an individual layer is supervised by labeled edges at its specific scal…

Cited by 551PDFcodeScholar
2019

Hierarchical Depthwise Graph Convolutional Neural Network for 3D Semantic Segmentation of Point Clouds

ICRA 2019poster

This paper proposes a hierarchical depthwise graph convolutional neural network (HDGCN) for point cloud semantic segmentation. The main chanllenge for learning on point clouds is to capture local structures or relationships. Graph convolution has the strong ability to extract local shape information…

Cited by 114SourceScholar
2019

SSAP: Single-Shot Instance Segmentation With Affinity Pyramid

ICCV 2019poster

Recently, proposal-free instance segmentation has received increasing attention due to its concise and efficient pipeline. Generally, proposal-free methods generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, pr…

Cited by 316PDFScholar
2018

BSN: Boundary Sensitive Network for Temporal Action Proposal Generation

ECCV 2018poster

Temporal action proposal generation is an important yet challenging problem, since temporal proposals with rich action content are indispensable for analysing real-world videos with long duration and high proportion irrelevant content. This problem requires methods not only generating proposals with…

2018

Conditional Generative Adversarial Network for Structured Domain Adaptation

CVPR 2018poster

In recent years, deep neural nets have triumphed over many computer vision problems, including semantic segmentation, which is a critical task in emerging autonomous driving and medical image diagnostics applications. In general, training deep neural nets requires a humongous amount of labeled data,…

Cited by 374SourcePDFScholar
2018

Deep Reinforcement Learning with Iterative Shift for Visual Tracking

ECCV 2018poster

Visual tracking is confronted by the dilemma to locate a target both}accurately and efficiently, and make decisions online whether and how to adapt the appearance model or even restart tracking. In this paper, we propose a deep reinforcement learning with iterative shift (DRL-IS) method for single o…

Cited by 79SourcePDFScholar
2018

Image Blind Denoising With Generative Adversarial Network Based Noise Modeling

CVPR 2018poster

In this paper, we consider a typical image blind denoising problem, which is to remove unknown noise from noisy images. As we all know, discriminative learning based methods, such as DnCNN, can achieve state-of-the-art denoising results, but they are not applicable to this problem due to the lack of…

Cited by 728SourcePDFScholar
2018

Instance-level Human Parsing via Part Grouping Network

ECCV 2018poster

Instance-level human parsing towards real-world human analysis scenarios is still under-explored due to the absence of sufficient data resources and technical difficulty in parsing multiple instances in a single pass. Several related works all follow the ``parsing-by-detection" pipeline that heavily…

2017

Gaussian mixture model-signature quadratic form distance based point set registration

IROS 2017poster

Point set registration is a long addressed problem in lots of pattern recognition tasks. This paper presents a robust point set registration algorithm based on optimization of distance between two probability distributions. A major problem encountered in the point to point algorithms is the definiti…

Cited by 7SourceScholar