← Search

Jun Li

122 accepted papers

2026

Accelerating Autoregressive Video Diffusion via History-Guided Cache and Residual Correction

CVPR 2026

Caching-based acceleration methods have recently driven significant progress in efficient video generation with diffusion models. However, we identify a critical limitation when directly applying these acceleration techniques to auto-regressive video diffusion models, which generate long videos by s

Cited by 0SourceScholar
2026

Adaptive Depth Lightweight RGB-T Tracking with Holistic Token Routing

CVPR 2026

The appeal of RGB-T tracking lies in its resilience when RGB fails under night scenes, glare, fog, and partial occlusion. Despite notable accuracy gains, recent architectures emphasize deep fusion and large parameter counts, driving up FLOPs and bandwidth. This computational burden constrains real-t

Cited by 0SourceScholar
2026

Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration

CVPR 2026

Text-to-image generative models have achieved impressive fidelity and diversity, but can inadvertently produce unsafe or undesirable content due to implicit biases embedded in large-scale training datasets.Existing concept erasure methods, whether text-only or image-assisted, face trade-offs: textua

Cited by 0SourceScholar
2026

CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model

CVPR 2026

Segment Anything Models (SAMs) are extensively used in computer vision for universal image segmentation, but deploying them on resource-constrained devices is challenging due to their high computational and memory demands. Post-Training Quantization (PTQ) is a widely used technique for model compres

Cited by 0SourceScholar
2026

ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure

ICML 2026poster

Large reasoning models (LRMs) typically solve reasoning-intensive tasks by generating long chain-of-thought (CoT) traces, leading to substantial inference overhead. We identify a reproducible inference-time phenomenon, termed \textbf{\emph{Self-Compression}}: when multiple independent and answerable…

Cited by 0SourceScholar
2026

DVAR: Dynamic Visual Autoregressive Modeling for Image Super-Resolution

CVPR 2026

Next-scale prediction paradigm visual autoregressive (VAR) models have demonstrated significant potential for image super-resolution. However, their practical application is constrained by a rigid, size-specific design. This limitation stems from their reliance on memorizing fixed, absolute scaling

Cited by 0SourcecodeScholar
2026

Depth Completion by Rescaling Monocular Depth Estimates Via Compressed Sensing

ICRA 2026poster

Depth completion is the challenge of recovering a dense depth map from an RGB image and corresponding sparse depth measurements. Many modern depth completion strategies often rely on deep neural networks, using a monocular depth estimation (MDE) backbone to generate an initial dense depth map from t…

Cited by 0Scholar
2026

Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases

ICML 2026poster

Clinical abnormality grounding for rare diseases is often hindered by data scarcity, rendering supervised fine-tuning infeasible and single-pass inference highly unstable. Thus, we propose Dynamic Decision Learning (DDL), a framework that enables frozen LVLMs to refine their decisions across languag…

Cited by 0SourceScholar
2026

Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval

CVPR 2026

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos based on text queries that describe only partial events. Existing methods suffer from incomplete global contextual perception, struggling with query ambiguity and local noise induced by spurious responses. To address these i

Cited by 0SourcecodeScholar
2026

MAPI-GNN: Multi-Activation Plane Interaction Graph Neural Network for Multimodal Medical Diagnosis

AAAI 2026technical

Graph neural networks are increasingly applied to multimodal medical diagnosis for their inherent relational modeling capabilities. However, their efficacy is often compromised by the prevailing reliance on a single, static graph built from indiscriminate features, hindering the ability to model pat

Cited by 0SourcePDFScholar
2026

No Way To Steal My Face: Proactive Defense Against Identity-Preserving Personalized Generation

CVPR 2026

Recent advances in diffusion models have enabled high-fidelity, identity-preserving image generation for personalized applications such as digital avatars and virtual try-on systems. However, their reliance on sensitive facial reference images raises growing privacy concerns. Existing defense mechan

Cited by 0SourceScholar
2026

RMLer: Synthesizing Novel Objects Across Diverse Categories via Reinforcement Mixing Learning

AAAI 2026technical

Novel object synthesis by integrating distinct textual concepts from diverse categories remains a significant challenge in text-to-image generation. Existing methods often suffer from insufficient concept mixing, lack of rigorous evaluation, and suboptimal outputs, resulting in conceptual imbalance,

Cited by 0SourcePDFScholar
2026

Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval

ICML 2026poster

Partially relevant video retrieval aims to retrieve untrimmed videos using text queries that describe only partial content. However, the inherent asymmetry between brief queries and rich video content inevitably introduces uncertainty into the retrieval process. In this setting, vague queries often …

Cited by 0SourceScholar
2026

Semantic-Level Conflict Traffic Scenario Generation Via Spatiotemporal Polygon Anchors

ICRA 2026poster

Autonomous Driving Systems (ADS) require rigorous and complex testing under diverse conditions to fulfill various demands and purposes of testing tasks, such as occlusion-triggered events, necessitating semantic-level control in scenario generation. Existing methods, reliant on low-level state contr…

Cited by 0Scholar
2026

Structure-Aware Riemannian Flow Matching for Registration and Fusion of Hyperspectral and Multispectral Images

ICML 2026poster

Precise alignment is a prerequisite for hyperspectral and multispectral image fusion, yet existing methods struggle with complex non-rigid deformations. Existing techniques either suffer from inter-task error accumulation by treating registration and fusion as disjoint processes or neglect the geome…

Cited by 0SourceScholar
2026

The Hidden Risk: Membership Inference Attacks on Multimodal Federated Learning via Modality Imbalance

ICML 2026poster

Federated learning (FL) faces significant challenges from modality heterogeneity, which motivates multimodal federated learning (MFL) to leverage complementary modalities across decentralized clients for improved performance. However, modality imbalance introduces a new attack surface, making MFL mo…

Cited by 0SourceScholar
2026

VMDiff: Visual Mixing Diffusion for Limitless Cross-Object Synthesis

ICLR 2026poster

Creating novel images by fusing visual cues from multiple sources is a fundamental yet underexplored problem in image-to-image generation, with broad applications in artistic creation, virtual reality and visual media. Existing methods often face two key challenges: coexistent generation, where mult…

Cited by 0SourceScholar
2025

Advancing Multiple Instance Learning with Continual Learning for Whole Slide Imaging

CVPR 2025highlight

Advances in medical imaging and deep learning have propelled progress in whole slide image (WSI) analysis, with multiple instance learning (MIL) showing promise for efficient and accurate diagnostics. However, conventional MIL models often lack adaptability to evolving datasets, as they rely on stat…

Cited by 0SourcePDFScholar
2025

AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing

CVPR 2025poster

Self-Supervised Video Hashing (SSVH) compresses videos into hash codes for efficient indexing and retrieval using unlabeled training videos. Existing approaches rely on random frame sampling to learn video features and treat all frames equally. This results in suboptimal hash codes, as it ignores fr…

2025

Completion as Enhancement: A Degradation-Aware Selective Image Guided Network for Depth Completion

CVPR 2025poster

In this paper, we introduce the Selective Image Guided Network (SigNet), a novel degradation-aware framework that transforms depth completion into depth enhancement for the first time. Moving beyond direct completion using convolutional neural networks (CNNs), SigNet initially densifies sparse dept…

Cited by 3SourcePDFScholar
2025

Convex Potential Mirror Langevin Algorithm for Efficient Sampling of Energy-Based Models

NeurIPS 2025poster

This paper introduces the Convex Potential Mirror Langevin Algorithm (CPMLA), a novel method to improve sampling efficiency for Energy-Based Models (EBMs). CPMLA uses mirror Langevin dynamics with a convex potential flow as a dynamic mirror map for EBM sampling. This dynamic mirror map enables targe…

Cited by 0SourceScholar
2025

Deep Unfolding Using Score-based Generative Networks for Automotive Radar Interference Mitigation

ICASSP 2025accepted

Automotive frequency-modulated continuous wave (FMCW) radars, essential in Advanced Driver Assistance Systems, encounter mutual interference issues that degrade their detection capabilities. Model-based algorithms, though widely used, rely heavily on predetermined assumptions about the statistical p…

Cited by 0SourceScholar
2025

Depth-Centric Dehazing and Depth-Estimation from Real-World Hazy Driving Video

AAAI 2025technical

In this paper, we study the challenging problem of simultaneously removing haze and estimating depth from real monocular hazy videos. These tasks are inherently complementary: enhanced depth estimation improves dehazing via the atmospheric scattering model (ASM), while superior dehazing contributes…

2025

DuCos: Duality Constrained Depth Super-Resolution via Foundation Model

ICCV 2025poster

We introduce DuCos, a novel depth super-resolution framework grounded in Lagrangian duality theory, offering a flexible integration of multiple constraints and reconstruction objectives to enhance accuracy and robustness. Our DuCos is the first to significantly improve generalization across diverse…

2025

Dynamic Dictionary Design for Localization in Automotive Radar Systems

ICASSP 2025accepted

This paper proposes a dynamic orthogonal matching pursuit (OMP)-based localization for automotive radar systems. At each time instant, three dictionaries are designed based on prior information on targets’ approximate locations. We use mutual coherence as the criterion for designing each dictionary.…

Cited by 0SourceScholar
2025

Efficient Self-Supervised Video Hashing with Selective State Spaces

AAAI 2025technical

Self-supervised video hashing (SSVH) is a practical task in video indexing and retrieval. Although Transformers are predominant in SSVH for their impressive temporal modeling capabilities, they often suffer from computational and memory inefficiencies. Drawing inspiration from Mamba, an advanced sta…

2025

Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

ICCV 2025poster

Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos…

2025

Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking

AAAI 2025technical

Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal re…

2025

Guided Real Image Dehazing Using YCbCr Color Space

AAAI 2025technical

Image dehazing, particularly with learning-based methods, has gained significant attention due to its importance in real-world applications. However, relying solely on the RGB color space often fall short, frequently leaving residual haze. This arises from two main issues: the difficulty in obtainin…

2025

Harmonious Music-driven Group Choreography with Trajectory-Controllable Diffusion

AAAI 2025technical

Creating group choreography from music is crucial in cultural entertainment and virtual reality, with a focus on generating harmonious movements. Despite growing interest, recent approaches often struggle with two major challenges: multi-dancer collisions and single-dancer foot sliding. To address…

Cited by 0SourcePDFScholar
2025

Learning Based MPC for Autonomous Driving Using a Low Dimensional Residual Model

ICRA 2025

In this paper, a learning based Model Predictive Control (MPC) using a low dimensional residual model is proposed for autonomous driving. One of the critical challenge in autonomous driving is the complexity of vehicle dynamics, which impedes the formulation of accurate vehicle model. Inaccurate veh

Cited by 8SourceScholar
2025

Learning Generalized Residual Exchange-Correlation-Uncertain Functional for Density Functional Theory

AAAI 2025technical

Density Functional Theory (DFT) stands as a widely used and efficient approach for addressing the many-electron Schrödinger equation across various domains such as physics, chemistry, and biology. However, a core challenge that persists over the long term pertains to refining the exchange-correlatio…

Cited by 0SourcePDFScholar
2025

NOVA: A Benchmark for Rare Anomaly Localization and Clinical Reasoning in Brain MRI

NeurIPS 2025oral

In many real-world applications, deployed models encounter inputs that differ from the data seen during training. Open-world recognition ensures that such systems remain robust as ever-emerging, previously _unknown_ categories appear and must be addressed without retraining. Foundation and vision-la…

Cited by 0SourceScholar
2025

Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose Estimation

AAAI 2025technical

With the rapid development of autonomous driving, LiDAR-based 3D Human Pose Estimation (3D HPE) is becoming a research focus. However, due to the noise and sparsity of LiDAR-captured point clouds, robust human pose estimation remains challenging. Most of the existing methods use temporal information…

2025

Quantum Multi-Path Communication Protocol Based on Maximum Flow Theory

ICASSP 2025accepted

Quantum networks are an actively researched and promising field, aiming to achieve efficient quantum information transmission by interconnecting quantum nodes. In large-scale quantum networks, end-to-end throughput is a critical factor that affects the overall performance of the network. The maximum…

Cited by 0SourceScholar
2025

RAGD: Regional-Aware Diffusion Model for Text-to-Image Generation

ICCV 2025poster

Regional prompting, or compositional generation, which enables fine-grained spatial control, has gained increasing attention for its practicality in real-world applications. However, previous methods either introduce additional trainable modules, thus only applicable to specific models, or manipulat…

2025

ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical Judgment

EMNLP 2025

Automatically generated radiology reports often receive high scores from existing evaluation metrics but fail to earn clinicians’ trust. This gap reveals fundamental flaws in how current metrics assess the quality of generated reports. We rethink the design and evaluation of these metrics and propos

2025

Reading Your Heart: Learning ECG Words and Sentences via Pre-training ECG Language Model

ICLR 2025poster

Electrocardiogram (ECG) is essential for the clinical diagnosis of arrhythmias and other heart diseases, but deep learning methods based on ECG often face limitations due to the need for high-quality annotations. Although previous ECG self-supervised learning (eSSL) methods have made significant pro…

2025

Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal Transport

CVPR 2025poster

Identifying multiple novel classes in an image, known as open-vocabulary multi-label recognition, is a challenging task in computer vision. Recent studies explore the transfer of powerful vision-language models such as CLIP. However, these approaches face two critical challenges: (1) The local seman…

2025

SFADNet: Spatio-temporal Fused Graph based on Attention Decoupling Network for Traffic Prediction

ICASSP 2025accepted

In recent years, traffic flow prediction has played a crucial role in the management of intelligent transportation systems. However, traditional prediction methods are often limited by static spatial modeling, making it difficult to accurately capture the dynamic and complex relationships between ti…

Cited by 0SourceScholar
2025

The Parallel Pneumatic Artificial Muscle Platform Based on RBF Neural Network Compensation

IROS 2025

A two-degree-of-freedom parallel mechanism control system based on an adaptive learning rate and radial basis function (RBF) neural network controller is studied in this paper. The mechanism is composed of four pneumatic artificial muscles(PAM), forming two pairs of antagonistic single-degree-of-fre

Cited by 0SourceScholar
2025

TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perception

CVPR 2025poster

Cooperative perception presents significant potential for enhancing the sensing capabilities of individual vehicles, however, inter-agent latency remains a critical challenge. Latencies cause misalignments in both spatial and semantic features, complicating the fusion of real-time observations from…

2025

UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic Programming

CVPR 2025award

Distributed learning is commonly used for training deep learning models, especially large models. In distributed learning, manual parallelism (MP) methods demand considerable human effort and have limited flexibility. Hence, automatic parallelism (AP) methods have recently been proposed for automati…

2025

Unveiling Environmental Sensitivity of Individual Gains in Influence Maximization

NeurIPS 2025poster

Influence Maximization (IM) seeks a seed set to maximize information dissemination in a network. Elegant IM algorithms could naturally extend to cases where each node is equipped with a specific weight, reflecting individual gains to measure its importance. In prevailing literature, these gains are…

Cited by 0SourceScholar
2025

V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative Perception

NeurIPS 2025spotlight

Modern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In…

Cited by 0SourceScholar
2024

A Spatial Calibration Method for Robust Cooperative Perception

RA-L 2024

Cooperative perception is a promising technique for intelligent and connected vehicles through vehicle-to-everything (V2X) cooperation, provided that accurate pose information and relative pose transforms are available. Nevertheless, obtaining precise positioning information often entails high costs

Cited by 37SourcecodeScholar
2024

AltNeRF: Learning Robust Neural Radiance Field via Alternating Depth-Pose Optimization

AAAI 2024technical

Neural Radiance Fields (NeRF) have shown promise in generating realistic novel views from sparse scene images. However, existing NeRF approaches often encounter challenges due to the lack of explicit 3D supervision and imprecise camera poses, resulting in suboptimal outcomes. To tackle these issues,…

Cited by 2SourcePDFScholar
2024

Bi2Lane: Bi-Directional Temporal Refinement with Bi-Level Feature Aggregation for 3D Lane Detection

ICRA 2024poster

Monocular 3D lane detection has recently received increasing research attention in autonomous driving due to its application effectiveness and simplicity. However, depending solely on the limited semantic information from a single image makes current monocular detection methods unable to deal with c…

Cited by 0SourceScholar
2024

Compound Text-Guided Prompt Tuning via Image-Adaptive Cues

AAAI 2024technical

Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable generalization capabilities to downstream tasks. However, existing prompt tuning based frameworks need to parallelize learnable textual inputs for all categories, suffering from massive GPU memory consumption when there is a lar…

2024

DCDepth: Progressive Monocular Depth Estimation in Discrete Cosine Domain

NeurIPS 2024poster

In this paper, we introduce DCDepth, a novel framework for the long-standing monocular depth estimation task. Moving beyond conventional pixel-wise depth estimation in the spatial domain, our approach estimates the frequency coefficients of depth patches after transforming them into the discrete cos…

2024

DMKD: Improving Feature-Based Knowledge Distillation for Object Detection Via Dual Masking Augmentation

ICASSP 2024accepted

Recent mainstream masked distillation methods function by reconstructing selectively masked areas of a student network from the feature map of its teacher counterpart. In these methods, the masked regions need to be properly selected, such that reconstructed features encode sufficient discrimination…

Cited by 0SourceScholar
2024

Diff-Reg: Diffusion Model in Doubly Stochastic Matrix Space for Registration Problem

ECCV 2024poster

"Establishing reliable correspondences is essential for 3D and 2D-3D registration tasks. Existing methods commonly leverage geometric or semantic point features to generate potential correspondences. However, these features may face challenges such as large deformation, scale inconsistency, and ambi…

2024

Driving-Video Dehazing with Non-Aligned Regularization for Safety Assistance

CVPR 2024poster

Real driving-video dehazing poses a significant challenge due to the inherent difficulty in acquiring precisely aligned hazy/clear video pairs for effective model training especially in dynamic driving scenarios with unpredictable weather conditions. In this paper we propose a pioneering approach th…

Cited by 10SourcePDFScholar
2024

DualAT: Dual Attention Transformer for End-to-End Autonomous Driving

ICRA 2024poster

The effective reasoning of integrated multimodal perception information is crucial for achieving enhanced end-to-end autonomous driving performance. In this paper, we introduce a novel multitask imitation learning framework for end-to-end autonomous driving that leverages a dual attention transforme…

Cited by 2SourceScholar
2024

Exploring Multi-Modal Control in Music-Driven Dance Generation

ICASSP 2024accepted

Existing music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating high-quality dance movements and supporting multi-modal cont…

Cited by 0SourceScholar
2024

Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation

COLING 2024main

Previous Sign Language Translation (SLT) methods achieve superior performance by relying on gloss annotations. However, labeling high-quality glosses is a labor-intensive task, which limits the further development of SLT. Although some approaches work towards gloss-free SLT through jointly training…

Cited by 14SourcePDFScholar
2024

MambaLLIE: Implicit Retinex-Aware Low Light Enhancement with Global-then-Local State Space

NeurIPS 2024poster

Recent advances in low light image enhancement have been dominated by Retinex-based learning framework, leveraging convolutional neural networks (CNNs) and Transformers. However, the vanilla Retinex theory primarily addresses global illumination degradation and neglects local issues such as noise an…

2024

Novel Object Synthesis via Adaptive Text-Image Harmony

NeurIPS 2024poster

In this paper, we study an object synthesis task that combines an object text with an object image to create a new object image. However, most diffusion models struggle with this task, \textit{i.e.}, often generating an object that predominantly reflects either the text or the image due to an imbala…

2024

Text2Avatar: Text to 3d Human Avatar Generation with Codebook-Driven Body Controllable Attribute

ICASSP 2024accepted

Generating 3D human models directly from text helps reduce the cost and time of character modeling. However, achieving multi-attribute controllable and realistic 3D human avatar generation is still challenging due to feature coupling and the scarcity of realistic 3D human avatar datasets. To address…

Cited by 0SourceScholar
2024

Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval

AAAI 2024technical

Video Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often sub-optimal since they ignore the modality imbalance problem, i.e., the semantic…

2024

Tri-Perspective View Decomposition for Geometry-Aware Depth Completion

CVPR 2024poster

Depth completion is a vital task for autonomous driving as it involves reconstructing the precise 3D geometry of a scene from sparse and noisy depth measurements. However most existing methods either rely only on 2D depth representations or directly incorporate raw 3D point clouds for compensation w…

Cited by 30SourcePDFScholar
2023

BEVHeight: A Robust Framework for Vision-Based Roadside 3D Object Detection

CVPR 2023poster

While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision…

2023

Creative Birds: Self-Supervised Single-View 3D Style Transfer

ICCV 2023poster

In this paper, we propose a novel method for single-view 3D style transfer that generates a unique 3D object with both shape and texture transfer. Our focus lies primarily on birds, a popular subject in 3D reconstruction, for which no existing single-view 3D transfer methods have been developed. T…

Cited by 7PDFcodeScholar
2023

Curriculum Temperature for Knowledge Distillation

AAAI 2023technical

Most existing distillation methods ignore the flexible role of the temperature in the loss function and fix it as a hyper-parameter that can be decided by an inefficient grid search. In general, the temperature controls the discrepancy between two distributions and can faithfully determine the diffi…

2023

DesNet: Decomposed Scale-Consistent Network for Unsupervised Depth Completion

AAAI 2023technical

Unsupervised depth completion aims to recover dense depth from the sparse one without using the ground-truth annotation. Although depth measurement obtained from LiDAR is usually sparse, it contains valid and real distance information, i.e., scale-consistent absolute depth values. Meanwhile, scale-a…

Cited by 32SourcePDFScholar
2023

Distortion and Uncertainty Aware Loss for Panoramic Depth Completion

ICML 2023poster

Standard MSE or MAE loss function is commonly used in limited field-of-vision depth completion, treating each pixel equally under a basic assumption that all pixels have same contribution during optimization. Recently, with the rapid rise of panoramic photography, panoramic depth completion (PDC) ha…

Cited by 17SourcePDFScholar
2023

Failure Detection for Motion Prediction of Autonomous Driving: An Uncertainty Perspective

ICRA 2023poster

Motion prediction is essential for safe and efficient autonomous driving. However, the inexplicability and uncertainty of complex artificial intelligence models may lead to unpredictable failures of the motion prediction module, which may mislead the system to make unsafe decisions. Therefore, it is…

Cited by 19SourceScholar
2023

Recurrent Structure Attention Guidance for Depth Super-resolution

AAAI 2023technical

Image guidance is an effective strategy for depth super-resolution. Generally, most existing methods employ hand-crafted operators to decompose the high-frequency (HF) and low-frequency (LF) ingredients from low-resolution depth maps and guide the HF ingredients by directly concatenating them with i…

2023

ScatterFormer: Locally-Invariant Scattering Transformer for Patient-Independent Multispectral Detection of Epileptiform Discharges

AAAI 2023technical

Patient-independent detection of epileptic activities based on visual spectral representation of continuous EEG (cEEG) has been widely used for diagnosing epilepsy. However, precise detection remains a considerable challenge due to subtle variabilities across subjects, channels and time points. Thus…

2023

Structure Flow-Guided Network for Real Depth Super-resolution

AAAI 2023technical

Real depth super-resolution (DSR), unlike synthetic settings, is a challenging task due to the structural distortion and the edge noise caused by the natural degradation in real-world low-resolution (LR) depth maps. These defeats result in significant structure inconsistency between the depth map an…

2022

A Tale of Two Flows: Cooperative Learning of Langevin Flow and Normalizing Flow Toward Energy-Based Model

ICLR 2022poster

This paper studies the cooperative learning of two generative flow models, in which the two models are iteratively updated based on the jointly synthesized examples. The first flow model is a normalizing flow that transforms an initial simple density to a target density by applying a sequence of inv…

Cited by 54SourcePDFScholar
2022

IPS300+: a Challenging multi-modal data sets for Intersection Perception System

ICRA 2022poster

Due to high complexity and occlusion, insufficient perception in the crowded urban intersection can be a serious safety risk for both human drivers and autonomous algorithms, whereas CVIS (Cooperative Vehicle Infrastructure System) is a proposed solution for full-participants perception under this s…

Cited by 2SourceScholar
2022

Industrial Style Transfer With Large-Scale Geometric Warping and Content Preservation

CVPR 2022poster

We propose a novel style transfer method to quickly create a new visual product with a nice appearance for industrial designers' reference. Given a source product, a target product, and an art style image, our method produces a neural warping field that warps the source shape to imitate the geometri…

Cited by 19PDFcodeScholar
2022

InterFusion: Interaction-based 4D Radar and LiDAR Fusion for 3D Object Detection

IROS 2022poster

Many recent works detect 3D objects by several sensor modalities for autonomous driving, where high-resolution cameras and high-line LiDARs are mostly used but relatively expensive. To achieve a balance between overall cost and detection accuracy, many multi-modal fusion techniques have been suggest…

Cited by 26SourceScholar
2022

Learning Contrastive Embedding in Low-Dimensional Space

NeurIPS 2022accept

Contrastive learning (CL) pretrains feature embeddings to scatter instances in the feature space so that the training data can be well discriminated. Most existing CL techniques usually encourage learning such feature embeddings in the highdimensional space to maximize the instance discrimination. H…

Cited by 15SourcePDFScholar
2022

Lightweight Projective Derivative Codes for Compressed Asynchronous Gradient Descent

ICML 2022spotlight

Coded distributed computation has become common practice for performing gradient descent on large datasets to mitigate stragglers and other faults. This paper proposes a novel algorithm that encodes the partial derivatives themselves and furthermore optimizes the codes by performing lossy compressio…

Cited by 5SourcePDFScholar
2022

Multi-modal Masked Pre-training for Monocular Panoramic Depth Completion

ECCV 2022poster

"In this paper, we formulate a potentially valuable panoramic depth completion (PDC) task as panoramic 3D cameras often produce 360° depth with missing data in complex scenes. Its goal is to recover dense panoramic depths from raw sparse ones and panoramic RGB images. To deal with the PDC task, we…

2022

Nested Collaborative Learning for Long-Tailed Visual Recognition

CVPR 2022poster

The networks trained on the long-tailed dataset vary remarkably, despite the same training settings, which shows the great uncertainty in long-tailed learning. To alleviate the uncertainty, we propose a Nested Collaborative Learning (NCL), which tackles the problem by collaboratively learning multip…

Cited by 121PDFcodeScholar
2022

RigNet: Repetitive Image Guided Network for Depth Completion

ECCV 2022poster

"Depth completion deals with the problem of recovering dense depth maps from sparse ones, where color images are often used to facilitate this task. Recent approaches mainly focus on image guided learning frameworks to predict dense depth. However, blurry guidance in the image and unclear structure…

Cited by 149SourcePDFScholar
2022

Spectral-Spatial Symmetrical Aggregation Cross-Linking Multi-Modal Data Fusion Network

ICASSP 2022accepted

In this paper, a spectral-spatial symmetrical aggregation cross-linking network (SACLNet) is developed for multi-modal data classification, which contains three modules as follows. First, the Spectro-Spatial Feature Learning Module is proposed, using the involution operation sliding over the spectra…

Cited by 0SourceScholar
2021

Cognitive Navigation for Indoor Environment Using Floorplan

IROS 2021poster

The recent years have seen the increasing importance of cognitive models for improved robot navigation. In this paper, a novel cognitive navigation package, which consists of topometric map representation and a three-level path planner, is proposed. The topometric maps are built from architectural f…

Cited by 7SourceScholar
2021

F-Net: Fusion Neural Network for Vehicle Trajectory Prediction in Autonomous Driving

ICASSP 2021accepted

Recent research has been remarkable in recurrent neural networks (RNNs) on sequence-to-sequence problems for image caption, and promising in convolutional neural networks (CNNs) on spatial analysis problems for image detection and sematic segmentation problems. In this paper, based on recurrent neur…

Cited by 0SourceScholar
2021

Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object Detection

CVPR 2021poster

Localization Quality Estimation (LQE) is crucial and popular in the recent advancement of dense object detectors since it can provide accurate ranking scores that benefit the Non-Maximum Suppression processing and improve detection performance. As a common practice, most existing methods predict LQE…

Cited by 324PDFcodeScholar
2021

Line-based Automatic Extrinsic Calibration of LiDAR and Camera

ICRA 2021poster

Reliable real-time extrinsic parameters of 3D Light Detection and Ranging (LiDAR) and camera are a key component of multi-modal perception systems. However, extrinsic transformation may drift gradually during operation, which can result in decreased accuracy of perception system. To solve this probl…

Cited by 54SourceScholar
2021

Regularizing Nighttime Weirdness: Efficient Self-Supervised Monocular Depth Estimation in the Dark

ICCV 2021poster

Monocular depth estimation aims at predicting depth from a single image or video. Recently, self-supervised methods draw much attention since they are free of depth annotations and achieve impressive performance on several daytime benchmarks. However, they produce weird outputs in more challenging n…

Cited by 91PDFcodeScholar
2021

Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud Registration

ICCV 2021poster

In this paper, by modeling the point cloud registration task as a Markov decision process, we propose an end-to-end deep model embedded with the cross-entropy method (CEM) for unsupervised 3D registration. Our model consists of a sampling network module and a differentiable CEM module. In our sampli…

Cited by 45PDFcodeScholar
2020

Drss-Based Localisation Using Weighted Instrumental Variables and Selective Power Measurement

ICASSP 2020accepted

Differential received signal strength (DRSS) provides a practical means of localisation for wireless sensor networks. Closed-form location estimators based on a linearised propagation path loss model are computationally efficient and hence suitable for wireless sensor nodes. However, most existing s…

Cited by 0SourceScholar
2020

Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection

NeurIPS 2020poster

One-stage detector basically formulates object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for one-stage detectors is…

2020

Learning Canonical Shape Space for Category-Level 6D Object Pose and Size Estimation

CVPR 2020poster

We present a novel approach to category-level 6D object pose and size estimation. To tackle intra-class shape variations, we learn canonical shape space (CASS), a unified representation for a large variety of instances of a certain object category. In particular, CASS is modeled as the latent space…

Cited by 218PDFScholar
2019

Cross-X Learning for Fine-Grained Visual Categorization

ICCV 2019poster

Recognizing objects from subcategories with very subtle differences remains a challenging task due to the large intra-class and small inter-class variation. Recent work tackles this problem in a weakly-supervised manner: object parts are first detected and the corresponding part-specific features ar…

Cited by 241PDFcodeScholar
2019

Dual Entangled Polynomial Code: Three-Dimensional Coding for Distributed Matrix Multiplication

ICML 2019oral

Matrix multiplication is a fundamental building block in various machine learning algorithms. When the matrix comes from a large dataset, the multiplication can be split into multiple tasks which calculate the multiplication of submatrices on different nodes. As some nodes may be stragglers, coding…

Cited by 19SourcePDFScholar
2019

Joint Design for MIMO Radar and Downlink Communication Systems Coexistence

ICASSP 2019accepted

This paper focuses on the problem of coexistence between a colocated multiple-in multiple-out (MIMO) radar and down-link communication systems. We propose an iterative algorithm to minimize the Cramer-Rao Bound (CRB) on the direction of arrival (DOA) of a target, accounting for the energy and simila…

Cited by 0SourceScholar
2019

Seeking the Analytical Approximation of the Stance Dynamics of the 3D Spring-Loaded Inverted Pendulum Model By Using Perturbation Approach

IROS 2019poster

The Spring-Loaded Inverted Pendulum (SLIP) has been widely exploited in both biomechanical and robotics research due to its simple form in mathematics and high accuracy in fitting experimental biology data. However the intrinsic nonlinearity of the SLIP dynamics makes accurate analytical representat…

Cited by 2SourceScholar
2018

Spectrally Compatible Waveform Design for MIMO Radar Transmit Beampattern with Par and Similarity Constraints

ICASSP 2018accepted

This paper investigates the problem of the spectrally compatible waveform design for multiple-in multiple-out (MIMO) radar transmit beampattern formation, subject to peak-to-average-power ratio (PAR) and waveform similarity constraints. Since the formulated optimization problem of minimizing the Int…

Cited by 0SourceScholar
2018

Towards Emergence of Tool Use in Robots: Automatic Tool Recognition and Use Without Prior Tool Learning

ICRA 2018poster

Humans are adept at tool use. We can intuitively and immediately improvise and use unknown objects in our environment as tools, to assist us in performing tasks. In this study, we provide similar cognition and capabilities to robots. Neuroscientific studies on tool use have suggested that human dext…

Cited by 23SourceScholar
2016

ReD-SFA: Relation Discovery Based Slow Feature Analysis for Trajectory Clustering

CVPR 2016poster

For spectral embedding/clustering, it is still an open problem on how to construct an relation graph to reflect the intrinsic structures in data. In this paper, we proposed an approach, named Relation Discovery based Slow Feature Analysis (ReD-SFA), for feature learning and graph construction simult…

Cited by 15PDFScholar