← Search

Jie Li

128 accepted papers

2026

AutoDebias: An Automated Framework for Detecting and Mitigating Backdoor Biases in Text-to-Image Models

CVPR 2026

Text-to-Image (T2I) models generate high-quality images but are vulnerable to malicious backdoor attacks that inject harmful biases (e.g., trigger-activated gender or racial stereotypes). Existing debiasing methods, often designed for natural statistical biases, struggle with these deliberate and su

Cited by 0SourcecodeScholar
2026

CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World

AAAI 2026technical

How far are deep models from real-world video anomaly understanding (VAU)? Current works typically emphasize detecting unexpected occurrences deviating from normal patterns or comprehending anomalous events with interpretable descriptions. However, they exhibit only a superficial comprehension of re

Cited by 0SourcePDFScholar
2026

DIFFA: Large Language Diffusion Models Can Listen and Understand

AAAI 2026technical

Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, large language diffusion models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context mode

Cited by 0SourcePDFScholar
2026

Data Agent: Learning to Select Data via End-to-End Dynamic Optimization

ICML 2026poster

Dynamic Data selection aims to accelerate training by prioritizing informative samples during online training. However, existing methods typically rely on task-specific handcrafted metrics or static/snapshot-based criteria to estimate sample importance, limiting scalability across learning paradigms…

Cited by 0SourceScholar
2026

DiasR: Dual-Modal Identity-Anchored Sparse Routing for Efficient Multi-Subject Video Generation

ICML 2026poster

Personalized multi-subject video generation is a promising direction within the field of controllable video generation; however, existing methods face challenges in maintaining cross-frame identity consistency and incur high computational overhead. To address these issues, we propose DiasR, an effic…

Cited by 0SourceScholar
2026

E-mem: Multi-Agent Based Episodic Context Reconstruction for LLM Agent Memory

ICML 2026poster

The evolution of Large Language Model (LLM) agents towards System~2 reasoning, characterized by deliberative, high-precision problem-solving, necessitates maintaining rigorous logical integrity over extended horizons. However, prevalent memory preprocessing paradigms incur destructive de-contextuali…

Cited by 0SourceScholar
2026

EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models

ICML 2026poster

Electroencephalography foundation models (EEG-FMs) have advanced brain signal analysis, but the lack of standardized evaluation benchmarks impedes model comparison and scientific progress. Current evaluations rely on inconsistent protocols that render cross-model comparisons unreliable, while a lack…

Cited by 0SourceScholar
2026

FIND: A Simple Yet Effective Baseline for Diffusion-Generated Image Detection

AAAI 2026technical

The remarkable realism of images generated by diffusion models poses critical detection challenges. Current methods utilize reconstruction error as a discriminative feature, exploiting the observation that real images exhibit higher reconstruction errors when processed through diffusion models. Howe

Cited by 0SourcePDFScholar
2026

FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training

AAAI 2026technical

Recent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, their massive parameter sizes lead to slow inference, high memory usage, and poor deployability. Existing acceleration meth

Cited by 0SourcePDFScholar
2026

Force Estimation and Position Control of a Hydraulic Folded Pouch Actuator for Soft Robotics

ICRA 2026poster

This paper investigates position control and force estimation for a hydraulic folded pouch actuator. First, experimental platforms are designed to characterize the actuator and the results show two key properties: (i) angular hysteresis when the motion direction reverses, and (ii) strong nonlinearit…

Cited by 0Scholar
2026

GhostEI-Bench: Do Mobile Agent Resilience to Environmental Injection in Dynamic On-Device Environments?

ICLR 2026poster

Vision-Language Models (VLMs) are increasingly deployed as autonomous agents to navigate mobile Graphical User Interfaces (GUIs). However, their operation within dynamic on-device ecosystems, which include notifications, pop-ups, and inter-app interactions, exposes them to a unique and underexplored…

Cited by 0SourceScholar
2026

LogART: Pushing the Limit of Efficient Logarithmic Post-Training Quantization

ICLR 2026poster

Efficient deployment of deep neural networks increasingly relies on Post-Training Quantization (PTQ). Logarithmic PTQ, in particular, promises multiplier-free hardware efficiency, but its performance is often limited by the nonlinear and symmetric quantization grid and standard rounding-to-nearest (…

Cited by 0SourcecodeScholar
2026

Revealing the Invisible: Latent Structure Modeling for Semantically Consistent Cloud Removal

AAAI 2026technical

Cloud removal (CR) in remote sensing imagery is a critical yet challenging task due to complex cloud patterns and diverse underlying ground structures. Despite recent progress in generative models such as diffusion models, CR remains limited by their inadequate capability to perceive and reconstruct

Cited by 0SourcePDFScholar
2026

SIGMA: An Agent-Based Modeling UAV Swarm Simulator for Swarm Intelligence Algorithms (I)

ICRA 2026poster

Swarm intelligence for uncrewed aerial vehicles (UAVs) significantly improves the success rate of executing intricate tasks using “distributed platforms and aggregated effects”. However, the high experimental costs and safety risks constrain its development. This paper introduces SIGMA (Swarm Intell…

Cited by 0Scholar
2026

SPR$^2$Q: Static Priority-based Rectifier Routing Quantization for Image Super-Resolution

ICLR 2026poster

Low-bit quantization has achieved significant progress in image super-resolution. However, existing quantization methods show evident limitations in handling the heterogeneity of different components. Particularly under extreme low-bit compression, the issue of information loss becomes especially pr…

Cited by 0SourceScholar
2026

Shrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision Alignment

AAAI 2026technical

Decoding visual features from EEG signals is a central challenge in neuroscience, with cross-modal alignment as the dominant approach. We argue that the relationship between visual and brain modalities is fundamentally asymmetric, characterized by two critical gaps: a Fidelity Gap (stemming from EEG

Cited by 0SourcePDFScholar
2026

StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental Learning

ICLR 2026poster

Video Class-Incremental Learning (VCIL) seeks to develop models that continuously learn new action categories over time without forgetting previously acquired knowledge. Unlike traditional Class-Incremental Learning (CIL), VCIL introduces the added complexity of spatiotemporal structures, making it…

Cited by 0SourceScholar
2026

TokenPowerBench: Benchmarking the Power Consumption of LLM Inference

AAAI 2026technical

Large language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little

Cited by 0SourcePDFScholar
2026

WENETSPEECH-CHUAN: A LARGE-SCALE SICHUANESE CORPUS WITH RICH ANNOTATION FOR DIALECTAL SPEECH PROCESSING

ICASSP 2026poster

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce WenetSpeech-Chuan, a 10,000-hour, richly annotated corpus constru…

Cited by 0SourcePDFScholar
2026

WenetSpeech-Yue: A Large-Scale Cantonese Speech Corpus with Multi-dimensional Annotation

AAAI 2026technical

The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximate

Cited by 0SourcePDFScholar
2025

A Blind Super-Resolution Method for Near-Field Channel Estimation with Angle-Range Recovery

ICASSP 2025accepted

As massive MIMO and extremely large antenna array rapidly evolve, wireless communication is gradually shifting from the far-field to the near-field region. To accurately capture the characteristics of the near-field channel, the spherical-wave assumption has to be applied, making channel estimation…

Cited by 0SourceScholar
2025

Asynchronous Harmony-based Decentralized Auctions Method for Scalable UAV Swarm

IROS 2025

Unmanned aerial vehicle (UAV) swarms find extensive applications in diverse fields, including search and rescue, logistics delivery, and environmental surveillance, necessitating meticulous task and temporal scheduling to meet intricate spatiotemporal requirements. A market-based strategy emerges as

Cited by 0SourceScholar
2025

Automated UAV-based Wind Turbine Blade Inspection: Blade Stop Angle Estimation and Blade Detail Prioritized Exposure Adjustment

IROS 2025

Unmanned aerial vehicles (UAVs) are critical in the automated inspection of wind turbine blades. Nevertheless, several issues persist in this domain. Firstly, existing inspection platforms encounter challenges in meeting the demands of automated inspection tasks and scenarios. Moreover, current blad

Cited by 0SourceScholar
2025

BoRe-Depth: Self-Supervised Monocular Depth Estimation with Boundary Refinement for Embedded Systems

IROS 2025

Depth estimation is one of the key technologies for realizing 3D perception in unmanned systems. Monocular depth estimation has been widely researched because of its low-cost advantage, but the existing methods face the challenges of poor depth estimation performance and blurred object boundaries on

Cited by 1SourcecodeScholar
2025

Bridging the Reality Gap: Communication-Aware Task Allocation with Multi-Objective Asynchronous Policy Learning

IROS 2025

Distributed task allocation in the UAV swarm is sensitive to excessive communication overhead and frequent transmissions. Combining reinforcement learning and task allocation demonstrates great potential in enhancing algorithm performance and optimizing communication. However, existing studies rely

Cited by 0SourceScholar
2025

Contrasting Adversarial Perturbations: The Space of Harmless Perturbations

AAAI 2025technical

Existing works have extensively studied adversarial examples, which are minimal perturbations that can mislead the output of deep neural networks (DNNs) while remaining imperceptible to humans. However, in this work, we reveal the existence of a harmless perturbation space, in which perturbations dr…

2025

DS-VLM: Diffusion Supervision Vision Language Model

ICML 2025poster

Vision-Language Models (VLMs) face two critical limitations in visual representation learning: degraded supervision due to information loss during gradient propagation, and the inherent semantic sparsity of textual supervision compared to visual data. We propose the Diffusion Supervision Vision-Lang…

Cited by 0SourcePDFScholar
2025

Double-Filter: Efficient Fine-tuning of Pre-trained Vision-Language Models via Patch&Layer Filtering

ICML 2025poster

In this paper, we present a novel approach, termed Double-Filter,to “slim down” the fine-tuning process of vision-language pre-trained (VLP) models via filtering redundancies in feature inputs and architectural components. We enhance the fine-tuning process using two approaches. First, we develop a…

Cited by 0SourcePDFScholar
2025

Enhancing Transferability of Targeted Adversarial Examples via Inverse Target Gradient Competition and Spatial Distance Stretching

ICCV 2025poster

In the field of AI security, deep neural networks (DNNs) are highly sensitive to adversarial examples (AEs), which can cause incorrect predictions with minimal input perturbations. Although AEs exhibit transferability across models, targeted attack success rates (TASRs) are low due to differences in…

Cited by 0SourcePDFScholar
2025

GD$^2$: Robust Graph Learning under Label Noise via Dual-View Prediction Discrepancy

NeurIPS 2025poster

Graph Neural Networks (GNNs) achieve strong performance in node classification tasks but exhibit substantial performance degradation under label noise. Despite recent advances in noise-robust learning, a principled approach that exploits the node-neighbor interdependencies inherent in graph data for…

Cited by 0SourceScholar
2025

GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

NeurIPS 2025poster

Aerial-Ground person re-identification (AG-ReID) is an emerging yet challenging task that aims to match pedestrian images captured from drastically different viewpoints, typically from unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task poses significant challenges due to…

Cited by 0SourceScholar
2025

Generative Diffusion Model-based Energy Management in Networked Energy Systems

ICASSP 2025accepted

In recent years, the proliferation of renewable energy sources has heightened the focus on networked energy systems. These systems face significant challenges due to the unpredictable nature of energy generation and consumption, as well as the complexity of managing numerous components and parameter…

Cited by 0SourceScholar
2025

HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation

CVPR 2025poster

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale videos with accurate captions for HOI. To address this issue,…

2025

Hierarchical Multimodal Decoupling-Fusion Framework for offline Multiple Appropriate Facial Reaction Generation

ICASSP 2025accepted

Facial reactions convey crucial emotional information and coordinating interpersonal relationships in human dyadic interactions. While existing Multiple Appropriate Facial Reaction Generation (MAFRG) methods focus on generating multiple reasonable facial reactions, none of these approaches combines…

Cited by 0SourceScholar
2025

Infrared and Visible Image Fusion with Hierarchical Human Perception

ICASSP 2025accepted

Image fusion combines images from multiple domains into one image, containing complementary information from source domains. Existing methods take pixel intensity, texture and high-level vision task information as the standards to determine preservation of information, lacking enhancement for human…

Cited by 0SourceScholar
2025

JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models

NeurIPS 2025poster

Vision-Language Models (VLMs) exhibit impressive performance, yet the integration of powerful vision encoders has significantly broadened their attack surface, rendering them increasingly susceptible to jailbreak attacks. However, lacking well-defined attack objectives, existing jailbreak methods of…

Cited by 0SourceScholar
2025

Leveraging Peer-Informed Label Consistency for Robust Graph Neural Networks with Noisy Labels

IJCAI 2025

Graph Neural Networks (GNNs) excel in many applications but struggle when trained with noisy labels, especially as noise can propagate through the graph structure. Despite recent progress in developing robust GNNs, few methods exploit the intrinsic properties of graph data to filter out noise. In th

Cited by 0SourcePDFScholar
2025

Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuning

EMNLP 2025

The effectiveness of instruction fine-tuning for Large Language Models is fundamentally constrained by the quality and efficiency of training datasets. This work introduces Low-Confidence Gold (LCG), a novel filtering framework that employs centroid-based clustering and confidence-guided selection f

2025

MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification

NeurIPS 2025spotlight

The challenge of inconsistent modalities in real-world applications presents significant obstacles to effective object re-identification (ReID). However, most existing approaches assume modality-matched conditions, significantly limiting their effectiveness in modality-mismatched scenarios. To overc…

Cited by 0SourceScholar
2025

Multi-Modal Object Re-identification via Sparse Mixture-of-Experts

ICML 2025poster

We present MFRNet, a novel network for multi-modal object re-identification that integrates multi-modal data features to effectively retrieve specific objects across different modalities. Current methods suffer from two principal limitations: (1) insufficient interaction between pixel-level semantic…

Cited by 0SourcePDFScholar
2025

Multi-modal Anchor Gated Transformer with Knowledge Distillation for Emotion Recognition in Conversation

IJCAI 2025

Emotion Recognition in Conversation (ERC) aims to detect the emotions of individual utterances within a conversation. Generating efficient and modality-specific representations for each utterance remains a significant challenge. Previous studies have proposed various models to integrate features ext

2025

SML: A Backdoor Defense for Non-Intrusive Speech Quality Assessment via Semi-Supervised and Multi-Task Learning

ICASSP 2025accepted

Non-intrusive speech quality assessment (NISQA) is widely used in speech downstream tasks due to its ability to predict the quality of speech without a reference speech. However, few researchers have focused on the backdoor security of NISQA. Despite the backdoor defenses have been extensively studi…

Cited by 0SourceScholar
2025

Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations

NeurIPS 2025poster

Generative models have recently gained attention in recommendation systems by directly predicting item identifiers from user interaction sequences. However, existing methods suffer from significant information loss due to the separation of stages such as quantization and sequence modeling, hindering…

Cited by 0SourceScholar
2025

Textual and Visual Prompt Fusion for Image Editing via Step-Wise Alignment

ICASSP 2025accepted

The use of denoising diffusion models is becoming increasingly popular in the field of image editing. However, current approaches often rely on either image-guided methods, which provide a visual reference but lack control over semantic consistency, or text-guided methods, which ensure alignment wit…

Cited by 0SourceScholar
2025

UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic Programming

CVPR 2025award

Distributed learning is commonly used for training deep learning models, especially large models. In distributed learning, manual parallelism (MP) methods demand considerable human effort and have limited flexibility. Hence, automatic parallelism (AP) methods have recently been proposed for automati…

2025

Variational Perturbation Personalized Federated Learning via Prior-Posterior Distance

ICASSP 2025accepted

Personalized Federated Learning (pFL) mitigates the impact of statistical heterogeneity on FL architecture to some extent by allowing participants to use personalized models based on local data distributions. The existing pFL methods optimize from the perspective of model structure, attempting to ad…

Cited by 0SourceScholar
2025

WonderTurbo: Generating Interactive 3D World in 0.72 Seconds

ICCV 2025poster

Interactive 3D generation is gaining momentum and capturing extensive attention for its potential to create immersive virtual experiences. However, a critical challenge in current 3D generation technologies lies in achieving real-time interactivity. To address this issue, we introduce WonderTurbo, t…

Cited by 0SourcePDFScholar
2024

A Riemannian-Based Joint Design Framework of Mimo Radar Transmit Waveform And Receive Filter Via Information Theory

ICASSP 2024accepted

In this paper, we explore the joint design of a transmit waveform and receive filter to enhance the detection performance of multiple-input multiple-output (MIMO) radar. Target echoes are assumed to be embedded in signal-dependent interference and colored Gaussian noise. As design metrics, we exploi…

Cited by 0SourceScholar
2024

Augmenting Lane Perception and Topology Understanding with Standard Definition Navigation Maps

ICRA 2024poster

Autonomous driving has traditionally relied heavily on costly and labor-intensive High Definition (HD) maps, hindering scalability. In contrast, Standard Definition (SD) maps are more affordable and have worldwide coverage, offering a scalable alternative. In this work, we systematically explore the…

Cited by 34SourcecodeScholar
2024

DiffFAS: Face Anti-Spoofing via Generative Diffusion Models

ECCV 2024poster

"Face anti-spoofing (FAS) plays a vital role in preventing face recognition (FR) systems from presentation attacks. Nowadays, FAS systems face the challenge of domain shift, impacting the generalization performance of existing FAS methods. In this paper, we rethink about the inherence of domain shif…

2024

InterpGNN: Understand and Improve Generalization Ability of Transdutive GNNs through the Lens of Interplay between Train and Test Nodes

ICLR 2024poster

Transductive node prediction has been a popular learning setting in Graph Neural Networks (GNNs). It has been widely observed that the shortage of information flow between the distant nodes and intra-batch nodes (for large-scale graphs) often hurt the generalization of GNNs which overwhelmingly adop…

Cited by 1SourcePDFScholar
2024

Multi-Granularity Graph-Convolution-Based Method for Weakly Supervised Person Search

IJCAI 2024poster

One-step Weakly Supervised Person Search (WSPS) jointly performs pedestrian detection and person Re-IDentification (ReID) only with bounding box annotations, which makes the traditional person ReID problem more suitable and efficient for real-world applications. However, this task is very challengin…

Cited by 0SourcePDFScholar
2024

Multi-Scene Generalized Trajectory Global Graph Solver with Composite Nodes for Multiple Object Tracking

AAAI 2024technical

The global multi-object tracking (MOT) system can consider interaction, occlusion, and other ``visual blur'' scenarios to ensure effective object tracking in long videos. Among them, graph-based tracking-by-detection paradigms achieve surprising performance. However, their fully-connected nature pos…

Cited by 4SourcePDFScholar
2024

Scheduling of Time-Constrained Single-Arm Cluster Tools with Two-Space Process Modules in Semiconductor Manufacturing

RA-L 2024

In semiconductor manufacturing, to improve the throughput of cluster tools, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">process modules</i> (PMs) are designed to have multiple spaces such that more than one wafers can be processed in a PM concurr

Cited by 4SourceScholar
2024

ShaSTA: Modeling Shape and Spatio-Temporal Affinities for 3D Multi-Object Tracking

RA-L 2024

Multi-object tracking (MOT) is a cornerstone capability of any robotic system. Tracking quality is largely dependent on the quality of input detections. In many applications, such as autonomous driving, it is preferable to over-detect objects to avoid catastrophic outcomes due to missed detections.

Cited by 44SourcecodeScholar
2024

TICOP: Time-Critical Coordinated Planning for Fixed-Wing UAVs in Unknown Unstructured Environments

RA-L 2024

Safe coordination of fixed-wing UAVs in unstructured environments poses challenges due to the intricate coupling of UAV cooperation, obstacle avoidance, and motion constraints. One task is time-critical coordination, which means that all UAVs can safely reach their destinations simultaneously. Exist

Cited by 3SourceScholar
2024

Towards Cross-View-Consistent Self-Supervised Surround Depth Estimation

IROS 2024poster

Depth estimation is a cornerstone for autonomous driving, yet acquiring per-pixel depth ground truth for supervised learning is challenging. Self-Supervised Surround Depth Estimation (SSSDE) from consecutive images offers an economical alternative. While previous SSSDE methods have proposed differen…

Cited by 0SourcecodeScholar
2024

Trades++: Enhancing Multi-Object Tracking of Real Low Confidence Targets Using a Pyramid-Like Self-Attention Model

ICASSP 2024accepted

In reality, multi-object tracking (MOT) is used in a wide range of scenarios. Maintaining the motion trajectory of the target, especially in high-density pedestrian scenarios, is often difficult. The tracking quality of most multi-object trackers correlates strongly with the detector quality and the…

Cited by 0SourceScholar
2023

A Distributed Scheduling Method for Networked UAV Swarm based on Computing for Communication

IROS 2023poster

UAV swarms have attracted much attention for post-disaster search and rescue, pollution monitoring and trace-ability, etc., where distributed scheduling is required to arrange careful tasks and time quickly. The market-based methods are widely favored but they rely on the environmentally influenced…

Cited by 1SourceScholar
2023

Depth Is All You Need for Monocular 3D Detection

ICRA 2023poster

A key contributor to recent progress in 3D detection from single images is monocular depth estimation. Existing methods focus on how to leverage depth explicitly, by generating pseudo-pointclouds or providing attention cues for image features. More recent works leverage depth prediction as a pretrai…

Cited by 11SourcecodeScholar
2023

Learning to Reconnect Interrupted Trajectories for Weakly Supervised Multi-Object Tracking

ICASSP 2023accepted

Recently, some weakly supervised multi-object tracking (MOT) methods learn identity embedding features with pseudo identity labels rather than the high-cost manual ones. However, these pseudo identity labels may contain many false or missing identities, which adversely affect the optimization of tra…

Cited by 0SourceScholar
2023

MRCN: A Novel Modality Restitution and Compensation Network for Visible-Infrared Person Re-identification

AAAI 2023technical

Visible-infrared person re-identification (VI-ReID), which aims to search identities across different spectra, is a challenging task due to large cross-modality discrepancy between visible and infrared images. The key to reduce the discrepancy is to filter out identity-irrelevant interference and ef…

Cited by 43SourcePDFScholar
2023

Simple-BEV: What Really Matters for Multi-Sensor BEV Perception?

ICRA 2023poster

Building 3D perception systems for autonomous vehicles that do not rely on high-density LiDAR is a critical research problem because of the expense of LiDAR systems compared to cameras and other sensors. Recent research has developed a variety of camera-only methods, where features are differentiabl…

Cited by 143SourceScholar
2023

Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking

CVPR 2023poster

This work proposes an end-to-end multi-camera 3D multi-object tracking (MOT) framework. It emphasizes spatio-temporal continuity and integrates both past and future reasoning for tracked objects. Thus, we name it "Past-and-Future reasoning for Tracking" (PF-Track). Specifically, our method adapts th…

2023

Towards Practical Edge Inference Attacks Against Graph Neural Networks

ICASSP 2023accepted

Graph Neural Networks (GNNs) have demonstrated superior performance in numerous real-world applications. Despite their success, recent studies have shown that GNNs are vulnerable under edge inference attacks aimed to infer the connectivity of a given pair of nodes. However, existing methods primaril…

Cited by 0SourceScholar
2023

Tracking Through Containers and Occluders in the Wild

CVPR 2023poster

Tracking objects with persistence in cluttered and dynamic environments remains a difficult challenge for computer vision systems. In this paper, we introduce TCOW, a new benchmark and model for visual tracking through heavy occlusion and containment. We set up a task where the goal is to, given a v…

2023

Viewpoint Equivariance for Multi-View 3D Object Detection

CVPR 2023poster

3D object detection from visual sensors is a cornerstone capability of robotic systems. State-of-the-art methods focus on reasoning and decoding object bounding boxes from multi-view camera input. In this work we gain intuition from the integral role of multi-view consistency in 3D scene understandi…

2022

Ada-STNet: A Dynamic AdaBoost Spatio-Temporal Network for Traffic Flow Prediction

ICASSP 2022accepted

Traffic flow prediction is of particular interest since its massive applications in intelligent transportation systems (ITS). The problem is challenging due to the complex spatio-temporal correlations and nonlinearities of traffic flows. However, existing methods based on the graph neural networks c…

Cited by 0SourceScholar
2022

An Improved In-flight Alignment Method Based on Backtracking Navigation for GPS-aided Low Cost SINS With Short Endurance

RA-L 2022

Aiming at the problem that the low-cost Strapdown Inertial Navigation System (SINS)/Global Positioning System (GPS) cannot effectively achieve fast in-flight alignment with short-endurance, an improved alignment method based on backtracking navigation is proposed. Firstly, the existing backtracking

Cited by 10SourceScholar
2022

Attention-Based Deep Driving Model for Autonomous Vehicles with Surround-View Cameras

IROS 2022poster

Experienced human drivers always make safe driving decisions by selectively observing the front, rear and side- view mirrors. Several end - to-end methods have been pro-posed to learn driving models with multi-view visual infor-mation. However, these benchmark methods lack semantic understanding of…

Cited by 0SourceScholar
2022

Black-Box Dissector: Towards Erasing-Based Hard-Label Model Stealing Attack

ECCV 2022poster

"Previous studies have verified that the functionality of black-box models can be stolen with full probability outputs. However, under the more practical hard-label setting, we observe that existing methods suffer from catastrophic performance degradation. We argue this is due to the lack of rich in…

2022

Domain Adaptation via Mutual Information Maximization for Handwriting Recognition

ICASSP 2022accepted

Deep learning models for handwriting recognition have been developed in recent years. To improve the model’s generalization ability for sequence modeling task, this paper proposes to use domain adaptation with statistical distribution alignment and entropy regularization. For statistical distributio…

Cited by 0SourceScholar
2022

Fully Attentional Network for Semantic Segmentation

AAAI 2022technical

Recent non-local self-attention methods have proven to be effective in capturing long-range dependencies for semantic segmentation. These methods usually form a similarity map of R^(CxC) (by compressing spatial dimensions) or R^(HWxHW) (by compressing channels) to describe the feature relations alon…

2022

Input-Specific Robustness Certification for Randomized Smoothing

AAAI 2022technical

Although randomized smoothing has demonstrated high certified robustness and superior scalability to other certified defenses, the high computational overhead of the robustness certification bottlenecks the practical applicability, as it depends heavily on the large sample approximation for estimati…

2022

On Collective Robustness of Bagging Against Data Poisoning

ICML 2022spotlight

Bootstrap aggregating (bagging) is an effective ensemble protocol, which is believed can enhance robustness by its majority voting mechanism. Recent works further prove the sample-wise robustness certificates for certain forms of bagging (e.g. partition aggregation). Beyond these particular forms, i…

2022

One-Bit Active Query With Contrastive Pairs

CVPR 2022poster

How to achieve better results with fewer labeling costs remains a challenging task. In this paper, we present a new active learning framework, which for the first time incorporates contrastive learning into recently proposed one-bit supervision. Here one-bit supervision denotes a simple Yes or No qu…

Cited by 9PDFcodeScholar
2022

SpOT: Spatiotemporal Modeling for 3D Object Tracking

ECCV 2022poster

"3D multi-object tracking aims to uniquely and consistently identify all mobile entities through time. Despite the rich spatiotemporal information available in this setting, current 3D tracking methods primarily rely on abstracted information and limited history, e.g. single-frame object bounding bo…

Cited by 13SourcePDFScholar
2021

A Circular-Structured Representation for Visual Emotion Distribution Learning

CVPR 2021poster

Visual Emotion Analysis (VEA) has attracted increasing attention recently with the prevalence of sharing images on social networks. Since human emotions are ambiguous and subjective, it is more reasonable to address VEA in a label distribution learning (LDL) paradigm rather than a single-label class…

Cited by 39PDFScholar
2021

Aha! Adaptive History-Driven Attack for Decision-Based Black-Box Models

ICCV 2021poster

The decision-based black-box attack means to craft adversarial examples with only the top-1 label of the victim model available. A common practice is to start from a large perturbation and then iteratively reduce it with a deterministic direction and a random one while keeping it adversarial. The li…

Cited by 21PDFScholar
2021

Detection Of Malicious DNS and Web Servers using Graph-Based Approaches

ICASSP 2021accepted

The DNS hijacking attack represents a significant threat to users. In this type of attack, a malicious DNS server redirects a victim domain to an attacker-controlled web server. Existing defenses are not scalable and have not been widely deployed. In this work, we propose both unsupervised and semi-…

Cited by 0SourceScholar
2021

Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style Transfer

CVPR 2021poster

Artistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize comple…

Cited by 119PDFcodeScholar
2021

Geometric Unsupervised Domain Adaptation for Semantic Segmentation

ICCV 2021poster

Simulators can efficiently generate large amounts of labeled synthetic data with perfect supervision for hard-to-label tasks like semantic segmentation. However, they introduce a domain gap that severely hurts real-world performance. We propose to use self-supervised monocular depth estimation as a…

Cited by 48PDFcodeScholar
2021

Hierarchical Lovasz Embeddings for Proposal-Free Panoptic Segmentation

CVPR 2021poster

Panoptic segmentation brings together two separate tasks: instance and semantic segmentation. Although they are related, unifying them faces an apparent paradox: how to learn simultaneously instance-specific and category-specific (i.e. instance-agnostic) representations jointly. Hence, state-of-the-…

Cited by 10PDFScholar
2021

IMENet: Joint 3D Semantic Scene Completion and 2D Semantic Segmentation through Iterative Mutual Enhancement

IJCAI 2021poster

3D semantic scene completion and 2D semantic segmentation are two tightly correlated tasks that are both essential for indoor scene understanding, because they predict the same semantic classes, using positively correlated high-level features. Current methods use 2D features extracted from early-fus…

Cited by 16SourcePDFScholar
2021

Inverse Simulation: Reconstructing Dynamic Geometry of Clothed Humans via Optimal Control

CVPR 2021poster

This paper studies the problem of inverse cloth simulation---to estimate shape and time-varying poses of the underlying body that generates physically plausible cloth motion, which matches to the point cloud measurements on the clothed humans. A key innovation is to represent the dynamics of the clo…

Cited by 16PDFScholar
2021

Is Pseudo-Lidar Needed for Monocular 3D Object Detection?

ICCV 2021poster

Recent progress in 3D object detection from single images leverages monocular depth estimation as a way to produce 3D pointclouds, turning cameras into pseudo-lidar sensors. These two-stage detectors improve with the accuracy of the intermediate depth estimation network, which can itself be improved…

Cited by 388PDFcodeScholar
2021

LocTex: Learning Data-Efficient Visual Representations From Localized Textual Supervision

ICCV 2021poster

Computer vision tasks such as object detection and semantic/instance segmentation rely on the painstaking annotation of large training datasets. In this paper, we propose LocTex that takes advantage of the low-cost localized textual annotations (i.e., captions and synchronized mouse-over gestures) t…

Cited by 14PDFScholar
2021

One-Shot Voice Conversion Based on Speaker Aware Module

ICASSP 2021accepted

Voice conversion (VC) is a task to convert the voice of speech while preserving its linguistic content. Although several methods have been proposed to enable VC with non-parallel data, it is still difficult to model the voice without a great number of data or an adaptive process. In this paper, we p…

Cited by 0SourceScholar
2021

Probabilistic 3D Multi-Modal, Multi-Object Tracking for Autonomous Driving

ICRA 2021poster

Multi-object tracking is an important ability for an autonomous vehicle to safely navigate a traffic scene. Current state-of-the-art follows the tracking-by-detection paradigm where existing tracks are associated with detected objects through some distance metric. Key challenges to increase tracking…

Cited by 306SourceScholar
2021

Riemannian Geometric Optimization Methods for Joint Design of Transmit Sequence and Receive Filter of MIMO Radar

ICASSP 2021accepted

To maximize the signal-to-interference-plus-noise ratio (SINR) under a constant-envelope constraint, an efficient joint design of the transmit waveform and the receive filter for multipleinput multiple-output (MIMO) radars is essential. In this paper, we propose a novel optimization framework to sol…

Cited by 0SourceScholar
2021

Simultaneous Control of Terrain Adaptation and Wheel Speed Allocation for a Planetary Rover With an Active Suspension System

RA-L 2021

Active suspensions are important features of many recent planetary rovers. For such a rover, the control strategy is crucial to its performance. This letter presents the control method for a planetary rover equipped with an active suspension system. The control algorithm is based on the estimation o

Cited by 9SourceScholar
2021

Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene Completion

AAAI 2021technical

LiDAR point cloud analysis is a core task for 3D computer vision, especially for autonomous driving. However, due to the severe sparsity and noise interference in the single sweep LiDAR point cloud, the accurate semantic segmentation is non-trivial to achieve. In this paper, we propose a novel spars…

2020

Binarized Neural Network for Single Image Super Resolution

ECCV 2020poster

Lighter model and faster inference are the focus of current single image super-resolution (SISR) research. However, existing methods are still hard to be applied in real-world applications due to the requirement of its heavy computation. Model quantization is an effective way to significantly reduce…

Cited by 90SourcePDFScholar
2020

Depth Based Semantic Scene Completion With Position Importance Aware Loss

RA-L 2020

Semantic scene completion (SSC) refers to the task of inferring the 3D semantic segmentation of a scene while simultaneously completing the 3D shapes. We propose PALNet, a novel hybrid network for SSC based on single depth. PALNet utilizes a two-stream network to extract both 2D and 3D features from

Cited by 71SourcecodeScholar
2020

PillarFlow: End-to-end Birds-eye-view Flow Estimation for Autonomous Driving

IROS 2020poster

In autonomous driving, accurately estimating the state of surrounding obstacles is critical for safe and robust path planning. However, this perception task is difficult, particularly for generic obstacles/objects, due to appearance and occlusion changes. To tackle this problem, we propose an end-to…

Cited by 26SourceScholar
2020

Projection & Probability-Driven Black-Box Attack

CVPR 2020poster

Generating adversarial examples in a black-box setting retains a significant challenge with vast practical application prospects. In particular, existing black-box attacks suffer from the need for excessive queries, as it is non-trivial to find an appropriate direction to optimize in the high-dimens…

Cited by 58PDFcodeScholar
2020

Real-Time Panoptic Segmentation From Dense Detections

CVPR 2020oral

Panoptic segmentation is a complex full scene parsing task requiring simultaneous instance and semantic segmentation at high resolution. Current state-of-the-art approaches cannot run in real-time, and simplifying these architectures to improve efficiency severely degrades their accuracy. In this pa…

Cited by 94PDFScholar
2020

Semantically-Guided Representation Learning for Self-Supervised Monocular Depth

ICLR 2020poster

Self-supervised learning is showing great promise for monocular depth estimation, using geometry as the only source of supervision. Depth networks are indeed capable of learning representations that relate visual appearance to 3D properties by implicitly leveraging category-level patterns. In this w…

Cited by 285SourcecodeScholar
2019

RGBD Based Dimensional Decomposition Residual Network for 3D Semantic Scene Completion

CVPR 2019poster

RGB images differentiate from depth as they carry more details about the color and texture information, which can be utilized as a vital complement to depth for boosting the performance of 3D semantic scene completion (SSC). SSC is composed of 3D shape completion (SC) and semantic scene labeling whi…

Cited by 97PDFScholar
2019

Robust Semi-Supervised Monocular Depth Estimation with Reprojected Distances

CoRL 2019

Dense depth estimation from a single image is a key problem in computer vision, with exciting applications in a multitude of robotic tasks. Initially viewed as a direct regression problem, requiring annotated labels as supervision at training time, in the past few years a substantial amount of work

Cited by 0SourcePDFScholar
2019

The Speechtransformer for Large-scale Mandarin Chinese Speech Recognition

ICASSP 2019accepted

Attention-based sequence-to-sequence architectures have made great progress in the speech recognition task. The SpeechTransformer, a no-recurrence encoder-decoder architecture, has shown promising results on small-scale speech recognition data sets in previous works. In this paper, we focus on a lar…

Cited by 0SourceScholar
2019

Two Stream Networks for Self-Supervised Ego-Motion Estimation

CoRL 2019

Learning depth and camera ego-motion from raw unlabeled RGB video streams is seeing exciting progress through self-supervision from strong geometric cues. To leverage not only appearance but also scene geometry, we propose a novel self-supervised two-stream network using RGB and inferred depth infor

Cited by 0SourcePDFScholar
2019

Universal Adversarial Perturbation via Prior Driven Uncertainty Approximation

ICCV 2019oral

Deep learning models have shown their vulnerabilities to universal adversarial perturbations (UAP), which are quasi-imperceptible. Compared to the conventional supervised UAPs that suffer from the knowledge of training data, the data-independent unsupervised UAPs are more applicable. Existing unsupe…

Cited by 119PDFScholar
2019

ViLiVO: Virtual LiDAR-Visual Odometry for an Autonomous Vehicle with a Multi-Camera System

IROS 2019poster

In this paper, we present a multi-camera visual odometry (VO) system for an autonomous vehicle. Our system mainly consists of a virtual LiDAR and a pose tracker. We use a perspective transformation method to synthesize a surroundview image from undistorted fisheye camera images. With a semantic segm…

Cited by 18SourceScholar
2018

Exploring Sequential Characteristics in Speaker Bottleneck Feature for Text-Dependent Speaker Verification

ICASSP 2018accepted

In this paper, given the speaker bottleneck feature vectors extracted with speaker discriminant neural networks, we focus on using the sequential speaker characteristics for text-dependent speaker verification. In each evaluation trial, speaker supervectors are used as the representations of the seq…

Cited by 0SourceScholar
2018

WaterGAN: Unsupervised Generative Network to Enable Real-Time Color Correction of Monocular Underwater Images

RA-L 2018

This letter reports on WaterGAN, a generative adversarial network (GAN) for generating realistic underwater images from in-air image and depth pairings in an unsupervised pipeline used for color correction of monocular underwater images. Cameras onboard autonomous and remotely operated vehicles can

Cited by 848SourcecodeScholar
2016

Utilizing high-dimensional features for real-time robotic applications: Reducing the curse of dimensionality for recursive Bayesian estimation

IROS 2016poster

Feature learning has become popular in robotics due to recent advances in machine learning. In this paper, we propose a novel method to utilize the high-dimensional features from these techniques as observations in Bayesian estimation problems in a real-time manner. We develop an approach that: 1) p…

Cited by 25SourceScholar