← Search

Liang Li

114 accepted papers

2026

AutoPercep: A Pipeline for Onboard Neighbor Position Estimation Toward Large-Scale Swarm Robotics

ICRA 2026poster

Autonomous mobile robots must know each other's positions to coordinate their actions and motion. Beyond collision avoidance, relative position estimation is essential for spatial coordination tasks such as collective motion, leader–follower dynamics, or formation control.To overcome the scalability…

Cited by 0codeScholar
2026

Boosting Vehicle-to-Vehicle Collaborative Perception in Bird's-Eye View by Attentive Feature Fusion and Robust Pose Correction

RA-L 2026

Collaborative perception enables Connected Autonomous Vehicles (CAVs) to share sensory data, and therefore presents a promising path towards long-range robust environmental understanding by overcoming individual perception limitations such as occlusions. The core challenge of collaborative perceptio

Cited by 0SourcecodeScholar
2026

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models

CVPR 2026

Prompt learning has emerged as an efficient alternative to fine-tuning pre-trained vision-language models (VLMs).Despite its promise, current methods still struggle to maintain tail-class discriminability when adapting to class-imbalanced datasets. In this work, we propose cluster-aware neural colla

Cited by 0SourceScholar
2026

Consistent Noisy Latent Rewards for Trajectory Preference Optimization in Diffusion Models

ICLR 2026poster

Recent advances in diffusion models for visual generation have sparked interest in human preference alignment, similar to developments in Large Language Models. While reward model (RM) based approaches enable trajectory-aware optimization by evaluating intermediate timesteps, they face two critical…

Cited by 0SourceScholar
2026

Efficient Feature-Free Initialization for Monocular Visual-Inertial Systems Using A Feed-Forward 3D Model

RSS 2026poster

Fast and reliable initialization is critical for monocular visual–inertial navigation systems (VINS), as it establishes the starting conditions for subsequent state estimation. Despite steady progress, most existing methods heavily rely on visual feature correspondences and require 3-4 seconds of se…

Cited by 0SourceScholar
2026

Enhancing Communication Compression via Discrepancy-aware Calibration for Federated Learning

ICLR 2026poster

Federated Learning (FL) offers a privacy-preserving paradigm for distributed model training by enabling clients to collaboratively learn a shared model without exchanging their raw data. However, the communication overhead associated with exchanging model updates remains a critical challenge, partic…

Cited by 0SourcecodeScholar
2026

EvalMuse-40K: A Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Alignment Evaluation

AAAI 2026technical

Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated methods emerge to evaluate the image-text alignment capabilities of generative models. However, the performance comparison among these automated methods is constrained by the limited scale o

Cited by 0SourcePDFScholar
2026

Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object Detection

CVPR 2026

Domain adaptive object detection (DAOD) aims to generalize an object detector trained on a source domain to a target domain, where the domain gap degrades the adaptability. Recently, large-scale vision foundation models (VFMs), pretrained on web-scale datasets, exhibit such powerful generalization c

Cited by 0SourceScholar
2026

Forgetting Knowledge Localization and Isolation for Continual Forgetting of Pre-trained Vision Models

AAAI 2026technical

Continual forgetting task aims to continuously remove multiple target knowledge subsets from pre-trained models while maintaining the integrity of remaining knowledge. Existing methods suffer from both incomplete forgetting of target knowledge and unintended forgetting of indistinguishable remaining

Cited by 0SourcePDFScholar
2026

GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine-grained geometric structures, constraining their ability of geometric understanding and visual reasoning. To address this, we propose GeoTikzBrid

Cited by 0SourcecodeScholar
2026

HAIC: Humanoid Agile Object Interaction Control via Dynamics-Aware World Model

RSS 2026poster

Humanoid robots exhibit significant potential for executing complex whole-body interaction tasks in unstructured environments. While recent advancements in Human-Object Interaction (HOI) have been substantial, prevailing methodologies predominantly address the manipulation of fully actuated objects,…

Cited by 0SourceScholar
2026

Hint2Gen: Bridging Understanding and Generation via Code-structured Hints

CVPR 2026

Recent unified models have made remarkable strides in generating high-quality images, yet they consistently fail on reasoning-intensive tasks, i.e., solving mazes, assembling tangrams. Intriguingly, we find that vision-language models (VLMs) and large language models (LLMs) can accurately solve thes

Cited by 0SourceScholar
2026

InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing

AAAI 2026technical

Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character’s visual performance. However, existing alignment approaches based on visual features face two key limitations: (1) they r

Cited by 0SourcePDFScholar
2026

Language Does Matter for Cross-Domain Few-Shot Visual Feature Enhancement

CVPR 2026

Cross-domain few-shot image interpretation (CD-FSII) has been significantly advanced by fine-tuning pre-trained visual feature models using limited labeled samples in target domains. However, profound cross-domain distribution discrepancies, along with inherent conflicts between extensive object vis

Cited by 0SourcecodeScholar
2026

Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models

CVPR 2026

Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-image (T2I) tasks. Despite their theoretical promise, a persistent capability gap exists: UMMs typically exhibit superior vi

Cited by 0SourcecodeScholar
2026

Metric, Inertially Aligned Monocular State Estimation Via Kinetodynamic Priors

ICRA 2026poster

Accurate state estimation for flexible robotic systems poses significant challenges, particularly for platforms with dynamically deforming structures that invalidate rigid-body assumptions. This paper addresses this problem and enables the extension of existing rigid-body pose estimation methods to …

2026

STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning Models

AAAI 2026technical

Large Reasoning Models (LRMs) have advanced automated multi-step reasoning, but their ability to generate complex Chain-of-Thought (CoT) trajectories introduces severe privacy risks, as sensitive information may be deeply embedded throughout the reasoning process. Existing Large Language Models (LLM

Cited by 0SourcePDFScholar
2026

TEMPORAL-AWARE HETEROGENEOUS GRAPH REASONING WITH MULTI-VIEW FUSION FOR TEMPORAL QUESTION ANSWERING

ICASSP 2026poster

Question Answering over Temporal Knowledge Graphs (TKGQA) has attracted growing interest for handling time-sensitive queries. However, existing methods still struggle with: 1) weak incorporation of temporal constraints in question representation, causing biased reasoning; 2) limited ability to perfo…

Cited by 0SourcePDFScholar
2026

Temporal Calibrating and Distilling for Scene-Text Aware Text-Video Retrieval

AAAI 2026technical

Existing text-video retrieval methods mainly focus on singlemodal video content (i.e., visual entities), often overlooking heterogeneous scene text that is ubiquitous in human environments. Although scene text in videos provides finegrained semantics for cross-modal retrieval, effectively utilizing

Cited by 0SourcePDFScholar
2025

AEQA-NAT : Adaptive End-to-end Quantization Alignment Training Framework for Non-autoregressive Machine Translation

ICML 2025poster

Non-autoregressive Transformers (NATs) have garnered significant attention due to their efficient decoding compared to autoregressive methods. However, existing conditional dependency modeling schemes based on masked language modeling introduce a *training-inference gap* in NATs. For instance, while…

Cited by 0SourcePDFScholar
2025

AF-RLIO: Adaptive Fusion of Radar-LiDAR-Inertial Information for Robust Odometry in Challenging Environments

ICRA 2025

In robotic navigation, maintaining precise pose estimation and navigation in complex and dynamic environments is crucial. However, environmental challenges such as smoke, tunnels, and adverse weather can significantly degrade the performance of single-sensor systems like LiDAR or GPS, compromising t

Cited by 3SourcecodeScholar
2025

Change Entity-guided Heterogeneous Representation Disentangling for Change Captioning

ACL 2025finding

Change captioning aims to describe differences between a pair of images using natural language. However, learning effective difference representations is highly challenging due to distractors such as illumination and viewpoint changes. To address this, we propose a change-entity-guided disentangleme…

2025

DCTMamba: Advancing JPEG Image Restoration Through Long-Sequence Modeling and Adaptive Frequency Strategy

AAAI 2025technical

Despite the advanced long-sequence modeling of Mamba, which has expanded its applications in image restoration, there remains a lack of exploration combining its strengths with the specific characteristics of JPEG image restoration, where high-frequency components are lost after the Discrete Cosine…

2025

DEFORM: Adaptive Formation Reconfiguration of Multi-Robot Systems in Confined Environments

RA-L 2025

Achieving desired formation patterns without collisions is rather challenging for multi-robot systems in unknown obstacle-rich and confined environments, especially in narrow corridor scenes containing large-volume obstacles. To address this, we propose an adaptive formation reconfiguration method t

Cited by 5SourceScholar
2025

DHC-ME: A Decentralized Hybrid Cooperative Approach for Multi-Robot Autonomous Exploration

IROS 2025

Multi-robot exploration in unknown environments is a fundamental task for multi-robot systems, which requires the coordination of the robots to avoid collisions and conflicts while performing task allocation. Existing exploration strategies improve the efficiency of multi-robot exploration by modeli

Cited by 0SourcecodeScholar
2025

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

CVPR 2025highlight

Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing methods have two primary deficiencies: (1) They struggle to simultaneously hold audio-visual sync and achieve clear pron…

2025

Frequency Dynamic Convolution for Dense Image Prediction

CVPR 2025poster

While Dynamic Convolution (DY-Conv) has shown promising performance by enabling adaptive weight selection through multiple parallel weights combined with an attention mechanism, the frequency response of these weights tends to exhibit high similarity, resulting in high parameter costs but limited ad…

2025

Generalizing Single-Frame Supervision to Event-Level Understanding for Video Anomaly Detection

NeurIPS 2025poster

Video Anomaly Detection (VAD) aims to identify abnormal frames from discrete events within video sequences. Existing VAD methods suffer from heavy annotation burdens in fully-supervised paradigm, insensitivity to subtle anomalies in semi-supervised paradigm, and vulnerability to noise in weakly-supe…

Cited by 0SourceScholar
2025

Heterogeneous Prompt-Guided Entity Inferring and Distilling for Scene-Text Aware Cross-Modal Retrieval

AAAI 2025technical

In cross-modal retrieval, comprehensive image understanding is vital while the scene text in images can provide fine-grained information to understand visual semantics. Current methods fail to make full use of scene text. They suffer from the semantic ambiguity of independent scene text and overlook…

Cited by 0SourcePDFScholar
2025

LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization

NeurIPS 2025poster

We present LongVPO, a novel two‑stage Direct Preference Optimization framework that enables short‑context vision‑language models to robustly understand ultra‑long videos without any long‑video annotations. In Stage 1, we synthesize preference triples by anchoring questions to individual short clips,…

Cited by 0SourceScholar
2025

MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation

ACL 2025long

With the rapid adoption of large language models (LLMs) in natural language processing, the ability to follow instructions has emerged as a key metric for evaluating their practical utility. However, existing evaluation methods often focus on single-language scenarios, overlooking the challenges and…

2025

Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain Adaptation

CVPR 2025poster

This paper explores the Class-Incremental Source-Free Unsupervised Domain Adaptation (CI-SFUDA) problem, where the unlabeled target data come incrementally without access to labeled source instances. This problem poses two challenges, the interference of similar source-class knowledge in target-clas…

Cited by 1SourcePDFScholar
2025

Multi-Robot Autonomous 3D Reconstruction Using Gaussian Splatting With Semantic Guidance

RA-L 2025

Implicit neural representations and 3D Gaussian splatting (3DGS) have shown great potential for scene reconstruction. Recent studies have expanded their applications in autonomous reconstruction through task assignment methods. However, these methods are mainly limited to a single robot, and rapid r

Cited by 4SourceScholar
2025

Online Video Understanding: OVBench and VideoChat-Online

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have significantly progressed in offline video understanding. However, applying these models to real-world scenarios, such as autonomous driving and human-computer interaction, presents unique challenges due to the need for real-time processing of continuous…

Cited by 0SourcePDFScholar
2025

Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning

ICCV 2025poster

Audio-visual multi-task incremental learning aims to continuously learn from multiple audio-visual tasks without the need for joint training on all tasks. The challenge of the problem is how to preserve the old task knowledge while facilitating the learning of new task with previous experiences. To…

2025

Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing

CVPR 2025poster

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip. This task demands the model bridge character performances and complicated proso…

2025

Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning

AAAI 2025technical

Video has emerged as a favored multimedia format on the internet. To better gain video contents, a new topic HIREST is presented, including video retrieval, moment retrieval, moment segmentation, and step-captioning. The pioneering work chooses the pre-trained CLIP-based model for video retrieval,…

2025

Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change Captioning

AAAI 2025technical

Change captioning aims to describe the differences between two similar images using natural language, significantly aiding in understanding and monitoring changes. This challenging task requires a fine-grained understanding of subtle changes while resisting disturbances like viewpoint shifts and ill…

Cited by 0SourcePDFScholar
2025

Reinforcement Learning for Adaptive Planner Parameter Tuning: A Perspective on Hierarchical Architecture

ICRA 2025

Automatic parameter tuning methods for planning algorithms, which integrate pipeline approaches with learning-based techniques, are regarded as promising due to their stability and capability to handle highly constrained environments. While existing parameter tuning methods have demonstrated conside

Cited by 2SourceScholar
2025

SSTrack: Sample-interval Scheduling for Lightweight Visual Object Tracking

IJCAI 2025

In recent years, CPU real-time object tracking has gained significant attention due to its broad applications such as UAV-tracking. To maintain computational efficiency, most existing CPU real-time object trackers rely on lightweight backbones and employ a single initial template image without inter

2025

Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question Answering

CVPR 2025poster

Knowledge-based visual question answering (KBVQA) separates image interpretation and knowledge retrieval into separate processes, motivated in part by the fact that they are very different tasks. In this paper, we transform the KBVQA into linguistic question-answering tasks so that we can leverage t…

Cited by 0SourcePDFScholar
2025

Tracking Tiny Drones against Clutter: Large-Scale Infrared Benchmark with Motion-Centric Adaptive Algorithm

ICCV 2025poster

Tracking flying drones in infrared videos is a crucial yet challenging task. Existing drone trackers and datasets have limitations in dealing with and characterizing tiny targets (<=20x20 pixels) against highly complex backgrounds. To tackle this issue, we have developed a large-scale benchmark for…

2025

Union Is Strength! Unite the Power of LLMs and MLLMs for Chart Question Answering

AAAI 2025technical

Chart Question Answering (CQA) requires models to perform chart perception and reasoning. Recent studies driven by Large Language Models (LLMs) have dominated CQA. These include employing more cognitively capable LLMs for indirectly reasoning over transformed charts, i.e., tables, and directly perce…

2025

WHALE-FL: Wireless and Heterogeneity Aware Latency Efficient Federated Learning over Mobile Devices via Adaptive Subnetwork Scheduling

AAAI 2025technical

As a popular distributed learning paradigm, federated learning (FL) over mobile devices fosters numerous applications, while their practical deployment is hindered by participating devices' computing and communication heterogeneity. Some pioneering research efforts proposed to extract subnetworks fr…

Cited by 0SourcePDFScholar
2024

A Consistency-Aware Spot-Guided Transformer for Versatile and Hierarchical Point Cloud Registration

NeurIPS 2024poster

Deep learning-based feature matching has shown great superiority for point cloud registration in the absence of pose priors. Although coarse-to-fine matching approaches are prevalent, the coarse matching of existing methods is typically sparse and loose without consideration of geometric consistency…

2024

Context-aware Difference Distilling for Multi-change Captioning

ACL 2024long

Multi-change captioning aims to describe complex and coupled changes within an image pair in natural language. Compared with single-change captioning, this task requires the model to have higher-level cognition ability to reason an arbitrary number of changes. In this paper, we propose a novel conte…

2024

Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning

ECCV 2024poster

"Change captioning aims to succinctly describe the semantic change between a pair of similar images, while being immune to distractors (illumination and viewpoint changes). Under these distractors, unchanged objects often appear pseudo changes about location and scale, and certain objects might over…

2024

Free your mouse! Command Large Language Models to Generate Code to Format Word Documents

EMNLP 2024main

Recently, LLMs have significantly improved code generation, making it increasingly accessible to users. As a result, LLM-powered code generation applications have sprung up, vastly boosting user productivity. This paper mainly explores how to improve the efficiency and experience of users in formatt…

2024

Hydrodynamic Interactions in Schooling Fish: Prioritizing Real Fish Kinematics Over Travelling-wavy Undulation

ICRA 2024poster

Hydrodynamic interactions are crucial for understanding fish movement, particularly within the realm of robotic applications. Traditionally, many studies have favoured simplified travelling-wavy undulations derived from observed real fish kinematics. This approach often neglects higher-order undulat…

Cited by 1SourceScholar
2024

KDD-LOAM: Jointly Learned Keypoint Detector and Descriptors Assisted LiDAR Odometry and Mapping

ICRA 2024poster

Sparse keypoint matching based on distinct 3D feature representations can improve the efficiency and robustness of point cloud registration. Existing learning-based 3D descriptors and keypoint detectors are either independent or loosely coupled, so they cannot fully adapt to each other. In this work…

Cited by 5SourceScholar
2024

LESS-Map: Lightweight and Evolving Semantic Map in Parking Lots for Long-term Self-Localization

ICRA 2024poster

Precise and long-term stable localization is essential in parking lots for tasks like autonomous driving or autonomous valet parking, etc. Existing methods rely on a fixed and memory-inefficient map, which lacks robust data association approaches. And it is not suitable for precise localization or l…

Cited by 2SourceScholar
2024

Leveraging Catastrophic Forgetting to Develop Safe Diffusion Models against Malicious Finetuning

NeurIPS 2024spotlight

Diffusion models (DMs) have demonstrated remarkable proficiency in producing images based on textual prompts. Numerous methods have been proposed to ensure these models generate safe images. Early methods attempt to incorporate safety filters into models to mitigate the risk of generating harmful im…

Cited by 1SourcePDFScholar
2024

LiDAR-PTQ: Post-Training Quantization for Point Cloud 3D Object Detection

ICLR 2024poster

Due to highly constrained computing power and memory, deploying 3D lidar-based detectors on edge devices equipped in autonomous vehicles and robots poses a crucial challenge. Being a convenient and straightforward model compression approach, Post-Training Quantization (PTQ) has been widely adopted i…

2024

Local and Global: Text Matching Via Syntax Graph Calibration

ICASSP 2024accepted

Pre-trained models such as BERT have achieved remarkable results in text matching tasks. However, existing models still suffer from the challenge of capturing local subtle differences when modeling complex semantic matching relationships. In this work, we find that the integration of local syntax aw…

Cited by 0SourceScholar
2024

Modeling Route Representation With Mixed-Scale Hierarchical Transformer

ICASSP 2024accepted

Modeling route representation aims to obtain contextual representations of an entire route for various traffic-related tasks. In reality, spatial-temporal data often exhibits multi-scale characteristics, which are utilized by many studies to enhance their performance. However, there is still a lack…

Cited by 0SourceScholar
2024

Norm Tweaking: High-Performance Low-Bit Quantization of Large Language Models

AAAI 2024technical

As the size of large language models (LLMs) continues to grow, model compression without sacrificing accuracy has become a crucial challenge for deployment. While some quantization methods, such as GPTQ, have made progress in achieving acceptable 4-bit weight-only quantization, attempts at lower-bi…

Cited by 34SourcePDFScholar
2024

Prompt-Enhanced Multiple Instance Learning for Weakly Supervised Video Anomaly Detection

CVPR 2024poster

Weakly-supervised Video Anomaly Detection (wVAD) aims to detect frame-level anomalies using only video-level labels in training. Due to the limitation of coarse-grained labels Multi-Instance Learning (MIL) is prevailing in wVAD. However MIL suffers from insufficiency of binary supervision to model d…

2024

Prompting Large Language Models with Fine-Grained Visual Relations from Scene Graph for Visual Question Answering

ICASSP 2024accepted

Visual Question Answering (VQA) is a task that requires models to comprehend both questions and images. An increasing number of works are leveraging the strong reasoning capabilities of Large Language Models (LLMs) to address VQA. These methods typically utilize image captions as visual text descrip…

Cited by 0SourceScholar
2024

R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation

ICLR 2024poster

Recent text-to-image (T2I) diffusion models have achieved remarkable progress in generating high-quality images given text-prompts as input. However, these models fail to convey appropriate spatial composition specified by a layout instruction. In this work, we probe into zero-shot grounded T2I gene…

2024

Segment then Match: Find the Carrier before Reasoning in Scene-Text VQA

ICASSP 2024accepted

Text-based Visual Question Answering (TextVQA) requires models to answer questions about the scene text in images by reasoning the context between the scene text and the question. Previous works demonstrated that clustering the scene text could help the model understand the context between different…

Cited by 0SourceScholar
2024

StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing

ACL 2024findings

Given a script, the challenge in Movie Dubbing (Visual Voice Cloning, V2C) is to generate speech that aligns well with the video in both time and emotion, based on the tone of a reference audio track. Existing state-of-the-art V2C models break the phonemes in the script according to the divisions be…

2024

SynSP: Synergy of Smoothness and Precision in Pose Sequences Refinement

CVPR 2024poster

Predicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily co…

2024

Trajectory set Empowered Hypergraph Transformer for Mobile Sensor Based Traffic Prediction

ICASSP 2024accepted

Traffic speed prediction is vital for intelligent transportation systems. However, most existing methods focus on costly static sensors. In contrast, utilizing GPS devices from vehicles as mobile sensors offers a cost-effective means to gather dynamic traffic data. Despite the presence of historical…

Cited by 0SourceScholar
2024

iBoW3D: Place Recognition Based on Incremental and General Bag of Words in 3D Scans

ICRA 2024poster

Existing methods for place recognition in 3D point clouds either ignore partial structure information by converting 3D scans to 2D images or construct constrained bag-of-words (BoW) representations reliant on specific feature extraction algorithms. In this paper, we propose a novel method based on i…

Cited by 0SourceScholar
2023

ACROSS: An Alignment-based Framework for Low-Resource Many-to-One Cross-Lingual Summarization

ACL 2023findings

This research addresses the challenges of Cross-Lingual Summarization (CLS) in low-resource scenarios and over imbalanced multilingual data. Existing CLS studies mostly resort to pipeline frameworks or multi-task methods in bilingual settings. However, they ignore the data imbalance in multilingual…

2023

CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High Quality

ACL 2023long

There are three problems existing in the popular data-to-text datasets. First, the large-scale datasets either contain noise or lack real application scenarios. Second, the datasets close to real applications are relatively small in size. Last, current datasets bias in the English language while lea…

Cited by 2SourcePDFScholar
2023

Hard Sample Aware Network for Contrastive Deep Graph Clustering

AAAI 2023technical

Contrastive deep graph clustering, which aims to divide nodes into disjoint groups via contrastive mechanisms, is a challenging research spot. Among the recent works, hard sample mining-based algorithms have achieved great attention for their promising performance. However, we find that the existing…

2023

Learning To Dub Movies via Hierarchical Prosody Models

CVPR 2023poster

Given a piece of text, a video clip and a reference audio, the movie dubbing (also known as visual voice clone, V2C) task aims to generate speeches that match the speaker's emotion presented in the video using the desired speaker voice as reference. V2C is more challenging than conventional text-to-…

2023

Let the Data Choose: Flexible and Diverse Anchor Graph Fusion for Scalable Multi-View Clustering

AAAI 2023technical

In the past few years, numerous multi-view graph clustering algorithms have been proposed to enhance the clustering performance by exploring information from multiple views. Despite the superior performance, the high time and space expenditures limit their scalability. Accordingly, anchor graph lear…

2023

Nasty-SFDA: Source Free Domain Adaptation from a Nasty Model

ICASSP 2023accepted

A challenging problem called Nasty Source Free Domain Adaptation (Nasty-SFDA) is proposed in this work, where only a nasty source model and unlabeled target samples are available for DA. Further, after DA, the target model is expected to be a nasty model. In order to deal with Nasty-SFDA, Nasty HypO…

Cited by 0SourceScholar
2023

PBACalib: Targetless Extrinsic Calibration for High-Resolution LiDAR-Camera System Based on Plane-Constrained Bundle Adjustment

RA-L 2023

The strategy of fusing multi-model data especially from cameras, light detection and ranging sensors (LiDAR), is frequently considered in robotics to enhance the performance of the perception and navigation tasks. Extrinsic calibration, which spatially aligns different sources into a unified coordin

Cited by 20SourceScholar
2023

Self-supervised Cross-view Representation Reconstruction for Change Captioning

ICCV 2023poster

Change captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused by viewpoint change. In this paper, we address this by proposing a self-supervised cross-view representation reconstruc…

Cited by 36PDFcodeScholar
2023

Text-Driven Generative Domain Adaptation with Spectral Consistency Regularization

ICCV 2023poster

Combined with the generative prior of pre-trained models and the flexibility of text, text-driven generative domain adaptation can generate images from a wide range of target domains. However, current methods still suffer from overfitting and the mode collapse problem. In this paper, we analyze the…

Cited by 8PDFcodeScholar
2023

ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual Tracking

NeurIPS 2023spotlight

Recently, the transformer has enabled the speed-oriented trackers to approach state-of-the-art (SOTA) performance with high-speed thanks to the smaller input size or the lighter feature extraction backbone, though they still substantially lag behind their corresponding performance-oriented versions.…

2022

Automatic Relation-Aware Graph Network Proliferation

CVPR 2022oral

Graph neural architecture search has sparked much attention as Graph Neural Networks (GNNs) have shown powerful reasoning capability in many relational tasks. However, the currently used graph search space overemphasizes learning node features and neglects mining hierarchical relational information.…

Cited by 12PDFcodeScholar
2022

Debiased Batch Normalization via Gaussian Process for Generalizable Person Re-identification

AAAI 2022technical

Generalizable person re-identification aims to learn a model with only several labeled source domains that can perform well on unseen domains. Without access to the unseen domain, the feature statistics of the batch normalization (BN) layer learned from a limited number of source domains is doubtles…

Cited by 36SourcePDFScholar
2022

FaceVerse: A Fine-Grained and Detail-Controllable 3D Face Morphable Model From a Hybrid Dataset

CVPR 2022poster

We present FaceVerse, a fine-grained 3D Neural Face Model, which is built from hybrid East Asian face datasets containing 60K fused RGB-D images and 2K high-fidelity 3D head scan models. A novel coarse-to-fine structure is proposed to take better advantage of our hybrid dataset. In the coarse module…

Cited by 110PDFcodeScholar
2022

Fast-R2D2: A Pretrained Recursive Neural Network based on Pruned CKY for Grammar Induction and Text Representation

EMNLP 2022main

Chart-based models have shown great potential in unsupervised grammar induction, running recursively and hierarchically, but requiring O(n³) time-complexity. The Recursive Transformer based on Differentiable Trees (R2D2) makes it possible to scale to large language model pretraining even with a comp…

2022

Few Shot Generative Model Adaption via Relaxed Spatial Structural Alignment

CVPR 2022poster

Training a generative adversarial network (GAN) with limited data has been a challenging task. A feasible solution is to start with a GAN well-trained on a large scale source domain and adapt it to the target domain with a few samples, termed as few shot generative model adaption. However, existing…

Cited by 90PDFcodeScholar
2022

Graph-to-Text Generation with Dynamic Structure Pruning

COLING 2022main

Most graph-to-text works are built on the encoder-decoder framework with cross-attention mechanism. Recent studies have shown that explicitly modeling the input graph structure can significantly improve the performance. However, the vanilla structural encoder cannot capture all specialized informati…

2022

Leveraging Inter-Layer Dependency for Post -Training Quantization

NeurIPS 2022accept

Prior works on Post-training Quantization (PTQ) typically separate a neural network into sub-nets and quantize them sequentially. This process pays little attention to the dependency across the sub-nets, hence is less optimal. In this paper, we propose a novel Network-Wise Quantization (NWQ) approac…

Cited by 23SourcePDFScholar
2022

Modality-Adaptive Mixup and Invariant Decomposition for RGB-Infrared Person Re-identification

AAAI 2022technical

RGB-infrared person re-identification is an emerging cross-modality re-identification task, which is very challenging due to significant modality discrepancy between RGB and infrared images. In this work, we propose a novel modality-adaptive mixup and invariant decomposition (MID) approach for RGB-i…

Cited by 105SourcePDFScholar
2022

Open-Vocabulary One-Stage Detection With Hierarchical Visual-Language Knowledge Distillation

CVPR 2022poster

Open-vocabulary object detection aims to detect novel object categories beyond the training set. The advanced open-vocabulary two-stage detectors employ instance-level visual-to-visual knowledge distillation to align the visual space of the detector with the semantic space of the Pre-trained Visual-…

Cited by 106PDFcodeScholar
2022

Think Beyond Words: Exploring Context-Relevant Visual Commonsense for Diverse Dialogue Generation

EMNLP 2022finding

Commonsense knowledge has been widely considered for building intelligent open-domain dialogue agents, aiming to generate meaningful and diverse responses. Previous works in this field usually lack the ability to effectively obtain and utilize auxiliary commonsense from the external visual world. In…

2022

Unsupervised Coherent Video Cartoonization with Perceptual Motion Consistency

AAAI 2022technical

In recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry. Different from image translations focusing on improving the st…

2021

Context-Aware Safe Reinforcement Learning for Non-Stationary Environments

ICRA 2021poster

Safety is a critical concern when deploying reinforcement learning agents for realistic tasks. Recently, safe reinforcement learning algorithms have been developed to optimize the agent’s performance while avoiding violations of safety constraints. However, few studies have addressed the nonstationa…

Cited by 45SourceScholar
2021

Globally Optimal Fetoscopic Mosaicking Based on Pose Graph Optimisation With Affine Constraints

RA-L 2021

Fetoscopic laser ablation surgery could be guided using a high-quality panorama of the operating site, representing a map of the placental vasculature. This can be achieved during the initial inspection phase of the procedure using image mosaicking techniques. Due to the lack of camera calibration i

Cited by 17SourceScholar
2021

Improving Encoder by Auxiliary Supervision Tasks for Table-to-Text Generation

ACL 2021long

Table-to-text generation aims at automatically generating natural text to help people conveniently obtain salient information in tables. Although neural models for table-to-text have achieved remarkable progress, some problems are still overlooked. Previous methods cannot deduce the factual results…

2021

Rethinking Graph Neural Architecture Search From Message-Passing

CVPR 2021poster

Graph neural networks (GNNs) emerged recently as a standard toolkit for learning from data on graphs. Current GNN designing works depend on immense human expertise to explore different message-passing mechanisms, and require manual enumeration to determine the proper message-passing depth. Inspired…

Cited by 66PDFcodeScholar
2021

Rˆ3Net:Relation-embedded Representation Reconstruction Network for Change Captioning

EMNLP 2021main

Change captioning is to use a natural language sentence to describe the fine-grained disagreement between two similar images. Viewpoint change is the most typical distractor in this task, because it changes the scale and location of the objects and overwhelms the representation of real change. In th…

2021

Structured Multi-Level Interaction Network for Video Moment Localization via Language Query

CVPR 2021poster

We address the problem of localizing a specific moment described by a natural language query. Existing works interact the query with either video frame or moment proposal, and neglect the inherent structure of moment construction for both cross-modal understanding and video content comprehension, wh…

Cited by 102PDFScholar
2021

Symbolic Music Generation with Transformer-GANs

AAAI 2021technical

Autoregressive models using Transformers have emerged as the dominant approach for music generation with the goal of synthesizing minute-long compositions that exhibit large-scale musical structure. These models are commonly trained by minimizing the negative log-likelihood (NLL) of the obse…

Cited by 84SourcePDFScholar
2021

TreeBERT: A tree-based pre-trained model for programming language

UAI 2021poster

Source code can be parsed into the abstract syntax tree (AST) based on defined syntax rules. However, in pre-training, little work has considered the incorporation of tree structure into the learning process. In this paper, we present TreeBERT, a tree-based pre-trained model for improving programmin…

2020

A Structured Latent Variable Recurrent Network With Stochastic Attention For Generating Weibo Comments

IJCAI 2020poster

Building intelligent agents to generate realistic Weibo comments is challenging. For such realistic Weibo comments, the key criterion is improving diversity while maintaining coherency. Considering that the variability of linguistic comments arises from multi-level sources, including both discourse-…

2020

Enhancing Urban Flow Maps via Neural ODEs

IJCAI 2020poster

Flow super-resolution (FSR) enables inferring fine-grained urban flows with coarse-grained observations and plays an important role in traffic monitoring and prediction. The existing FSR solutions rely on deep CNN models (e.g., ResNet) for learning spatial correlation, incurring excessive memory cos…

2020

Parsing-Based View-Aware Embedding Network for Vehicle Re-Identification

CVPR 2020poster

Vehicle Re-Identification is to find images of the same vehicle from various views in the cross-camera scenario. The main challenges of this task are the large intra-instance distance caused by different views and the subtle inter-instance discrepancy caused by similar vehicles. In this paper, we pr…

Cited by 255PDFcodeScholar
2020

Real-World Person Re-Identification via Degradation Invariance Learning

CVPR 2020poster

Person re-identification (Re-ID) in real-world scenarios usually suffers from various degradation factors, e.g., low-resolution, weak illumination, blurring and adverse weather. On the one hand, these degradations lead to severe discriminative information loss, which significantly obstructs identity…

Cited by 89PDFScholar
2020

Retinal Vessel Segmentation via a Semantics and Multi-Scale Aggregation Network

ICASSP 2020accepted

Precise segmentation of retinal vessels is crucial for a computer-aided diagnosis system of retinal fundus images. However, this task remains challenging due to large variations in scales and poor segmentation of capillary vessels. In this paper, we propose a semantics and multi-scale aggregation ne…

Cited by 0SourceScholar
2020

State-Relabeling Adversarial Active Learning

CVPR 2020oral

Active learning is to design label-efficient algorithms by sampling the most representative samples to be labeled by an oracle. In this paper, we propose a state relabeling adversarial active learning model (SRAAL), that leverages both the annotation and the labeled/unlabeled state information for d…

Cited by 160PDFScholar
2020

Towards Discriminability and Diversity: Batch Nuclear-Norm Maximization Under Label Insufficient Situations

CVPR 2020oral

The learning of the deep networks largely relies on the data with human-annotated labels. In some label insufficient situations, the performance degrades on the decision boundary with high data density. A common solution is to directly minimize the Shannon Entropy, but the side effect caused by entr…

Cited by 489PDFcodeScholar
2019

Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding

ICCV 2019poster

Weakly supervised referring expression grounding aims at localizing the referential object in an image according to the linguistic query, where the mapping between the referential object and query is unknown in the training stage. To address this problem, we propose a novel end-to-end adaptive recon…

Cited by 110PDFcodeScholar
2017

A Graph Regularized Deep Neural Network for Unsupervised Image Representation Learning

CVPR 2017poster

Deep Auto-Encoder (DAE) has shown its promising power in high-level representation learning. From the perspective of manifold learning, we propose a graph regularized deep neural network (GR-DNN) to endue traditional DAEs with the ability of retaining local geometric structure. A deep-structured reg…

Cited by 37PDFcodeScholar
2017

Gaussian mixture model-signature quadratic form distance based point set registration

IROS 2017poster

Point set registration is a long addressed problem in lots of pattern recognition tasks. This paper presents a robust point set registration algorithm based on optimization of distance between two probability distributions. A major problem encountered in the point to point algorithms is the definiti…

Cited by 7SourceScholar
2017

Human-inspired compliant strategy for peg-in-hole assembly using environmental constraint and coarse force information

IROS 2017poster

Automated assembly, especially peg-in-hole insertion, is a common task in manufacturing. In particular, the high-precision assembly is achieved by high-precision manipulator and sensing system. However, uncertainty and various parts for assembly are still challenges for robotic assembly, especially…

Cited by 29SourceScholar