← Search

Bing Li

132 accepted papers

2026

AutoQVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization

ICLR 2026poster

The advent of Vision-Language-Action (VLA) models represents a significant leap for embodied intelligence, yet their immense computational demands critically hinder deployment on resource-constrained robotic platforms. Intuitively, low-bit quantization is a prevalent and preferred technique for larg…

Cited by 0SourcecodeScholar
2026

Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval

CVPR 2026

Composed Image Retrieval (CIR) has demonstrated significant potential by enabling flexible, multimodal queries that combine a reference image and modification text. However, CIR inherently prioritizes semantic matching, struggling to reliably retrieve a user-specified instance across contexts. In pr

Cited by 0SourceScholar
2026

Beyond Sequences: A Dynamic Hierarchical Heterogeneous Spatio-Temporal Graph for Irregular Multivariate Time Series Forecasting

IJCAI 2026

Irregular Multivariate Time Series (IMTS) analysis is a challenging task as asynchronous irregular sampling disrupts intra-variable temporal consistency and cross-variable alignment. Most existing methods model multivariate correlations at either the variable or observation level in static ways. The

Cited by 0Scholar
2026

CHESS: Chebyshev Spectral Synthesis for Trajectory Condensation

ICML 2026poster

Learning from continuous-time trajectories requires modeling multivariate sensor measurements generated by underlying physical or dynamical processes. Under extreme data compression and heterogeneous sampling, directly optimizing synthetic signals as discrete sample values becomes fundamentally misa…

Cited by 0SourceScholar
2026

Constant Degree Matrix-Driven Incomplete Multi-View Clustering via Connectivity-Structure and Embedding Tensor Learning

ICLR 2026poster

Tensor-based incomplete multi-view clustering has attracted significant research attention due to its capability to exploit high-order correlations across different views for revealing underlying cluster structures from partially observed multi-view data. However, most existing approaches construct…

Cited by 0SourceScholar
2026

Dual-Branch Representations with Dynamic Gated Fusion and Triple-Granularity Alignment for Deep Multi-View Clustering

ICLR 2026poster

Multi-view clustering seeks to exploit complementary information across different views to enhance clustering performance, where both semantic and structural information are crucial. However, existing approaches often bias toward one type of information while treating the other as auxiliary, overloo…

Cited by 0SourceScholar
2026

Ev-iCRF: Self-supervised Event-guided iCRF Estimation for HDR Image Reconstruction

AAAI 2026technical

In this paper, we present Ev-iCRF, a novel self-supervised pipeline for high dynamic range (HDR) image reconstruction from a single-exposure low dynamic range (LDR) image, guided by asynchronous event streams generated by a bio-inspired event camera. The highlight of Ev-iCRF lies in its formulation

Cited by 0SourcePDFScholar
2026

Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New Thinking

AAAI 2026technical

In this paper, we rethink model agent behaviors from a geometric structure perspective in multi-agent reinforcement learning. Modeling agent behaviors is essential for understanding how agents interact and facilitating effective decisions. The key lies in capturing the dependencies and sequential re

Cited by 0SourcePDFScholar
2026

Improving Autoregressive Video Modeling with History Understanding

ICLR 2026poster

Video autoregressive generation (VideoAR) sequentially predicts future frames conditioned on history frames. Despite the advance of recent diffusion-based VideoAR, the role of conditioning signal—internal representations of history frames—remains underexplored. Inspired by the success of strong cond…

Cited by 0SourceScholar
2026

Light but Sharp: SlimSTAD for Real-Time Action Detection from Sensor Data

AAAI 2026technical

Sensory Temporal Action Detection (STAD) aims to localize and classify human actions within long, untrimmed sequences captured by non-visual sensors such as WiFi or inertial measurement units (IMUs). Unlike video-based TAD, STAD poses unique challenges due to the low-dimensional, noisy, and heteroge

Cited by 0SourcePDFScholar
2026

MMhops-R1: Multimodal Multi-hop Reasoning

AAAI 2026technical

The ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex real-world challenges. However, existing Multi-modal Large Language Models (MLLMs) are predominantly limited to single-ste

Cited by 0SourcePDFScholar
2026

Misclassification-Aware Robust Learning from Multiple Human Labelers (Student Abstract)

AAAI 2026technical

Adversarial training is an effective technique for enhancing the robustness of deep neural networks (DNNs). Prior research shows that misclassified examples influence final adversarial robustness much more than correctly classified examples. Ignoring this difference during training can hurt model pe

Cited by 0SourcePDFScholar
2026

NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge Devices

CVPR 2026

Vision Transformers (ViTs) often need to be compressed for deployment on resource-constrained edge devices like drones and smart vehicles. However, existing model compression methods ignore that many edge devices only require the knowledge of specific classes for their applications. As a result, the

Cited by 0SourcecodeScholar
2026

OBJVanish: Prompt-Driven Generation of Physically Realizable 3D LiDAR-Invisible Objects

ICML 2026poster

LiDAR-based 3D object detectors are fundamental to autonomous driving, where missed detections pose severe safety risks. While adversarial attacks are crucial for evaluating the robustness of these detectors, existing point-level perturbation methods rarely cause complete object disappearance and pr…

Cited by 0SourceScholar
2026

Resolving the Timestep Scaling Paradox in Spiking Neural Networks with a Timestep-Scalable Neuron Model

ICML 2026poster

Spiking Neural Networks (SNNs) have garnered increasing attention for their biological plausibility, energy efficiency, and temporal modeling capability. Due to the non-differentiability of spike generation, a widely used supervised training method for SNNs is backpropagation through time with surro…

Cited by 0SourceScholar
2026

SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Names

CVPR 2026

Open-vocabulary object detection (OVD) aims to detect objects described by arbitrary text, but most existing methods operate at a coarse category level and struggle with fine-grained, attribute-sensitive queries. We address this from both model and data perspectives. We propose a Semantic-Retrieval-

Cited by 0SourceScholar
2026

SiMO: Single-Modality-Operable Multimodal Collaborative Perception

ICLR 2026poster

Collaborative perception integrates multi-agent perspectives to enhance the sensing range and overcome occlusion issues. While existing multimodal approaches leverage complementary sensors to improve performance, they are highly prone to failure—especially when a key sensor like LiDAR is unavailable…

Cited by 0SourcecodeScholar
2026

Sim2Real Diffusion: Leveraging Foundation Vision Language Models for Adaptive Automated Driving

RA-L 2026

Simulation-based design, optimization, and validation of autonomous vehicles have proven to be crucial for their improvement over the years. Nevertheless, the ultimate measure of effectiveness is their successful transition from simulation to reality (sim2real). However, existing sim2real transfer m

Cited by 3SourcecodeScholar
2026

SpikeClouds: Streaming Spike-Based Processing of LiDAR for Fast and Efficient Object Detection

ICRA 2026poster

LiDAR sensors are used to provide three-dimensional information about the environment in many robotics applications. The information, accumulated in 3D point clouds, is first acquired by the sensor and then processed further, which leads to high end-to-end latencies and large memory footprints. Stre…

Cited by 0SourceScholar
2026

WiTTA-Bench: Benchmarking Test-Time Adaptation for WiFi Sensing

CVPR 2026

WiFi sensing offers passive and privacy-preserving perception that complements vision-based sensing, but its performance degrades sharply under domain shifts caused by changes in environment, subjects, or hardware. This challenge is exacerbated in real-world deployments where source data are unavail

Cited by 0SourcecodeScholar
2025

4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

ICCV 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive 2D image/video understanding capabilities.However, there are no publicly standardized benchmarks to assess the abilities of MLLMs in understanding the 4D objects.In this paper, we introduce 4D-Bench, the first benchmark to evaluat…

2025

An Empirical Study of Federated Prompt Learning for Vision Language Model

IJCAI 2025

The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt learning with VLM in federated learning (FL) scenarios remains underexplored. Th

2025

An Inflatable Deployable Origami Grasper for Adaptive and High-Load Grasping

IROS 2025

Robotic graspers are essential for enhancing the efficiency and versatility of robots in grasping tasks. In this paper, we propose a novel inflatable deployable origami grasper with a rigid-flexible coupling structure. The proposed grasper can achieve multiple deployment configurations under a singl

Cited by 0SourceScholar
2025

Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model Compression

ICLR 2025poster

Large Language Models (LLMs) have achieved remarkable breakthroughs. However, the huge number of parameters in LLMs require significant amount of memory storage in inference, which prevents their practical deployment in many applications. To reduce memory storage of LLMs, singular value decompositio…

2025

Cockroach's Turning Strategy Enhanced Hexapod Robot with Flexible Torso

IROS 2025

The design and control of hexapod robots have become an active research field due to the ability to achieve adaptive and stable multi-terrain locomotion. However, existing hexapod robots focus on the integration of flexible pitch joints to enhance their obstacle-crossing and slope-climbing abilities

Cited by 0SourceScholar
2025

Codar: Complex-valued Neural Network for Crossing-Floor Intrusion Detection via WiFi

ICASSP 2025accepted

WiFi systems offer enormous potential for device-free human intrusion detection. Current methods often require routers to be deployed in multiple adjacent rooms on the same floor, which is redundant and costly. To solve this, we introduce the first work on intrusion detection in the crossing-floor s…

Cited by 0SourceScholar
2025

D-RAG: Differentiable Retrieval-Augmented Generation for Knowledge Graph Question Answering

EMNLP 2025

Knowledge Graph Question Answering (KGQA) aims to answer natural language questions based on knowledge graphs.Recent approaches apply the Retrieval-Augmented Generation (RAG) paradigm to incorporate Large Language Models (LLMs) to this task, where a retriever selects a question-related subgraph and

Cited by 0SourcePDFScholar
2025

DeepTAGE: Deep Temporal-Aligned Gradient Enhancement for Optimizing Spiking Neural Networks

ICLR 2025poster

Spiking Neural Networks (SNNs), with their biologically inspired spatio-temporal dynamics and spike-driven processing, are emerging as a promising low-power alternative to traditional Artificial Neural Networks (ANNs). However, the complex neuronal dynamics and non-differentiable spike communication…

Cited by 0SourcePDFScholar
2025

Federated Recommendation with Explicitly Encoding Item Bias

AAAI 2025technical

With the development of federated learning techniques and the increased need for user privacy protection, the federated recommendation has become a new recommendation paradigm. However, most existing works focus on user-level federated recommendation, leaving platform-level federated recommendation…

Cited by 0SourcePDFScholar
2025

MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural Networks

NeurIPS 2025poster

Brain-inspired spiking neural networks (SNNs) provide energy-efficient computation through event-driven processing. However, the shared weights across multiple timesteps lead to serious temporal feature redundancy, limiting both efficiency and performance. This issue is further aggravated when proce…

Cited by 0SourcecodeScholar
2025

Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations

EMNLP 2025

Large Vision-Language Models (LVLMs) suffer from serious hallucination problems, where the model-generated responses are inconsistent with the visual inputs. Existing hallucination mitigation methods are mainly based on preference alignment and require external human annotations or auxiliary models

2025

Multi-Drone-Truck Collaborative Delivery with En Route Operations: A Hierarchical MARL-Based Approach

ICRA 2025

The multi-drone-truck collaborative delivery, where unmanned trucks serve as mobile supply stations for drones, effectively combines the strengths of both vehicles and presents wide application prospects. But the majority of existing literature restricts drone launch and retrieve operations (LARO) t

Cited by 1SourceScholar
2025

Multimodal Large Language Model-Guided ISP Hyperparameter Optimization with Dynamic Preference Learning

ICCV 2025poster

The image signal processing (ISP) pipeline is responsible for converting the RAW images collected from the sensor into high-quality RGB images. It contains a series of image processing modules and associated ISP hyperparameters. Recent learning-based approaches aim to automate ISP hyperparameter opt…

Cited by 0SourcePDFScholar
2025

OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions

NeurIPS 2025poster

In this paper, we introduce Online Multimodal Conversational Response Generation (OMCRG), a novel task designed to produce synchronized verbal and non-verbal listener feedback online, based on the speaker's multimodal inputs. OMCRG captures natural dyadic interactions and introduces new challenges i…

Cited by 0SourcecodeScholar
2025

One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention Head

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in tasks requiring multimodal understanding. However, recent studies indicate that LVLMs are more vulnerable than LLMs to unsafe inputs and prone to generating harmful content. Existing defense strategies primarily includ…

Cited by 0SourcecodeScholar
2025

Point Cloud Self-supervised Learning via 3D to Multi-view Masked Learner

ICCV 2025poster

Recently, multi-modal masked autoencoders (MAE) has been introduced in 3D self-supervised learning, offering enhanced feature learning by leveraging both 2D and 3D data to capture richer cross-modal representations. However, these approaches have two limitations: (1) they inefficiently require both…

Cited by 0SourcePDFScholar
2025

Reversing Flow for Image Restoration

CVPR 2025poster

Image restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, wh…

Cited by 0SourcePDFScholar
2025

SSTrack: Sample-interval Scheduling for Lightweight Visual Object Tracking

IJCAI 2025

In recent years, CPU real-time object tracking has gained significant attention due to its broad applications such as UAV-tracking. To maintain computational efficiency, most existing CPU real-time object trackers rely on lightweight backbones and employ a single initial template image without inter

2025

Similar Modality Enhancement and Action Consistency Learning for Weakly Supervised Temporal Action Localization

AAAI 2025technical

Weakly-supervised temporal action localization (WTAL) aims to identify and localize action instances in untrimmed videos using only video-level labels. Existing methods typically rely on original features from frozen pre-trained encoders designed for trimmed action classification (TAC) tasks, which…

2025

SpikeClouds: Streaming Spike-Based Processing of LiDAR for Fast and Efficient Object Detection

RA-L 2025

LiDAR sensors are used to provide three-dimensional information about the environment in many robotics applications. The information, accumulated in 3D point clouds, is first acquired by the sensor and then processed further, which leads to high end-to-end latencies and large memory footprints. Stre

Cited by 1SourceScholar
2025

SynCL: A Synergistic Training Strategy with Instance-Aware Contrastive Learning for End-to-End Multi-Camera 3D Tracking

NeurIPS 2025poster

While existing query-based 3D end-to-end visual trackers integrate detection and tracking via the *tracking-by-attention* paradigm, these two chicken-and-egg tasks encounter optimization difficulties when sharing the same parameters. Our findings reveal that these difficulties arise due to two inher…

Cited by 0SourcecodeScholar
2025

SynFER: Towards Boosting Facial Expression Recognition with Synthetic Data

ICCV 2025poster

Facial expression datasets remain limited in scale due to privacy concerns, the subjectivity of annotations, and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foun…

Cited by 0SourcePDFScholar
2025

Towards More Discriminative Feature Learning in SNNs with Temporal-Self-Erasing Supervision

AAAI 2025technical

Spiking Neural Networks (SNNs) are biologically inspired models that process visual inputs over multiple time steps. However, they often struggle with limited feature discrimination along the temporal dimension due to inherent spatiotemporal invariance. This limitation arises from the redundant acti…

Cited by 0SourcePDFScholar
2025

Union Is Strength! Unite the Power of LLMs and MLLMs for Chart Question Answering

AAAI 2025technical

Chart Question Answering (CQA) requires models to perform chart perception and reasoning. Recent studies driven by Large Language Models (LLMs) have dominated CQA. These include employing more cognitively capable LLMs for indirectly reasoning over transformed charts, i.e., tables, and directly perce…

2025

VisionMath: Vision-Form Mathematical Problem-Solving

ICCV 2025poster

Mathematical problems in real-world scenarios are often presented in a purely vision-form, where textual problem statement and accompanying math figures, e.g., geometry figures and functional graphs, are integrated into a single image. This vision-form problem-solving task requires precise comprehen…

2025

Visual-Instructed Degradation Diffusion for All-in-One Image Restoration

CVPR 2025poster

Image restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that…

2025

WiFi CSI Based Temporal Activity Detection via Dual Pyramid Network

AAAI 2025technical

We address the challenge of WiFi-based temporal activity detection and propose an efficient Dual Pyramid Network that integrates Temporal Signal Semantic Encoders and Local Sensitive Response Encoders. The Temporal Signal Semantic Encoder splits feature learning into high and low-frequency componen…

2024

A Hybrid CNN-Transformer for Focal Liver Lesion Classification

ICASSP 2024accepted

The early diagnosis of focal liver lesions (FLLs) plays a key role in the successful treatment of liver cancer. To effectively diagnose focal liver lesions, we used contrast-enhanced ultrasound (CEUS) to diagnose FLLs. A hybrid CNN and Transformer network is used to extract local and global spatio-t…

Cited by 0SourceScholar
2024

Benchmarking Segmentation Models with Mask-Preserved Attribute Editing

CVPR 2024poster

When deploying segmentation models in practice it is critical to evaluate their behaviors in varied and complex scenes. Different from the previous evaluation paradigms only in consideration of global attribute variations (e.g. adverse weather) we investigate both local and global attribute variatio…

2024

Diagnosis of Autism Spectrum Disorder Based on Contrastive Functional Connectivity Graph Learning Network

ICASSP 2024accepted

To reduce the dependence on tagged data, we proposed a Contrastive Functional Connectivity Graph Learning Network (CFCG-Net) for the diagnosis of autism spectrum disorder. CFCG-Net is mainly composed of three parts: construction of contrastive Functional Connection (FC) graphs, learning of contrasti…

Cited by 0SourceScholar
2024

EA-VTR: Event-Aware Video-Text Retrieval

ECCV 2024poster

"Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and the widely adopted video-level cross-modal contrastive learning also struggles to…

Cited by 3SourcePDFScholar
2024

How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?

CVPR 2024poster

Dominant dual-encoder models enable efficient image-text retrieval but suffer from limited accuracy while the cross-encoder models offer higher accuracy at the expense of efficiency. Distilling cross-modality matching knowledge from cross-encoder to dual-encoder provides a natural approach to harnes…

Cited by 2SourcePDFScholar
2024

Invisible Backdoor Attack against 3D Point Cloud Classifier in Graph Spectral Domain

AAAI 2024technical

3D point cloud has been wildly used in security crucial domains, such as self-driving and 3D face recognition. Backdoor attack is a serious threat that usually destroy Deep Neural Networks (DNN) in the training stage. Though a few 3D backdoor attacks are designed to achieve guaranteed attack efficie…

2024

MIBench: Evaluating Multimodal Large Language Models over Multiple Images

EMNLP 2024main

Built on the power of LLMs, numerous multimodal large language models (MLLMs) have recently achieved remarkable performance on various vision-language tasks. However, most existing MLLMs and benchmarks primarily focus on single-image input scenarios, leaving the performance of MLLMs when handling re…

Cited by 10SourcePDFScholar
2024

Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors

ICLR 2024poster

We present ``Magic123'', a two-stage coarse-to-fine approach for high-quality, textured 3D mesh generation from a single image in the wild using *both 2D and 3D priors*. In the first stage, we optimize a neural radiance field to produce a coarse geometry. In the second stage, we adopt a memory-effic…

2024

NARUTO: Neural Active Reconstruction from Uncertain Target Observations

CVPR 2024poster

We present NARUTO a neural active reconstruction system that combines a hybrid neural representation with uncertainty learning enabling high-fidelity surface reconstruction. Our approach leverages a multi-resolution hash-grid as the mapping backbone chosen for its exceptional convergence speed and c…

2024

Prompting Large Language Models with Fine-Grained Visual Relations from Scene Graph for Visual Question Answering

ICASSP 2024accepted

Visual Question Answering (VQA) is a task that requires models to comprehend both questions and images. An increasing number of works are leveraging the strong reasoning capabilities of Large Language Models (LLMs) to address VQA. These methods typically utilize image captions as visual text descrip…

Cited by 0SourceScholar
2024

RL-SeqISP: Reinforcement Learning-Based Sequential Optimization for Image Signal Processing

AAAI 2024technical

Hardware image signal processing (ISP), aiming at converting RAW inputs to RGB images, consists of a series of processing blocks, each with multiple parameters. Traditionally, ISP parameters are manually tuned in isolation by imaging experts according to application-specific quality and performance…

Cited by 4SourcePDFScholar
2024

SAM-Guided Masked Token Prediction for 3D Scene Understanding

NeurIPS 2024poster

Foundation models have significantly enhanced 2D task performance, and recent works like Bridge3D have successfully applied these models to improve 3D scene understanding through knowledge distillation, marking considerable advancements. Nonetheless, challenges such as the misalignment between 2D an…

Cited by 1SourcePDFScholar
2024

Segment then Match: Find the Carrier before Reasoning in Scene-Text VQA

ICASSP 2024accepted

Text-based Visual Question Answering (TextVQA) requires models to answer questions about the scene text in images by reasoning the context between the scene text and the question. Previous works demonstrated that clustering the scene text could help the model understand the context between different…

Cited by 0SourceScholar
2024

Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training

COLING 2024main

In vision-language pre-training (VLP), masked image modeling (MIM) has recently been introduced for fine-grained cross-modal alignment. However, in most existing methods, the reconstruction targets for MIM lack high-level semantics, and text is not sufficiently involved in masked modeling. These two…

Cited by 0SourcePDFScholar
2024

Set Prediction Guided by Semantic Concepts for Diverse Video Captioning

AAAI 2024technical

Diverse video captioning aims to generate a set of sentences to describe the given video in various aspects. Mainstream methods are trained with independent pairs of a video and a caption from its ground-truth set without exploiting the intra-set relationship, resulting in low diversity of generated…

Cited by 3SourcePDFScholar
2024

TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks

ECCV 2024poster

"Neural radiance fields (NeRFs) generally require many images with accurate poses for accurate novel view synthesis, which does not reflect realistic setups where views can be sparse and poses can be noisy. Previous solutions for learning NeRFs with sparse views and noisy poses only consider local g…

2024

Tune-An-Ellipse: CLIP Has Potential to Find What You Want

CVPR 2024highlight

Visual prompting of large vision language models such as CLIP exhibits intriguing zero-shot capabilities. A manually drawn red circle commonly used for highlighting can guide CLIP's attention to the surrounding region to identify specific objects within an image. Without precise object proposals how…

2024

Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval

COLING 2024main

In video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representations for videos and texts. However, they face challenges in capturing fine-grained semantic concepts. In this work, we prop…

2024

Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing

ECCV 2024poster

"Diffusion models have recently been investigated as powerful generative solvers for image dehazing, owing to their remarkable capability to model the data distribution. However, the massive computational burden imposed by the retraining of diffusion models, coupled with the extensive sampling steps…

Cited by 1SourcePDFScholar
2024

VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization

NeurIPS 2024poster

Bird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion, unfavourable imaging conditions and low resolution, \emph{g…

2024

Vivid-ZOO: Multi-View Video Generation with Diffusion Model

NeurIPS 2024poster

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of massive captioned multi-view videos and the complexity of modeli…

Cited by 11SourcePDFScholar
2023

AUNet: Learning Relations Between Action Units for Face Forgery Detection

CVPR 2023poster

Face forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Recent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same domain. However, the problem remains chall…

Cited by 56SourcePDFScholar
2023

AdaptiveMix: Improving GAN Training via Feature Space Shrinkage

CVPR 2023poster

Due to the outstanding capability for data generation, Generative Adversarial Networks (GANs) have attracted considerable attention in unsupervised learning. However, training GANs is difficult, since the training distribution is dynamic for the discriminator, leading to unstable image representatio…

2023

An Origami-Based Miniature Jumping Robot with Adjustable Jumping Trajectory and Enhanced Intermittent Jumps

IROS 2023poster

A small-scale jumping robot can reach obstacles much larger than its size. It is important for a jumping robot to perform intermittent jumps to cross through rough terrains. However, the limitations of conventional structures hinder the further integration of functions to a miniature (sub-50 g) jump…

Cited by 1SourceScholar
2023

Automatic Animation of Hair Blowing in Still Portrait Photos

ICCV 2023poster

We propose a novel approach to animate human hair in a still portrait photo. Existing work has largely studied the animation of fluid elements such as water and fire. However, hair animation for a real image remains underexplored, which is a challenging problem, due to the high complexity of hair st…

Cited by 10PDFcodeScholar
2023

Bridging the Domain Gap: Self-Supervised 3D Scene Understanding with Foundation Models

NeurIPS 2023poster

Foundation models have achieved remarkable results in 2D and language tasks like image segmentation, object detection, and visual-language understanding. However, their potential to enrich 3D scene representation learning is largely untapped due to the existence of the domain gap. In this work, we p…

2023

CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High Quality

ACL 2023long

There are three problems existing in the popular data-to-text datasets. First, the large-scale datasets either contain noise or lack real application scenarios. Second, the datasets close to real applications are relatively small in size. Last, current datasets bias in the English language while lea…

Cited by 2SourcePDFScholar
2023

CVRecon: Rethinking 3D Geometric Feature Learning For Neural Reconstruction

ICCV 2023poster

Recent advances in neural reconstruction using posed image sequences have made remarkable progress. However, due to the lack of depth information, existing volumetric-based techniques simply duplicate 2D image features of the object surface along the entire camera ray. We contend this duplication in…

Cited by 22PDFScholar
2023

Combating Mode Collapse via Offline Manifold Entropy Estimation

AAAI 2023technical

Generative Adversarial Networks (GANs) have shown compelling results in various tasks and applications in recent years. However, mode collapse remains a critical problem in GANs. In this paper, we propose a novel training pipeline to address the mode collapse issue of GANs. Different from existing m…

2023

Dipping PLMs Sauce: Bridging Structure and Text for Effective Knowledge Graph Completion via Conditional Soft Prompting

ACL 2023findings

Knowledge Graph Completion (KGC) often requires both KG structural and textual information to be effective. Pre-trained Language Models (PLMs) have been used to learn the textual information, usually under the fine-tune paradigm for the KGC task. However, the fine-tuned PLMs often overwhelmingly foc…

2023

Dynamically Masked Discriminator for GANs

NeurIPS 2023poster

Training Generative Adversarial Networks (GANs) remains a challenging problem. The discriminator trains the generator by learning the distribution of real/generated data. However, the distribution of generated data changes throughout the training process, which is difficult for the discriminator to…

2023

Exploiting Contextual Objects and Relations for 3D Visual Grounding

NeurIPS 2023poster

3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information…

2023

FourStr: When Multi-sensor Fusion Meets Semi-supervised Learning

ICRA 2023poster

This research proposes a novel semi-supervised learning framework FourStr (Four-Stream formed by two two-stream models) that focuses on the improvement of fusion and labeling efficiency for 3D multi-sensor detector. FourStr adopts a multi-sensor single-stage detector named adaptive fusion network (A…

Cited by 1SourceScholar
2023

Learning To Exploit the Sequence-Specific Prior Knowledge for Image Processing Pipelines Optimization

CVPR 2023poster

The hardware image signal processing (ISP) pipeline is the intermediate layer between the imaging sensor and the downstream application, processing the sensor signal into an RGB image. The ISP is less programmable and consists of a series of processing modules. Each processing module handles a subta…

Cited by 7SourcePDFScholar
2023

Learning from the Raw Domain: Cross Modality Distillation for Compressed Video Action Recognition

ICASSP 2023accepted

Video action recognition is faced with the challenges of both huge computation burden and performance requirements. Using compressed domain data, which saves much decoding computation, is a possible solution. Unfortunately, existing compressed-domain-based (CD) methods fail to obtain high performanc…

Cited by 0SourceScholar
2023

Learning to Identify Critical States for Reinforcement Learning from Videos

ICCV 2023poster

Recent work on deep reinforcement learning (DRL) has pointed out that algorithmic information about good policies can be extracted from offline data which lack explicit information about executed actions. For example, videos of humans or robots may convey a lot of implicit information about rewardin…

Cited by 12PDFcodeScholar
2023

NewsNet: A Novel Dataset for Hierarchical Temporal Segmentation

CVPR 2023poster

Temporal video segmentation is the get-to-go automatic video analysis, which decomposes a long-form video into smaller components for the following-up understanding tasks. Recent works have studied several levels of granularity to segment a video, such as shot, event, and scene. Those segmentations…

2023

Order-Prompted Tag Sequence Generation for Video Tagging

ICCV 2023poster

Video Tagging intends to infer multiple tags spanning relevant content for a given video. Typically, video tags are freely defined and uploaded by a variety of users, so they have two characteristics: abundant in quantity and disordered intra-video. It is difficult for the existing multi-label class…

Cited by 4PDFScholar
2023

ViLEM: Visual-Language Error Modeling for Image-Text Retrieval

CVPR 2023poster

Dominant pre-training works for image-text retrieval adopt "dual-encoder" architecture to enable high efficiency, where two encoders are used to extract image and text representations and contrastive learning is employed for global alignment. However, coarse-grained global alignment ignores detailed…

Cited by 14SourcePDFScholar
2023

ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual Tracking

NeurIPS 2023spotlight

Recently, the transformer has enabled the speed-oriented trackers to approach state-of-the-art (SOTA) performance with high-speed thanks to the smaller input size or the lighter feature extraction backbone, though they still substantially lag behind their corresponding performance-oriented versions.…

2022

A Modular Lockable Mechanism for Tendon-Driven Robots: Design, Modeling and Characterization

RA-L 2022

Surgical robots with variable stiffness can provide more stable configurations during minimally-invasive surgery.This letter presents a new design of modular lockable mechanism which can be used to change the stiffness of tendon-driven surgical robots. Locking and unlocking are simply actuated by pu

Cited by 24SourceScholar
2022

Attention-Aware Learning for Hyperparameter Prediction in Image Processing Pipelines

ECCV 2022poster

"Between the imaging sensor and the image applications, the hardware image signal processing (ISP) pipelines reconstruct an RGB image from the sensor signal and feed it into downstream tasks. The processing blocks in ISPs depend on a set of tunable hyperparameters that have a complex interaction wit…

Cited by 13SourcePDFScholar
2022

Design and Experimental Validation of a Shock-Absorption Mechanism Inspired From the Frog's Forelimbs

RA-L 2022

Frogs reveal superior shock-absorption ability, which is mainly brought by the forelimbs. Frogs touch the ground with their forelimbs first in landing, followed by the elbow joints’ compression. In this process, the muscles and tendons of the forelimbs are pulled to store energy. However, muscles an

Cited by 0SourceScholar
2022

Design and Kinematic Modeling of In-Situ Torsionally-Steerable Flexible Surgical Robots

RA-L 2022

Flexible robots have been widely used in minimally-invasive surgery because of their dexterity and accuracy. However, the torsion of end-effector under bending state usually produces unnecessary concomitant motion of the robot arm. This letter proposes a novel design of <i>in-situ</i> torsionally-st

Cited by 10SourceScholar
2022

Dexterity Analysis and Motion Optimization of In-Situ Torsionally-Steerable Flexible Surgical Robots

RA-L 2022

Flexible robots with <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">in-situ</i> torsion can be used in laryngeal endoscopic surgery which can maintain the position and approach vector of the end-effector during the operation. However, the inherent e

Cited by 16SourceScholar
2022

Disentangling Object Motion and Occlusion for Unsupervised Multi-Frame Monocular Depth

ECCV 2022poster

"Conventional self-supervised monocular depth prediction methods are based on a static environment assumption, which leads to accuracy degradation in dynamic scenes due to the mismatch and occlusion problems introduced by object motions. Existing dynamic-object-focused methods only partially solved…

2022

EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding Matching

CVPR 2022poster

Current metrics for video captioning are mostly based on the text-level comparison between reference and candidate captions. However, they have some insuperable drawbacks, e.g., they cannot handle videos without references, and they may result in biased evaluation due to the one-to-many nature of vi…

Cited by 43PDFcodeScholar
2022

FocusTR: Focusing on Valuable Feature by Multiple Transformers for Fusing Feature Pyramid on Object Detection

IROS 2022poster

The feature pyramid, which is a vital component of the convolutional neural networks, plays a significant role in several perception tasks, including object detection for autonomous driving. However, how to better fuse multi-level and multi-sensor feature pyramids is still a significant challenge, e…

Cited by 3SourceScholar
2022

Improving Visual Grounding With Visual-Linguistic Verification and Iterative Reasoning

CVPR 2022poster

Visual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They base the visual grounding on the features from pre-generated proposals or anchors, and fuse these features with the text em…

Cited by 147PDFcodeScholar
2022

Knowledge Is Flat: A Seq2Seq Generative Framework for Various Knowledge Graph Completion

COLING 2022main

Knowledge Graph Completion (KGC) has been recently extended to multiple knowledge graph (KG) structures, initiating new research directions, e.g. static KGC, temporal KGC and few-shot KGC. Previous works often design KGC models closely coupled with specific graph structures, which inevitably results…

2022

Learning Target-aware Representation for Visual Tracking via Informative Interactions

IJCAI 2022poster

We introduce a novel backbone architecture to improve target-perception ability of feature representation for tracking. Having observed de facto frameworks perform feature matching simply using the backbone outputs for target localization, there is no direct feedback from the matching module to the…

Cited by 61SourcePDFScholar
2022

Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video Classification

IJCAI 2022poster

Compared with image few-shot learning, most of the existing few-shot video classification methods perform worse on feature matching, because they fail to sufficiently exploit the temporal information and relation. Specifically, frames are usually evenly sampled, which may miss important frames. On t…

Cited by 16SourcePDFScholar
2022

MDERank: A Masked Document Embedding Rank Approach for Unsupervised Keyphrase Extraction

ACL 2022findings

Keyphrase extraction (KPE) automatically extracts phrases in a document that provide a concise summary of the core content, which benefits downstream information retrieval and NLP tasks. Previous state-of-the-art methods select candidate keyphrases based on the similarity between learned representat…

2022

One More Check: Making “Fake Background” Be Tracked Again

AAAI 2022technical

The one-shot multi-object tracking, which integrates object detection and ID embedding extraction into a unified network, has achieved groundbreaking results in recent years. However, current one-shot trackers solely rely on single-frame detections to predict candidate bounding boxes, which may be u…

2022

Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement

IJCAI 2022poster

Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues (e.g., motion vectors and residuals). However, this task severely suffers from t…

Cited by 19SourcePDFScholar
2022

SCTN: Sparse Convolution-Transformer Network for Scene Flow Estimation

AAAI 2022technical

We propose a novel scene flow estimation approach to capture and infer 3D motions from point clouds. Estimating 3D motions for point clouds is challenging, since a point cloud is unordered and its density is significantly non-uniform. Such unstructured data poses difficulties in matching correspondi…

2022

The Feedback Trajectory Control of a SMA-Driven Miniature Jumping Robot

ICRA 2022poster

Jumping motion is an effective way to overcome large obstacles, especially for the miniature robots. However, controlling of the jumping trajectory on a centimeter scale robot is not easy due to the limitation of size and payload. None of the jumping robots lighter than 90 g achieved the feedback co…

Cited by 8SourceScholar
2021

A Lightweight Soft Gripper Driven by Self-Sensing Super-Coiled Polymer Actuator

RA-L 2021

This study proposes a soft gripper driven by a simple, lightweight, low-cost, self-sensing super-coiled polymer (SCP) actuator with a high power-to-weight ratio. SCP generates an untwisting motion when heated, and therefore, is suitable for designing the actuator. The actuator is fabricated with a m

Cited by 22SourceScholar
2021

Advancing Self-supervised Monocular Depth Learning with Sparse LiDAR

CoRL 2021poster

Self-supervised monocular depth prediction provides a cost-effective solution to obtain the 3D location of each pixel. However, the existing approaches usually lead to unsatisfactory accuracy, which is critical for autonomous robots. In this paper, we propose FusionDepth, a novel two-stage network t…

Cited by 35SourceScholar
2021

Automatic Construction of Enterprise Knowledge Base

EMNLP 2021system demonstrations

In this paper, we present an automatic knowledge base construction system from large scale enterprise documents with minimal efforts of human intervention. In the design and deployment of such a knowledge mining system for enterprise, we faced several challenges including data distributional shift,…

Cited by 6SourcePDFScholar
2021

Channel-Wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition

ICCV 2021poster

Graph convolutional networks (GCNs) have been widely used and achieved remarkable results in skeleton-based action recognition. In GCNs, graph topology dominates feature aggregation and therefore is the key to extracting representative features. In this work, we propose a novel Channel-wise Topology…

Cited by 872PDFcodeScholar
2021

DPFPS: Dynamic and Progressive Filter Pruning for Compressing Convolutional Neural Networks from Scratch

AAAI 2021technical

Filter pruning is a commonly used method for compressing Convolutional Neural Networks (ConvNets), due to its friendly hardware supporting and flexibility. However, existing methods mostly need a cumbersome procedure, which brings many extra hyper-parameters and training epochs. This is because only…

2021

High Quality Disparity Remapping With Two-Stage Warping

ICCV 2021poster

A high quality disparity remapping method that preserves 2D shapes and 3D structures, and adjusts disparities of important objects in stereo image pairs is proposed. It is formulated as a constrained optimization problem, whose solution is challenging, since we need to meet multiple requirements of…

Cited by 2PDFScholar
2021

Hybrid Adaptive Control Strategy for Continuum Surgical Robot Under External Load

RA-L 2021

Natural orifice transluminal endoscopic surgery (NOTES) has received significant attentions due to its minimal incision trauma compared with traditional multi-port robot assisted surgery. Continuum robot can be used in NOTES due to its high flexibility which can adapt to circuitous paths. However, t

Cited by 45SourceScholar
2021

Improving the Efficiency and Effectiveness for BERT-based Entity Resolution

AAAI 2021technical

BERT has set a new state-of-the-art performance on entity resolution (ER) task, largely owed to fine-tuning pre-trained language models and the deep pair-wise interaction. Albeit being remarkably effective, it comes with a steep increase in computational cost, as the deep-interaction requires to exh…

2021

Learn To Match: Automatic Matching Network Design for Visual Tracking

ICCV 2021poster

Siamese tracking has achieved groundbreaking performance in recent years, where the essence is the efficient matching operator cross-correlation and its variants. Besides the remarkable success, it is important to note that the heuristic matching network design relies heavily on expert experience. M…

Cited by 239PDFcodeScholar
2021

Open-Book Video Captioning With Retrieve-Copy-Generate Network

CVPR 2021poster

In this paper, we convert traditional video captioning task into a new paradigm, i.e., Open-book Video Captioning, which generates natural language under the prompts of video-content-relevant sentences, not limited to the video itself. To address the open-book video captioning problem, we propose a…

Cited by 125PDFScholar
2021

Sent2Span: Span Detection for PICO Extraction in the Biomedical Text without Span Annotations

EMNLP 2021finding

The rapid growth in published clinical trials makes it difficult to maintain up-to-date systematic reviews, which require finding all relevant trials. This leads to policy and practice decisions based on out-of-date, incomplete, and biased subsets of available clinical evidence. Extracting and then…

2021

Two-Stream Convolution Augmented Transformer for Human Activity Recognition

AAAI 2021technical

Recognition of human activities is an important task due to its far-reaching applications such as healthcare system, context-aware applications, and security monitoring. Recently, WiFi based human activity recognition (HAR) is becoming ubiquitous due to its non-invasiveness. Existing WiFi-based HAR…

2020

Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model

ECCV 2020poster

Recently, video streams have occupied a large proportion of Internet traffic, most of which contain human faces. Hence, it is necessary to predict saliency on multiple-face videos, which can provide attention cues for many content based applications. However, most of multiple-face prediction works o…

2020

Object Relational Graph With Teacher-Recommended Learning for Video Captioning

CVPR 2020poster

Taking full advantage of the information from both vision and language is critical for the video captioning task. Existing models lack adequate visual representation due to the neglect of interaction between object, and sufficient training for content-related words due to long-tailed problems. In th…

Cited by 387PDFScholar
2019

Deep Neural Network based Visual Inspection with 3D Metric Measurement of Concrete Defects using Wall-climbing Robot

IROS 2019poster

This paper presents a novel metric inspection robot system using a deep neural network to detect and measure surface flaws (i.e., crack and spalling) on concrete structures performed by a wall-climbing robot. The system consists of four modules: robotics data collection module to obtain RGB-D images…

Cited by 28SourceScholar
2019

Knowledge Distillation via Instance Relationship Graph

CVPR 2019poster

The key challenge of knowledge distillation is to extract general, moderate and sufficient knowledge from a teacher network to guide a student network. In this paper, a novel Instance Relationship Graph (IRG) is proposed for knowledge distillation. It models three kinds of knowledge, including insta…

Cited by 371PDFScholar
2018

Interaction-aware Spatio-temporal Pyramid Attention Networks for Action Classification

ECCV 2018poster

Local features at neighboring spatial positions in feature maps have high correlation since their receptive fields are often overlapped. Self-attention usually uses the weighted sum (or other functions) with internal elements of each local feature to obtain its weight score, which ignores interactio…

Cited by 118SourcePDFScholar
2017

Spatio-Temporal Self-Organizing Map Deep Network for Dynamic Object Detection From Videos

CVPR 2017poster

In dynamic object detection, it is challenging to construct an effective model to sufficiently characterize the spatial-temporal properties of the background. This paper proposes a new Spatio-Temporal Self-Organizing Map (STSOM) deep network to detect dynamic objects in complex scenarios. The propos…

Cited by 18PDFScholar
2015

Design and analysis of parallel robots for a flexible fixturing system with performance atlases

IROS 2015poster

According to the automobile industry's flexible manufacturing requirements, a novel flexible fixturing system for sheet metal assembly is proposed with parallel robots. A methodology of the structure synthesis is presented by taking account simultaneously several performance indices. Taking the 3UPU…

Cited by 2SourceScholar
Bing Li — accepted AI-conference papers · AIConfPaper