← Search

Li Zhang

243 accepted papers

2026

A Hybrid Magnetic Actuation System for Hybrid Microrobotic Targeted Delivery

ICRA 2026poster

Magnetic microrobots hold great promise for biomedical applications. However, achieving flexible magnetic field adjustment with a magnetic actuation system (MAS) to actuate diverse microrobots remains a significant challenge. In this work, we propose an Electromagnetic-Permanent Magnet Actuation (EP…

Cited by 0Scholar
2026

ACO-MoE-LoRA: Evolving-while-Training for Adapting Segment Anything Model 2 to Specialized Domains

ICML 2026poster

Static fine-tuning paradigms impose rigid structural constraints on foundation models like the Segment Anything Model 2 (SAM2), limiting their adaptability to the varying complexity of specialized downstream tasks. To overcome this limitation, we propose **ACO-MoE-LoRA**, a dynamic framework that in…

Cited by 0SourceScholar
2026

Anatomical Prior-Driven Framework for Autonomous Robotic Cardiac Ultrasound Standard View Acquisition

ICRA 2026poster

Cardiac ultrasound diagnosis is critical for cardiovascular disease assessment, but acquiring standard views remains highly operator-dependent. Existing medical segmentation models often yield anatomically inconsistent results in images with poor textural differentiation between distinct feature cla…

2026

Bi-directional Autoregressive Diffusion for Large Complex Motion Interpolation

CVPR 2026

Despite recent progress, diffusion-based video frame interpolation methods still struggle with large, complex motions, resulting in discontinuous motions and inconsistent object appearances across frames. We observe that these limitations arise from both the current full-sequence interpolation strat

Cited by 0SourceScholar
2026

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation (VLN) requires the agent to navigate based on natural instructions. This task is challenging due to partial observability, which makes it difficult to align perception with language. Recent methods mitigate this by imagining future scenes, yet they rely on vision-based

Cited by 0SourcecodeScholar
2026

DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces

CVPR 2026

Articulated object pose estimation is a core task in embodied AI and computer vision. Existing methods typically regress poses in a continuous space, but often struggle with 1) navigating a large, complex search space and 2) failing to incorporate intrinsic kinematic constraints. In this paper, we i

Cited by 0SourceScholar
2026

DNA: Uncovering Universal Latent Forgery Knowledge

ICML 2026poster

As generative AI achieves hyper-realism, superficial artifact detection has become obsolete. While prevailing methods rely on resource-intensive fine-tuning of black-box backbones, we propose that forgery detection capability is already encoded within pre-trained models rather than requiring end-to-…

Cited by 0SourceScholar
2026

Demystifying the Optimal Fair Classifier in Multi-Class Classification

ICML 2026poster

Ensuring fair and equitable treatment across diverse groups, particularly in multi-class classification tasks, poses a significant challenge due to the persistent biases inherent in machine learning models. Most existing bias mitigation techniques are tailored to binary settings, and the presence of…

Cited by 0SourceScholar
2026

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

ICLR 2026poster

Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma–misalignment of 2D image forecasting and 3D action prediction. Besides, such a vision-action entangled training manner limits model learning from large-scale, action-free web video…

Cited by 0SourcecodeScholar
2026

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding

ICML 2026poster

While speculative decoding improves inference throughput for multi-batch long-context Large Language Models (LLMs), its efficiency is often limited by a verification bottleneck where Key-Value (KV) cache loading dominates latency. Existing compression methods fail in this regime: static eviction inc…

Cited by 0SourceScholar
2026

ELSPR: Evaluator LLM Training Data Self-Purification on Non-Transitive Preferences via Tournament Graph Reconstruction

AAAI 2026technical

Pairwise evaluation of large language models (LLMs) has become the dominant paradigm for benchmarking open-ended tasks, yet non-transitive preferences—where evaluators prefer A over B, B over C, but C over A—fundamentally undermine ranking reliability. We show that this critical issue stems largely

Cited by 0SourcePDFScholar
2026

EvoFMVC: Trusted Federated Multi-View Clustering with Evolutionary Fusion

AAAI 2026technical

With the growing demand for decentralized collaborative analysis of privacy-sensitive data, federated multi-view clustering (FMVC) has attracted widespread attention due to its ability to balance privacy protection and collaborative modeling. However, current methods still face the following challen

Cited by 0SourcePDFScholar
2026

Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds

AAAI 2026technical

Articulated objects are prevalent in daily life and robotic manipulation tasks. However, compared to rigid objects, pose tracking for articulated objects remains an underexplored problem due to their inherent kinematic constraints. To address these challenges, this work proposes a novel point-pair-b

Cited by 0SourcePDFScholar
2026

F2SST: Frequency-to-Spatial Semantic Transfer for Few-Shot Image Classification

AAAI 2026technical

Few-shot image classification (FSIC) aims to recognize novel categories from only a few labeled examples, making it inherently challenging under limited supervision. Existing approaches have attempted to alleviate this issue by incorporating explicit semantics like class names or knowledge graphs to

Cited by 0SourcePDFScholar
2026

Fine-Grained Activation Steering: Steering Less, Achieving More

ICLR 2026poster

Activation steering has emerged as a cost-effective paradigm for modifying large language model (LLM) behaviors. Existing methods typically intervene at the block level, steering the bundled activations of selected attention heads, feedforward networks, or residual streams. However, we reveal that b…

Cited by 0SourcecodeScholar
2026

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

ICML 2026poster

Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined futures for action execution remains challenging. Existing approaches either condition policies on predicted frames or dire…

Cited by 0SourceScholar
2026

GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning

ICLR 2026poster

Training effective Vision-Language Models (VLMs) for GUI agents typically depends on large-scale annotated datasets, whose collection is both labor-intensive and error-prone. We introduce K-step GUI Transition, a self-supervised inverse dynamics task in which VLMs learn GUI dynamics by predicting th…

Cited by 0SourcecodeScholar
2026

GauMVC: Generative Decoupled Gaussian Representation for Human-centric Multi-view Video Compression

CVPR 2026

Human-centric multi-view video has a clear semantic structure: a static background and dynamic human motion. We propose a generative compression framework that explicitly decouples these components. The background is modeled once with 3D Gaussian Splatting, while the human is represented by a person

Cited by 0SourceScholar
2026

GeoTeacher: Geometry-Guided Semi-Supervised 3D Object Detection

ICRA 2026poster

Semi-supervised 3D object detection (SS3D), aiming to explore unlabeled data for boosting 3D object detectors, has emerged as an active research area in recent years. Some previous methods have shown substantial improvements by either employing heterogeneous teacher models to provide high-quality ps…

2026

Harmonic Principle-Based Adjustable Constant-Force Mechanism for Stable Human-Robot Interaction in Massage Robotics

RA-L 2026

Stable and precise force contact is critical in human-robot interaction, typically requiring precise sensors and advanced algorithms. This paper presents a novel adjustable constant-force mechanism (CFM) serving as a force generator for massage robotic systems, aiming to reduce hardware costs and sy

Cited by 0SourceScholar
2026

High Resolution Neural Video Coding with Bi-directional Confidence-Guided Reference Information Modeling

CVPR 2026

Exploiting bi-directional context prediction has long been recognized as a key direction for improving compression efficiency in neural video coding. However, existing neural B-frame codecs still exhibit limited performance gains, particularly in high-resolution videos with large motion, where optic

Cited by 0SourceScholar
2026

IEBGL:An Interpretability-Enhanced Brain Graph Learning Framework with LLM-Instructed Topology and Literature-Augmented Semantics

CVPR 2026

Resting-state functional MRI (rs-fMRI) provides rich information for modeling brain connectivity in disease diagnosis. However, most existing brain graph learning methods rely solely on imaging data, leading to limited biological interpretability and poor integration of external medical knowledge. T

Cited by 0SourcecodeScholar
2026

ImagiDrive: A Unified Imagination-And-Planning Framework for Autonomous Driving

ICRA 2026poster

Autonomous driving requires rich contextual comprehension and precise predictive reasoning to navigate dynamic and complex environments safely. Vision-Language Models (VLMs) and Driving World Models (DWMs) have independently emerged as powerful recipes addressing different aspects of this challenge.…

2026

Improved Adversarial Diffusion Compression for Real-World Video Super-Resolution

ICLR 2026poster

While many diffusion models have achieved impressive results in real-world video super-resolution (Real-VSR) by generating rich and realistic details, their reliance on multi-step sampling leads to slow inference. One-step networks like SeedVR2, DOVE, and DLoRAL alleviate this through condensing gen…

Cited by 0SourceScholar
2026

Leveraging Machine Unlearning for Cost-Efficient Preference Alignment

ICML 2026poster

Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human feedback face notable challenges. These approaches require high-quality datasets of positive preference examples, which are costly to obtain and computationally i…

Cited by 0SourceScholar
2026

MIDB: Multilingual Instruction Data Booster for Enhancing Cultural Equality in Multilingual Instruction Synthesis

AAAI 2026technical

Despite doubts on data quality, instruction synthesis has been widely applied into instruction tuning (IT) of LLMs as an economic and rapid alternative. Recent endeavors focus on improving data quality for synthesized instruction pairs in English and have facilitated IT of English-centric LLMs. Howe

Cited by 0SourcePDFScholar
2026

Omni-Attack: Adversarial Attacks on Open-Ended VQA in Black-Box Multimodal LLMs

CVPR 2026

Multimodal large language models (MLLMs) have achieved remarkable success across diverse applications, from autonomous driving to document understanding. As these models are deployed in safety-critical contexts, understanding their adversarial robustness becomes crucial. However, current evaluations

Cited by 0SourcecodeScholar
2026

Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving

AAAI 2026technical

End-to-end autonomous driving has achieved remarkable advancements in recent years. Existing methods primarily follow a perception–planning paradigm, where perception and planning are executed sequentially within a fully differentiable framework for planning-oriented optimization. We further advance

Cited by 0SourcePDFScholar
2026

Position: Creating High-Fidelity Synthetic Training Data Should Employ Multi-level Optimization

ICML 2026poster

The reliance of machine learning (ML) models on large-scale, high-quality labeled training data incurs significant challenges in specialized domains where such data is expensive and difficult to obtain. A promising solution is the automatic creation of synthetic training data. However, current appro…

Cited by 0SourceScholar
2026

Realtime Video Frame Interpolation using One-Step Diffusion Sampling

ICLR 2026poster

Recent research on video Frame Interpolation (VFI) shows that a pretrained Video Diffusion Model (VDM) can solve many challenging scenarios, including large or complex motion. However, VDMs require tedious diffusion sampling, making the inference slow. One possible way to accelerate is to distill a…

Cited by 0SourceScholar
2026

Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment

ICLR 2026oral

Reasoning-based image quality assessment (IQA) models trained through reinforcement learning (RL) exhibit exceptional generalization, yet the underlying mechanisms and critical factors driving this capability remain underexplored in current research. Moreover, despite their superior performance, the…

Cited by 0SourcecodeScholar
2026

Reasoning in Space via Grounding in the World

ICLR 2026poster

In this paper, we claim that 3D visual grounding is the cornerstone of spatial reasoning and introduce the $\textit{Grounded-Spatial Reasoner (GS-Reasoner)}$ to explore the effective spatial representations that bridge the gap between them. Existing 3D LLMs suffer from the absence of a unified 3D re…

Cited by 0SourcecodeScholar
2026

Relative Position Matters: Trajectory Prediction and Planning with Polar Representation

ICRA 2026poster

Trajectory prediction and planning in autonomous driving are highly challenging due to the complexity of predicting surrounding agents' movements and planning the ego agent's actions in dynamic environments. Existing methods encode map and agent positions and decode future trajectories in Cartesian …

2026

Robust Training of Neural Networks at Arbitrary Precision and Sparsity

ICLR 2026poster

The discontinuous operations inherent in quantization and sparsification introduce a long-standing obstacle to backpropagation, particularly in ultra-low precision and sparse regimes. While the community has long viewed quantization as unfriendly to gradient descent due to its lack of smoothness, we…

Cited by 0SourceScholar
2026

SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving

CVPR 2026

Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inherently trained as generalist models, lacking specialized understanding of driving-specific reasoning in 3D space and time.

Cited by 0SourcecodeScholar
2026

SIAM: Towards Generalizable Articulated Object Modeling via Single Robot-Object Interaction

AAAI 2026technical

Articulated object modeling, which represents interconnected rigid bodies with their geometry, part segmentation, articulation tree, and physical properties, is crucial for robotic perception and manipulation. Recently existing methods like SAGCI leverage Interactive Perception (IP) to refine models

Cited by 0SourcePDFScholar
2026

Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking

ICLR 2026poster

Improving vision-language models (VLMs) in the post-training stage typically relies on supervised fine-tuning or reinforcement learning, methods that necessitate costly, human-annotated data. While self-supervised techniques such as self-consistency have proven effective for enhancing reasoning cap…

Cited by 0SourceScholar
2026

Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution

AAAI 2026technical

Omnidirectional videos (ODVs) provide an immersive visual experience by capturing the 360° scene. With the rapid advancements in virtual/augmented reality, metaverse, and generative artificial intelligence, the demand for high-quality ODVs is surging. However, ODVs often suffer from low resolution d

Cited by 0SourcePDFScholar
2026

TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language Models

AAAI 2026technical

Efficient and lightweight adaptation of pre-trained Vision-Language Models (VLMs) to downstream tasks through collaborative interactions between local clients and a central server is a rapidly emerging research topic in federated learning. Existing adaptation algorithms are typically trained iterati

Cited by 0SourcePDFScholar
2026

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

CVPR 2026

Enhancing temporal understanding of MLLMs is essential for long-form video analysis, supporting tasks such as temporal localization and time-sensitive question answering. While reinforcement learning (RL) has been explored for temporal reasoning, existing approaches are often limited to specific tas

Cited by 0SourceScholar
2026

TokenTrace: Multi-Concept Attribution through Watermarked Token Recovery

CVPR 2026

Generative AI models pose a significant challenge to intellectual property (IP), as they can replicate unique artistic styles and concepts without attribution. While watermarking offers a potential solution, existing methods often fail in complex scenarios where multiple concepts (e.g., an object an

Cited by 0SourceScholar
2026

UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement

CVPR 2026

Image quality assessment (IQA) and image restoration are fundamental problems in low-level vision. Although IQA and restoration are closely connected conceptually, most existing work treats them in isolation. Recent advances in unified multimodal understanding-generation models demonstrate promising

Cited by 0SourcecodeScholar
2026

Uncertainty-Guided View-Strength-Aware Feature Utilization for Multi-View Classification

AAAI 2026technical

In multi-view classification tasks (MVC), each view provides an unique perspective on the data, offering complementary information that can improve classification performance when properly integrated. However, traditional methods typically adopt a uniform processing strategy for all views before fus

Cited by 0SourcePDFScholar
2026

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding

ICLR 2026poster

Despite the impressive progress on understanding and generating images shown by the recent unified architectures, the integration of 3D tasks remains challenging and largely unexplored. In this paper, we introduce UniUGG, the first unified understanding and generation framework for 3D modalities. Ou…

Cited by 0SourcecodeScholar
2026

VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning

AAAI 2026technical

Recent advances in AI-generated content (AIGC) have led to the emergence of powerful text-to-video generation models. Despite these successes, evaluating the quality of AIGC-generated videos remains challenging due to limited generalization, lack of temporal awareness, heavy reliance on large-scale

Cited by 0SourcePDFScholar
2026

VistaDepth: Improving Far-Range Depth Estimation With Spectral Modulation and Adaptive Reweighting

RA-L 2026

Monocular depth estimation infers per-pixel depth from a single RGB image. It remains particularly challenging in far-range regions, where sparse observations and long-tailed depth distributions bias learning toward near-range content. Diffusion models offer a promising alternative to discriminative

Cited by 0SourceScholar
2025

4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration

NeurIPS 2025poster

Leveraging diverse robotic data for pretraining remains a critical challenge. Existing methods typically model the dataset’s action distribution using simple observations as inputs. However, these inputs are often incomplete, resulting in a dispersed conditional action distribution—an issue we refer…

Cited by 0SourceScholar
2025

A Magnetically-Actuated Ultrasound Capsule Endoscope (MUSCE) for Endoluminal Imaging in Tubular Environments

RA-L 2025

Endoscopic ultrasound (EUS) has the ability to image tissue in and beyond the wall of the gastrointestinal (GI) tract, assisting in the early diagnosis of digestive diseases. However, traditional EUS based on flexible endoscopes could make the operation procedure traumatic and intolerable to patient

Cited by 9SourceScholar
2025

An Information-Theoretic Regularizer for Lossy Neural Image Compression

ICCV 2025poster

Lossy image compression networks aim to minimize the latent entropy of images while adhering to specific distortion constraints. However, optimizing the neural network can be challenging due to its nature of learning quantized latent representations. In this paper, our key finding is that minimizing…

Cited by 0SourcePDFScholar
2025

BezierGS: Dynamic Urban Scene Reconstruction with Bezier Curve Gaussian Splatting

ICCV 2025poster

The realistic reconstruction of street scenes is critical for developing real-world simulators in autonomous driving. Most existing methods rely on object pose annotations, using these poses to reconstruct dynamic objects and move them during the rendering process. This dependence on high-precision…

2025

Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising

ICASSP 2025accepted

Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power and extensive data, leading most video diffusion models to be limited to a small number of frames. Existing training-free…

Cited by 0SourceScholar
2025

CMFNThinker: A Novel Cross-source Multi-modal Fake News Detection Model

ICASSP 2025accepted

The rapid development of social media platforms has accelerated the generation and spread of fake news. News on different platforms varies significantly in content and audience. It makes most existing fake news detection models, which rely on single-source datasets, struggle to perform well on news…

Cited by 0SourceScholar
2025

Calibrating Large Language Models with Sample Consistency

AAAI 2025technical

Accurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive mod…

2025

Controllable Unlearning for Image-to-Image Generative Models via $\epsilon$-Constrained Optimization

ICLR 2025poster

While generative models have made significant advancements in recent years, they also raise concerns such as privacy breaches and biases. Machine unlearning has emerged as a viable solution, aiming to remove specific training data, e.g., containing private information and bias, from models. In this…

Cited by 1SourcePDFScholar
2025

Cross-Component Residual Prediction for Geometry-Based Point Cloud Compression

ICASSP 2025accepted

Point cloud compression is pivotal for the success of immersive multimedia applications. For attribute compression in geometry-based point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. Inspired by the significant impact of cross-component pr…

Cited by 0SourceScholar
2025

Data Interpreter: An LLM Agent for Data Science

ACL 2025finding

Large Language Model (LLM)-based agents have excelled in various domains but face significant challenges when applied to data science workflows due to their complex, multi-stage nature. Current LLM-based agents struggle with non-linear relationships, recursive dependencies, implicit data- and logic-…

2025

Diffusion$^2$: Dynamic 3D Content Generation via Score Composition of Video and Multi-view Diffusion Models

ICLR 2025poster

Recent advancements in 3D generation are predominantly propelled by improvements in 3D-aware image diffusion models. These models are pretrained on Internet-scale image data and fine-tuned on massive 3D data, offering the capability of producing highly consistent multi-view images. However, due to t…

2025

DroidCall: A Dataset for LLM-powered Android Intent Invocation

EMNLP 2025

The growing capabilities of large language models in natural language understanding significantly strengthen existing agentic systems. To power performant on-device mobile agents for better data privacy, we introduce DroidCall, the first training and testing dataset for accurate Android Intent invoc

2025

Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans?

CVPR 2025poster

To detect prohibited items in challenging categories, human inspectors typically rely on images from two distinct views (vertical and side). Can AI detect prohibited items from dual-view X-ray images in the same way humans do? Existing X-ray datasets often suffer from limitations, such as single-vie…

2025

DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation

ACL 2025finding

The rapid advancement of large language models (LLMs) has significantly improved their performance in code generation tasks. However, existing code benchmarks remain static, consisting of fixed datasets with predefined problems. This makes them vulnerable to memorization during training, where LLMs…

2025

ECVC: Exploiting Non-Local Correlations in Multiple Frames for Contextual Video Compression

CVPR 2025poster

In Learned Video Compression (LVC), improving inter prediction, such as enhancing temporal context mining and mitigating accumulated errors, is crucial for boosting rate-distortion performance. Existing LVCs mainly focus on mining the temporal movements while neglecting non-local correlations among…

2025

FedFACT: A Provable Framework for Controllable Group-Fairness Calibration in Federated Learning

NeurIPS 2025poster

With emerging application of Federated Learning (FL) in decision-making scenarios, it is imperative to regulate model fairness to prevent disparities across sensitive groups (e.g., female, male). Current research predominantly focuses on two concepts of group fairness within FL: *Global Fairness* (o…

Cited by 0SourceScholar
2025

FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise

ICLR 2025poster

Text-driven video generation has advanced significantly due to developments in diffusion models. Beyond the training and sampling phases, recent studies have investigated noise priors of diffusion models, as improved noise priors yield better generation results. One recent approach employs the Fouri…

Cited by 0SourcePDFScholar
2025

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

NeurIPS 2025poster

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations into models to improve spatial understanding, we aim to unlo…

Cited by 0SourceScholar
2025

Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution

NeurIPS 2025poster

End-to-end autonomous driving methods aim to directly map raw sensor inputs to future driving actions such as planned trajectories, bypassing traditional modular pipelines. While these approaches have shown promise, they often operate under a one-shot paradigm that relies heavily on the current scen…

Cited by 0SourcecodeScholar
2025

GS-LiDAR: Generating Realistic LiDAR Point Clouds with Panoramic Gaussian Splatting

ICLR 2025poster

LiDAR novel view synthesis (NVS) has emerged as a novel task within LiDAR simulation, offering valuable simulated point cloud data from novel viewpoints to aid in autonomous driving systems. However, existing LiDAR NVS methods typically rely on neural radiance fields (NeRF) as their 3D representatio…

2025

GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D Reconstruction

CVPR 2025poster

Garments are common in daily life and are important for embodied intelligence community. Current category-level garments pose tracking works focus on predicting point-wise canonical correspondence and learning a shape deformation in point cloud sequences. In this paper, motivated by the 2D warping s…

Cited by 0SourcePDFScholar
2025

Generalizable Articulated Object Perception with Superpoints

ICASSP 2025accepted

Manipulating articulated objects with robotic arms is challenging due to the complex kinematic structure, which requires precise part segmentation for efficient manipulation. In this work, we introduce a novel superpoint-based perception method designed to improve part segmentation in 3D point cloud…

Cited by 4SourceScholar
2025

Generalizable and Actionable Part Detection and Manipulation with SAM-rectified Segmentation and Iterative Pose Refinement

IROS 2025

The ability to perform cross-category object perception and manipulation is highly desirable in building intelligent robots. One promising approach is to define the concept of Generalizable and Actionable Parts (GAParts), such as buttons and handles, on both seen and unseen object categories. Howeve

Cited by 0SourceScholar
2025

Generative Adversarial Network with Structured Semantic Prompts Constrainting Clip for Text-to-Image

ICASSP 2025accepted

Rapidly synthesizing text-relevant images has long been a significant challenge. Introducing pre-trained models into GANs can significantly enhance model performance, enabling the fast generation of high-quality images. Previous work has primarily focused on improving visual quality, with limited fi…

Cited by 0SourceScholar
2025

Improving Gaussian Splatting with Localized Points Management

CVPR 2025highlight

Point management is critical for optimizing 3D Gaussian Splatting models, as point initiation (e.g., via structure from motion) is often distributionally inappropriate. Typically, Adaptive Density Control (ADC) algorithm is adopted, leveraging view-averaged gradient magnitude thresholding for point…

Cited by 0SourcePDFScholar
2025

Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognition

CVPR 2025highlight

Parameter-efficient fine-tuning (PEFT) has attracted significant attention due to the growth of pre-trained model sizes and the need to fine-tune (FT) them for superior downstream performance. Despite a surge in new PEFT methods, a systematic study to understand their performance and suitable applic…

2025

LoGoFair: Post-Processing for Local and Global Fairness in Federated Learning

AAAI 2025technical

Federated learning (FL) has garnered considerable interest for its capability to learn from decentralized data sources. Given the increasing application of FL in decision-making scenarios, addressing fairness issues across different sensitive groups (e.g., female, male) in FL is crucial. Current res…

2025

MPDG-SLAM: Motion Probability-Based 3DGS-SLAM in Dynamic Environment

IROS 2025

We present MPDG-SLAM, a novel 3D Gaussian point cloud rendering SLAM method based on Motion Probability (MP) for dynamic interference handling. Current 3DGSSLAM approaches for dynamic environments often rely on optical flow estimation masks. However, these deep learning-based optical flow models are

Cited by 1SourceScholar
2025

On the Stability of Graph Convolutional Neural Networks: A Probabilistic Perspective

NeurIPS 2025poster

Graph convolutional neural networks (GCNNs) have emerged as powerful tools for analyzing graph-structured data, achieving remarkable success across diverse applications. However, the theoretical understanding of the stability of these models, i.e., their sensitivity to small changes in the graph str…

Cited by 0SourceScholar
2025

Optimized Design and Calibration of a Human-Eye-Sized Active Binocular Vision System Based on Spherical Parallel Mechanism

RA-L 2025

The Active Binocular Vision System (ABVS), resembling the human eye, demonstrates potential for improving visual perception in robotic systems, especially in dynamic and complex environments. In this letter, we present an optimized design of a three degree-of-freedom (DoF) Active Monocular Vision Sy

Cited by 2SourceScholar
2025

Pre-defined Keypoints Promote Category-level Articulation Pose Estimation via Multi-Modal Alignment

IJCAI 2025

Articulations are essential in everyday interactions, yet traditional RGB-based pose estimation methods often struggle with issues such as lighting variations and shadows. To overcome these challenges, we propose a novel Pre-defined keypoint based framework for category-level articulation pose estim

Cited by 0SourcePDFScholar
2025

Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

NeurIPS 2025spotlight

Image quality assessment (IQA) focuses on the perceptual visual quality of images, playing a crucial role in downstream tasks such as image reconstruction, compression, and generation. The rapid advancement of multi-modal large language models (MLLMs) has significantly broadened the scope of IQA, mo…

Cited by 0SourcecodeScholar
2025

R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render Strategy

AAAI 2025technical

Human life is filled with articulated objects. Previous works for estimating the pose of category-level articulated objects rely on costly 3D point clouds or RGB-D images. In this paper, our goal is to estimate category-level articulation poses from a single RGB image, where we propose R2-Art, a nov…

Cited by 0SourcePDFScholar
2025

Reinforcement Learning-Based Microrobotic Swarm Navigation and Obstacle Avoidance in Partially Observable Environments

IROS 2025

Microrobotic swarms have shown promising features due to their collective and flexible behaviours, while achieving precise swarm control and autonomous navigation in complex environments remains a challenge. Here, we propose a Transformer-based reinforcement learning strategy that integrates Proxima

Cited by 0SourceScholar
2025

Rethinking Layered Graphic Design Generation with a Top-Down Approach

ICCV 2025poster

Graphic design is crucial for conveying ideas and messages. Designers usually organize their work into objects, backgrounds, and vectorized text layers to simplify editing. However, this workflow demands considerable expertise. With the rise of GenAI methods, an endless supply of high-quality graphi…

Cited by 0SourcePDFScholar
2025

ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents

ICLR 2025poster

Recent advancements in integrating large language models (LLMs) with application programming interfaces (APIs) have gained significant interest in both academia and industry. Recent work demonstrates that these API-based agents exhibit relatively strong autonomy and planning capabilities. However, t…

2025

Towards Robust Category-level Articulation Pose Estimation via Integrated Differentiable Rendering

ICASSP 2025accepted

Accurate object pose estimation is crucial for embodied intelligence tasks such as manipulation, grasping, and human-robot interaction. However, due to the inherent characteristics of articulated objects, such as kinematic constraints and self-occlusion, pose estimation for articulated objects has r…

Cited by 1SourceScholar
2025

Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization

NeurIPS 2025poster

Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, which is a crucial capability for tasks like visual storytelling and step-by-step visual reasoning. In this work, we prop…

Cited by 0SourceScholar
2025

Training Language Model to Critique for Better Refinement

ACL 2025finding

Large language models (LLMs) have demonstrated remarkable evaluation and critique capabilities, providing insightful feedback and identifying flaws in various tasks. However, limited research has explored which types of critiques are most effective for improving model responses or how to generate su…

2025

TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games

EMNLP 2025

This paper introduces TurnaboutLLM, a novel framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa. The framework tasks LLMs with identifying contradictions between

Cited by 0SourcePDFScholar
2025

UniMotion: A Unified Motion Framework for Simulation, Prediction and Planning

NeurIPS 2025poster

Motion simulation, prediction and planning are foundational tasks in autonomous driving, each essential for modeling and reasoning about dynamic traffic scenarios. While often addressed in isolation due to their differing objectives, such as generating diverse motion states or estimating optimal tra…

Cited by 0SourceScholar
2025

UniScene: Unified Occupancy-centric Driving Scene Generation

CVPR 2025poster

Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles…

2024

A Hybrid Approach for Cross-Modality Pose Estimation Between Image and Point Cloud

RA-L 2024

Cross-modality pose estimation/localization is a critical challenge for multi-sensor-based perception systems, with applications spanning vehicle localization and online calibrations. In this paper, we introduce a hybrid approach to estimate the camera pose with respect to a point cloud with co-visi

Cited by 1SourceScholar
2024

A Magnetic Continuum Robot with In-situ Magnetic Reprogramming Capability

ICRA 2024poster

Magnetic continuum robots (MCR) have shown great potential in minimally invasive interventions because they can be actively and remotely navigated through complex in vivo environments. However, the deformation capability of current MCRs is limited by fixed magnetization congurations, preventing them…

Cited by 2SourceScholar
2024

A Tri-Dynamic Preprocessing Framework for UGC Video Compression

ICASSP 2024accepted

In recent years, user generated content (UGC) has become the dominant force in internet traffic. However, UGC videos exhibit a higher degree of variability and diverse characteristics compared to traditional encoding test videos. This variance challenges the effectiveness of data-driven machine lear…

Cited by 0SourceScholar
2024

BLO-SAM: Bi-level Optimization Based Finetuning of the Segment Anything Model for Overfitting-Preventing Semantic Segmentation

ICML 2024poster

The Segment Anything Model (SAM), a foundation model pretrained on millions of images and segmentation masks, has significantly advanced semantic segmentation, a fundamental task in computer vision. Despite its strengths, SAM encounters two major challenges. Firstly, it struggles with segmenting spe…

2024

CURE4Rec: A Benchmark for Recommendation Unlearning with Deeper Influence

NeurIPS 2024poster

With increasing privacy concerns in artificial intelligence, regulations have mandated the right to be forgotten, granting individuals the right to withdraw their data from models. Machine unlearning has emerged as a potential solution to enable selective forgetting in models, particularly in recomm…

2024

CatmullRom Splines-Based Regression for Image Forgery Localization

AAAI 2024technical

IFL (Image Forgery Location) helps secure digital media forensics. However, many methods suffer from false detections (i.e., FPs) and inaccurate boundaries. In this paper, we proposed the CatmullRom Splines-based Regression Network (CSR-Net), which first rethinks the IFL task from the perspective of…

Cited by 14SourcePDFScholar
2024

Causal-IQA: Towards the Generalization of Image Quality Assessment Based on Causal Inference

ICML 2024poster

Due to the high cost of Image Quality Assessment (IQA) datasets, achieving robust generalization remains challenging for prevalent deep learning-based IQA methods. To address this, this paper proposes a novel end-to-end blind IQA method: Causal-IQA. Specifically, we first analyze the causal mechanis…

Cited by 4SourcePDFScholar
2024

Choice-75: A Dataset on Decision Branching in Script Learning

COLING 2024main

Script learning studies how daily events unfold. It enables machines to reason about narratives with implicit information. Previous works mainly consider a script as a linear sequence of events while ignoring the potential branches that arise due to people’s circumstantial choices. We hence propose…

2024

Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video

ICLR 2024poster

In this paper, we present Consistent4D, a novel approach for generating 4D dynamic objects from uncalibrated monocular videos. Uniquely, we cast the 360-degree dynamic object reconstruction as a 4D generation problem, eliminating the need for tedious multi-view data collection and camera calibration…

2024

DG-SLAM: Robust Dynamic Gaussian Splatting SLAM with Hybrid Pose Optimization

NeurIPS 2024poster

Achieving robust and precise pose estimation in dynamic scenes is a significant research challenge in Visual Simultaneous Localization and Mapping (SLAM). Recent advancements integrating Gaussian Splatting into SLAM systems have proven effective in creating high-quality renderings using explicit 3D…

Cited by 5SourcePDFScholar
2024

DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic States

NeurIPS 2024poster

Accurate motion forecasting for traffic agents is crucial for ensuring the safety and efficiency of autonomous driving systems in dynamically changing environments. Mainstream methods adopt a one-query-one-trajectory paradigm, where each query corresponds to a unique trajectory for predicting multi-…

2024

Effective Lymph Nodes Detection in CT Scans Using Location Debiased Query Selection and Contrastive Query Representation in Transformer

ECCV 2024poster

"Lymph node (LN) assessment is a critical yet very challenging task in the routine clinical workflow of radiology and oncology. Accurate LN analysis is essential for cancer diagnosis, staging and treatment planning. Finding scatteredly distributed, low-contrast clinically relevant LNs in 3D CT is di…

2024

EfficientCAPER: An End-to-End Framework for Fast and Robust Category-Level Articulated Object Pose Estimation

NeurIPS 2024poster

Human life is populated with articulated objects. Pose estimation for category-level articulated objects is a significant challenge due to their inherent complexity and diverse kinematic structures. Current methods for this task usually meet the problems of insufficient consideration of kinematic co…

Cited by 0SourcePDFScholar
2024

Enhancing High-Resolution 3D Generation through Pixel-wise Gradient Clipping

ICLR 2024poster

High-resolution 3D object generation remains a challenging task primarily due to the limited availability of comprehensive annotated training data. Recent advancements have aimed to overcome this constraint by harnessing image generative models, pretrained on extensive curated web datasets, using kn…

2024

FastGAT: Simple and Efficient Graph Attention Neural Network with Global-Aware Adaptive Computational Node Attention

ICASSP 2024accepted

Graph attention neural network (GAT) stands as a fundamental model within graph neural networks, extensively employed across various applications. It assigns different weights to different nodes for feature aggregation by comparing the similarity of features between nodes. However, as the amount and…

Cited by 0SourceScholar
2024

FrameQuant: Flexible Low-Bit Quantization for Transformers

ICML 2024poster

Transformers are the backbone of powerful foundation models for many Vision and Natural Language Processing tasks. But their compute and memory/storage footprint is large, and so, serving such models is expensive often requiring high-end hardware. To mitigate this difficulty, Post-Training Quantizat…

2024

Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts

NAACL 2024long

Pretrained Language Models (PLMs) have advanced Natural Language Processing (NLP) tasks significantly, but finetuning PLMs on low-resource datasets poses significant challenges such as instability and overfitting. Previous methods tackle these issues by finetuning a strategically chosen subnetwork o…

2024

LVC-LGMC: Joint Local and Global Motion Compensation for Learned Video Compression

ICASSP 2024accepted

Existing learned video compression models employ flow net or deformable convolutional networks (DCN) to estimate motion information. However, the limited receptive fields of flow net and DCN inherently direct their attentiveness towards the local contexts. Global contexts, such as large-scale motion…

Cited by 0SourceScholar
2024

LaneGraph2Seq: Lane Topology Extraction with Language Model via Vertex-Edge Encoding and Connectivity Enhancement

AAAI 2024technical

Understanding road structures is crucial for autonomous driving. Intricate road structures are often depicted using lane graphs, which include centerline curves and connections forming a Directed Acyclic Graph (DAG). Accurate extraction of lane graphs relies on precisely estimating vertex and edge i…

2024

Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model Disentanglement

NeurIPS 2024poster

A longstanding goal of artificial general intelligence is highly capable generalists that can learn from diverse experiences and generalize to unseen tasks. The language and vision communities have seen remarkable progress toward this trend by scaling up transformer-based models trained on massive d…

2024

MoE-DiffIR: Task-customized Diffusion Priors for Universal Compressed Image Restoration

ECCV 2024poster

"We present MoE-DiffIR, an innovative universal compressed image restoration (CIR) method with task-customized diffusion priors. This intends to handle two pivotal challenges in the existing CIR methods: (i) lacking adaptability and universality for different image codecs, , JPEG and WebP; (ii) poor…

2024

NeRF-LiDAR: Generating Realistic LiDAR Point Clouds with Neural Radiance Fields

AAAI 2024technical

Labelling LiDAR point clouds for training autonomous driving is extremely expensive and difficult. LiDAR simulation aims at generating realistic LiDAR data with labels for training and verifying self-driving algorithms more efficiently. Recently, Neural Radiance Fields (NeRF) have been proposed for…

2024

OpenOcc: Open Vocabulary 3D Scene Reconstruction via Occupancy Representation

IROS 2024poster

3D reconstruction has been widely used in autonomous navigation fields of mobile robotics. However, the former research can only provide the basic geometry structure without the capability of open-world scene understanding, limiting advanced tasks like human interaction and visual navigation. Moreov…

Cited by 2SourcecodeScholar
2024

Personalized Video Comment Generation

EMNLP 2024finding

Generating personalized responses, particularly in the context of video, poses a unique challenge for language models. This paper introduces the novel task of Personalized Video Comment Generation (PVCG), aiming to predict user comments tailored to both the input video and the user’s comment history…

2024

Private Learning with Public Features

AISTATS 2024poster

We study a class of private learning problems in which the data is a join of private and public features. This is often the case in private personalization tasks such as recommendation or ad prediction, in which features related to individuals are sensitive, while features related to items (the movi…

Cited by 7SourcePDFScholar
2024

Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting

ICLR 2024poster

Reconstructing dynamic 3D scenes from 2D images and generating diverse views over time is challenging due to scene complexity and temporal dynamics. Despite advancements in neural implicit models, limitations persist: (i) Inadequate Scene Structure: Existing methods struggle to reveal the spatial an…

2024

Region-Adaptive Video Sharpening Via Rate-Perception Optimization

ICASSP 2024accepted

Sharpening is a widely adopted video enhancement technique. However, uniform sharpening intensity ignores texture variations, degrading video quality. Sharpening also increases bitrate, and there’s a lack of techniques to optimally allocate these additional bits across diverse regions. Thus, this pa…

Cited by 0SourceScholar
2024

Rethinking 3D Convolution in $\ell_p$-norm Space

NeurIPS 2024spotlight

Convolution is a fundamental operation in the 3D backbone. However, under certain conditions, the feature extraction ability of traditional convolution methods may be weakened. In this paper, we introduce a new convolution method based on $\ell_p$-norm. For theoretical support, we prove the univer…

Cited by 9SourcePDFScholar
2024

RoDyn-SLAM: Robust Dynamic Dense RGB-D SLAM With Neural Radiance Fields

RA-L 2024

Leveraging neural implicit representation to conduct dense RGB-D SLAM has been studied in recent years. However, this approach relies on a static environment assumption and does not work robustly within a dynamic environment due to the inconsistent observation of geometry and photometry. To address

Cited by 53SourcecodeScholar
2024

STEntConv: Predicting Disagreement between Reddit Users with Stance Detection and a Signed Graph Convolutional Network

COLING 2024main

The rise of social media platforms has led to an increase in polarised online discussions, especially on political and socio-cultural topics such as elections and climate change. We propose a simple and entirely novel unsupervised method to better predict whether the authors of two posts agree or di…

2024

SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM

NeurIPS 2024poster

Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze video data. However, current Vid-LLMs struggle to simultaneously retain high-quality frame-level semantic information (i.e.…

Cited by 2SourcePDFScholar
2024

Tailoring with Targeted Precision: Edit-Based Agents for Open-Domain Procedure Customization

ACL 2024findings

How-to procedures, such as how to plant a garden, are now used by millions of users, but sometimes need customizing to meet a user’s specific needs, e.g., planting a garden without pesticides. Our goal is to measure and improve an LLM’s ability to perform such customization. Our approach is to test…

Cited by 0SourcePDFScholar
2024

TimeR4 : Time-aware Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering

EMNLP 2024main

Temporal Knowledge Graph Question Answering (TKGQA) aims to answer temporal questions using knowledge in Temporal Knowledge Graphs (TKGs). Previous works employ pre-trained TKG embeddings or graph neural networks to incorporate the knowledge of TKGs. However, these methods fail to fully understand t…

2024

TimeSiam: A Pre-Training Framework for Siamese Time-Series Modeling

ICML 2024poster

Time series pre-training has recently garnered wide attention for its potential to reduce labeling expenses and benefit various downstream tasks. Prior methods are mainly based on pre-training techniques well-acknowledged in vision or language, such as masked modeling and contrastive learning. Howev…

2024

UPDP: A Unified Progressive Depth Pruner for CNN and Vision Transformer

AAAI 2024technical

Traditional channel-wise pruning methods by reducing network channels struggle to effectively prune efficient CNN models with depth-wise convolutional layers and certain efficient modules, such as popular inverted residual blocks. Prior depth pruning methods by reducing network depths are not suitab…

Cited by 13SourcePDFScholar
2023

DQN-based on-line Path Planning Method for Automatic Navigation of Miniature Robots

ICRA 2023poster

Untethered magnetic microrobots with control-lable locomotion property and multiple functions have attracted lots of attention in recent years. Owing to the small scale, micro-robots with automatic navigation possess a promising perspec-tive for biomedical applications including precise delivery and…

Cited by 10SourceScholar
2023

Devil Is in the Queries: Advancing Mask Transformers for Real-World Medical Image Segmentation and Out-of-Distribution Localization

CVPR 2023highlight

Real-world medical image segmentation has tremendous long-tailed complexity of objects, among which tail conditions correlate with relatively rare diseases and are clinically significant. A trustworthy medical AI algorithm should demonstrate its effectiveness on tail conditions to avoid clinically d…

Cited by 28SourcePDFScholar
2023

Distance-Based Weight Transfer for Fine-Tuning From Near-Field to Far-Field Speaker Verification

ICASSP 2023accepted

The scarcity of labeled far-field speech is a constraint for training superior far-field speaker verification systems. In general, fine-tuning the model pre-trained on large-scale near- field speech through a small amount of far-field speech substantially outperforms training from scratch. However,…

Cited by 0SourceScholar
2023

FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks

CVPR 2023highlight

In the fashion domain, there exists a variety of vision-and-language (V+L) tasks, including cross-modal retrieval, text-guided image retrieval, multi-modal classification, and image captioning. They differ drastically in each individual input/output format and dataset size. It has been common to des…

2023

Multi-Task Differential Privacy Under Distribution Skew

ICML 2023poster

We study the problem of multi-task learning under user-level differential privacy, in which n users contribute data to m tasks, each involving a subset of users. One important aspect of the problem, that can significantly impact quality, is the distribution skew among tasks. Tasks that have much few…

Cited by 5SourcePDFScholar
2023

PARTNER: Level up the Polar Representation for LiDAR 3D Object Detection

ICCV 2023poster

Recently, polar-based representation has shown promising properties in perceptual tasks. In addition to Cartesian-based approaches, which separate point clouds unevenly, representing point clouds as polar grids has been recognized as an alternative due to (1) its advantage in robust performance unde…

Cited by 10PDFcodeScholar
2023

Panoramic Video Salient Object Detection with Ambisonic Audio Guidance

AAAI 2023technical

Video salient object detection (VSOD), as a fundamental computer vision problem, has been extensively discussed in the last decade. However, all existing works focus on addressing the VSOD problem in 2D scenarios. With the rapid development of VR devices, panoramic videos have been a promising alter…

Cited by 16SourcePDFScholar
2023

PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer

AAAI 2023technical

3D object detection in autonomous driving aims to reason “what” and “where” the objects of interest present in a 3D world. Following the conventional wisdom of previous 2D object detection, existing methods often adopt the canonical Cartesian coordinate system with perpendicular axis. However, we co…

2023

QuadMag: A Mobile-Coil System With Enhanced Magnetic Actuation Efficiency and Dexterity

ICRA 2023poster

Magnetic field is a favorable power source for actuation and control of micro-/nanorobots. To overcome the fast decay of magnetic field for large-workspace microrobotic actuation, mobile field source-based systems have been proposed. In this work, we report a new mobile-coil system, i.e., QuadMag. I…

Cited by 6SourceScholar
2023

SUIT: Learning Significance-Guided Information for 3D Temporal Detection

IROS 2023poster

3D object detection from LiDAR point cloud is of critical importance for autonomous driving and robotics. While sequential point cloud has the potential to enhance 3D perception through temporal information, utilizing these temporal features effectively and efficiently remains a challenging problem.…

Cited by 3SourceScholar
2023

Safety Verification of Nonlinear Systems with Bayesian Neural Network Controllers

AAAI 2023technical

Bayesian neural networks (BNNs) retain NN structures with a probability distribution placed over their weights. With the introduced uncertainties and redundancies, BNNs are proper choices of robust controllers for safety-critical control systems. This paper considers the problem of verifying the saf…

2023

SeaFormer: Squeeze-enhanced Axial Transformer for Mobile Semantic Segmentation

ICLR 2023poster

Since the introduction of Vision Transformers, the landscape of many computer vision tasks (e.g., semantic segmentation), which has been overwhelmingly dominated by CNNs, recently has significantly revolutionized. However, the computational cost and memory requirement render these methods unsuitable…

Cited by 189SourcePDFScholar
2023

Self-Asymmetric Invertible Network for Compression-Aware Image Rescaling

AAAI 2023technical

High-resolution (HR) images are usually downscaled to low-resolution (LR) ones for better display and afterward upscaled back to the original size to recover details. Recent work in image rescaling formulates downscaling and upscaling as a unified task and learns a bijective mapping between HR and L…

2023

SimMTM: A Simple Pre-Training Framework for Masked Time-Series Modeling

NeurIPS 2023spotlight

Time series analysis is widely used in extensive areas. Recently, to reduce labeling expenses and benefit various tasks, self-supervised pre-training has attracted immense interest. One mainstream paradigm is masked modeling, which successfully pre-trains deep models by learning to reconstruct the m…

2023

Translating Images to Road Network: A Non-Autoregressive Sequence-to-Sequence Approach

ICCV 2023oral

The extraction of road network is essential for the generation of high-definition maps since it enables the precise localization of road landmarks and their interconnections. However, generating road network poses a significant challenge due to the conflicting underlying combination of Euclidean (e.…

Cited by 8PDFScholar
2022

Accelerating Score-Based Generative Models with Preconditioned Diffusion Sampling

ECCV 2022poster

"Score-based generative models (SGMs) have recently emerged as a promising class of generative models. However, a fundamental limitation is that their inference is very slow due to a need for many (e.g., 2000) iterations of sequential computations. An intuitive acceleration method is to reduce the s…

2022

CelebV-HQ: A Large-Scale Video Facial Attributes Dataset

ECCV 2022poster

"Large-scale datasets played an indispensable role in the recent success of face generation/editing and significantly facilitate the advances of emerging research fields. However, the academic community still lacks a video dataset with diverse facial attribute annotations, which is crucial for face-…

2022

DeepInteraction: 3D Object Detection via Modality Interaction

NeurIPS 2022accept

Existing top-performance 3D object detectors typically rely on the multi-modal fusion strategy. This design is however fundamentally restricted due to overlooking the modality-specific useful information and finally hampering the model performance. To address this limitation, in this work we introdu…

2022

FashionViL: Fashion-Focused Vision-and-Language Representation Learning

ECCV 2022poster

"Large-scale Vision-and-Language (V+L) pre-training for representation learning has proven to be effective in boosting various downstream V+L tasks. However, when it comes to the fashion domain, existing V+L methods are inadequate as they overlook the unique characteristics of both fashion V+L data…

2022

Hierarchical Road Topology Learning for Urban Mapless Driving

IROS 2022poster

The majority of current approaches in autonomous driving rely on High-Definition (HD) maps which detail the road geometry and surrounding area. Yet, this reliance is one of the obstacles to mass deployment of autonomous vehicles due to poor scalability of such prior maps. In this paper, we tackle th…

Cited by 14SourceScholar
2022

Is “My Favorite New Movie” My Favorite Movie? Probing the Understanding of Recursive Noun Phrases

NAACL 2022long

Recursive noun phrases (NPs) have interesting semantic properties. For example, “my favorite new movie” is not necessarily my favorite movie, whereas “my new favorite movie” is. This is common sense to humans, yet it is unknown whether language models have such knowledge. We introduce the Recursive…

2022

Learning Ego 3D Representation As Ray Tracing

ECCV 2022poster

"A self-driving perception model aims to extract 3D semantic representations from multiple cameras collectively into the bird’s-eye-view (BEV) coordinate frame of the ego car in order to ground downstream planner. Existing perception methods often rely on error-prone depth estimation of the whole sc…

2022

Learning from Mistakes – a Framework for Neural Architecture Search

AAAI 2022technical

Learning from one's mistakes is an effective human learning technique where the learners focus more on the topics where mistakes were made, so as to deepen their understanding. In this paper, we investigate if this human learning strategy can be applied in machine learning. We propose a novel machin…

Cited by 14SourcePDFScholar
2022

Magnetic Micro-Driller System for Nasolacrimal Duct Recanalization

RA-L 2022

Primary acquired nasolacrimal duct obstruction (PANDO) is the commonest cause of obstructive tearing and can lead to infections including dacryocystitis, cellulitis and postoperative endophthalmitis. External or endoscopic dacryocystorhinostomy (DCR) is the current standard treatment of PANDO. Howev

Cited by 8SourceScholar
2022

ONCE-3DLanes: Building Monocular 3D Lane Detection

CVPR 2022poster

We present ONCE-3DLanes, a real-world autonomous driving dataset with lane layout annotation in 3D space. Conventional 2D lane detection from a monocular image yields poor performance of following planning and control tasks in autonomous driving due to the case of uneven road. Predicting the 3D lane…

Cited by 76PDFcodeScholar
2022

RCLane: Relay Chain Prediction for Lane Detection

ECCV 2022poster

"Lane detection is an important component of many real-world autonomous systems. Despite a wide variety of lane detection approaches have been proposed, reporting steady benchmark improvements over time, lane detection remains a largely unsolved problem. This is because most of the existing lane det…

Cited by 32SourcePDFScholar
2022

Real-Time Navigation of an Untethered Miniature Robot Using Mobile Ultrasound Imaging and Magnetic Actuation Systems

RA-L 2022

Remotely actuated small-scale robots have shown promising capabilities in targeted delivery, micromanipulation, and biosensing. However, challenges remain in imaging and control of a robot in a large workspace, especially when conducting targeted navigation in hard-to-reach and tortuous regions of a

Cited by 21SourceScholar
2022

Region-Aware Metric Learning for Open World Semantic Segmentation via Meta-Channel Aggregation

IJCAI 2022poster

As one of the most challenging and practical segmentation tasks, open-world semantic segmentation requires the model to segment the anomaly regions in the images and incrementally learn to segment out-of-distribution (OOD) objects, especially under a few-shot condition. The current state-of-the-art…

2022

SGM3D: Stereo Guided Monocular 3D Object Detection

RA-L 2022

Monocular 3D object detection aims to predict the object location, dimension and orientation in 3D space alongside the object category given only a monocular image. It poses a great challenge due to its ill-posed property, which is a critical lack of depth information in the 2D image plane. While ex

Cited by 39SourcecodeScholar
2022

Show Me More Details: Discovering Hierarchies of Procedures from Semi-structured Web Data

ACL 2022long

Procedures are inherently hierarchical. To “make videos”, one may need to “purchase a camera”, which in turn may require one to “set a budget”. While such hierarchical knowledge is critical for reasoning about complex procedures, most existing work has treated procedures as shallow structures withou…

2022

Torque-Actuated Multimodal Locomotion of Ferrofluid Robot With Environment and Task Adaptability

IROS 2022poster

Soft microrobotics have recently been an active field that advances microrobotics with new robot design, locomotion, and applications. In this paper, we study the ferrofluid robot (FR), which has soft nature and exhibits paramagnetism. Currently, the FR locomotion is usually realized by magnetic for…

Cited by 2SourceScholar
2022

Unsupervised Entity Linking with Guided Summarization and Multiple-Choice Selection

EMNLP 2022main

Entity linking, the task of linking potentially ambiguous mentions in texts to corresponding knowledge-base entities, is an important component for language understanding. We address two challenge in entity linking: how to leverage wider contexts surrounding a mention, and how to deal with limited t…

Cited by 7SourcePDFScholar
2021

Boundary-Sensitive Pre-Training for Temporal Localization in Videos

ICCV 2021poster

Many video analysis tasks require temporal localization for the detection of content changes. However, most existing models developed for these tasks are pre-trained on general video action classification tasks. This is due to large scale annotation of temporal boundaries in untrimmed videos being e…

Cited by 76PDFcodeScholar
2021

Closed-Loop Control of a Helmholtz Coil System for Accurate Actuation of Magnetic Microrobot Swarms

RA-L 2021

Magnetic microrobot swarms have attracted lots of research interests in the robotics field. Since the remote actuation of swarm can be affected by the accuracy of the external magnetic field, electro-magnetic systems capable of generating precise magnetic fields have great value. In this work, we pr

Cited by 29SourceScholar
2021

DAST: Unsupervised Domain Adaptation in Semantic Segmentation Based on Discriminator Attention and Self-Training

AAAI 2021technical

Unsupervised domain adaption has recently been used to reduce the domain shift, which would ultimately improve the performance of the semantic segmentation on unlabeled real-world data. In this paper, we follow the trend to propose a novel method to reduce the domain shift using strategies of discri…

2021

Delving into Data: Effectively Substitute Training for Black-box Attack

CVPR 2021poster

Deep models have shown their vulnerability when processing adversarial samples. As for the black-box attack, without access to the architecture and weights of the attacked model, training a substitute model for adversarial attacks has attracted wide attention. Previous substitute training approaches…

Cited by 90PDFScholar
2021

Depth-Conditioned Dynamic Message Propagation for Monocular 3D Object Detection

CVPR 2021poster

The objective of this paper is to learn context- and depth-aware feature representation to solve the problem of monocular 3D object detection. We make following contributions: (i) rather than appealing to the complicated pseudo-LiDAR based approach, we propose a depth-conditioned dynamic message pro…

Cited by 156PDFcodeScholar
2021

Hybrid Magnetic Force and Torque Actuation of Miniature Helical Robots Using Mobile Coils to Accelerate Blood Clot Removal

IROS 2021poster

Mechanical rubbing of blood clot using miniature magnetic helical robots is a potential way for thrombolysis. In this paper, we report a new strategy for this issue based on mobile coils. Previously, we proposed the concept of magnetic actuation with parallel mobile coils, in which multiple coils ca…

Cited by 2SourceScholar
2021

Learning Dynamic Alignment via Meta-Filter for Few-Shot Learning

CVPR 2021poster

Few-shot learning (FSL), which aims to recognise new classes by adapting the learned knowledge with extremely limited few-shot (support) examples, remains an important open problem in computer vision. Most of the existing methods for feature alignment in few-shot learning only consider image-level o…

Cited by 150PDFScholar
2021

Learning a Few-shot Embedding Model with Contrastive Learning

AAAI 2021technical

Few-shot learning (FSL) aims to recognize target classes by adapting the prior knowledge learned from source classes. Such knowledge usually resides in a deep embedding model for a general matching purpose of the support and query image pairs. The objective of this paper is to repurpose the contrast…

Cited by 217SourcePDFScholar
2021

Magnetic Control of a Steerable Guidewire Under Ultrasound Guidance Using Mobile Electromagnets

RA-L 2021

Endovascular surgery has become a popular minimally invasive approach to diagnose and treat various vascular diseases. However, manipulating conventional passive guidewires and catheters still has technical challenges, such as long duration and undesired trauma. In addition, radiation exposure induc

Cited by 61SourceScholar
2021

MoViNets: Mobile Video Networks for Efficient Video Recognition

CVPR 2021poster

We present Mobile Video Networks (MoViNets), a family of computation and memory efficient video networks that can operate on streaming video for online inference. 3D convolutional neural networks (CNNs) are accurate at video recognition but require large computation and memory budgets and do not sup…

Cited by 324PDFcodeScholar
2021

On-Demand Assembly and Disassembly of a 3D Swimming Magnetic Mini-Propeller With Two Modules

RA-L 2021

Various morphologies and motions have been investigated extensively for robots on substrates or interfaces. However, the locomotion of modular robots in 3D free spaces remains challenging owing to zero dependence on boundaries, in particular at small scales. Here we propose simply modularized miniat

Cited by 4SourceScholar
2021

Parallel Actuation of Nanorod Swarm and Nanoparticle Swarm to Different Targets

ICRA 2021poster

After years of development, various swarms of robots have been proposed for many complicated tasks, such as forming patterns, cooperative locomotion, and adapting to different environments. However, controlling microrobotic swarms is still a challenging task owing to the lacking of integrated device…

Cited by 1SourceScholar
2021

Partial-Label and Structure-constrained Deep Coupled Factorization Network

AAAI 2021technical

In this paper, we technically propose an enriched prior guided framework, called Dual-constrained Deep Semi-Supervised Coupled Factorization Network (DS2CF-Net), for discovering hierarchical coupled data representation. To extract hidden deep features, DS2CF-Net is formulated as a partial-label and…

Cited by 6SourcePDFScholar
2021

Private Alternating Least Squares: Practical Private Matrix Completion with Tighter Rates

ICML 2021oral

We study the problem of differentially private (DP) matrix completion under user-level privacy. We design a joint differentially private variant of the popular Alternating-Least-Squares (ALS) method that achieves: i) (nearly) optimal sample complexity for matrix completion (in terms of number of ite…

Cited by 23SourcePDFScholar
2021

Progressive Coordinate Transforms for Monocular 3D Object Detection

NeurIPS 2021poster

Recognizing and localizing objects in the 3D space is a crucial ability for an AI agent to perceive its surrounding environment. While significant progress has been achieved with expensive LiDAR point clouds, it poses a great challenge for 3D object detection given only a monocular image. While ther…

2021

Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With Transformers

CVPR 2021poster

Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution and learns more abstract/semantic visual concepts with larger receptive fields. Since context modeling is critical for se…

Cited by 4009PDFcodeScholar
2021

Robust and Accurate Object Detection via Adversarial Learning

CVPR 2021poster

Data augmentation has become a de facto component for training high-performance deep image classifiers, but its potential is under-explored for object detection. Noting that most state-of-the-art object detectors benefit from fine-tuning a pre-trained classifier, we first study how the classifiers'…

Cited by 100PDFcodeScholar
2021

SOFT: Softmax-free Transformer with Linear Complexity

NeurIPS 2021spotlight

Vision transformers (ViTs) have pushed the state-of-the-art for various visual recognition tasks by patch-wise image tokenization followed by self-attention. However, the employment of self-attention modules results in a quadratic complexity in both computation and memory usage. Various attempts on…

Cited by 198SourcePDFScholar
2021

Simpler Is Better: Few-Shot Semantic Segmentation With Classifier Weight Transformer

ICCV 2021poster

A few-shot semantic segmentation model is typically composed of a CNN encoder, a CNN decoder and a simple classifier (separating foreground and background pixels). Most existing methods meta-learn all three model components for fast adaptation to a new class. However, given that as few as a single s…

Cited by 229PDFcodeScholar
2021

Simultaneous Actuation and Localization of Magnetic Robots Using Mobile Coils and Eye-In-Hand Hall-Effect Sensors

IROS 2021poster

Large workspace localization of magnetic robots is important for medical applications. This paper presents a novel localization strategy to achieve simultaneous localization and actuation of magnetic robots using hall-effect sensors. We integrate 25 sensors into a sensing probe and mount it on to th…

Cited by 7SourceScholar
2021

The Devil Is in the Task: Exploiting Reciprocal Appearance-Localization Features for Monocular 3D Object Detection

ICCV 2021poster

Low-cost monocular 3D object detection plays a fundamental role in autonomous driving, whereas its accuracy is still far from satisfactory. Our objective is to dig into the 3D object detection task and reformulate it as the sub-tasks of object localization and appearance perception, which benefits t…

Cited by 58PDFScholar
2021

Ultrasound Doppler Imaging and Navigation of Collective Magnetic Cell Microrobots in Blood

ICRA 2021poster

We propose ultrasound Doppler imaging and magnetic navigation of collective cell microrobots in whole blood. Cell microrobots are cultured using stem cells and iron microparticles, they have spheroidal structures and can be actuated under external magnetic fields. A collective of cell microrobots ca…

Cited by 3SourceScholar