← Search

CHEN CHEN

319 accepted papers

2026

A Novel Human-Machine Dual-Task Gaming Framework for Visual-Attention Training

ICRA 2026poster

Efficient brain functional training with rehabilitation robots has been an important and challenging topic in the human-machine interaction (HMI) field. Adjusting the interaction and gaming behaviors between human and machine to effectively activate the brain’s functional behavior is still a substan…

Cited by 0Scholar
2026

A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering

ICLR 2026poster

Effectively applying Vision-Language Models (VLMs) to Video Question Answering (VideoQA) hinges on selecting a concise yet comprehensive set of frames, as processing entire videos is computationally infeasible. However, current frame selection methods face a critical trade-off: approaches relying on…

Cited by 0SourceScholar
2026

ASTRAEA: A Token-wise Acceleration Framework for Video Diffusion Transformers

ICLR 2026poster

Video diffusion transformers (vDiTs) have made tremendous progress in text-to-video generation, but their high computational demands pose a major challenge for practical deployment. While existing studies propose acceleration methods to reduce workload at various granularities, they often rely on he…

Cited by 0SourceScholar
2026

An Intention-Aware Robust Safety Framework for Robot Teleoperation: Unifying Object Interaction and Obstacle Avoidance

ICRA 2026poster

Control barrier functions (CBFs) have proven to be effective for obstacle avoidance in robot teleoperation systems. However, for classical CBF, model uncertainties and external disturbances can significantly degrade the robustness of safety control. Moreover, the fixed safety boundary lacks adaptabi…

Cited by 0SourceScholar
2026

Attributing Response to Context: A Jensen–Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation

ICLR 2026poster

Retrieval-Augmented Generation (RAG) leverages large language models (LLMs) combined with external contexts to enhance the accuracy and reliability of generated responses. However, reliably attributing generated content to specific context segments, context attribution, remains challenging due to th…

Cited by 0SourceScholar
2026

Bridging Dynamics and Data: A Unified Diffusion Framework for Mechanistically-Informed Epidemic Forecasting

ICML 2026poster

Reliable epidemic forecasting is critical for public health decision-making yet remains challenging due to data sparsity and the non-stationary nature of disease dynamics. While recent hybrid models attempt to integrate mechanistic principles with data-driven approaches, they often relegate mechanis…

Cited by 0SourceScholar
2026

CMAR-Search: Commonsense and Memory Augmented Reasoning for Object Search in Dynamic Interactive Environments

ICRA 2026poster

Dynamic interactive object search in large-scale human environments presents substantial challenges for existing methods. Current scene representations like 3D Scene Graphs (3DSG) only provide coarse-grained spatial segmentation and cannot identify functional areas such as storage or leisure areas. …

Cited by 0Scholar
2026

Complex Instruction Following with Diverse Style Policies in Football Games

AAAI 2026technical

Despite advancements in language-controlled reinforcement learning (LC-RL) for basic domains and straightforward commands (e.g., object manipulation and navigation), effectively extending LC-RL to comprehend and execute high-level or abstract instructions in complex, multi-agent environments, such a

Cited by 0SourcePDFScholar
2026

Design, Modeling and Direction Control of a Wire-Driven Robotic Fish Based on a 2-DoF Crank–Slider Mechanism

ICRA 2026poster

Robotic fish have attracted growing attention in recent years owing to their biomimetic design and potential applications in environmental monitoring and biological surveys. Among robotic fish employing the Body–Caudal Fin (BCF) locomotion pattern, motor-driven actuation is widely adopted. Some appr…

2026

Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking

ICML 2026poster

Vision–language models (VLMs) now support both direct Instruct and explicit-reasoning Thinking modes, but practitioners lack principled ways to decide when reasoning helps or how much computation to allocate at test time. We investigate whether VLMs encode meta-cognitive signals for adaptive inferen…

Cited by 0SourceScholar
2026

Does a Hybrid Space-Aware Randomized Defense Improve Empirical and Certified Adversarial Robustness?

ICML 2026poster

We introduce Hybrid Space-aware Stochastic Convolution Attention Noise (HySCAN), a hybrid randomized defense that helps close the long-standing gap between provable robustness under ℓ2 certificates and empirical robustness against strong ℓ∞ attacks, while maintaining strong generalization across div…

Cited by 0SourceScholar
2026

DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models

CVPR 2026

Current story visualization methods tend to position subjects solely by text and face challenges in maintaining artistic consistency. To address these limitations, we introduce DreamingComics, a layout-aware story visualization framework. We build upon a pretrained video diffusion-transformer (DiT)

Cited by 0SourceScholar
2026

EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer

AAAI 2026technical

Most existing spatial reasoning benchmarks focus on static or globally observable environments, failing to capture the challenges of long-horizon reasoning and memory utilization under partial observability and dynamic changes. We introduce two dynamic spatial benchmarks—locally observable maze navi

Cited by 0SourcePDFScholar
2026

FG-HOCBF: Safe Operation Area Extension and Obstacle Avoidance Direction Guidance for Surface Detection in Narrow Environments

ICRA 2026poster

High-order control barrier functions (HOCBFs) that can achieve strict safety guarantees are widely used in robot safety control. However, robot obstacle avoidance in narrow environments with curved surfaces, as represented by aircraft blade detection, is still a challenge. Considering the narrow spa…

Cited by 0Scholar
2026

FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models

ICLR 2026poster

Bidirectional language models (LMs) consistently show stronger context understanding than unidirectional models, yet the theoretical reason remains unclear. We present a simple information bottleneck (IB) perspective: bidirectional representations preserve more mutual information (MI) about both the…

Cited by 0SourcecodeScholar
2026

From Swept Contact to Pose: Probe-Aware Registration Via Complementary-Shape Docking

ICRA 2026poster

Accurate registration between a prior model and the real scene is essential for high-precision robotic manipulation, yet optical methods suffer from long calibration chains, line-of-sight constraints, and fabrication errors. We propose a calibration-free alternative that reformulates contact registr…

2026

GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization Geometry

ICML 2026poster

Targeted data selection has emerged as a crucial paradigm for efficient instruction tuning, aiming to identify a small yet influential subset of training examples for a specific target task. In practice, influence is often measured through the effect of an example on parameter updates. To make selec…

Cited by 0SourceScholar
2026

GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents

AAAI 2026technical

Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot visual inputs and passive perception, lacking the ability to adaptively determine when, whether, and how to observe the int

Cited by 0SourcePDFScholar
2026

Geo2: Geometry-Guided Cross-view Geo-Localization and Image Synthesis

CVPR 2026

Cross-view geo-spatial learning consists of two important tasks: Cross-View Geo-Localization (CVGL) and Cross-View Image Synthesis (CVIS), both of which rely on establishing geometric correspondences between ground and aerial views. Recent Geometric Foundation Models (GFMs) have demonstrated strong

Cited by 0SourceScholar
2026

GeoFlow: Real-Time Fine-Grained Cross-View Geolocalization via Iterative Flow Prediction

CVPR 2026

Accurate and fast localization is vital for safe autonomous navigation in GPS-denied areas. Fine-Grained Cross-View Geolocalization (FG-CVG) aims to estimate the precise 2-Degree-of-Freedom (2-DoF) location of a ground image relative to a satellite image. However, current methods force a difficult t

Cited by 0SourcecodeScholar
2026

HGATSolver: A Heterogeneous Graph Attention Solver for Fluid–Structure Interaction

AAAI 2026technical

Fluid–structure interaction (FSI) systems involve distinct physical domains, fluid and solid, governed by different partial differential equations and coupled at a dynamic interface. While learning-based solvers offer a promising alternative to costly numerical simulations, existing methods struggle

Cited by 0SourcePDFScholar
2026

Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization

ICML 2026poster

Multimodal hallucination remains a persistent challenge for Vision-Language Models (VLMs). Standard textual Direct Preference Optimization (DPO) often fails to mitigate it due to a lack of explicit visual supervision. While existing works introduce visual preference DPO by contrasting original image…

Cited by 0SourceScholar
2026

Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization

ICLR 2026poster

Human visual preferences are inherently multi-dimensional, encompassing aspects of aesthetics, detail fidelity, and semantic alignment. However, existing open-source preference datasets provide only single, holistic annotations, resulting in severe label noise—images that excel in some dimensions (e…

Cited by 0SourceScholar
2026

LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning

ICML 2026spotlight

MoE-PEFT methods combine Mixture of Experts with parameter-efficient fine-tuning for multi-task adaptation, but require separate adapters per expert—causing trainable parameters to scale linearly with expert count and limiting applicability to adapter-based architectures. We propose LiME (Lightweigh…

Cited by 0SourceScholar
2026

LiveGesture: Streamable Co-Speech Gesture Generation Model

CVPR 2026

We propose LiveGesture, the first fully streamable, speech-driven full-body gesture generation framework that operates with zero look-ahead and supports arbitrary sequence length. Unlike existing co-speech gesture methods--which are designed for offline generation and either treat body regions indep

Cited by 0SourceScholar
2026

LogCD: Local-to-global Consistency Distillation for Few-step Image Generation

CVPR 2026

Distilling latent diffusion models (LDMs)/rectified flow models (RFMs) into ones that are fast to sample from conditions is attracting huge interest. However, the majority of existing methods either need significant training resources or lead to quality degradation, especially in text-image alignmen

Cited by 0SourceScholar
2026

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

ICLR 2026poster

Unified multimodal Large Language Models (LLMs) that can both understand and generate visual content hold immense potential. However, existing open-source models often suffer from a performance trade-off between these capabilities. We present Manzano, a simple and scalable unified framework that sub…

Cited by 0SourceScholar
2026

Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers

CVPR 2026

Diffusion Transformers (DiTs) have shown exceptional performance in image generation, yet their large parameter counts incur high computational costs, impeding deployment in resource-constrained settings. To address this, we propose Pluggable Pruning with Contiguous Layer Distillation (PPCL), a flex

Cited by 0SourcecodeScholar
2026

Position: Video LLMs Must Not Ignore the Pixel Dynamics in Plain Sight

ICML 2026poster

The essence of video lies in pixel dynamics: motion, state transitions, and the flow of visual information across frames. Video Large Language Models (LLMs) have rapidly become the dominant paradigm for video understanding in computer vision, sophisticated multimodal reasoning over complex, long-for…

Cited by 0SourceScholar
2026

PsyPARSE: Retrieval-Augmented Slow Thinking for Personalized Empathetic Counseling

AAAI 2026technical

The escalating global demand for mental health services highlights the potential of Large Language Models (LLMs) in psychological counseling. However, current LLM-based approaches, particularly fine-tuned models, are constrained by data distribution biases, leading to limited therapeutic diversity a

Cited by 0SourcePDFScholar
2026

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling

ICLR 2026poster

Test-time scaling has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs) by allocating additional computational resources during inference. However, this paradigm is inherently inefficient due to the generation of redundant and repetitive reasonin…

Cited by 0SourcecodeScholar
2026

SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models

ICML 2026poster

Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and challenging frontier. Current video benchmarks predominantly target general scenarios where perception/recognition is heav…

Cited by 0SourceScholar
2026

SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks

ICML 2026poster

This paper presents a theoretical framework that explains why fine-tuning small, randomly selected subnetworks (slices) within pre-trained models is sufficient for downstream adaptation. We prove that pretrained networks exhibit a universal winning slice property, arising from two phenomena: (1) spe…

Cited by 0SourceScholar
2026

Spatial Retrieval Augmented Autonomous Driving

CVPR 2026

Existing autonomous driving systems rely on onboard sensors (cameras, LiDAR, IMU, etc) for environmental perception. However, this paradigm is limited by the drive-time perception horizon and often fails under limited view scope, occlusion or extreme conditions such as darkness and rain. In contrast

Cited by 0SourcecodeScholar
2026

Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression

ICML 2026poster

The deployment of Large Language Models is constrained by the memory and bandwidth demands of static weights and dynamic Key-Value cache. SVD-based compression provides a hardware-friendly solution to reduce these costs. However, existing methods suffer from two key limitations: some are suboptimal …

Cited by 0SourceScholar
2026

Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting Code

ICML 2026poster

Multimodal geometry reasoning requires models to jointly understand visual diagrams and perform structured symbolic inference, yet current vision--language models struggle with complex geometric constructions due to limited training data and weak visual--symbolic alignment. We propose a pipeline for…

Cited by 0SourceScholar
2026

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

ICML 2026poster

During conversational interactions, humans subconsciously engage in concurrent thinking while listening to a speaker. Although this internal cognitive processing may not always manifest as explicit linguistic structures, it is instrumental in formulating high-quality responses. Inspired by this cogn…

Cited by 0SourceScholar
2026

Uncovering Latent Communication Patterns in Brain Networks via Adaptive Flow Routing

ICML 2026poster

Unraveling how macroscopic cognitive phenotypes emerge from microscopic neuronal connectivity remains one of the core pursuits of neuroscience. To this end, researchers typically leverage multi-modal information from structural connectivity (SC) and functional connectivity (FC) to complete downstrea…

Cited by 0SourceScholar
2026

UniCompress: Token Compression for Unified Vision-Language Understanding and Generation

CVPR 2026

Unified models aim to support both understanding and generation by encoding images into discrete tokens and processing them alongside text within a single autoregressive framework. This unified design offers architectural simplicity and cross-modal synergy, which facilitates shared parameterization,

Cited by 0SourceScholar
2026

UniSTFormer: Unified Spatio-Temporal Lightweight Transformer for Efficient Skeleton-Based Action Recognition

ICASSP 2026oral

Skeleton-based action recognition (SAR) has achieved impressive progress with transformer architectures. However, existing methods often rely on complex module compositions and heavy designs, leading to increased parameter counts, high computational costs, and limited scalability. In this paper, we…

Cited by 0SourcePDFScholar
2026

X2Edit: Revisiting Arbitrary-Instruction Image Editing Through Self-Constructed Data and Task-Aware Representation Learning

AAAI 2026technical

Existing open-source datasets for arbitrary-instruction image editing remain suboptimal, while a plug-and-play editing module compatible with community-prevalent generative models is notably absent. In this paper, we first introduce the X2Edit Dataset, a comprehensive dataset covering 14 diverse edi

Cited by 0SourcePDFScholar
2026

iGRPO: Fast Online RL for Flow Matching Model with Dense Reward

ICML 2026poster

Conventional practice assumes that online reinforcement learning for flow-matching models requires sampling full denoising trajectories to compute rewards. This assumption underlies methods such as Group Relative Policy Optimization (GRPO), where the policy must traverse the entire reverse process b…

Cited by 0SourceScholar
2025

3D Vision-Language Gaussian Splatting

ICLR 2025poster

Recent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality. However, current multi-modal scene understanding approaches h…

Cited by 16SourcePDFScholar
2025

A VisuoMotor Human-Robot Interaction Framework for Attention-Motion-Integrated Training

IROS 2025

Focus of attention is one of the most influential factors facilitating motor training performance. Most of robotic training methods have not well solved the negative effect of divided-attention on motor execution performance, resulting in limited rehabilitation efficiency for motor-cognitive dysfunc

Cited by 0SourceScholar
2025

ACEBench: A Comprehensive Evaluation of LLM Tool Usage

EMNLP 2025

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex problems. However, existing benchmarks for evaluating LLMs’ tool usage face several limitations: (1) limited evaluation

Cited by 0SourcePDFScholar
2025

Adversarial Preference Learning for Robust LLM Alignment

ACL 2025finding

Modern language models often rely on Reinforcement Learning from Human Feedback (RLHF) to encourage safe behaviors. However, they remain vulnerable to adversarial attacks due to three key limitations: (1) the inefficiency and high cost of human annotation, (2) the vast diversity of potential adversa…

2025

An Adversarial Learning Framework for Reliable Myoelectric Force Estimation Under Fatigue

ICRA 2025

Electromyography (EMG) signals are widely used as control inputs for myoelectric exoskeletons. However, muscle fatigue, which can result from prolonged use or heavy loads, significantly affects muscle activation patterns, leading to reduced estimation accuracy. To address this challenge, we propose

Cited by 0SourceScholar
2025

Argus: A Compact and Versatile Foundation Model for Vision

CVPR 2025poster

While existing vision and multi-modal foundation models can handle multiple computer vision tasks, they often suffer from significant limitations, including huge demand for data and computational resources during training and inconsistent performance across vision tasks at deployment time. To addres…

Cited by 0SourcePDFScholar
2025

AttentionPredictor: Temporal Patterns Matter for KV Cache Compression

NeurIPS 2025poster

With the development of large language models (LLMs), efficient inference through Key-Value (KV) cache compression has attracted considerable attention, especially for long-context generation. To compress the KV cache, recent methods identify critical KV tokens through static modeling of attention s…

Cited by 0SourcecodeScholar
2025

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators

ICLR 2025poster

An ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs remain unaware of the quality of the speech they process. Th…

Cited by 1SourcePDFScholar
2025

Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context Learning

ACL 2025long

Large language models (LLMs) have shown impressive few-shot generalization on many tasks via in-context learning (ICL). Despite their success in showing such emergent abilities, the scale and complexity of larger models also lead to unprecedentedly high computational demands and deployment challenge…

Cited by 0SourcePDFScholar
2025

BrainMAP: Learning Multiple Activation Pathways in Brain Networks

AAAI 2025technical

Functional Magnetic Resonance Image (fMRI) is commonly employed to study human brain activity, since it offers insight into the relationship between functional fluctuations and human behavior. To enhance analysis and comprehension of brain activity, Graph Neural Networks (GNNs) have been widely appl…

2025

CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching

NeurIPS 2025spotlight

Conditional generative modeling aims to learn a conditional data distribution from samples containing data-condition pairs. For this, diffusion and flow-based methods have attained compelling results. These methods use a learned (flow) model to transport an initial standard Gaussian noise that ignor…

Cited by 0SourceScholar
2025

CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling

EMNLP 2025

Mixture-of-Experts (MoE) models are crucial for scaling model capacity while controlling inference costs. While integrating MoE into multimodal models like CLIP improves performance, training these models is notoriously challenging and expensive. We propose CLIP-Upcycling (CLIP-UP), an efficient alt

Cited by 0SourcePDFScholar
2025

CPO: Condition Preference Optimization for Controllable Image Generation

NeurIPS 2025poster

To enhance controllability in text-to-image generation, ControlNet introduces image-based control signals, while ControlNet++ improves pixel-level cycle consistency between generated images and the input control signal. To avoid the prohibitive cost of back-propagating through the sampling process,…

Cited by 0SourcecodeScholar
2025

DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers

CVPR 2025poster

Visual Prompt Tuning (VPT) has become a promising solution for Parameter-Efficient Fine-Tuning (PEFT) approach for Vision Transformer (ViT) models by partially fine-tuning learnable tokens while keeping most model parameters frozen. Recent research has explored modifying the connection structures of…

2025

Dive into Aerial Remote Sensing Underwater Depth Estimation with Hyperspectral Imagery

AAAI 2025technical

Visible spectrum images capture limited information from just three discrete bands, often resulting in suboptimal performance in underwater depth estimation (UDE) due to significant information loss from water absorption. In contrast, HSIs, which include hundreds of continuous bands, provide abunda…

2025

DuPI: Dual-resolution Pseudo-label Integration for Semi-supervised Instance Segmentation

ICASSP 2025accepted

The role of high-quality pseudo-labels is pivotal in semi-supervised instance segmentation (SSIS). However, existing SSIS frameworks predominantly produce pseudo-labels at a single resolution, which can introduce noise that adversely affects the quality of learning at both the pixel level and in ter…

Cited by 0SourceScholar
2025

EGGS: Exchangeable 2D/3D Gaussian Splatting for Geometry-Appearance Balanced Novel View Synthesis

NeurIPS 2025spotlight

Novel view synthesis (NVS) is crucial in computer vision and graphics, with wide applications in AR, VR, and autonomous driving. While 3D Gaussian Splatting (3DGS) enables real-time rendering with high appearance fidelity, it suffers from multi-view inconsistencies, limiting geometric accuracy. In c…

Cited by 0SourceScholar
2025

Enhancing Foundation Models with Federated Domain Knowledge Infusion

ICML 2025poster

Vision foundation models (FMs) like CLIP have exhibited exceptional capabilities in visual and linguistic understanding, particularly in zero-shot inference tasks. However, these models struggle with data that significantly deviates from their training samples, necessitating fine-tuning, which is of…

Cited by 0SourcePDFScholar
2025

Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models

CVPR 2025poster

Text-to-image diffusion models have demonstrated remarkable capabilities in creating images highly aligned with user prompts, yet their proclivity for memorizing training set images has sparked concerns about the originality of the generated images and privacy issues, potentially leading to legal co…

Cited by 0SourcePDFScholar
2025

Exploit Gradient Skewness to Circumvent Byzantine Defenses for Federated Learning

AAAI 2025technical

Federated Learning (FL) is notorious for its vulnerability to Byzantine attacks. Most current Byzantine defenses share a common inductive bias: among all the gradients, the densely distributed ones are more likely to be honest. However, such a bias is a poison to Byzantine robustness due to a newly…

2025

Exploring Local Memorization in Diffusion Models via Bright Ending Attention

ICLR 2025spotlight

Text-to-image diffusion models have achieved unprecedented proficiency in generating realistic images. However, their inherent tendency to memorize and replicate training data during inference raises significant concerns, including potential copyright infringement. In response, various methods have…

Cited by 2SourcePDFScholar
2025

FGDGNN: Fine-Grained Dynamic Graph Neural Network for Rumor Detection on Social Media

ACL 2025finding

Detecting rumors on social media has become a crucial issue.Propagation structure-based methods have recently attracted increasing attention.When the propagation structure is represented by the dynamic graph, temporal information is considered.However, existing rumor detection models using dynamic g…

Cited by 0SourcePDFScholar
2025

Frequency-Semantic Enhanced Variational Autoencoder for Zero-Shot Skeleton-based Action Recognition

ICCV 2025poster

Zero-shot skeleton-based action recognition aims to develop models capable of identifying actions beyond the categories encountered during training. Previous approaches have primarily focused on aligning visual and semantic representations but often overlooked the importance of fine-grained action p…

2025

From Coarse to Fine: A Matching and Alignment Framework for Unsupervised Cross-View Geo-Localization

AAAI 2025technical

Cross-view geo-localization aims at determining the geographic location of a query image by matching the reference images. The matching pairs can be captured from diverse perspectives, such as those from satellites and drones. Most existing methods are supervised that require input of location-label…

Cited by 0SourcePDFScholar
2025

Fusion Meets Diverse Conditions: A High-diversity Benchmark and Baseline for UAV-based Multimodal Object Detection with Condition Cues

ICCV 2025poster

Unmanned aerial vehicles (UAV)-based object detection with visible (RGB) and infrared (IR) images facilitates robust around-the-clock detection, driven by advancements in deep learning techniques and the availability of high-quality dataset. However, the existing dataset struggles to fully capture r…

Cited by 0SourcePDFScholar
2025

GenHMR: Generative Human Mesh Recovery

AAAI 2025technical

Human mesh recovery (HMR) is crucial in many computer vision applications; from health to arts and entertainment. HMR from monocular images has predominantly been addressed by deterministic methods that output a single prediction for a given 2D image. However, HMR from a single image is an ill-posed…

Cited by 0SourcePDFScholar
2025

GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling

ICLR 2025poster

Semantic information refers to the meaning conveyed through words, phrases, and contextual relationships within a given linguistic structure. Humans can leverage semantic information, such as familiar linguistic patterns and contextual cues, to reconstruct incomplete or masked speech signals in nois…

Cited by 1SourcePDFScholar
2025

GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language Models

AAAI 2025technical

Posters serve an essential function in marketing and advertising by improving visual communication and brand visibility, thus significantly contributing to industrial design. With the latest developments in controllable T2I diffusion models, research interest has surged in text rendering within synt…

2025

How to Evaluate and Mitigate IP Infringement in Visual Generative AI?

ICML 2025poster

The popularity of visual generative AI models like DALL-E 3, Stable Diffusion XL, Stable Video Diffusion, and Sora has been increasing. Through extensive evaluation, we discovered that the state-of-the-art visual generative models can generate content that bears a striking resemblance to characters…

Cited by 0SourcePDFScholar
2025

InstructGEC: Enhancing Unsupervised Grammatical Error Correction with Instruction Tuning

COLING 2025main

Recent works have proposed methods of generating synthetic data automatically for unsupervised Grammatical Error Correction (GEC). Although a large amount of synthetic data is generated at a low cost, it is unrealistic and of poor quality. The copying phenomenon of synthetic data prevents GEC models…

2025

Interaction-Driven Updates: 3D Scene Graph Maintenance During Robot Task Execution

ICRA 2025

Robots powered by large language model (LLM) demonstrate significant research and application potential by effectively interpreting scene information to respond to human commands. However, when robots rely on static scene information during task execution, they face difficulties in adapting to chang

Cited by 0SourceScholar
2025

Local Policies Enable Zero-Shot Long-Horizon Manipulation

ICRA 2025

Sim2real for robotic manipulation is difficult due to the challenges of simulating complex contacts and generating realistic task distributions. To tackle the latter problem, we introduce ManipGen, which leverages a new class of policies for sim2real transfer: local policies. Locality enables a vari

Cited by 31SourcecodeScholar
2025

MaskControl: Spatio-Temporal Control for Masked Motion Synthesis

ICCV 2025poster

Recent advances in motion diffusion models have enabled spatially controllable text-to-motion generation. However, these models struggle to achieve high-precision control while maintaining high-quality motion generation. To address these challenges, we propose MaskControl, the first approach to intr…

2025

MixA: A Mixed Attention approach with Stable Lightweight Linear Attention to enhance Efficiency of Vision Transformers at the Edge

ICCV 2025poster

Vision transformers (ViTs) have become widely popular due to their strong performance across various computer vision tasks. However, deploying ViTs on edge devices remains a persistent challenge due to their high computational demands primarily caused by the over use of self-attention layers with qu…

Cited by 0SourcePDFScholar
2025

Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models

ICLR 2025poster

Federated prompt learning benefits federated learning with CLIP-like Vision-Language Model's (VLM's) robust representation learning ability through prompt learning. However, current federated prompt learning methods are habitually restricted to the traditional FL paradigm, where the participating cl…

2025

Model-Free Catheter Delivery Strategy for Robotic Transcatheter Tricuspid Valve Replacement

IROS 2025

Transcatheter tricuspid valve replacement (TTVR) has emerged as a promising minimally invasive procedure for treating severe tricuspid regurgitation (TR). However, accurate catheter delivery remains a significant challenge, primarily due to the reliance on 2D vision feedback, complex catheter kinema

Cited by 0SourceScholar
2025

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level

CVPR 2025poster

In this paper, we introduce Motion-Grounded Video Reasoning, a new motionunderstanding task that requires generating visual answers (video segmentationmasks) according to the input question, and hence needs implicit spatiotemporalreasoning and grounding. This task extends existing spatiotemporal gro…

Cited by 3SourcePDFScholar
2025

MultiCAT: Multimodal Communication Annotations for Teams

NAACL 2025findings

Successful teamwork requires team members to understand each other and communicate effectively, managing multiple linguistic and paralinguistic tasks at once. Because of the potential for interrelatedness of these tasks, it is important to have the ability to make multiple types of predictions on th…

Cited by 0SourcePDFScholar
2025

Online-HMM with Two-Layer Bayesian Method for Operator's Expected Speed Estimation in Teleoperated Gluing Tasks *

IROS 2025

For direct teleoperation tasks, the follower robot accomplishes tasks by strictly executing the inputs from the operator. However, the operator's physiological tremor seriously reduces the smoothness of the trajectory, especially in tasks relying on operator’s experience such as gluing, while the ra

Cited by 0SourceScholar
2025

Out-of-Distribution Generalization on Graphs via Progressive Inference

AAAI 2025technical

The development and evaluation of graph neural networks (GNNs) generally follow the independent and identically distributed (i.i.d.) assumption. Yet this assumption is often untenable in practice due to the uncontrollable data generation mechanism. In particular, when the data distribution shows a s…

2025

Predicting Through Generation: Why Generation Is Better for Prediction

ACL 2025long

This paper argues that generating output tokens is more effective than using pooled representations for prediction tasks because token-level generation retains more mutual information. Since LLMs are trained on massive text corpora using next-token prediction, generation aligns naturally with their…

2025

ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models

COLING 2025main

Activation sparsity refers to the existence of considerable weakly-contributed elements among activation outputs, serving as a promising paradigm for accelerating model inference. Nevertheless, most large language models (LLMs) adopt activation functions without intrinsic activation sparsity (e.g.,…

2025

Question-Aware Knowledge Graph Prompting for Enhancing Large Language Models

ACL 2025finding

Large Language Models (LLMs) often struggle with tasks requiring external knowledge, such as knowledge-intensive Multiple Choice Question Answering (MCQA). Integrating Knowledge Graphs (KGs) can enhance reasoning; however, existing methods typically demand costly fine-tuning or retrieve noisy KG inf…

2025

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

NeurIPS 2025poster

Transformer-based Large Language Models (LLMs) have become increasingly important. However, scaling LLMs to longer contexts incurs slow inference speed and high GPU memory consumption for caching key-value (KV) vectors. This paper presents RetrievalAttention, a training-free approach to both acceler…

Cited by 0SourcecodeScholar
2025

Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

ICLR 2025poster

Recent advancements in multimodal models highlight the value of rewritten captions for improving performance, yet key challenges remain. For example, while synthetic captions often provide superior quality and image-text alignment, it is not clear whether they can fully replace AltTexts: the role of…

Cited by 4SourcePDFScholar
2025

Revisiting Graph Contrastive Learning on Anomaly Detection: A Structural Imbalance Perspective

AAAI 2025technical

The superiority of graph contrastive learning (GCL) has prompted its application to anomaly detection tasks for more powerful risk warning systems. Unfortunately, existing GCL-based models tend to excessively prioritize overall detection performance while neglecting robustness to structural imbalanc…

2025

RoCoFT: Efficient Finetuning of Large Language Models with Row-Column Updates

ACL 2025long

We propose Row-Column Fine-Tuning(RoCoFT), a parameter-efficient fine-tuning method for large language models based on updating only a few rows and columns of the weight matrices in transformers. Through extensive experiments with medium-sized LMs like RoBERTa and DeBERTa, and larger LMs like Bloom-…

2025

Robotic In-Hand Manipulation for Large-Range Precise Object Movement: The RGMC Champion Solution

RA-L 2025

In-hand manipulation using multiple dexterous fingers is a critical robotic skill that can reduce the reliance on large arm motions, thereby saving space and energy. This letter focuses on in-grasp object movement, which refers to manipulating an object to a desired pose through only finger motions

Cited by 9SourceScholar
2025

SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real World

ICCV 2025poster

Existing vision-based 3D occupancy prediction methods are inherently limited in accuracy due to their exclusive reliance on street-view imagery, neglecting the potential benefits of incorporating satellite views. We propose SA-Occ, the first Satellite-Assisted 3D occupancy prediction model, which le…

2025

SCott: Accelerating Diffusion Models with Stochastic Consistency Distillation

AAAI 2025technical

The iterative sampling procedure employed by diffusion models (DMs) often leads to significant latency. To address this, we propose Stochastic Consistency Distillation (SCott) to enable accelerated text-to-image generation, where high-quality generations can be achieved with just 2-4 sampling steps…

Cited by 2SourcePDFScholar
2025

SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning

NeurIPS 2025poster

Existing diffusion-based super-resolution approaches often exhibit semantic ambiguities due to inaccuracies and incompleteness in their text conditioning, coupled with the inherent tendency for cross-attention to divert towards irrelevant pixels. These limitations can lead to semantic misalignment a…

Cited by 0SourceScholar
2025

SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

ICASSP 2025accepted

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the g…

Cited by 0SourceScholar
2025

ST-FiT: Inductive Spatial-Temporal Forecasting with Limited Training Data

AAAI 2025technical

Spatial-temporal graphs are widely used in a variety of real-world applications. Spatial-Temporal Graph Neural Networks (STGNNs) have emerged as a powerful tool to extract meaningful insights from this data. However, in real-world applications, most nodes may not possess any available temporal data…

2025

STIV: Scalable Text and Image Conditioned Video Generation

ICCV 2025poster

We present a simple and scalable text and image conditioned video generation method. Our approach, named STIV, integrates a variable number of image conditions into a Diffusion Transformer (DiT) through frame replacement. This design enables STIV to perform both text-to-video (T2V) and text-image-to…

2025

SWAM: Adaptive Sliding Window and Memory-Augmented Attention Model for Rumor Detection

EMNLP 2025

Detecting rumors on social media has become a critical task in combating misinformation. Existing propagation-based rumor detection methods often focus on the static propagation graph, overlooking that rumor propagation is inherently dynamic and incremental in the real world. Recently propagation-ba

Cited by 0SourcePDFScholar
2025

SafeInt: Shielding Large Language Models from Jailbreak Attacks via Safety-Aware Representation Intervention

EMNLP 2025

With the widespread real-world deployment of large language models (LLMs), ensuring their behavior complies with safety standards has become crucial. Jailbreak attacks exploit vulnerabilities in LLMs to induce undesirable behavior, posing a significant threat to LLM safety. Previous defenses often f

2025

SemStereo: Semantic-Constrained Stereo Matching Network for Remote Sensing

AAAI 2025technical

Semantic segmentation and 3D reconstruction are two fundamental tasks in remote sensing, typically treated as separate or loosely coupled tasks. Despite attempts to integrate them into a unified network, the constraints between the two heterogeneous tasks are not explicitly modeled, since the pionee…

Cited by 0SourcePDFScholar
2025

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding

CVPR 2025poster

Temporal awareness is essential for video large language models (LLMs) to understand and reason about events within long videos, enabling applications like dense video captioning and temporal video grounding in a unified system. However, the scarcity of long videos with detailed captions and precise…

Cited by 1SourcePDFScholar
2025

SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality

ICCV 2025poster

In this paper, we propose SimMLM, a simple yet powerful framework for multimodal learning with missing modalities. Unlike existing approaches that rely on sophisticated network architectures or complex data imputation techniques, SimMLM provides a generic and effective solution that can adapt to var…

2025

SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing

ICCV 2025poster

Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signals caused by the mismatch between editing instructions and original-edited image pairs. Recent efforts attempt to improve…

2025

TARFVAE: Efficient One-Step Generative Time Series Forecasting via TARFLOW based VAE

NeurIPS 2025poster

Time series data is ubiquitous, with forecasting applications spanning from finance to healthcare. Beyond popular deterministic methods, generative models are gaining attention due to advancements in areas like image synthesis and video generation, as well as their inherent ability to provide probab…

Cited by 0SourcecodeScholar
2025

Towards Global-level Mechanistic Interpretability: A Perspective of Modular Circuits of Large Language Models

ICML 2025poster

Mechanistic interpretability (MI) research aims to understand large language models (LLMs) by identifying computational circuits, subgraphs of model components with associated functional interpretations, that explain specific behaviors. Current MI approaches focus on discovering task-specific circui…

Cited by 0SourcePDFScholar
2025

UCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-Identification

CVPR 2025poster

Cross-Modality Re-Identification (VI-ReID) aims to achieve around-the-clock target matching, benefiting from the strengths of both RGB and infrared (IR) modalities. However, the field is hindered by limited datasets, particularly for vehicle VI-ReID, and by challenges such as modality bias training…

Cited by 0SourcePDFScholar
2025

UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing

ICCV 2025poster

Text-to-Image (T2I) diffusion models have shown impressive results in generating visually compelling images following user prompts. Building on this, various methods further fine-tune the pre-trained T2I model for specific tasks. However, this requires separate model architectures, training designs,…

2025

Universal Online Temporal Calibration for Optimization-Based Visual-Inertial Navigation Systems

ICRA 2025

6-Degree of Freedom (6DoF) motion estimation with a combination of visual and inertial sensors is a growing area with numerous real-world applications. However, precise calibration of the time offset between these two sensor types is a prerequisite for accurate and robust tracking. To address this,

Cited by 0SourcecodeScholar
2025

Virtual Nodes Can Help: Tackling Distribution Shifts in Federated Graph Learning

AAAI 2025technical

Federated Graph Learning (FGL) enables multiple clients to jointly train powerful graph learning models, e.g., Graph Neural Networks (GNNs), without sharing their local graph data for graph-related downstream tasks, such as graph property prediction. In the real world, however, the graph data can su…

2025

Wasserstein Heterogeneous Graph Neural Networks for Uncertainty-Aware Anomaly Detection

ICASSP 2025accepted

Graph anomaly detection, a critical topic in graph mining, has garnered significant research interest and found applications across diverse domains such as attack event detection, spam review identification, and financial fraud prevention. Graph Neural Networks (GNNs) have emerged as the dominant ap…

Cited by 0SourceScholar
2025

X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation

ICCV 2025poster

Text-to-image (T2I) models are well known for their ability to produce highly realistic images, while multimodal large language models (MLLMs) are renowned for their proficiency in understanding and integrating multiple modalities. However, currently there is no straightforward and efficient framewo…

2024

3D Ultrasound Image Acquisition and Diagnostic Analysis of the Common Carotid Artery with a Portable Robotic Device

IROS 2024poster

Ultrasound (US) imaging of the carotid artery (CA) is a non-invasive diagnostic tool widely used in the medical field to assess the condition of the carotid artery, thereby predicting the risk of cardiovascular and cerebrovascular diseases. However, implementing this method in primary healthcare can…

Cited by 0SourceScholar
2024

A Dual-Augmentor Framework for Domain Generalization in 3D Human Pose Estimation

CVPR 2024poster

3D human pose data collected in controlled laboratory settings present challenges for pose estimators that generalize across diverse scenarios. To address this domain generalization is employed. Current methodologies in domain generalization for 3D human pose estimation typically utilize adversarial…

2024

A Simple Background Augmentation Method for Object Detection with Diffusion Model

ECCV 2024poster

"In computer vision, it is well-known that a lack of data diversity will impair model performance. In this study, we address the challenges of enhancing the dataset diversity problem in order to benefit various downstream tasks such as object detection and instance segmentation. We propose a simple…

Cited by 5SourcePDFScholar
2024

A Unified Interaction Control Framework for Safe Robotic Ultrasound Scanning with Human-Intention-Aware Compliance

IROS 2024

The ultrasound scanning robot operates in environments where frequent human-robot interactions occur. Most existing control methods for ultrasound scanning address only one specific interaction situation or implement hard switches between controllers for different situations, which compromises both

Cited by 6SourceScholar
2024

Adaptive FSS: A Novel Few-Shot Segmentation Framework via Prototype Enhancement

AAAI 2024technical

The Few-Shot Segmentation (FSS) aims to accomplish the novel class segmentation task with a few annotated images. Current FSS research based on meta-learning focuses on designing a complex interaction mechanism between the query and support feature. However, unlike humans who can rapidly learn new t…

2024

Advancing Video Anomaly Detection: A Concise Review and a New Dataset

NeurIPS 2024poster

Video Anomaly Detection (VAD) finds widespread applications in security surveillance, traffic monitoring, industrial monitoring, and healthcare. Despite extensive research efforts, there remains a lack of concise reviews that provide insightful guidance for researchers. Such reviews would serve as q…

Cited by 14SourcePDFScholar
2024

Adversarial Attacks on Fairness of Graph Neural Networks

ICLR 2024poster

Fairness-aware graph neural networks (GNNs) have gained a surge of attention as they can reduce the bias of predictions on any demographic group (e.g., female) in graph-based applications. Although these methods greatly improve the algorithmic fairness of GNNs, the fairness can be easily corrupted b…

2024

BAMM: Bidirectional Autoregressive Motion Model

ECCV 2024poster

"Generating human motion from text has been dominated by denoising motion models either through diffusion or generative masking process. However, these models face great limitations in usability by requiring prior knowledge of the motion length. Conversely, autoregressive motion models address this…

2024

COALA: A Practical and Vision-Centric Federated Learning Platform

ICML 2024poster

We present COALA, a vision-centric Federated Learning (FL) platform, and a suite of benchmarks for practical FL scenarios, which we categorize as task, data, and model levels. At the task level, COALA extends support from simple classification to 15 computer vision tasks, including object detection,…

2024

Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models

AAAI 2024technical

Recent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts. However, they fail to semantically align the generated images with the prompts due to their limited compositional capabilities, leading to attribute leakage, ent…

2024

ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback

ECCV 2024poster

"To enhance the controllability of text-to-image diffusion models, existing efforts like ControlNet incorporated image-based conditional controls. In this paper, we reveal that existing methods still face significant challenges in generating images that align with the image conditional controls. To…

2024

Cross-Modality and Within-Modality Regularization for Audio-Visual Deepfake Detection

ICASSP 2024accepted

Audio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal targets face challenges due to uncertainties and inconsistencies in learned representations caused by independent modali…

Cited by 0SourceScholar
2024

DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models

ICLR 2024poster

Recent text-to-image diffusion models have shown surprising performance in generating high-quality images. However, concerns have arisen regarding the unauthorized data usage during the training or fine-tuning process. One example is when a model trainer collects a set of images created by a particu…

2024

Decouple Content and Motion for Conditional Image-to-Video Generation

AAAI 2024technical

The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text. The previous cI2V generation methods conventionally perform in RGB pixel space, with limitations in modeling motion consistency and visual continuit…

Cited by 5SourcePDFScholar
2024

Detecting, Explaining, and Mitigating Memorization in Diffusion Models

ICLR 2024oral

Recent breakthroughs in diffusion models have exhibited exceptional image-generation capabilities. However, studies show that some outputs are merely replications of training data. Such replications present potential legal challenges for model owners, especially when the generated content contains p…

2024

Few-shot Knowledge Graph Relational Reasoning via Subgraph Adaptation

NAACL 2024long

Few-shot Knowledge Graph (KG) Relational Reasoning aims to predict unseen triplets (i.e., query triplets) for rare relations in KGs, given only several triplets of these relations as references (i.e., support triplets). This task has gained significant traction due to the widespread use of knowledge…

2024

Free-Editor: Zero-shot Text-driven 3D Scene Editing

ECCV 2024poster

"Text-to-Image (T2I) diffusion models have recently gained traction for their versatility and user-friendliness in 2D content generation and editing. However, training a diffusion model specifically for 3D scene editing is challenging due to the scarcity of large-scale datasets. Currently, editing 3…

2024

GCNext: Towards the Unity of Graph Convolutions for Human Motion Prediction

AAAI 2024technical

The past few years has witnessed the dominance of Graph Convolutional Networks (GCNs) over human motion prediction. Various styles of graph convolutions have been proposed, with each one meticulously designed and incorporated into a carefully-crafted network architecture. This paper breaks the limit…

2024

GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators

ACL 2024long

Recent advances in large language models (LLMs) have stepped forward the development of multilingual speech and machine translation by its reduced representation errors and incorporated external knowledge. However, both translation tasks typically utilize beam search decoding and top-1 hypothesis se…

2024

How to Trace Latent Generative Model Generated Images without Artificial Watermark?

ICML 2024poster

Latent generative models (e.g., Stable Diffusion) have become more and more popular, but concerns have arisen regarding potential misuse related to images generated by these models. It is, therefore, necessary to analyze the origin of images by inferring if a particular image was generated by a spec…

2024

In-Context Learning with Iterative Demonstration Selection

EMNLP 2024finding

Spurred by advancements in scale, large language models (LLMs) have demonstrated strong few-shot learning ability via in-context learning (ICL). However, the performance of ICL has been shown to be highly sensitive to the selection of few-shot demonstrations. Selecting the most suitable examples as…

Cited by 44SourcePDFScholar
2024

InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions

NAACL 2024long

Instruction tuning effectively optimizes Large Language Models (LLMs) for downstream tasks. Due to the changing environment in real-life applications, LLMs necessitate continual task-specific adaptation without catastrophic forgetting. Considering the heavy computational cost, replay-based Continual…

Cited by 35SourcePDFScholar
2024

It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition

ICLR 2024poster

Recent studies have successfully shown that large language models (LLMs) can be successfully used for generative error correction (GER) on top of the automatic speech recognition (ASR) output. Specifically, an LLM is utilized to carry out a direct mapping from the N-best hypotheses list generated by…

Cited by 27SourcePDFScholar
2024

Large Language Models are Efficient Learners of Noise-Robust Speech Recognition

ICLR 2024spotlight

Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which leverages the rich linguistic knowledge and powerful reasoning ability of LLMs to improve recognition results. The latest work proposes a GER benchmark with "…

2024

LatentEditor: Text Driven Local Editing of 3D Scenes

ECCV 2024poster

"While neural fields have made significant strides in view synthesis and scene reconstruction, editing them poses a formidable challenge due to their implicit encoding of geometry and texture information from multi-view inputs. In this paper, we introduce LatentEditor, an innovative framework design…

2024

Learning Semantic Proxies from Visual Prompts for Parameter-Efficient Fine-Tuning in Deep Metric Learning

ICLR 2024poster

Deep Metric Learning (DML) has long attracted the attention of the machine learning community as a key objective. Existing solutions concentrate on fine-tuning the pre-trained models on conventional image datasets. As a result of the success of recent pre-trained models derived from larger-scale dat…

2024

Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models

ACL 2024findings

Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which aims to predict the ground-truth transcription from the decoded N-best hypotheses. Thanks to the strong language generation ability of LLMs and rich informati…

2024

MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval

NAACL 2024findings

Due to the success of large-scale visual-language pretraining (VLP) models and the widespread use of image-text retrieval in industry areas, it is now critically necessary to reduce the model size and streamline their mobile-device deployment. Single- and dual-stream model structures are commonly us…

2024

MOFI: Learning Image Representations from Noisy Entity Annotated Images

ICLR 2024poster

We present MOFI, Manifold OF Images, a new vision foundation model designed to learn image representations from noisy entity annotated images. MOFI differs from previous work in two key aspects: 1. pre-training data, and 2. training recipe. Regarding data, we introduce a new approach to automaticall…

2024

Multi-Signal Fusion of Social Diffusion Graph with Bi-Directional Semantic Consistency

ICASSP 2024accepted

Devising diffusion graph to learn user representations is a crucial step in studying information propagation prediction. However, previous works mainly focused on structural and temporal features. To better incorporate content features, we introduce the Backward Decomposition and Forward Preservatio…

Cited by 0SourceScholar
2024

Multi-View Attentive Contextualization for Multi-View 3D Object Detection

CVPR 2024poster

We present Multi-View Attentive Contextualization (MvACon) a simple yet effective method for improving 2D-to-3D feature lifting in query-based multi-view 3D (MV3D) object detection. Despite remarkable progress witnessed in the field of query-based MV3D object detection prior art often suffers from e…

Cited by 2SourcePDFScholar
2024

Noise-Aware Speech Separation with Contrastive Learning

ICASSP 2024accepted

Recently, speech separation (SS) task has achieved remarkable progress driven by deep learning technique. However, it is still challenging to separate target speech from noisy mixture, as the neural model is vulnerable to assign background noise to each speaker. In this paper, we propose a noise-awa…

Cited by 0SourceScholar
2024

OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition

CVPR 2024poster

Due to the resource-intensive nature of training vision-language models on expansive video data a majority of studies have centered on adapting pre-trained image-language models to the video domain. Dominant pipelines propose to tackle the visual discrepancies with additional temporal learners while…

2024

Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression

NeurIPS 2024poster

In offline reinforcement learning (RL), addressing the out-of-distribution (OOD) action issue has been a focus, but we argue that there exists an OOD state issue that also impairs performance yet has been underexplored. Such an issue describes the scenario when the agent encounters states out of the…

2024

Overcoming Catastrophic Forgetting by Exemplar Selection in Task-oriented Dialogue System

ACL 2024findings

Intelligent task-oriented dialogue systems (ToDs) are expected to continuously acquire new knowledge, also known as Continual Learning (CL), which is crucial to fit ever-changing user needs. However, catastrophic forgetting dramatically degrades the model performance in face of a long streamed curri…

Cited by 0SourcePDFScholar
2024

Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective

ECCV 2024poster

"This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposal selection framework that utilizes contrastive learning and reconstruction paradigm for scoring the pre-defined moment proposals. Although they have…

Cited by 16SourcePDFScholar
2024

Robust and Scalable Model Editing for Large Language Models

COLING 2024main

Large language models (LLMs) can make predictions using *parametric knowledge* – knowledge encoded in the model weights – or *contextual knowledge* – knowledge presented in the context. In many scenarios, a desirable behavior is that LLMs give precedence to contextual knowledge when it conflicts wit…

2024

SEPT: Towards Efficient Scene Representation Learning for Motion Prediction

ICLR 2024poster

Motion prediction is crucial for autonomous vehicles to operate safely in complex traffic environments. Extracting effective spatiotemporal relationships among traffic elements is key to accurate forecasting. Inspired by the successful practice of pretrained large language models, this paper present…

Cited by 33SourcePDFScholar
2024

STL-SLAM: A Structured-Constrained RGB-D SLAM Approach to Texture-Limited Environments

IROS 2024poster

Most RGB-D-based SLAM methods assume texture-rich environments, making them susceptible to significant tracking errors or complete failures in the absence of texture features. Moreover, many existing methods encounter substantial rotation estimation errors, leading to long-term drift in tracking. Th…

Cited by 1SourceScholar
2024

Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models

NeurIPS 2024poster

We propose an unsupervised adaptation framework, Self-TAught Recognizer (STAR), which leverages unlabeled data to enhance the robustness of automatic speech recognition (ASR) systems in diverse target domains, such as noise and accents. STAR is developed for prevalent speech foundation models based…

2024

Sparse Points to Dense Clouds: Enhancing 3D Detection with Limited LiDAR Data

IROS 2024

3D detection is a critical task that enables machines to identify and locate objects in three-dimensional space. It has a broad range of applications in several fields, including autonomous driving, robotics and augmented reality. Monocular 3D detection is attractive as it requires only a single cam

Cited by 6SourceScholar
2024

Taming Cross-Domain Representation Variance in Federated Prototype Learning with Heterogeneous Data Domains

NeurIPS 2024poster

Federated learning (FL) allows collaborative machine learning training without sharing private data. While most FL methods assume identical data domains across clients, real-world scenarios often involve heterogeneous data domains. Federated Prototype Learning (FedPL) addresses this issue, using mea…

Cited by 7SourcePDFScholar
2024

The Ladder in Chaos: Improving Policy Learning by Harnessing the Parameter Evolving Path in A Low-dimensional Space

NeurIPS 2024poster

Knowing the learning dynamics of policy is significant to unveiling the mysteries of Reinforcement Learning (RL). It is especially crucial yet challenging to Deep RL, from which the remedies to notorious issues like sample inefficiency and learning instability could be obtained. In this paper, we st…

Cited by 1SourcePDFScholar
2024

Towards Diverse Device Heterogeneous Federated Learning via Task Arithmetic Knowledge Integration

NeurIPS 2024poster

Federated Learning (FL) has emerged as a promising paradigm for collaborative machine learning, while preserving user data privacy. Despite its potential, standard FL algorithms lack support for diverse heterogeneous device prototypes, which vary significantly in model and dataset sizes---from small…

2024

Towards Improved Proxy-Based Deep Metric Learning via Data-Augmented Domain Adaptation

AAAI 2024technical

Deep Metric Learning (DML) plays an important role in modern computer vision research, where we learn a distance metric for a set of image representations. Recent DML techniques utilize the proxy to interact with the corresponding image samples in the embedding space. However, existing proxy-based D…

2024

Towards Multi-modal Transformers in Federated Learning

ECCV 2024poster

"Multi-modal transformers mark significant progress in different domains, but privacy concerns on high-quality data hinder their further improvement. Federated learning (FL) has emerged as a promising privacy-preserving paradigm for training models without direct access to the raw data held by diffe…

2024

Towards Surveillance Video-and-Language Understanding: New Dataset Baselines and Challenges

CVPR 2024poster

Surveillance videos are important for public security. However current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to detecting and classifying the predefined events with unsatisfactory semantic understanding although they have o…

Cited by 18SourcePDFScholar
2024

Visual Attention Based Cognitive Human–Robot Collaboration for Pedicle Screw Placement in Robot-Assisted Orthopedic Surgery

IROS 2024poster

Current orthopedic robotic systems largely focus on navigation, aiding surgeons in positioning a guiding tube but still requiring manual drilling and screw placement. The automation of this task not only demands high precision and safety due to the intricate physical interactions between the surgica…

Cited by 1SourceScholar
2024

Weakly Misalignment-free Adaptive Feature Alignment for UAVs-based Multimodal Object Detection

CVPR 2024poster

Visible-infrared (RGB-IR) image fusion has shown great potentials in object detection based on unmanned aerial vehicles (UAVs). However the weakly misalignment problem between multimodal image pairs limits its performance in object detection. Most existing methods often ignore the modality gap and e…

Cited by 6SourcePDFScholar
2023

A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action Recognition

ICCV 2023poster

The goal of building a benchmark (suite of datasets) is to provide a unified protocol for fair evaluation and thus facilitate the evolution of a specific area. Nonetheless, we point out that existing protocols of action recognition could yield partial evaluations due to several limitations. To compr…

Cited by 19PDFcodeScholar
2023

A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose Estimation

NeurIPS 2023poster

The dominant paradigm in 3D human pose estimation that lifts a 2D pose sequence to 3D heavily relies on long-term temporal clues (i.e., using a daunting number of video frames) for improved accuracy, which incurs performance saturation, intractable computation and the non-causal problem. This can be…

2023

AIM: Adapting Image Models for Efficient Video Action Recognition

ICLR 2023poster

Recent vision transformer based video models mostly follow the ``image pre-training then finetuning" paradigm and have achieved great success on multiple video benchmarks. However, fully finetuning such a video model could be computationally expensive and unnecessary, given the pre-trained image tra…

2023

AlignDet: Aligning Pre-training and Fine-tuning in Object Detection

ICCV 2023poster

The paradigm of large-scale pre-training followed by downstream fine-tuning has been widely employed in various object detection algorithms. In this paper, we reveal discrepancies in data, model, and task between the pre-training and fine-tuning procedure in existing practices, which implicitly limi…

Cited by 22PDFcodeScholar
2023

An Empirical Study of Frame Selection for Text-to-Video Retrieval

EMNLP 2023long findings

Text-to-video retrieval (TVR) aims to find the most relevant video in a large video gallery given a query text. The intricate and abundant context of the video challenges the performance and efficiency of TVR. To handle the serialized video contexts, existing methods typically select a subset of fra…

Cited by 0SourceScholar
2023

Byzantine-Robust Learning on Heterogeneous Data via Gradient Splitting

ICML 2023poster

Federated learning has exhibited vulnerabilities to Byzantine attacks, where the Byzantine attackers can send arbitrary gradients to a central server to destroy the convergence and performance of the global model. A wealth of robust AGgregation Rules (AGRs) have been proposed to defend against Byzan…

2023

CEFHRI: A Communication Efficient Federated Learning Framework for Recognizing Industrial Human-Robot Interaction

IROS 2023poster

Human-robot interaction (HRI) is a rapidly growing field that encompasses social and industrial applications. Machine learning plays a vital role in industrial HRI by enhancing the adaptability and autonomy of robots in complex environments. However, data privacy is a crucial concern in the interact…

Cited by 10SourcecodeScholar
2023

CN-CVS: A Mandarin Audio-Visual Dataset for Large Vocabulary Continuous Visual to Speech Synthesis

ICASSP 2023accepted

Research on Video to Speech Synthesis (VTS) surges recently and the focus is gradually shifting from small-vocabulary short-phrase VTS to large-vocabulary continuous VTS (LVC-VTS). A large-scale dataset with sufficient speakers and utterances is a prerequisite for such research, and the database is…

Cited by 0SourceScholar
2023

Combating Unknown Bias with Effective Bias-Conflicting Scoring and Gradient Alignment

AAAI 2023technical

Models notoriously suffer from dataset biases which are detrimental to robustness and generalization. The identify-emphasize paradigm shows a promising effect in dealing with unknown biases. However, we find that it is still plagued by two challenges: A, the quality of the identified bias-conflictin…

Cited by 9SourcePDFScholar
2023

Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement Learning

NeurIPS 2023poster

Offline multi-agent reinforcement learning is challenging due to the coupling effect of both distribution shift issue common in offline setting and the high dimension issue common in multi-agent setting, making the action out-of-distribution (OOD) and value overestimation phenomenon excessively seve…

2023

Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

IJCAI 2023poster

Audio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-invariant visual information. However, most existing AVSR approaches simply fuse the audio and visual features by concaten…

2023

Delving into the Adversarial Robustness of Federated Learning

AAAI 2023technical

In Federated Learning (FL), models are as fragile as centrally trained models against adversarial examples. However, the adversarial robustness of federated learning remains largely unexplored. This paper casts light on the challenge of adversarial robustness of federated learning. To facilitate a b…

Cited by 38SourcePDFScholar
2023

Dipping PLMs Sauce: Bridging Structure and Text for Effective Knowledge Graph Completion via Conditional Soft Prompting

ACL 2023findings

Knowledge Graph Completion (KGC) often requires both KG structural and textual information to be effective. Pre-trained Language Models (PLMs) have been used to learn the textual information, usually under the fine-tune paradigm for the KGC task. However, the fine-tuned PLMs often overwhelmingly foc…

2023

Dynamic Ensemble of Low-Fidelity Experts: Mitigating NAS “Cold-Start”

AAAI 2023technical

Predictor-based Neural Architecture Search (NAS) employs an architecture performance predictor to improve the sample efficiency. However, predictor-based NAS suffers from the severe ``cold-start'' problem, since a large amount of architecture-performance data is required to get a working predictor.…

2023

Dynamic Graph Learning With Content-Guided Spatial-Frequency Relation Reasoning for Deepfake Detection

CVPR 2023poster

With the springing up of face synthesis techniques, it is prominent in need to develop powerful face forgery detection methods due to security concerns. Some existing methods attempt to employ auxiliary frequency-aware information combined with CNN backbones to discover the forged clues. Due to the…

Cited by 110SourcePDFScholar
2023

Efficient Distribution Similarity Identification in Clustered Federated Learning via Principal Angles between Client Data Subspaces

AAAI 2023technical

Clustered federated learning (FL) has been shown to produce promising results by grouping clients into clusters. This is especially effective in scenarios where separate groups of clients have significant differences in the distributions of their local data. Existing clustered FL algorithms are esse…

2023

FedPerfix: Towards Partial Model Personalization of Vision Transformers in Federated Learning

ICCV 2023poster

Personalized Federated Learning (PFL) represents a promising solution for decentralized learning in heterogeneous data environments. Partial model personalization has been proposed to improve the efficiency of PFL by selectively updating local model parameters instead of aggregating all of them. How…

Cited by 22PDFcodeScholar
2023

Gaitmixer: Skeleton-Based Gait Representation Learning Via Wide-Spectrum Multi-Axial Mixer

ICASSP 2023accepted

Most existing gait recognition methods are appearance-based, which rely on the silhouettes extracted from the video data of human walking activities. The less-investigated skeleton-based gait recognition methods directly learn the gait dynamics from 2D/3D human skeleton sequences, which are theoreti…

Cited by 0SourceScholar
2023

Gradient Remedy for Multi-Task Learning in End-to-End Noise-Robust Speech Recognition

ICASSP 2023accepted

Speech enhancement (SE) is proved effective in reducing noise from noisy speech signals for downstream automatic speech recognition (ASR), where multi-task learning strategy is employed to jointly optimize these two tasks. However, the enhanced speech learned by SE objective may not always yield goo…

Cited by 0SourceScholar
2023

Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition

ACL 2023long

Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most existing efforts still focus on audio modality to improve robustness considering its dominance in AVSR task, with noise adap…

2023

Hokoff: Real Game Dataset from Honor of Kings and its Offline Reinforcement Learning Benchmarks

NeurIPS 2023poster

The advancement of Offline Reinforcement Learning (RL) and Offline Multi-Agent Reinforcement Learning (MARL) critically depends on the availability of high-quality, pre-collected offline datasets that represent real-world complexities and practical applications. However, existing datasets often fall…

2023

HyPoradise: An Open Baseline for Generative Speech Recognition with Large Language Models

NeurIPS 2023poster

Advancements in deep neural networks have allowed automatic speech recognition (ASR) systems to attain human parity on several publicly available clean speech datasets. However, even state-of-the-art ASR systems experience performance degradation when confronted with adverse conditions, as a well-tr…

2023

Ideology Prediction from Scarce and Biased Supervision: Learn to Disregard the “What” and Focus on the “How”!

ACL 2023long

We propose a novel supervised learning approach for political ideology prediction (PIP) that is capable of predicting out-of-distribution inputs. This problem is motivated by the fact that manual data-labeling is expensive, while self-reported labels are often scarce and exhibit significant selectio…

Cited by 5SourcePDFScholar
2023

Is Heterogeneity Notorious? Taming Heterogeneity to Handle Test-Time Shift in Federated Learning

NeurIPS 2023poster

Federated learning (FL) is an effective machine learning paradigm where multiple clients can train models based on heterogeneous data in a decentralized manner without accessing their private data. However, existing FL systems undergo performance deterioration due to feature-level test-time shifts,…

Cited by 26SourcePDFScholar