← Search

Meng Wang

163 accepted papers

2026

A Theoretical Analysis of Mamba’s Training Dynamics: Filtering Relevant Features for Generalization in State Space Models

ICLR 2026poster

The recent empirical success of Mamba and other selective state space models (SSMs) has renewed interest in non-attention architectures for sequence modeling, yet their theoretical foundations remain underexplored. We present a first-step analysis of generalization and learning dynamics for a simpli…

Cited by 3SourceScholar
2026

AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection

AAAI 2026technical

Universal visual anomaly detection aims to identify anomalies from novel or unseen vision domains without additional fine-tuning, which is critical in open scenarios. Recent studies have demonstrated that pre-trained vision-language models like CLIP exhibit strong generalization with just zero or a

Cited by 0SourcePDFScholar
2026

Authorize-on-Demand: Dynamic Authorization with Legality-Aware Intellectual Property Protection for VLMs

CVPR 2026

The rapid adoption of vision-language models (VLMs) has heightened the demand for robust intellectual property (IP) protection of these high-value pretrained models. Effective IP protection should proactively confine model deployment within authorized domains and prevent unauthorized transfers. Howe

Cited by 0SourcecodeScholar
2026

Benchmarking Trustworthiness in Multimodal LLMs for Video Understanding

AAAI 2026technical

Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability.

Cited by 0SourcePDFScholar
2026

Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training

CVPR 2026

Recent advances in video generation models have achieved impressive results. However, these models heavily rely on the use of high-quality data that combines both high visual quality and high motion quality. In this paper, we identify a key challenge in video data curation: the Motion-Vision Quality

Cited by 0SourceScholar
2026

Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees

ICLR 2026poster

Sparse Mixture-of-Experts (MoE) allows scaling of language and vision models efficiently by activating only a small subset of experts per input. While this reduces computation, the large number of parameters still incurs substantial memory overhead during inference. Post-training quantization has be…

Cited by 0SourcecodeScholar
2026

Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition

CVPR 2026

Recently, masked skeleton reconstruction models have emerged as strong action representation learners, driving significant progress in self-supervised skeleton-based action recognition. However, existing state-of-the-art methods must predict an exceedingly large number of spatiotemporal patches, sig

Cited by 0SourcecodeScholar
2026

FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion

AAAI 2026technical

Current video generation models perform well at single-shot synthesis but struggle with multi-shot videos, facing critical challenges in maintaining character and background consistency across shots and flexibly generating videos of arbitrary length and shot count. To address these limitations, we i

Cited by 0SourcePDFScholar
2026

GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping

CVPR 2026

Recently, GRPO-based reinforcement learning has shown remarkable progress in optimizing flow-matching models, effectively improving their alignment with task-specific rewards. Within these frameworks, the policy update relies on importance-ratio clipping to constrain overconfident positive and negat

Cited by 0SourcecodeScholar
2026

GeCo: Geometry-Consistent Regularization for Domain Generalized Semantic Segmentation

CVPR 2026

Vision Foundation Models (VFMs) provide rich and transferable representations through large-scale pretraining, yet their high-capacity representations remain underutilized when adapted to downstream tasks. In Domain Generalization Semantic Segmentation (DGSS), parameter-efficient fine-tuning (PEFT)

Cited by 0SourcecodeScholar
2026

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

CVPR 2026

While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominantly focus on models without safety alignment. This critical oversight ignores the fact that real-world MLLMs inherently require such mechanisms to mitig

Cited by 0SourcecodeScholar
2026

How Can Mamba Learn In Context with Outliers and Generalize Provably?

ICML 2026poster

The Mamba model has gained significant attention for its computational advantages over Transformer-based models, while achieving comparable performance across a wide range of language tasks. Like Transformers, Mamba exhibits in-context learning (ICL) capabilities, i.e., making predictions for new ta…

Cited by 0SourceScholar
2026

Layer Consistency Matters: Elegant Latent Transition Discrepancy for Generalizable Synthetic Image Detection

CVPR 2026

Recent rapid advancement of generative models has significantly improved the fidelity and accessibility of AI-generated synthetic images. While enabling various innovative applications, the unprecedented realism of these synthetics makes them increasingly indistinguishable from authentic photographs

Cited by 0SourcecodeScholar
2026

Learning Spatial-Temporal Consistency for 3D Semantic Scene Completion

CVPR 2026

Camera-based Semantic Scene Completion (SSC) is able to comprehensively understand the entire scene, but it suffers from ambiguous predictions due to occlusions and incomplete information. Temporal SSC alleviates this issue, but existing models simply stack multi-frame temporal features, which can l

Cited by 0SourceScholar
2026

Learning-Based Observer for Coupled Disturbance

ICRA 2026poster

Achieving high-precision control for robotic systems is hindered by the low-fidelity dynamical model and external disturbances. Especially, the intricate coupling between internal uncertainties and external disturbances further exacerbates this challenge. This study introduces an effective and conve…

2026

Lookahead-GCG: Improving Multi-Model Gradient-Based Jailbreaking Attacks via Nesterov Momentum

ICML 2026poster

Transferable jailbreaking attacks enable red-teaming of black-box large language models by optimizing adversarial prompts on open-source surrogates. A natural approach to improve transferability is multi-model training---optimizing against multiple source models simultaneously. Yet this approach has…

Cited by 0SourceScholar
2026

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory

CVPR 2026

In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation.Naively applying MLLMs leads to unreliable navigation due to weak directional grounding and the lack of explicit spatial memory.To address the

Cited by 0SourcecodeScholar
2026

PureCC: Pure Learning for Text-to-Image Concept Customization

CVPR 2026

Existing concept customization methods have achieved remarkable outcomes in high-fidelity and multi-concept customization. However, they often neglect the influence on the original model's behavior and capabilities when learning new personalized concepts. To address this issue, we propose PureCC. Pu

Cited by 0SourcecodeScholar
2026

RecCocktail: A Generalizable and Efficient Framework for LLM-Based Recommendation

AAAI 2026technical

Large Language Models (LLMs) have achieved remarkable success in recent years, owing to their impressive generalization capabilities and rich world knowledge. To capitalize on the potential of using LLMs as recommender systems, mainstream approaches typically focus on two paradigms. The first paradi

Cited by 0SourcePDFScholar
2026

SAM 3: Segment Anything with Concepts

ICLR 2026poster

We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short noun phrases (e.g., “yellow school bus”), image exemplars, or a combination of both. Promptable Concept Segmentation (P…

Cited by 687SourcecodeScholar
2026

See the Emotion: A Facial Emoji Proxy Modeling for EEG Emotion Recognition

ICML 2026poster

Despite the high accuracy of EEG-based emotion recognition, existing models remain opaque "black boxes", lacking semantic grounding between abstract neural features and human-interpretable states. In this paper, we reframe EEG explainability as a cross-modal generation task, shifting the paradigm fr…

Cited by 0SourceScholar
2026

Semantic Alignment of Malicious Question Based on Contrastive Semantic Networks and Data Augmentation (Abstract Reprint)

AAAI 2026technical

The identification and filtration of malicious texts in social media environments represent a significant technical challenge aimed at protecting users from online violence and disinformation. This complexity stems from the diversity and innovativeness of social media texts, which include unique exp

Cited by 0SourcePDFScholar
2026

Sparse-Scale Transformer with Bidirectional Awareness for Time Series Forecasting

AAAI 2026technical

Time series forecasting (TSF) plays a crucial role in many real-world applications, such as weather prediction and economic planning. While Transformer-based models have shown strong capabilities in modeling long-range dependencies, effectively capturing the multi-scale temporal dynamics inherent in

Cited by 0SourcePDFScholar
2026

TacTape: Real-Time High-Accuracy Tactile Fiducial System with Structured 3D Texture for Vision-Based Tactile Sensors

ICRA 2026poster

Vision-based tactile sensors enable high-resolution tactile perception by capturing image-based contact data. However, their utility in tactile localization is limited by their inherently small and local sensing area, as well as their dependence on distinct object surface features. We propose TacTap…

Cited by 0Scholar
2026

Theoretical Analysis of Contrastive Learning under Imbalanced Data: From Training Dynamics to a Pruning Solution

ICLR 2026poster

Contrastive learning has emerged as a powerful framework for learning generalizable representations, yet its theoretical understanding remains limited, particularly under imbalanced data distributions that are prevalent in real-world applications. Such an imbalance can degrade representation quality…

Cited by 0SourceScholar
2026

Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model

CVPR 2026

Recent advances in video reward models and post-training strategies have improved text-to-video (T2V) generation. While these models typically assess visual quality, motion quality, and text alignment, they often overlook key structural distortions, such as abnormal object appearances and interactio

Cited by 0SourceScholar
2026

Towards Foundation Models for Zero-Shot Time Series Anomaly Detection: Leveraging Synthetic Data and Relative Context Discrepancy

ICML 2026poster

TSAD is a critical task, but developing models that generalize to unseen data in a zero-shot manner remains a major challenge. Prevailing foundation models for TSAD predominantly rely on reconstruction-based objectives, which suffer from a fundamental objective mismatch and representation conflict: …

Cited by 0SourceScholar
2026

Towards Non-Stationary Time Series Forecasting with Temporal Stabilization and Frequency Differencing

AAAI 2026technical

Time series forecasting is critical for decision making across dynamic domains such as energy, finance, transportation, and cloud computing. However, real-world time series often exhibit non-stationarity, including temporal distribution shifts and spectral variability, which poses significant challe

Cited by 0SourcePDFScholar
2026

Unified Meta-Representation and Feedback Calibration for General Disturbance Estimation

ICRA 2026poster

Precise control in modern robotic applications is always an open issue due to unknown time-varying disturbances. Existing meta-learning-based approaches require a shared representation of environmental structures, which lack flexibility for realistic non-structural disturbances. Besides, representat…

2026

Urban-GS: A Unified 3D Gaussian Splatting Framework for Compact and High-Fidelity Aerial-to-Street Reconstruction

CVPR 2026

Recently, 3D Gaussian Splatting (3DGS) has revolutionized radiance field reconstruction, enabling efficient and high-fidelity novel view synthesis. However, seamless integration of both aerial and street view images to model urban scenes remains a significant challenge for 3DGS. This joint setting s

Cited by 0SourceScholar
2025

3D Gaussian Splatting based Scene-independent Relocalization with Unidirectional and Bidirectional Feature Fusion

NeurIPS 2025poster

Visual localization is a critical component across various domains. The recent emergence of novel scene representations, such as 3D Gaussian Splatting (3D GS), introduces new opportunities for advancing localization pipelines. In this paper, we propose a novel 3D GS-based framework for RGB based, sc…

Cited by 0SourceScholar
2025

An Information-Theoretic Regularizer for Lossy Neural Image Compression

ICCV 2025poster

Lossy image compression networks aim to minimize the latent entropy of images while adhering to specific distortion constraints. However, optimizing the neural network can be challenging due to its nature of learning quantized latent representations. In this paper, our key finding is that minimizing…

Cited by 0SourcePDFScholar
2025

Bi-perspective Splitting Defense: Achieving Clean-Seed-Free Backdoor Security

ICML 2025poster

Backdoor attacks have seriously threatened deep neural networks (DNNs) by embedding concealed vulnerabilities through data poisoning. To counteract these attacks, training benign models from poisoned data garnered considerable interest from researchers. High-performing defenses often rely on additio…

Cited by 0SourcePDFScholar
2025

Boosting Adversarial Transferability via Residual Perturbation Attack

ICCV 2025poster

Deep neural networks are susceptible to adversarial examples while suffering from incorrect predictions via imperceptible perturbations. Transfer-based attacks create adversarial examples for surrogate models and transfer these examples to target models under black-box scenarios. Recent studies reve…

2025

Cognitive Bias and Reassignment: Who Can Contribute High Quality LLM Data

AAAI 2025technical

In recent years, the rapid development of Large Language Models has highlighted the urgent need for large-scale, high-quality, and diverse data. We have launched an LLM data co-creation platform aimed at bringing together a wide range of participants to contribute data. Within six months, the platfo…

Cited by 0SourcePDFScholar
2025

Contrastive Learning with Data Misalignment: Feature Purity, Training Dynamics and Theoretical Generalization Guarantees

NeurIPS 2025poster

Contrastive learning is a powerful framework for learning discriminative representations from image-text pairs. Despite its success, its theoretical foundations, especially when the image-text pair exhibits misalignment, remain underexplored. This paper provides the first theoretical analysis of c…

Cited by 0SourceScholar
2025

DistillDrive: End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning Model

ICCV 2025poster

End-to-end autonomous driving has been recently seen rapid development, exerting a profound influence on both industry and academia. However, the existing work places excessive focus on ego-vehicle status as their sole learning objectives and lacks of planning-oriented understanding, which limits th…

2025

EgoBlind: Towards Egocentric Visual Assistance for the Blind

NeurIPS 2025poster

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from the daily lives of blind and visually impaired individuals. It…

Cited by 0SourcecodeScholar
2025

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

CVPR 2025poster

We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user needs in outdoor driving and indoor house-keeping activities. The questions are d…

2025

FakeDiffer: Distributional Disparity Learning on Differentiated Reconstruction for Face Forgery Detection

AAAI 2025technical

Existing face forgery detection methods achieve promising performance when training and testing forgery data are from identical manipulation types, while they fail to generalize well to unseen samples. In this paper, we experimentally investigate and find that the poor generalization of the methods…

Cited by 0SourcePDFScholar
2025

Feedback Favors the Generalization of Neural ODEs

ICLR 2025oral

The well-known generalization problem hinders the application of artificial neural networks in continuous-time prediction tasks with varying latent dynamics. In sharp contrast, biological systems can neatly adapt to evolving environments benefiting from real-time feedback mechanisms. Inspired by the…

Cited by 1SourcePDFScholar
2025

From End-to-end to Step-by-step: Learning to Abstract via Abductive Reinforcement Learning

IJCAI 2025

Abstraction is a critical technique in general problem-solving, allowing complex tasks to be decomposed into smaller, manageable sub-tasks. While traditional symbolic planning relies on predefined primitive symbols to construct structured abstractions, its reliance on formal representations limits a

2025

GT-Mean Loss: A Simple Yet Effective Solution for Brightness Mismatch in Low-Light Image Enhancement

ICCV 2025poster

Low-light image enhancement (LLIE) aims to improve the visual quality of images captured under poor lighting conditions. In supervised LLIE tasks, there exists a significant yet often overlooked inconsistency between the overall brightness of an enhanced image and its ground truth counterpart, refer…

Cited by 0SourcePDFScholar
2025

High-Fidelity Stereoscopic Image Rain Removal with Texture Integrity and Disparity Consistency

ICASSP 2025accepted

This paper tackles the challenge of stereoscopic image rain removal by focusing on enhancing texture integrity and disparity consistency. Existing stereoscopic rain removal techniques often fall short due to 1) disruptions in texture coherence caused by complex rain streaks, and 2) inaccuracies in d…

Cited by 0SourceScholar
2025

Knowledge Swapping via Learning and Unlearning

ICML 2025poster

We introduce Knowledge Swapping, a novel task designed to selectively regulate knowledge of a pretrained model by enabling the forgetting of user-specified information, retaining essential knowledge, and acquiring new knowledge simultaneously. By delving into the analysis of knock-on feature hierar…

2025

Learning Temporal 3D Semantic Scene Completion via Optical Flow Guidance

NeurIPS 2025poster

3D Semantic Scene Completion (SSC) provides comprehensive scene geometry and semantics for autonomous driving perception, which is crucial for enabling accurate and reliable decision-making. However, existing SSC methods are limited to capturing sparse information from the current frame or naively s…

Cited by 0SourceScholar
2025

MMAD: Multi-label Micro-Action Detection in Videos

ICCV 2025poster

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising applications in human emotion analysis. In real-world scenari…

2025

MOL-Mamba: Enhancing Molecular Representation with Structural & Electronic Insights

AAAI 2025technical

Molecular representation learning plays a crucial role in various downstream tasks, such as molecular property prediction and drug design. To accurately represent molecules, Graph Neural Networks (GNNs) and Graph Transformers (GTs) have shown potential in the realm of self-supervised pretraining. Ho…

2025

Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning

ICML 2025oral

Class-incremental learning (CIL) seeks to enable a model to sequentially learn new classes while retaining knowledge of previously learned ones. Balancing flexibility and stability remains a significant challenge, particularly when the task ID is unknown. To address this, our study reveals that the…

2025

One Stone with Two Birds: A Null-Text-Null Frequency-Aware Diffusion Models for Text-Guided Image Inpainting

NeurIPS 2025poster

Text-guided image inpainting aims at reconstructing the masked regions as per text prompts, where the longstanding challenges lie in the preservation for unmasked regions, while achieving the semantics consistency between unmasked and inpainted masked regions. Previous arts failed to address both of…

Cited by 0SourcecodeScholar
2025

PhysDiff: Physiology-based Dynamicity Disentangled Diffusion Model for Remote Physiological Measurement

AAAI 2025technical

Recent works on remote PhotoPlethysmoGraphy (rPPG) estimation typically use techniques like CNNs and Transformers to encode implicit features from facial videos for prediction. These methods learn to directly map facial videos to the static values of rPPG signals, overlooking the inherent dynamic ch…

2025

Precise Localization of Memories: A Fine-grained Neuron-level Knowledge Editing Technique for LLMs

ICLR 2025poster

Knowledge editing aims to update outdated information in Large Language Models (LLMs). A representative line of study is locate-then-edit methods, which typically employ causal tracing to identify the modules responsible for recalling factual knowledge about entities. However, we find these methods…

2025

Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition

AAAI 2025technical

Micro-Action Recognition (MAR) has gained increasing attention due to its crucial role as a form of non-verbal communication in social interactions, with promising potential for applications in human communication and emotion analysis. However, current approaches often overlook the inherent ambiguit…

2025

R-Tac0: A Rounded High-Frequency Transferable Monochrome Vision-based Tactile Sensor for Shape Reconstruction

IROS 2025

Endowing the curved surfaces of rounded vision-based tactile fingers is essential for dexterous robotic manipulation, as they offer more sufficient contact with the environment. However, current rounded designs are constrained by a low sensing frequency (30–60 Hz) and the need for recalibration when

Cited by 1SourceScholar
2025

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

CVPR 2025poster

Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centric Transformer architecture, where object queries are derived from audio features. However, audio-centric Transformers su…

Cited by 0SourcePDFScholar
2025

SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning

ICCV 2025poster

Visual instruction tuning (VIT) enables multimodal large language models (MLLMs) to effectively handle a wide range of vision tasks by framing them as language-based instructions. Building on this, continual visual instruction tuning (CVIT) extends the capability of MLLMs to incrementally learn new…

2025

SuperMag: Vision-based Tactile Data Guided High-resolution Tactile Shape Reconstruction for Magnetic Tactile Sensors

IROS 2025

Magnetic-based tactile sensors (MBTS) combine the advantages of compact design and high-frequency operation but suffer from limited spatial resolution due to their sparse taxel arrays. This paper proposes SuperMag, a tactile shape reconstruction method that addresses this limitation by leveraging hi

Cited by 0SourceScholar
2025

Swift Pursuer: A Topology-Accelerated and Robust Approach for Pursuing an Evader in Obstacle Environments With State Measurement Uncertainty

RA-L 2025

This letter presents a topology-accelerated and robust pursuit framework for environments with obstacles considering state measurement uncertainty. Our framework consists of three primary components: the selection of virtual target points using topological heuristic method to encourage path diversit

Cited by 9SourceScholar
2025

TASAR: Transfer-based Attack on Skeletal Action Recognition

ICLR 2025poster

Skeletal sequence data, as a widely employed representation of human actions, are crucial in Human Activity Recognition (HAR). Recently, adversarial attacks have been proposed in this area, which exposes potential security concerns, and more importantly provides a good tool for model robustness test…

2025

Thinking in Granularity: Dynamic Quantization for Image Super-Resolution by Intriguing Multi-Granularity Clues

AAAI 2025technical

Dynamic quantization has attracted rising attention in image super-resolution (SR) as it expands the potential of heavy SR models onto mobile devices while preserving competitive performance. Most current methods explore layer-to-bit configuration upon varying local regions, adaptively allocating th…

2025

Towards Efficient General Feature Prediction in Masked Skeleton Modeling

ICCV 2025poster

Recent advances in the masked autoencoder (MAE) paradigm have significantly propelled self-supervised skeleton-based action recognition. However, most existing approaches limit reconstruction targets to raw joint coordinates or their simple variants, resulting in computational redundancy and limited…

Cited by 0SourcePDFScholar
2025

Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach

IJCAI 2025

Micro-Action Recognition (MAR) aims to classify subtle human actions in video. However, annotating MAR datasets is particularly challenging due to the subtlety of actions. To this end, we introduce the setting of Semi-Supervised MAR (SSMAR), where only a part of samples are labeled. We first evaluat

2025

Towards Open-Vocabulary Audio-Visual Event Localization

CVPR 2025poster

The Audio-Visual Event Localization (AVEL) task aims to temporally locate and classify video events that are both audible and visible.Most research in this field assumes a closed-set setting, which restricts these models' ability to handle test data containing event categories absent (unseen) during…

2025

Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis

ICLR 2025poster

Chain-of-Thought (CoT) is an efficient prompting method that enables the reasoning ability of large language models by augmenting the query using multiple examples with multiple intermediate steps. Despite the empirical success, the theoretical understanding of how to train a Transformer to achieve…

Cited by 3SourcePDFScholar
2025

VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion

AAAI 2025technical

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and perspective distortion. Existing methods often lack explici…

2025

Vision-Language Model IP Protection via Prompt-based Learning

CVPR 2025poster

Vision-language models (VLMs) like CLIP (Contrastive Language-Image Pre-Training) have seen remarkable success in visual recognition, highlighting the increasing need to safeguard the intellectual property (IP) of well-trained models. Effective IP protection extends beyond ensuring authorized usage;…

2025

Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language Models

ICCV 2025poster

Existing knowledge editing works for MultiModal Large Language Models primarily focus on text-oriented, coarse-grained scenarios, where modifying textual content alone is sufficient. As a result, they fail to capture the unique challenges of multimodal editing, particularly when visual information i…

2025

When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers

ICLR 2025oral

Task arithmetic refers to editing the pre-trained model by adding a weighted sum of task vectors, each of which is the weight update from the pre-trained model to fine-tuned models for certain tasks. This approach recently gained attention as a computationally efficient inference method for model ed…

Cited by 0SourcePDFScholar
2024

A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking

AAAI 2024technical

Multimodal Entity Linking (MEL) aims at linking ambiguous mentions with multimodal information to entity in Knowledge Graph (KG) such as Wikipedia, which plays a key role in many applications. However, existing methods suffer from shortcomings, including modality impurity such as noise in raw image…

2024

A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts

ICML 2024poster

The sparsely gated mixture of experts (MoE) architecture sends different inputs to different subnetworks (experts), through trainable routers. MoE reduces the training computation significantly for large models, but its deployment can be still memory/computation expensive for some downstream tasks.…

Cited by 3SourcePDFScholar
2024

Adaptive Group Personalization for Federated Mutual Transfer Learning

ICML 2024poster

Mutual transfer learning aims to improve prediction with knowledge from related domains. Recently, federated learning is applied in this field to address the communication and privacy concerns. However, previous clustered federated learning (CFL) solutions lack theoretical guarantee of learnability…

Cited by 0SourcePDFScholar
2024

DURRNET: Deep Unfolded Single Image Reflection Removal Network with Joint Prior

ICASSP 2024accepted

Single image reflection removal (SIRR) problem can be interpreted as a canonical blind source separation problem and is highly ill-posed. A parameter effective, fast learning and interpretable reflection removal algorithm is essential for many vision analysis applications. In this paper, we propose…

Cited by 0SourceScholar
2024

EulerMormer: Robust Eulerian Motion Magnification via Dynamic Filtering within Transformer

AAAI 2024technical

Video Motion Magnification (VMM) aims to break the resolution limit of human visual perception capability and reveal the imperceptible minor motion that contains valuable information in the macroscopic domain. However, challenges arise in this task due to photon noise inevitably introduced by photog…

2024

FasMe: Fast and Sample-efficient Meta Estimator for Precision Matrix Learning in Small Sample Settings

NeurIPS 2024poster

Precision matrix estimation is a ubiquitous task featuring numerous applications such as rare disease diagnosis and neural connectivity exploration. However, this task becomes challenging in small sample settings, where the number of samples is significantly less than the number of dimensions, leadi…

Cited by 0SourcePDFScholar
2024

Finding and Editing Multi-Modal Neurons in Pre-Trained Transformers

ACL 2024findings

Understanding the internal mechanisms by which multi-modal large language models (LLMs) interpret different modalities and integrate cross-modal representations is becoming increasingly critical for continuous improvements in both academia and industry. In this paper, we propose a novel method to id…

2024

Flight Structure Optimization of Modular Reconfigurable UAVs

IROS 2024poster

This paper presents a Genetic Algorithm (GA) designed to reconfigure a large group of modular Unmanned Aerial Vehicles (UAVs), each with different weights and inertia parameters, into an over-actuated flight structure with improved dynamic properties. Previous research efforts either utilized expert…

Cited by 9SourceScholar
2024

Frequency Decoupling for Motion Magnification via Multi-Level Isomorphic Architecture

CVPR 2024poster

Video Motion Magnification (VMM) aims to reveal subtle and imperceptible motion information of objects in the macroscopic world. Prior methods directly model the motion field from the Eulerian perspective by Representation Learning that separates shape and texture or Multi-domain Learning from phase…

2024

How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?

ICML 2024poster

Transformer-based large language models have displayed impressive in-context learning capabilities, where a pre-trained model can handle new tasks without fine-tuning by simply augmenting the query with some input-output examples from that task. Despite the empirical success, the mechanics of how to…

Cited by 14SourcePDFScholar
2024

Invertible Mosaic Image Hiding Network for Very Large Capacity Image Steganography

ICASSP 2024accepted

The existing image steganography methods either sequentially conceal secret images or conceal a concatenation of multiple images. In such ways, the interference of information among multiple images will become increasingly severe when the number of secret images becomes larger, thus restrict the dev…

Cited by 0SourceScholar
2024

KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose Tracking

AAAI 2024technical

Our life is populated with articulated objects. Current category-level articulation estimation works largely focus on predicting part-level 6D poses on static point cloud observations. In this paper, we tackle the problem of category-level online robust and real-time 6D pose tracking of articulated…

2024

Knowledge-augmented Financial Market Analysis and Report Generation

EMNLP 2024industry

Crafting a convincing financial market analysis report necessitates a wealth of market information and the expertise of financial analysts, posing a highly challenging task. While large language models (LLMs) have enabled the automated generation of financial market analysis text, they still face is…

Cited by 2SourcePDFScholar
2024

Large-scale Deployment of Vision-based Tactile Sensors on Multi-fingered Grippers

IROS 2024

Vision-based Tactile Sensors (VBTSs) show significant promise in that they can leverage image measurements to provide high-spatial-resolution human-like performance. However, current VBTS designs, typically confined to the fingertips of robotic grippers, prove somewhat inadequate, as many grasping a

Cited by 6SourceScholar
2024

Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis

CVPR 2024poster

Recent works in implicit representations such as Neural Radiance Fields (NeRF) have advanced the generation of realistic and animatable head avatars from video sequences. These implicit methods are still confronted by visual artifacts and jitters since the lack of explicit geometric constraints pose…

2024

OSIC: A New One-Stage Image Captioner Coined

IJCAI 2024poster

Mainstream image captioning models are usually two-stage captioners, i.e., encoding the region features by a pre-trained detector and then feeding them into a language model to generate the captions. However, such a two-stage procedure will lead to a task-based information gap that decreases the per…

Cited by 7SourcePDFScholar
2024

Object-Aware Adaptive-Positivity Learning for Audio-Visual Question Answering

AAAI 2024technical

This paper focuses on the Audio-Visual Question Answering (AVQA) task that aims to answer questions derived from untrimmed audible videos. To generate accurate answers, an AVQA model is expected to find the most informative audio-visual clues relevant to the given questions. In this paper, we propos…

2024

Real-time Dynamic-consistent Motion Planning for Over-actuated UAVs

ICRA 2024poster

Existing motion planning approaches for over-actuated unmanned aerial vehicle (UAV) platforms can achieve online planning without considering dynamics. However, in many envisioned application areas such as aerial manipulation, payload delivery, and moving target tracking, it is critical to ensure dy…

Cited by 4SourceScholar
2024

Revisiting the Power of Prompt for Visual Tuning

ICML 2024spotlight

Visual prompt tuning (VPT) is a promising solution incorporating learnable prompt tokens to customize pre-trained models for downstream tasks. However, VPT and its variants often encounter challenges like prompt initialization, prompt length, and subpar performance in self-supervised pretraining, hi…

2024

SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning

ICML 2024poster

This paper studies the transfer reinforcement learning (RL) problem where multiple RL problems have different reward functions but share the same underlying transition dynamics. In this setting, the Q-function of each RL problem (task) can be decomposed into a successor feature (SF) and a reward map…

Cited by 2SourcePDFScholar
2024

StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models

ECCV 2024poster

"Despite the burst of innovative methods for controlling the diffusion process, effectively controlling image styles in text-to-image generation remains a challenging task. Many adapter-based methods impose image representation conditions on the denoising process to accomplish image control. However…

2024

Temporal Sentence Grounding with Relevance Feedback in Videos

NeurIPS 2024poster

As a widely explored multi-modal task, Temporal Sentence Grounding in videos (TSG) endeavors to retrieve a specific video segment matched with a given query text from a video. The traditional paradigm for TSG generally assumes that relevant segments always exist within a given video. However, this a…

2024

Towards Proactive Interactions for In-Vehicle Conversational Assistants Utilizing Large Language Models

IJCAI 2024poster

Research demonstrates that the proactivity of in-vehicle conversational assistants (IVCAs) can help to reduce distractions and enhance driving safety, better meeting users' cognitive needs. However, existing IVCAs struggle with user intent recognition and context awareness, which leads to suboptimal…

2024

Training A Small Emotional Vision Language Model for Visual Art Comprehension

ECCV 2024poster

"This paper develops small vision language models to understand visual art, which, given an art work, aims to identify its emotion category and explain this prediction with natural language. While small models are computationally efficient, their capacity is much limited compared with large models.…

2024

Variance Reduction Can Improve Trade-Off in Multi-Objective Learning

ICASSP 2024accepted

Many machine learning problems today have multiple objective functions, which are often tackled by the multi-objective learning (MOL) framework. Albeit many encouraging results are obtained by MOL algorithms, a recent theoretical study [1] revealed that these gradient-based MOL methods (e.g., MGDA,…

Cited by 0SourceScholar
2024

What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding

ICML 2024poster

Graph Transformers, which incorporate self-attention and positional encoding, have recently emerged as a powerful architecture for various graph learning tasks. Despite their impressive performance, the complex non-convex interactions across layers and the recursive graph structure have made it chal…

Cited by 16SourcePDFScholar
2023

A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity

ICLR 2023poster

Vision Transformers (ViTs) with self-attention modules have recently achieved great empirical success in many vision tasks. Due to non-convex interactions across layers, however, the theoretical learning and generalization analysis is mostly elusive. Based on a data model characterizing both label-r…

Cited by 83SourcePDFScholar
2023

Aggregating Single-Wheeled Mobile Robots for Omnidirectional Movements

IROS 2023poster

This paper presents a novel modular robot system that can self-reconfigure to achieve omnidirectional movements for collaborative object transportation. Each robotic module is equipped with a steerable omni-wheel for navigation and is shaped as a regular icositetragon with a permanent magnet install…

Cited by 2SourceScholar
2023

DC-Former: Diverse and Compact Transformer for Person Re-identification

AAAI 2023technical

In person re-identification (ReID) task, it is still challenging to learn discriminative representation by deep learning, due to limited data. Generally speaking, the model will get better performance when increasing the amount of data. The addition of similar classes strengthens the ability of the…

2023

Disentangling Cognitive Diagnosis with Limited Exercise Labels

NeurIPS 2023poster

Cognitive diagnosis is an important task in intelligence education, which aims at measuring students’ proficiency in specific knowledge concepts. Given a fully labeled exercise-concept matrix, most existing models focused on mining students' response records for cognitive diagnosis. Despite their su…

2023

Domain Generalized Stereo Matching via Hierarchical Visual Transformation

CVPR 2023poster

Recently, deep Stereo Matching (SM) networks have shown impressive performance and attracted increasing attention in computer vision. However, existing deep SM networks are prone to learn dataset-dependent shortcuts, which fail to generalize well on unseen realistic datasets. This paper takes a step…

Cited by 26SourcePDFScholar
2023

Fair Representation Learning for Recommendation: A Mutual Information Perspective

AAAI 2023technical

Recommender systems have been widely used in recent years. By exploiting historical user-item interactions, recommender systems can model personalized potential interests of users and have been widely applied to a wide range of scenarios. Despite their impressive performance, most of them may be sub…

Cited by 20SourcePDFScholar
2023

Fine-Grained Audible Video Description

CVPR 2023poster

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of each object, the actions of moving objects, and the sounds i…

2023

Joint Edge-Model Sparse Learning is Provably Efficient for Graph Neural Networks

ICLR 2023poster

Due to the significant computational challenge of training large-scale graph neural networks (GNNs), various sparse learning techniques have been exploited to reduce memory and storage costs. Examples include graph sparsification that samples a subgraph to reduce the amount of data aggregation and m…

Cited by 20SourcePDFScholar
2023

L${3}$ F-TOUCH: A Wireless GelSight With Decoupled Tactile and Three-Axis Force Sensing

RA-L 2023

GelSight sensors that estimate contact geometry and force by reconstructing the deformation of their soft elastomer from images would yield poor force measurements when the elastomer deforms uniformly or reaches deformation saturation. Here we present an L <inline-formula xmlns:mml="http://www.w3.or

Cited by 37SourceScholar
2023

LP-DIF: Learning Local Pattern-Specific Deep Implicit Function for 3D Objects and Scenes

CVPR 2023poster

Deep Implicit Function (DIF) has gained much popularity as an efficient 3D shape representation. To capture geometry details, current mainstream methods divide 3D shapes into local regions and then learn each one with a local latent code via a decoder, where the decoder shares the geometric similari…

2023

MCL: Multi-Granularity Contrastive Learning Framework for Chinese NER

AAAI 2023technical

Recently, researchers have applied the word-character lattice framework to integrated word information, which has become very popular for Chinese named entity recognition (NER). However, prior approaches fuse word information by different variants of encoders such as Lattice LSTM or Flat-Lattice…

2023

Model Barrier: A Compact Un-Transferable Isolation Domain for Model Intellectual Property Protection

CVPR 2023poster

As the scientific and technological achievements produced by human intellectual labor and computation cost, model intellectual property (IP) protection, which refers to preventing the usage of the well-trained model on an unauthorized domain, deserves further attention, so as to effectively mobilize…

2023

On the Convergence and Sample Complexity Analysis of Deep Q-Networks with $\epsilon$-Greedy Exploration

NeurIPS 2023poster

This paper provides a theoretical understanding of deep Q-Network (DQN) with the $\varepsilon$-greedy exploration in deep reinforcement learning. Despite the tremendous empirical achievement of the DQN, its theoretical characterization remains underexplored. First, the exploration strategy is either…

Cited by 27SourcePDFScholar
2023

Patch-level Routing in Mixture-of-Experts is Provably Sample-efficient for Convolutional Neural Networks

ICML 2023oral

In deep learning, mixture-of-experts (MoE) activates one or few experts (sub-networks) on a per-sample or per-token basis, resulting in significant computation reduction. The recently proposed patch-level routing in MoE (pMoE) divides each input into $n$ patches (or tokens) and sends $l$ patches ($l…

2023

Prompting Large Language Models With Answer Heuristics for Knowledge-Based Visual Question Answering

CVPR 2023poster

Knowledge-based visual question answering (VQA) requires external knowledge beyond the image to answer the question. Early studies retrieve required knowledge from explicit knowledge bases (KBs), which often introduces irrelevant information to the question, hence restricting the performance of thei…

2023

Sequential Manipulation Planning for Over-Actuated Unmanned Aerial Manipulators

IROS 2023poster

We investigate the sequential manipulation planning problem for unmanned aerial manipulators (UAMs). Unlike prior work that primarily focuses on one-step manipulation tasks, sequential manipulations require coordinated motions of a UAM's floating base, the manipulator, and the object being manipulat…

Cited by 17SourceScholar
2023

Towards Efficient Pre-Trained Language Model via Feature Correlation Distillation

NeurIPS 2023poster

Knowledge Distillation (KD) has emerged as a promising approach for compressing large Pre-trained Language Models (PLMs). The performance of KD relies on how to effectively formulate and transfer the knowledge from the teacher model to the student model. Prior arts mainly focus on directly aligning…

Cited by 4SourcePDFScholar
2022

Audio—Visual Segmentation

ECCV 2022poster

"We propose to explore a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first audio-visual segmentation benchmark (AVSBench), provid…

2022

Deep Color Consistent Network for Low-Light Image Enhancement

CVPR 2022poster

Low-light image enhancement focus on refining the illumination and keep naturalness to obtain the normal-light image. Current low-light image enhancement methods can well improve the illumination. However, there is still color difference between the enhanced image and the ground-truth image. To alle…

Cited by 165PDFcodeScholar
2022

Downwash-aware Control Allocation for Over-actuated UAV Platforms

IROS 2022poster

Tracking position and orientation independently affords more agile maneuver for over-actuated multirotor Unmanned Aerial Vehicles (UAVs) while introducing undesired downwash effects; downwash flows generated by thrust generators may counteract others due to close proximity, which significantly threa…

Cited by 17SourceScholar
2022

Generalization Guarantee of Training Graph Convolutional Networks with Graph Topology Sampling

ICML 2022spotlight

Graph convolutional networks (GCNs) have recently achieved great empirical success in learning graph-structured data. To address its scalability issue due to the recursive embedding of neighboring features, graph topology sampling has been proposed to reduce the memory and computational cost of trai…

Cited by 31SourcePDFScholar
2022

How unlabeled data improve generalization in self-training? A one-hidden-layer theoretical analysis

ICLR 2022poster

Self-training, a semi-supervised learning algorithm, leverages a large amount of unlabeled data to improve learning when the labeled data are limited. Despite empirical successes, its theoretical characterization remains elusive. To the best of our knowledge, this work establishes the first theoreti…

Cited by 32SourcePDFScholar
2022

Multi-modal Contrastive Representation Learning for Entity Alignment

COLING 2022main

Multi-modal entity alignment aims to identify equivalent entities between two different multi-modal knowledge graphs, which consist of structural triples and images associated with entities. Most previous works focus on how to utilize and encode information from different modalities, while it is not…

2021

Discrimination-Aware Mechanism for Fine-Grained Representation Learning

CVPR 2021poster

Recently, with the emergence of retrieval requirements for certain individual in the same superclass, e.g., birds, persons, cars, fine-grained recognition task has attracted a significant amount of attention from academia and industry. In fine-grained recognition scenario, the inter-class difference…

Cited by 24PDFScholar
2021

Leveraging Table Content for Zero-shot Text-to-SQL with Meta-Learning

AAAI 2021technical

Single-table text-to-SQL aims to transform a natural language question into a SQL query according to one single table. Recent work has made promising progress on this task by pre-trained language models and a multi-submodule framework. However, zero-shot table, that is, the invisible table in the t…

2021

Making the Relation Matters: Relation of Relation Learning Network for Sentence Semantic Matching

AAAI 2021technical

Sentence semantic matching is one of the fundamental tasks in natural language processing, which requires an agent to determine the semantic relation among input sentences. Recently, deep neural networks have achieved impressive performance in this area, especially BERT. Despite the effectiveness of…

2021

Motion Prediction Using Trajectory Cues

ICCV 2021poster

Predicting human motion from a historical pose sequence is at the core of many applications in computer vision. Current state-of-the-art methods concentrate on learning motion contexts in the pose space, however, the high dimensionality and complex nature of human pose invoke inherent difficulties i…

Cited by 64PDFcodeScholar
2021

On Fast Adversarial Robustness Adaptation in Model-Agnostic Meta-Learning

ICLR 2021poster

Model-agnostic meta-learning (MAML) has emerged as one of the most successful meta-learning techniques in few-shot learning. It enables us to learn a $\textit{meta-initialization}$ of model parameters (that we call $\textit{meta-model}$) to rapidly adapt to new tasks using a small amount of labeled…

2021

Partial-Label and Structure-constrained Deep Coupled Factorization Network

AAAI 2021technical

In this paper, we technically propose an enriched prior guided framework, called Dual-constrained Deep Semi-Supervised Coupled Factorization Network (DS2CF-Net), for discovering hierarchical coupled data representation. To extract hidden deep features, DS2CF-Net is formulated as a partial-label and…

Cited by 6SourcePDFScholar
2021

Positive Sample Propagation Along the Audio-Visual Event Line

CVPR 2021poster

Visual and audio signals often coexist in natural environments, forming audio-visual events (AVEs). Given a video, we aim to localize video segments containing an AVE and identify its category. In order to learn discriminative features for a classifier, it is pivotal to identify the helpful (or posi…

Cited by 127PDFcodeScholar
2021

Positive-Congruent Training: Towards Regression-Free Model Updates

CVPR 2021poster

Reducing inconsistencies in the behavior of different versions of an AI system can be as important in practice as reducing its overall error. In image classification, sample-wise inconsistencies appear as "negative flips": A new model incorrectly predicts the output for a test sample that was correc…

Cited by 63PDFScholar
2021

Single View Physical Distance Estimation Using Human Pose

ICCV 2021poster

We propose a fully automated system that simultaneously estimates the camera intrinsics, the ground plane, and physical distances between people from a single RGB image or video captured by a camera viewing a 3-D scene from a fixed vantage point. To automate camera calibration and distance estimatio…

Cited by 11PDFScholar
2021

Why Lottery Ticket Wins? A Theoretical Perspective of Sample Complexity on Sparse Neural Networks

NeurIPS 2021poster

The lottery ticket hypothesis (LTH) states that learning on a properly pruned network (the winning ticket) has improved test accuracy over the original unpruned network. Although LTH has been justified empirically in a broad range of deep neural network (DNN) involved applications like computer visi…

Cited by 38SourcePDFScholar
2020

A Multi-Scaled Receptive Field Learning Approach for Medical Image Segmentation

ICASSP 2020accepted

Biomedical image segmentation has been widely studied, and lots of methods have been proposed. Among these methods, attention U-Net has achieved a promising performance. However, it has drawbacks of extracting the multi-scaled receptive field features at the high-level feature maps, resulting in the…

Cited by 0SourceScholar
2020

Detail-recovery Image Deraining via Context Aggregation Networks

CVPR 2020poster

This paper looks at this intriguing question: are single images with their details lost during deraining, reversible to their artifact-free status? We propose an end-to-end detail-recovery image deraining network (termed a DRDNet) to solve the problem. Unlike existing image deraining approaches that…

Cited by 217PDFcodeScholar
2020

Enhanced Blind Face Restoration With Multi-Exemplar Images and Adaptive Spatial Feature Fusion

CVPR 2020oral

In many real-world face restoration applications, e.g., smartphone photo albums and old films, multiple high-quality (HQ) images of the same person usually are available for a given degraded low-quality (LQ) observation. However, most existing guided face restoration methods are based on single HQ e…

Cited by 129PDFcodeScholar
2020

Fast Learning of Graph Neural Networks with Guaranteed Generalizability: One-hidden-layer Case

ICML 2020poster

Although graph neural networks (GNNs) have made great progress recently on learning from graph-structured data in practice, their theoretical guarantee on generalizability remains elusive in the literature. In this paper, we provide a theoretically-grounded generalizability analysis of GNNs with one…

Cited by 39SourcePDFScholar
2020

Large-Scale Few-Shot Learning via Multi-Modal Knowledge Discovery

ECCV 2020poster

Large-scale few-shot learning aims at identifying hundreds of novel object categories where each category has only a few samples. It is a challenging problem since (1) the identifying process is susceptible to over-fitting with limited samples of an object, and (2) the sample imbalance between a bas…

Cited by 43SourcePDFScholar
2020

Learning to Discretely Compose Reasoning Module Networks for Video Captioning

IJCAI 2020poster

Generating natural language descriptions for videos, i.e., video captioning, essentially requires step-by-step reasoning along the generation process. For example, to generate the sentence “a man is shooting a basketball”, we need to first locate and describe the subject “man”, next reason out the m…

2020

More Grounded Image Captioning by Distilling Image-Text Matching Model

CVPR 2020poster

Visual attention not only improves the performance of image captioners, but also serves as a visual interpretation to qualitatively measure the caption rationality and model transparency. Specifically, we expect that a captioner can fix its attentive gaze on the correct objects while generating the…

Cited by 179PDFcodeScholar
2020

Multi-Scale Spatial-Temporal Integration Convolutional Tube for Human Action Recognition

IJCAI 2020poster

Applying multi-scale representations leads to consistent performance improvements on a wide range of image recognition tasks. However, with the addition of the temporal dimension in video domain, directly obtaining layer-wise multi-scale spatial-temporal features will add a lot extra computational c…

Cited by 0SourcePDFScholar
2020

Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases

ECCV 2020poster

When the training data are maliciously tampered, the predictions of the acquired deep neural network (DNN) can be manipulated by an adversary known as the Trojan attack (or poisoning backdoor attack). The lack of robustness of DNNs against Trojan attacks could significantly harm real-life machine le…

2020

Quadratic Sparse Gaussian Graphical Model Estimation Method for Massive Variables

IJCAI 2020poster

We consider the problem of estimating a sparse Gaussian Graphical Model with a special graph topological structure and more than a million variables. Most previous scalable estimators still contain expensive calculation steps (e.g., matrix inversion or Hessian matrix calculation) and become infeasib…

Cited by 0SourcePDFScholar
2020

Unsupervised Vehicle Re-identification with Progressive Adaptation

IJCAI 2020poster

Vehicle re-identification (reID) aims at identifying vehicles across different non-overlapping cameras views. The existing methods heavily relied on well-labeled datasets for ideal performance, which inevitably causes fateful drop due to the severe domain bias between the training domain and the rea…

Cited by 0SourcePDFScholar
2019

Adaptive Transfer Network for Cross-Domain Person Re-Identification

CVPR 2019poster

Recent deep learning based person re-identification approaches have steadily improved the performance for benchmarks, however they often fail to generalize well from one domain to another. In this work, we propose a novel adaptive transfer network (ATNet) for effective cross-domain person re-identif…

Cited by 350PDFScholar
2019

ClusterNet: Deep Hierarchical Cluster Network With Rigorously Rotation-Invariant Representation for Point Cloud Analysis

CVPR 2019poster

Current neural networks for 3D object recognition are vulnerable to 3D rotation. Existing works mostly rely on massive amounts of rotation-augmented data to alleviate the problem, which lacks solid guarantee of the 3D rotation invariance. In this paper, we address the issue by introducing a novel po…

Cited by 217PDFScholar
2019

Graphonomy: Universal Human Parsing via Graph Transfer Learning

CVPR 2019poster

Prior highly-tuned human parsing models tend to fit towards each dataset in a specific domain or with discrepant label granularity, and can hardly be adapted to other human parsing tasks without extensive re-training. In this paper, we aim to learn a single universal human parsing model that can tac…

Cited by 228PDFcodeScholar
2019

Robust Unsupervised Flexible Auto-weighted Local-coordinate Concept Factorization for Image Clustering

ICASSP 2019accepted

We investigate the high-dimensional data clustering problem by proposing a novel and unsupervised representation learning model called Robust Flexible Auto-weighted Local-coordinate Concept Factorization (RFA-LCF). RFA-LCF integrates the robust flexible CF, robust sparse local-coordinate coding and…

Cited by 0SourceScholar
2018

Multi-Cue Correlation Filters for Robust Visual Tracking

CVPR 2018poster

In recent years, many tracking algorithms achieve impressive performance via fusing multiple types of features, however, most of them fail to fully explore the context among the adopted multiple features and the strength of them. In this paper, we propose an efficient multi-cue analysis framework fo…

2016

Impedance control of a cable-driven series elastic actuator with the 2-DOF control structure

IROS 2016poster

Series elastic actuators (SEAs) are growingly important in physical human-robot interaction (HRI) due to their inherent safety and compliance. Cable-driven SEAs also allow flexible installation and remote torque transmission, etc. However, there are still challenges for the impedance control of cabl…

Cited by 7SourceScholar
2016

Nonlinear disturbance observer based torque control for series elastic actuator

IROS 2016poster

This paper presents a practical control approach for series elastic actuators(SEAs) to generate the desired torque. Specifically, the controller is applicable to both linear and nonlinear SEAs and it works well even in the presence of unknown payload parameters and external disturbances. Via the ana…

Cited by 11SourceScholar
2015

Interaction Part Mining: A Mid-Level Approach for Fine-Grained Action Recognition

CVPR 2015poster

Modeling human-object interactions and manipulating motions lies in the heart of fine-grained action recognition. Previous methods heavily rely on explicit detection of the object being interacted, which requires intensive human labour on object annotation. To bypass this constraint and achieve bett…

Cited by 103SourcePDFScholar