← Search

Jingjing Chen

45 accepted papers

2026

Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward

AAAI 2026technical

Existing reinforcement learning (RL) methods struggle with long-horizon robotic manipulation tasks, particularly those involving sparse rewards. While action chunking is a promising paradigm for robotic manipulation, using RL to directly learn continuous action chunks in a stable and data-efficient

Cited by 7SourcePDFScholar
2026

Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferences

CVPR 2026

Narrative image generation aims to create images featuring multiple distinct characters while capturing their interrelationships, posing significant challenges for current text-to-image diffusion models. As a result, general personalized methods often suffer from poor semantic alignment, identity bl

Cited by 0SourceScholar
2026

Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Prediction

CVPR 2026

The research frontier in human pose prediction (HPP) is advancing toward continual test-time adaptation (TTA), where models must self-adapt to dynamic test distributions. To date, the homeostatic continual TTA remains the sole viable solution, which isolates the model parameters and update domain-se

Cited by 0SourceScholar
2026

Copyright Infringement Detection in Text-to-Image Diffusion Models via Differential Privacy

AAAI 2026technical

The widespread deployment of large vision models such as Stable Diffusion raises significant legal and ethical concerns, as these models can memorize and reproduce copyrighted content without authorization. Existing detection approaches often lack robustness and fail to provide rigorous theoretical

Cited by 0SourcePDFScholar
2026

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

ICML 2026poster

Audio-Visual Large Language Models (AV-LLMs) grapple with the prohibitive computational costs of processing massive, redundant audio and video tokens. Existing unimodal compression techniques fail to capture the heterogeneous and mutually influential information density of joint audio-visual signals…

Cited by 0SourceScholar
2026

Force Policy: Learning Hybrid Force-Position Control Policy under Interaction Frame for Contact-Rich Manipulation

RSS 2026poster

Contact-rich manipulation demands human-like integration of perception and force feedback: vision should guide task progress, while high-frequency interaction control must stabilize contact under uncertainty. Existing learning-based policies often entangle these roles in a monolithic network, tradin…

Cited by 0SourceScholar
2026

FreeInpaint: Tuning-free Prompt Alignment and Visual Rationality Enhancement in Image Inpainting

AAAI 2026technical

Text-guided image inpainting endeavors to generate new content within specified regions of images using textual prompts from users. The primary challenge is to accurately align the inpainted areas with the user-provided prompts while maintaining a high degree of visual fidelity. While existing inpai

Cited by 0SourcePDFScholar
2026

GrOCE : Graph-Guided Online Concept Erasure for Text-to-Image Diffusion Models

CVPR 2026

Concept erasure aims to remove harmful, inappropriate, or copyrighted content from text-to-image diffusion models while preserving non-target semantics. However, existing methods either rely on costly fine-tuning or apply coarse semantic separation, often degrading unrelated concepts and lacking ada

Cited by 0SourcecodeScholar
2026

History-Aware Visuomotor Policy Learning Via Point Tracking

ICRA 2026poster

Many manipulation tasks require memory beyond the current observation, yet most visuomotor policies rely on the Markov assumption and thus struggle with repeated states or long-horizon dependencies. Existing methods attempt to extend observation horizons but remain insufficient for diverse memory re…

2026

Identity-Aware Vision-Language Model for Explainable Face Forgery Detection

AAAI 2026technical

Recent advances in generative artificial intelligence have enabled the creation of highly realistic image forgeries, raising significant concerns about digital media authenticity. While existing detection methods demonstrate promising results on benchmark datasets, they face critical limitations in

Cited by 0SourcePDFScholar
2026

Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning

ICLR 2026poster

Large Language Models (LLMs) are widely used as judges to evaluate response quality, providing a scalable alternative to human evaluation. However, most LLM judges operate solely on intrinsic text-based reasoning, limiting their ability to verify complex constraints or perform accurate computation.…

Cited by 0SourceScholar
2026

Learning Dexterous Manipulation with Quantized Hand State

ICRA 2026poster

Dexterous robotic hands enable robots to perform complex manipulations that require fine-grained control and adaptability. Achieving such manipulation is challenging because the high degrees of freedom tightly couple hand and arm motions, making learning and control difficult. Successful dexterous m…

2026

Object Fusion via Diffusion Time-step for Customized Image Editing with Single Example

AAAI 2026technical

We tackle the task of customized image editing using a text-conditioned Diffusion Model (DM). The goal is to fuse the subject in a reference image (e.g., sunglasses) with a source one (e.g., a boy), while retaining the fidelity of them both (e.g., the boy wearing the sunglasses). An intuitive approa

Cited by 0SourcePDFScholar
2026

Predicting Future KV Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction

ICML 2026poster

Given the quadratic complexity of attention, KV cache eviction is vital to accelerate model inference. Current KV cache eviction methods typically rely on instantaneous heuristic metrics, implicitly assuming that score magnitudes are consistent proxies for importance across all heads. However, this …

Cited by 0SourceScholar
2026

RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation

CVPR 2026

Recent advances in Vision-Language-Action (VLA) models have enabled robots to execute increasingly complex tasks. However, VLA models trained through imitation learning struggle to operate reliably in dynamic environments and often fail under Out-of-Distribution (OOD) conditions. To address this iss

Cited by 0SourceScholar
2025

AirExo-2: Scaling up Generalizable Robotic Imitation Learning with Low-Cost Exoskeletons

CoRL 2025oral

Scaling up robotic imitation learning for real-world applications requires efficient and scalable demonstration collection methods. While teleoperation is effective, it depends on costly and inflexible robot platforms. In-the-wild demonstrations offer a promising alternative, but existing collection…

Cited by 0SourceScholar
2025

DLPO: Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Learning Perspective

EMNLP 2025

Large Language Models (LLMs) have achieved remarkable success across diverse tasks, largely driven by well-designed prompts. However, crafting and selecting such prompts often requires considerable human effort, significantly limiting its scalability. To mitigate this, recent studies have explored a

2025

DuMo: Dual Encoder Modulation Network for Precise Concept Erasure

AAAI 2025technical

The exceptional generative capability of text-to-image models has raised substantial safety concerns regarding the generation of Not-Safe-For-Work (NSFW) content and potential copyright infringement. To address these concerns, previous methods safeguard the models by eliminating inappropriate concep…

2025

FaceA-Net: Facial Attribute-Driven ID Preserving Image Generation Network

AAAI 2025technical

Recent advances in diffusion-based generative models have demonstrated superior performance in subject-driven image generation. Identity (ID) preserving image generation, as a subtask of subject-driven image generation, aims to generate customized images for specific human identity and has broad app…

Cited by 0SourcePDFScholar
2025

FoAR: Force-Aware Reactive Policy for Contact-Rich Robotic Manipulation

RA-L 2025

Contact-rich tasks present significant challenges for robotic manipulation policies due to the complex dynamics of contact and the need for precise control. Vision-based policies often struggle with the skill required for such tasks, as they typically lack critical contact feedback modalities like f

Cited by 40SourceScholar
2025

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning

ICCV 2025poster

Efficient Visual Instruction Fine-Tuning (EVIT) seeks to adapt Multimodal Large Language Models (MLLMs) to downstream tasks with minimal computational overhead. However, as task diversity and complexity increase, EVIT faces significant challenges in resolving data conflicts. To address this limitati…

2025

TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Models

CVPR 2025poster

Large pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated excellent zero-shot generalizability across various downstream tasks. However, recent studies have shown that the inference performance of CLIP can be greatly degraded by small adversarial perturbations, especially its vi…

2025

Towards Effective Utilization of Mixed-Quality Demonstrations in Robotic Manipulation via Segment-Level Selection and Optimization

ICRA 2025

Data is crucial for robotic manipulation, as it underpins the development of robotic systems for complex tasks. While high-quality, diverse datasets enhance the performance and adaptability of robotic manipulation policies, collecting extensive expert-level data is resource-intensive. Consequently,

Cited by 6SourceScholar
2024

AirExo: Low-Cost Exoskeletons for Learning Whole-Arm Manipulation in the Wild

ICRA 2024poster

While humans can use parts of their arms other than the hands for manipulations like gathering and supporting, whether robots can effectively learn and perform the same type of operations remains relatively unexplored. As these manipulations require joint-level control to regulate the complete poses…

Cited by 40SourcecodeScholar
2024

Doubly Abductive Counterfactual Inference for Text-based Image Editing

CVPR 2024poster

We study text-based image editing (TBIE) of a single image by counterfactual inference because it is an elegant formulation to precisely address the requirement: the edited image should retain the fidelity of the original one. Through the lens of the formulation we find that the crux of TBIE is that…

2024

Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting Learning

AAAI 2024technical

Camera-based bird-eye-view (BEV) perception paradigm has made significant progress in the autonomous driving field. Under such a paradigm, accurate BEV representation construction relies on reliable depth estimation for multi-camera images. However, existing approaches exhaustively predict depths fo…

2024

Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models

NeurIPS 2024poster

Large Multimodal Model (LMM) is a hot research topic in the computer vision area and has also demonstrated remarkable potential across multiple disciplinary fields. A recent trend is to further extend and enhance the perception capabilities of LMMs. The current methods follow the paradigm of adaptin…

2024

NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario

AAAI 2024technical

We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to traditional VQA tasks, VQA in autonomous driving scenario presents more challenges. Firstly, the raw visual data are mul…

2024

Open-Vocabulary Video Relation Extraction

AAAI 2024technical

A comprehensive understanding of videos is inseparable from describing the action with its contextual action-object interactions. However, many current video understanding tasks prioritize general action classification and overlook the actors and relationships that shape the nature of the action, re…

2024

Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image

ECCV 2024poster

"Open-vocabulary 3D object detection (OV-3DDet) aims to localize and recognize both seen and previously unseen object categories within any new 3D scene. While language and vision foundation models have achieved success in handling various open-vocabulary tasks with abundant training data, OV-3DDet…

Cited by 5SourcePDFScholar
2023

Enhancing the Self-Universality for Transferable Targeted Attacks

CVPR 2023poster

In this paper, we propose a novel transfer-based targeted attack method that optimizes the adversarial perturbations without any extra training efforts for auxiliary networks on training data. Our new attack method is proposed based on the observation that highly universal adversarial perturbations…

2023

MSMDFusion: Fusing LiDAR and Camera at Multiple Scales With Multi-Depth Seeds for 3D Object Detection

CVPR 2023poster

Fusing LiDAR and camera information is essential for accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features from two drastically different modalities. Recent approaches aim at e…

2023

SVFormer: Semi-Supervised Video Transformer for Action Recognition

CVPR 2023poster

Semi-supervised action recognition is a challenging but critical task due to the high cost of video annotations. Existing approaches mainly use convolutional neural networks, yet current revolutionary vision transformer models have been less explored. In this paper, we investigate the use of transfo…

2023

Unifying Cross-Lingual and Cross-Modal Modeling Towards Weakly Supervised Multilingual Vision-Language Pre-training

ACL 2023long

Multilingual Vision-Language Pre-training (VLP) is a promising but challenging topic due to the lack of large-scale multilingual image-text pairs. Existing works address the problem by translating English data into other languages, which is intuitive and the generated data is usually limited in form…

2022

Attacking Video Recognition Models with Bullet-Screen Comments

AAAI 2022technical

Recent research has demonstrated that Deep Neural Networks (DNNs) are vulnerable to adversarial patches which introduce perceptible but localized changes to the input. Nevertheless, existing approaches have focused on generating adversarial patches on images, their counterparts in videos have been l…

2022

Balanced Contrastive Learning for Long-Tailed Visual Recognition

CVPR 2022poster

Real-world data typically follow a long-tailed distribution, where a few majority categories occupy most of the data while most minority categories contain a limited number of samples. Classification models minimizing cross-entropy struggle to represent and classify the tail classes. Although the pr…

Cited by 259PDFcodeScholar
2022

Boosting the Transferability of Video Adversarial Examples via Temporal Translation

AAAI 2022technical

Although deep-learning based video recognition models have achieved remarkable success, they are vulnerable to adversarial examples that are generated by adding human-imperceptible perturbations on clean video samples. As indicated in recent studies, adversarial examples are transferable, which make…

2022

MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes

ECCV 2022poster

"3D dense captioning is a recently-proposed novel task, where point clouds contain more geometric information than the 2D counterpart. However, it is also more challenging due to the higher complexity and wider variety of inter-object relations contained in point clouds. Existing methods only treat…

2022

ObjectFormer for Image Manipulation Detection and Localization

CVPR 2022poster

Recent advances in image editing techniques have posed serious challenges to the trustworthiness of multimedia data, which drives the research of image tampering detection. In this paper, we propose ObjectFormer to detect and localize image manipulations. To capture subtle manipulation traces that a…

Cited by 190PDFScholar
2022

Towards Transferable Adversarial Attacks on Vision Transformers

AAAI 2022technical

Vision transformers (ViTs) have demonstrated impressive performance on a series of computer vision tasks, yet they still suffer from adversarial examples. In this paper, we posit that adversarial attacks on transformers should be specially tailored for their architecture, jointly considering both pa…

2021

VideoLT: Large-Scale Long-Tailed Video Recognition

ICCV 2021poster

Label distributions in real-world are oftentimes long-tailed and imbalanced, resulting in biased models towards dominant labels. While long-tailed recognition has been extensively studied for image classification tasks, limited effort has been made for video domain. In this paper, we introduce Video…

Cited by 53PDFcodeScholar
2020

Clean-Label Backdoor Attacks on Video Recognition Models

CVPR 2020poster

Deep neural networks (DNNs) are vulnerable to backdoor attacks which can hide backdoor triggers in DNNs by poisoning training data. A backdoored model behaves normally on clean test images, yet consistently predicts a particular target class for any test examples that contain the trigger pattern. As…

Cited by 346PDFcodeScholar
2020

Hyperbolic Visual Embedding Learning for Zero-Shot Recognition

CVPR 2020poster

This paper proposes a Hyperbolic Visual Embedding Learning Network for zero-shot recognition. The network learns image embeddings in hyperbolic space, which is capable of preserving the hierarchical structure of semantic classes in low dimensions. Comparing with existing zero-shot learning approache…

Cited by 178PDFcodeScholar
2019

R2GAN: Cross-Modal Recipe Retrieval With Generative Adversarial Network

CVPR 2019poster

Representing procedure text such as recipe for crossmodal retrieval is inherently a difficult problem, not mentioning to generate image from recipe for visualization. This paper studies a new version of GAN, named Recipe Retrieval Generative Adversarial Network (R2GAN), to explore the feasibility of…

Cited by 159PDFcodeScholar