← Search

JING DONG

41 accepted papers

2026

$\alpha$-DPO: Robust Preference Alignment for Diffusion Models via $\alpha$ Divergence

ICLR 2026poster

Diffusion models have demonstrated remarkable success in high-fidelity image generation, yet aligning them with human preferences remains challenging. Direct Preference Optimization (DPO) offers a promising framework, but its effectiveness is critically hindered by noisy data arising from mislabeled…

Cited by 0SourceScholar
2026

RADAR: Defending RAG Dynamically against Retrieval Corruption

ICML 2026poster

While RAG systems are increasingly deployed in dynamic web search, temporal volatility amplifies their vulnerability to adversarial attacks. Existing static-oriented defenses struggle to handle evolving threats and incur prohibitive storage costs in dynamic settings. We propose RADAR, a framework th…

Cited by 0SourceScholar
2026

Revisiting MLLM Based Image Quality Assessment: Errors and Remedy

AAAI 2026technical

The rapid progress of multi-modal large language models (MLLMs) has boosted the task of image quality assessment (IQA). However, a key challenge arises from the inherent mismatch between the discrete token outputs of MLLMs and the continuous nature of quality scores required by IQA tasks. This discr

Cited by 0SourcePDFScholar
2026

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination

CVPR 2026

Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely attributes this to insufficient visual attention. However, in this work, we are surprised to find that both real and hallucinated objects receive equally

Cited by 0SourcecodeScholar
2025

Adaptive Median Smoothing: Adversarial Defense for Unlearned Text-to-Image Diffusion Models at Inference Time

ICML 2025poster

Text-to-image (T2I) diffusion models have raised concerns about generating inappropriate content, such as "*nudity*". Despite efforts to erase undesirable concepts through unlearning techniques, these unlearned models remain vulnerable to adversarial inputs that can potentially regenerate such conte…

Cited by 0SourcePDFScholar
2025

Image-level Memorization Detection via Inversion-based Inference Perturbation

ICLR 2025poster

Recent studies have discovered that widely used text-to-image diffusion models can replicate training samples during image generation, a phenomenon known as memorization. Existing detection methods primarily focus on identifying memorized prompts. However, in real-world scenarios, image owners may n…

Cited by 0SourcePDFScholar
2025

Learning Imperfect Information Extensive-form Games with Last-iterate Convergence under Bandit Feedback

ICML 2025poster

We investigate learning approximate Nash equilibrium (NE) policy profiles in two-player zero-sum imperfect information extensive-form games (IIEFGs) with last-iterate convergence guarantees. Existing algorithms either rely on full-information feedback or provide only asymptotic convergence rates. In…

Cited by 0SourcePDFScholar
2025

Partial Reconstruction Error for Deepfake Detection

ICASSP 2025accepted

The rapid development of deepfake technology poses a formidable challenge to personal privacy and security, underscoring the urgent need for deepfake detection. Recently, the methods based on the reconstruction error, such as DIRE and RECCE, achieve impressive performance in forgery detection. Howev…

Cited by 0SourceScholar
2025

Towards Black-Box Membership Inference Attack for Diffusion Models

ICML 2025poster

Given the rising popularity of AI-generated art and the associated copyright concerns, identifying whether an artwork was used to train a diffusion model is an important research topic. The work approaches this problem from the membership inference attack (MIA) perspective. We first identify the lim…

Cited by 4SourcePDFScholar
2024

AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head Synthesis

AAAI 2024technical

Audio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies. However, when it comes to few-shot talking head generation, a p…

Cited by 10SourcePDFScholar
2024

Convergence to Nash Equilibrium and No-regret Guarantee in (Markov) Potential Games

AISTATS 2024poster

In this work, we study potential games and Markov potential games under stochastic cost and bandit feedback. We propose a variant of the Frank-Wolfe algorithm with sufficient exploration and recursive gradient estimation, which provably converges to the Nash equilibrium while attaining sublinear reg…

Cited by 0SourcePDFScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

Learning Dense Correspondence for NeRF-Based Face Reenactment

AAAI 2024technical

Face reenactment is challenging due to the need to establish dense correspondence between various face representations for motion transfer. Recent studies have utilized Neural Radiance Field (NeRF) as fundamental representation, which further enhanced the performance of multi-view face reenactment i…

Cited by 11SourcePDFScholar
2024

Online Control with Adversarial Disturbance for Continuous-time Linear Systems

NeurIPS 2024poster

We study online control for continuous-time linear systems with finite sampling rates, where the objective is to design an online procedure that learns under non-stochastic noise and performs comparably to a fixed optimal linear controller. We present a novel two-level online algorithm, by integrat…

Cited by 0SourcePDFScholar
2024

Online Policy Optimization for Robust Markov Decision Process

UAI 2024poster

Reinforcement learning (RL) has exceeded human performance in many synthetic settings such as video games and Go. However, real-world deployment of end-to-end RL models is less common, as RL models can be very sensitive to perturbations in the environment. The robust Markov decision process (MDP) fr…

2024

QGym: Scalable Simulation and Benchmarking of Queuing Network Controllers

NeurIPS 2024poster

Queuing network control allows allocation of scarce resources to manage congestion, a fundamental problem in manufacturing, communications, and healthcare. Compared to standard RL problems, queueing problems are distinguished by unique challenges: i) a system operating in continuous time, ii) high…

2024

Robust Indoor Localization with Ranging-IMU Fusion

ICRA 2024poster

Indoor wireless ranging localization is a promising approach for low-power and high-accuracy localization of wearable devices. A primary challenge in this domain stems from non-line of sight propagation of radio waves. This study tackles a fundamental issue in wireless ranging: the unpredictability…

Cited by 4SourceScholar
2024

S^3D-NeRF: Single-Shot Speech-Driven Neural Radiance Field for High Fidelity Talking Head Synthesis

ECCV 2024poster

"Talking head synthesis is a practical technique with wide applications. Current Neural Radiance Field (NeRF) based approaches have shown their superiority on driving one-shot talking heads with videos or signals regressed from audio. However, most of them failed to take the audio as driven informat…

Cited by 3SourcePDFScholar
2023

CFFT-GAN: Cross-Domain Feature Fusion Transformer for Exemplar-Based Image Translation

AAAI 2023technical

Exemplar-based image translation refers to the task of generating images with the desired style, while conditioning on certain input image. Most of the current methods learn the correspondence between two input domains and lack the mining of information within the domain. In this paper, we propose a…

Cited by 7SourcePDFScholar
2023

DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning

IJCAI 2023poster

Communication lays the foundation for cooperation in human society and in multi-agent reinforcement learning (MARL). Humans also desire to maintain their privacy when communicating with others, yet such privacy concern has not been considered in existing works in MARL. We propose the differentially…

2023

Decentralization and Acceleration Enables Large-Scale Bundle Adjustment

RSS 2023poster

Scaling to arbitrarily large bundle adjustment problems requires data and compute to be distributed across multiple devices. Centralized methods in prior works are only able to solve small or medium size problems due to overhead in computation and communication. In this paper, we present a fully dec…

2023

Semantic 3D-Aware Portrait Synthesis and Manipulation Based on Compositional Neural Radiance Field

AAAI 2023technical

Recently 3D-aware GAN methods with neural radiance field have developed rapidly. However, current methods model the whole image as an overall neural radiance field, which limits the partial semantic editability of synthetic results. Since NeRF renders an image pixel by pixel, it is possible to split…

2022

Theseus: A Library for Differentiable Nonlinear Optimization

NeurIPS 2022accept

We present Theseus, an efficient application-agnostic open source library for differentiable nonlinear least squares (DNLS) optimization built on PyTorch, providing a common framework for end-to-end structured learning in robotics and vision. Existing DNLS implementations are application specific an…

Cited by 107SourcePDFScholar
2022

iSDF: Real-Time Neural Signed Distance Fields for Robot Perception

RSS 2022poster

We present iSDF, a continual learning system for real-time signed distance field (SDF) reconstruction. Given a stream of posed depth images from a moving camera, it trains a randomly initialised neural network to map input 3D coordinate to approximate signed distance. The model is self-supervised by…

2021

Distributed Client-Server Optimization for SLAM with Limited On-Device Resources

ICRA 2021poster

Simultaneous localization and mapping (SLAM) is a crucial functionality for exploration robots and virtual/augmented reality (VR/AR) devices. However, some of such devices with limited resources cannot afford the computational or memory cost to run full SLAM algorithms. We propose a general client-s…

Cited by 12SourceScholar
2021

MR-iSAM2: Incremental Smoothing and Mapping with Multi-Root Bayes Tree for Multi-Robot SLAM

IROS 2021poster

We present multi-robot iSAM2 (MR-iSAM2), an efficient incremental smoothing and mapping (iSAM) algorithm to solve multi-robot simultaneous localization and mapping (SLAM) inference problems. MR-iSAM2 is based on a novel data structure multi-root Bayes tree (MRBT), which packs multiple Bayes trees wi…

Cited by 14SourceScholar
2021

MUST-GAN: Multi-Level Statistics Transfer for Self-Driven Person Image Generation

CVPR 2021poster

Pose-guided person image generation usually involves using paired source-target images to supervise the training, which significantly increases the data preparation effort and limits the application of the models. To deal with this problem, we propose a novel multi-level statistics transfer model, w…

Cited by 36PDFScholar
2020

TLIO: Tight Learned Inertial Odometry

RA-L 2020

In this letter we propose a tightly-coupled Extended Kalman Filter framework for IMU-only state estimation. Strap-down IMU measurements provide relative state estimates based on IMU kinematic motion model. However the integration of measurements is sensitive to sensor bias and noise, causing signifi

Cited by 241SourcecodeScholar
2019

ACCELERATING NONCONVEX LEARNING VIA REPLICA EXCHANGE LANGEVIN DIFFUSION

ICLR 2019poster

Langevin diffusion is a powerful method for nonconvex optimization, which enables the escape from local minima by injecting noise into the gradient. In particular, the temperature parameter controlling the noise level gives rise to a tradeoff between ``global exploration'' and ``local exploitation''…

Cited by 44SourcePDFScholar
2018

Sparse Gaussian Processes on Matrix Lie Groups: A Unified Framework for Optimizing Continuous-Time Trajectories

ICRA 2018poster

Continuous-time trajectories are useful for reasoning about robot motion in a wide range of tasks. Sparse Gaussian processes (GPs) can be used as a non-parametric representation for trajectory distributions that enables fast trajectory optimization by sparse GP regression. However, most previous app…

Cited by 34SourceScholar
2017

4D crop monitoring: Spatio-temporal reconstruction for agriculture

ICRA 2017poster

Autonomous crop monitoring at high spatial and temporal resolution is a critical problem in precision agriculture. While Structure from Motion and Multi-View Stereo algorithms can finely reconstruct the 3D structure of a field with low-cost image sensors, these algorithms fail to capture the dynamic…

Cited by 140SourceScholar
2017

Simultaneous Trajectory Estimation and Planning via Probabilistic Inference

RSS 2017poster

We provide a unified probabilistic framework for trajectory estimation and planning. The key idea is to view these two problems, usually considered separately, as a single problem. At each time-step the robot is tasked with finding the complete continuous-time trajectory from start to goal. This can…

2016

Motion Planning as Probabilistic Inference using Gaussian Processes and Factor Graphs

RSS 2016poster

With the increased use of high degree-of-freedom robots that must perform tasks in real-time, there is a need for fast algorithms for motion planning. In this work, we view motion planning from a probabilistic perspective. We consider smooth continuous-time trajectories as samples from a Gaussian pr…

Cited by 171SourcePDFScholar
2015

Distributed real-time cooperative localization and mapping using an uncertainty-aware expectation maximization approach

ICRA 2015poster

We demonstrate distributed, online, and real-time cooperative localization and mapping between multiple robots operating throughout an unknown environment using indirect measurements. We present a novel Expectation Maximization (EM) based approach to efficiently identify inlier multi-robot loop clos…

Cited by 98SourceScholar