← Search

Dong Xu

62 accepted papers

2026

MLLMSplat: A 2D MLLM-Powered Framework for 3D Gaussian Splatting Understanding, Generation, and Editing

CVPR 2026

3D Gaussian Splatting (3DGS) has emerged as a mainstream representation for 3D scenes, drawing increasing research attention to its understanding, generation, and editing. However, existing studies remain limited to low-level perception, low-quality generation, and low-efficiency editing, lagging fa

Cited by 0SourceScholar
2026

TORM: Transparent Objects Reconstruction and Manipulation With Multi-View Segmentation

RA-L 2026

Transparent objects are common in daily life and industry, necessitating that robots be able to perceive and manipulate them. The physical properties of reflection and refraction pose challenges for accurately reconstructing the 3D geometry of transparent objects. Conventional methods, which rely on

Cited by 0SourcecodeScholar
2026

TORM: Transparent Objects Reconstruction and Manipulation with Multi-View Segmentation

ICRA 2026poster

Transparent objects are common in daily life and industry, necessitating that robots be able to perceive and manipulate them. The physical properties of reflection and refraction pose challenges for accurately reconstructing the 3D geometry of transparent objects. Conventional methods, which rely on…

Cited by 0SourceScholar
2026

VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation

ICML 2026poster

Scalable Vector Graphics (SVG) animation generation is pivotal for professional design due to their structural editability and resolution independence. However, this task remains challenging as it requires bridging discrete code representations with continuous visual dynamics. Existing optimization-…

Cited by 0SourceScholar
2025

CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward

NeurIPS 2025poster

In this work, we introduce CAD-Coder, a novel framework that reformulates text-to-CAD as the generation of CadQuery scripts—a Python-based, parametric CAD language. This representation enables direct geometric validation, a richer modeling vocabulary, and seamless integration with existing LLMs. To…

Cited by 0SourceScholar
2025

Decentralized Cooperative Localization: A Communication-Efficient Dual-Fusion Consistent Approach

RA-L 2025

Decentralized cooperative localization poses significant challenges in managing inter-robot correlations, especially in environments with limited communication capacity and unreliable network connectivity. In this letter, we propose a communication-efficient decentralized consistent cooperative loca

Cited by 3SourceScholar
2025

Empowering LLMs to Understand and Generate Complex Vector Graphics

CVPR 2025poster

The unprecedented advancements in Large Language Models (LLMs) have profoundly impacted natural language processing but have yet to fully embrace the realm of scalable vector graphics (SVG) generation. While LLMs encode partial knowledge of SVG data from web pages during training, recent findings su…

2025

Improving Long-Text Alignment for Text-to-Image Diffusion Models

ICLR 2025poster

The rapid advancement of text-to-image (T2I) diffusion models has enabled them to generate unprecedented results from given texts. However, as text inputs become longer, existing encoding methods like CLIP face limitations, and aligning the generated images with long texts becomes challenging. To ta…

2025

On-Device Diffusion Transformer Policy for Efficient Robot Manipulation

ICCV 2025poster

Diffusion Policies have significantly advanced robotic manipulation tasks via imitation learning, but their application on resource-constrained mobile platforms remains challenging due to computational inefficiency and extensive memory footprint. In this paper, we propose LightDP, a novel framework…

Cited by 0SourcePDFScholar
2025

SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification

NeurIPS 2025poster

Large language models with Mixture-of-Experts (MoE) architectures achieve efficiency and scalability, yet their routing mechanisms introduce safety alignment challenges insufficiently addressed by techniques developed for dense models. In this work, the MoE-specific safety risk of positional vulnera…

Cited by 0SourceScholar
2025

Scalable MARL for Cooperative Exploration with Dynamic Robot Populations via Graph-Based Information Aggregation

IROS 2025

This study addresses the challenge of multi-robot cooperative exploration under limited local observations in environments with dynamic robot populations. To achieve efficient area coverage within constrained timeframes, we propose the Multi-Robot Informative Planner (MIP), a novel reinforcement lea

Cited by 0SourceScholar
2025

TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine

NeurIPS 2025poster

Traditional Chinese Medicine (TCM), as an effective alternative medicine, has been receiving increasing attention. In recent years, the rapid development of large language models (LLMs) tailored for TCM has highlighted the urgent need for an objective and comprehensive evaluation framework to assess…

Cited by 0SourcecodeScholar
2024

A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing

CVPR 2024poster

This paper presents a video inversion approach for zero-shot video editing which models the input video with low-rank representation during the inversion process. The existing video editing methods usually apply the typical 2D DDIM inversion or naive spatial-temporal DDIM inversion before editing wh…

2024

Data-Free Generalized Zero-Shot Learning

AAAI 2024technical

Deep learning models have the ability to extract rich knowledge from large-scale datasets. However, the sharing of data has become increasingly challenging due to concerns regarding data copyright and privacy. Consequently, this hampers the effective transfer of knowledge from existing data to novel…

2024

Multi-Modality Affinity Inference for Weakly Supervised 3D Semantic Segmentation

AAAI 2024technical

3D point cloud semantic segmentation has a wide range of applications. Recently, weakly supervised point cloud segmentation methods have been proposed, aiming to alleviate the expensive and laborious manual annotation process by leveraging scene-level labels. However, these methods have not effectiv…

2024

SVGDreamer: Text Guided SVG Generation with Diffusion Model

CVPR 2024poster

Recently text-guided scalable vector graphics (SVGs) synthesis has shown promise in domains such as iconography and sketch. However existing text-to-SVG generation methods lack editability and struggle with visual quality and result diversity. To address these limitations we propose a novel text-gui…

2024

UFDA: Universal Federated Domain Adaptation with Practical Assumptions

AAAI 2024technical

Conventional Federated Domain Adaptation (FDA) approaches usually demand an abundance of assumptions, which makes them significantly less feasible for real-world situations and introduces security hazards. This paper relaxes the assumptions from previous FDAs and studies a more practical scenario na…

2023

CS-Isolate: Extracting Hard Confident Examples by Content and Style Isolation

NeurIPS 2023poster

Label noise widely exists in large-scale image datasets. To mitigate the side effects of label noise, state-of-the-art methods focus on selecting confident examples by leveraging semi-supervised learning. Existing research shows that the ability to extract hard confident examples, which are close to…

2023

Conflict-Based Cross-View Consistency for Semi-Supervised Semantic Segmentation

CVPR 2023poster

Semi-supervised semantic segmentation (SSS) has recently gained increasing research interest as it can reduce the requirement for large-scale fully-annotated training data. The current methods often suffer from the confirmation bias from the pseudo-labelling process, which can be alleviated by the c…

2023

DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models

NeurIPS 2023poster

Even though trained mainly on images, we discover that pretrained diffusion models show impressive power in guiding sketch synthesis. In this paper, we present DiffSketcher, an innovative algorithm that creates \textit{vectorized} free-hand sketches using natural language input. DiffSketcher is deve…

2023

VL-SAT: Visual-Linguistic Semantics Assisted Training for 3D Semantic Scene Graph Prediction in Point Cloud

CVPR 2023highlight

The task of 3D semantic scene graph (3DSSG) prediction in the point cloud is challenging since (1) the 3D point cloud only captures geometric structures with limited semantics compared to 2D images, and (2) long-tailed relation distribution inherently hinders the learning of unbiased prediction. Sin…

2022

3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point Clouds

CVPR 2022oral

Observing that the 3D captioning task and the 3D grounding task contain both shared and complementary information in nature, in this work, we propose a unified framework to jointly solve these two distinct but closely related tasks in a synergistic fashion, which consists of both shared task-agnosti…

Cited by 120PDFScholar
2022

Coarse-To-Fine Deep Video Coding With Hyperprior-Guided Mode Prediction

CVPR 2022poster

The previous deep video compression approaches only use the single scale motion compensation strategy and rarely adopt the mode prediction technique from the traditional standards like H.264/H.265 for both motion and residual compression. In this work, we first propose a coarse-to-fine (C2F) deep vi…

Cited by 107PDFScholar
2022

Improving RGB-D Point Cloud Registration by Learning Multi-Scale Local Linear Transformation

ECCV 2022poster

"Point cloud registration aims at estimating the geometric transformation between two point cloud scans, in which accurate correspondence estimation is the key to its success. In addition to previous methods that seek correspondences by hand-crafted or learnt geometric features, recent point cloud r…

2022

SketchSampler: Sketch-Based 3D Reconstruction via View-Dependent Depth Sampling

ECCV 2022poster

"Reconstructing a 3D shape based on a single sketch image is challenging due to the large domain gap between a sparse, irregular sketch and a regular, dense 3D shape. Existing works try to employ the global feature extracted from sketch to directly predict the 3D coordinates, but they usually suffer…

2022

Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis

COLING 2022main

Naturalness and expressiveness are crucial for audiobook speech synthesis, but now are limited by the averaged global-scale speaking style representation. In this paper, we propose an unsupervised multi-scale context-sensitive text-to-speech model for audiobooks. A multi-scale hierarchical context e…

Cited by 9SourcePDFScholar
2021

Back-Tracing Representative Points for Voting-Based 3D Object Detection in Point Clouds

CVPR 2021poster

3D object detection in point clouds is a challenging vision task that benefits various applications for understanding the 3D visual world. Lots of recent research focuses on how to exploit end-to-end trainable Hough voting for generating object proposals. However, the current voting strategy can onl…

Cited by 126PDFcodeScholar
2021

Enhance Curvature Information by Structured Stochastic Quasi-Newton Methods

CVPR 2021poster

In this paper, we consider stochastic second-order methods for minimizing a finite summation of nonconvex functions. One important key is to find an ingenious but cheap scheme to incorporate local curvature information. Since the true Hessian matrix is often a combination of a cheap part and an expe…

Cited by 10PDFScholar
2021

Inception Convolution With Efficient Dilation Search

CVPR 2021poster

As a variant of standard convolution, a dilated convolution can control effective receptive fields and handle large scale variance of objects without introducing additional computational costs. To fully explore the potential of dilated convolution, we proposed a new type of dilated convolution (refe…

Cited by 45PDFcodeScholar
2021

Reinforcement Learning Control of A Novel Magnetic Actuated Flexible-joint Robotic Camera System for Single Incision Laparoscopic Surgery

ICRA 2021poster

This paper describes the control of a novel Magnetic Actuated Flexible-joint Robotic Surgical (MAFRS) camera system with four degrees of freedom (4-DOF) for single incision laparoscopic surgery. Based on the idea of motion decoupling, we designed a novel MAFRS system which is consists of an external…

Cited by 6SourceScholar
2021

SRDAN: Scale-Aware and Range-Aware Domain Adaptation Network for Cross-Dataset 3D Object Detection

CVPR 2021poster

Geometric characteristic plays an important role in the representation of an object in 3D point clouds. For example, large objects often contain more points, while small ones contain fewer points. The point clouds of objects near the capture device are denser, while those of distant objects are spar…

Cited by 54PDFcodeScholar
2021

StyleFormer: Real-Time Arbitrary Style Transfer via Parametric Style Composition

ICCV 2021poster

In this work, we propose a new feed-forward arbitrary style transfer method, referred to as StyleFormer, which can simultaneously fulfill fine-grained style diversity and semantic content coherency. Specifically, our transformer-inspired feature-level stylization method consists of three modules: (a…

Cited by 123PDFcodeScholar
2020

Content Adaptive and Error Propagation Aware Deep Video Compression

ECCV 2020poster

Recently, learning based video compression methods attract increasing attention. However, previous works suffer from error propagation, which stems from the accumulation of reconstructed error in inter predictive coding. Meanwhile, previous learning based video codecs are also not adaptive to differ…

Cited by 160SourcePDFScholar
2020

Improving Deep Video Compression by Resolution-adaptive Flow Coding

ECCV 2020poster

In the learning based video compression approaches, it is an essential issue to compress pixel-level optical flow maps by developing new motion vector (MV) encoders. In this work, we propose a new framework called Resolution-adaptive Flow Coding (RaFC) to effectively compress the flow maps globally…

Cited by 147SourcePDFScholar
2019

DVC: An End-To-End Deep Video Compression Framework

CVPR 2019oral

Conventional video compression approaches use the predictive coding architecture and encode the corresponding motion information and residual information. In this paper, taking advantage of both classical architecture in the conventional video compression method and the powerful non-linear represent…

Cited by 832PDFcodeScholar
2018

Collaborative and Adversarial Network for Unsupervised Domain Adaptation

CVPR 2018poster

In this paper, we propose a new unsupervised domain adaptation approach called Collaborative and Adversarial Network (CAN) through domain-collaborative and domain-adversarial training of neural networks. We use several domain classifiers on multiple CNN feature extraction layers/blocks, in which eac…

2018

Deep Kalman Filtering Network for Video Compression Artifact Reduction

ECCV 2018poster

When lossy video compression algorithms are applied, compression artifacts often appear in videos, making decoded videos unpleasant for human visual systems. In this paper, we model the video artifact reduction task as a Kalman filtering procedure and restore decoded frames through a deep Kalman fil…

Cited by 120SourcePDFScholar
2017

Complex Event Detection by Identifying Reliable Shots From Untrimmed Videos

ICCV 2017poster

The goal of complex event detection is to automatically detect whether an event of interest happens in temporally untrimmed long videos which usually consist of multiple video shots. Observing some video shots in positive (resp. negative) videos are irrelevant (resp. relevant) to the given event cla…

Cited by 55PDFScholar
2017

SPFTN: A Self-Paced Fine-Tuning Network for Segmenting Objects in Weakly Labelled Videos

CVPR 2017poster

Object segmentation in weakly labelled videos is an interesting yet challenging task, which aims at learning to perform category-specific video object segmentation by only using video-level tags. Existing works in this research area might still have some limitations, e.g., lack of effective DNN-base…

Cited by 61PDFScholar
2016

Proximal Riemannian Pursuit for Large-Scale Trace-Norm Minimization

CVPR 2016poster

Trace-norm regularization plays an important role in many areas such as machine learning and computer vision. Solving trace-norm regularized Trace-norm regularization plays an important role in many areas such as computer vision and machine learning. When solving general large-scale trace-norm regul…

Cited by 4PDFcodeScholar
2015

Visual Recognition by Learning From Web Data: A Weakly Supervised Domain Generalization Approach

CVPR 2015poster

In this work, we formulate a new weakly supervised domain generalization problem for the visual recognition task by using loosely labeled web images/videos as training data. Specifically, we aim to address two challenging issues when learning robust classifiers: 1) enhancing the generalization capab…

Cited by 94SourcePDFScholar