← Search

Bing Deng

19 accepted papers

2026

AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References

CVPR 2026

Identity-preserving video generation offers powerful tools for creative expression, allowing users to customize videos featuring their beloved characters. However, prevailing methods are typically designed and optimized for a single identity reference. This underlying assumption restricts creative f

Cited by 0SourcecodeScholar
2026

EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer

ICLR 2026poster

Video generation models have advanced significantly, yet they still struggle to synthesize complex human movements due to the high degrees of freedom in human articulation. This limitation stems from the intrinsic constraints of pixel-only training objectives, which inherently bias models toward app…

Cited by 0SourceScholar
2026

Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter (I)

ICRA 2026poster

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation…

Cited by 0codeScholar
2026

Illuminating Visual Identity in Universal Multimodal Embeddings

CVPR 2026

Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has witnessed substantial progress driven by the development of Multimodal Large Language Models (MLLMs). However, a crucial capability, visual identity di

Cited by 0SourcecodeScholar
2026

Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs

ICLR 2026oral

Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques that refine reasoning paths for subsequent fine-tuning. However, extending these language-based self-improving approaches to vision language models (V…

Cited by 0SourcecodeScholar
2025

EchoShot: Multi-Shot Portrait Video Generation

NeurIPS 2025poster

Video diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world applications urge for multiple shots with identity consistency and…

Cited by 0SourcecodeScholar
2025

Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions

CVPR 2025poster

Grounding 3D object affordance is a task that locates objects in 3D space where they can be manipulated, which links perception and action for embodied intelligence. For example, for an intelligent robot, it is necessary to accurately ground the affordance of an object and grasp it according to huma…

Cited by 0SourcePDFScholar
2025

PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Model

ICCV 2025poster

Controllable generation is considered a potentially vital approach to address the challenge of annotating 3D data, and the precision of such controllable generation becomes particularly imperative in the context of data production for autonomous driving. Existing methods focus on the integration of…

2025

TAU-106K: A New Dataset for Comprehensive Understanding of Traffic Accident

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in general visual understanding tasks. However, their potential for high-level, fine-grained comprehension, such as anomaly understanding, remains unexplored. Focusing on traffic accidents, a critical and practical sce…

2024

Enhancing Closed-Loop Performance in Learning-Based Vehicle Motion Planning by Integrating Rule-Based Insights

RA-L 2024

This letter introduces an innovative vehicle motion planning method that leverages the integration of rule-based insights to significantly improve closed-loop performance within a learning-based framework. We first employ rule-based methods to heuristically search and generate a diverse set of traje

Cited by 2SourceScholar
2024

OTVIC: A Dataset with Online Transmission for Vehicle-to-Infrastructure Cooperative 3D Object Detection

IROS 2024poster

Vehicle-to-infrastructure cooperative 3D object detection (VIC3D) is a task that leverages both vehicle and roadside sensors to jointly perceive the surrounding environment. However, considering the high speed of vehicles, the real-time requirements, and the limitations of communication bandwidth, r…

Cited by 1SourceScholar
2024

RoScenes: A Large-scale Multi-view 3D Dataset for Roadside Perception

ECCV 2024poster

"We introduce RoScenes, the largest multi-view roadside perception dataset, which aims to shed light on the development of vision-centric Bird’s Eye View (BEV) approaches for more challenging traffic scenes. The highlights of RoScenes include significantly large perception area, full scene coverage…

2022

Balanced and Hierarchical Relation Learning for One-Shot Object Detection

CVPR 2022poster

Instance-level feature matching is significantly important to the success of modern one-shot object detectors. Recently, the methods based on the metric-learning paradigm have achieved an impressive process. Most of these works only measure the relations between query and target objects on a single…

Cited by 31PDFcodeScholar
2022

Rethinking IoU-Based Optimization for Single-Stage 3D Object Detection

ECCV 2022poster

"Since Intersection-over-Union (IoU) based optimization maintains the consistency of the final IoU prediction metric and losses, it has been widely used in both regression and classification branches of single-stage 2D object detectors. Recently, several 3D object detection methods adopt IoU-based o…

2021

DCT-Mask: Discrete Cosine Transform Mask Representation for Instance Segmentation

CVPR 2021poster

Binary grid mask representation is broadly used in instance segmentation. A representative instantiation is Mask R-CNN which predicts masks on a 28*28 binary grid. Generally, a low-resolution grid is not sufficient to capture the details, while a high-resolution grid dramatically increases the train…

Cited by 88PDFcodeScholar
2021

Improving 3D Object Detection With Channel-Wise Transformer

ICCV 2021poster

Though 3D object detection from point clouds has achieved rapid progress in recent years, the lack of flexible and high-performance proposal refinement remains a great hurdle for existing state-of-the-art two-stage detectors. Previous works on refining 3D proposals have relied on human-designed comp…

Cited by 298PDFcodeScholar
2021

Revisiting Knowledge Distillation: An Inheritance and Exploration Framework

CVPR 2021poster

Knowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the t…

Cited by 41PDFcodeScholar
2021

The Blessings of Unlabeled Background in Untrimmed Videos

CVPR 2021poster

Weakly-supervised Temporal Action Localization (WTAL) aims to detect the action segments with only video-level action labels in training. The key challenge is how to distinguish the action of interest segments from the background, which is unlabelled even on the video-level. While previous works tre…

Cited by 50PDFcodeScholar