← Search

Jiajun Deng

43 accepted papers

2026

EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models

ICML 2026spotlight

While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulation, motivating parameter sparsification. However, as the environment evolves during VLA execution, the optimal sparsity…

Cited by 0SourceScholar
2026

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding

CVPR 2026

Finetuning Large Vision-Language Models with reinforcement learning has emerged as a promising approach to enhance their capability in object-level grounding. However, existing methods, mainly based on GRPO, assign rewards at the response level. Such sparse reward leads to minimal learning signals w

Cited by 0SourcecodeScholar
2026

Goal-VLA: Image-Generative VLMs As Object-Centric World Models Empowering Zero-Shot Robot Manipulation

ICRA 2026poster

Generalization remains a fundamental challenge in robotic manipulation. To tackle this challenge, recent Vision-Language-Action (VLA) models build policies on top of Vision-Language Models (VLMs), seeking to transfer their open-world semantic knowledge. However, their zero-shot capability lags signi…

2026

LISN: Language-Instructed Social Navigation with VLM-Based Controller Modulating

ICRA 2026poster

Towards human-robot coexistence, socially aware navigation is significant for mobile robots. Yet existing studies on this area focus mainly on path efficiency and pedestrian collision avoidance, which are essential but represent only a fraction of social navigation. Beyond these basics, robots must …

2026

Learning Surgical Robotic Manipulation with 3D Spatial Priors

CVPR 2026

Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or enhance multi-view features by adding wrist-mounted cameras to supplement the d

Cited by 0SourceScholar
2026

RaCFusion: Improving Camera-Based 3D Object Detection via Radar-Assisted Hierarchical Refinement

RA-L 2026

Cameras and radar sensors are complementary in 3D object detection in that cameras specialize in capturing an object's visual information while radar provides spatial information and velocity hints. Existing radar-camera fusion methods often employ a symmetrical architecture that processes inputs fr

Cited by 0SourcecodeScholar
2025

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

CVPR 2025poster

Current 3D Large Multimodal Models (3D LMMs) have shown tremendous potential in 3D-vision-based dialogue and reasoning. However, how to further enhance 3D LMMs to achieve fine-grained scene understanding and facilitate flexible human-agent interaction remains a challenging problem. In this work, we…

2025

Effective and Efficient Mixed Precision Quantization of Speech Foundation Models

ICASSP 2025accepted

This paper presents a novel mixed-precision quantization approach for speech foundation models that tightly integrates mixed-precision learning and quantized model parameter estimation into one single model compression stage. Experiments conducted on LibriSpeech dataset with fine-tuned wav2vec2.0-ba…

Cited by 7SourceScholar
2025

GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions

ICCV 2025poster

Flexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities of the large language models (LLMs) to establish mappings between expressions and targets, allowing robots to comprehend…

2025

Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots

ICML 2025poster

Autoregressive models have emerged as a powerful generative paradigm for visual generation. The current de-facto standard of next token prediction commonly operates over a single-scale sequence of dense image tokens, and is incapable of utilizing global context especially for early tokens prediction…

2025

Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition

ICASSP 2025accepted

Discrete tokens provide compact and domain-adaptable representations of speech features. However, their application to disordered speech, characterized by articulation imprecision and significant mismatch with normal voice, remains unexplored. To this end, this paper proposes novel phone-purity guid…

Cited by 0SourceScholar
2025

RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion

CVPR 2025poster

We propose Radar-Camera fusion transformer (RaCFormer) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception is capped by the image-to-BEV transformation-if the depth of pixels is not accurately estimated, the naive combination…

2025

S3R-GS: Streamlining the Pipeline for Large-Scale Street Scene Reconstruction

ICCV 2025poster

Recently, 3D Gaussian Splatting (3DGS) has reshaped the field of photorealistic 3D reconstruction, achieving impressive rendering quality and speed. However, when applied to large-scale street scenes, existing methods suffer from rapidly escalating per-viewpoint reconstruction costs as scene size in…

2025

Self-Classification Enhancement and Correction for Weakly Supervised Object Detection

IJCAI 2025

In recent years, weakly supervised object detection (WSOD) has attracted much attention due to its low labeling cost. The success of recent WSOD models is often ascribed to the two-stage multi-class classification (MCC) task, i.e., multiple instance learning and online classification refinement. Des

Cited by 0SourcePDFScholar
2025

SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images

ICCV 2025poster

A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods…

Cited by 0SourcePDFScholar
2025

SynTag: Enhancing the Geometric Robustness of Inversion-based Generative Image Watermarking

ICCV 2025poster

Robustness is significant for generative image watermarking, typically achieved by injecting distortion-invariant watermark features. The leading paradigm, i.e., inversion-based framework, excels against non-geometric distortions but struggles with geometric ones. To address this, we propose SynTag,…

Cited by 0SourcePDFScholar
2024

End-to-End Rate-Distortion Optimized 3D Gaussian Representation

ECCV 2024poster

"3D Gaussian Splatting (3DGS) has become an emerging technique with remarkable potential in 3D representation and image rendering. However, the substantial storage overhead of 3DGS significantly impedes its practical applications. In this work, we formulate the compact 3D Gaussian learning as an end…

2024

Hierarchical Temporal Context Learning for Camera-based Semantic Scene Completion

ECCV 2024poster

"Camera-based 3D semantic scene completion (SSC) is pivotal for predicting complicated 3D layouts with limited 2D image observations. The existing mainstream solutions generally leverage temporal information by roughly stacking history frames to supplement the current frame, such straightforward tem…

2024

Revisiting Open-Set Panoptic Segmentation

AAAI 2024technical

In this paper, we focus on the open-set panoptic segmentation (OPS) task to circumvent the data explosion problem. Different from the close-set setting, OPS targets to detect both known and unknown categories, where the latter is not annotated during training. Different from existing work that only…

Cited by 0SourcePDFScholar
2024

Towards Automatic Data Augmentation for Disordered Speech Recognition

ICASSP 2024accepted

Automatic recognition of disordered speech remains a highly challenging task to date due to data scarcity. This paper presents a reinforcement learning (RL) based on-the-fly data augmentation approach for training state-of-the-art PyChain TDNN and end-to-end Conformer ASR systems on such data. The h…

Cited by 11SourceScholar
2024

Towards High-Performance and Low-Latency Feature-Based Speaker Adaptation of Conformer Speech Recognition Systems

ICASSP 2024accepted

Practical application of model-based speaker adaptation techniques to end-to-end ASR systems is hindered by speaker-level data scarcity and latency in speaker-dependent (SD) parameters update. To this end, data-efficient and low-latency rapid feature-based speaker adaptation approaches are proposed…

Cited by 0SourceScholar
2023

3DPPE: 3D Point Positional Encoding for Transformer-based Multi-Camera 3D Object Detection

ICCV 2023poster

Transformer-based methods have swept the benchmarks on 2D and 3D detection on images. Because tokenization before the attention mechanism drops the spatial information, positional encoding becomes critical for those methods. Recent works found that encodings based on samples of the 3D viewing rays c…

Cited by 25PDFcodeScholar
2023

Adversarial Data Augmentation Using VAE-GAN for Disordered Speech Recognition

ICASSP 2023accepted

Automatic recognition of disordered speech remains a highly challenging task to date. The underlying neuro-motor conditions, often compounded with co-occurring physical disabilities, lead to the difficulty in collecting large quantities of impaired speech required for ASR system development. This pa…

Cited by 0SourceScholar
2023

Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object Detection

CVPR 2023poster

LiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear.…

2023

CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI Detection

NeurIPS 2023poster

Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories. A strong zero-shot HOI detector is supposed to be not only capable of discriminating novel interactions but also robust to positional distribution discrepancy between seen and unseen categories w…

Cited by 21SourcePDFScholar
2023

CluB: Cluster Meets BEV for LiDAR-Based 3D Object Detection

NeurIPS 2023poster

Currently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors. BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution l…

Cited by 6SourcePDFScholar
2023

Cyclic-Bootstrap Labeling for Weakly Supervised Object Detection

ICCV 2023poster

Recent progress in weakly supervised object detection is featured by a combination of multiple instance detection networks (MIDN) and ordinal online refinement. However, with only image-level annotation, MIDN inevitably assigns high scores to some unexpected region proposals when generating pseudo l…

Cited by 12PDFcodeScholar
2023

Exploiting Prompt Learning with Pre-Trained Language Models for Alzheimer's Disease Detection

ICASSP 2023accepted

Early diagnosis of Alzheimer’s disease (AD) is crucial in facilitating preventive care and to delay further progression. Speech based automatic AD screening systems provide a non-intrusive and more scalable alternative to other clinical screening techniques. Textual embedding features produced by pr…

Cited by 0SourceScholar
2023

Exploring Self-Supervised Pre-Trained ASR Models for Dysarthric and Elderly Speech Recognition

ICASSP 2023accepted

Automatic recognition of disordered and elderly speech remains a highly challenging task to date due to the difficulty in collecting such data in large quantities. This paper explores a series of approaches to integrate domain adapted Self-Supervised Learning (SSL) pre-trained models into TDNN and C…

Cited by 0SourceScholar
2023

Invariant Training 2D-3D Joint Hard Samples for Few-Shot Point Cloud Recognition

ICCV 2023poster

We tackle the data scarcity challenge in few-shot point cloud recognition of 3D objects by using a joint prediction from a conventional 3D model and a well-pretrained 2D model. Surprisingly, such an ensemble, though seems trivial, has hardly been shown effective in recent 2D-3D models. We find out t…

Cited by 13PDFcodeScholar
2023

Masked Motion Predictors are Strong 3D Action Representation Learners

ICCV 2023poster

In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. In this work, we show that ins…

Cited by 46PDFcodeScholar
2023

SimFIR: A Simple Framework for Fisheye Image Rectification with Self-supervised Representation Learning

ICCV 2023poster

In fisheye images, rich distinct distortion patterns are regularly distributed in the image plane. These distortion patterns are independent of the visual content and provide informative cues for rectification. To make the best of such rectification cues, we introduce SimFIR, a simple framework for…

Cited by 31PDFScholar
2022

${\mathsf{EZFusion}}$: A Close Look at the Integration of LiDAR, Millimeter-Wave Radar, and Camera for Accurate 3D Object Detection and Tracking

RA-L 2022

A recent trend is to combine multiple sensors ( <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i.e.</i> , cameras, LiDARs and millimeter-wave Radars) to achieve robust multi-modal perception for autonomous systems such as self-driving vehicles. Alth

Cited by 11SourceScholar
2022

Audio-Visual Multi-Channel Speech Separation, Dereverberation and Recognition

ICASSP 2022accepted

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverberation remains a highly challenging task to date. Motivated by the invariance of v…

Cited by 0SourceScholar
2022

CMD: Self-Supervised 3D Action Representation Learning with Cross-Modal Mutual Distillation

ECCV 2022poster

"In 3D action recognition, there exists rich complementary information between skeleton modalities. Nevertheless, how to model and utilize this information remains a challenging problem for self-supervised 3D action representation learning. In this work, we formulate the cross-modal interaction as a…

2022

Geometric Representation Learning for Document Image Rectification

ECCV 2022poster

"In document image rectification, there exist rich geometric constraints between the distorted image and the ground truth one. How- ever, such geometric constraints are largely ignored in existing advanced solutions, which limits the rectification performance. To this end, we present DocGeoNet for d…

2021

Instance Mining with Class Feature Banks for Weakly Supervised Object Detection

AAAI 2021technical

Recent progress on weakly supervised object detection (WSOD) is characterized by formulating WSOD as a Multiple Instance Learning (MIL) problem and taking online refinement with the selected region proposals from MIL. However, MIL inclines to select the most discriminative part rather than the entir…

Cited by 39SourcePDFScholar
2021

TransVG: End-to-End Visual Grounding With Transformers

ICCV 2021poster

In this paper, we present a neat yet effective transformer-based framework for visual grounding, namely TransVG, to address the task of grounding a language query to the corresponding region onto an image. The state-of-the-art methods, including two-stage or one-stage ones, rely on a complex module…

Cited by 408PDFcodeScholar
2021

Voxel R-CNN: Towards High Performance Voxel-based 3D Object Detection

AAAI 2021technical

Recent advances on 3D object detection heavily rely on how the 3D data are represented, i.e., voxel-based or point-based representation. Many existing high performance 3D detectors are point-based because this structure can better retain precise point positions. Nevertheless, point-level features le…

2019

Relation Distillation Networks for Video Object Detection

ICCV 2019poster

It has been well recognized that modeling object-to-object relations would be helpful for object detection. Nevertheless, the problem is not trivial especially when exploring the interactions between objects to boost video object detectors. The difficulty originates from the aspect that reliable obj…

Cited by 280PDFScholar