← Search

Li Jiang

59 accepted papers

2026

FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging

ICLR 2026oral

Although Video Large Language Models (VLLMs) have shown remarkable capabilities in video understanding, they are required to process high volumes of visual tokens, causing significant computational inefficiency. Existing VLLMs acceleration frameworks usually compress spatial and temporal redundancy…

Cited by 0SourcecodeScholar
2026

GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation

CVPR 2026

Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require precise 3D reasoning. We propose GeoPredict, a geometry-aware VLA framework that augments a continuous-action policy with

Cited by 0SourcecodeScholar
2026

LIVE: Long-horizon Interactive Video World Modeling

ICML 2026poster

Autoregressive video world models predict future visual observations conditioned on actions. While effective over short horizons, these models often struggle with long-horizon generation, as small prediction errors accumulate over time. Prior methods alleviate this by introducing pre-trained teacher…

Cited by 0SourceScholar
2026

SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization

AAAI 2026technical

The emergence of accurate open large language models (LLMs) has sparked a push for advanced quantization techniques to enable efficient deployment on end-user devices. In this paper, we revisit the challenge of extreme LLM compression---targeting ultra-low-bit quantization for both activations and w

Cited by 0SourcePDFScholar
2026

SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

CVPR 2026

Controlling both camera motion and object dynamics is essential for coherent and expressive video generation, yet current methods typically handle only one motion type or rely on ambiguous 2D cues that entangle camera-induced parallax with true object movement. We present SymphoMotion, a unified mot

Cited by 1SourceScholar
2026

UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction

ICLR 2026poster

Feed-forward 3D reconstruction for autonomous driving has advanced rapidly, yet existing methods struggle with the joint challenges of sparse, non-overlapping camera views and complex scene dynamics. We present UniSplat, a general feed-forward framework that learns robust dynamic scene reconstructi…

Cited by 0SourceScholar
2025

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

CVPR 2025poster

Current 3D Large Multimodal Models (3D LMMs) have shown tremendous potential in 3D-vision-based dialogue and reasoning. However, how to further enhance 3D LMMs to achieve fine-grained scene understanding and facilitate flexible human-agent interaction remains a challenging problem. In this work, we…

2025

Are Expressive Models Truly Necessary for Offline RL?

AAAI 2025technical

Among various branches of offline reinforcement learning (RL) methods, goal-conditioned supervised learning (GCSL) has gained increasing popularity as it formulates the offline RL problem as a sequential modeling task, therefore bypassing the notoriously difficult credit assignment challenge of valu…

2025

DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving

ICCV 2025poster

Data-driven learning has advanced autonomous driving, yet task-specific models struggle with out-of-distribution scenarios due to their narrow optimization objectives and reliance on costly annotated data. We present DriveX, a self-supervised world model that learns generalizable scene dynamics and…

Cited by 0SourcePDFScholar
2025

Edit360: 2D Image Edits to 3D Assets from Any Angle

ICCV 2025poster

Recent advances in diffusion models have significantly improved image generation and editing, but extending these capabilities to 3D assets remains challenging, especially for fine-grained edits that require multi-view consistency. Existing methods typically restrict editing to predetermined viewing…

Cited by 0SourcePDFScholar
2025

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

ICCV 2025poster

Recent advances in point cloud perception have demonstrated remarkable progress in scene understanding through vision-language alignment leveraging large language models (LLMs). However, existing methods may still encounter challenges in handling complex instructions that require accurate spatial re…

Cited by 0SourcePDFScholar
2025

FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization

EMNLP 2025

The rapid advancement of large language models (LLMs) has exacerbated the memory bottleneck due to the widening gap between model parameter scaling and hardware capabilities. While post-training quantization techniques effectively reduce memory overhead, existing methods predominantly rely on static

Cited by 0SourcePDFScholar
2025

JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World Data

CVPR 2025poster

Deep-learning-based autonomous driving (AD) perception introduces a promising picture for safe and environment-friendly transportation. However, the over-reliance on real labeled data in LiDAR perception limits the scale of on-road attempts. 3D real world data is notoriously time-and-energy-consumin…

2025

LiON: Learning Point-Wise Abstaining Penalty for LiDAR Outlier DetectioN Using Diverse Synthetic Data

AAAI 2025technical

LiDAR-based semantic scene understanding is an important module in the modern autonomous driving perception stack. However, identifying outlier points in a LiDAR point cloud is challenging as LiDAR point clouds lack semantically-rich information. While former SOTA methods adopt heuristic architectur…

2025

Mitigating Object Hallucinations via Sentence-Level Early Intervention

ICCV 2025poster

Multimodal large language models (MLLMs) have revolutionized cross-modal understanding but continue to struggle with hallucinations - fabricated content contradicting visual inputs. Existing hallucination mitigation methods either incur prohibitive computational costs or introduce distribution misma…

2024

Design of A Rigid-soft Hybrid Robotic Glove with Force Sensing Function

ICRA 2024poster

Soft robotic gloves can not only provide timely, effective, safe and cheap rehabilitation training for patients with impaired movement function of hand, but also assist in completing daily grasping activities. However, most soft robotic gloves are completely composed of flexible structures. Although…

Cited by 1SourceScholar
2024

Force Perception for Rigid-Soft Finger Without Force Sensors: Theoretical Analysis, and Model Transfer

RA-L 2024

Force perception is important for the manipulation of soft robotic hands. Multiple-direction interactions between fingers and objects occur predominantly at the fingertips during manipulation. Integrating physical multi-dimensional force sensors for soft fingertips poses stringent demands on the man

Cited by 3SourceScholar
2024

FreePoint: Unsupervised Point Cloud Instance Segmentation

CVPR 2024poster

Instance segmentation of point clouds is a crucial task in 3D field with numerous applications that involve localizing and segmenting objects in a scene. However achieving satisfactory results requires a large number of manual annotations which is time-consuming and expensive. To alleviate dependenc…

2024

GiT: Towards Generalist Vision Transformer through Universal Language Interface

ECCV 2024oral

"This paper proposes a simple, yet effective framework, called , simultaneously applicable for various vision tasks only with a vanilla ViT. Motivated by the universality of the Multi-layer Transformer architecture (e.g., GPT) widely used in large language models (LLMs), we seek to broaden its scope…

2024

GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding

CVPR 2024poster

Self-supervised 3D representation learning aims to learn effective representations from large-scale unlabeled point clouds. Most existing approaches adopt point discrimination as the pretext task which assigns matched points in two distinct views as positive pairs and unmatched points as negative pa…

2024

MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment

ECCV 2024poster

"Recent approaches have shown that large-scale vision-language models such as CLIP can improve semantic segmentation performance. These methods typically aim for pixel-level vision-language alignment, but often rely on low-resolution image features from CLIP, resulting in class ambiguities along bou…

Cited by 5SourcePDFScholar
2024

OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation

CVPR 2024poster

The booming of 3D recognition in the 2020s began with the introduction of point cloud transformers. They quickly overwhelmed sparse CNNs and became state-of-the-art models especially in 3D semantic segmentation. However sparse CNNs are still valuable networks due to their efficiency treasure and eas…

2024

Point Transformer V3: Simpler Faster Stronger

CVPR 2024poster

This paper is not motivated to seek innovation within the attention mechanism. Instead it focuses on overcoming the existing trade-offs between accuracy and efficiency within the context of point cloud processing leveraging the power of scale. Drawing inspiration from recent advances in 3D large-sca…

Cited by 981SourcePDFScholar
2024

Training Vision Transformers for Semi-Supervised Semantic Segmentation

CVPR 2024poster

We present S4Former a novel approach to training Vision Transformers for Semi-Supervised Semantic Segmentation (S4). At its core S4Former employs a Vision Transformer within a classic teacher-student framework and then leverages three novel technical ingredients: PatchShuffle as a parameter-free per…

2024

Viiat-Hand: A Reach-and-Grasp Restoration System Integrating Voice Interaction, Computer Vision, Auditory and Tactile Feedback for Non-Sighted Amputees

RA-L 2024

For non-sighted and visually impaired (BVI) amputees, the combined loss of vision and grasping abilities turns the seemingly simple task of reaching and grasping into a significant challenge. This letter introduces a novel multi-sensory prosthesis system designed for BVI amputees to assist in percep

Cited by 9SourceScholar
2023

Learning Context-Aware Classifier for Semantic Segmentation

AAAI 2023technical

Semantic segmentation is still a challenging task for parsing diverse contexts in different scenes, thus the fixed classifier might not be able to well address varying feature distributions during testing. Different from the mainstream literature where the efficacy of strong backbones and effective…

2023

Look Beneath the Surface: Exploiting Fundamental Symmetry for Sample-Efficient Offline RL

NeurIPS 2023poster

Offline reinforcement learning (RL) offers an appealing approach to real-world tasks by learning policies from pre-collected datasets without interacting with the environment. However, the performance of existing offline RL algorithms heavily depends on the scale and state-action space coverage of d…

2023

Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

ICLR 2023top-5%

Most offline reinforcement learning (RL) methods suffer from the trade-off between improving the policy to surpass the behavior policy and constraining the policy to limit the deviation from the behavior policy as computing $Q$-values using out-of-distribution (OOD) actions will suffer from errors d…

2023

Self-Supervised Pre-Training With Masked Shape Prediction for 3D Scene Understanding

CVPR 2023poster

Masked signal modeling has greatly advanced self-supervised pre-training for language and 2D images. However, it is still not fully explored in 3D scene understanding. Thus, this paper introduces Masked Shape Prediction (MSP), a new framework to conduct masked signal modeling in 3D scenes. MSP uses…

2022

A Policy-Guided Imitation Approach for Offline Reinforcement Learning

NeurIPS 2022accept

Offline reinforcement learning (RL) methods can generally be categorized into two types: RL-based and Imitation-based. RL-based methods could in principle enjoy out-of-distribution generalization but suffer from erroneous off-policy evaluation. Imitation-based methods avoid off-policy evaluation but…

2022

A Unified Query-Based Paradigm for Point Cloud Understanding

CVPR 2022poster

3D point cloud understanding is an important component in autonomous driving and robotics. In this paper, we present a novel Embedding-Querying paradigm (EQ- Paradigm) for 3D understanding tasks including detection, segmentation and classification. EQ-Paradigm is a unified paradigm that enables comb…

Cited by 55PDFcodeScholar
2022

DODA: Data-Oriented Sim-to-Real Domain Adaptation for 3D Semantic Segmentation

ECCV 2022poster

"Deep learning approaches achieve prominent success in 3D semantic segmentation. However, collecting densely annotated real-world 3D datasets is extremely time-consuming and expensive. Training models on synthetic data and generalizing on real-world scenarios becomes an appealing alternative, but un…

2022

Generalized Few-Shot Semantic Segmentation

CVPR 2022poster

Training semantic segmentation models requires a large amount of finely annotated data, making it hard to quickly adapt to novel classes not satisfying this condition. Few-Shot Segmentation (FS-Seg) tackles this problem with many constraints. In this paper, we introduce a new benchmark, called Gener…

Cited by 109PDFcodeScholar
2022

Motion Transformer with Global Intention Localization and Local Movement Refinement

NeurIPS 2022accept

Predicting multimodal future behavior of traffic participants is essential for robotic vehicles to make safe decisions. Existing works explore to directly predict future trajectories based on latent features or utilize dense goal candidates to identify agent's destinations, where the former strategy…

2022

Point Transformer V2: Grouped Vector Attention and Partition-based Pooling

NeurIPS 2022accept

As a pioneering work exploring transformer architecture for 3D point cloud understanding, Point Transformer achieves impressive results on multiple highly competitive benchmarks. In this work, we analyze the limitations of the Point Transformer and propose our powerful and efficient Point Transforme…

2022

SpikeConverter: An Efficient Conversion Framework Zipping the Gap between Artificial Neural Networks and Spiking Neural Networks

AAAI 2022technical

Spiking Neural Networks (SNNs) have recently attracted enormous research interest since their event-driven and brain-inspired structure enables low-power computation. In image recognition tasks, the best results are achieved by SNN so far utilizing ANN-SNN conversion methods that replace activation…

Cited by 56SourcePDFScholar
2022

Stratified Transformer for 3D Point Cloud Segmentation

CVPR 2022poster

3D point cloud segmentation has made tremendous progress in recent years. Most current methods focus on aggregating local features, but fail to directly model long-range dependencies. In this paper, we propose Stratified Transformer that is able to capture long-range contexts and demonstrates strong…

Cited by 520PDFcodeScholar
2021

A Model-Free Synchronous Control of Humanoid Robot Finger

ICRA 2021poster

For a multi-fingered robot hand, the individual control over single joints cannot guarantee their fine collaboration. For achieving a high-precision synchronization, a theory of synchronous control is introduced to multi-fingered robot hands. This paper introduced a new model-free and cross-coupling…

Cited by 2SourceScholar
2021

Bidirectional Projection Network for Cross Dimension Scene Understanding

CVPR 2021poster

2D image representations are in regular grids and can be processed efficiently, whereas 3D point clouds are unordered and scattered in 3D space. The information inside these two visual domains is well complementary, e.g., 2D images have fine-grained texture while 3D point clouds contain plentiful ge…

Cited by 147PDFcodeScholar
2021

CIA-SSD: Confident IoU-Aware Single-Stage Object Detector From Point Cloud

AAAI 2021technical

Existing single-stage detectors for locating objects in point clouds often treat object localization and category classification as separate tasks, so the localization accuracy and classification confidence may not well align. To address this issue, we present a new single-stage detector named the C…

2021

Guided Point Contrastive Learning for Semi-Supervised Point Cloud Semantic Segmentation

ICCV 2021poster

Rapid progress in 3D semantic segmentation is inseparable from the advances of deep network models, which highly rely on large-scale annotated data for training. To address the high cost and challenges of 3D point-level labeling, we present a method for semi-supervised point cloud semantic segmentat…

Cited by 160PDFScholar
2021

Improving Neural Network Efficiency via Post-Training Quantization With Adaptive Floating-Point

ICCV 2021poster

Model quantization has emerged as a mandatory technique for efficient inference with advanced Deep Neural Networks (DNN). It converts the model parameters in full precision (32-bit floating point) to the hardware friendly data representation with shorter bit-width, to not only reduce the model size…

Cited by 58PDFcodeScholar
2021

SE-SSD: Self-Ensembling Single-Stage Object Detector From Point Cloud

CVPR 2021poster

We present Self-Ensembling Single-Stage object Detector (SE-SSD) for accurate and efficient 3D object detection in outdoor point clouds. Our key focus is on exploiting both soft and hard targets with our formulated constraints to jointly optimize the model, without introducing extra computation in t…

Cited by 460PDFcodeScholar
2021

Semi-Supervised Semantic Segmentation With Directional Context-Aware Consistency

CVPR 2021poster

Semantic segmentation has made tremendous progress in recent years. However, satisfying performance highly depends on a large number of pixel-level annotations. Therefore, in this paper, we focus on the semi-supervised segmentation problem where only a small set of labeled data is provided with a mu…

Cited by 277PDFcodeScholar
2020

PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection

CVPR 2020poster

We present a novel and high-performance 3D object detection framework, named PointVoxel-RCNN (PV-RCNN), for accurate 3D object detection from point clouds. Our proposed method deeply integrates both 3D voxel Convolutional Neural Network (CNN) and PointNet-based set abstraction to learn more discrimi…

Cited by 2428PDFcodeScholar
2020

PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation

CVPR 2020oral

Instance segmentation is an important task for scene understanding. Compared to the fully-developed 2D, 3D instance segmentation for point clouds have much room to improve. In this paper, we present PointGroup, a new end-to-end bottom-up architecture, specifically focused on better grouping the poin…

Cited by 519PDFScholar
2019

Hierarchical Point-Edge Interaction Network for Point Cloud Semantic Segmentation

ICCV 2019poster

We achieve 3D semantic scene labeling by exploring semantic relation between each point and its contextual neighbors through edges. Besides an encoder-decoder branch for predicting point labels, we construct an edge branch to hierarchically integrate point features and generate edge features. To inc…

Cited by 245PDFScholar
2019

PointWeb: Enhancing Local Neighborhood Features for Point Cloud Processing

CVPR 2019poster

This paper presents PointWeb, a new approach to extract contextual features from local neighborhood in a point cloud. Unlike previous work, we densely connect each point with every other in a local neighborhood, aiming to specify feature of each point based on the local region characteristics for be…

Cited by 953PDFcodeScholar
2018

GAL: Geometric Adversarial Loss for Single-View 3D-Object Reconstruction

ECCV 2018poster

In this paper, we present a framework for reconstructing a point-based 3D model of an object from a single view image. Distance metrics, like Chamfer distance, were used in previous work to measure the difference of two point sets and serve as the loss function in point-based reconstruction. However…

Cited by 153SourcePDFScholar
2017

A novel actuation configuration of robotic hand and the mechanical implementation via postural synergies

ICRA 2017poster

How to design a robotic hand for reproducing the move characteristics of human hand joints is a big challenge in robotics. In this paper, we present an approach to determine the actuation configuration based on the statistical results of hand joint angle in different grasps. A relationship between t…

Cited by 8SourceScholar