← Search

Jun Cen

18 accepted papers

2026

Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective

ICLR 2026poster

Autoregressive large language models (LLMs) have unified a vast range of language tasks, inspiring preliminary efforts in autoregressive (AR) video generation. Existing AR video generators either diverge from standard LLM architectures, depend on bulky external text encoders, or incur prohibitive la…

Cited by 0SourcecodeScholar
2026

RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation

ICRA 2026poster

This paper presents RynnVLA-001, a vision-language-action (VLA) model built upon large-scale video generative pretraining from human demonstrations. We propose a novel two-stage pretraining methodology. The first stage, Ego-Centric Video Generative Pretraining, trains an Image-to-Video model to pred…

2026

Towards Affordance-Aware Robotic Dexterous Grasping with Human-like Priors

AAAI 2026technical

A dexterous hand capable of generalizable grasping objects is fundamental for the development of general-purpose embodied AI. However, previous methods focus narrowly on low-level grasp stability metrics, neglecting affordance-aware positioning and human-like poses which are crucial for downstream m

Cited by 0SourcePDFScholar
2025

$\textit{HiMaCon:}$ Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data

NeurIPS 2025poster

Effective generalization in robotic manipulation requires representations that capture invariant patterns of interaction across environments and tasks. We present a self-supervised framework for learning hierarchical manipulation concepts that encode these invariant patterns through cross-modal sens…

Cited by 0SourceScholar
2024

CMDFusion: Bidirectional Fusion Network With Cross-Modality Knowledge Distillation for LiDAR Semantic Segmentation

RA-L 2024

2D RGB images and 3D LIDAR point clouds provide complementary knowledge for the perception system of autonomous vehicles. Several 2D and 3D fusion methods have been explored for the LIDAR semantic segmentation task, but they suffer from different problems. 2D-to-3D fusion methods require strictly pa

Cited by 21SourcecodeScholar
2024

Self-Supervised Class-Agnostic Motion Prediction with Spatial and Temporal Consistency Regularizations

CVPR 2024poster

The perception of motion behavior in a dynamic environment holds significant importance for autonomous driving systems wherein class-agnostic motion prediction methods directly predict the motion of the entire point cloud. While most existing methods rely on fully-supervised learning the manual labe…

2024

Using Left and Right Brains Together: Towards Vision and Language Planning

ICML 2024poster

Large Language Models (LLMs) and Large Multi-modality Models (LMMs) have demonstrated remarkable decision masking capabilities on a variety of tasks. However, they inherently operate planning within the language space, lacking the vision and spatial imagination ability. In contrast, humans utilize b…

Cited by 5SourcePDFScholar
2023

4D Panoptic Scene Graph Generation

NeurIPS 2023spotlight

We are living in a three-dimensional space while moving forward through a fourth dimension: time. To allow artificial intelligence to develop a comprehensive understanding of such a 4D environment, we introduce **4D Panoptic Scene Graph (PSG-4D)**, a new representation that bridges the raw visual da…

Cited by 16SourcePDFScholar
2023

Enlarging Instance-Specific and Class-Specific Information for Open-Set Action Recognition

CVPR 2023poster

Open-set action recognition is to reject unknown human action cases which are out of the distribution of the training set. Existing methods mainly focus on learning better uncertainty scores but dismiss the importance of feature representations. We find that features with richer semantic diversity c…

2023

Segment Any Point Cloud Sequences by Distilling Vision Foundation Models

NeurIPS 2023spotlight

Recent advancements in vision foundation models (VFMs) have opened up new possibilities for versatile and efficient visual perception. In this work, we introduce Seal, a novel framework that harnesses VFMs for segmenting diverse automotive point cloud sequences. Seal exhibits three appealing propert…

Cited by 66SourcePDFScholar
2023

The Devil is in the Wrongly-classified Samples: Towards Unified Open-set Recognition

ICLR 2023poster

Open-set Recognition (OSR) aims to identify test samples whose classes are not seen during the training process. Recently, Unified Open-set Recognition (UOSR) has been proposed to reject not only unknown samples but also known but wrongly classified samples, which tends to be more practical in real-…

2022

Learning a Condensed Frame for Memory-Efficient Video Class-Incremental Learning

NeurIPS 2022accept

Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to the limited memory. To address this problem, we propose FrameMaker, a memory-efficient video class-incremental learning…

Cited by 20SourcePDFScholar
2022

Open-World Semantic Segmentation for LIDAR Point Clouds

ECCV 2022poster

"Classical LIDAR semantic segmentation is not robust for real-world applications, e.g., autonomous driving, since it is closed-set and static. The closed-set network is only able to output labels of trained classes, even for objects never seen before, while a static network cannot update its knowled…

2022

Real-Time Collision-Free Grasp Pose Detection With Geometry-Aware Refinement Using High-Resolution Volume

RA-L 2022

In this letter, we proposea novel vision-based grasp system for closed-loop 6-degrees of freedom grasping of unknown objects in cluttered environments. The key factor in our system is that we make the most of a geometry-aware scene representation based on a truncated signed distance function (TSDF)

Cited by 22SourceScholar
2021

BORM: Bayesian Object Relation Model for Indoor Scene Recognition

IROS 2021poster

Scene recognition is a fundamental task in robotic perception. For human beings, scene recognition is reasonable because they have abundant object knowledge of the real world. The idea of transferring prior object knowledge from humans to scene recognition is significant but still less exploited. In…

Cited by 23SourcecodeScholar