← Search

Junwei Liang

27 accepted papers

2026

EgoTraj-Bench: Towards Robust Trajectory Prediction under Ego-View Noisy Observations

ICRA 2026poster

Reliable trajectory prediction from an ego-centric perspective is crucial for robotic navigation in human-centric environments. However, existing methods typically assume noiseless observation histories, failing to account for the perceptual artifacts inherent in first-person vision, such as occlusi…

2026

Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration

RA-L 2026

Deployable service and delivery robots struggle to navigate multi-floor buildings to reach object goals, as existing systems fail due to single-floor assumptions and requirements for offline, globally consistent maps. Multi-floor environments pose unique challenges including cross-floor transitions

Cited by 3SourcecodeScholar
2025

Exploring the Limits of Vision-Language-Action Manipulation in Cross-task Generalization

NeurIPS 2025poster

The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored. To address this…

Cited by 0SourceScholar
2025

From Cognition to Precognition: A Future-Aware Framework for Social Navigation

ICRA 2025

To navigate safely and efficiently in crowded spaces, robots should not only perceive the current state of the environment but also anticipate future human movements. In this paper, we propose a reinforcement learning architecture, namely Falcon, to tackle socially-aware navigation by explicitly pre

Cited by 13SourcecodeScholar
2025

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation

CoRL 2025poster

Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence—particularly through the lens of *actionable affordances*. However, transferring such knowledge remains challenging due to: 1) a lack of large-scale data…

Cited by 0SourceScholar
2025

GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs

ICCV 2025poster

Estimating physical properties for visual data is a crucial task in computer vision, graphics, and robotics, underpinning applications such as augmented reality, physical simulation, and robotic grasping. However, this area remains under-explored due to the inherent ambiguities in physical property…

Cited by 0SourcePDFScholar
2025

Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation

CVPR 2025poster

Learning generalizable visual representations across different embodied environments is essential for effective robotic manipulation in real-world scenarios. However, the limited scale and diversity of robot demonstration data pose a significant challenge. Recent research has explored leveraging lar…

Cited by 8SourcePDFScholar
2025

Omni-Perception: Omnidirectional Collision Avoidance of Legged Robots in Dynamic Environments

CoRL 2025oral

Agile locomotion in complex 3D environments requires robust spatial awareness to safely avoid diverse obstacles such as aerial clutter, uneven terrain, and dynamic agents. Depth-based perception approaches often struggle with sensor noise, lighting variability, computational overhead from intermedia…

Cited by 0SourceScholar
2025

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

CVPR 2025poster

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and predefined object categories, limiting scalability and adaptability. To overcome…

2024

An Examination of the Compositionality of Large Generative Vision-Language Models

NAACL 2024long

With the success of Large Language Models (LLMs), many Generative Vision-Language Models (GVLMs) have been constructed via multimodal instruction tuning. However, the performance of GVLMs in multimodal compositional reasoning remains under-explored. In this paper, we examine both the evaluation metr…

2024

Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation

CoRL 2024poster

Developing robots capable of executing various manipulation tasks, guided by natural language instructions and visual observations of intricate real-world environments, remains a significant challenge in robotics. Such robot agents need to understand linguistic commands and distinguish between the…

Cited by 11SourceScholar
2024

DragTraffic: Interactive and Controllable Traffic Scene Generation for Autonomous Driving

IROS 2024

Evaluating and training autonomous driving systems require diverse and scalable corner cases. However, most existing scene generation methods lack controllability, accuracy, and versatility, resulting in unsatisfactory generation results. Inspired by DragGAN in image generation, we propose DragTraff

Cited by 6SourcecodeScholar
2024

FinTextQA: A Dataset for Long-form Financial Question Answering

ACL 2024long

Accurate evaluation of financial question answering (QA) systems necessitates a comprehensive dataset encompassing diverse question types and contexts. However, current financial QA datasets lack scope diversity and question complexity. This work introduces FinTextQA, a novel dataset for long-form q…

Cited by 10SourcePDFScholar
2024

Improving Gloss-free Sign Language Translation by Reducing Representation Density

NeurIPS 2024poster

Gloss-free sign language translation (SLT) aims to develop well-performing SLT systems with no requirement for the costly gloss annotations, but currently still lags behind gloss-based approaches significantly. In this paper, we identify **a representation density problem** that could be a bottlenec…

2024

Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models

ECCV 2024poster

"In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel method, namely Diff2Scene, which leverages frozen representations from text-image generative models, along with salien…

Cited by 4SourcePDFScholar
2023

STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action Recognition

CVPR 2023poster

We study the problem of human action recognition using motion capture (MoCap) sequences. Unlike existing techniques that take multiple manual steps to derive standardized skeleton representations as model input, we propose a novel Spatial-Temporal Mesh Transformer (STMT) to directly model the mesh s…

2022

Multi-dataset Training of Transformers for Robust Action Recognition

NeurIPS 2022accept

We study the task of robust feature representations, aiming to generalize well on multiple datasets for action recognition. We build our method on Transformers for its efficacy. Although we have witnessed great progress for video action recognition in the past decade, it remains challenging yet valu…

2022

Text-Adaptive Multiple Visual Prototype Matching for Video-Text Retrieval

NeurIPS 2022accept

Cross-modal retrieval between videos and texts has gained increasing interest because of the rapid emergence of videos on the web. Generally, a video contains rich instance and event information and the query text only describes a part of the information. Thus, a video can have multiple different…

Cited by 32SourcePDFScholar
2021

Weakly Supervised 3D Semantic Segmentation Using Cross-Image Consensus and Inter-Voxel Affinity Relations

ICCV 2021poster

We propose a novel weakly supervised approach for 3D semantic segmentation on volumetric images. Unlike most existing methods that require voxel-wise densely labeled training data, our weakly-supervised CIVA-Net is the first model that only needs image-level class labels as guidance to learn accurat…

Cited by 20PDFcodeScholar
2020

SimAug: Learning Robust Representations from Simulation for Trajectory Prediction

ECCV 2020poster

This paper studies the problem of predicting future trajectories of people in unseen cameras of novel scenarios and views. We approach this problem through the real-data-free setting in which the model is trained only on 3D simulation data and applied out-of-the-box to a wide variety of real cameras…

2020

The Garden of Forking Paths: Towards Multi-Future Trajectory Prediction

CVPR 2020poster

This paper studies the problem of predicting the distribution over multiple possible future paths of people as they move through various visual scenes. We make two main contributions. The first contribution is a new dataset, created in a realistic 3D simulator, which is based on real world trajector…

Cited by 200PDFcodeScholar
2019

Peeking Into the Future: Predicting Future Person Activities and Locations in Videos

CVPR 2019poster

Deciphering human behaviors to predict their future paths/trajectories and what they would do from videos is important in many applications. Motivated by this idea, this paper studies predicting a pedestrian's future path jointly with future activities. We propose an end-to-end, multi-task learning…

Cited by 504PDFcodeScholar
2018

Focal Visual-Text Attention for Visual Question Answering

CVPR 2018poster

Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collections such as personal photos, we have to look at whole collections with sequences…

2017

Synchronization for multi-perspective videos in the wild

ICASSP 2017accepted

In the era of social media, a large number of user-generated videos are uploaded to the Internet every day, capturing events all over the world. Reconstructing the event truth based on information mined from these videos has been an emerging challenging task. Temporal alignment of videos “in the wil…

Cited by 0SourceScholar
2017

Temporal localization of audio events for conflict monitoring in social media

ICASSP 2017accepted

With the explosion in the availability of user-generated videos documenting any conflicts and human rights abuses around the world, analysts and researchers increasingly find themselves overwhelmed with massive amounts of video data to acquire and analyze useful information. In this paper, we develo…

Cited by 0SourceScholar
2015

Detecting semantic concepts in consumer videos using audio

ICASSP 2015accepted

With the increasing use of audio sensors in user generated content collection, how to detect semantic concepts using audio streams has become an important research problem. In this paper, we present a semantic concept annotation system using soundtracks/ audio of the video. We investigate three diff…

Cited by 0SourceScholar