← Search

Changyin Sun

13 accepted papers

2025

Boosting Efficient Reinforcement Learning for Vision-and-Language Navigation With Open-Sourced LLM

RA-L 2025

Vision-and-Language Navigation (VLN) requires an agent to navigate in photo-realistic environments based on language instructions. Existing methods typically employ imitation learning to train agents. However, approaches based on recurrent neural networks suffer from poor generalization, while trans

Cited by 13SourceScholar
2025

EDCFlow: Exploring Temporally Dense Difference Maps for Event-based Optical Flow Estimation

CVPR 2025poster

Recent learning-based methods for event-based optical flow estimation utilize cost volumes for pixel matching but suffer from redundant computations and limited scalability to higher resolutions for flow refinement. In this work, we take advantage of the complementarity between temporally dense feat…

Cited by 0SourcePDFScholar
2024

ACAMDA: Improving Data Efficiency in Reinforcement Learning through Guided Counterfactual Data Augmentation

AAAI 2024technical

Data augmentation plays a crucial role in improving the data efficiency of reinforcement learning (RL). However, the generation of high-quality augmented data remains a significant challenge. To overcome this, we introduce ACAMDA (Adversarial Causal Modeling for Data Augmentation), a novel framework…

Cited by 6SourcePDFScholar
2024

Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill

ICRA 2024poster

Zero-shot object navigation is a challenging task for home-assistance robots. This task emphasizes visual grounding, commonsense inference and locomotion abilities, where the first two are inherent in foundation models. But for the locomotion part, most works still depend on map-based planning appro…

Cited by 39SourcecodeScholar
2024

Point-to-Spike Residual Learning for Energy-Efficient 3D Point Cloud Classification

AAAI 2024technical

Spiking neural networks (SNNs) have revolutionized neural learning and are making remarkable strides in image analysis and robot control tasks with ultra-low power consumption advantages. Inspired by this success, we investigate the application of spiking neural networks to 3D point cloud processing…

Cited by 12SourcePDFScholar
2024

Temporal Correlation Vision Transformer for Video Person Re-Identification

AAAI 2024technical

Video Person Re-Identification (Re-ID) is a task of retrieving persons from multi-camera surveillance systems. Despite the progress made in leveraging spatio-temporal information in videos, occlusion in dense crowds still hinders further progress. To address this issue, we propose a Temporal Correla…

Cited by 5SourcePDFScholar
2023

Robust Navigation with Cross-Modal Fusion and Knowledge Transfer

ICRA 2023poster

Recently, learning-based approaches show promising results in navigation tasks. However, the poor generalization capability and the simulation-reality gap prevent a wide range of applications. We consider the problem of improving the generalization of mobile robots and achieving sim-to-real transfer…

Cited by 1SourcecodeScholar
2022

Learning Temporally Causal Latent Processes from General Temporal Data

ICLR 2022poster

Our goal is to recover time-delayed latent causal variables and identify their relations from measured temporal data. Estimating causally-related latent variables from observations is particularly challenging as the latent variables are not uniquely recoverable in the most general case. In this work…

2017

Semantic Regularisation for Recurrent Image Annotation

CVPR 2017poster

The "CNN-RNN" design pattern is increasingly widely applied in a variety of image annotation tasks including multi-label classification and captioning. Existing models use the weakly semantic CNN hidden layer or its transform as the image embedding that provides the interface between the CNN and RN…

Cited by 136PDFScholar