← Search

Yang Jin

20 accepted papers

2026

A Unified Self-Regulating Training Framework for Federated Deep Reinforcement Learning

AAAI 2026technical

Federated Deep Reinforcement Learning (FDRL) aims to enable distributed collaborative training of multiple DRL models while preserving privacy. Existing FDRL methods function in static client environments, but real-world scenarios often involve dynamic state transitions, such as noise, which render

Cited by 0SourcePDFScholar
2026

SOE: Sample-Efficient Robot Policy Self-Improvement Via On-Manifold Exploration

ICRA 2026poster

Intelligent agents progress by continually refining their capabilities through actively exploring environments. Yet robot policies often lack sufficient exploration capability due to action mode collapse. Existing methods that encourage exploration typically rely on random perturbations, which are u…

2025

DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language Model

IROS 2025

Generating robot demonstrations through simulation is widely recognized as an effective way to scale up robot data. Previous work often trained reinforcement learning agents to generate expert policies, but this approach lacks sample efficiency. Recently, a line of work has attempted to generate rob

Cited by 3SourceScholar
2025

Enhancing Consistency of Flow-Based Image Editing through Kalman Control

NeurIPS 2025poster

Flow-based generative models have gained popularity for image generation and editing. For instruction-based image editing, it is critical to ensure that modifications are confined to the targeted regions. Yet existing methods often fail to maintain consistency in non-targeted regions between the ori…

Cited by 0SourceScholar
2025

Granularity-Adaptive Spatial Evidence Tokenization for Video Question Answering

AAAI 2025technical

Video question answering plays a vital role in computer vision, and recent advances in large language models have further propelled the development of this field. However, existing video question answering techniques often face limitations in grasping fine-grained video content in spatial dimensions…

Cited by 0SourcePDFScholar
2025

Knowledge-Driven Imitation Learning: Enabling Generalization Across Diverse Conditions

IROS 2025

Imitation learning has emerged as a powerful paradigm in robot manipulation, yet its generalization capability remains constrained by object-specific dependencies in limited expert demonstrations. To address this challenge, we propose knowledge-driven imitation learning, a framework that leverages e

Cited by 1SourcecodeScholar
2025

Pyramidal Flow Matching for Efficient Video Generative Modeling

ICLR 2025poster

Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational de…

2025

SIME: Enhancing Policy Self-Improvement with Modal-level Exploration

IROS 2025

Self-improvement requires robotic systems to initially learn from human-provided data and then gradually enhance their capabilities through interaction with the environment. This is similar to how humans improve their skills through continuous practice. However, achieving effective self-improvement

Cited by 4SourcecodeScholar
2024

Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model

ECCV 2024poster

"Gaze object prediction (GOP) aims to predict the category and location of the object that a human is looking at. Previous methods utilized box-level supervision to identify the object that a person is looking at, but struggled with semantic ambiguity, , a single box may contain several items since…

2024

CTA-LO: Accurate and Robust LiDAR Odometry Using Continuous-Time Adaptive Estimation

ICRA 2024poster

Accurate and robust LiDAR odometry is a crucial technology for robot localization. However, motion distortion and ranging error make it a bottleneck. Most existing methods are limited in accuracy and robustness because they simply compensate for motion distortion by constant velocity motion assumpti…

Cited by 0SourceScholar
2024

Harder Task Needs More Experts: Dynamic Routing in MoE Models

ACL 2024long

In this paper, we introduce a novel dynamic expert selection framework for Mixture of Experts (MoE) models, aiming to enhance computational efficiency and model performance by adjusting the number of activated experts based on input difficulty. Unlike existing MoE approaches that rely on fixed TopK…

2024

RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance

NeurIPS 2024poster

Customizing diffusion models to generate identity-preserving images from user-provided reference images is an intriguing new problem. The prevalent approaches typically require training on extensive domain-specific images to achieve identity preservation, which lacks flexibility across different use…

2024

TransGOP: Transformer-Based Gaze Object Prediction

AAAI 2024technical

Gaze object prediction aims to predict the location and category of the object that is watched by a human. Previous gaze object prediction works use CNN-based object detectors to predict the object's location. However, we find that Transformer-based object detectors can predict more accurate object…

2024

Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization

ICLR 2024poster

Recently, the remarkable advance of the Large Language Model (LLM) has inspired researchers to transfer its extraordinary reasoning capability to both vision and language data. However, the prevailing approaches primarily regard the visual input as a prompt and focus exclusively on optimizing the te…

2024

Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization

ICML 2024oral

In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for effective large-scale pre-training due to the modeling of its…

2023

Learning Instance-Level Representation for Large-Scale Multi-Modal Pretraining in E-Commerce

CVPR 2023poster

This paper aims to establish a generic multi-modal foundation model that has the scalable capability to massive downstream applications in E-commerce. Recently, large-scale vision-language pretraining approaches have achieved remarkable advances in the general domain. However, due to the significant…

Cited by 13SourcePDFScholar
2022

Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding

NeurIPS 2022accept

Spatio-Temporal video grounding (STVG) focuses on retrieving the spatio-temporal tube of a specific object depicted by a free-form textual expression. Existing approaches mainly treat this complicated task as a parallel frame-grounding problem and thus suffer from two types of inconsistency drawback…

2020

Beyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global Context

CVPR 2020poster

Video visual relation detection (VidVRD) aims to describe all interacting objects in a video. Different from relationships in static images, videos contain an addition temporal channel. A majority of existing works divide a video into short segments, predict relationships in each segment, and merge…

Cited by 92PDFScholar