← Search

Jin Chen

16 accepted papers

2026

Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval

AAAI 2026technical

In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in ali

Cited by 0SourcePDFScholar
2026

Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration

RSS 2026poster

Human demonstrations offer rich environmental diversity and scale naturally, making them an appealing alternative to robot teleoperation. While this paradigm has advanced robot-arm manipulation, its potential for the more challenging, data-hungry problem of humanoid loco-manipulation remains largely…

Cited by 0SourceScholar
2026

VUDG: A Dataset for Video Understanding Domain Generalization

ICLR 2026poster

Video understanding has made remarkable progress in recent years, largely driven by advances in deep models and the availability of large-scale annotated datasets. However, the robustness of these models to domain shifts encountered in real-world video applications remains a critical yet underexplor…

Cited by 0SourceScholar
2026

WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control

ICLR 2026poster

Humanoid robots require precise locomotion and dexterous manipulation to per- form challenging locomanipulation tasks. Yet existing approaches, modular or end-to-end, are deficient in manipulation-aware locomotion. This confines the robot to a limited workspace, preventing it from performing large-s…

Cited by 0SourcecodeScholar
2025

Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning

NeurIPS 2025poster

Open-vocabulary learning requires modeling the data distribution in open environments, which consists of both seen-class and unseen-class data. Existing methods estimate the distribution in open environments using seen-class data, where the absence of unseen classes makes the estimation error inhe…

Cited by 0SourceScholar
2025

Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification

AAAI 2025technical

Multi-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly…

2025

ViCo: A Multitask Video-enhanced and Cognition-preserving Modality Alignment Training Framework

ICASSP 2025accepted

The rapid development of multimodal large language models (MLLMs) has brought significant breakthroughs to this field. However, current MLLMs typically rely on vision instruction tuning based on large language models (LLMs) to endow them with multimodal capabilities, which may lead to low video util…

Cited by 0SourceScholar
2024

Learning-Efficient Yet Generalizable Collaborative Filtering for Item Recommendation

ICML 2024poster

The weighted squared loss is a common component in several Collaborative Filtering (CF) algorithms for item recommendation, including the representative implicit Alternating Least Squares (iALS). Despite its widespread use, this loss function lacks a clear connection to ranking objectives such as Di…

Cited by 4SourcePDFScholar
2023

Knowledge Distillation for High Dimensional Search Index

NeurIPS 2023poster

Lightweight compressed models are prevalent in Approximate Nearest Neighbor Search (ANNS) and Maximum Inner Product Search (MIPS) owing to their superiority of retrieval efficiency in large-scale datasets. However, results given by compressed methods are less accurate due to the curse of dimension a…

Cited by 7SourcePDFScholar
2022

Adaptive Image-to-Video Scene Graph Generation via Knowledge Reasoning and Adversarial Learning

AAAI 2022technical

Scene graph in a video conveys a wealth of information about objects and their relationships in the scene, thus benefiting many downstream tasks such as video captioning and visual question answering. Existing methods of scene graph generation require large-scale training videos annotated with objec…

Cited by 1SourcePDFScholar
2022

Cache-Augmented Inbatch Importance Resampling for Training Recommender Retriever

NeurIPS 2022accept

Recommender retrievers aim to rapidly retrieve a fraction of items from the entire item corpus when a user query requests, with the representative two-tower model trained with the log softmax loss. For efficiently training recommender retrievers on modern hardwares, inbatch sampling, where the items…

Cited by 12SourcePDFScholar
2021

Efficient Optimal Selection for Composited Advertising Creatives with Tree Structure

AAAI 2021technical

Ad creatives are one of the prominent mediums for online e-commerce advertisements. Ad creatives with enjoyable visual appearance may increase the click-through rate (CTR) of products. Ad creatives are typically handcrafted by advertisers and then delivered to the advertising platforms for advertise…

2021

Spatial-temporal Causal Inference for Partial Image-to-video Adaptation

AAAI 2021technical

Image-to-video adaptation leverages off-the-shelf learned models in labeled images to help classification in unlabeled videos, thus alleviating the high computation overhead of training a video classifier from scratch. This task is very challenging since there exist two types of domain shifts betwee…

2020

Adaptive Graph Convolutional Network With Attention Graph Clustering for Co-Saliency Detection

CVPR 2020poster

Co-saliency detection aims to discover the common and salient foregrounds from a group of relevant images. For this task, we present a novel adaptive graph convolutional network with attention graph clustering (GCAGC). Three major contributions have been made, and are experimentally shown to have su…

Cited by 127PDFScholar