← Search

Xu Jia

46 accepted papers

2026

Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos

CVPR 2026

Understanding physical transformation processes is crucial for both human cognition and artificial intelligence systems, particularly from an egocentric perspective, which serves as a key bridge between humans and machines in action modeling. We define this modeling process as Egocentric Instructed

Cited by 0SourceScholar
2026

MultiShotMaster: A Controllable Multi-Shot Video Generation Framework

CVPR 2026

Current video generation techniques excel at single-shot clips but struggle to produce narrative multi-shot videos, which require flexible shot arrangement, coherent narrative, and controllability beyond text prompts. To tackle these challenges, we propose MultiShotMaster, a framework for highly con

Cited by 0SourcecodeScholar
2026

RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer

ICASSP 2026oral

Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamic…

Cited by 0SourcePDFScholar
2026

Reinforcing Video Object Segmentation to Think before it Segments

CVPR 2026

Video reasoning segmentation (VRS) endeavors to delineate referred objects in videos guided by implicit instructions that encapsulate human intent and temporal logic. Previous approaches leverage large vision language models (LVLMs) to encode object semantics into \SEG tokens for mask prediction. Ho

Cited by 0SourceScholar
2026

Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment

AAAI 2026technical

As super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical f

Cited by 0SourcePDFScholar
2025

CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

ICCV 2025poster

Recent advances in 3D reconstruction techniques and vision-language models have fueled significant progress in 3D semantic understanding, a capability critical to robotics, autonomous driving, and virtual/augmented reality. However, methods that rely on 2D priors are prone to a critical challenge: c…

2025

CPRM: A LLM-based Continual Pre-training Framework for Relevance Modeling in Commercial Search

NAACL 2025industry

Relevance modeling between queries and items stands as a pivotal component in commercial search engines, directly affecting the user experience. Given the remarkable achievements of large language models (LLMs) in various natural language processing (NLP) tasks, LLM-based relevance modeling is gradu…

2025

EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images

ICCV 2025poster

3D Gaussian Splatting (3D-GS) has demonstrated exceptional capabilities in synthesizing novel views of 3D scenes. However, its training is heavily reliant on high-quality images and precise camera poses. Meeting these criteria can be challenging in non-ideal real-world conditions, where motion-blurr…

Cited by 0SourcePDFScholar
2025

ReNeg: Learning Negative Embedding with Reward Guidance

CVPR 2025highlight

In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In…

2025

Towards Survivability in Complex Motion Scenarios: RGB-Event Object Tracking via Historical Trajectory Prompting

ICRA 2025

Event data has recently emerged as a valuable complement to object tracking, offering dense temporal resolution and a high dynamic range. However, existing RGB-Event trackers struggle with targets exhibiting complex motion trajectories, where RGB features alone fail to provide sufficient discriminat

Cited by 4SourcecodeScholar
2025

Using Review Combination and Pseudo-Tokens for Aspect Sentiment Quad Prediction

NAACL 2025findings

Aspect Sentiment Quad Prediction (ASQP) aims to identify quadruples consisting of an aspect term, aspect category, opinion term, and sentiment polarity from a given sentence, which is the most representative and challenging task in aspect-based sentiment analysis. A major challenge arises when impli…

2025

VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

ICCV 2025accepted

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos…

2024

EvSign: Sign Language Recognition and Translation with Streaming Events

ECCV 2024poster

"Sign language is one of the most effective communication tools for people with hearing difficulties. Most existing works focus on improving the performance of sign language tasks on RGB videos, which may suffer from degraded recording conditions, such as fast movement of hands with motion blur and…

2024

GenTKG: Generative Forecasting on Temporal Knowledge Graph with Large Language Models

NAACL 2024findings

The rapid advancements in large language models (LLMs) have ignited interest in the temporal knowledge graph (tKG) domain, where conventional embedding-based and rule-based methods dominate. The question remains open of whether pre-trained LLMs can understand structured temporal relational data and…

2024

Retrieval and Reasoning on KGs: Integrate Knowledge Graphs into Large Language Models for Complex Question Answering

EMNLP 2024finding

Despite Large Language Models (LLMs) have performed impressively in various Natural Language Processing (NLP) tasks, their inherent hallucination phenomena severely challenge their credibility in complex reasoning. Combining explainable Knowledge Graphs (KGs) with LLMs is a promising path to address…

Cited by 7SourcePDFScholar
2024

SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning

ECCV 2024poster

"Parameter-efficient transfer learning (PETL) has emerged as a flourishing research field for adapting large pre-trained models to downstream tasks, greatly reducing trainable parameters while grappling with memory challenges during fine-tuning. To address it, memory-efficient series (METL) avoid ba…

2024

UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory

CVPR 2024poster

Parameter-efficient transfer learning (PETL) i.e. fine-tuning a small portion of parameters is an effective strategy for adapting pre-trained models to downstream domains. To further reduce the memory demand recent PETL works focus on the more valuable memory-efficient characteristic. In this paper…

2023

A Benchmark for Evaluating Robustness of Spoken Language Understanding Models in Slot Filling

ICASSP 2023accepted

Slot filling is a major problem in spoken language understanding (SLU) task. However, the current SLU models may experience performance degradation when encountering unfamiliar data in different datasets. Meanwhile, as recent models are becoming more complex, retraining the model in a new applicatio…

Cited by 0SourceScholar
2023

Compression-Aware Video Super-Resolution

CVPR 2023poster

Videos stored on mobile devices or delivered on the Internet are usually in compressed format and are of various unknown compression parameters, but most video super-resolution (VSR) methods often assume ideal inputs resulting in large performance gap between experimental settings and real-world app…

2023

Dual Memory Aggregation Network for Event-Based Object Detection with Learnable Representation

AAAI 2023technical

Event-based cameras are bio-inspired sensors that capture brightness change of every pixel in an asynchronous manner. Compared with frame-based sensors, event cameras have microsecond-level latency and high dynamic range, hence showing great potential for object detection under high-speed motion and…

2023

GM-NeRF: Learning Generalizable Model-Based Neural Radiance Fields From Multi-View Images

CVPR 2023poster

In this work, we focus on synthesizing high-fidelity novel view images for arbitrary human performers, given a set of sparse multi-view images. It is a challenging task due to the large variation among articulated body poses and heavy self-occlusions. To alleviate this, we introduce an effective gen…

2023

Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion

ACL 2023findings

Temporal Knowledge graph completion (TKGC) is a crucial task that involves reasoning at known timestamps to complete the missing part of facts and has attracted more and more attention in recent years. Most existing methods focus on learning representations based on graph neural networks while inacc…

2022

AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-Time Image Enhancement

CVPR 2022poster

The 3D Lookup Table (3D LUT) is a highly-efficient tool for real-time image enhancement tasks, which models a non-linear 3D color transform by sparsely sampling it into a discretized 3D lattice. Previous works have made efforts to learn image-adaptive output color values of LUTs for flexible enhance…

Cited by 82PDFcodeScholar
2022

Class-Balanced Pixel-Level Self-Labeling for Domain Adaptive Semantic Segmentation

CVPR 2022poster

Domain adaptive semantic segmentation aims to learn a model with the supervision of source domain data, and produce satisfactory dense predictions on unlabeled target domain. One popular solution to this challenging task is self-training, which selects high-scoring predictions on target samples as p…

Cited by 111PDFcodeScholar
2022

Look Back and Forth: Video Super-Resolution With Explicit Temporal Difference Modeling

CVPR 2022poster

Temporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the model complexity and might fail in case of occlusion or com…

Cited by 60PDFcodeScholar
2022

TimeReplayer: Unlocking the Potential of Event Cameras for Video Interpolation

CVPR 2022poster

Recording fast motion in a high FPS (frame-per-second) requires expensive high-speed cameras. As an alternative, interpolating low-FPS videos from commodity cameras has attracted significant attention. If only low-FPS videos are available, motion assumptions (linear or quadratic) are necessary to in…

Cited by 37PDFScholar
2021

FoV-Net: Field-of-View Extrapolation Using Self-Attention and Uncertainty

RA-L 2021

The ability to make educated predictions about their surroundings, and associate them with certain confidence, is important for intelligent systems, like autonomous vehicles and robots. It allows them to plan early and decide accordingly. Motivated by this observation, in this letter we utilize info

Cited by 7SourcecodeScholar
2021

Multi-Source Domain Adaptation With Collaborative Learning for Semantic Segmentation

CVPR 2021poster

Multi-source unsupervised domain adaptation (MSDA) aims at adapting models trained on multiple labeled source domains to an unlabeled target domain. In this paper, we propose a novel multi-source domain adaptation framework based on collaborative learning for semantic segmentation. Firstly, a simple…

Cited by 106PDFScholar
2021

Multi-Target Domain Adaptation With Collaborative Consistency Learning

CVPR 2021poster

Recently unsupervised domain adaptation for the semantic segmentation task has become more and more popular due to the high-cost of pixel-level annotation on real-world images. However, most domain adaptation methods are only restricted to single-source-single-target pair, and can not be directly ex…

Cited by 108PDFcodeScholar
2021

Neighbor2Neighbor: Self-Supervised Denoising From Single Noisy Images

CVPR 2021poster

In the last few years, image denoising has benefited a lot from the fast development of neural networks. However, the requirement of large amounts of noisy-clean image pairs for supervision limits the wide use of these models. Although there have been a few attempts in training an image denoising mo…

Cited by 451PDFcodeScholar
2021

Semi-Supervised Domain Adaptation Based on Dual-Level Domain Mixing for Semantic Segmentation

CVPR 2021poster

Data-driven based approaches, in spite of great success in many tasks, have poor generalization when applied to unseen image domains, and require expensive cost of annotation especially for dense pixel prediction tasks such as semantic segmentation. Recently, both unsupervised domain adaptation (UDA…

Cited by 79PDFScholar
2021

T-SVDNet: Exploring High-Order Prototypical Correlations for Multi-Source Domain Adaptation

ICCV 2021poster

Most existing domain adaptation methods focus on adaptation from only one source domain, however, in practice there are a number of relevant sources that could be leveraged to help improve performance on target domain. We propose a novel approach named T-SVDNet to address the task of Multi-source Do…

Cited by 57PDFcodeScholar
2020

More Classifiers, Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning

ECCV 2020poster

Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning","Overcoming catastrophic forgetting in neural networks is a long-standing and core research objective for incremental learning. Notable studies have shown regularization strategies enable the network to remember previousl…

2020

Unsupervised Model Personalization While Preserving Privacy and Scalability: An Open Problem

CVPR 2020poster

This work investigates the task of unsupervised model personalization, adapted to continually evolving, unlabeled local user images. We consider the practical scenario where a high capacity server interacts with a myriad of resource-limited edge devices, imposing strong requirements on scalability a…

Cited by 35PDFcodeScholar
2020

Video Super-Resolution With Temporal Group Attention

CVPR 2020poster

Video super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, has recently drawn increasing attention. In this work, we propose a novel method that can effectively incorporate temporal information in a hierarchical way. The input sequence is d…

Cited by 220PDFcodeScholar
2020

Video Super-Resolution with Recurrent Structure-Detail Network

ECCV 2020poster

Most video super-resolution methods super-resolve a single reference frame with the help of neighboring frames in a temporal sliding window. They are less efficient compared to the recurrent-based methods. In this work, we propose a novel recurrent video super-resolution method which is both effecti…

2019

Co-Evolutionary Compression for Unpaired Image Translation

ICCV 2019poster

Generative adversarial networks (GANs) have been successfully used for considerable computer vision tasks, especially the image-to-image translation. However, generators in these networks are of complicated architectures with large number of parameters and huge computational complexities. Existing m…

Cited by 93PDFScholar
2019

Exemplar Guided Unsupervised Image-to-Image Translation with Semantic Consistency

ICLR 2019poster

Image-to-image translation has recently received significant attention due to advances in deep learning. Most works focus on learning either a one-to-one mapping in an unsupervised way or a many-to-many mapping in a supervised way. However, a more practical setting is many-to-many mapping in an unsu…

Cited by 165SourcePDFScholar
2019

Video Generation From Single Semantic Label Map

CVPR 2019poster

This paper proposes the novel task of video generation conditioned on a SINGLE semantic label map, which provides a good balance between flexibility and quality in the generation process. Different from typical end-to-end approaches, which model both scene content and dynamics in a single step, we p…

Cited by 128PDFcodeScholar
2017

Pose Guided Person Image Generation

NeurIPS 2017poster

This paper proposes the novel Pose Guided Person Generation Network (PG$^2$) that allows to synthesize person images in arbitrary poses, based on an image of that person and a novel pose. Our generation framework PG$^2$ utilizes the pose information explicitly and consists of two key stages: pose in…

Cited by 1080SourcePDFScholar
2015

Guiding the Long-Short Term Memory Model for Image Caption Generation

ICCV 2015poster

In this work we focus on the problem of image caption generation. We propose an extension of the long short term memory (LSTM) model, which we coin gLSTM for short. In particular, we add semantic information extracted from the image as extra input to each unit of the LSTM block, with the aim of gui…

Cited by 587PDFcodeScholar