← Search

Xin Lu

56 accepted papers

2026

CompEvent: Complex-valued Event-RGB Fusion for Low-light Video Enhancement and Deblurring

AAAI 2026technical

Low-light video deblurring poses significant challenges in applications like nighttime surveillance and autonomous driving due to dim lighting and long exposures. While event cameras offer potential solutions with superior low-light sensitivity and high temporal resolution, existing fusion methods t

Cited by 0SourcePDFScholar
2026

Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

CVPR 2026

Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking the essential global illumination in images and the inherent noise sensitivity of event signals in real-world scenarios. To address these issues, we

Cited by 0SourcecodeScholar
2026

GGBall: Graph Generative Model on Poincaré Ball

ICLR 2026poster

Generating graphs with hierarchical structures remains a fundamental challenge due to the limitations of Euclidean geometry in capturing exponential complexity. Here we introduce GGBall, a novel hyperbolic framework for graph generation that integrates geometric inductive biases with modern generati…

Cited by 0SourcecodeScholar
2026

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

CVPR 2026

Embodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specific RL training. However, existing zero-shot methods that build explicit 3D scene

Cited by 0SourceScholar
2026

Pose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth Priors

CVPR 2026

Omnidirectional 3D Gaussian Splatting with panoramas is a key technique for 3D scene representation, and existing methods typically rely on slow SfM to provide camera poses and sparse points priors. In this work, we propose a pose-free omnidirectional 3DGS method, named PFGS360, that reconstructs 3D

Cited by 0SourcecodeScholar
2026

SMAP: Semantic Route Planning with Map-Grounded Multimodal Alignment

CVPR 2026

Semantic route planning involves generating itineraries that align with user intent while respecting real-world spatial constraints. However, text-only large language models (LLMs) often hallucinate geographically implausible routes due to poor spatial grounding. Inspired by how humans use maps for

Cited by 0SourcecodeScholar
2025

Channel and space-based joint rate allocation algorithm

ICASSP 2025accepted

Rate control is a critical component for image and video compression Particularly under limited network bandwidth conditions, bitrate control is essential to ensure efficient image transmission by effectively allocation channel resources. In this research, since both Channel and Spatial have relatio…

Cited by 0SourceScholar
2025

How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation

NeurIPS 2025poster

Pre-trained language models represented by the Transformer have been proven to possess strong base capabilities, and the representative self-attention mechanism in the Transformer has become a classic in sequence modeling architectures. Different from the work of proposing sequence modeling architec…

Cited by 0SourceScholar
2025

InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity

ICCV 2025poster

Achieving flexible and high-fidelity identity-preserved image generation remains formidable, particularly with advanced Diffusion Transformers (DiTs) like FLUX. We introduce InfiniteYou (InfU), one of the earliest robust frameworks leveraging DiTs for this task. InfU addresses significant issues of…

2025

Learnable Frequency Decomposition for Image Forgery Detection and Localization

IJCAI 2025

Concern for image authenticity spurs research in image forgery detection and localization (IFDL). Most deep learning-based methods focus primarily on spatial domain modeling and have not fully explored frequency domain strategies. In this paper, we observe and analyze the frequency characteristic ch

Cited by 0SourcePDFScholar
2025

PanComplex: Leveraging Complex-Valued Neural Networks for Enhanced Pansharpening

IJCAI 2025

Pansharpening combines panchromatic and low-resolution multispectral images to generate high-resolution multispectral images. Previous studies have explored the connection between pansharpening and the frequency domain, but mostly in the real-valued domain, leaving the complex domain relatively unex

2025

ProDiff: Prototype-Guided Diffusion for Minimal Information Trajectory Imputation

ICML 2025poster

Trajectory data is crucial for various applications but often suffers from incompleteness due to device limitations and diverse collection scenarios. Existing imputation methods rely on sparse trajectory or travel information, such as velocity, to infer missing points. However, these approaches assu…

2025

PychoAgent: Psychology-driven LLM Agents for Explainable Panic Prediction on Social Media during Sudden Disaster Events

EMNLP 2025

Accurately predicting public panic sentiment on social media is crucial for proactive governance and crisis management. Current efforts on this problem face three main challenges: lack of finely annotated data hinders emotion prediction studies, unmodeled risk perception causes prediction inaccuraci

2025

Self-Foveate: Enhancing Diversity and Difficulty of Synthesized Instructions from Unsupervised Text via Multi-Level Foveation

ACL 2025finding

Large language models (LLMs) with instruction following capabilities have demonstrated impressive problem-solving abilities. While synthesizing instructional data from unsupervised text has become a common approach for training such models, conventional methods rely heavily on human effort for data…

2025

Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models

ACL 2025finding

Although large language models (LLMs) achieve effective safety alignment at the time of release, they still face various safety challenges. A key issue is that fine-tuning often compromises the safety alignment of LLMs. To address this issue, we propose a method named IRR (Identify, Remove, and Reca…

2025

SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters

NAACL 2025long

The widespread applications of large language models (LLMs) have brought about concerns regarding their potential misuse. Although aligned with human preference data before release, LLMs remain vulnerable to various malicious attacks. In this paper, we adopt a red-teaming strategy to enhance LLM saf…

2025

UHD-processer: Unified UHD Image Restoration with Progressive Frequency Learning and Degradation-aware Prompts

CVPR 2025poster

We introduce UHD-Processor, a unified and robust framework for all-in-one image restoration, which is particularly resource-efficient for Ultra-High-Definition (UHD) images. To address the limitations of traditional all-in-one methods that rely on complex restoration backbones, our strategy employs…

2024

Are U a Joke Master? Pun Generation via Multi-Stage Curriculum Learning towards a Humor LLM

ACL 2024findings

Although large language models (LLMs) acquire extensive world knowledge and some reasoning abilities, their proficiency in generating humorous sentences remains a challenge. Previous research has demonstrated that the humor generation capabilities of ChatGPT are confined to producing merely 25 uniqu…

2024

Design and Control of Rapid In-Air Reconfiguration for Modular Quadrotors With Full Controllable Degrees of Freedom

RA-L 2024

This letter introduces a novel reconfigurable modular aerial robot capable of in-air self-disassembly and achieving full-actuation control. The proposed robot utilizes a unique modular design, each module incorporates a vector tilting structure and an active undocking mechanism. The design of the ve

Cited by 15SourceScholar
2024

Dr. Bokeh: DiffeRentiable Occlusion-aware Bokeh Rendering

CVPR 2024poster

Bokeh is widely used in photography to draw attention to the subject while effectively isolating distractions in the background. Computational methods can simulate bokeh effects without relying on a physical camera lens but the inaccurate lens modeling in existing filtering-based methods leads to ar…

Cited by 8SourcePDFScholar
2024

Fast Intra Mode Prediction Algorithms for SCBS in VVC SCC

ICASSP 2024accepted

Versatile Video Coding (VVC) now supports Screen Content Coding (SCC) by integrating two efficient coding modes: Intra Block Copy (IBC) and Palette (PLT). However, the numerous modes and the Quad-Tree Plus Multi-Type Tree (QTMT) structure inherent to VVC contribute to a very high coding complexity.…

Cited by 0SourceScholar
2024

How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers

NeurIPS 2024poster

Pre-trained language models have been proven to possess strong base capabilities, which not only excel in in-distribution language modeling but also show powerful abilities in out-of-distribution language modeling, transfer learning and few-shot learning. Unlike existing work focusing on the influen…

Cited by 0SourcePDFScholar
2024

Real-Oriented Object Detection Driven by Intelligent Stockbreeding

ICASSP 2024accepted

Detecting objects with inherent orientations has numerous applications in the context of livestock reproduction. In this new scenario, the inherent orientations in the range [0, 2π) of target objects are detected alongside their bounding boxes to produce real-oriented bounding boxes. Due to the 0-to…

Cited by 0SourceScholar
2023

A Novel Mode Selection-Based Fast Intra Prediction Algorithm for Spatial SHVC

ICASSP 2023accepted

Due to multi-layer encoding and Inter-layer prediction, Spatial Scalable High-Efficiency Video Coding (SSHVC) has extremely high coding complexity. It is very crucial to improve its coding speed so as to promote widespread and cost-effective SSHVC applications. In this paper, we have proposed a nove…

Cited by 0SourceScholar
2023

A Topic-Enhanced Approach for Emotion Distribution Forecasting in Conversations

ICASSP 2023accepted

Emotion Forecasting in Conversations (EFC), the task aims to predict the emotion of next utterance (yet to come), has received more and more attention in recent years. However, this task ignores the one-to-many feature of dialogue and its prediction target is emotion label, which is flawed in most c…

Cited by 0SourceScholar
2023

Don’t Lose Yourself! Empathetic Response Generation via Explicit Self-Other Awareness

ACL 2023findings

As a critical step to achieve human-like chatbots, empathetic response generation has attained increasing interests. Previous attempts are incomplete and not sufficient enough to elicit empathy because they only stay on the initial stage of empathy to automatically sense and simulate the feelings an…

2023

On Efficient Transformer-Based Image Pre-training for Low-Level Vision

IJCAI 2023poster

Pre-training has marked numerous state of the arts in high-level computer vision, while few attempts have ever been made to investigate how pre-training acts in image processing systems. In this paper, we tailor transformer-based pre-training regimes that boost various low-level tasks. To comprehens…

2023

Precognition in Contextual Spoken Language Understanding via Knowledge Distillation

ICASSP 2023accepted

Task-oriented dialogue systems have become overwhelmingly popular in recent researches. Spoken Language Understanding (SLU) is widely used to extract the semantics frame of user queries and comprehend users’ intent/emotion/dialogue state in task-oriented dialogue systems. Most previous works on such…

Cited by 0SourceScholar
2022

A Unified Model for Multi-class Anomaly Detection

NeurIPS 2022accept

Despite the rapid advance of unsupervised anomaly detection, existing methods require to train separate models for different objects. In this work, we present UniAD that accomplishes anomaly detection for multiple classes with a unified framework. Under such a challenging setting, popular reconstruc…

2022

CauAIN: Causal Aware Interaction Network for Emotion Recognition in Conversations

IJCAI 2022poster

Emotion Recognition in Conversations has attained increasing interest in the natural language processing community. Many neural-network based approaches endeavor to solve the challenge of emotional dynamics in conversations and gain appealing results. However, these works are limited in capturing de…

Cited by 72SourcePDFScholar
2022

Dynamic Binary Neural Network by Learning Channel-Wise Thresholds

ICASSP 2022accepted

Binary neural networks (BNNs) constrain weights and activations to +1 or -1 with limited storage and computational cost, which is hardware-friendly for portable devices. Recently, BNNs have achieved remarkable progress and been adopted into various fields. However, the performance of BNNs is sensiti…

Cited by 0SourceScholar
2022

Improving Image Restoration by Revisiting Global Information Aggregation

ECCV 2022poster

"Global operations, such as global average pooling, are widely used in top-performance image restorers. They aggregate global information from input features along entire spatial dimensions but behave differently during training and inference in image restoration tasks: they are based on different r…

2021

Equalization Loss v2: A New Gradient Balance Approach for Long-Tailed Object Detection

CVPR 2021poster

Recently proposed decoupled training methods emerge as a dominant paradigm for long-tailed object detection. But they require an extra fine-tuning stage, and the disjointed optimization of representation and classifier might lead to suboptimal results. However, end-to-end training methods, like equa…

Cited by 211PDFcodeScholar
2021

RefineMask: Towards High-Quality Instance Segmentation With Fine-Grained Features

CVPR 2021poster

The two-stage methods for instance segmentation, e.g. Mask R-CNN, have achieved excellent performance recently. However, the segmented masks are still very coarse due to the downsampling operations in both the feature pyramid and the instance-wise pooling process, especially for large objects. In th…

Cited by 154PDFcodeScholar
2021

Retrieve, Discriminate and Rewrite: A Simple and Effective Framework for Obtaining Affective Response in Retrieval-Based Chatbots

EMNLP 2021finding

Obtaining affective response is a key step in building empathetic dialogue systems. This task has been studied a lot in generation-based chatbots, but the related research in retrieval-based chatbots is still in the early stage. Existing works in retrieval-based chatbots are based on Retrieve-and-Re…

2020

An Iterative Emotion Interaction Network for Emotion Recognition in Conversations

COLING 2020main

Emotion recognition in conversations (ERC) has received much attention recently in the natural language processing community. Considering that the emotions of the utterances in conversations are interactive, previous works usually implicitly model the emotion interaction between utterances by modeli…

2020

MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object Detection

ECCV 2020poster

Modern object detection methods can be divided into one-stage approaches and two-stage ones. One-stage detectors are more efficient owing to straightforward architectures, but the two-stage detectors still take the lead in accuracy. Although recent work try to improve the one-stage detectors by imit…

Cited by 95SourcePDFScholar
2019

Free-Form Image Inpainting With Gated Convolution

ICCV 2019oral

We present a generative image inpainting system to complete images with free-form mask and guidance. The system is based on gated convolutions learned from millions of images without additional labelling efforts. The proposed gated convolution solves the issue of vanilla convolution that treats all…

Cited by 2386PDFcodeScholar
2019

Grid R-CNN

CVPR 2019poster

This paper proposes a novel object detection framework named Grid R-CNN, which adopts a grid guided localization mechanism for accurate object detection. Different from the traditional regression based methods, the Grid R-CNN captures the spatial information explicitly and enjoys the position sensit…

Cited by 607PDFScholar
2019

Semantic Component Decomposition for Face Attribute Manipulation

CVPR 2019poster

Deep neural network-based methods were proposed for face attribute manipulation. There still exist, however, two major issues, i.e., insufficient visual quality (or resolution) of the results and lack of user control. They limit the applicability of existing methods since users may have different ed…

Cited by 47PDFScholar
2018

Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias

ECCV 2018poster

While machine learning approaches to visual emotion recognition offer great promise, current methods consider training and testing models on small scale datasets covering limited visual emotion concepts. Our analysis identifies an important but long overlooked issue of existing visual emotion benchm…

Cited by 106SourcePDFScholar
2018

Flow-Grounded Spatial-Temporal Video Prediction from Still Images

ECCV 2018poster

Existing video prediction methods mainly rely on observing multiple historical frames or focus on predicting the next one-frame. In this work, we study the problem of generating consecutive multiple future frames by observing one single still image only. We formulate the multi-frame prediction task…

Cited by 159SourcePDFScholar
2018

Generative Image Inpainting With Contextual Attention

CVPR 2018poster

Recent deep learning based approaches have shown promising results for the challenging task of inpainting large missing regions in an image. These methods can generate visually plausible image structures and textures, but often create distorted structures or blurry textures inconsistent with surroun…

2018

MAttNet: Modular Attention Network for Referring Expression Comprehension

CVPR 2018poster

In this paper, we address referring expression comprehension: localizing an image region described by a natural language expression. While most recent work treats expressions as a single unit, we propose to decompose them into three modular components related to subject appearance, location, and re…

2018

Rethinking the Smaller-Norm-Less-Informative Assumption in Channel Pruning of Convolution Layers

ICLR 2018poster

Model pruning has become a useful technique that improves the computational efficiency of deep learning, making it possible to deploy solutions in resource-limited scenarios. A widely-used practice in relevant work assumes that a smaller-norm parameter or feature plays a less informative role at the…

2017

Diversified Texture Synthesis With Feed-Forward Networks

CVPR 2017spotlight

Recent progresses on deep discriminative and generative modeling have shown promising results on texture synthesis. However, existing feed-forward based methods trade off generality for efficiency, which suffer from many issues, such as shortage of generality (i.e., build one network per texture), l…

Cited by 341PDFScholar
2017

High-Resolution Image Inpainting Using Multi-Scale Neural Patch Synthesis

CVPR 2017poster

Recent advances in deep learning have shown exciting promise in filling large holes in natural images with semantically plausible and context aware details, impacting fundamental image manipulation tasks such as object removal. While these learning-based methods are significantly more effective in c…

Cited by 1114PDFScholar
2017

Recurrent Multimodal Interaction for Referring Image Segmentation

ICCV 2017poster

In this paper we are interested in the problem of image segmentation given natural language descriptions, i.e. referring expressions. Existing works tackle this problem by first modeling images and sentences independently and then segment images by combining these two types of representations. We ar…

Cited by 296PDFcodeScholar
2017

Scene Parsing With Global Context Embedding

ICCV 2017poster

We present a scene parsing method that utilizes global context information based on both the parametric and non-parametric models. Compared to previous methods that only exploit the local relationship between objects, we train a context network based on scene similarities to generate feature represe…

Cited by 70PDFcodeScholar
2017

Universal Style Transfer via Feature Transforms

NeurIPS 2017poster

Universal style transfer aims to transfer arbitrary visual styles to content images. Existing feed-forward based methods, while enjoying the inference efficiency, are mainly limited by inability of generalizing to unseen styles or compromised visual quality. In this paper, we present a simple yet ef…

2015

Deep Multi-Patch Aggregation Network for Image Style, Aesthetics, and Quality Estimation

ICCV 2015poster

This paper investigates problems of image style, aesthetics, and quality estimation, which require fine-grained details from high-resolution images, utilizing deep neural network training approach. Existing deep convolutional neural networks mostly extracted one patch such as a down-sized crop from…

Cited by 399PDFcodeScholar