← Search

Xin Jin

119 accepted papers

2026

ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning

CVPR 2026

The introduction of negative labels (NLs) has proven effective in enhancing Out-of-Distribution (OOD) detection. However, existing methods often lack an understanding of OOD images, making it difficult to construct an accurate negative space. Furthermore, the absence of negative labels semantically

Cited by 0SourcecodeScholar
2026

DR-GGAD: Dual Residual Centering for Mitigating Anomaly Non‑Discriminativity in Generalist Graph Anomaly Detection

ICLR 2026poster

Generalist Graph Anomaly Detection (GGAD) seeks a unified representation learning model to detect anomalies in unseen graphs, but cross-domain transfer often entangles the learned anomalous and normal representations. We formalize this degradation as Anomaly non-Discriminativity (AnD) and define a n…

Cited by 0SourceScholar
2026

Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield

ICLR 2026poster

Diffusion model distillation has emerged as a powerful technique for creating efficient few-step and single-step generators. Among these, Distribution Matching Distillation (DMD) and its variants stand out for their impressive performance, which is widely attributed to their core mechanism of matchi…

Cited by 0SourceScholar
2026

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

ICLR 2026poster

Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma–misalignment of 2D image forecasting and 3D action prediction. Besides, such a vision-action entangled training manner limits model learning from large-scale, action-free web video…

Cited by 0SourcecodeScholar
2026

EarlyTom: Early Token Compression Completes Fast Video Understanding

CVPR 2026

Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficiency introduced by processing massive amounts of visual tokens. Although recent approaches achieve extremely low token ret

Cited by 0SourceScholar
2026

GI-GCN: Global Interacted Graph Convolutional Networks via Dominant Sets for Graph Classification

ICML 2026poster

Graph Convolutional Networks (GCNs) are defined based on aggregating the node information of adjacent nodes, that are usually treated as equally important as each other, limiting the representational power of existing GCNs for graph classification. To address this shortcoming, we propose a novel Glo…

Cited by 0SourceScholar
2026

Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning

CVPR 2026

Reinforcement Learning (RL) has achieved remarkable success in various domains, yet it often relies on carefully designed programmatic reward functions to guide agent behavior. Designing such reward functions can be challenging and may not generalize well across different tasks. To address this limi

Cited by 0SourceScholar
2026

Hierarchical Dual-Domain Fusion with Frequency-Guided Spatial Modeling for Pan-Sharpening

AAAI 2026technical

Pan-sharpening aims to generate high-resolution multispectral images by integrating the spectral richness of low-resolution multispectral images with the spatial details of high-resolution panchromatic images. Although frequency-domain modeling shows great potential in this field, most existing meth

Cited by 0SourcePDFScholar
2026

ImagiDrive: A Unified Imagination-And-Planning Framework for Autonomous Driving

ICRA 2026poster

Autonomous driving requires rich contextual comprehension and precise predictive reasoning to navigate dynamic and complex environments safely. Vision-Language Models (VLMs) and Driving World Models (DWMs) have independently emerged as powerful recipes addressing different aspects of this challenge.…

2026

InteractComp: Evaluating Search Agents With Ambiguous Queries

ICML 2026poster

Language agents have demonstrated remarkable potential in web search and information retrieval. However, these search agents assume user queries are complete and unambiguous, an assumption that diverges from reality where users begin with incomplete queries requiring clarification through interactio…

Cited by 0SourceScholar
2026

Learning AND–OR Templates for Compositional Representation in Art and Design

ICLR 2026poster

This work proposes a compositional AND–OR template for art and design that encodes the part–relation–geometry organization of images in a structured and interpretable form. Within a maximum-entropy log-linear model, we define a unified consistency score as log-likelihood gain against a reference dis…

Cited by 0SourceScholar
2026

MFmamba: A Multi-function Network for Panchromatic Image Resolution Restoration Based on State-Space Model

AAAI 2026technical

Remote sensing images are becoming increasingly widespread in military, earth resource exploration. Because of the limitation of a single sensor, we can obtain high spatial resolution grayscale panchromatic (PAN) images and low spatial resolution color multispectral (MS) images. Therefore, an import

Cited by 0SourcePDFScholar
2026

MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding

ICLR 2026poster

Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language models (MLLMs) in the post-training stage, supervised fine-tuning (SFT) is a stable choice but requires human annotations…

Cited by 0SourcecodeScholar
2026

ORV: 4D Occupancy-centric Robot Video Generation

CVPR 2026

Recent embodied intelligence suffers from data scarcity, while conventional simulators lack visual realism. Controllable video generation is emerging as a promising data engine, yet current action-conditioned methods still fall short: generated videos are limited in fidelity and temporal consistency

Cited by 0SourcecodeScholar
2026

OmniFood8K: Single-Image Nutrition Estimation via Hierarchical Frequency-Aligned Fusion

CVPR 2026

Accurate estimation of food nutrition plays a vital role in promoting healthy dietary habits and personalized diet management. Most existing food datasets primarily focus on Western cuisines and lack sufficient coverage of Chinese dishes, which restricts accurate nutritional estimation for Chinese m

Cited by 0SourcecodeScholar
2026

PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations

CVPR 2026

Achieving efficient and robust whole-body control (WBC) is essential for enabling humanoid robots to perform complex tasks in dynamic environments. Despite the success of reinforcement learning (RL) in this domain, its sample inefficiency remains a significant challenge due to the intricate dynamics

Cited by 0SourcecodeScholar
2026

Reasoning in Space via Grounding in the World

ICLR 2026poster

In this paper, we claim that 3D visual grounding is the cornerstone of spatial reasoning and introduce the $\textit{Grounded-Spatial Reasoner (GS-Reasoner)}$ to explore the effective spatial representations that bridge the gap between them. Existing 3D LLMs suffer from the absence of a unified 3D re…

Cited by 0SourcecodeScholar
2026

Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning

ICLR 2026poster

Recent breakthroughs in reasoning language models have significantly advanced text-based reasoning. On the other hand, Multi-modal Large Language Models (MLLMs) still lag behind, hindered by their outdated internal LLMs. Upgrading these is often prohibitively expensive, as it requires complete visio…

Cited by 0SourcecodeScholar
2026

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

CVPR 2026

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D understanding or video generation, a significant gap remains in

Cited by 0SourceScholar
2026

Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment

AAAI 2026technical

As super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical f

Cited by 0SourcePDFScholar
2026

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

CVPR 2026

The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics to encompass aesthetic appeal. However, existing benchmarks remain largely focused on technical fidelity, leaving a sig

Cited by 0SourceScholar
2026

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

IJCAI 2026

The rapid advancement of AIGC video generation calls for evaluation frameworks that move beyond technical fidelity and incorporate human-centered aesthetic assessment. Existing benchmarks often overlook fine-grained perceptual qualities such as visual aesthetics, artistic style, and human preference

Cited by 0Scholar
2026

Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation

AAAI 2026technical

Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal an

Cited by 0SourcePDFScholar
2026

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving

ICML 2026poster

Deploying multiple models within shared GPU clusters is a key strategy to improve resource efficiency in large language model (LLM) serving. Existing multi-LLM serving systems improve GPU utilization at the cost of degraded inference performance, particularly time-to-first-token (TTFT). We attribute…

Cited by 0SourceScholar
2026

When MLLMs Meets Compression Distortion: A Coding Paradigm Tailored to MLLMs

ICLR 2026poster

The increasing deployment of powerful Multimodal Large Language Models (MLLMs), typically hosted on cloud platforms, urgently requires effective compression techniques to efficiently transmit signal inputs (e.g., images, videos) from edge devices with minimal bandwidth usage. However, conventional i…

Cited by 0SourcecodeScholar
2025

Classic Video Denoising in a Machine Learning World: Robust, Fast, and Controllable

CVPR 2025poster

Denoising is a crucial step in many video processing pipelines such as in interactive editing, where high quality, speed, and user control are essential. While recent approaches achieve significant improvements in denoising quality by leveraging deep learning, they are prone to unexpected failures d…

Cited by 0SourcePDFScholar
2025

DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution

ICCV 2025poster

Large-scale pre-trained diffusion models are becoming increasingly popular in solving the Real-World Image Super-Resolution (Real-ISR) problem because of their rich generative priors. The recent development of diffusion transformer (DiT) has witnessed overwhelming performance over the traditional UN…

Cited by 0SourcePDFScholar
2025

Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior

NeurIPS 2025poster

Image compression methods are usually optimized isolatedly for human perception or machine analysis tasks. We reveal fundamental commonalities between these objectives: preserving accurate semantic information is paramount, as it directly dictates the integrity of critical information for intelligen…

Cited by 0SourceScholar
2025

DiffRetouch: Using Diffusion to Retouch on the Shoulder of Experts

AAAI 2025technical

Image retouching aims to enhance the visual quality of photos. Considering the different aesthetic preferences of users, the target of retouching is subjective. However, current retouching methods mostly adopt deterministic models, which not only neglects the style diversity in the expert-retouched…

Cited by 0SourcePDFScholar
2025

Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning

ICCV 2025poster

Training visual reinforcement learning (RL) in practical scenarios presents a significant challenge, i.e., RL agents suffer from low sample efficiency in environments with variations. While various approaches have attempted to alleviate this issue by disentangled representation learning, these metho…

Cited by 0SourcePDFScholar
2025

Dis²Booth: Learning Image Distribution with Disentangled Features for Text-to-Image Diffusion Models

AAAI 2025technical

Personalized image generation enables customized content creation based on the text-to-image diffusion models.However, existing personalization methods focus on fine-tuning generative models to learn to generate specific single individuals or concepts, such as an image of a specific Corgi, but are u…

Cited by 0SourcePDFScholar
2025

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

NeurIPS 2025poster

Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation. However, existing methods are limited to challenging image-based forecasting, which suffers from redundant i…

Cited by 0SourcecodeScholar
2025

EDBench: Large-Scale Electron Density Data for Molecular Modeling

NeurIPS 2025poster

Existing molecular machine learning force fields (MLFFs) generally focus on the learning of atoms, molecules, and simple quantum chemical properties (such as energy and force), but ignore the importance of electron density (ED) $\rho(r)$ in accurately understanding molecular force fields (MFFs). ED…

Cited by 0SourcecodeScholar
2025

Electron Density-enhanced Molecular Geometry Learning

IJCAI 2025

Electron density (ED), which describes the probability distribution of electrons in space, is crucial for accurately understanding the energy and force distribution in molecular force fields (MFF). Existing machine learning force fields (MLFF) focus on mining appropriate physical quantities from the

2025

Exploring Simple Siamese Network for High-Resolution Video Quality Assessment

ICASSP 2025accepted

In the research of video quality assessment (VQA), two-branch network [1] has emerged as a promising solution. It decouples VQA with separate technical and aesthetic branches to measure the perception of low-level distortions and high-level semantics respectively. However, we argue that while techni…

Cited by 0SourceScholar
2025

GCTAM: Global and Contextual Truncated Affinity Combined Maximization Model For Unsupervised Graph Anomaly Detection

IJCAI 2025

Anomalies often occur in real-world information networks/graphs, such as malevolent users, malicious comments, banned users, and fake news in social graphs. The latest graph anomaly detection methods use a novel mechanism called truncated affinity maximization (TAM) to detect anomaly nodes without u

2025

GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-based Transformer

ICCV 2025poster

Lidar-based 3D detection is one of the most popular research fields in autonomous driving. 3D detectors typically detect specific targets in a scene according to the pattern formed by the spatial distribution of point clouds. However, existing voxel-based methods usually adopt MLP and global pooling…

2025

Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation

ICCV 2025poster

Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this challenge, we propose Hybrid-depth, a novel framework that systematically integrates foundation models (e.g., CLIP and DINO…

Cited by 0SourcePDFScholar
2025

KLMN: Knowledge distillation based lightweight multi-clue image forgery detection and localization

ICASSP 2025accepted

Current image forensics methods often utilize image features from various frequency domains. However, the effective use of these features frequently depends on complex network architectures and a large number of parameters. In this paper, we introduce a lightweight Multi-Clue image forgery detection…

Cited by 0SourceScholar
2025

Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to…

2025

Multimodal Language Models See Better When They Look Shallower

EMNLP 2025

Multimodal large language models (MLLMs) typically extract visual features from the final layers of a pretrained Vision Transformer (ViT). This widespread deep-layer bias, however, is largely driven by empirical convention rather than principled analysis. While prior studies suggest that different V

2025

Open-World Reinforcement Learning over Long Short-Term Imagination

ICLR 2025oral

Training visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be “short-sighted”, as they are typically trained on short snip…

2025

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

ICCV 2025poster

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and ignore that AI assistants perceive and act based on first-per…

2025

RLLTE: Long-Term Evolution Project of Reinforcement Learning

AAAI 2025technical

We present RLLTE: a long-term evolution, extremely modular, and open-source framework for reinforcement learning (RL) research and application. Beyond delivering top-notch algorithm implementations, RLLTE also serves as a toolkit for developing algorithms. More specifically, RLLTE decouples the RL a…

2025

SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation

NeurIPS 2025spotlight

While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation—a key factor in 6-DoF fine-grained manipulation. Traditional pose representations rely on pre-defined frames or templates, limiting generalization and semantic grounding. In this pap…

Cited by 0SourceScholar
2025

Taming LLMs with Gradient Grouping

ACL 2025long

Training large language models (LLMs) poses challenges due to their massive scale and heterogeneous architectures. While adaptive optimizers like AdamW help address gradient variations, they still struggle with efficient and effective parameter-wise learning rate estimation, resulting in training in…

2025

Towards RAW Object Detection in Diverse Conditions

CVPR 2025highlight

Existing object detection methods often consider sRGB input, which was compressed from RAW data using ISP originally designed for visualization. However, such compression might lose crucial information for detection, especially under complex light and weather conditions. We introduce the AODRaw data…

2025

ULTHO: Ultra-Lightweight yet Efficient Hyperparameter Optimization in Deep Reinforcement Learning

ICCV 2025poster

Hyperparameter optimization (HPO) is a billion-dollar problem in machine learning, which significantly impacts the training efficiency and model performance. However, achieving efficient and robust HPO in deep reinforcement learning (RL) is consistently challenging due to its high non-stationarity a…

2025

UltraLED: Learning to See Everything in Ultra-High Dynamic Range Scenes

NeurIPS 2025poster

Ultra-high dynamic range (UHDR) scenes exhibit pronounced exposure disparities between bright and dark regions. Such conditions are Ultra-high dynamic range (UHDR) scenes exhibit significant exposure disparities between bright and dark regions. Such conditions are commonly encountered in nighttime s…

Cited by 0SourcecodeScholar
2025

UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object Detection

CVPR 2025poster

Recent advances in LiDAR 3D detection have demonstrated the effectiveness of Transformer-based frameworks in capturing the global dependencies from point cloud spaces, which serialize the 3D voxels into the flattened 1D sequence for iterative self-attention. However, the spatial structure of 3D voxe…

Cited by 0SourcePDFScholar
2025

UniScene: Unified Occupancy-centric Driving Scene Generation

CVPR 2025poster

Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles…

2025

Unified Arbitrary-Time Video Frame Interpolation and Prediction

ICASSP 2025accepted

Video frame interpolation and prediction aim to synthesize frames in-between and subsequent to existing frames, respectively. Despite being closely-related, these two tasks are traditionally studied with different model architectures, or same architecture but individually trained weights. Furthermor…

Cited by 0SourceScholar
2025

VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations

NeurIPS 2025poster

Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and multimodal fusion challenges hinder direct application of image-bas…

Cited by 0SourcecodeScholar
2024

APDDv2: Aesthetics of Paintings and Drawings Dataset with Artist Labeled Scores and Comments

NeurIPS 2024poster

Datasets play a pivotal role in training visual models, facilitating the development of abstract understandings of visual features through diverse image samples and multidimensional attributes. However, in the realm of aesthetic evaluation of artistic images, datasets remain relatively scarce. Exist…

2024

Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion

IJCAI 2024poster

3D semantic scene completion (SSC) is an ill-posed perception task that requires inferring a dense 3D scene from limited observations. Previous camera-based methods struggle to predict accurate semantic scenes due to inherent geometric ambiguity and incomplete observations. In this paper, we resort…

2024

Dynamic Video Frame Interpolation with Integrated Difficulty Pre-Assessment

ICASSP 2024accepted

Video frame interpolation (VFI) has witnessed great progress in recent years. However, existing VFI models still struggle to achieve a good trade-off between accuracy and efficiency. Accurate VFI models typically rely on heavy compute to process all samples, ignoring the fact that easy samples with…

Cited by 0SourceScholar
2024

Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models

NeurIPS 2024poster

Disentangled representation learning (DRL) aims to identify and decompose underlying factors behind observations, thus facilitating data perception and generation. However, current DRL approaches often rely on the unrealistic assumption that semantic factors are statistically independent. In reality…

Cited by 2SourcePDFScholar
2024

HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects

ECCV 2024poster

"Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of multiple objects. Thus, we propose HIMO, a large-scale MoCap da…

Cited by 3SourcePDFScholar
2024

Hierarchical Temporal Context Learning for Camera-based Semantic Scene Completion

ECCV 2024poster

"Camera-based 3D semantic scene completion (SSC) is pivotal for predicting complicated 3D layouts with limited 2D image observations. The existing mainstream solutions generally leverage temporal information by roughly stacking history frames to supplement the current frame, such straightforward tem…

2024

In Pursuit of Causal Label Correlations for Multi-label Image Recognition

NeurIPS 2024poster

Multi-label image recognition aims to predict all objects present in an input image. A common belief is that modeling the correlations between objects is beneficial for multi-label recognition. However, this belief has been recently challenged as label correlations may mislead the classifier in test…

Cited by 0SourcePDFScholar
2024

Inter-X: Towards Versatile Human-Human Interaction Analysis

CVPR 2024poster

The analysis of the ubiquitous human-human interactions is pivotal for understanding humans as social beings. Existing human-human interaction datasets typically suffer from inaccurate body motions lack of hand gestures and fine-grained textual descriptions. To better perceive and generate human-hum…

2024

Language-Image Pre-training with Long Captions

ECCV 2024poster

"Language-image pre-training largely relies on how precisely and thoroughly a text describes its paired image. In practice, however, the contents of an image can be so rich that well describing them requires lengthy captions (e.g., with 10 sentences), which are usually missing in existing datasets.…

2024

Lighting Every Darkness with 3DGS: Fast Training and Real-Time Rendering for HDR View Synthesis

NeurIPS 2024poster

Volumetric rendering-based methods, like NeRF, excel in HDR view synthesis from RAW images, especially for nighttime scenes. They suffer from long training times and cannot perform real-time rendering due to dense sampling requirements. The advent of 3D Gaussian Splatting (3DGS) enables real-time re…

2024

Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning

NeurIPS 2024poster

Training offline RL models using visual inputs poses two significant challenges, *i.e.*, the overfitting problem in representation learning and the overestimation bias for expected future rewards. Recent work has attempted to alleviate the overestimation bias by encouraging conservative behaviors. T…

2024

Multi-Dimensional Geometric Feature-Based Calibration Method for LiDAR and Camera Fusion

ICASSP 2024accepted

Extrinsic calibration between LiDAR and camera has become an indispensable task across diverse domains, including autonomous vehicles, robotics, and surveillance systems. However, existing methods suffer from limited precision due to the inaccurate and insufficient detected features caused by the sp…

Cited by 0SourceScholar
2024

Multi-Prompts Learning with Cross-Modal Alignment for Attribute-Based Person Re-identification

AAAI 2024technical

The fine-grained attribute descriptions can significantly supplement the valuable semantic information for person image, which is vital to the success of person re-identification (ReID) task. However, current ReID algorithms typically failed to effectively leverage the rich contextual information av…

2024

One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene Perception

AAAI 2024technical

Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes directly with geometric correspondence, attempting to fully address…

Cited by 3SourcePDFScholar
2024

Paintings and Drawings Aesthetics Assessment with Rich Attributes for Various Artistic Categories

IJCAI 2024poster

Image aesthetic evaluation is a highly prominent research domain in the field of computer vision. In recent years, there has been a proliferation of datasets and corresponding evaluation methodologies for assessing the aesthetic quality of photographic works, leading to the establishment of a relati…

2024

Rate-Distortion-Cognition Controllable Versatile Neural Image Compression

ECCV 2024poster

"Recently, the field of Image Coding for Machines (ICM) has garnered heightened interest and significant advances thanks to the rapid progress of learning-based techniques for image compression and analysis. Previous studies often require training separate codecs to support various bitrate levels, m…

Cited by 4SourcePDFScholar
2024

ReGenNet: Towards Human Action-Reaction Synthesis

CVPR 2024poster

Humans constantly interact with their surrounding environments. Current human-centric generative models mainly focus on synthesizing humans plausibly interacting with static scenes and objects while the dynamic human action-reaction synthesis for ubiquitous causal human-human interactions is less ex…

2024

Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation

NeurIPS 2024spotlight

There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a…

Cited by 2SourcePDFScholar
2024

StyDeSty: Min-Max Stylization and Destylization for Single Domain Generalization

ICML 2024poster

Single domain generalization (single DG) aims at learning a robust model generalizable to unseen domains from only one training domain, making it a highly ambitious and challenging task. State-of-the-art approaches have mostly relied on data augmentations, such as adversarial perturbation and style…

2024

SwiftPillars: High-Efficiency Pillar Encoder for Lidar-Based 3D Detection

AAAI 2024technical

Lidar-based 3D Detection is one of the significant components of Autonomous Driving. However, current methods over-focus on improving the performance of 3D Lidar perception, which causes the architecture of networks becoming complicated and hard to deploy. Thus, the methods are difficult to apply in…

Cited by 4SourcePDFScholar
2023

A Unified Pyramid Recurrent Network for Video Frame Interpolation

CVPR 2023poster

Flow-guided synthesis provides a common framework for frame interpolation, where optical flow is estimated to guide the synthesis of intermediate frames between consecutive inputs. In this paper, we present UPR-Net, a novel Unified Pyramid Recurrent Network for frame interpolation. Cast in a flexibl…

2023

Accurate Implicit Neural Mapping With More Compact Representation in Large-Scale Scenes Using Ranging Data

RA-L 2023

Large-scale 3D mapping nowadays is a research hotspot in robotics. A greatly concerning issue is reconstructing high-accuracy maps in a hardware environment with limited memory. To address this problem, we propose a novel implicit neural mapping approach with higher accuracy and less memory. It firs

Cited by 13SourceScholar
2023

ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation

ICCV 2023poster

We present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion TransFormer (ActFormer) under a GAN training scheme, equippe…

Cited by 74PDFScholar
2023

Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement Learning

ICML 2023poster

We present AIRS: **A**utomatic **I**ntrinsic **R**eward **S**haping that intelligently and adaptively provides high-quality intrinsic rewards to enhance exploration in reinforcement learning (RL). More specifically, AIRS selects shaping function from a predefined set based on the estimated task retu…

2023

DNF: Decouple and Feedback Network for Seeing in the Dark

CVPR 2023highlight

The exclusive properties of RAW data have shown great potential for low-light image enhancement. Nevertheless, the performance is bottlenecked by the inherent limitations of existing architectures in both single-stage and multi-stage methods. Mixed mapping across two different domains, noise-to-clea…

2023

Discrete Point-Wise Attack Is Not Enough: Generalized Manifold Adversarial Attack for Face Recognition

CVPR 2023poster

Classical adversarial attacks for Face Recognition (FR) models typically generate discrete examples for target identity with a single state image. However, such paradigm of point-wise attack exhibits poor generalization against numerous unknown states of identity and can be easily defended. In this…

2023

Generalized Lightness Adaptation with Channel Selective Normalization

ICCV 2023poster

Lightness adaptation is vital to the success of image processing to avoid unexpected visual deterioration, which covers multiple aspects, e.g., low-light image enhancement, image retouching, and inverse tone mapping. Existing methods typically work well on their trained lightness conditions but perf…

Cited by 20PDFcodeScholar
2023

Learning Distortion Invariant Representation for Image Restoration From a Causality Perspective

CVPR 2023poster

In recent years, we have witnessed the great advancement of Deep neural networks (DNNs) in image restoration. However, a critical limitation is that they cannot generalize well to real-world degradations with different degrees or types. In this paper, we are the first to propose a novel training str…

2023

Lighting Every Darkness in Two Pairs: A Calibration-Free Pipeline for RAW Denoising

ICCV 2023poster

Calibration-based methods have dominated RAW image denoising under extremely low-light environments. However, these methods suffer from several main deficiencies: 1) the calibration procedure is laborious and time-consuming, 2) denoisers for different cameras are difficult to transfer, and 3) the di…

Cited by 23PDFScholar
2023

Multi-Layer Seasonal Perception Network for Time Series Forecasting

ICASSP 2023accepted

Seasonal time series contain rich long-term dependencies. How to make good use of the seasonal information to predict the future is still a challenging problem. In this paper, we propose a neural network model called Multilayer Seasonal Perception Network (MSPNet) to predict seasonal time series. Fi…

Cited by 0SourceScholar
2023

NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigation

ICCV 2023poster

3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. This task is currently under-explored and poses great challenges: (i) the 3D representations are complex and in general c…

Cited by 11PDFcodeScholar
2023

PLPL-VIO: A Novel Probabilistic Line Measurement Model for Point-Line-Based Visual-Inertial Odometry

IROS 2023poster

Point and line features are complementary in Visual-Inertial Odometry (VIO) or Visual-Inertial Simultaneous Localization And Mapping (VI-SLAM) systems. The advantage of combining these two types of features relies on their proper weighting in the cost function, usually set by their uncertainty. Comp…

Cited by 5SourceScholar
2023

Semantically Structured Image Compression via Irregular Group-Based Decoupling

ICCV 2023poster

Image compression techniques typically focus on compressing rectangular images for human consumption, however, resulting in transmitting redundant content for downstream applications. To overcome this limitation, some previous works propose to semantically structure the bitstream, which can meet spe…

Cited by 13PDFScholar
2023

Underwater Ranker: Learn Which Is Better and How to Be Better

AAAI 2023technical

In this paper, we present a ranking-based underwater image quality assessment (UIQA) method, abbreviated as URanker. The URanker is built on the efficient conv-attentional image Transformer. In terms of underwater images, we specially devise (1) the histogram prior that embeds the color distribution…

2022

Attribute Group Editing for Reliable Few-Shot Image Generation

CVPR 2022poster

Few-shot image generation is a challenging task even using the state-of-the-art Generative Adversarial Networks (GANs). Due to the unstable GAN training process and the limited training data, the generated images are often of low quality and low diversity. In this work, we propose a new "editing-bas…

Cited by 36PDFcodeScholar
2022

Cloth-Changing Person Re-Identification From a Single Image With Gait Prediction and Regularization

CVPR 2022poster

Cloth-Changing person re-identification (CC-ReID) aims at matching the same person across different locations over a long-duration, e.g., over days, and therefore inevitably has cases of changing clothing. In this paper, we focus on handling well the CC-ReID problem under a more challenging setting,…

Cited by 179PDFcodeScholar
2022

Deliberated Domain Bridging for Domain Adaptive Semantic Segmentation

NeurIPS 2022accept

In unsupervised domain adaptation (UDA), directly adapting from the source to the target domain usually suffers significant discrepancies and leads to insufficient alignment. Thus, many UDA works attempt to vanish the domain gap gradually and softly via various intermediate spaces, dubbed domain bri…

2022

Image Coding for Machines with Omnipotent Feature Learning

ECCV 2022poster

"Image Coding for Machines (ICM) aims to compress images for AI tasks analysis rather than meeting human perception. Learning a kind of feature that is both general (for AI tasks) and compact (for compression) is pivotal for its success. In this paper, we attempt to develop an ICM framework by learn…

2022

Reusing the Task-Specific Classifier as a Discriminator: Discriminator-Free Adversarial Domain Adaptation

CVPR 2022poster

Adversarial learning has achieved remarkable performances for unsupervised domain adaptation (UDA). Existing adversarial UDA methods typically adopt an additional discriminator to play the min-max game with a feature extractor. However, most of these methods failed to effectively leverage the predic…

Cited by 201PDFcodeScholar
2022

SADN: Learned Light Field Image Compression with Spatial-Angular Decorrelation

ICASSP 2022accepted

Light field image becomes one of the most promising media types for immersive video applications. In this paper, we propose a novel end-to-end spatial-angular-decorrelated network (SADN) for high-efficiency light field image compression. Different from the existing methods that exploit either spatia…

Cited by 0SourceScholar
2022

Unleashing Potential of Unsupervised Pre-Training With Intra-Identity Regularization for Person Re-Identification

CVPR 2022poster

Existing person re-identification (ReID) methods typically directly load the pre-trained ImageNet weights for initialization. However, as a fine-grained classification task, ReID is more challenging and exists a large domain gap between ImageNet classification. Inspired by the great success of self-…

Cited by 47PDFScholar
2022

Unsupervised Coherent Video Cartoonization with Perceptual Motion Consistency

AAAI 2022technical

In recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry. Different from image translations focusing on improving the st…

2021

Dense Interaction Learning for Video-Based Person Re-Identification

ICCV 2021poster

Video-based person re-identification (re-ID) aims at matching the same person across video clips. Efficiently exploiting multi-scale fine-grained features while building the structural interaction among them is pivotal for its success. In this paper, we propose a hybrid framework, Dense Interaction…

Cited by 67PDFcodeScholar
2021

Learning Omni-Frequency Region-adaptive Representations for Real Image Super-Resolution

AAAI 2021technical

Traditional single image super-resolution (SISR) methods that focus on solving single and uniform degradation (i.e., bicubic down-sampling), typically suffer from poor performance when applied into real-world low-resolution (LR) images due to the complicated realistic degradations. The key to solvin…

Cited by 44SourcePDFScholar
2021

Re-Energizing Domain Discriminator With Sample Relabeling for Adversarial Domain Adaptation

ICCV 2021poster

Many unsupervised domain adaptation (UDA) methods exploit domain adversarial training to align the features to reduce domain gap, where a feature extractor is trained to fool a domain discriminator in order to have aligned feature distributions. The discrimination capability of the domain classifier…

Cited by 18PDFScholar
2020

Exploring Categorical Regularization for Domain Adaptive Object Detection

CVPR 2020poster

In this paper, we tackle the domain adaptive object detection problem, where the main challenge lies in significant domain gaps between source and target domains. Previous work seeks to plainly align image-level and instance-level shifts to eventually minimize the domain discrepancy. However, they s…

Cited by 378PDFcodeScholar
2020

Global Distance-distributions Separation for Unsupervised Person Re-identification

ECCV 2020poster

Supervised person re-identification (ReID) often has poor scalability and usability in real-world deployments due to domain gaps and the lack of annotations for the target domain data. Unsupervised person ReID through domain adaptation is attractive yet challenging. Existing unsupervised ReID approa…

Cited by 90SourcePDFScholar
2020

Hierarchical Context Embedding for Region-based Object Detection

ECCV 2020poster

State-of-the-art two-stage object detectors apply a classifier to a sparse set of object proposals, relying on region-wise features extracted by RoIPool or RoIAlign as inputs. The region-wise features, in spite of aligning well with the proposal locations, may still lack the crucial context informat…

Cited by 35SourcePDFScholar
2020

Learning Disentangled Feature Representation for Hybrid-distorted Image Restoration

ECCV 2020poster

Hybrid-distorted image restoration (HD-IR) is dedicated to restore real distorted image that is degraded by multiple distortions. Existing HD-IR approaches usually ignore the inherent interference among hybrid distortions which compromises the restoration performance. To decompose such interference,…

Cited by 55SourcePDFScholar
2020

Relation-Aware Global Attention for Person Re-Identification

CVPR 2020poster

For person re-identification (re-id), attention mechanisms have become attractive as they aim at strengthening discriminative features and suppressing irrelevant ones, which matches well the key of re-id, i.e., discriminative feature learning. Previous approaches typically learn attention using loca…

Cited by 713PDFcodeScholar
2020

Style Normalization and Restitution for Generalizable Person Re-Identification

CVPR 2020poster

Existing fully-supervised person re-identification (ReID) methods usually suffer from poor generalization capability caused by domain gaps. The key to solving this problem lies in filtering out identity-irrelevant interference and learning domain-invariant person representations. In this paper, we a…

Cited by 448PDFcodeScholar
2019

Light Field Image Compression Using Depth-based CNN in Intra Prediction

ICASSP 2019accepted

Recently, light field images have received extensive attention due to their potential applications. Since they take up a huge memory because of its super-high resolution, efficient compression methods are fundamentally required. In this paper, we propose a novel intra prediction mode by using depth-…

Cited by 0SourceScholar
2018

High-Speed Light Field Image Formation Analysis Using Wavefield Modeling with Flexible Sampling

ICASSP 2018accepted

Understanding the image formation inside plenoptic cameras is significant for the investigations of improving the low spatial resolution. Most researches explore the image formation from the perspective of geometric optics. However, as the hardware components in combination with low-aperture optical…

Cited by 0SourceScholar
2015

Design of a cable-driven active leg exoskeleton (C-ALEX) and gait training experiments with human subjects

ICRA 2015poster

Robotic rehabilitation devices are attractive to physical therapists. Various leg exoskeletons have been developed during the past decade and have been used in gait training. Traditional exoskeletons usually have a complex structure and add extra inertia to the wearer's leg, which may change their n…

Cited by 136SourceScholar