← Search

Yu LIU

264 accepted papers

2026

AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization

ICLR 2026poster

Large Vision-Language Models (LVLMs), such as GPT-4o and LLaVA, have recently witnessed remarkable advancements and are increasingly being deployed in real-world applications. However, inheriting the sensitivity of visual neural networks, LVLMs remain vulnerable to adversarial attacks, which can re…

Cited by 0SourceScholar
2026

Adaptive Dynamic Dehazing via Instruction-Driven and Task-Feedback Closed-Loop Optimization for Diverse Downstream Task Adaptation

AAAI 2026technical

In real-world vision systems, haze removal is required not only to enhance image visibility but also to meet the specific needs of diverse downstream tasks. To address this challenge, we propose a novel adaptive dynamic dehazing framework that incorporates a closed-loop optimization mechanism. It en

Cited by 0SourcePDFScholar
2026

Adaptor: Advancing Assistive Teleoperation with Few-Shot Learning and Cross-Operator Generalization

ICRA 2026poster

Assistive teleoperation enhances efficiency via shared control, yet inter-operator variability, stemming from diverse habits and expertise, induces highly heterogeneous trajectory distributions that undermine intent recognition stability. We present Adaptor, a few-shot framework for robust cross-ope…

2026

BioX-Bridge: Model Bridging for Unsupervised Cross-Modal Knowledge Transfer across Biosignals

ICLR 2026oral

Biosignals offer valuable insights into the physiological states of the human body. Although biosignal modalities differ in functionality, signal fidelity, sensor comfort, and cost, they are often intercorrelated, reflecting the holistic and interconnected nature of human physiology. This opens up t…

Cited by 0SourcecodeScholar
2026

CLoD-GS: Continuous Level-of-Detail via 3D Gaussian Splatting

ICLR 2026poster

Level of Detail (LoD) is a fundamental technique in real-time computer graphics for managing the rendering costs of complex scenes while preserving visual fidelity. Traditionally, LoD is implemented using discrete levels (DLoD), where multiple, distinct versions of a model are swapped out at differe…

Cited by 0SourcecodeScholar
2026

Causality-Aligned Semantic Recovery for Incomplete Cross-Modal Retrieval

AAAI 2026technical

Incomplete cross-modal retrieval (ICMR) requires models to recover missing modalities and robustly align heterogeneous ones for effective retrieval. Existing methods, however, fall short in both aspects. They often rely on limited semantic cues, such as single samples or coarse category prototypes,

Cited by 0SourcePDFScholar
2026

CircuitPrint: Mechanistic Circuit Fingerprints for Large Language Models

ICML 2026poster

Large language models (LLMs) are trained at significant computational and data cost, making them valuable intellectual property (IP). Existing IP verification methods primarily rely either on invasive watermarking that degrades model utility, or on superficial behavioral signatures disrupted by fine…

Cited by 0SourceScholar
2026

Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models

AAAI 2026technical

Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-lab

Cited by 0SourcePDFScholar
2026

Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image Fusion

CVPR 2026

Infrared-visible image fusion aims to integrate complementary information for robust visual understanding, but existing fusion methods struggle with simultaneously adapting to multiple downstream tasks. To address this issue, we propose a Closed-Loop Dynamic Network (CLDyN) that can adaptively respo

Cited by 0SourcecodeScholar
2026

Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios

CVPR 2026

Complex degradations like noise, blur, and low resolution are typical challenges in real-world image fusion tasks, limiting the performance and practicality of existing methods. End-to-end neural network-based approaches are generally simple to design and highly efficient in inference, but their bla

Cited by 0SourcecodeScholar
2026

DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment

ICLR 2026poster

Recent GRPO-based approaches built on flow matching models have shown remarkable improvements in human preference alignment for text-to-image generation. Nevertheless, they still suffer from the sparse reward problem: the terminal reward of the entire denoising trajectory is applied to all intermedi…

Cited by 0SourceScholar
2026

Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning

AAAI 2026technical

Emotion Recognition in Conversation (ERC) is a crucial task for understanding human emotions and enabling natural human-computer interaction. Although Large Language Models (LLMs) have recently shown great potential in this field, their ability to capture the intrinsic connections between explicit a

Cited by 0SourcePDFScholar
2026

Dual-Level Hypergraph Generation for Addressing Feature Scarcity in Whole-Slide Image Classification

CVPR 2026

Lymph node metastasis diagnosis in pathological images is a highly challenging four-class classification task, comprising macrometastasis, micrometastasis, isolated tumor cells (ITC), and negative lesions.Unlike conventional classification settings, this four-class scenario simultaneously suffers fr

Cited by 0SourcecodeScholar
2026

EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision

AAAI 2026technical

Egocentric visual query localization is vital for embodied AI and VR/AR, yet remains challenging due to camera motion, viewpoint changes, and appearance variations. We present EAGLE, a novel framework that leverages episodic appearance- and geometry-aware memory to achieve unified 2D-3D visual query

Cited by 0SourcePDFScholar
2026

FDP: A Frequency-Decomposition Preprocessing Pipeline for Unsupervised Anomaly Detection in Brain MRI

AAAI 2026technical

Due to the diversity of brain anatomy and the scarcity of annotated data, supervised anomaly detection for brain MRI remains challenging, driving the development of unsupervised anomaly detection (UAD) approaches. Current UAD methods typically utilize synthetically generated noise perturbations on

Cited by 0SourcePDFScholar
2026

From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection

CVPR 2026

Zero-shot anomaly detection (ZSAD) aims to identify unseen anomalies without abnormal supervision, which is essential for open-world scenarios. Recent vision-language models such as CLIP enable anomaly reasoning through shared visual-textual embeddings, but existing methods often rely on coarse prom

Cited by 0SourceScholar
2026

From Observations to Events: Event-Aware World Models for Reinforcement Learning

ICLR 2026poster

While model-based reinforcement learning (MBRL) improves sample efficiency by learning world models from raw observations, existing methods struggle to generalize across structurally similar scenes and remain vulnerable to spurious variations such as textures or color shifts. From a cognitive scienc…

Cited by 0SourcecodeScholar
2026

From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction

ICML 2026poster

By processing electronic health records (EHRs) as natural language sequences, large language models (LLMs) have shown potential in clinical prediction tasks such as mortality prediction and phenotyping. However, longitudinal or highly frequent EHRs often yield excessively long token sequences that r…

Cited by 0SourceScholar
2026

G2C-MT: Graph-Guided Context Selection for Document-Level Machine Translation

IJCAI 2026

Effective document-level machine translation (DocMT) requires capturing long-range discourse dependencies. Recent work has explored retrieval-based and discourse-aware context selection. However, these approaches often lack an explicit mechanism for modeling structured discourse dependencies between

Cited by 0Scholar
2026

G4Splat: Geometry-Guided Gaussian Splatting with Generative Prior

ICLR 2026poster

Despite recent advances in leveraging generative prior from pre-trained diffusion models for 3D scene reconstruction, existing methods still face two critical limitations. First, due to the lack of reliable geometric supervision, they struggle to produce high-quality reconstructions even in observed…

Cited by 0SourcecodeScholar
2026

GVF-MPC Depth Control Framework Motivated by Restoring-Moment Mechanism Analysis for Bionic Robotic Fish

RA-L 2026

Longitudinal oscillations are commonly observed in depth control of low-speed underwater platforms such as bio-inspired robotic fish. To investigate this issue, this letter conducts a quantitative analysis of the vertical-plane dynamics. It reveals the speed-dependent coupling between the depth and

Cited by 0SourceScholar
2026

Gracefully Air-Written: Enhancing the Legibility and Style Consistency of In-Air Handwriting

AAAI 2026technical

Space computing devices expand handwritten input from two-dimensional screens into three-dimensional space, providing an unrestricted interactive experience. Due to the high degree of freedom and lack of tactile feedback in in-air handwriting, handwritten characters not only become less legible but

Cited by 0SourcePDFScholar
2026

High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning

CVPR 2026

Face swapping aims to seamlessly transfer a source facial identity onto a target while preserving target attributes such as pose and expression. Diffusion models, known for their superior generative capabilities, have recently shown promise in advancing face-swapping quality. This paper addresses tw

Cited by 0SourceScholar
2026

IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments

AAAI 2026technical

Vision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robots or outdoor Unmanned Aerial Vehicles (UAVs), indoor UAV-based VLN remains unde

Cited by 0SourcePDFScholar
2026

Intention-Aware Diffusion Model for Pedestrian Trajectory Prediction

AAAI 2026technical

Predicting pedestrian motion trajectories is critical for the path planning and motion control of autonomous vehicles. Recent diffusion-based models have shown promising results in capturing the inherent stochasticity of pedestrian behavior for trajectory prediction. However, the absence of explicit

Cited by 0SourcePDFScholar
2026

Learning 3D Occupancy from Beam Overlap in 2D Rotating mmWave Radar

AAAI 2026technical

Robust 3D perception under adverse weather is critical for autonomous systems. While mmWave Radars are inherently weather-resistant, conventional 2D rotating Radar sensors lack direct elevation resolution, limiting their 3D perception ability. Although 4D imaging radars can provide elevation informa

Cited by 0SourcePDFScholar
2026

Learning to Learn Weight Generation via Local Consistency Diffusion

CVPR 2026

Diffusion-based algorithms have emerged as promising techniques for weight generation. However, existing solutions are limited by two challenges: generalizability and missing local supervision targets. The first challenge stems from the inherent lack of cross-task transferability in existing single-

Cited by 0SourceScholar
2026

LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette Alignment

CVPR 2026

We present LoD-Loc v3, a novel method for generalized aerial visual localization in dense urban environments. While prior work LoD-Loc v2 achieves localization through semantic building silhouette alignment with low-detail city models, it suffers from two key limitations: poor cross-scene generaliza

Cited by 0SourcecodeScholar
2026

Local Precise Refinement: A Dual-Gated Mixture-of-Experts for Enhancing Foundation Model Generalization against Spectral Shifts

CVPR 2026

Domain Generalization Semantic Segmentation (DGSS) in spectral remote sensing is severely challenged by spectral shifts across diverse acquisition conditions, which cause significant performance degradation for models deployed in unseen domains. While fine-tuning foundation models is a promising dir

Cited by 0SourceScholar
2026

MPMA: Preference Manipulation Attack Against Model Context Protocol

AAAI 2026technical

Model Context Protocol (MCP) standardizes interface mapping for large language models (LLMs) to access external data and tools, which revolutionizes the paradigm of tool selection and facilitates the rapid expansion of the LLM agent tool ecosystem. However, as the MCP is increasingly adopted, third-

Cited by 0SourcePDFScholar
2026

Missing No More: Dictionary-Guided Cross-Modal Image Fusion under Missing Infrared

CVPR 2026

Infrared-visible (IR-VIS) image fusion is vital for perception and security, yet most methods rely on the availability of both modalities during training and inference. When the infrared modality is absent, pixel-space generative substitutes become hard to control and inherently lack interpretabilit

Cited by 0SourcecodeScholar
2026

Mitigating Gradient Pathology in PINNs through Aligned Constraint

ICML 2026poster

While Physics-Informed Neural Networks (PINNs) are powerful for solving Partial Differential Equations (PDEs), their training is often paralyzed by gradient pathology. The gradients from PDE residuals and boundary constraints oppose each other, trapping the model in local minima. Current solutions, …

Cited by 0SourceScholar
2026

Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models

CVPR 2026

Group Relative Policy Optimization (GRPO) has shown promise in aligning image and video generative models with human preferences. However, applying it to modern flow matching models is challenging because of its deterministic sampling paradigm. Current methods address this issue by converting Ordina

Cited by 0SourceScholar
2026

OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval

AAAI 2026technical

Recent advances in large language models (LLMs) and dense retrievers have driven significant progress in retrieval-augmented generation (RAG). However, existing approaches face significant challenges in complex reasoning-oriented multi-hop retrieval tasks: 1) Ineffective reasoning-oriented planning:

Cited by 0SourcePDFScholar
2026

Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection

ICML 2026poster

Open-Vocabulary Aerial Detection (OVAD) and Remote Sensing Visual Grounding (RSVG) have emerged as two key paradigms for aerial scene understanding. However, each paradigm suffers from inherent limitations when operating in isolation: OVAD is restricted to coarse category-level semantics, while RSVG…

Cited by 0SourceScholar
2026

PathMind: A Retrieve-Prioritize-Reason Framework for Knowledge Graph Reasoning with Large Language Models

AAAI 2026technical

Knowledge graph reasoning (KGR) is the task of inferring new knowledge by performing logical deductions on knowledge graphs. Recently, large language models (LLMs) have demonstrated remarkable performance in complex reasoning tasks. Despite promising success, current LLM-based KGR methods still fac

Cited by 0SourcePDFScholar
2026

PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localization

CVPR 2026

We present PiLoT, a unified framework that tackles UAV-based ego and target geo-localization. Conventional approaches rely on decoupled pipelines that fuse GNSS and Visual-Inertial Odometry (VIO) for ego-pose estimation, and active sensors like laser rangefinders for target localization. However, th

Cited by 0SourcecodeScholar
2026

PrivCode++ : Latent-Conditioned Differentially Private Code Generation for Comprehensive Guarantees

ICML 2026poster

Large language models fine-tuned on instruction–code pairs may memorize and subsequently leak sensitive training data. Existing differentially private (DP) code generation methods primarily protect code snippets while assuming prompts are public, which fails in realistic scenarios where prompts may …

Cited by 0SourceScholar
2026

RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels

AAAI 2026technical

Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their i

Cited by 0SourcePDFScholar
2026

Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance

ICLR 2026poster

Mixture-of-Experts (MoE) has emerged as a powerful paradigm for scaling model capacity while preserving computational efficiency. Despite its notable success in large language models (LLMs), existing attempts to apply MoE to Diffusion Transformers (DiTs) have yielded limited gains. We attribute this…

Cited by 0SourcecodeScholar
2026

Selective Actuation for Microrobots Based on Distributed Magnetic Field Design

ICRA 2026poster

Mechanical stimulation is essential for regulating cellular processes such as proliferation, differentiation, and apoptosis. Magnetic microrobot swarms offer a promising platform for delivering targeted mechanical stimulation to cells via remote actuation under rotating magnetic fields. However, mag…

Cited by 0Scholar
2026

Time Series Class-Incremental Learning via Confidence-guided Mask Distillation and Prototype-guided Contrastive Learning

AAAI 2026technical

Class-incremental learning (CIL) has recently gained great attention in the field of time series classification. Existing CIL methods based on knowledge distillation exhibit impressive ability to retain prior knowledge and overcome catastrophic forgetting, however, their effectiveness faces major c

Cited by 0SourcePDFScholar
2026

Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model

ICLR 2026poster

Autoregressive image generation aims to predict the next token based on previous ones. However, this process is challenged by the bidirectional dependencies inherent in conventional image tokenizations, which creates a fundamental misalignment with the unidirectional nature of autoregressive models.…

Cited by 0SourcecodeScholar
2026

UniCA: Unified Covariate Adaptation for Time Series Foundation Model

ICLR 2026poster

Time Series Foundation Models (TSFMs) have achieved remarkable success through large-scale pretraining. However, their design primarily targets real-valued series, limiting their ability to handle general forecasting tasks involving diverse and often \emph{heterogeneous covariates}—such as categoric…

Cited by 0SourcecodeScholar
2026

Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training

CVPR 2026

Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only single-modality outputs. This challenge of producing interleaved content is mainly due to training data scarcity and the dif

Cited by 0SourceScholar
2025

A Novel Split Deep Unfolding Transformer for Pan-Sharpening

ICASSP 2025accepted

Pan-sharpening is a commonly employed strategy to obtain high-resolution multispectral (HRMS) images. Existing deep unfolding networks for pan-sharpening suffer from ineffectively establishing the relationship between panchromatic (PAN) images and generated noisy HRMS (GN-HRMS) images in PAN-guided…

Cited by 0SourceScholar
2025

ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer

ICLR 2025poster

Diffusion models have emerged as a powerful generative technology and have been found to be applicable in various scenarios. Most existing foundational diffusion models are primarily designed for text-guided visual generation and do not support multi-modal conditions, which are essential for many vi…

Cited by 10SourcePDFScholar
2025

ArenaSim: A High-Performance Simulation Platform for Multi-Robot Self-Play Learning

RA-L 2025

In this letter, we introduce <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">ArenaSim</i>, a novel simulation platform designed for realistic and efficient self-play learning in multi-robot cooperative-competitive games. Compared to previous simulati

Cited by 0SourceScholar
2025

As Pseudo-Label Free as Possible: Leveraging Adaptive Feature Generation for Sparsely Annotated Object Detection

AAAI 2025technical

Compared to fully supervised object detection, training with sparse annotations typically leads to a decline in performance due to insufficient feature diversity. Existing sparsely annotated object detection (SAOD) methods often rely on pseudo-labeling strategies, but these pseudo-labels tend to int…

2025

Aspect-Based Sentiment Analysis with Syntax-Opinion-Sentiment Reasoning Chain

COLING 2025main

Despite the impressive capabilities of large language models (LLMs) in aspect-based sentiment analysis (ABSA), the role of syntactic information remains underexplored in LLMs. Syntactic structures are known to be crucial for capturing aspect-opinion relationships. To explore whether LLMs can effecti…

2025

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

CVPR 2025poster

Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions often carry lengthy, intertwined contexts that are difficult to parse and frequently overlook essential cues, posing a…

Cited by 0SourcePDFScholar
2025

Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting

ICLR 2025poster

Building interactable replicas of articulated objects is a key challenge in computer vision. Existing methods often fail to effectively integrate information across different object states, limiting the accuracy of part-mesh reconstruction and part dynamics modeling, particularly for complex multi-p…

Cited by 0SourcePDFScholar
2025

Decompositional Neural Scene Reconstruction with Generative Diffusion Prior

CVPR 2025poster

Decompositional reconstruction of 3D scenes, with complete shapes and detailed texture of all objects within, is intriguing for downstream applications but remains challenging, particularly with sparse views as input. Recent approaches incorporate semantic or geometric regularization to address this…

2025

DiffDoctor: Diagnosing Image Diffusion Models Before Treating

ICCV 2025poster

In spite of recent progress, image diffusion models still produce artifacts. A common solution is to leverage the feedback provided by quality assessment systems or human annotators to optimize the model, where images are generally rated in their entirety. In this work, we believe problem-solving st…

Cited by 0SourcePDFScholar
2025

DoseSurv: Predicting Personalized Survival Outcomes under Continuous-Valued Treatments

NeurIPS 2025poster

Estimating heterogeneous treatment effects (HTEs) of continuous-valued interventions on survival, that is, time-to-event (TTE) outcomes, is crucial in various fields, notably in clinical decision-making and in driving the advancement of next-generation clinical trials. However, while HTE estimation…

Cited by 0SourceScholar
2025

EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM

ICML 2025poster

Significant achievements in personalization of diffusion models have been witnessed. Conventional tuning-free methods mostly encode multiple reference images by averaging or concatenating their image embeddings as the injection condition, but such an image-independent operation cannot perform intera…

Cited by 6SourcePDFScholar
2025

EfficientPIE: Real-Time Prediction on Pedestrian Crossing Intention with Sole Observation

IJCAI 2025

Present Advanced Driving Assistance System (ADAS) responds to the dangerous crossing of pedestrians after the occurrence of the incident, occasionally causing severe accidents due to the stringent response window. Inference of pedestrian crossing intention may help vehicles operate in advance and en

2025

Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuning

EMNLP 2025

Knowledge graph completion (KGC) aims to infer new knowledge and make predictions from knowledge graphs. Recently, large language models (LLMs) have exhibited remarkable reasoning capabilities. LLM-enhanced KGC methods primarily focus on designing task-specific instructions, achieving promising adva

Cited by 0SourcePDFScholar
2025

Enhancing Semantic Clarity: Discriminative and Fine-grained Information Mining for Remote Sensing Image-Text Retrieval

IJCAI 2025

Remote sensing image-text retrieval is a fundamental task in remote sensing multimodal analysis, promoting the alignment of visual and language representations. The mainstream approaches commonly focus on capturing shared semantic representations between visual and textual modalities. However, the i

Cited by 0SourcePDFScholar
2025

How Distributed Collaboration Influences the Diffusion Model Training? A Theoretical Perspective

ICML 2025poster

This paper examines the theoretical performance of distributed diffusion models in environments where computational resources and data availability vary significantly among workers. Traditional models centered on single-worker scenarios fall short in such distributed settings, particularly when some…

Cited by 0SourcePDFScholar
2025

ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing

ICCV 2025poster

Image generation has witnessed significant advancements in the past few years. However, evaluating the performance of image generation models remains a formidable challenge. In this paper, we propose ICE-Bench, a unified and comprehensive benchmark designed to rigorously assess image generation mode…

2025

IDEA-Bench: How Far are Generative Models from Professional Designing?

CVPR 2025poster

Recent advancements in image generation models enable the creation of high-quality images and targeted modifications based on textual instructions. Some models even support multimodal complex guidance and demonstrate robust task generalization capabilities. However, they still fall short of meeting…

2025

Improved Video VAE for Latent Video Diffusion Model

CVPR 2025poster

Variational Autoencoder (VAE) aims to compress pixel data into low-dimensional latent space, playing an important role in OpenAI's Sora and other latent video diffusion generation models. While most existing video VAEs inflate a pre-trained image VAE into the 3D causal structure for temporal-spatial…

2025

Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy

ICCV 2025poster

Meta-learning is a powerful paradigm for tackling few-shot tasks. However, recent studies indicate that models trained with the whole-class training strategy can achieve comparable performance to those trained with meta-learning in few-shot classification tasks. To demonstrate the value of meta-lear…

Cited by 0SourcePDFScholar
2025

Knocking on IP: Unveiling Websites through Cache-Aware Fingerprinting

ICASSP 2025accepted

As user privacy becomes increasingly critical in the digital landscape, traditional methods of website fingerprinting (WF) face significant challenges, particularly in caching scenarios. Existing WF studies are limited by the assumption of disabled caching. Recently, only a few have explored how to…

Cited by 0SourceScholar
2025

LoD-Loc v2: Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette Alignment

ICCV 2025poster

We propose a novel method for aerial visual localization over low Level-of-Detail (LoD) city models. Previous wireframe-alignment-based method LoD-Loc [99] has shown promising localization results leveraging LoD models. However, LoD-Loc mainly relies on high-LoD (LoD3 or LoD2) city models, but the m…

2025

MMET: A Multi-Input and Multi-Scale Transformer for Efficient PDEs Solving

IJCAI 2025

Partial Differential Equations (PDEs) are fundamental for modeling physical systems, yet solving them in a generic and efficient manner using machine learning-based approaches remains challenging due to limited multi-input and multi-scale generalization capabilities, as well as high computational co

2025

MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines

ICLR 2025poster

The advent of Large Language Models (LLMs) has paved the way for AI search engines, e.g., SearchGPT, showcasing a new paradigm in human-internet interaction. However, most current AI search engines are limited to text-only settings, neglecting the multimodal user queries and the text-image interleav…

Cited by 0SourcePDFScholar
2025

MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes

CVPR 2025poster

Repurposing pre-trained diffusion models has been proven to be effective for NVS. However, these methods are mostly limited to a single object; directly applying such methods to compositional multi-object scenarios yields inferior results, especially incorrect object placement and inconsistent shape…

2025

MPBR: Multimodal Progressive Bidirectional Reasoning for Open-Set Fine-Grained Recognition

ICCV 2025poster

Open-set fine-grained recognition (OSFGR) is the core exploration of building open-world intelligent systems. The challenge lies in the gradual semantic drift during the transition from coarse-grained to fine-grained categories. However, although existing methods leverage hierarchical representation…

Cited by 0SourcePDFScholar
2025

MangaNinja: Line Art Colorization with Precise Reference Following

CVPR 2025highlight

Derived from diffusion models, MangaNinja specializes in the task of reference-guided line art colorization. We incorporate two thoughtful designs to ensure precise character detail transcription, including a patch shuffling module to facilitate correspondence learning between the reference color im…

Cited by 3SourcePDFScholar
2025

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

NeurIPS 2025poster

This work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule-based reinforcement learning for Vision-Language Models (VLMs). However, such methods typically rely on manually curated question-answer pairs, which c…

Cited by 0SourceScholar
2025

NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on Thermodynamics

CVPR 2025poster

Thermal infrared imaging enables a non-invasive measurement of the surface temperature of objects with all-weather applicability. Leveraging such techniques for 3D reconstruction can accurately reflect the temperature distribution of a scene, thereby supporting applications such as building monitori…

Cited by 1SourcePDFScholar
2025

OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection

IJCAI 2025

Out-of-distribution (OOD) detection is crucial for ensuring the reliability and safety of machine learning models in real-world applications. While zero-shot OOD detection, which requires no training on in-distribution (ID) data, has become feasible with the emergence of vision-language models like

Cited by 0SourcePDFScholar
2025

Occlusion-Aware 6D Pose Estimation with Visual Observation Guided Diffusion Model

IROS 2025

Category-level 6D pose estimation in cluttered and occluded environments is a challenging task. Most existing methods rely on deterministic point-based correspondences to estimate target poses, which cannot consider the uncertainty for occluded objects, and thus result in inferior performance. In th

Cited by 0SourceScholar
2025

OpenCarbon: A Contrastive Learning-based Cross-Modality Neural Approach for High-Resolution Carbon Emission Prediction Using Open Data

IJCAI 2025

Accurately estimating high-resolution carbon emissions is crucial for effective emission governance and mitigation planning. While conventional methods for precise carbon accounting are hindered by substantial data collection efforts, the rise of open data and advanced learning techniques offers a p

2025

PDUDT: Provable Decentralized Unlearning under Dynamic Topologies

ICML 2025poster

This paper investigates decentralized unlearning, aiming to eliminate the impact of a specific client on the whole decentralized system. However, decentralized communication characterizations pose new challenges for effective unlearning: the indirect connections make it difficult to trace the specif…

Cited by 0SourcePDFScholar
2025

Pretrained Reversible Generation as Unsupervised Visual Representation Learning

ICCV 2025poster

Recent generative models based on score matching and flow matching have significantly advanced generation tasks, but their potential in discriminative tasks remains underexplored. Previous approaches, such as generative classifiers, have not fully leveraged the capabilities of these models for discr…

2025

Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning

ICRA 2025

Multimodal task specification is essential for enhanced robotic performance, where Cross-modality Alignment enables the robot to holistically understand complex task instructions. Directly annotating multimodal instructions for model training proves impractical, due to the sparsity of paired multimo

Cited by 5SourceScholar
2025

Safety and Naturalness Perceptions of Robot-to-Human Handovers Performed by Data-Driven Robotic Mimicry of Human Givers

ICRA 2025

We study human perceptions of a robot that performs robot-to-human (R2H) handovers controlled to grasp, transport, and transfer 34 objects by mimicking human givers in human-human (H2H) handover data. Recognizing the importance of human-like robotic behavior for successful collaboration, R2H studies

Cited by 1SourceScholar
2025

See Further When Clear: Curriculum Consistency Model

CVPR 2025poster

Significant advances have been made in the sampling efficiency of diffusion and flow matching models, driven by Consistency Distillation (CD), which trains a student model to mimic the output of a teacher model at a later timestep. However, we found that the knowledge discrepancy between student and…

Cited by 0SourcePDFScholar
2025

SmartPretrain: Model-Agnostic and Dataset-Agnostic Representation Learning for Motion Prediction

ICLR 2025poster

Predicting the future motion of surrounding agents is essential for autonomous vehicles (AVs) to operate safely in dynamic, human-robot-mixed environments. However, the scarcity of large-scale driving datasets has hindered the development of robust and generalizable motion prediction models, limitin…

2025

TACO: Taming Diffusion for in-the-wild Video Amodal Completion

ICCV 2025poster

Humans can infer complete shapes and appearances of objects from limited visual cues, relying on extensive prior knowledge of the physical world. However, completing partially observable objects while ensuring consistency across video frames remains challenging for existing models, especially for un…

Cited by 0SourcePDFScholar
2025

TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster

NeurIPS 2025poster

Large Language Models (LLMs) and Foundation Models (FMs) have recently become prevalent for time series forecasting tasks. While fine-tuning LLMs enables domain adaptation, they often struggle to generalize across diverse and unseen datasets. Moreover, existing Time Series Foundation Models (TSFMs)…

Cited by 0SourcecodeScholar
2025

The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control

NeurIPS 2025poster

We present The Matrix, a foundational realistic world simulator capable of generating infinitely long 720p high-fidelity real-scene video streams with real-time, responsive control in both first- and third-person perspectives. Trained on limited supervised data from video games like Forza Horizon 5…

Cited by 0SourceScholar
2025

ThermalLoc: A Vision Transformer-Based Approach for Robust Thermal Camera Relocalization in Large-Scale Environments

IROS 2025

Thermal cameras capture environmental data through heat emission, a fundamentally different mechanism compared to visible light cameras, which rely on pinhole imaging. As a result, traditional visual relocalization methods designed for visible light images are not directly applicable to thermal imag

Cited by 0SourceScholar
2025

ThinkAnswer Loss: Balancing Semantic Similarity and Exact Matching for LLM Reasoning Enhancement

EMNLP 2025

Knowledge distillation for large language models often uses Chain-of-Thought (CoT) and answer pairs, but existing methods struggle with appropriate supervision signals. Uniform constraints (e.g., cross-entropy) on CoT can enforce literal, verbose reasoning and suppress expressive diversity, while so

Cited by 0SourcePDFScholar
2025

UniFuse: A Unified All-in-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and Misalignments

ICCV 2025poster

Current multimodal medical image fusion typically assumes that source images are of high quality and perfectly aligned at the pixel level. Its effectiveness heavily relies on these conditions and often deteriorates when handling misaligned or degraded medical images. To address this, we propose UniF…

2025

Universal Actions for Enhanced Embodied Foundation Models

CVPR 2025poster

Training on diverse, internet-scale data is a key factor in the success of recent large foundation models. Yet, using the same recipe for building embodied agents has faced noticeable difficulties. Despite the availability of many crowd-sourced embodied datasets, their action spaces often exhibit si…

2025

VACE: All-in-One Video Creation and Editing

ICCV 2025poster

Diffusion Transformer has demonstrated powerful capability and scalability in generating high-quality images and videos. Further pursuing the unification of generation and editing tasks has yielded significant progress in the domain of image content creation. However, due to the intrinsic demands fo…

Cited by 0SourcePDFScholar
2025

ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models

NeurIPS 2025poster

Panoramic video generation aims to synthesize 360-degree immersive videos, holding significant importance in the fields of VR, world models, and spatial intelligence. Existing works fail to synthesize high-quality panoramic videos due to the inherent modality gap between panoramic data and perspecti…

Cited by 0SourceScholar
2025

VividFace: A Robost and High-Fidelity Video Face Swapping Framework

NeurIPS 2025poster

Video face swapping has seen increasing adoption in diverse applications, yet existing methods primarily trained on static images struggle to address temporal consistency and complex real-world scenarios. To overcome these limitations, we propose the first video face swapping framework, VividFace,…

Cited by 0SourceScholar
2025

Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance

NeurIPS 2025poster

We present Wan-Move, a simple and scalable framework that brings motion control to video generative models. Existing motion-controllable methods typically suffer from coarse control granularity and limited scalability, leaving their outputs insufficient for practical use. We narrow this gap by achie…

Cited by 0SourceScholar
2024

A Novel Medical Image Fusion Framework Integrating Multi-scale Encoder-Decoder with Discrete Wavelet Decomposition

ICASSP 2024accepted

In recent years, many fusion algorithms based on multi-scale transform or neural networks have been proposed to improve medical image fusion (MIF) performance. However, there is still enormous potential to explore the combination of different fusion theories. In this paper, we propose a novel MIF fr…

Cited by 0SourceScholar
2024

A Perspective of Q-value Estimation on Offline-to-Online Reinforcement Learning

AAAI 2024technical

Offline-to-online Reinforcement Learning (O2O RL) aims to improve the performance of offline pretrained policy using only a few online samples. Built on offline RL algorithms, most O2O methods focus on the balance between RL objective and pessimism, or the utilization of offline and online samples.…

2024

AnyDoor: Zero-shot Object-level Image Customization

CVPR 2024poster

This work presents AnyDoor a diffusion-based image generator with the power to teleport target objects to new scenes at user-specified locations with desired shapes. Instead of tuning parameters for each object our model is trained only once and effortlessly generalizes to diverse object-scene combi…

Cited by 268SourcePDFScholar
2024

Attribution-Based Scanline Perturbation Attack on 3d Detectors of Lidar Point Clouds

ICASSP 2024accepted

LiDAR point cloud data is widely utilized in autonomous driving systems and has significantly improved the 3D detection performance with well-designed deep neural network models. However, due to the complexity of real-world environments and model vulnerability, false detections or malicious attacks…

Cited by 0SourceScholar
2024

Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation

ECCV 2024poster

"Video outpainting is a challenging task, aiming at generating video content outside the viewport of the input video while maintaining inter-frame and intra-frame consistency. Existing methods fall short in either generation quality or flexibility. We introduce (Mastering Video Outpainting Through I…

2024

CCM: Real-Time Controllable Visual Content Creation Using Text-to-Image Consistency Models

ICML 2024poster

Consistency Models (CMs) have showed a promise in creating high-quality images with few steps. However, the way to add new conditional controls to the pre-trained CMs has not been explored. In this paper, we explore the pivotal subject of leveraging the generative capacity and efficiency of consiste…

Cited by 4SourcePDFScholar
2024

CI-STHPAN: Pre-trained Attention Network for Stock Selection with Channel-Independent Spatio-Temporal Hypergraph

AAAI 2024technical

Quantitative stock selection is one of the most challenging FinTech tasks due to the non-stationary dynamics and complex market dependencies. Existing studies rely on channel mixing methods, exacerbating the issue of distribution shift in financial time series. Additionally, complex model structures…

Cited by 17SourcePDFScholar
2024

CPGA: Coding Priors-Guided Aggregation Network for Compressed Video Quality Enhancement

CVPR 2024poster

Recently numerous approaches have achieved notable success in compressed video quality enhancement (VQE). However these methods usually ignore the utilization of valuable coding priors inherently embedded in compressed videos such as motion vectors and residual frames which carry abundant temporal a…

2024

CSCNet: Class-Specified Cascaded Network for Compositional Zero-Shot Learning

ICASSP 2024accepted

Attribute and object (A-O) disentanglement is a fundamental and critical problem for Compositional Zero-shot Learning (CZSL), whose aim is to recognize novel A-O compositions based on foregone knowledge. Existing methods based on disentangled representation learning lose sight of the contextual depe…

Cited by 0SourceScholar
2024

Causality-Inspired Invariant Representation Learning for Text-Based Person Retrieval

AAAI 2024technical

Text-based Person Retrieval (TPR) aims to retrieve relevant images of specific pedestrians based on the given textual query. The mainstream approaches primarily leverage pretrained deep neural networks to learn the mapping of visual and textual modalities into a common latent space for cross-modalit…

Cited by 17SourcePDFScholar
2024

Check Locate Rectify: A Training-Free Layout Calibration System for Text-to-Image Generation

CVPR 2024poster

Diffusion models have recently achieved remarkable progress in generating realistic images. However challenges remain in accurately understanding and synthesizing the layout requirements in the textual prompts. To align the generated image with layout instructions we present a training-free layout c…

2024

CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

NeurIPS 2024poster

Diffusion models have demonstrated great success in the field of text-to-image generation. However, alleviating the misalignment between the text prompts and images is still challenging. We break down the problem into two causes: concept ignorance and concept mismapping. To tackle the two challenges…

2024

Critic-Guided Decision Transformer for Offline Reinforcement Learning

AAAI 2024technical

Recent advancements in offline reinforcement learning (RL) have underscored the capabilities of Return-Conditioned Supervised Learning (RCSL), a paradigm that learns the action distribution based on target returns for each state in a supervised manner. However, prevailing RCSL methods largely focus…

2024

Data-Assisted Dynamic Modeling of Bionic Robotic Fish and Its Precise Speed Control

RA-L 2024

Dynamic modeling is essential for comprehending physical mechanisms and devising control strategies in bionic robot research. This letter introduces a novel dynamic modeling method that combines Lagrangian dynamics and data-assisted techniques for robotic fish with bionic morphology, multi-joint str

Cited by 5SourceScholar
2024

DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning

ICML 2024poster

Multimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: $1)$ extracting both local and global task progressions; $2)$ enforcing temporal consistency of visual representation; $3)$ capturing trajectory-level language grounding. Most ex…

2024

Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models

ECCV 2024poster

"Optimizing a text-to-image diffusion model with a given reward function is an important but underexplored research area. In this study, we propose Deep Reward Tuning (DRTune), an algorithm that directly supervises the final output image of a text-to-image diffusion model and back-propagates through…

Cited by 14SourcePDFScholar
2024

Depth-Guided Dominant Plane Perception for Unsupervised Homography Estimation

ICASSP 2024accepted

Homography describes the mapping relations of the same plane across views. In scenarios with multiple planes, single homography estimation aims to obtain the optimal solution generated by the largest consistent plane to obey the coplanar constraints. However, existing methods typically consider all…

Cited by 0SourceScholar
2024

DreamClean: Restoring Clean Image Using Deep Diffusion Prior

ICLR 2024poster

Image restoration poses a garners substantial interest due to the exponential surge in demands for recovering high-quality images from diverse mobile camera devices, adverse lighting conditions, suboptimal shooting environments, and frequent image compression for efficient transmission purposes. Yet…

Cited by 9SourcePDFScholar
2024

DreamVideo: Composing Your Dream Videos with Customized Subject and Motion

CVPR 2024poster

Customized generation using diffusion models has made impressive progress in image generation but remains unsatisfactory in the challenging video generation task as it requires the controllability of both subjects and motions. To that end we present DreamVideo a novel approach to generating personal…

2024

ESCP: Enhancing Emotion Recognition in Conversation with Speech and Contextual Prefixes

COLING 2024main

Emotion Recognition in Conversation (ERC) aims to analyze the speaker’s emotional state in a conversation. Fully mining the information in multimodal and historical utterances plays a crucial role in the performance of the model. However, recent works in ERC focus on historical utterances modeling a…

Cited by 0SourcePDFScholar
2024

EasyDrag: Efficient Point-based Manipulation on Diffusion Models

CVPR 2024poster

Generative models are gaining increasing popularity and the demand for precisely generating images is on the rise. However generating an image that perfectly aligns with users' expectations is extremely challenging. The shapes of objects the poses of animals the structures of landscapes and more may…

2024

Effect Size Estimation for Duration Recommendation in Online Experiments: Leveraging Hierarchical Models and Objective Utility Approaches

AAAI 2024technical

The selection of the assumed effect size (AES) critically determines the duration of an experiment, and hence its accuracy and efficiency. Traditionally, experimenters determine AES based on domain knowledge. However, this method becomes impractical for online experimentation services managing numer…

Cited by 1SourcePDFScholar
2024

Elegantly Written: Disentangling Writer and Character Styles for Enhancing Online Chinese Handwriting

ECCV 2024poster

"The electronic writing tools, while enhancing convenience, sacrifice the readability and efficiency of handwritten content. Balancing high efficiency with readable handwriting poses a challenging research task. In this paper, we propose a method sequence-based models to beautify user handwritten tr…

2024

Estimating On-Road Transportation Carbon Emissions from Open Data of Road Network and Origin-Destination Flow Data

AAAI 2024technical

Accounting for over 20% of the total carbon emissions, the precise estimation of on-road transportation carbon emissions is crucial for carbon emission monitoring and efficient mitigation policy formulation. However, existing estimation methods typically depend on hard-to-collect individual statisti…

2024

Event-Triggered Adaptive Fault-Tolerant Boundary Control for Flexible Bionic Fish Tail With Output Constraint

RA-L 2024

The article focuses on the tracking issue of a flexible bionic fish tail system under boundary control. The flexible bionic fish tail system is modeled as an Euler-Bernoulli beam with non-uniform parameters, where its actuator is a DC motor located at the front end of the tail. Firstly, the problem

Cited by 3SourceScholar
2024

Exploring Guided Sampling of Conditional GANs

ECCV 2024poster

"Guided sampling serves as a widely used inference technique in diffusion models to trade off sample fidelity and diversity. In this work, we confirm that generative adversarial networks (GANs) can also benefit from guided sampling, not even requiring to pre-prepare a classifier (, classifier guidan…

2024

Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models

NeurIPS 2024poster

Large language models based on decoder-only transformers have demonstrated superior text understanding capabilities compared to CLIP and T5-series models. However, the paradigm for utilizing current advanced LLMs in text-to-image diffusion models remains to be explored. We observed an unusual phenom…

Cited by 17SourcePDFScholar
2024

Fast Context-Based Low-Light Image Enhancement via Neural Implicit Representations

ECCV 2024poster

"Current deep learning-based low-light image enhancement methods often struggle with high-resolution images, and fail to meet the practical demands of visual perception across diverse and unseen scenarios. In this paper, we introduce a novel approach termed CoLIE, which redefines the enhancement pro…

2024

FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis

ECCV 2024poster

"In this study, we delve into the generation of high-resolution images from pre-trained diffusion models, addressing persistent challenges, such as repetitive patterns and structural distortions, that emerge when models are applied beyond their trained resolutions. To address this issue, we introduc…

2024

From Pixels to Progress: Generating Road Network from Satellite Imagery for Socioeconomic Insights in Impoverished Areas

IJCAI 2024poster

The Sustainable Development Goals (SDGs) aim to resolve societal challenges, such as eradicating poverty and improving the lives of vulnerable populations in impoverished areas. Those areas rely on road infrastructure construction to promote accessibility and economic development. Although publicly…

2024

GMP-AR: Granularity Message Passing and Adaptive Reconciliation for Temporal Hierarchy Forecasting

AAAI 2024technical

Time series forecasts of different temporal granularity are widely used in real-world applications, e.g., sales prediction in days and weeks for making different inventory plans. However, these tasks are usually solved separately without ensuring coherence, which is crucial for aligning downstream d…

Cited by 0SourcePDFScholar
2024

Improving Multi-Modal Emotion Recognition Using Entropy-Based Fusion and Pruning-Based Network Architecture Optimization

ICASSP 2024accepted

In this study, we aim to improve our recent hierarchical information fusion system for multi-modal emotion recognition challenge (MER 2023) in both efficiency and performance. Specifically, we extract robust acoustic and visual representations from pre-trained models and fuse them together in differ…

Cited by 0SourceScholar
2024

Instruction-Guided Visual Masking

NeurIPS 2024poster

Instruction following is crucial in contemporary LLM. However, when extended to multimodal setting, it often suffers from misalignment between specific textual instruction and targeted local region of an image. To achieve more accurate and nuanced multimodal instruction following, we introduce Instr…

2024

LMDrive: Closed-Loop End-to-End Driving with Large Language Models

CVPR 2024poster

Despite significant recent progress in the field of autonomous driving modern methods still struggle and can incur serious accidents when encountering long-tail unforeseen events and challenging urban scenarios. On the one hand large language models (LLM) have shown impressive reasoning capabilities…

2024

Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation

CVPR 2024poster

This study focuses on a novel task in text-to-image (T2I) generation namely action customization. The objective of this task is to learn the co-existing action from limited data and generalize it to unseen humans or even animals. Experimental results show that existing subject-driven customization m…

2024

Lipschitz Singularities in Diffusion Models

ICLR 2024oral

Diffusion models, which employ stochastic differential equations to sample images through integrals, have emerged as a dominant class of generative models. However, the rationality of the diffusion process itself receives limited attention, leaving the question of whether the problem is well-posed a…

Cited by 10SourcePDFScholar
2024

LivePhoto: Real Image Animation with Text-guided Motion Control

ECCV 2024poster

"Despite the recent progress in text-to-video generation, existing studies usually overlook the issue that only spatial contents but not temporal motions in synthesized videos are under the control of text. Towards such a challenge, this work presents a practical system, named , which allows users t…

2024

LoD-Loc: Aerial Visual Localization using LoD 3D Map with Neural Wireframe Alignment

NeurIPS 2024poster

We propose a new method named LoD-Loc for visual localization in the air. Unlike existing localization algorithms, LoD-Loc does not rely on complex 3D representations and can estimate the pose of an Unmanned Aerial Vehicle (UAV) using a Level-of-Detail (LoD) 3D map. LoD-Loc mainly achieves this goal…

2024

Long-term Detection and Monitory of Chinese Urban Village Using Satellite Imagery

IJCAI 2024poster

Urban villages are areas filled with rural-like improvised structures in Chinese cities, usually housing the most vulnerable groups. Under the guidance of the Sustainable Development Goals (SDGs), the Chinese government initiated renewal and redevelopment projects, underscoring the meticulous mapp…

2024

MoVA: Adapting Mixture of Vision Experts to Multimodal Context

NeurIPS 2024poster

As the key component in multimodal large language models (MLLMs), the ability of the visual encoder greatly affects MLLM's understanding on diverse image content. Although some large-scale pretrained vision encoders such as vision encoders in CLIP and DINOv2 have brought promising performance, we fo…

2024

Multi-agent Collaborative Perception via Motion-aware Robust Communication Network

CVPR 2024poster

Collaborative perception allows for information sharing between multiple agents such as vehicles and infrastructure to obtain a comprehensive view of the environment through communication and fusion. Current research on multi-agent collaborative perception systems often assumes ideal communication a…

Cited by 4SourcePDFScholar
2024

MultiGen: Zero-shot Image Generation from Multi-modal Prompts

ECCV 2024poster

"The field of text-to-image generation has witnessed substantial advancements in the preceding years, allowing the generation of high-quality images based solely on text prompts. However, accurately describing objects through text alone is challenging, necessitating the integration of additional mod…

Cited by 0SourcePDFScholar
2024

Not Just Object, But State: Compositional Incremental Learning without Forgetting

NeurIPS 2024poster

Most incremental learners excessively prioritize object classes while neglecting various kinds of states (e.g. color and material) attached to the objects. As a result, they are limited in the ability to model state-object compositionality accurately. To remedy this limitation, we propose a novel ta…

2024

Novel Class Discovery for Ultra-Fine-Grained Visual Categorization

CVPR 2024highlight

Ultra-fine-grained visual categorization (Ultra-FGVC) aims at distinguishing highly similar sub-categories within fine-grained objects such as different soybean cultivars. Compared to traditional fine-grained visual categorization Ultra-FGVC encounters more hurdles due to the small inter-class and l…

2024

Phased Consistency Models

NeurIPS 2024poster

Consistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of…

2024

Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following

CVPR 2024poster

Existing text-to-image (T2I) diffusion models usually struggle in interpreting complex prompts especially those with quantity object-attribute binding and multi-subject descriptions. In this work we introduce a semantic panel as the middleware in decoding texts to images supporting the generator to…

Cited by 48SourcePDFScholar
2024

ReDiffuser: Reliable Decision-Making Using a Diffuser with Confidence Estimation

ICML 2024poster

The diffusion model has demonstrated impressive performance in offline reinforcement learning. However, non-deterministic sampling in diffusion models can lead to unstable performance. Furthermore, the lack of confidence measurements makes it difficult to evaluate the reliability and trustworthiness…

Cited by 4SourcePDFScholar
2024

Rethinking the Spatial Inconsistency in Classifier-Free Diffusion Guidance

CVPR 2024poster

Classifier-Free Guidance (CFG) has been widely used in text-to-image diffusion models where the CFG scale is introduced to control the strength of text guidance on the whole image space. However we argue that a global CFG scale results in spatial inconsistency on varying semantic strengths and subop…

2024

Semantics Driven Multi-View Knowledge Graph Embedding for Cross-Lingual Entity Alignment

ICASSP 2024accepted

Cross-lingual entity alignment (EA) is a critical step in the integration of multilingual knowledge, which aims to match entities with the same meaning in different knowledge graphs (KGs). Recently, based on GCN models and pre-trained language models (PLMs), EA has achieved breakthrough performance…

Cited by 0SourceScholar
2024

Sketch-Based 3D Shape Retrieval With Multi-View Fusion Transformer

ICASSP 2024accepted

Sketch-based 3D shape retrieval aims to retrieve similar 3D shapes given a 2D sketch query. Although this task has been studied for years, the inherent cross-modal gap and data imbalance between 2D sketches and 3D shapes remain challenging. To address the problems, we propose a simple and effective…

Cited by 0SourceScholar
2024

SlotLifter: Slot-guided Feature Lifting for Learning Object-Centric Radiance Fields

ECCV 2024poster

"The ability to distill object-centric abstractions from intricate visual scenes underpins human-level generalization. Despite the significant progress in object-centric learning methods, learning object-centric representations in the 3D physical world remains a crucial challenge. In this work, we p…

Cited by 3SourcePDFScholar
2024

SmartRefine: A Scenario-Adaptive Refinement Framework for Efficient Motion Prediction

CVPR 2024poster

Predicting the future motion of surrounding agents is essential for autonomous vehicles (AVs) to operate safely in dynamic human-robot-mixed environments. Context information such as road maps and surrounding agents' states provides crucial geometric and semantic information for motion behavior pred…

2024

StrokeNUWA—Tokenizing Strokes for Vector Graphic Synthesis

ICML 2024poster

To leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model’s ability to capture the true semantic representation of visual scenes. This paper posits that an alternative represent…

Cited by 12SourcePDFScholar
2024

The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models

ICLR 2024poster

Pre-trained Language models (PLMs) have been acknowledged to contain harmful information, such as social biases, which may cause negative social impacts or even bring catastrophic results in application. Previous works on this problem mainly focused on using black-box methods such as probing to dete…

Cited by 15SourcePDFScholar
2024

Three Things We Need to Know About Transferring Stable Diffusion to Visual Dense Prediciton Tasks

ECCV 2024poster

"In this paper, we investigate how to conduct transfer learning to adapt Stable Diffusion to downstream visual dense prediction tasks such as semantic segmentation and depth estimation. We focus on fine-tuning the Stable Diffusion model, which has demonstrated impressive abilities in modeling image…

Cited by 3SourcePDFScholar
2024

UV-SAM: Adapting Segment Anything Model for Urban Village Identification

AAAI 2024technical

Urban villages, defined as informal residential areas in or around urban centers, are characterized by inadequate infrastructures and poor living conditions, closely related to the Sustainable Development Goals (SDGs) on poverty, adequate housing, and sustainable cities. Traditionally, governments h…

2024

Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

NeurIPS 2024spotlight

Multi-Modal Large Language Models (MLLMs) have demonstrated impressive performance in various VQA tasks. However, they often lack interpretability and struggle with complex visual inputs, especially when the resolution of the input image is high or when the interested region that could provide key i…

2024

Zero-shot Image Editing with Reference Imitation

NeurIPS 2024poster

Image editing serves as a practical yet challenging task considering the diverse demands from users, where one of the hardest parts is to precisely describe how the edited image should look like. In this work, we present a new form of editing, termed imitative editing, to help users exercise their c…

Cited by 24SourcePDFScholar
2024

ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model

ECCV 2024oral

"Although video generation has made great progress in capacity and controllability and is gaining increasing attention, currently available video generation models still make minimal progress in the video length they can generate. Due to the lack of well-annotated long video data, high training/infe…

2023

3D Semantic Subspace Traverser: Empowering 3D Generative Model with Shape Editing Capability

ICCV 2023poster

Shape generation is the practice of producing 3D shapes as various representations for 3D content creation. Previous studies on 3D shape generation have focused on shape quality and structure, without or less considering the importance of semantic information. Consequently, such generative models of…

Cited by 3PDFcodeScholar
2023

ACE: Cooperative Multi-Agent Q-learning with Bidirectional Action-Dependency

AAAI 2023technical

Multi-agent reinforcement learning (MARL) suffers from the non-stationarity problem, which is the ever-changing targets at every iteration when multiple agents update their policies at the same time. Starting from first principle, in this paper, we manage to solve the non-stationarity problem by pro…

2023

AIRA-DA: Adversarial Image Reconstruction Alignments for Unsupervised Domain Adaptive Object Detection

RA-L 2023

Unsupervised domain adaptive object detection is a challenging perception task where object detectors are adapted from a label-rich source domain to an unlabeled target domain, playing a vital role in autonomous driving and robot navigation. Since the camera settings, weather, and light conditions v

Cited by 7SourceScholar
2023

Accelerating Reinforcement Learning for Autonomous Driving Using Task-Agnostic and Ego-Centric Motion Skills

IROS 2023poster

Efficient and effective exploration in continuous space is a central problem in applying reinforcement learning (RL) to autonomous driving. Skills learned from expert demonstrations or designed for specific tasks can benefit the exploration, but they are usually costly-collected, unbalanced/suboptim…

Cited by 13SourceScholar
2023

Arbitrary Virtual Try-on Network: Characteristics Representation and Trade-off between Body and Clothing

ICLR 2023poster

Deep learning based virtual try-on system has achieved some encouraging progress recently, but there still remain several big challenges that need to be solved, such as trying on arbitrary clothes of all types, trying on the clothes from one category to another and generating image-realistic results…

Cited by 0SourcePDFScholar
2023

COCA: COllaborative CAusal Regularization for Audio-Visual Question Answering

AAAI 2023technical

Audio-Visual Question Answering (AVQA) is a sophisticated QA task, which aims at answering textual questions over given video-audio pairs with comprehensive multimodal reasoning. Through detailed causal-graph analyses and careful inspections of their learning processes, we reveal that AVQA models ar…

Cited by 21SourcePDFScholar
2023

Centimeter-Scale Underwater Robot With High-Speed Inspired by Jellyfish

RA-L 2023

Centimeter-scale underwater robots have important applications in underwater resource exploration, environmental monitoring, equipment fault diagnosis, and military applications. However, developing small underwater robots remains a challenge because of the limitation of miniaturized structures. In

Cited by 13SourceScholar
2023

Composer: Creative and Controllable Image Synthesis with Composable Conditions

ICML 2023poster

Recent large-scale generative models learned on big data are capable of synthesizing incredible images yet suffer from limited controllability. This work offers a new generation paradigm that allows flexible control of the output image, such as spatial layout and palette, while maintaining the synth…

2023

Cones: Concept Neurons in Diffusion Models for Customized Generation

ICML 2023oral

Human brains respond to semantic features of presented stimuli with different neurons. This raises the question of whether deep neural networks admit a similar behavior pattern. To investigate this phenomenon, this paper identifies a small cluster of neurons associated with a specific subject in a d…

Cited by 119SourcePDFScholar
2023

Customizable Image Synthesis with Multiple Subjects

NeurIPS 2023poster

Synthesizing images with user-specified subjects has received growing attention due to its practical applications. Despite the recent success in single subject customization, existing algorithms suffer from high training cost and low success rate along with increased number of subjects. Towards cont…

Cited by 84SourcePDFScholar
2023

Decoupled DETR: Spatially Disentangling Localization and Classification for Improved End-to-End Object Detection

ICCV 2023poster

The introduction of DETR represents a new paradigm for object detection. However, its decoder conducts classification and box localization using shared queries and cross-attention layers, leading to suboptimal results. We observe that different regions of interest in the visual feature map are sui…

Cited by 23PDFScholar
2023

Deep Active Contours for Real-time 6-DoF Object Tracking

ICCV 2023poster

This paper solves the problem of real-time 6-DoF object tracking from an RGB video. Prior optimization-based methods optimize the object pose by aligning the projected model to the image based on handcrafted features, which are prone to suboptimal solutions. Recent learning-based methods use neural…

Cited by 15PDFcodeScholar
2023

Dimensionality-Varying Diffusion Process

CVPR 2023poster

Diffusion models, which learn to reverse a signal destruction process to generate new data, typically require the signal at each step to have the same dimension. We argue that, considering the spatial redundancy in image signals, there is no need to maintain a high dimensionality in the evolution pr…

2023

Efficient Reinforcement Learning for Autonomous Driving with Parameterized Skills and Priors

RSS 2023poster

When autonomous vehicles are deployed on public roads, they will encounter countless and diverse driving situations. Many manually designed driving policies are difficult to scale to the real world. Fortunately, reinforcement learning has shown great success in many tasks by automatic trial and erro…

2023

Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision Transformers

ICCV 2023poster

Vector-quantized image modeling has shown great potential in synthesizing high-quality images. However, generating high-resolution images remains a challenging task due to the quadratic computational overhead of the self-attention process. In this study, we seek to explore a more efficient two-stage…

Cited by 22PDFScholar
2023

Generating Dynamic Kernels via Transformers for Lane Detection

ICCV 2023poster

State-of-the-art lane detection methods often rely on specific knowledge about lanes -- such as straight lines and parametric curves -- to detect lane lines. While the specific knowledge can ease the modeling process, it poses challenges in handling lane lines with complex topologies (e.g., dense, f…

Cited by 28PDFcodeScholar
2023

GeoMIM: Towards Better 3D Knowledge Transfer via Masked Image Modeling for Multi-view 3D Understanding

ICCV 2023poster

Multi-view camera-based 3D detection is a challenging problem in computer vision. Recent works leverage a pretrained LiDAR detection model to transfer knowledge to a camera-based student network. However, we argue that there is a major domain gap between the LiDAR BEV features and the camera-based B…

Cited by 16PDFcodeScholar
2023

GoBigger: A Scalable Platform for Cooperative-Competitive Multi-Agent Interactive Simulation

ICLR 2023poster

The emergence of various multi-agent environments has motivated powerful algorithms to explore agents' cooperation or competition. Even though this has greatly promoted the development of multi-agent reinforcement learning (MARL), it is still not enough to support further exploration on the behavio…

2023

LightZero: A Unified Benchmark for Monte Carlo Tree Search in General Sequential Decision Scenarios

NeurIPS 2023spotlight

Building agents based on tree-search planning capabilities with learned models has achieved remarkable success in classic decision-making problems, such as Go and Atari. However, it has been deemed challenging or even infeasible to extend Monte Carlo Tree Search (MCTS) based algorithms to diverse re…

2023

Long-Term Visual Localization With Mobile Sensors

CVPR 2023poster

Despite the remarkable advances in image matching and pose estimation, image-based localization of a camera in a temporally-varying outdoor environment is still a challenging problem due to huge appearance disparity between query and reference images caused by illumination, seasonal and structural c…

2023

MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers

CVPR 2023poster

In this paper, we propose Mixed and Masked AutoEncoder (MixMAE), a simple but efficient pretraining method that is applicable to various hierarchical Vision Transformers. Existing masked image modeling (MIM) methods for hierarchical Vision Transformers replace a random subset of input tokens with a…

2023

RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths

NeurIPS 2023poster

Text-to-image generation has recently witnessed remarkable achievements. We introduce a text-conditional image diffusion model, termed RAPHAEL, to generate highly artistic images, which accurately portray the text prompts, encompassing multiple nouns, adjectives, and verbs. This is achieved by stac…

2023

ReasonNet: End-to-End Driving With Temporal and Global Reasoning

CVPR 2023poster

The large-scale deployment of autonomous vehicles is yet to come, and one of the major remaining challenges lies in urban dense traffic scenarios. In such cases, it remains challenging to predict the future evolution of the scene and future behaviors of objects, and to deal with rare adverse events…

Cited by 94SourcePDFScholar
2023

SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on Hierarchies

AAAI 2023technical

Multivariate time series forecasting with hierarchical structure is widely used in real-world applications, e.g., sales predictions for the geographical hierarchy formed by cities, states, and countries. The hierarchical time series (HTS) forecasting includes two sub-tasks, i.e., forecasting and rec…

Cited by 4SourcePDFScholar
2023

Style-Content Metric Learning for Multidomain Remote Sensing Object Recognition

AAAI 2023technical

Previous remote sensing recognition approaches predominantly perform well on the training-testing dataset. However, due to large style discrepancies not only among multidomain datasets but also within a single domain, they suffer from obvious performance degradation when applied to unseen domains. I…

2023

Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object Prediction

ICCV 2023poster

In this paper, we propose a new paradigm, named Historical Object Prediction (HoP) for multi-view 3D detection to leverage temporal information more effectively. The HoP approach is straightforward: given the current timestamp t, we generate a pseudo Bird's-Eye View (BEV) feature of timestamp t-k fr…

Cited by 37PDFcodeScholar
2023

UniKD: Universal Knowledge Distillation for Mimicking Homogeneous or Heterogeneous Object Detectors

ICCV 2023poster

Knowledge distillation (KD) has become a standard method to boost the performance of lightweight object detectors. Most previous works are feature-based, where students mimic the features of homogeneous teacher detectors. However, distilling the knowledge from the heterogeneous teacher fails in this…

Cited by 9PDFScholar
2023

Video Diffusion Models with Local-Global Context Guidance

IJCAI 2023poster

Diffusion models have emerged as a powerful paradigm in video synthesis tasks including prediction, generation, and interpolation. Due to the limitation of the computational budget, existing methods usually implement conditional diffusion models with an autoregressive inference pipeline, in which th…

2022

"UniNet: Unified Architecture Search with Convolution, Transformer, and MLP"

ECCV 2022poster

"Recently, transformer and multi-layer perceptron (MLP) architectures have achieved impressive results on various vision tasks. However, how to effectively combine those operators to form high-performance hybrid visual architectures still remains a challenge. In this work, we study the learnable com…