← Search

Gang Xu

21 accepted papers

2026

DC-MAMBER: A DUAL CHANNEL PREDICTION MODEL BASED ON MAMBA AND LINEAR TRANSFORMER FOR MULTIVARIATE TIME SERIES FORECASTING

ICASSP 2026poster

In multivariate time series forecasting (MTSF), existing strategies for processing sequences are typically categorized as channel-independent and channel-mixing. The former treats all temporal information of each variable as a token, focusing on capturing local temporal features of individual variab…

Cited by 0SourcePDFScholar
2026

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

ICML 2026poster

Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerable to multimodal jailbreak attacks. Existing defenses predominantly rely on safety fine-tuning or \textit{aggressive} token manipulations, incurring subs…

Cited by 0SourceScholar
2026

StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback

AAAI 2026technical

The advancement of intelligent agents has revolutionized problem-solving across diverse domains, yet solutions for personalized fashion styling remain underexplored, which holds immense promise for promoting shopping experiences. In this work, we present StyleTailor, the first collaborative agent fr

Cited by 0SourcePDFScholar
2026

SwarmNav: Swarm Robotics Navigation in Dynamic and Dense Environments Via Reinforcement Learning

ICRA 2026poster

Collision avoidance and navigation in dynamic and dense environments remain highly challenging for swarm robotics. To address this, we propose SwarmNav, a novel goal-region amplification navigation policy that leverages LiDAR-based position data to generate velocity commands guiding robots toward th…

Cited by 0Scholar
2026

Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion

CVPR 2026

Realistic 3D city generation is fundamental to a wide range of applications, including virtual reality and digital twins. However, most existing methods rely on training a single diffusion model, which limits their ability to generate personalized and boundless city-scale scenes. In this paper, we p

Cited by 0SourceScholar
2025

Detail-Preserving Latent Diffusion for Stable Shadow Removal

CVPR 2025poster

Achieving high-quality shadow removal with strong generalizability is challenging in scenes with complex global illumination. Due to the limited diversity in shadow removal datasets, current methods are prone to overfitting training data, often leading to reduced performance on unseen cases. To addr…

Cited by 0SourcePDFScholar
2025

Efficient Multi-Robot Task and Path Planning in Large-Scale Cluttered Environments

RA-L 2025

As the potential of multi-robot systems continues to be explored and validated across various real-world applications, such as package delivery, search and rescue, and autonomous exploration, the need to improve the efficiency and quality of task and path planning has become increasingly urgent, par

Cited by 4SourceScholar
2025

FLARE: Fast Large-Scale Autonomous Exploration Guided by Unknown Regions

RA-L 2025

Autonomous exploration is a critical foundation for unmanned aerial vehicle (UAV) applications such as search and rescue. However, existing methods typically focus only on known spaces or frontiers without considering unknown regions or providing further guidance for the global path, which results i

Cited by 2SourceScholar
2025

Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction

IJCAI 2025

Recent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive obje

2025

MARF: Cooperative Multi-Agent Path Finding with Reinforcement Learning and Frenet Lattice in Dynamic Environments

ICRA 2025

Multi-agent path finding (MAPF) in dynamic and complex environments is a highly challenging task. Recent research has focused on the scalability of agent numbers or the complexity of the environment. Usually, they disregard the agents' physical constraints or use a differential-driven model. However

Cited by 1SourceScholar
2025

OmniSR: Shadow Removal Under Direct and Indirect Lighting

AAAI 2025technical

Shadows can originate from occlusions in both direct and indirect illumination. Although most current shadow removal research focuses on shadows caused by direct illumination, shadows from indirect illumination are often just as pervasive, particularly in indoor scenes. A significant challenge in re…

2025

PVChat: Personalized Video Chat with One-Shot Learning

ICCV 2025poster

Video large language models (ViLLMs) excel in general video understanding, e.g., recognizing activities like talking and eating, but struggle with identity-aware comprehension, such as "Wilson is receiving chemotherapy" or "Tom is discussing with Sarah", limiting their applicability in smart healthc…

Cited by 0SourcePDFScholar
2025

WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models

ICML 2025poster

Invisible watermarking is widely used to protect digital images from unauthorized use. Accurate assessment of watermarking efficacy is crucial for advancing algorithmic development. However, existing statistical metrics, such as PSNR, rely on access to original images, which are often unavailable in…

2024

Hierarchical Search-Based Cooperative Motion Planning

IROS 2024poster

Cooperative path planning, a crucial aspect of multi-agent systems research, serves a variety of sectors, including military, agriculture, and industry. Many existing algorithms, however, come with certain limitations, such as simplified kinematic models and inadequate support for multiple group sce…

Cited by 0SourcecodeScholar
2024

Trusted Deep Domain Adaptation with Uncertainty Measure Based on Evidence Theory

ICASSP 2024accepted

Domain adaptation aims to build up an adaptive model for the learning tasks on target domain by using the data from source domain. Existing domain adaptation methods focused on reducing the gap between the data distributions of source domain and target domain but neglected the uncertainty of source…

Cited by 0SourceScholar
2023

Masked and Adaptive Transformer for Exemplar Based Image Translation

CVPR 2023poster

We present a novel framework for exemplar based image translation. Recent advanced methods for this task mainly focus on establishing cross-domain semantic correspondence, which sequentially dominates image generation in the manner of local style control. Unfortunately, cross domain semantic matchin…

2023

Semantic-Aware Generation of Multi-View Portrait Drawings

IJCAI 2023poster

Neural radiance fields (NeRF) based methods have shown amazing performance in synthesizing 3D-consistent photographic images, but fail to generate multi-view portrait drawings. The key is that the basic assumption of these methods -- a surface point is consistent when rendered from different views -…

2023

Shunted Collision Avoidance for Multi-UAV Motion Planning with Posture Constraints

ICRA 2023poster

This paper investigates the problem of fixed-wing unmanned aerial vehicles (UAV s) motion planning with posture constraints and the problem of the more general symmetrical situations where UAVs have more than one optimal solution. In this paper, the posture constraints are formulated in the 3D Dubin…

Cited by 4SourcecodeScholar
2023

TraCo: Learning Virtual Traffic Coordinator for Cooperation with Multi-Agent Reinforcement Learning

CoRL 2023poster

Multi-agent reinforcement learning (MARL) has emerged as a popular technique in diverse domains due to its ability to automate system controller design and facilitate continuous intelligence learning. For instance, traffic flow is often trained with MARL to enable intelligent simulations for autonom…

Cited by 2SourceScholar
2021

Temporal Modulation Network for Controllable Space-Time Video Super-Resolution

CVPR 2021poster

Space-time video super-resolution (STVSR) aims to increase the spatial and temporal resolutions of low-resolution and low-frame-rate videos. Recently, deformable convolution based methods have achieved promising STVSR performance, but they could only infer the intermediate frame pre-defined in the t…

Cited by 113PDFcodeScholar
2019

CAMEL: A Weakly Supervised Learning Framework for Histopathology Image Segmentation

ICCV 2019accepted

Histopathology image analysis plays a critical role in cancer diagnosis and treatment. To automatically segment the cancerous regions, fully supervised segmentation algorithms require labor-intensive and time-consuming labeling at the pixel level. In this research, we propose CAMEL, a weakly supervi…