← Search

Xin Chen

123 accepted papers

2026

AERO-MPPI: Anchor-Guided Ensemble Trajectory Optimization for Agile Mapless Drone Navigation

ICRA 2026poster

Agile mapless navigation in cluttered 3D environments poses significant challenges for autonomous drones. Conventional mapping–planning–control pipelines incur high computational cost and propagate estimation errors. We present AERO-MPPI, a fully GPU-accelerated framework that unifies perception and…

2026

Bridging Your Imagination with Audio-Video Generation via a Unified Director

ICML 2026poster

Existing AI-driven video creation systems typically treat script drafting and key-shot design as two disjoint tasks: the former relies on large language models, while the latter depends on image generation models. We argue that these two tasks should be unified within a single framework, as logical …

Cited by 2SourceScholar
2026

D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool Use

ICML 2026poster

Effective tool use and reasoning are essential capabilities for large reasoning models (LRMs) to address complex real-world problems. Through empirical analysis, we identify a prevalent "Lazy Reasoning" phenomenon, where LRMs frequently engage in repetitive and meaningless reflective reasoning. This…

Cited by 0SourceScholar
2026

EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuning

ICML 2026poster

Domain-specific large language models (LLMs), typically developed by fine-tuning a pre-trained general-purpose LLM on specialized datasets, represent a significant advancement in applied AI. A common strategy in LLM fine-tuning is curriculum learning, which pre-orders training samples based on metri…

Cited by 0SourceScholar
2026

Evolving Interdependent Operators with Large Language Models for Multi-Objective Combinatorial Optimization

ICML 2026poster

Neighborhood search operators are critical to the performance of Multi-Objective Evolutionary Algorithms (MOEAs) and rely heavily on expert design. Although recent LLM-based Automated Heuristic Design (AHD) methods have made notable progress, they primarily optimize individual heuristics or componen…

Cited by 0SourceScholar
2026

InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs

CVPR 2026

Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for multi-agent interactions. To bridge this gap, we propose InterAg

Cited by 0SourcecodeScholar
2026

LRHDR: Learning Representation-enhanced HDR Video Reconstruction

CVPR 2026

Reconstructing High Dynamic Range (HDR) video from alternately exposed Low Dynamic Range (LDR) frames is challenged by large motion, exposure-induced photometric inconsistency, and information loss in saturated or under-exposed regions. Prior HDR video pipelines typically follow an alignment-reconst

Cited by 0SourceScholar
2026

MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models

CVPR 2026

Post-training quantization (PTQ) with computational equivalence for Large Language Models (LLMs) have demonstrated remarkable advances, however, their application to Multimodal Large Language Models (MLLMs) presents substantial challenges. In this paper, we analyze SmoothQuant as a case study and id

Cited by 0SourcecodeScholar
2026

MotionGPT3: Human Motion as a Second Modality

ICLR 2026poster

With the rapid progress of large language models (LLMs), multimodal frameworks that unify understanding and generation have become promising, yet they face increasing complexity as the number of modalities and tasks grows. We observe that motion quantization introduces approximation errors that cap…

Cited by 0SourceScholar
2026

Neuron-Aware Data Selection in Instruction Tuning for Large Language Models

ICLR 2026poster

Instruction Tuning (IT) has been proven to be an effective approach to unlock the powerful capabilities of large language models (LLMs). Recent studies indicate that excessive IT data can degrade LLMs performance, while carefully selecting a small subset of high-quality IT data can significantly en…

Cited by 0SourceScholar
2026

OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation

CVPR 2026

We introduce OneCAT, a unified multimodal model that seamlessly integrates understanding, generation, and editing within a single decoder-only transformer architecture. OneCAT uniquely eliminates the need for external components such as Vision Transformers (ViT) or vision tokenizer during inference,

Cited by 0SourcecodeScholar
2026

Parallelizable Riemannian Alternating Direction Method of Multipliers for Non-convex Pose Graph Optimization

AAAI 2026technical

Pose graph optimization (PGO) is fundamental to robot perception and navigation systems, serving as the mathematical backbone for solving simultaneous localization and mapping (SLAM). Existing solvers suffer from polynomial growth in computational complexity with graph size, hindering real-time depl

Cited by 0SourcePDFScholar
2026

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

ICML 2026oral

We argue that many Anthropomorphized Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for critical safety decisions, such as model deployment and regulation. By evaluating failure modes across different misalignment concepts, such as dece…

Cited by 0SourceScholar
2026

Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs

ICLR 2026poster

Multimodal Large Language Models (MLLMs) pose critical safety challenges, as they are susceptible not only to adversarial attacks such as jailbreaking but also to inadvertently generating harmful content for benign users. While internal safety alignment via Supervised Fine-Tuning (SFT) and Reinforce…

Cited by 0SourceScholar
2026

PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis

AAAI 2026technical

Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during pa

Cited by 0SourcePDFScholar
2026

RELO: Reinforcement Learning to Localize for Visual Object Tracking

ICML 2026poster

Existing one-stream Transformer-based visual trackers localize targets by training a classification head with a handcrafted spatial prior encoded as a heatmap. However, this heuristic supervision merely serves as a surrogate objective, which misaligns with evaluation metrics such as IoU and AUC. To …

Cited by 0SourceScholar
2026

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers

ICML 2026poster

The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placement of normalization layers, leading to a fundamental trade-off: the ''PreNorm'' architecture ensures training stability at the cost of potential perform…

Cited by 0SourceScholar
2026

TGTrack: Temporal Generative Learning for Unified Single Object Tracking

CVPR 2026

Existing single object trackers typically treat temporal modeling superficially by passing limited inter-frame information, such as propagated tokens or template updates, without intrinsic temporal supervision learning. To address this limitation, we propose TGTrack, a new unified tracking framework

Cited by 0SourcecodeScholar
2026

Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion Transformers

CVPR 2026

Existing video frame interpolation (VFI) methods often adopt a frame-centric approach, processing videos as independent short segments (e.g., triplets), which leads to temporal inconsistencies and motion artifacts. To overcome this, we propose a holistic, video-centric paradigm named Local Diffusion

Cited by 0SourcecodeScholar
2026

UETrack: A Unified and Efficient Framework for Single Object Tracking

CVPR 2026

With growing real-world demands, efficient tracking has received increasing attention. However, most existing methods are limited to RGB inputs and struggle in multi-modal scenarios. Moreover, current multi-modal tracking approaches typically use complex designs, making them too heavy and slow for r

Cited by 0SourcecodeScholar
2026

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

CVPR 2026

Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human behavior as a translation task--co-speech gesture or text-to-motion that maps a fixed utterance to motion clips--without requiring agentic decision-maki

Cited by 0SourceScholar
2026

Video-As-Prompt: Unified Semantic Control for Video Generation

ICLR 2026poster

Unified, generalizable semantic control in video generation remains a critical open challenge. Existing methods either introduce artifacts by enforcing inappropriate pixel-wise priors from structure-based controls, or rely on non-generalizable, condition-specific finetuning or task-specific architec…

Cited by 0SourcecodeScholar
2025

Disentangled Graph Spectral Domain Adaptation

ICML 2025poster

The distribution shifts and the scarcity of labels prevent graph learning methods, especially graph neural networks (GNNs), from generalizing across domains. Compared to Unsupervised Domain Adaptation (UDA) with embedding alignment, Unsupervised Graph Domain Adaptation (UGDA) becomes more challengin…

Cited by 0SourcePDFScholar
2025

DoDo-Code: an Efficient Levenshtein Distance Embedding-based Code for 4-ary IDS Channel

NeurIPS 2025poster

With the emergence of new storage and communication methods, the insertion, deletion, and substitution (IDS) channel has attracted considerable attention. However, many topics on the IDS channel and the associated Levenshtein distance remain open, making the invention of a novel IDS-correcting code…

Cited by 0SourceScholar
2025

ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded Gaps

AAAI 2025technical

Solving jigsaw puzzles has been extensively studied. While most existing models focus on solving either small-scale puzzles or puzzles with no gap between fragments, solving large-scale puzzles with gaps presents distinctive challenges in both image understanding and combinatorial optimization. To t…

Cited by 0SourcePDFScholar
2025

ESCNet:Edge-Semantic Collaborative Network for Camouflaged Object Detection

ICCV 2025poster

Camouflaged object detection (COD) faces unique challenges where target boundaries are intrinsically ambiguous due to their textural similarity to backgrounds. Existing methods relying on single-modality features often produce fragmented predictions due to insufficient boundary constraints.To addres…

2025

Efficient Motion Prompt Learning for Robust Visual Tracking

ICML 2025poster

Due to the challenges of processing temporal information, most trackers depend solely on visual discriminability and overlook the unique temporal coherence of video data. In this paper, we propose a lightweight and plug-and-play motion prompt tracking method. It can be easily integrated into existin…

2025

Exploring Enhanced Contextual Information for Video-Level Object Tracking

AAAI 2025technical

Contextual information at the video level has become increasingly crucial for visual object tracking. However, existing methods typically use only a few tokens to convey this information, which can lead to information loss and limit their ability to fully capture the context. To address this issue,…

2025

From Generalist to Specialist: A Survey of Large Language Models for Chemistry

COLING 2025main

Large Language Models (LLMs) have significantly transformed our daily life and established a new paradigm in natural language processing (NLP). However, the predominant pretraining of LLMs on extensive web-based texts remains insufficient for advanced scientific discovery, particularly in chemistry.…

2025

La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation

ICML 2025poster

Activation sparsity can reduce the computational overhead and memory transfers during the forward pass of Large Language Model (LLM) inference. Existing methods face limitations, either demanding time-consuming recovery training that hinders real-world adoption, or relying on empirical magnitude-bas…

Cited by 0SourcePDFScholar
2025

Learning Dynamic Collaborative Network for Semi-supervised 3D Vessel Segmentation

CVPR 2025poster

In this paper, we present a new dynamic collaborative network for semi-supervised 3D vessel segmentation, termed DiCo. Conventional mean teacher (MT) methods typically employ a static approach, where the roles of the teacher and student models are fixed. However, due to the complexity of 3D vessel d…

2025

Leveraging Unpaired Feedback for Long-Term LLM-based Recommendation Tuning

EMNLP 2025

Most recommender systems focus on short-term objectives such as click-through rate, often at the expense of long-term user satisfaction. This can lead to echo chambers, where users are repeatedly exposed to redundant content. While recent efforts integrate Large Language Models (LLMs) into recommend

2025

MS-BART: Unified Modeling of Mass Spectra and Molecules for Structure Elucidation

NeurIPS 2025poster

Mass spectrometry (MS) plays a critical role in molecular identification, significantly advancing scientific discovery. However, structure elucidation from MS data remains challenging due to the scarcity of annotated spectra. While large-scale pretraining has proven effective in addressing data scar…

Cited by 0SourcecodeScholar
2025

MikuDance: Animating Character Art with Mixed Motion Dynamics

ICCV 2025poster

We propose MikuDance, a diffusion-based pipeline incorporating mixed motion dynamics to animate stylized character art. MikuDance consists of two key techniques: Mixed Motion Modeling and Mixed-Control Diffusion, to address the challenges of high-dynamic motion and reference-guidance misalignment in…

Cited by 0SourcePDFScholar
2025

PNetGPT: Proprietary Protocol Network Traffic Generation with Pre-trained Transformer

ICASSP 2025accepted

Generative pre-trained transformers are exceedingly effective as generative models and classifiers, widely used in natural language processing and computer vision. This work contributes to the exploration of generative pre-trained transformer-based models in the proprietary protocol network traffic.…

Cited by 0SourceScholar
2025

SUTrack: Towards Simple and Unified Single Object Tracking

AAAI 2025technical

In this paper, we propose a simple yet unified single object tracking (SOT) framework, dubbed SUTrack. It consolidates five SOT tasks (RGB-based, RGB-Depth, RGB-Thermal, RGB-Event, RGB-Language Tracking) into a unified model trained in a single session. Due to the distinct nature of the data, curren…

2025

Swin-VasMamba: A Topologically Constrained Model For 3D Vascular Segmentation

ICASSP 2025accepted

Accurate 3D vascular segmentation is essential for diagnosing and treating vascular diseases. This task remains challenging due to the complexity of the 3D data and the morphological diversity of blood vessels. In recent years, state space models (SSMs) have received a great attention for its good p…

Cited by 0SourceScholar
2025

Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs

EMNLP 2025

As large language models (LLMs) often generate plausible but incorrect content, error detection has become increasingly critical to ensure truthfulness.However, existing detection methods often overlook a critical problem we term as **self-consistent error**, where LLMs repeatedly generate the same

2025

Two-stream Beats One-stream: Asymmetric Siamese Network for Efficient Visual Tracking

AAAI 2025technical

Efficient tracking has garnered attention for its ability to operate on resource-constrained platforms for real-world deployment beyond desktop GPUs. Current efficient trackers mainly follow precision-oriented trackers, adopting a one-stream framework with lightweight modules. However, blindly adher…

2025

X-Dancer: Expressive Music to Human Dance Video Generation

ICCV 2025poster

We present X-Dancer, a novel zero-shot music-driven image animation pipeline that creates diverse and long-range lifelike human dance videos from a single static image. As its core, we introduce a unified transformer-diffusion framework, featuring an autoregressive transformer model that synthesize…

Cited by 0SourcePDFScholar
2025

YOLO-TCT: An Effective Network For Long-Tailed Cervical Cell Detection

ICASSP 2025accepted

The Thinprep Cytologic Test (TCT) is a vital component in the early detection of cervical cancer. However, conventional manual screening methods are hindered by inefficiencies and high levels of subjectivity. This study presents YOLO-TCT, an enhanced YOLOv9 network designed for the automated detecti…

Cited by 0SourceScholar
2024

3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection

NeurIPS 2024poster

Transformer-based architectures have been proven successful in detecting 3D objects from point clouds. However, the quadratic complexity of the attention mechanism struggles to encode rich information as point cloud resolution increases. Recently, state space models (SSM) such as Mamba have gained g…

Cited by 0SourcePDFScholar
2024

Achieving Optimal Clustering in Gaussian Mixture Models with Anisotropic Covariance Structures

NeurIPS 2024oral

We study clustering under anisotropic Gaussian Mixture Models (GMMs), where covariance matrices from different clusters are unknown and are not necessarily the identity matrix. We analyze two anisotropic scenarios: homogeneous, with identical covariance matrices, and heterogeneous, with distinct mat…

Cited by 1SourcePDFScholar
2024

CycleINR: Cycle Implicit Neural Representation for Arbitrary-Scale Volumetric Super-Resolution of Medical Data

CVPR 2024poster

In the realm of medical 3D data such as CT and MRI images prevalent anisotropic resolution is characterized by high intra-slice but diminished inter-slice resolution. The lowered resolution between adjacent slices poses challenges hindering optimal viewing experiences and impeding the development of…

Cited by 2SourcePDFScholar
2024

Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models

NeurIPS 2024poster

Diffusion models (DMs) have achieved remarkable success in text-to-image generation, but they also pose safety risks, such as the potential generation of harmful content and copyright violations. The techniques of machine unlearning, also known as concept erasing, have been developed to address thes…

2024

Dual Contrastive Learning Guided Pathological Image Re-Staining

ICASSP 2024accepted

Pathological virtual re-staining is a valuable research topic in AI-aided diagnosis, as it reduces the need for costly and time-consuming physical staining. However, existing methods still suffer from the insufficient ability to preserve tissue microstructure and cellular details, making the generat…

Cited by 0SourceScholar
2024

Exploiting Spatial-Temporal Data for Sleep Stage Classification via Hypergraph Learning

ICASSP 2024accepted

Sleep stage classification is crucial for detecting patients’ health conditions. Existing models, which mainly use Convolutional Neural Networks (CNN) for modelling Euclidean data and Graph Convolution Networks (GNN) for modelling non-Euclidean data, are unable to consider the heterogeneity and inte…

Cited by 0SourceScholar
2024

Gradient descent in matrix factorization: Understanding large initialization

UAI 2024poster

Gradient Descent (GD) has been proven effective in solving various matrix factorization problems. However, its optimization behavior with large initial values remains less understood. To address this gap, this paper presents a novel theoretical framework for examining the convergence trajectory of G…

Cited by 2SourcePDFScholar
2024

LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding Reasoning and Planning

CVPR 2024poster

Recent progress in Large Multimodal Models (LMM) has opened up great possibilities for various applications in the field of human-machine interactions. However developing LMMs that can comprehend reason and plan in complex and diverse 3D environments remains a challenging topic especially considerin…

2024

Learning Safety Constraints from Demonstrations with Unknown Rewards

AISTATS 2024poster

We propose Convex Constraint Learning for Reinforcement Learning (CoCoRL), a novel approach for inferring shared constraints in a Constrained Markov Decision Process (CMDP) from a set of safe demonstrations with possibly different reward functions. While previous work is limited to demonstrations wi…

2024

Learning to Maximize Mutual Information for Chain-of-Thought Distillation

ACL 2024findings

Knowledge distillation, the technique of transferring knowledge from large, complex models to smaller ones, marks a pivotal step towards efficient AI deployment. Distilling Step-by-Step (DSS), a novel method utilizing chain-of-thought (CoT) distillation, has demonstrated promise by imbuing smaller m…

2024

LocMoE: A Low-overhead MoE for Large Language Model Training

IJCAI 2024poster

The Mixtures-of-Experts (MoE) model is a widespread distributed and integrated learning method for large language models (LLM), which is favored due to its ability to sparsify and expand models efficiently. However, the performance of MoE is limited by load imbalance and high latency of All-to-All c…

Cited by 13SourcePDFScholar
2024

M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions

ECCV 2024poster

"Recently, the understanding of the 3D world has garnered increased attention, facilitating autonomous agents to perform further decision-making. However, the majority of existing 3D vision-language datasets and methods are often limited to specific tasks, limiting their applicability in diverse sce…

Cited by 0SourcePDFScholar
2024

MeshXL: Neural Coordinate Field for Generative 3D Foundation Models

NeurIPS 2024poster

The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately,…

2024

MotionChain: Conversational Motion Controllers via Multimodal Prompts

ECCV 2024poster

"Recent advancements in language models have demonstrated their adeptness in conducting multi-turn dialogues and retaining conversational context. However, this proficiency remains largely unexplored in other multimodal generative models, particularly in human motion models. By integrating multi-tur…

2024

OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers

CVPR 2024poster

We have recently seen tremendous progress in realistic text-to-motion generation. Yet the existing methods often fail or produce implausible motions with unseen text inputs which limits the applications. In this paper we present OMG a novel framework which enables compelling motion generation from z…

2024

PM-INR: Prior-Rich Multi-Modal Implicit Large-Scale Scene Neural Representation

AAAI 2024technical

Recent advancements in implicit neural representations have contributed to high-fidelity surface reconstruction and photorealistic novel view synthesis. However, with the expansion of the scene scale, such as block or city level, existing methods will encounter challenges because traditional samplin…

Cited by 2SourcePDFScholar
2024

Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models

CVPR 2024poster

This paper presents Paint3D a novel coarse-to-fine generative framework that is capable of producing high-resolution lighting-less and diverse 2K UV texture maps for untextured 3D meshes conditioned on text or image inputs. The key challenge addressed is generating high-quality textures without embe…

2024

Plug-In Diffusion Model for Sequential Recommendation

AAAI 2024technical

Pioneering efforts have verified the effectiveness of the diffusion models in exploring the informative uncertainty for recommendation. Considering the difference between recommendation and image synthesis tasks, existing methods have undertaken tailored refinements to the diffusion and reverse proc…

2024

QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

ICLR 2024poster

Recently years have witnessed a rapid development of large language models (LLMs). Despite the strong ability in many language-understanding tasks, the heavy computational burden largely restricts the application of LLMs especially when one needs to deploy them onto edge devices. In this paper, we p…

2024

REGLO: Provable Neural Network Repair for Global Robustness Properties

AAAI 2024technical

We present REGLO, a novel methodology for repairing pretrained neural networks to satisfy global robustness and individual fairness properties. A neural network is said to be globally robust with respect to a given input region if and only if all the input points in the region are locally robust. Th…

2024

Safety-First Tracker: A Trajectory Planning Framework for Omnidirectional Robot Tracking

IROS 2024poster

This paper introduces a Safety-First Tracker (SF-Tracker) designed for omnidirectional autonomous tracking robots. The position and orientation of omnidirectional robots are decoupled for stepwise planning to ensure trajectory safety and maintain target visibility. SF-Tracker puts the trajectory saf…

Cited by 0SourcecodeScholar
2024

TapMo: Shape-aware Motion Generation of Skeleton-free Characters

ICLR 2024poster

Previous motion generation methods are limited to the pre-rigged 3D human model, hindering their applications in the animation of various non-rigged characters. In this work, we present TapMo, a Text-driven Animation PIpeline for synthesizing Motion in a broad spectrum of skeleton-free 3D characters…

Cited by 11SourcePDFScholar
2024

To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now

ECCV 2024poster

"The recent advances in diffusion models (DMs) have revolutionized the generation of realistic and complex images. However, these models also introduce potential safety hazards, such as producing harmful content and infringing data copyrights. Despite the development of safety-driven unlearning tech…

2024

Unbounded-GS: Extending 3D Gaussian Splatting With Hybrid Representation for Unbounded Large-Scale Scene Reconstruction

RA-L 2024

Modeling large-scale scenes from multi-view images is challenging due to the trade-off dilemma between visual quality and computational cost. Existing NeRF-based methods have made advancements in neural implicit representation through volumetric ray-marching, but still struggle to deal with cubicall

Cited by 7SourceScholar
2023

A Large-Scale Outdoor Multi-Modal Dataset and Benchmark for Novel View Synthesis and Implicit Scene Reconstruction

ICCV 2023poster

Neural Radiance Fields (NeRF) has achieved impressive results in single object scene reconstruction and novel view synthesis, as demonstrated on many single modality and single object focused indoor scene datasets like DTU, BMVS, and NeRF Synthetic. However, the study of NeRF on large-scale outdoor…

Cited by 29PDFScholar
2023

CancerUniT: Towards a Single Unified Model for Effective Detection, Segmentation, and Diagnosis of Eight Major Cancers Using a Large Collection of CT Scans

ICCV 2023poster

Human readers or radiologists routinely perform full-body multi-organ multi-disease detection and diagnosis in clinical practice, while most medical AI systems are built to focus on single organs with a narrow list of a few diseases. This might severely limit AI's clinical adoption. A certain number…

Cited by 12PDFScholar
2023

Devil Is in the Queries: Advancing Mask Transformers for Real-World Medical Image Segmentation and Out-of-Distribution Localization

CVPR 2023highlight

Real-world medical image segmentation has tremendous long-tailed complexity of objects, among which tail conditions correlate with relatively rare diseases and are clinically significant. A trustworthy medical AI algorithm should demonstrate its effectiveness on tail conditions to avoid clinically d…

Cited by 28SourcePDFScholar
2023

End-to-End 3D Dense Captioning With Vote2Cap-DETR

CVPR 2023poster

3D dense captioning aims to generate multiple captions localized with their associated object regions. Existing methods follow a sophisticated "detect-then-describe" pipeline equipped with numerous hand-crafted components. However, these hand-crafted components would yield suboptimal performance giv…

2023

Executing Your Commands via Motion Diffusion in Latent Space

CVPR 2023poster

We study a challenging task, conditional human motion generation, which produces plausible human motion sequences according to various conditional inputs, such as action classes or textual descriptors. Since human motions are highly diverse and have a property of quite different distribution from co…

2023

Exploring Lightweight Hierarchical Vision Transformers for Efficient Visual Tracking

ICCV 2023poster

Transformer-based visual trackers have demonstrated significant progress owing to their superior modeling capabilities. However, existing trackers are hampered by low speed, limiting their applicability on devices with limited computational power. To alleviate this problem, we propose HiT, a new fam…

Cited by 70PDFcodeScholar
2023

Fan-Beam Binarization Difference Projection (FB-BDP): A Novel Local Object Descriptor for Fine-Grained Leaf Image Retrieval

ICCV 2023poster

Fine-grained leaf image retrieval (FGLIR) aims to search similar leaf images in subspecies level which involves very high interclass visual similarity and accordingly poses great challenges to leaf image description. In this study, we introduce a new concept, named fan-beam binarization difference p…

Cited by 4PDFcodeScholar
2023

Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation

NeurIPS 2023poster

We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone to producing inconsistent results with the conditions becaus…

2023

MotionGPT: Human Motion as a Foreign Language

NeurIPS 2023poster

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multimodal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to human language, often perc…

2023

PDF: Point Diffusion Implicit Function for Large-scale Scene Neural Representation

NeurIPS 2023poster

Recent advances in implicit neural representations have achieved impressive results by sampling and fusing individual points along sampling rays in the sampling space. However, due to the explosively growing sampling space, finely representing and synthesizing detailed textures remains a challenge f…

Cited by 5SourcePDFScholar
2023

SeqTrack: Sequence to Sequence Learning for Visual Object Tracking

CVPR 2023poster

In this paper, we present a new sequence-to-sequence learning framework for visual tracking, dubbed SeqTrack. It casts visual tracking as a sequence generation problem, which predicts object bounding boxes in an autoregressive fashion. This is different from prior Siamese trackers and transformer tr…

2023

Sketched Ridgeless Linear Regression: The Role of Downsampling

ICML 2023poster

Overparametrization often helps improve the generalization performance. This paper presents a dual view of overparametrization suggesting that downsampling may also help generalize. Focusing on the proportional regime $m\asymp n \asymp p$, where $m$ represents the sketching size, $n$ is the sample s…

2023

Text-Visual Prompting for Efficient 2D Temporal Video Grounding

CVPR 2023poster

In this paper, we study the problem of temporal video grounding (TVG), which aims to predict the starting/ending time points of moments described by a text sentence within a long untrimmed video. Benefiting from fine-grained 3D visual features, the TVG techniques have achieved remarkable progress in…

2022

Anisotropic Fourier Features for Neural Image-Based Rendering and Relighting

AAAI 2022technical

Recent neural rendering techniques have greatly benefited image-based modeling and relighting tasks. They provide a continuous, compact, and parallelable representation by modeling the plenoptic function as multilayer perceptrons (MLPs). However, vanilla MLPs suffer from spectral biases on multidime…

Cited by 7SourcePDFScholar
2022

Arch-Graph: Acyclic Architecture Relation Predictor for Task-Transferable Neural Architecture Search

CVPR 2022poster

Neural Architecture Search (NAS) aims to find efficient models for multiple tasks. Beyond seeking solutions for a single task, there are surging interests in transferring network design knowledge across multiple tasks. In this line of research, effectively modeling task correlations is vital yet hig…

Cited by 25PDFcodeScholar
2022

Improve Single-Point Zeroth-Order Optimization Using High-Pass and Low-Pass Filters

ICML 2022spotlight

Single-point zeroth-order optimization (SZO) is useful in solving online black-box optimization and control problems in time-varying environments, as it queries the function value only once at each time step. However, the vanilla SZO method is known to suffer from a large estimation variance and slo…

Cited by 23SourcePDFScholar
2022

Pruning-as-Search: Efficient Neural Architecture Search via Channel Pruning and Structural Reparameterization

IJCAI 2022poster

Neural architecture search (NAS) and network pruning are widely studied efficient AI techniques, but not yet perfect. NAS performs exhaustive candidate architecture search, incurring tremendous search cost. Though (structured) pruning can simply shrink model dimension, it remains unclear how to de…

Cited by 48SourcePDFScholar
2022

Wassertrain: An Adversarial Training Framework Against Wasserstein Adversarial Attacks

ICASSP 2022accepted

This paper presents an adversarial training framework WasserTrain for improving model robustness against the adversarial attacks in terms of the Wasserstein distance. First, an effective attack method WasserAttack is introduced with a novel encoding of the optimization problem, which directly finds…

Cited by 0SourceScholar
2022

Wavelet Knowledge Distillation: Towards Efficient Image-to-Image Translation

CVPR 2022poster

Remarkable achievements have been attained with Generative Adversarial Networks (GANs) in image-to-image translation. However, due to a tremendous amount of parameters, state-of-the-art GANs usually suffer from low efficiency and bulky memory usage. To tackle this challenge, firstly, this paper inve…

Cited by 104PDFScholar
2021

An Empirical Investigation of Representation Learning for Imitation

NeurIPS 2021poster

Imitation learning often needs a large demonstration set in order to handle the full range of situations that an agent might find itself in during deployment. However, collecting expert demonstrations can be expensive. Recent work in vision, reinforcement learning, and NLP has shown that auxiliary r…

Cited by 33SourceScholar
2021

ChallenCap: Monocular 3D Capture of Challenging Human Performances Using Multi-Modal References

CVPR 2021poster

Capturing challenging human motions is critical for numerous applications, but it suffers from complex motion patterns and severe self-occlusion under the monocular setting. In this paper, we propose ChallenCap --- a template-based approach to capture challenging 3D human motions using a single RGB…

Cited by 28PDFScholar
2021

Correlation-Aware Heuristic Search for Intelligent Virtual Machine Provisioning in Cloud Systems

AAAI 2021technical

The optimization of resource is crucial for the operation of public cloud systems such as Microsoft Azure, as well as servers dedicated to the workloads of large customers such as Microsoft 365. Those optimization tasks often need to take unknown parameters into consideration and can be formulated a…

2021

Exploring Geometry-Aware Contrast and Clustering Harmonization for Self-Supervised 3D Object Detection

ICCV 2021poster

Current 3D object detection paradigms highly rely on extensive annotation efforts, which makes them not practical in many real-world industrial applications. Inspired by that a human driver can keep accumulating experiences from self-exploring the roads without any tutor's guidance, we first step fo…

Cited by 86PDFcodeScholar
2021

FFA-IR: Towards an Explainable and Reliable Medical Report Generation Benchmark

NeurIPS 2021poster

The automatic generation of long and coherent medical reports given medical images (e.g. Chest X-ray and Fundus Fluorescein Angiography (FFA)) has great potential to support clinical practice. Researchers have explored advanced methods from computer vision and natural language processing to incorpor…

Cited by 48SourcecodeScholar
2021

Few-shot Neural Human Performance Rendering from Sparse RGBD Videos

IJCAI 2021poster

Recent neural rendering approaches for human activities achieve remarkable view synthesis results, but still rely on dense input views or dense training with all the capture frames, leading to deployment difficulty and inefficient training overload. However, existing advances will be ill-posed if th…

Cited by 17SourcePDFScholar
2021

Fitting the Search Space of Weight-sharing NAS with Graph Convolutional Networks

AAAI 2021technical

Neural architecture search has attracted wide attentions in both academia and industry. To accelerate it, researchers proposed weight-sharing methods which first train a super-network to reuse computation among different operators, from which exponentially many sub-networks can be sampled and effici…

Cited by 20SourcePDFScholar
2021

TransNAS-Bench-101: Improving Transferability and Generalizability of Cross-Task Neural Architecture Search

CVPR 2021poster

Recent breakthroughs of Neural Architecture Search (NAS) extend the field's research scope towards a broader range of vision tasks and more diversified search spaces. While existing NAS methods mostly design architectures on a single task, algorithms that look beyond single-task search are surging t…

Cited by 82PDFScholar
2020

Biased Stochastic First-Order Methods for Conditional Stochastic Optimization and Applications in Meta Learning

NeurIPS 2020poster

Conditional stochastic optimization covers a variety of applications ranging from invariant learning and causal inference to meta-learning. However, constructing unbiased gradient estimators for such problems is challenging due to the composition structure. As an alternative, we propose a biased sto…

Cited by 73SourcePDFScholar
2020

CATCH: Context-based Meta Reinforcement Learning for Transferrable Architecture Search

ECCV 2020poster

Neural Architecture Search (NAS) achieved many breakthroughs in recent years. In spite of its remarkable progress, many algorithms are restricted to particular search spaces. They also lack efficient mechanisms to reuse knowledge when confronting multiple tasks. These challenges preclude their appli…

Cited by 24SourcePDFScholar
2020

Circumventing Outliers of AutoAugment with Knowledge Distillation

ECCV 2020poster

AutoAugment has been a powerful algorithm that improves the accuracy of many vision tasks, yet it is sensitive to the operator space as well as hyper-parameters, and an improper setting may degenerate network optimization. This paper delves deep into the working mechanism, and reveals that AutoAugme…

Cited by 76SourcePDFScholar
2020

Graph Stochastic Neural Networks for Semi-supervised Learning

NeurIPS 2020poster

Graph Neural Networks (GNNs) have achieved remarkable performance in the task of the semi-supervised node classification. However, most existing models learn a deterministic classification function, which lack sufficient flexibility to explore better choices in the presence of kinds of imperfect ob…

2020

Intelligent Virtual Machine Provisioning in Cloud Computing

IJCAI 2020poster

Virtual machine (VM) provisioning is a common and critical problem in cloud computing. In industrial cloud platforms, there are a huge number of VMs provisioned per day. Due to the complexity and resource constraints, it needs to be carefully optimized to make cloud platforms effectively utilize the…

2020

Max orientation coverage: efficient path planning to avoid collisions in the CNC milling of 3D objects

IROS 2020poster

Most path planning algorithms for covering a complex 3D object ignore physical limitations or constraints on a robot's motion. Adhering to such constraints for a given path can slow down the time to cover the path because the motion may need to be adjusted. This work considers a scenario in computer…

Cited by 2SourceScholar
2020

Navigating Discrete Difference Equation Governed WMR by Virtual Linear Leader Guided HMPC

ICRA 2020poster

In this paper, we revisit model predictive control (MPC) for the classical wheeled mobile robot (WMR) navigation problem. We prove that the reachable set based hierarchical MPC (HMPC), a state-of-the-art MPC, cannot handle WMR navigation in theory due to the non-existence of non-trivial linear syste…

Cited by 2SourceScholar
2020

PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search

ICLR 2020spotlight

Differentiable architecture search (DARTS) provided a fast solution in finding effective network architectures, but suffered from large memory and computing overheads in jointly training a super-net and searching for an optimal architecture. In this paper, we present a novel approach, namely Partia…

Cited by 920SourcecodeScholar
2020

Path Planning Under MIMO Network Constraints for Throughput Enhancement in Multi-robot Data Aggregation Tasks

IROS 2020poster

Under line-of-sight (LOS) network conditions, multi-input multi-output (MIMO) wireless communications can increase the channel capacity between a team of robots and a multi-antenna array at a stationary base station. This increased capacity can result in greater data throughput, shortening the time…

Cited by 6SourceScholar
2020

ReachFlow: An Online Safety Assurance Framework for Waypoint-Following of Self-driving Cars

IROS 2020poster

Learning-enabled components have been widely deployed in autonomous systems. However, due to the weak interpretability and the prohibitively high complexity of large-scale machine learning models such as neural networks, reliability has been a crucial concern for safety-critical autonomous systems.…

Cited by 15SourceScholar
2019

Adaptive Deep Path: Efficient Coverage of a Known Environment under Various Configurations

IROS 2019poster

Coverage path planning of a known environment sees a variety of applications, including cleaning, surveillance, agriculture and 3D printing. Most approaches employ hard-coded heuristics or other application-specific requirements, making them hard to extend to other problem scenarios or “configuratio…

Cited by 24SourceScholar
2019

Enhancing Low Light Videos by Exploring High Sensitivity Camera Noise

ICCV 2019poster

Enhancing low light videos, which consists of denoising and brightness adjustment, is an intriguing but knotty problem. Under low light condition, due to high sensitivity camera setting, commonly negligible noises become obvious and severely deteriorate the captured videos. To recover high quality v…

Cited by 63PDFScholar
2019

Online Optimal Control with Linear Dynamics and Predictions: Algorithms and Regret Analysis

NeurIPS 2019poster

This paper studies the online optimal control problem with time-varying convex stage costs for a time-invariant linear dynamical system, where a finite lookahead window of accurate predictions of the stage costs are available at each time. We design online algorithms, Receding Horizon Gradient-based…

Cited by 115SourcePDFScholar
2019

Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and Evaluation

ICCV 2019oral

Recently, differentiable search methods have made major progress in reducing the computational costs of neural architecture search. However, these approaches often report lower accuracy in evaluating the searched architecture or transferring it to another dataset. This is arguably due to the large g…

Cited by 847PDFcodeScholar
2019

Robustness Verification of Classification Deep Neural Networks via Linear Programming

CVPR 2019poster

There is a pressing need to verify robustness of classification deep neural networks (CDNNs) as they are embedded in many safety-critical applications. Existing robustness verification approaches rely on computing the over-approximation of the output set, and can hardly scale up to practical CDNNs,…

Cited by 49PDFScholar
2018

Sparse Photometric 3D Face Reconstruction Guided by Morphable Models

CVPR 2018poster

We present a novel 3D face reconstruction technique that leverages sparse photometric stereo (PS) and latest advances on face registration / modeling from a single image. We observe that 3D morphable faces approach provides a reasonable geometry proxy for light position calibration. Specifically, we…

Cited by 41SourcePDFScholar
2017

A regularized on-line sequential extreme learning machine with forgetting property for fast dynamic hysteresis modeling

IROS 2017poster

Piezoelectric ceramics(PZT)actuator has been widely used in flexure-guided nanopositioning stage because of their high resolution. However, it is quite hard to achieve high-rate precision positioning control because of the complex hysteresis nonlinearity effect of PZT actuator. Thus, an online RELM…

Cited by 2SourceScholar
2015

A novel pooling strategy for Full Reference Image Quality Assessment based on harmonic means

ICASSP 2015accepted

The most perceptual Full Reference Image Quality Assessment metrics (FR-IQA) shared a common two-step model; local quality measurement, and pooling. In this letter, a novel pooling strategy based on harmonic mean is proposed to predict the final quality score in FR-IQA. In contrast to arithmetic mea…

Cited by 0SourceScholar