← Search

Kai Wang

158 accepted papers

2026

A Fully First-Order Layer for Differentiable Optimization

ICML 2026spotlight

Differentiable optimization studies how to embed a mathematical program as a differentiable layer in machine learning pipelines. However, existing approaches typically rely on implicit differentiation, involving expensive Hessian computation while differentiating through optimality conditions. To ad…

Cited by 1SourceScholar
2026

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models

CVPR 2026

In Vision-Language-Action (VLA) models, action chunking (i.e., executing a sequence of actions without intermediate replanning) is a key technique to improve robotic manipulation abilities. However, a large chunk size reduces the model's responsiveness to new information, while a small one increases

Cited by 0SourcecodeScholar
2026

Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3D

CVPR 2026

Current 2D-to-3D conversion methods achieve geometric accuracy but are artistically deficient, failing to replicate the immersive and emotionally resonant experience of professional 3D cinema. This is because "geometric reconstruction" paradigms mistake deliberate artistic intent--such as strategic

Cited by 0SourceScholar
2026

Boost the Identity-Preserving Embedding for Consistent Visual Generation

ICML 2026poster

Text-to-image models have advanced high-fidelity content generation, but their inability to maintain subject consistency hampers realistic applications. Existing training-based methods rely on heavy computation and large datasets; while training-free approaches demand excessive memory or complex aux…

Cited by 0SourceScholar
2026

Diffusion-DFL: Decision-focused Diffusion Models for Stochastic Optimization

ICLR 2026poster

Decision-focused learning (DFL) integrates predictive modeling and optimization by training predictors to optimize the downstream decision target rather than merely minimizing prediction error. To date, existing DFL methods typically rely on deterministic point predictions, which are often insuffici…

Cited by 0SourcecodeScholar
2026

EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision

AAAI 2026technical

Egocentric visual query localization is vital for embodied AI and VR/AR, yet remains challenging due to camera motion, viewpoint changes, and appearance variations. We present EAGLE, a novel framework that leverages episodic appearance- and geometry-aware memory to achieve unified 2D-3D visual query

Cited by 0SourcePDFScholar
2026

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models

ICML 2026poster

Visual Autoregressive (VAR) modeling departs from the next-token prediction paradigm of traditional Autoregressive (AR) models through next-scale prediction, enabling high-quality image generation. However, the VAR paradigm suffers from sharply increased computational complexity and running time at …

Cited by 0SourceScholar
2026

Fractal Camouflage: A Bio-Inspired Approach for Multi-Scale Adversarial Attacks in the Infrared Domain

CVPR 2026

Infrared pedestrian detection is crucial in safety-critical systems but remains vulnerable to adversarial attacks. Existing physical attacks often rely on fixed, static patterns. However, they often lack robustness across scales, as their hand-crafted or uniformly generated structures are fundamenta

Cited by 0SourceScholar
2026

GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation

CVPR 2026

Recent years have seen impressive advances in text-to-image generation, with image generative or unified models, generating high-quality images from text. Yet these models still struggle with fine-grained color control, often failing to accurately match colors specified in text prompts. While existi

Cited by 0SourcecodeScholar
2026

Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization

CVPR 2026

Generating high-fidelity audio that is both semantically meaningful and temporally synchronized with silent videos remains a challenging problem in video-to-audio generation. Existing approaches often fail to capture fine-grained temporal correspondence between visual events and audio dynamics, lead

Cited by 0SourcecodeScholar
2026

HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment

AAAI 2026technical

Contrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-fo

Cited by 0SourcePDFScholar
2026

HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image

CVPR 2026

In this paper, we present HumanNOVA, a photorealistic, universal, and rapid model for generating 3D human avatars from a single RGB image. Achieving both photorealism and generalization is challenging due to the scarcity of diverse, high-quality 3D human data. To address this, we build a scalable da

Cited by 0SourcecodeScholar
2026

JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

ICLR 2026poster

Recent AIGC advances have rapidly expanded from text-to-image generation toward high-quality multimodal synthesis across video and audio. Within this context, joint audio-video generation (JAVG) has emerged as a fundamental task that produces synchronized and semantically aligned sound and vision fr…

Cited by 0SourcecodeScholar
2026

MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models

AAAI 2026technical

Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured kn

Cited by 0SourcePDFScholar
2026

MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference

ICLR 2026poster

We present MeanCache, a training-free caching framework for efficient Flow Matching inference. Existing caching methods reduce redundant computation but typically rely on instantaneous velocity information (e.g., feature caching), which often leads to severe trajectory deviations and error accumulat…

Cited by 0SourceScholar
2026

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

CVPR 2026

In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds across diverse domains (e.g., ambient events, musical instruments, and human speech). Prior video-conditioned audio genera

Cited by 0SourcecodeScholar
2026

RAPID$^3$: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer

ICLR 2026poster

Diffusion Transformers (DiTs) excel at visual generation yet remain hampered by slow sampling. Existing training-free accelerators—step reduction, feature caching, and sparse attention—enhance inference speed but typically rely on a uniform heuristic or manually designed adaptive strategy for all i…

Cited by 0SourceScholar
2026

SDNet: LiDAR Semantic Scene Completion with Sparse-Dense Fusion and Input-Aware Label Refinement

AAAI 2026technical

LiDAR Semantic Scene Completion (SSC) in autonomous driving requires predicting both dense occupancy and semantic labels from sparse input point cloud. Existing methods typically adopt cascaded architecture for feature dilation and semantic abstraction, which blurs distinctive geometric patterns and

Cited by 0SourcePDFScholar
2026

UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy

ICLR 2026poster

Computational replication of Chinese calligraphy, a cornerstone of cultural heritage, remains challenging. Existing methods split into two flawed camps: some render high-quality isolated characters yet miss page-level aesthetics (ligatures, spacing, scale), while others attempt page/column synthesis…

Cited by 0SourcecodeScholar
2025

$InterLCM$: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration

ICLR 2025poster

Diffusion priors have been used for blind face restoration (BFR) by fine-tuning diffusion models (DMs) on restoration datasets to recover low-quality images. However, the naive application of DMs presents several key limitations. (i) The diffusion prior has inferior semantic consistency (e.g., ID,…

Cited by 1SourcePDFScholar
2025

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

CVPR 2025poster

Vision-language models (VLMs) have shown remarkable success across various multi-modal tasks, yet large VLMs encounter significant efficiency challenges due to processing numerous visual tokens. A promising approach to accelerating large VLM inference is using partial information, such as attention…

2025

AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction

ICCV 2025poster

Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views, especially when there is a significant camera pose difference, leading to poor-quality 3D geometries and textures. We…

2025

Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing

ICCV 2025poster

Text-to-image generation has seen groundbreaking advancements with diffusion models, enabling high-fidelity synthesis and precise image editing through cross-attention manipulation. Recently, autoregressive (AR) models have re-emerged as powerful alternatives, leveraging next-token generation to mat…

2025

CALLIC: Content Adaptive Learning for Lossless Image Compression

AAAI 2025technical

Learned lossless image compression has achieved significant advancements in recent years. However, existing methods often rely on training amortized generative models on massive datasets, resulting in sub-optimal probability distribution estimation for specific testing images during encoding process…

Cited by 1SourcePDFScholar
2025

Covariances for Free: Exploiting Mean Distributions for Training-free Federated Learning

NeurIPS 2025poster

Using pre-trained models has been found to reduce the effect of data heterogeneity and speed up federated learning algorithms. Recent works have explored training-free methods using first- and second-order statistics to aggregate local client data distributions at the server and achieve high perform…

Cited by 0SourcecodeScholar
2025

DcDsDiff: Dual-Conditional and Dual-Stream Diffusion Model for Generative Image Tampering Localization

IJCAI 2025

Generative Image Tampering (GIT), due to its high diversity and realism, poses a significant challenge to traditional image tampering localization techniques. Consequently, this paper introduces a denoising diffusion probabilistic model-based DcDsDiff, which comprises a Dual-View Conditional Network

2025

Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights

NeurIPS 2025poster

Modern Parameter-Efficient Fine-Tuning (PEFT) methods such as low-rank adaptation (LoRA) reduce the cost of customizing large language models (LLMs), yet still require a separate optimization run for every downstream dataset. We introduce \textbf{Drag-and-Drop LLMs (\textit{DnD})}, a prompt-conditio…

Cited by 0SourcecodeScholar
2025

Drawing Informative Gradients from Sources: A One-stage Transfer Learning Framework for Cross-city Spatiotemporal Forecasting

AAAI 2025technical

Spatiotemporal forecasting (STF) is pivotal in urban computing, yet data scarcity in developing cities hampers robust model training. Addressing this, recent studies leverage transfer learning to migrate knowledge from data-rich (source) to data-poor (target) cities. This strategy, while effective,…

Cited by 0SourcePDFScholar
2025

Dynamic Diffusion Transformer

ICLR 2025poster

Diffusion Transformer (DiT), an emerging diffusion model for image generation, has demonstrated superior performance but suffers from substantial computational costs. Our investigations reveal that these costs stem from the static inference paradigm, which inevitably introduces redundant computation…

2025

EA-Vit: Efficient Adaptation for Elastic Vision Transformer

ICCV 2025poster

Vision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming…

2025

ElaD-Net: An Elastic Semantic Decoupling Network for Lesion Segmentation in Breast Ultrasound Images

IJCAI 2025

Breast diseases pose a significant threat to women’s health. Automatic lesion segmentation in breast ultrasound images (BUSI) plays a crucial role in fast diagnosis. While various enhanced U-Net-based models have achieved success in multi-scale feature analysis and handling blurred boundaries, two k

Cited by 0SourcePDFScholar
2025

Emphasizing Discriminative Features for Dataset Distillation in Complex Scenarios

CVPR 2025poster

Dataset distillation has demonstrated strong performance on simple datasets like CIFAR, MNIST, and TinyImageNet but struggles to achieve similar results in more complex scenarios. In this paper, we propose EDF (emphasizes the discriminative features), a dataset distillation method that enhances key…

2025

FilterTS: Comprehensive Frequency Filtering for Multivariate Time Series Forecasting

AAAI 2025technical

Multivariate time series forecasting is crucial across various industries, where accurate extraction of complex periodic and trend components can significantly enhance prediction performance. However, existing models often struggle to capture these intricate patterns. To address these challenges, we…

2025

Free-Lunch Color-Texture Disentanglement for Stylized Image Generation

NeurIPS 2025poster

Recent advances in Text-to-Image (T2I) diffusion models have transformed image generation, enabling significant progress in stylized generation using only a few style reference images. However, current diffusion-based methods struggle with \textit{fine-grained} style customization due to challenges…

Cited by 0SourceScholar
2025

From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face Aging

NeurIPS 2025poster

Face aging has become a crucial task in computer vision, with applications ranging from entertainment to healthcare. However, existing methods struggle with achieving a realistic and seamless transformation across the entire lifespan, especially when handling large age gaps or extreme head poses. Th…

Cited by 0SourcecodeScholar
2025

Fuzzy Reasoning Chain (FRC): An Innovative Reasoning Framework from Fuzziness to Clarity

EMNLP 2025

With the rapid advancement of large language models (LLMs), natural language processing (NLP) has achieved remarkable progress. Nonetheless, significant challenges remain in handling texts with ambiguity, polysemy, or uncertainty. We introduce the Fuzzy Reasoning Chain (FRC) framework, which integra

Cited by 0SourcePDFScholar
2025

GIPD: Global Intent Prediction and Decomposition of Cooperative Multi-Robot System in Non-Communication Environments

IROS 2025

In complex multi-robot application scenarios, particularly in dynamically adversarial, hazardous, or disaster environments, traditional cooperation paradigms face significant challenges due to unreliable or absent communication links. Achieving efficient cooperation in the absence of communication h

Cited by 0SourceScholar
2025

HANet: A Harmonic Attention-Based Network for Singing Melody Extraction from Polyphonic Music

ICASSP 2025accepted

Singing melody extraction from polyphonic music is a complex but important task in music information retrieval. Harmonic relationships have been shown to be crucial in this task, but most existing models based on Convolutional Neural Networks (CNNs) struggle to capture long-range harmonic dependenci…

Cited by 0SourceScholar
2025

IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance

ICCV 2025poster

Ensuring precise multimodal alignment between diffusion-generated images and input prompts has been a long-standing challenge. Earlier works finetune diffusion weight using high-quality preference data, which tends to be limited and difficult to scale up. Recent editing-based methods further refine…

2025

Info-Coevolution: An Efficient Framework for Data Model Coevolution

ICML 2025poster

Machine learning relies heavily on data, yet the continuous growth of real-world data poses challenges for efficient dataset construction and training. A fundamental yet unsolved question is: given our current model and data, does a new data (sample/batch) need annotation/learning? Conventional appr…

2025

InpDiffusion: Image Inpainting Localization via Conditional Diffusion Models

AAAI 2025technical

As artificial intelligence advances rapidly, particularly with the advent of GANs and diffusion models, the accuracy of Image Inpainting Localization (IIL) has become increasingly challenging. Current IIL methods face two main challenges: a tendency towards overconfidence, leading to incorrect predi…

2025

LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation

NeurIPS 2025spotlight

We present LeMiCa, a training-free and efficient acceleration framework for diffusion-based video generation. While existing caching strategies primarily focus on reducing local heuristic errors, they often overlook the accumulation of global errors, leading to noticeable content degradation between…

Cited by 0SourceScholar
2025

MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

ICLR 2025poster

We present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users. Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal task…

2025

MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification

ACL 2025finding

Reasoning is an essential capacity for large language models (LLMs) to address complex tasks, whereas the identification of process errors is vital for improving this ability. Recently, process-level reward models (PRMs) were proposed to provide step-wise rewards that facilitate reinforcement learni…

Cited by 0SourcePDFScholar
2025

ORAL: Prompting Your Large-Scale LoRAs via Conditional Recurrent Diffusion

EMNLP 2025

Parameter generation has emerged as a novel paradigm for neural network development, offering an alternative to traditional neural network training by synthesizing high-quality model weights directly. In the context of Low-Rank Adaptation (LoRA) for evolving ( i.e, constantly updated) large language

2025

One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

ICLR 2025spotlight

Text-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem typically require extensive training in large datasets or additiona…

2025

One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models

CVPR 2025poster

Text-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling ste…

2025

Optimizing for the Shortest Path in Denoising Diffusion Model

CVPR 2025highlight

In this research, we propose a novel denoising diffusion model based on shortest-path modeling that optimizes residual propagation to enhance both denoising efficiency and quality. Drawing on Denoising Diffusion Implicit Models (DDIM) and insights from graph theory, our model, termed the Shortest Pa…

2025

Permitted Knowledge Boundary: Evaluating the Knowledge-Constrained Responsiveness of Large Language Models

EMNLP 2025

With the advancement of large language models (LLMs), recent research has raised concerns about their controllability.. In this paper, we argue for the importance of Knowledge-Constrained Responsiveness (KCR), ensuring that LLMs comply with human-defined constraints. However, KCR is an implicit and

2025

Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths

NeurIPS 2025poster

Mamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's cha…

Cited by 0SourceScholar
2025

REPA Works Until It Doesn’t: Early-Stopped, Holistic Alignment Supercharges Diffusion Training

NeurIPS 2025poster

Diffusion Transformers (DiTs) deliver state-of-the-art image quality, yet their training remains notoriously slow. A recent remedy---representation alignment (REPA) that matches DiT hidden features to those of a non-generative teacher (e.g., DINO)---dramatically accelerates the early epochs but plat…

Cited by 0SourcecodeScholar
2025

Real-Time Video Generation with Pyramid Attention Broadcast

ICLR 2025poster

We present Pyramid Attention Broadcast (PAB), a real-time, high quality and training-free approach for DiT-based video generation. Our method is founded on the observation that attention difference in the diffusion process exhibits a U-shaped pattern, indicating significant redundancy. We mitigate t…

2025

Robust and Efficient Text-based Speech Editing using Noise Conditioning and Rectified Flow

ICASSP 2025accepted

Significant advancements have been made in text-based speech editing (TSE) for clear speech, but effectively editing the noise-contaminated speech remains a challenge. Background noise degrades the quality of generated speech, and edited speech that fails to maintain noise context consistency often…

Cited by 0SourceScholar
2025

Scaling Up Parameter Generation: A Recurrent Diffusion Approach

NeurIPS 2025poster

Parameter generation has long struggled to match the scale of today's large vision and language models, curbing its broader utility. In this paper, we introduce Recurrent Diffusion for Large-Scale Parameter Generation (RPG), a novel framework that generates full neural network parameters—up to hundr…

Cited by 0SourceScholar
2025

Self-Improvement in Multimodal Large Language Models: A Survey

EMNLP 2025

Recent advancements in self-improvement for Large Language Models (LLMs) have efficiently enhanced model capabilities without significantly increasing costs, particularly in terms of human effort. While this area is still relatively young, its extension to the multimodal domain holds immense potenti

2025

Single-View Graph Contrastive Learning with Soft Neighborhood Awareness

AAAI 2025technical

Most graph contrastive learning (GCL) methods heavily rely on cross-view contrast, thus facing several concomitant challenges, such as the complexity of designing effective augmentations, the potential for information loss between views, and increased computational costs. To mitigate reliance on cro…

2025

StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training

NeurIPS 2025poster

Training Transformer models on long sequences in a distributed setting poses significant challenges in terms of efficiency and scalability. Current methods are either constrained by the number of attention heads or excessive communication overheads. To address this problem, we propose StarTrail, a m…

Cited by 0SourceScholar
2025

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

ICCV 2025poster

We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQVAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabiliti…

2025

The Art of Deception: Color Visual Illusions and Diffusion Models

CVPR 2025poster

Visual illusions in humans arise when interpreting out-of-distribution stimuli: if the observer is adapted to certain statistics, perception of outliers deviates from reality. Recent studies have shown that artificial neural networks (ANNs) can also be deceived by visual illusions.This revelation ra…

Cited by 0SourcePDFScholar
2025

Time-Frequency Disentanglement Boosted Pre-Training: A Universal Spatio-Temporal Modeling Framework

IJCAI 2025

Current spatio-temporal modeling techniques largely rely on the abundant data and the design of task-specific models. However, many cities lack well-established digital infrastructures, making data scarcity and the high cost of model development significant barriers to application deployment. Theref

Cited by 0SourcePDFScholar
2025

Time-Space-Interlaced Spatiotemporal Graph Forecasting via Two-Stage Summarized Attention

ICASSP 2025accepted

Typical spatiotemporal graph forecasting methods process graph-structured spatiotemporal data respectively from spatial and temporal perspectives with the idea of divide and conquer. Existing works are incapable of capturing long-term transdimensional correlations among different spatial points in d…

Cited by 0SourceScholar
2025

Towards Graph Foundation Models: Training on Knowledge Graphs Enables Transferability to General Graphs

NeurIPS 2025poster

Inspired by the success of large language models, there is a trend toward developing graph foundation models to conduct diverse downstream tasks in various domains. However, current models often require extra fine-tuning to apply their learned structural and semantic representations to new graphs, w…

Cited by 0SourceScholar
2025

Unsupervised Learning for Class Distribution Mismatch

ICML 2025poster

Class distribution mismatch (CDM) refers to the discrepancy between class distributions in training data and target tasks. Previous methods address this by designing classifiers to categorize classes known during training, while grouping unknown or new classes into an "other" category. However, they…

2025

What is the Right Notion of Distance between Predict-then-Optimize Tasks?

UAI 2025

Comparing datasets is a fundamental task in machine learning, essential for various learning paradigms-from evaluating train and test datasets for model generalization to using dataset similarity for detecting data drift. While traditional notions of dataset distances offer principled measures of si

2025

X-Field: A Physically Informed Representation for 3D X-ray Reconstruction

NeurIPS 2025spotlight

X-ray imaging is indispensable in medical diagnostics, yet its use is tightly regulated due to radiation exposure. Recent research borrows representations from the 3D reconstruction area to complete two tasks with reduced radiation dose: X-ray Novel View Synthesis (NVS) and Computed Tomography (CT)…

Cited by 0SourceScholar
2024

A Large Vision-Language Model based Environment Perception System for Visually Impaired People

IROS 2024poster

It is a challenging task for visually impaired people to perceive their surrounding environment due to the complexity of the natural scenes. Their personal and social activities are thus highly limited. This paper introduces a Large Vision-Language Model(LVLM) based environment perception system whi…

Cited by 0SourceScholar
2024

A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis

ECCV 2024poster

"While conversational generative AI has shown considerable potential in enhancing decision-making for agricultural professionals, its exploration has predominantly been anchored in text-based interactions. The evolution of multimodal conversational AI, leveraging vast amounts of image-text data from…

2024

ADIFT: Zero-Shot Generative Model Adaption Via Adaptive Domain-Invariant Feature Transfer

ICASSP 2024accepted

CLIP-guided zero-shot image generative model adaption methods only require textual domain labels without any target domain images, but there are some dilemmas remain unsolved, such as identity degradation and pattern overfitting. To address these issues, an adaptive domain-invariant feature transfer…

Cited by 0SourceScholar
2024

Aligning Large Language Models with Representation Editing: A Control Perspective

NeurIPS 2024poster

Aligning large language models (LLMs) with human objectives is crucial for real-world applications. However, fine-tuning LLMs for alignment often suffers from unstable training and requires substantial computing resources. Test-time alignment techniques, such as prompting and guided decoding, do not…

2024

Automatic Captioning based on Visible and Infrared Images

ICRA 2024poster

In this paper, we tackle the task of image captioning with the complementarity of visible light images and infrared images. To address this problem, we propose an RGBIR image fusion captioning model, which can take full advantage of visible light images and infrared images under different conditions…

Cited by 1SourceScholar
2024

Can We Evaluate Domain Adaptation Models Without Target-Domain Labels?

ICLR 2024poster

Unsupervised domain adaptation (UDA) involves adapting a model trained on a label-rich source domain to an unlabeled target domain. However, in real-world scenarios, the absence of target-domain labels makes it challenging to evaluate the performance of UDA models. Furthermore, prevailing UDA method…

Cited by 13SourcePDFScholar
2024

Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion Model

NeurIPS 2024poster

Spatio-temporal (ST) prediction has garnered a De facto attention in earth sciences, such as meteorological prediction, human mobility perception. However, the scarcity of data coupled with the high expenses involved in sensor deployment results in notable data imbalances. Furthermore, models that a…

Cited by 2SourcePDFScholar
2024

ColorPeel: Color Prompt Learning with Diffusion Models via Color and Shape Disentanglement

ECCV 2024poster

"Text-to-Image (T2I) generation has made significant advancements with the advent of diffusion models. These models exhibit remarkable abilities to produce images based on textual prompts. Current T2I models allow users to specify object colors using linguistic color names. However, these labels enc…

2024

DiffAug: Enhance Unsupervised Contrastive Learning with Domain-Knowledge-Free Diffusion-based Data Augmentation

ICML 2024poster

Unsupervised Contrastive learning has gained prominence in fields such as vision, and biology, leveraging predefined positive/negative samples for representation learning. Data augmentation, categorized into hand-designed and model-based methods, has been identified as a crucial component for enhanc…

2024

Dual Level Intent-Slot Interaction for Improved Multi-Intent Spoken Language Understanding

ICASSP 2024accepted

Multi-intent spoken language understanding consists of two typical subtasks: multi-intent detection and slot filling. Existing approach suffers from two limitations: (1) It fails to explicitly model the information transfer between slots associated within the same intent clause; (2) Using a co-occur…

Cited by 0SourceScholar
2024

Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation

NeurIPS 2024poster

Existing parameter-efficient fine-tuning (PEFT) methods have achieved significant success on vision transformers (ViTs) adaptation by improving parameter efficiency. However, the exploration of enhancing inference efficiency during adaptation remains underexplored. This limits the broader applicatio…

2024

Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning

NeurIPS 2024poster

Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, exist…

2024

EnMatch: Matchmaking for Better Player Engagement via Neural Combinatorial Optimization

AAAI 2024technical

Matchmaking is a core task in e-sports and online games, as it contributes to player engagement and further influences the game's lifecycle. Previous methods focus on creating fair games at all times. They divide players into different tiers based on skill levels and only select players from the sam…

Cited by 3SourcePDFScholar
2024

Exemplar-free Continual Representation Learning via Learnable Drift Compensation

ECCV 2024poster

"Exemplar-free class-incremental learning using a backbone trained from scratch and starting from a small first task presents a significant challenge for continual representation learning. Prototype-based approaches, when continually updated, face the critical issue of semantic drift due to which th…

2024

First-Order Methods for Linearly Constrained Bilevel Optimization

NeurIPS 2024poster

Algorithms for bilevel optimization often encounter Hessian computations, which are prohibitive in high dimensions. While recent works offer first-order methods for unconstrained bilevel problems, the constrained setting remains relatively underexplored. We present first-order linearly constrained…

Cited by 18SourcePDFScholar
2024

GDeR: Safeguarding Efficiency, Balancing, and Robustness via Prototypical Graph Pruning

NeurIPS 2024poster

Training high-quality deep models necessitates vast amounts of data, resulting in overwhelming computational and memory demands. Recently, data pruning, distillation, and coreset selection have been developed to streamline data volume by \textit{retaining}, \textit{synthesizing}, or \textit{selectin…

2024

InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data Pruning

ICLR 2024oral

Data pruning aims to obtain lossless performances with less overall cost. A common approach is to filter out samples that make less contribution to the training. This could lead to gradient expectation bias compared to the original data. To solve this problem, we propose InfoBatch, a novel framework…

2024

LLM as Prompter: Low-resource Inductive Reasoning on Arbitrary Knowledge Graphs

ACL 2024findings

Knowledge Graph (KG) inductive reasoning, which aims to infer missing facts from new KGs that are not seen during training, has been widely adopted in various applications. One critical challenge of KG inductive reasoning is handling low-resource scenarios with scarcity in both textual and structura…

2024

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

NeurIPS 2024spotlight

In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance…

Cited by 269SourcePDFScholar
2024

MOMA: Mixture-of-Modality-Adaptations for Transferring Knowledge from Image Models Towards Efficient Audio-Visual Action Recognition

ICASSP 2024accepted

In this work, we investigate how to transfer learned knowledge from pre-trained image models for the audio-visual domain without relying on a full finetuning paradigm. To achieve this objective, we propose a novel parameter-efficient scheme called Mixture-of-Modality-Adaptations (MoMA) for audio-vis…

Cited by 0SourceScholar
2024

NGEL-SLAM: Neural Implicit Representation-based Global Consistent Low-Latency SLAM System

ICRA 2024poster

Neural implicit representations have emerged as a promising solution for providing dense geometry in Simultaneous Localization and Mapping (SLAM). However, existing methods in this direction fall short in terms of global consistency and low latency. This paper presents NGEL-SLAM to tackle the above…

Cited by 29SourceScholar
2024

Navigating Complexity: Toward Lossless Graph Condensation via Expanding Window Matching

ICML 2024poster

Graph condensation aims to reduce the size of a large-scale graph dataset by synthesizing a compact counterpart without sacrificing the performance of Graph Neural Networks (GNNs) trained on it, which has shed light on reducing the computational cost for training GNNs. Nevertheless, existing methods…

2024

NuwaDynamics: Discovering and Updating in Causal Spatio-Temporal Modeling

ICLR 2024spotlight

Spatio-temporal (ST) prediction plays a pivotal role in earth sciences, such as meteorological prediction, urban computing. Adequate high-quality data, coupled with deep models capable of inference, are both indispensable and prerequisite for achieving meaningful results. However, the sparsity of da…

Cited by 12SourcePDFScholar
2024

Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability, Reproducibility, and Practicality

NeurIPS 2024poster

Recent text-to-video (T2V) technology advancements, as demonstrated by models such as Gen2, Pika, and Sora, have significantly broadened its applicability and popularity. Despite these strides, evaluating these models poses substantial challenges. Primarily, due to the limitations inherent in auto…

2024

Summarizing Stream Data for Memory-Constrained Online Continual Learning

AAAI 2024technical

Replay-based methods have proved their effectiveness on online continual learning by rehearsing past samples from an auxiliary memory. With many efforts made on improving training schemes based on the memory, however, the information carried by each sample in the memory remains under-investigated. U…

2024

Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis

NeurIPS 2024poster

Although text-to-image (T2I) models exhibit remarkable generation capabilities, they frequently fail to accurately bind semantically related objects or attributes in the input prompts; a challenge termed semantic binding. Previous approaches either involve intensive fine-tuning of the entire T2I mod…

2024

Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory Matching

ICLR 2024poster

The ultimate goal of Dataset Distillation is to synthesize a small synthetic dataset such that a model trained on this synthetic set will perform equally well as a model trained on the full, real dataset. Until now, no method of Dataset Distillation has reached this completely lossless goal, in part…

2024

Two Heads Are Better Than One: Boosting Graph Sparse Training via Semantic and Topological Awareness

ICML 2024poster

Graph Neural Networks (GNNs) excel in various graph learning tasks but face computational challenges when applied to large-scale graphs. A promising solution is to remove non-essential edges to reduce the computational overheads in GNN. Previous literature generally falls into two categories: topolo…

Cited by 16SourcePDFScholar
2024

VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

EMNLP 2024main

The recent years have witnessed great advances in video generation. However, the development of automatic video metrics is lagging significantly behind. None of the existing metric is able to provide reliable scores over generated videos. The main barrier is the lack of large-scale human-annotated d…

2024

ν-DBA: Neural Implicit Dense Bundle Adjustment Enables Image-Only Driving Scene Reconstruction

IROS 2024poster

The joint optimization of the sensor trajectory and 3D map is a crucial characteristic of bundle adjustment (BA), essential for autonomous driving. This paper presents ν-DBA, a novel framework implementing geometric dense bundle adjustment (DBA) using 3D neural implicit surfaces for map parametrizat…

Cited by 0SourceScholar
2023

A Spatio-Temporal Decomposition Network for Compressed Video Quality Enhancement

ICASSP 2023accepted

Compressed video quality enhancement has always been a widely concerned research. However, existing methods rarely build models from the consideration of object motion diversity and feature frequency distribution. In this paper, we propose a Spatio-Temporal Decomposition Network (STDN) to reduce the…

Cited by 0SourceScholar
2023

BiCro: Noisy Correspondence Rectification for Multi-Modality Data via Bi-Directional Cross-Modal Similarity Consistency

CVPR 2023poster

As one of the most fundamental techniques in multimodal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal data…

2023

CORE: Co-planarity Regularized Monocular Geometry Estimation with Weak Supervision

ICCV 2023poster

The ill-posed nature of monocular 3D geometry (depth map and surface normals) estimation makes it rely mostly on data-driven approaches such as Deep Neural Networks (DNN). However, data acquisition of surface normals, especially the reliable normals, is acknowledged difficult. Commonly, reconstructi…

Cited by 0PDFScholar
2023

DREAM: Efficient Dataset Distillation by Representative Matching

ICCV 2023poster

Dataset distillation aims to synthesize small datasets with little information loss from original large-scale ones for reducing storage and training costs. Recent state-of-the-art methods mainly constrain the sample synthesis process by matching synthetic images and the original ones regarding gradi…

Cited by 91PDFcodeScholar
2023

Divide to Adapt: Mitigating Confirmation Bias for Domain Adaptation of Black-Box Predictors

ICLR 2023top-25%

Domain Adaptation of Black-box Predictors (DABP) aims to learn a model on an unlabeled target domain supervised by a black-box predictor trained on a source domain. It does not require access to both the source-domain data and the predictor parameters, thus addressing the data privacy and portabilit…

2023

Does Graph Distillation See Like Vision Dataset Counterpart?

NeurIPS 2023poster

Training on large-scale graphs has achieved remarkable results in graph representation learning, but its cost and storage have attracted increasing concerns. Existing graph condensation methods primarily focus on optimizing the feature matrices of condensed graphs while overlooking the impact of the…

Cited by 42SourcePDFScholar
2023

Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image Editing

NeurIPS 2023poster

Large-scale text-to-image generative models have been a ground-breaking development in generative AI, with diffusion models showing their astounding ability to synthesize convincing images following an input text prompt. The goal of image editing research is to give users control over the generated…

2023

Expanding Small-Scale Datasets with Guided Imagination

NeurIPS 2023poster

The power of DNNs relies heavily on the quantity and quality of training data. However, collecting and annotating data on a large scale is often expensive and time-consuming. To address this issue, we explore a new task, termed dataset expansion, aimed at expanding a ready-to-use small dataset by au…

2023

MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID

CVPR 2023poster

Neural Architecture Search (NAS) has been increasingly appealing to the society of object Re-Identification (ReID), for that task-specific architectures significantly improve the retrieval performance. Previous works explore new optimizing targets and search spaces for NAS ReID, yet they neglect the…

2023

Optimistic Whittle Index Policy: Online Learning for Restless Bandits

AAAI 2023technical

Restless multi-armed bandits (RMABs) extend multi-armed bandits to allow for stateful arms, where the state of each arm evolves restlessly with different transitions depending on whether that arm is pulled. Solving RMABs requires information on transition dynamics, which are often unknown upfront. T…

2023

PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated Learning.

NeurIPS 2023poster

Classical federated learning (FL) enables training machine learning models without sharing data for privacy preservation, but heterogeneous data characteristic degrades the performance of the localized model. Personalized FL (PFL) addresses this by synthesizing personalized models from a global mode…

2023

Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language Models

ICCV 2023poster

Continual learning (CL) can help pre-trained vision-language models efficiently adapt to new or under-trained data distributions without re-training. Nevertheless, during the continual training of the Contrastive Language-Image Pre-training (CLIP) model, we observe that the model's zero-shot transfe…

Cited by 100PDFcodeScholar
2023

Scalable Decision-Focused Learning in Restless Multi-Armed Bandits with Application to Maternal and Child Health

AAAI 2023technical

This paper studies restless multi-armed bandit (RMAB) problems with unknown arm transition dynamics but with known correlated arm features. The goal is to learn a model to predict transition dynamics given features, where the Whittle index policy solves the RMAB problems using predicted transitions.…

Cited by 30SourcePDFScholar
2023

Scenario Diffusion: Controllable Driving Scenario Generation With Diffusion

NeurIPS 2023poster

Automated creation of synthetic traffic scenarios is a key part of scaling the safety validation of autonomous vehicles (AVs). In this paper, we propose Scenario Diffusion, a novel diffusion-based architecture for generating traffic scenarios that enables controllable scenario generation. We combine…

Cited by 35SourcePDFScholar
2023

Smoothed Online Combinatorial Optimization Using Imperfect Predictions

AAAI 2023technical

Smoothed online combinatorial optimization considers a learner who repeatedly chooses a combinatorial decision to minimize an unknown changing cost function with a penalty on switching decisions in consecutive rounds. We study smoothed online combinatorial optimization problems when an imperfect pre…

Cited by 2SourcePDFScholar
2023

Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation

ICASSP 2023accepted

Existing speech separation models based on deep learning typically generalize poorly due to domain mismatch. In this paper, we propose SpeakerAugment (SA), a data augmentation method for generalizable speech separation that aims to increase the diversity of speaker identity in training data, to miti…

Cited by 0SourceScholar
2023

Specialist Diffusion: Plug-and-Play Sample-Efficient Fine-Tuning of Text-to-Image Diffusion Models To Learn Any Unseen Style

CVPR 2023poster

Diffusion models have demonstrated impressive capability of text-conditioned image synthesis, and broader application horizons are emerging by personalizing those pretrained diffusion models toward generating some specialized target object or style. In this paper, we aim to learn an unseen style by…

2023

Versatile Diffusion: Text, Images and Variations All in One Diffusion Model

ICCV 2023poster

Recent advances in diffusion models have set an impressive milestone in many generation tasks, and trending works such as DALL-E2, Imagen, and Stable Diffusion have attracted great interest. Despite the rapid landscape changes, recent new approaches focus on extensions and performance rather than ca…

Cited by 187PDFcodeScholar
2023

Zero-Shot Generative Model Adaptation via Image-Specific Prompt Learning

CVPR 2023poster

Recently, CLIP-guided image synthesis has shown appealing performance on adapting a pre-trained source-domain generator to an unseen target domain. It does not require any target-domain samples but only the textual domain labels. The training is highly efficient, e.g., a few minutes. However, existi…

2022

An Efficient Training Approach for Very Large Scale Face Recognition

CVPR 2022poster

Face recognition has achieved significant progress in deep learning era due to the ultra-large-scale and welllabeled datasets. However, training on the outsize datasets is time-consuming and takes up a lot of hardware resource. Therefore, designing an efficient training approach is indispensable. Th…

Cited by 39PDFcodeScholar
2022

Attracting and Dispersing: A Simple Approach for Source-free Domain Adaptation

NeurIPS 2022accept

We propose a simple but effective source-free domain adaptation (SFDA) method. Treating SFDA as an unsupervised clustering problem and following the intuition that local neighbors in feature space should have more similar predictions than other features, we propose to optimize an objective of predic…

2022

CAFE: Learning To Condense Dataset by Aligning Features

CVPR 2022poster

Dataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one. State-of-the-art approaches largely rely on learning the synthetic data by matching the gradients between the real and synthetic data batches. Despite the intu…

Cited by 277PDFcodeScholar
2022

Coordinating Followers to Reach Better Equilibria: End-to-End Gradient Descent for Stackelberg Games

AAAI 2022technical

A growing body of work in game theory extends the traditional Stackelberg game to settings with one leader and multiple followers who play a Nash equilibrium. Standard approaches for computing equilibria in these games reformulate the followers' best response as constraints in the leader's optimizat…

Cited by 31SourcePDFScholar
2022

Crafting Better Contrastive Views for Siamese Representation Learning

CVPR 2022oral

Recent self-supervised contrastive learning methods greatly benefit from the Siamese structure that aims at minimizing distances between positive pairs. For high performance Siamese representation learning, one of the keys is to design good contrastive pairs. Most previous works simply apply random…

Cited by 140PDFcodeScholar
2022

DLME: Deep Local-Flatness Manifold Embedding

ECCV 2022poster

"Manifold learning (ML) aims to seek low-dimensional embedding from high-dimensional data. The problem is challenging on real-world datasets, especially with under-sampling data, and we find that previous methods perform poorly in this case. Generally, ML methods first transform input data into a lo…

2022

Decision-Focused Learning without Decision-Making: Learning Locally Optimized Decision Losses

NeurIPS 2022accept

Decision-Focused Learning (DFL) is a paradigm for tailoring a predictive model to a downstream optimization task that uses its predictions in order to perform better \textit{on that specific task}. The main technical challenge associated with DFL is that it requires being able to differentiate throu…

Cited by 52SourcePDFScholar
2022

Instance-Guided Prompt Learning for Few-Shot Text Matching

EMNLP 2022finding

Few-shot text matching is a more practical technique in natural language processing (NLP) to determine whether two texts are semantically identical. They primarily design patterns to reformulate text matching into a pre-trained task with uniform prompts across all instances. But they fail to take in…

2022

MSDN: Mutually Semantic Distillation Network for Zero-Shot Learning

CVPR 2022poster

The key challenge of zero-shot learning (ZSL) is how to infer the latent semantic knowledge between visual and attribute features on seen classes, and thus achieving a desirable knowledge transfer to unseen classes. Prior works either simply align the global features of an image with its associated…

Cited by 177PDFcodeScholar
2022

Mining Hard Samples Locally And Globally For Improved Speech Separation

ICASSP 2022accepted

Speech separation dataset typically consists of hard and non-hard samples, and the former is minority and latter majority. The data imbalance problem biases the model towards non-hard samples and weakens the generalization capability. Given that the average separation performance is sufficiently goo…

Cited by 0SourceScholar
2022

Modeling Motion With Multi-Modal Features for Text-Based Video Segmentation

CVPR 2022poster

Text-based video segmentation aims to segment the target object in a video based on a describing sentence. Incorporating motion information from optical flow maps with appearance and linguistic modalities is crucial yet has been largely ignored by previous work. In this paper, we design a method to…

Cited by 28PDFcodeScholar
2022

Point-to-Box Network for Accurate Object Detection via Single Point Supervision

ECCV 2022poster

"Object detection using single point supervision has received increasing attention over the years. However, the performance gap between point supervised object detection (PSOD) and bounding box supervised detection remains large. In this paper, we attribute such a large performance gap to the failur…

2022

Robust Speaker Verification with Joint Self-Supervised and Supervised Learning

ICASSP 2022accepted

Supervised learning and self-supervised learning address different facets. Supervised learning achieves high accuracy, but it requires numerous expensive labeled data indeed. Correspondingly, self-supervised learning, makes use of abundant unlabeled data to learn, but the performance lags behind tha…

Cited by 0SourceScholar
2022

The Shape Part Slot Machine: Contact-Based Reasoning for Generating 3D Shapes from Parts

ECCV 2022poster

"We present the Shape Part Slot Machine, a new method for assembling novel 3D shapes from existing parts by performing contact-based reasoning. Our method represents each shape as a graph of ""slots,"" where each slot is a region of contact between two shape parts. Based on this representation, we d…

Cited by 12SourcePDFScholar
2021

Encoder-Decoder Based Pitch Tracking and Joint Model Training for Mandarin Tone Classification

ICASSP 2021accepted

We pursue an interpretable pitch tracking model and a jointly trained tone model for Mandarin tone classification. For pitch tracking, present deep learning based pitch model structure seldom considers the Viterbi decoding commonly implemented in prevalent manually designed pitch tracking algorithms…

Cited by 0SourceScholar
2021

Hyperbolic Geometry is Not Necessary: Lightweight Euclidean-Based Models for Low-Dimensional Knowledge Graph Embeddings

EMNLP 2021finding

Recent knowledge graph embedding (KGE) models based on hyperbolic geometry have shown great potential in a low-dimensional embedding space. However, the necessity of hyperbolic space in KGE is still questionable, because the calculation based on hyperbolic geometry is much more complicated than Eucl…

2021

Interpretable Visual Reasoning via Induced Symbolic Space

ICCV 2021poster

We study the problem of concept induction in visual reasoning, i.e., identifying concepts and their hierarchical relationships from question-answer pairs associated with images; and achieve an interpretable model via working on the induced symbolic concept space. To this end, we first design a new f…

Cited by 22PDFcodeScholar
2021

Labeling Trick: A Theory of Using Graph Neural Networks for Multi-Node Representation Learning

NeurIPS 2021poster

In this paper, we provide a theory of using graph neural networks (GNNs) for multi-node representation learning (where we are interested in learning a representation for a set of more than one node, such as link). We know that GNN is designed to learn single-node representations. When we want to lea…

2021

Learning MDPs from Features: Predict-Then-Optimize for Sequential Decision Making by Reinforcement Learning

NeurIPS 2021spotlight

In the predict-then-optimize framework, the objective is to train a predictive model, mapping from environment features to parameters of an optimization problem, which maximizes decision quality when the optimization is subsequently solved. Recent work on decision-focused learning shows that embeddi…

Cited by 38SourcePDFScholar
2021

Neighborhood Intervention Consistency: Measuring Confidence for Knowledge Graph Link Prediction

IJCAI 2021poster

Link prediction based on knowledge graph embeddings (KGE) has recently drawn a considerable momentum. However, existing KGE models suffer from insufficient accuracy and hardly evaluate the confidence probability of each predicted triple. To fill this critical gap, we propose a novel confidence measu…

Cited by 11SourcePDFScholar
2021

Reinforcement Learning with a Disentangled Universal Value Function for Item Recommendation

AAAI 2021technical

In recent years, there are great interests as well as many challenges in applying reinforcement learning (RL) to recommendation systems (RS). In this paper, we summarize three key practical challenges of large-scale RL-based recommender systems: massive state and action spaces, high-variance environ…

2021

TSTNN: Two-Stage Transformer Based Neural Network for Speech Enhancement in the Time Domain

ICASSP 2021accepted

In this paper, we propose a transformer-based architecture, called two-stage transformer neural network (TSTNN) for end-to-end speech denoising in the time domain. The proposed model is composed of an encoder, a two-stage transformer module (TSTM), a masking module and a decoder. The encoder maps in…

Cited by 0SourceScholar
2020

Automatically Learning Compact Quality-aware Surrogates for Optimization Problems

NeurIPS 2020spotlight

Solving optimization problems with unknown parameters often requires learning a predictive model to predict the values of the unknown parameters and then solving the problem using these values. Recent work has shown that including the optimization problem as a layer in the model training pipeline re…

2020

Robust Spatial-Temporal Incident Prediction

UAI 2020poster

Spatio-temporal incident prediction is a central issue in law enforcement, with applications in fighting crimes like poaching, human trafficking, illegal fishing, burglaries and smuggling. However, state of the art approaches fail to account for evasion in response to predictive models, a common fo…

Cited by 6SourcePDFScholar
2020

Semantic Drift Compensation for Class-Incremental Learning

CVPR 2020poster

Class-incremental learning of deep networks sequentially increases the number of classes to be classified. During training, the network has only access to data of one task at a time, where each task contains several classes. In this setting, networks suffer from catastrophic forgetting which refers…

Cited by 413PDFcodeScholar
2020

Suppressing Mislabeled Data via Grouping and Self-Attention

ECCV 2020poster

Deep networks achieve excellent results on large-scale clean data but degrade significantly when learning from noisy labels. To suppressing the impact of mislabeled data, this paper proposes a conceptually simple yet efficient training block, termed as Attentive Feature Mixup (AFM), which allows pay…

2020

Suppressing Uncertainties for Large-Scale Facial Expression Recognition

CVPR 2020poster

Annotating a qualitative large-scale facial expression dataset is extremely difficult due to the uncertainties caused by ambiguous facial expressions, low-quality facial images, and the subjectiveness of annotators. These uncertainties suspend the progress of large-scale Facial Expression Recognitio…

Cited by 783PDFcodeScholar
2019

A Robust Local Spectral Descriptor for Matching Non-Rigid Shapes With Incompatible Shape Structures

CVPR 2019poster

Constructing a robust and discriminative local descriptor for 3D shape is a key component of many computer vision applications. Although existing learning-based approaches can achieve good performance in some specific benchmarks, they usually fail to learn enough information from shapes with differe…

Cited by 25PDFScholar
2019

A Unified Framework for Mutual Improvement of SLAM and Semantic Segmentation

ICRA 2019poster

This paper presents a novel framework for simultaneously implementing localization and segmentation, which are two of the most important vision-based tasks for robotics. While the goals and techniques used for them were considered to be different previously, we show that by making use of the interme…

Cited by 44SourceScholar
2019

Towards More Realistic Human-Robot Conversation: A Seq2Seq-based Body Gesture Interaction System

IROS 2019poster

This paper presents a novel system that enables intelligent robots to exhibit realistic body gestures while communicating with humans. The proposed system consists of a listening model and a speaking model used in corresponding conversational phases. Both models are adapted from the sequence-to-sequ…

Cited by 13SourceScholar
2018

Sub-GAN: An Unsupervised Generative Model via Subspaces

ECCV 2018poster

The recent years have witnessed significant growth in constructing robust generative models to capture informative distributions of natural data. However, it is difficult to fully exploit the distribution of complex data, like images and videos, due to the high dimensionality of ambient space. Seque…

Cited by 24SourcePDFScholar
2018

Widely Linear CLMS Based Cancelation of Nonlinear Self -Interference in Full-Duplex Direct-Conversion Transceivers

ICASSP 2018accepted

An augmented nonlinear complex LMS (ANCLMS) algorithm is proposed to adaptively mitigate both the linear and nonlinear self-interference (SI) components in a full-duplex direct-conversion transceiver (DCT). A data prewhitening scheme, which exploits the known SI signal distributions, is also adopted…

Cited by 0SourceScholar