← Search

Xin Yuan

84 accepted papers

2026

3One2: One-Step Regression plus One-Step Diffusion for One-Hot Modulation in Dual-Path Video Snapshot Compressive Imaging

AAAI 2026technical

Video snapshot compressive imaging (SCI) captures dynamic scene sequences through a two-dimensional (2D) snapshot, fundamentally relying on optical modulation for hardware compression and the corresponding software reconstruction. While mainstream video SCI using random binary modulation has demonst

Cited by 0SourcePDFScholar
2026

BOLT: Decision‑Aligned Distillation and Budget-Aware Routing for Constrained Multimodal QA on Robots

ICLR 2026poster

Robotic systems can require multimodal reasoning under stringent constraints of latency, memory, and energy. Standard instruction tuning and token-level distillation fail to deliver decision quality, reliability, and interpretability under these constraints. We introduce BOLT, a decision-aligned dis…

Cited by 0SourceScholar
2026

Breaking Measurement Barriers: From Compressed Sensing to Deep Reconstruction

AAAI 2026technical

Deep learning methods have achieved remarkable success in image compressed sensing (CS) task, namely reconstructing a high-fidelity image from its compressed measurement. However, existing methods are deficient in incoherent compressed measurement at sensing phase and implicit measurement representa

Cited by 0SourcePDFScholar
2026

Budget-Efficient Attacks and Robustness Training for Cooperative MARL

ICML 2026poster

Cooperative multi-agent reinforcement learning (CMARL) policies are vulnerable to action hijacking even when only a few timesteps are compromised. Recent adversarial attacks and adversarial training methods have been explored, but under an explicit attack budget, existing attacks often fail to accur…

Cited by 0SourceScholar
2026

DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive Imaging

CVPR 2026

Video snapshot compressive imaging (SCI) offers a promising alternative to high-speed cameras by encoding multiple frames into a single 2D measurement. However, SCI requires algorithms to reconstruct the high-speed video, and as resolution increases, reconstruction becomes computationally expensive

Cited by 0SourceScholar
2026

Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment

ICLR 2026poster

The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based IQA methods primarily rely on general contextual descriptions, sometimes limitin…

Cited by 0SourcecodeScholar
2026

High-Speed FHD Full-Color Video Computer-Generated Holography

AAAI 2026technical

Computer-generated holography (CGH) is a promising technology for next-generation displays. However, generating high-speed, high-quality holographic video requires both high frame rate display and efficient computation, but is constrained by two key limitations: (i) Learning-based models often produ

Cited by 0SourcePDFScholar
2026

Joint Spectral Image Reconstruction and Semantic Segmentation with Cooperative Unfolding

CVPR 2026

Coded Aperture Snapshot Spectral Imaging (CASSI) is an emerging hyperspectral image (HSI) acquisition technique for downstream semantic segmentation. Due to the ill-posedness nature of CASSI systems, typical solutions are compelled to conduct a two-stage reconstruction-then-segmentation pipeline, na

Cited by 0SourcecodeScholar
2026

MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution

AAAI 2026technical

Chinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hindering archival efforts. Although space-time video super-resolution (STVSR) has adva

Cited by 0SourcePDFScholar
2026

Phy-CoSF: Physics-Guided Continuous Spectral Fields Reconstruction and Spectral Super-Resolution for Snapshot Compressive Imaging

ICML 2026poster

Recent advances have demonstrated that coded aperture snapshot spectral imaging (CASSI) systems show great potential for capturing 3D hyperspectral images (HSIs) from a single 2D measurement. Despite the inherent spectral continuity of scenes captured by CASSI, most existing reconstruction methods a…

Cited by 0SourceScholar
2026

Realism Control One-step Diffusion for Real-world Image Super Resolution

AAAI 2026technical

Pre-trained diffusion models have shown great potential in real-world image super-resolution (Real-ISR) tasks by enabling high-resolution reconstructions. While one-step diffusion (OSD) methods significantly improve efficiency compared to traditional multi-step approaches, they still have limitation

Cited by 0SourcePDFScholar
2026

SLCFormer: Spectral-Local Context Transformer with Physics-Grounded Flare Synthesis for Nighttime Flare Removal

AAAI 2026technical

Lens flare is a common nighttime artifact caused by strong light sources scattering within camera lenses, leading to hazy streaks, halos, and glare that degrade visual quality. However, existing methods usually fail to effectively address nonuniform scattered flares, which severely reduces their app

Cited by 0SourcePDFScholar
2026

TI-3DGS: 3D Thermal Reconstruction Via Thermal Imaging-Guided 3D Gaussian Splatting

ICRA 2026poster

Thermal imaging, with its all-weather capabilities and strong penetration, enables 3D reconstruction in low- light and adverse conditions. In this paper, we investigate RGB-independent pure 3D thermal reconstruction, aiming to overcome the challenges of 3D reconstruction in extreme environments wher…

Cited by 0Scholar
2025

DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution

NeurIPS 2025poster

Diffusion models have demonstrated promising performance in real-world video super-resolution (VSR). However, the dozens of sampling steps they require, make inference extremely slow. Sampling acceleration techniques, particularly single-step, provide a potential solution. Nonetheless, achieving one…

Cited by 0SourcecodeScholar
2025

Detail Matters: Mamba-Inspired Joint Unfolding Network for Snapshot Spectral Compressive Imaging

AAAI 2025technical

In the coded aperture snapshot spectral imaging system, Deep Unfolding Networks (DUNs) have made impressive progress in recovering 3D hyperspectral images (HSIs) from a single 2D measurement. However, the inherent nonlinear and ill-posed characteristics of HSI reconstruction still pose challenges t…

2025

Dual-branch Graph Feature Learning for NLOS Imaging

AAAI 2025technical

The domain of non-line-of-sight (NLOS) imaging is advancing rapidly, offering the capability to reveal occluded scenes that are not directly visible. However, contemporary NLOS systems face several significant challenges: (1) The computational and storage requirements are profound due to the inheren…

Cited by 0SourcePDFScholar
2025

Event-based Video Person Re-identification via Cross-Modality and Temporal Collaboration

ICASSP 2025accepted

Video-based person re-identification (ReID) has become increasingly important due to its applications in video surveillance applications. By employing events in video-based person ReID, more motion information can be provided between continuous frames to improve recognition accuracy. Previous approa…

Cited by 5SourceScholar
2025

HydraRAG: Structured Cross-Source Enhanced Large Language Model Reasoning

EMNLP 2025

Retrieval-augmented generation (RAG) enhances large language models (LLMs) by incorporating external knowledge. Current hybrid RAG system retrieves evidence from both knowledge graphs (KGs) and text documents to support LLM reasoning. However, it faces challenges like handling multi-hop reasoning, m

2025

KaRF: Weakly-Supervised Kolmogorov-Arnold Networks-based Radiance Fields for Local Color Editing

NeurIPS 2025poster

Recent advancements have suggested that neural radiance fields (NeRFs) show great potential in color editing within the 3D domain. However, most existing NeRF-based editing methods continue to face significant challenges in local region editing, which usually lead to imprecise local object boundarie…

Cited by 0SourcecodeScholar
2025

Prior-guided Hierarchical Harmonization Network for Efficient Image Dehazing

AAAI 2025technical

Image dehazing is a crucial task that involves the enhancement of degraded images to recover their sharpness and textures. While vision Transformers have exhibited impressive results in diverse dehazing tasks, their quadratic complexity and lack of dehazing priors pose significant drawbacks for real…

Cited by 0SourcePDFScholar
2025

Proximal Algorithm Unrolling: Flexible and Efficient Reconstruction Networks for Single-Pixel Imaging

CVPR 2025poster

Deep-unrolling and plug-and-play (PnP) approaches have become the de-facto standard solvers for single-pixel imaging (SPI) inverse problem. PnP approaches, a class of iterative algorithms where regularization is implicitly performed by an off-the-shelf deep denoiser, are flexible for varying compres…

2025

SCI-Gaussian: Optimizing 3D Gaussian Radiance Fields from a Snapshot Compressive Image

ICASSP 2025accepted

Snapshot compressive imaging (SCI) is a compressed sensing (CS)-based high-speed imaging modality. Recent efforts have explored the underlying 3D representation from only an SCI image using neural radiance fields (NeRF), yet the training time, rendering computation cost, and reconstruction quality l…

Cited by 0SourceScholar
2025

Spectral Compressive Imaging via Chromaticity-Intensity Decomposition

NeurIPS 2025poster

In coded aperture snapshot spectral imaging (CASSI), the captured measurement entangles spatial and spectral information, posing a severely ill-posed inverse problem for hyperspectral images (HSIs) reconstruction. Moreover, the captured radiance inherently depends on scene illumination, making it di…

Cited by 0SourcecodeScholar
2025

Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator

NeurIPS 2025poster

Diffusion models have demonstrated excellent performance for real-world image super-resolution (Real-ISR), albeit at high computational costs. Most existing methods are trying to derive one-step diffusion models from multi-step counterparts through knowledge distillation (KD) or variational score di…

Cited by 0SourceScholar
2025

pFedRAG: A Personalized Federated Retrieval-Augmented Generation System with Depth-Adaptive Tiered Embedding Tuning

EMNLP 2025

Large Language Models (LLMs) can undergo hallucinations in specialized domains, and standard Retrieval-Augmented Generation (RAG) often falters due to general-purpose embeddings ill-suited for domain-specific terminology. Though domain-specific fine-tuning enhances retrieval, centralizing data intro

Cited by 0SourcePDFScholar
2024

2DQuant: Low-bit Post-Training Quantization for Image Super-Resolution

NeurIPS 2024poster

Low-bit quantization has become widespread for compressing image super-resolution (SR) models for edge deployment, which allows advanced SR models to enjoy compact low-bit parameters and efficient integer/bitwise constructions for storage compression and inference acceleration, respectively. However…

2024

A Simple Low-bit Quantization Framework for Video Snapshot Compressive Imaging

ECCV 2024oral

"Video Snapshot Compressive Imaging (SCI) aims to use a low-speed 2D camera to capture high-speed scene as snapshot compressed measurements, followed by a reconstruction algorithm to reconstruct the high-speed video frames. State-of-the-art (SOTA) deep learning-based algorithms have achieved impress…

2024

Binarized Diffusion Model for Image Super-Resolution

NeurIPS 2024poster

Advanced diffusion models (DMs) perform impressively in image super-resolution (SR), but the high memory and computational costs hinder their deployment. Binarization, an ultra-compression algorithm, offers the potential for effectively accelerating DMs. Nonetheless, due to the model structure and t…

2024

Cooperative Hardware-Prompt Learning for Snapshot Compressive Imaging

NeurIPS 2024poster

Existing reconstruction models in snapshot compressive imaging systems (SCI) are trained with a single well-calibrated hardware instance, making their perfor- mance vulnerable to hardware shifts and limited in adapting to multiple hardware configurations. To facilitate cross-hardware learning, previ…

2024

Latent Diffusion Prior Enhanced Deep Unfolding for Snapshot Spectral Compressive Imaging

ECCV 2024oral

"Snapshot compressive spectral imaging reconstruction aims to reconstruct three-dimensional spatial-spectral images from a single-shot two-dimensional compressed measurement. Existing state-of-the-art methods are mostly based on deep unfolding structures but have intrinsic performance bottlenecks: i…

2024

SCINeRF: Neural Radiance Fields from a Snapshot Compressive Image

CVPR 2024highlight

In this paper we explore the potential of Snapshot Com- pressive Imaging (SCI) technique for recovering the under- lying 3D scene representation from a single temporal com- pressed image. SCI is a cost-effective method that enables the recording of high-dimensional data such as hyperspec- tral or te…

2024

Untrained Neural Nets for Snapshot Compressive Imaging: Theory and Algorithms

NeurIPS 2024poster

Snapshot compressive imaging (SCI) recovers high-dimensional (3D) data cubes from a single 2D measurement, enabling diverse applications like video and hyperspectral imaging to go beyond standard techniques in terms of acquisition speed and efficiency. In this paper, we focus on SCI recovery algorit…

2023

Accelerated Training via Incrementally Growing Neural Networks using Variance Transfer and Learning Rate Adaptation

NeurIPS 2023poster

We develop an approach to efficiently grow neural networks, within which parameterization and optimization strategies are designed by considering their effects on the training dynamics. Unlike existing growing methods, which follow simple replication heuristics or utilize auxiliary gradient-based l…

Cited by 7SourcePDFScholar
2023

Accurate Image Restoration with Attention Retractable Transformer

ICLR 2023top-25%

Recently, Transformer-based image restoration networks have achieved promising improvements over convolutional neural networks due to parameter-independent global interactions. To lower computational cost, existing works generally limit self-attention computation within non-overlapping windows. Howe…

2023

Binarized Spectral Compressive Imaging

NeurIPS 2023poster

Existing deep learning models for hyperspectral image (HSI) reconstruction achieve good performance but require powerful hardwares with enormous memory and computational resources. Consequently, these methods can hardly be deployed on resource-limited mobile devices. In this paper, we propose a nove…

2023

EfficientSCI: Densely Connected Network With Space-Time Factorization for Large-Scale Video Snapshot Compressive Imaging

CVPR 2023poster

Video snapshot compressive imaging (SCI) uses a two-dimensional detector to capture consecutive video frames during a single exposure time. Following this, an efficient reconstruction algorithm needs to be designed to reconstruct the desired video frames. Although recent deep learning-based state-of…

2023

Hierarchical Integration Diffusion Model for Realistic Image Deblurring

NeurIPS 2023spotlight

Diffusion models (DMs) have recently been introduced in image deblurring and exhibited promising performance, particularly in terms of details reconstruction. However, the diffusion model requires a large number of inference iterations to recover the clean image from pure Gaussian noise, which consu…

2023

Hyperspectral Image Denoising Via Nonlocal Rank Residual Modeling

ICASSP 2023accepted

Nonlocal low-rank (LR) tensor modeling has shown great potential in hyperspectral image (HSI) denoising, which first uses the nonlocal self-similarity (NSS) prior to search for many similar full-band patches to form three-dimensional nonlocal full-band groups (tensors), and then usually enforces an…

Cited by 0SourceScholar
2023

Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation

EMNLP 2023long main

Mis- and disinformation online have become a major societal problem as major sources of online harms of different kinds. One common form of mis- and disinformation is out-of-context (OOC) information, where different pieces of information are falsely associated, e.g., a real image combined with a fa…

Cited by 0SourcecodeScholar
2023

Unfolding Framework with Prior of Convolution-Transformer Mixture and Uncertainty Estimation for Video Snapshot Compressive Imaging

ICCV 2023poster

We consider the problem of video snapshot compressive imaging (SCI), where sequential high-speed frames are modulated by different masks and captured by a single measurement. The underlying principle of reconstructing multi-frame images from only one single measurement is to solve an ill-posed probl…

Cited by 7PDFcodeScholar
2022

Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction

ECCV 2022poster

"Many learning-based algorithms have been developed to solve the inverse problem of coded aperture snapshot spectral imaging (CASSI). However, CNN-based methods show limitations in capturing long-range dependencies. Previous Transformer-based methods densely sample tokens, some of which are uninform…

2022

Cross Aggregation Transformer for Image Restoration

NeurIPS 2022accept

Recently, Transformer architecture has been introduced into image restoration to replace convolution neural network (CNN) with surprising results. Considering the high computational complexity of Transformer with global attention, some methods use the local square window to limit the scope of self-a…

2022

Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive Imaging

NeurIPS 2022accept

In coded aperture snapshot spectral compressive imaging (CASSI) systems, hyperspectral image (HSI) reconstruction methods are employed to recover the spatial-spectral signal from a compressed measurement. Among these algorithms, deep unfolding methods demonstrate promising performance but suffer fro…

2022

Ensemble Learning Priors Driven Deep Unfolding for Scalable Video Snapshot Compressive Imaging

ECCV 2022poster

"Snapshot compressive imaging (SCI) can record the 3D datacube by a 2D measurement and from this 2D measurement to reconstruct the desired 3D information by algorithms. The reconstruction algorithm thus plays a vital role in SCI. Recently, deep learning (DL) has demonstrated outstanding performance…

Cited by 34SourcePDFScholar
2022

HDNet: High-Resolution Dual-Domain Learning for Spectral Compressive Imaging

CVPR 2022poster

The rapid development of deep learning provides a better solution for the end-to-end reconstruction of hyperspectral image (HSI). However, existing learning-based methods have two major defects. Firstly, networks with self-attention usually sacrifice internal resolution to balance model performance…

Cited by 188PDFcodeScholar
2022

Mask-Guided Spectral-Wise Transformer for Efficient Hyperspectral Image Reconstruction

CVPR 2022poster

Hyperspectral image (HSI) reconstruction aims to recover the 3D spatial-spectral signal from a 2D measurement in the coded aperture snapshot spectral imaging (CASSI) system. The HSI representations are highly similar and correlated across the spectral dimension. Modeling the inter-spectra interactio…

Cited by 333PDFcodeScholar
2022

Modeling Mask Uncertainty in Hyperspectral Image Reconstruction

ECCV 2022poster

"Recently, hyperspectral imaging (HSI) has attracted increasing research attention, especially for the ones based on a coded aperture snapshot spectral imaging (CASSI) system. Existing deep HSI reconstruction models are generally trained on paired data to retrieve original signals upon 2D compressed…

2022

Not All Bits have Equal Value: Heterogeneous Precisions via Trainable Noise

NeurIPS 2022accept

We study the problem of training deep networks while quantizing parameters and activations into low-precision numeric representations, a setting central to reducing energy consumption and inference time of deployed models. We propose a method that learns different precisions, as measured by bits in…

Cited by 8SourcePDFScholar
2022

Simultaneous Nonlocal Low-Rank And Deep Priors For Poisson Denoising

ICASSP 2022accepted

Poisson noise is a common electronic noise, which has widely occurred in various photo-limited imaging systems. However, due to signal-dependent and multiplicative characteristics for Poisson noise, Poisson denoising is still an open problem. In this paper, we propose a novel approach using simultan…

Cited by 0SourceScholar
2021

Deep Gaussian Scale Mixture Prior for Spectral Compressive Imaging

CVPR 2021poster

In coded aperture snapshot spectral imaging (CASSI) system, the real-world hyperspectral image (HSI) can be reconstructed from the captured compressive image in a snapshot. Model-based HSI reconstruction methods employed hand-crafted priors to solve the reconstruction problem, but most of which achi…

Cited by 185PDFScholar
2021

Dian: Duration Informed Auto-Regressive Network for Voice Cloning

ICASSP 2021accepted

In this paper, we propose a novel end-to-end speech synthesis approach, Duration Informed Auto-regressive Network (DIAN), which consists of an acoustic model and a separate duration model. Un-like other auto-regressive TTS methods, the duration information of phonemes is provided as part of the inpu…

Cited by 0SourceScholar
2021

Growing Efficient Deep Networks by Structured Continuous Sparsification

ICLR 2021oral

We develop an approach to growing deep network architectures over the course of training, driven by a principled combination of accuracy and sparsity objectives. Unlike existing pruning or architecture search techniques that operate on full-sized models or supernet architectures, our method can sta…

Cited by 69SourcePDFScholar
2021

Memory-Efficient Network for Large-Scale Video Compressive Sensing

CVPR 2021poster

Video snapshot compressive imaging (SCI) captures a sequence of video frames in a single shot using a 2D detector. The underlying principle is that during one exposure time, different masks are imposed on the high-speed scene to form a compressed measurement. With the knowledge of masks, optimizatio…

Cited by 93PDFcodeScholar
2021

MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive Sensing

CVPR 2021poster

To capture high-speed videos using a two-dimensional detector, video snapshot compressive imaging (SCI) is a promising system, where the video frames are coded by different masks and then compressed to a snapshot measurement. Following this, efficient algorithms are desired to reconstruct the high-s…

Cited by 68PDFcodeScholar
2021

Multimodal Contrastive Training for Visual Representation Learning

CVPR 2021poster

We develop an approach to learning visual representations that embraces multimodal data, driven by a combination of intra- and inter-modal similarity preservation objectives. Unlike existing visual pre-training methods, which solve a proxy prediction task in a single domain, our method exploits intr…

Cited by 215PDFcodeScholar
2021

Self-Supervised Neural Networks for Spectral Snapshot Compressive Imaging

ICCV 2021poster

We consider using untrained neural networks to solve the reconstruction problem of snapshot compressive imaging (SCI), which uses a two-dimensional (2D) detector to capture a high-dimensional (usually 3D) data-cube in a compressed manner. Various SCI systems have been built in recent years to captur…

Cited by 126PDFcodeScholar
2021

Universal and Flexible Optical Aberration Correction Using Deep-Prior Based Deconvolution

ICCV 2021poster

High quality imaging usually requires bulky and expensive lenses to compensate geometric and chromatic aberrations. This poses high constraints on the optical hash or low cost applications. Although one can utilize algorithmic reconstruction to remove the artifacts of low-end lenses, the degeneratio…

Cited by 30PDFcodeScholar
2020

BIRNAT: Bidirectional Recurrent Neural Networks with Adversarial Training for Video Snapshot Compressive Imaging

ECCV 2020poster

We consider the problem of video snapshot compressive imaging (SCI), where multiple high-speed frames are coded by different masks and then summed to a single measurement. This measurement and the modulation masks are fed into our Recurrent Neural Network (RNN) to reconstruct the desired high-speed…

2020

End-to-End Low Cost Compressive Spectral Imaging with Spatial-Spectral Self-Attention

ECCV 2020poster

Coded aperture snapshot spectral imaging (CASSI) is an effective tool to capture real-world 3D hyperspectral images. While a number of existing work has been conducted for hardware and algorithm design, we make a step towards the low-cost solution that enjoys video-rate high-quality reconstruction.…

2019

Online Hyper-Parameter Learning for Auto-Augmentation Strategy

ICCV 2019poster

Data augmentation is critical to the success of modern deep learning techniques. In this paper, we propose Online Hyper-parameter Learning for Auto-Augmentation (OHL-Auto-Aug), an economical solution that learns the augmentation policy distribution along with network training. Unlike previous method…

Cited by 109PDFScholar
2019

l-Net: Reconstruct Hyperspectral Images From a Snapshot Measurement

ICCV 2019poster

We propose the l-net, which reconstructs hyperspectral images (e.g., with 24 spectral channels) from a single shot measurement. This task is usually termed snapshot compressive-spectral imaging (SCI), which enjoys low cost, low bandwidth and high-speed sensing rate via capturing the three-dimensiona…

Cited by 283PDFcodeScholar
2018

Deep Reinforcement Learning with Iterative Shift for Visual Tracking

ECCV 2018poster

Visual tracking is confronted by the dilemma to locate a target both}accurately and efficiently, and make decisions online whether and how to adapt the appearance model or even restart tracking. In this paper, we propose a deep reinforcement learning with iterative shift (DRL-IS) method for single o…

Cited by 79SourcePDFScholar
2018

Group Sparsity Residual with Non-Local Samples for Image Denoising

ICASSP 2018accepted

Inspired by group-based sparse coding, recently proposed group sparsity residual (GSR) scheme has demonstrated superior performance in image processing. However, one challenge in GSR is to estimate the residual by using a proper reference of the group-based sparse coding (GSC), which is desired to b…

Cited by 0SourceScholar
2016

A general framework for reconstruction and classification from compressive measurements with side information

ICASSP 2016accepted

We develop a general framework for compressive linear-projection measurements with side information. Side information is an additional signal correlated with the signal of interest. We investigate the impact of side information on classification and signal recovery from low-dimensional measurements.…

Cited by 0SourceScholar
2016

A new array geometry for DOA estimation with enhanced degrees of freedom

ICASSP 2016accepted

This work presents a new array geometry, which is capable of providing O(M2N2) degrees of freedom (DOF) using only MN physical sensors via utilizing the second-order statistics of the received data. This new array is composed of multiple, identical minimum redundancy subarrays, whose positions follo…

Cited by 0SourceScholar
2016

Variational Autoencoder for Deep Learning of Images, Labels and Captions

NeurIPS 2016poster

A novel variational autoencoder is developed to model images, as well as associated labels or captions. The Deep Generative Deconvolutional Network (DGDN) is used as a decoder of the latent image features, and a deep Convolutional Neural Network (CNN) is used as an image encoder; the CNN is used to…

Cited by 1096SourcePDFScholar
2015

Non-Gaussian Discriminative Factor Models via the Max-Margin Rank-Likelihood

ICML 2015poster

We consider the problem of discriminative factor analysis for data that are in general non-Gaussian. A Bayesian model based on the ranks of the data is proposed. We first introduce a max-margin version of the rank-likelihood. A discriminative factor model is then developed, integrating the new max-m…

Cited by 8SourcePDFScholar
2015

Polynomial-phase signal direction-finding and source-tracking with a single acoustic vector sensor

ICASSP 2015accepted

This paper introduces a new ESPRIT-based algorithm to estimate the direction-of-arrival of an arbitrary degree polynomial-phase signal with a single acoustic vector-sensor. The proposed time-invariant ESPRIT algorithm is based on a matrix-pencil pair derived from the time-delayed data-sets collected…

Cited by 0SourceScholar