← Search

Fan Zhang

128 accepted papers

2026

BEYOND VISUAL REALISM: TOWARD RELIABLE FINANCIAL TIME SERIES GENERATION

ICASSP 2026poster

Generative models for financial time series often create data that look realistic and even reproduce stylized facts such as fat tails or volatility clustering. However, these apparent successes break down under trading backtests: models like GANs or WGAN-GP frequently collapse, yielding extreme and…

Cited by 0SourcePDFScholar
2026

Being More Lightweight and Practical: Mini-sized Contrastive Learning Pre-trained Models for Fine-grained Traffic Task

ICML 2026poster

Fine-grained traffic prediction is critically important for mitigating traffic congestion in key urban areas and for providing lane-change guidance in autonomous vehicles and navigation systems. However, task-specific models are not efficient enough, city-scale pre-trained models often overlook fine…

Cited by 0SourceScholar
2026

Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training

ICLR 2026poster

Zero-Shot Super-Resolution Spatiotemporal Forecasting requires a deep learning model to be trained on low-resolution data and deployed for inference on high-resolution. Existing studies consider **maintaining** similar error across different resolutions as indicative of successful multi-resolution g…

Cited by 0SourceScholar
2026

CASE-Net: Deep Spatio-Temporal Representation Learning via Causal Attention and Channel Recalibration for Multivariate Time Series Classification

IJCAI 2026

Multivariate time series (MTS) classification is foundational to pervasive computing and financial analysis, yet existing multi-scale paradigms are often constrained by suboptimal representation fidelity. We identify two critical bottlenecks: temporal non-causality in standard encoders that induces

Cited by 0Scholar
2026

CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information Decoding

AAAI 2026technical

Medical Visual Question Answering (Med-VQA) aims to generate accurate answers for clinical questions grounded in medical images, which has attracted increasing research attention due to its potential to streamline diagnostics and reduce clinical burden. Recent advances in Large Vision-Language Model

Cited by 0SourcePDFScholar
2026

Constructing Industrial-Scale Optimization Modeling Benchmark

ICML 2026poster

Optimization modeling underpins decision-making in logistics, manufacturing, energy, and finance, yet translating natural-language requirements into correct optimization formulations and solver-executable code remains labor-intensive. Although large language models (LLMs) have been explored for this…

Cited by 0SourceScholar
2026

Continuous Gaussian Process Pre-Optimization for Asynchronous Event-Inertial Odometry

RA-L 2026

Event cameras, as bio-inspired sensors, are asynchronously triggered with high-temporal resolution compared to intensity cameras. Recent work has focused on fusing the event measurements with inertial measurements to enable ego-motion estimation in high-speed and HDR environments. However, existing

Cited by 5SourcecodeScholar
2026

Continuous Gaussian Process Pre-Optimization for Asynchronous Event-Inertial Odometry

ICRA 2026poster

Event cameras, as bio-inspired sensors, are asynchronously triggered with high-temporal resolution compared to intensity cameras. Recent work has focused on fusing the event measurements with inertial measurements to enable ego-motion estimation in high-speed and HDR environments. However, existing …

2026

Decoding with Structured Awareness: Integrating Directional, Frequency-Spatial, and Structural Attention for Medical Image Segmentation

AAAI 2026technical

To address the limitations of Transformer decoders in capturing edge details, recognizing local textures and modeling spatial continuity, this paper proposes a novel decoder framework specifically designed for medical image segmentation, comprising three core modules. First, the Adaptive Cross-Fusio

Cited by 0SourcePDFScholar
2026

DiMA: Distinguishing Resident and Tourist Preferences via Multi-Modal LLM Alignment for Out-of-Town Cross-Domain Recommendation

AAAI 2026technical

Out-of-Town (OOT) recommendation aims to provide personalized suggestions for users in unfamiliar cities. However, OOT recommendation faces two fundamental challenges: the difficulty of reasoning across modalities, as preference signals in disparate formats such as images and text are hard to compar

Cited by 0SourcePDFScholar
2026

EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?

ICLR 2026poster

Descriptive Multimodal Emotion Recognition (DMER) has garnered increasing research attention. Unlike traditional discriminative paradigms that rely on predefined emotion taxonomies, DMER aims to describe human emotional state using free-form natural language, enabling finer-grained and more interpre…

Cited by 0SourcecodeScholar
2026

GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning

ICRA 2026poster

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing methods often rely on multi-stage pipelines that separate obje…

2026

Gloria: Consistent Character Video Generation via Content Anchors

CVPR 2026

Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to preserve identity or leverage non-character-centric information

Cited by 0SourceScholar
2026

ICDiffAD: Implicit Conditioning Diffusion Model for Time Series Anomaly Detection

ICLR 2026poster

Time series anomaly detection (TSAD) faces critical challenges from intrinsic data noisiness and temporal heterogeneity, which undermine the reconstruction fidelity of prevailing generative approaches. While diffusion models offer theoretical advantages in capturing complex temporal dynamics, their…

Cited by 0SourceScholar
2026

IdealTSF: Can Non-Ideal Data Contribute to Enhancing the Performance of Time Series Forecasting Models?

AAAI 2026technical

Deep learning has shown strong performance in time series forecasting tasks. However, issues such as missing values and anomalies in sequential data hinder its further development in prediction tasks. Previous research has primarily focused on extracting feature information from sequence data or add

Cited by 0SourcePDFScholar
2026

InstantRetouch: Efficient and High-Fidelity Instruction-Guided Image Retouching with Bilateral Space

CVPR 2026

Language-guided photo retouching aims to adjust color and tone while preserving geometry and texture. Recently, diffusion-based retouching shows a superior visual quality, but often struggles with both fidelity issues due to its generative nature and efficiency because of its iterative sampling proc

Cited by 0SourceScholar
2026

LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World

CVPR 2026

Tracking 3D human motion from egocentric, multi-camera devices is challenged by severe egomotion and partial visibility or occlusions. Existing methods are designed for monocular video often recorded from static or slowly-moving cameras and cannot easily leverage multi-view, calibrated and localized

Cited by 0SourcecodeScholar
2026

Learning to Memorize with Attributive and Associative Memory for Online Test-Time Adaptation of Vision-Language Models

ICML 2026poster

Memory-based test-time adaptation (TTA) assigns streaming test samples into class-specific memory slots based on pseudo-labels predicted by models like CLIP, and retrieves them to facilitate subsequent predictions under distribution shift. However, this process introduces two challenges: ❶ **Each sa…

Cited by 0SourceScholar
2026

MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human–Robot Interaction

ICRA 2026poster

We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human–robot group interactions. Effective collaboration in such settings requires consistent situational awareness, built on persistent representations of people and objects and an episodic abstraction o…

2026

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a)…

Cited by 0SourcecodeScholar
2026

Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step Entropy

ICLR 2026poster

Large Language Models (LLMs) using Chain-of-Thought (CoT) prompting excel at complex reasoning but generate verbose thought processes with considerable redundancy, leading to increased inference costs and reduced efficiency. We introduce a novel CoT compression framework based on step entropy, a met…

Cited by 0SourcecodeScholar
2026

MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models

AAAI 2026technical

Answering complex medical questions requires not only domain expertise and patient-specific information, but also structured and multi-perspective reasoning. Existing multi-agent approaches often rely on fixed roles or shallow interaction prompts, limiting their ability to detect and resolve fine-gr

Cited by 0SourcePDFScholar
2026

PESD-TSF: A Period-Aware and Explicit Structured Decomposition Framework for Long-Term Time Series Forecasting

ICML 2026poster

Deep forecasting models often suffer from attenuated periodic perception and entangled trend–noise representations as network depth increases. Moreover, the widely adopted channel-independent paradigm, while improving training stability, disrupts intrinsic dynamic coordination among variables, hinde…

Cited by 0SourceScholar
2026

PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting

ICML 2026poster

Coupled spatiotemporal forecasting is important for predicting the future evolution of multiple interacting dynamical systems, such as in climate models. However, existing methods are severely constrained by the persistent bottleneck of compounding errors. In coupled systems, errors from each subsys…

Cited by 0SourceScholar
2026

Prune Wisely, Reconstruct Sharply: Compact 3D Gaussian Splatting via Adaptive Pruning and Difference-of-Gaussian Primitives

CVPR 2026

Recent significant advances in 3D scene representation have been driven by 3D Gaussian Splatting (3DGS), which has enabled real-time rendering with photorealistic quality. 3DGS often requires a large number of primitives to achieve high fidelity, leading to redundant representations and high resourc

Cited by 0SourcecodeScholar
2026

Resilient UAV Swarm with Fast Connectivity Recovery and Extensive Coverage

AAAI 2026technical

To address partial node failures in unmanned aerial vehicle swarms, self-healing communication techniques are commonly employed to restore backbone connectivity while preserving area coverage. However, existing heuristic methods struggle to scale under large-scale failures and dynamic conditions, wh

Cited by 0SourcePDFScholar
2026

S2R-HDR: A Large-Scale Rendered Dataset for HDR Fusion

ICLR 2026poster

The generalization of learning-based high dynamic range (HDR) fusion is often limited by the availability of training data, as collecting large-scale HDR images from dynamic scenes is both costly and technically challenging. To address these challenges, we propose S2R-HDR, the first large-scale high…

Cited by 0SourcecodeScholar
2026

SEMC: Structure-Enhanced Mixture-of-Experts Contrastive Learning for Ultrasound Standard Plane Recognition

AAAI 2026technical

Ultrasound standard plane recognition is essential for clinical tasks such as disease screening, organ evaluation, and biometric measurement. However, existing methods fail to effectively exploit shallow structural information and struggle to capture fine-grained semantic differences through contras

Cited by 0SourcePDFScholar
2026

SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation

ICLR 2026poster

Multi-track music generation has garnered significant research interest due to its precise mixing and remixing capabilities. However, existing models often overlook essential attributes such as rhythmic stability and synchronization, leading to a focus on differences between tracks rather than their…

Cited by 0SourceScholar
2026

S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm Detection

AAAI 2026technical

Multimodal sarcasm detection (MSD) aims to identify sarcasm polarity from diverse modalities (i.e., image–text pairs), a task that has received increasing attention. While significant progress has been made, existing approaches still face two major issues: lack of explainability and weak generalizab

Cited by 0SourcePDFScholar
2026

Trajectory-aware Shifted State Space Models for Online Video Super-Resolution

ICLR 2026poster

Online video super-resolution (VSR) is an important technique for many real-world video processing applications, which aims to restore the current high-resolution video frame based on temporally previous frames. Most of the existing online VSR methods solely employ one neighboring previous frame to…

Cited by 0SourcecodeScholar
2026

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

ICML 2026spotlight

Uniform Discrete Diffusion (UDM) has recently emerged as a promising paradigm for discrete generative modeling; however, its integration with reinforcement learning remains largely unexplored. We observe that naively adapting GRPO to UDM leads to unstable training and marginal performance. To addres…

Cited by 0SourceScholar
2026

Uniform Discrete Diffusion with Metric Path for Video Generation

ICLR 2026poster

Continuous-space video generation has advanced rapidly, while discrete approaches lag behind due to error accumulation and long-context inconsistency. In this work, we revisit discrete generative modeling and present Uniform discRete diffuSion with metric pAth (URSA), a simple yet powerful framework…

Cited by 0SourcecodeScholar
2025

A Generative Framework for Personalized Sticker Retrieval

EMNLP 2025

Formulating information retrieval as a variant of generative modeling, specifically using autoregressive models to generate relevant identifiers for a given query, has recently attracted considerable attention. However, its application to personalized sticker retrieval remains largely unexplored and

Cited by 0SourcePDFScholar
2025

A Survey on Foundation Language Models for Single-cell Biology

ACL 2025long

The recent advancements in language models have significantly catalyzed progress in computational biology. A growing body of research strives to construct unified foundation models for single-cell biology, with language models serving as the cornerstone. In this paper, we systematically review the d…

Cited by 0SourcePDFScholar
2025

A Survey on Multi-modal Intent Recognition: Recent Advances and New Frontiers

EMNLP 2025

Multi-modal intent recognition (MIR) requires integrating non-verbal cues from real-world contexts to enhance human intention understanding, which has attracted substantial research attention in recent years. Despite promising advancements, a comprehensive survey summarizing recent advances and new

2025

AdaptiveAE: An Adaptive Exposure Strategy for HDR Capturing in Dynamic Scenes

ICCV 2025poster

Mainstream high dynamic range imaging techniques typically rely on fusing multiple images captured with different exposure setups (shutter speed and ISO). A good balance between shutter speed and ISO is crucial for achieving high-quality HDR, as high ISO values introduce significant noise, while lon…

Cited by 0SourcePDFScholar
2025

Blind Video Super-Resolution based on Implicit Kernels

ICCV 2025poster

Blind video super-resolution (BVSR) is a low-level vision task which aims to generate high-resolution videos from low-resolution counterparts in unknown degradation scenarios. Existing approaches typically predict blur kernels that are spatially invariant in each video frame or even the entire video…

2025

CCDP: Composition of Conditional Diffusion Policies with Guided Sampling

IROS 2025

Imitation Learning offers a promising approach to learn directly from data without requiring explicit models, simulations, or detailed task definitions. During inference, actions are sampled from the learned distribution and executed on the robot. However, sampled actions may fail for various reason

Cited by 2SourcecodeScholar
2025

Can We Trust AI Doctors? A Survey of Medical Hallucination in Large Language and Large Vision-Language Models

ACL 2025finding

Hallucination has emerged as a critical challenge for large language models (LLMs) and large vision-language models (LVLMs), particularly in high-stakes medical applications. Despite its significance, dedicated research on medical hallucination remains unexplored. In this survey, we first provide a…

Cited by 0SourcePDFScholar
2025

CellVerse: Do Large Language Models Really Understand Cell Biology?

NeurIPS 2025poster

Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis ta…

Cited by 0SourcecodeScholar
2025

DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal Retrieval

AAAI 2025technical

With the burst of big data, 2D-3D cross-modal retrieval has received increasing attention, which aims to retrieve relevant data from one modality given the query from the other modality. In this paper, we study an underexplored yet practical problem of semi-supervised 2D-3D cross-modal retrieval, wh…

Cited by 0SourcePDFScholar
2025

Decoupled Feature Matching for Few-shot Counting and Localization

ICASSP 2025accepted

Few-shot counting (FSC) aims to train a generalized visual counting model that can count any novel category given a small number of support samples. Current prevalent approaches treat FSC as a feature-matching task, leveraging attention to aggregate information from all other query patches or suppor…

Cited by 0SourceScholar
2025

Diffusion Feedback Helps CLIP See Better

ICLR 2025poster

Contrastive Language-Image Pre-training (CLIP), which excels at abstracting open-world representations across domains and modalities, has become a foundation for a variety of vision and multimodal tasks. However, recent studies reveal that CLIP has severe visual shortcomings, such as which can hardl…

2025

Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models

ACL 2025long

Recent advancements in visual generative models have enabled high-quality image and video generation, opening diverse applications. However, evaluating these models often demands sampling hundreds or thousands of images or videos, making the process computationally expensive, especially for diffusio…

2025

Fully Connected Tensor Network based Brain Structural Feature Extraction for Early Alzheimer's Disease Detection

ICASSP 2025accepted

Alzheimer’s disease (AD) is an incurable neurodegenerative disease that involves structural changes in the brain. Early diagnosis of AD helps provide timely treatment and delay its progressive process. Many studies have been conducted based on brain images to detect AD. However, these works are most…

Cited by 0SourceScholar
2025

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

NeurIPS 2025poster

Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers,…

Cited by 0SourcecodeScholar
2025

GauUpdate: New Object Insertion in 3D Gaussian Fields with Consistent Global Illumination

ICCV 2025poster

3D Gaussian Splatting (3DGS) is a prevailing technique to reconstruct large-scale 3D scenes from multiview images for novel view synthesis, like a room, a block, and even a city. Such large-scale scenes are not static with changes constantly happening in these scenes, like a new building being built…

Cited by 0SourcePDFScholar
2025

HIIF: Hierarchical Encoding based Implicit Image Function for Continuous Super-resolution

CVPR 2025poster

Recent advances in implicit neural representations (INRs) have shown significant promise in modeling visual signals for various low-vision tasks including image super-resolution (ISR). INR-based ISR methods typically learn continuous representations, providing flexibility for generating high-resolut…

2025

HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos

CVPR 2025highlight

We introduce HOT3D, a publicly available dataset for egocentric hand and object tracking in 3D. The dataset offers over 833 minutes (3.7M+ images) of recordings that feature 19 subjects interacting with 33 diverse rigid objects. In addition to simple pick-up, observe, and put-down actions, the subje…

2025

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly

ICCV 2025poster

Numerous synthesized videos from generative models, especially human-centric ones that simulate realistic human actions, pose significant threats to human information security and authenticity. While progress has been made in binary forgery video detection, the lack of fine-grained understanding of…

Cited by 0SourcePDFScholar
2025

IntelliCockpitBench: A Comprehensive Benchmark to Evaluate VLMs for Intelligent Cockpit

ACL 2025finding

The integration of sophisticated Vision-Language Models (VLMs) in vehicular systems is revolutionizing vehicle interaction and safety, performing tasks such as Visual Question Answering (VQA). However, a critical gap persists due to the lack of a comprehensive benchmark for multimodal VQA models in…

2025

OneGT: One-Shot Geometry-Texture Neural Rendering for Head Avatars

ICCV 2025poster

Existing solutions for creating high-fidelity digital head avatars encounter various obstacles. Traditional rendering tools offer realistic results, while heavily requiring expert skills. Neural rendering methods are more efficient but often compromise between the generated fidelity and flexibility.…

Cited by 0SourcePDFScholar
2025

PreGenie: An Agentic Framework for High-quality Visual Presentation Generation

EMNLP 2025

Visual presentations are vital for effective communication. Early attempts to automate their creation using deep learning often faced issues such as poorly organized layouts, inaccurate text summarization, and a lack of image understanding, leading to mismatched visuals and text. These limitations r

Cited by 0SourcePDFScholar
2025

SGTC: Semantic-Guided Triplet Co-training for Sparsely Annotated Semi-Supervised Medical Image Segmentation

AAAI 2025technical

Although semi-supervised learning has made significant advances in the field of medical image segmentation, fully annotating a volumetric sample slice by slice remains a costly and time-consuming task. Even worse, most of the existing approaches pay much attention to image-level information and igno…

2025

ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

NeurIPS 2025poster

Recent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicat…

Cited by 0SourceScholar
2025

Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning

CVPR 2025poster

Heterogeneous Federated Learning (HFL) has received widespread attention due to its adaptability to different models and data. The HFL approach utilizing auxiliary models for knowledge transfer enhances flexibility. However, existing frameworks face the challenges of aggregation bias and local over…

2025

Super Capacity SRS Design for 5G and beyond using Channel In-painting

ICASSP 2025accepted

Reliable communication of data in modern wireless systems requires accurate channel state information (CSI). Sounding Reference Signal (SRS) based CSI acquisition enables the estimation of the channel between the base station and user equipment through the uplink transmission of known SRS by the use…

Cited by 0SourceScholar
2025

UltraFusion: Ultra High Dynamic Imaging using Exposure Fusion

CVPR 2025highlight

Capturing high dynamic range (HDR) scenes is one of the most important issues in camera design. Majority of cameras use exposure fusion, which fuses images captured by different exposure levels, to increase dynamic range. However, this approach can only handle images with limited exposure difference…

2025

Understanding PII Leakage in Large Language Models: A Systematic Survey

IJCAI 2025

Large Language Models (LLMs) have demonstrated exceptional success across a variety of tasks, particularly in natural language processing, leading to their growing integration into numerous facets of daily life. However, this widespread deployment has raised substantial privacy concerns, especially

2025

X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding

EMNLP 2025

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activity analysis, and personalized assistive technologies. However, existing benchmark

2024

A Cross Search Method for Data Augmentation in Neural Machine Translation

ICASSP 2024accepted

Large language models (LLMs) have shown excellent performance on general machine translation. However, LLMs suffer from high deployment cost and unsatisfying quality on low-resource domains. To this end, we explore to build base translation models with LLM-enhanced data augmentation. For data augmen…

Cited by 0SourceScholar
2024

A Lightweight Mixture-of-Experts Neural Machine Translation Model with Stage-wise Training Strategy

NAACL 2024findings

Dealing with language heterogeneity has always been one of the challenges in neural machine translation (NMT).The idea of using mixture-of-experts (MoE) naturally excels in addressing this issue by employing different experts to take responsibility for different problems.However, the parameter-ineff…

Cited by 2SourcePDFScholar
2024

Asynchronous Event-Inertial Odometry using a Unified Gaussian Process Regression Framework

IROS 2024poster

Recent works have combined monocular event camera and inertial measurement unit to estimate the SE(3) trajectory. However, the asynchronicity of event cameras brings a great challenge to conventional fusion algorithms. In this paper, we present an asynchronous event-inertial odometry under a unified…

Cited by 2SourceScholar
2024

Atlantis: Enabling Underwater Depth Estimation with Stable Diffusion

CVPR 2024highlight

Monocular depth estimation has experienced significant progress on terrestrial images in recent years thanks to deep learning advancements. But it remains inadequate for underwater scenes primarily due to data scarcity. Given the inherent challenges of light attenuation and backscatter in water acqu…

2024

CapsFusion: Rethinking Image-Text Data at Scale

CVPR 2024poster

Large multimodal models demonstrate remarkable generalist ability to perform diverse multimodal tasks in a zero-shot manner. Large-scale web-based image-text pairs contribute fundamentally to this success but suffer from excessive noise. Recent studies use alternative captions synthesized by caption…

2024

DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

NeurIPS 2024poster

Existing Multimodal Large Language Models (MLLMs) increasingly emphasize complex understanding of various visual elements, including multiple objects, text information, spatial relations. Their development for comprehensive visual perception hinges on the availability of high-quality image-text data…

2024

Directional Gain Based Noise Covariance Matrix Estimation for MVDR Beamforming

ICASSP 2024accepted

This paper is devoted to the problem of noise covariance matrix (NCM) estimation. It proposes a time-frequency masking based approach. We first present an optimal mask function based on the mean-squared error criterion. To estimate this mask, we employ the recently developed directional gain method…

Cited by 0SourceScholar
2024

DualDn: Dual-domain Denoising via Differentiable ISP

ECCV 2024poster

"Image denoising is a critical component in a camera’s Image Signal Processing (ISP) pipeline. There are two typical ways to inject a denoiser into the ISP pipeline: applying a denoiser directly to captured raw frames (raw domain) or to the ISP’s output sRGB images (sRGB domain). However, both appro…

2024

Emu: Generative Pretraining in Multimodality

ICLR 2024poster

We present Emu, a multimodal foundation model that seamlessly generates images and text in multimodal context. This omnivore model can take in any single-modality or multimodal data input indiscriminately (e.g., interleaved image, text and video) through a one-model-for-all autoregressive training p…

2024

Generative Multimodal Models are In-Context Learners

CVPR 2024poster

Humans can easily solve multimodal tasks in context with only a few demonstrations or simple instructions which current multimodal systems largely struggle to imitate. In this work we demonstrate that by effectively scaling up generative multimodal models their task-agnostic in-context learning capa…

2024

Global Terminal Sliding Mode Control of Tethered Satellites Formation with Chattering Reduction via PID Laws

ICRA 2024poster

This paper researches a novel global terminal sliding mode control(GTSMC) on a tethered satellites system(TSS) under outer disturbances, and the effect of PI/PD compensation in restraining chattering on sliding surface is appended. By taking advantage of the finite-time convergence of traditional te…

Cited by 0SourceScholar
2024

Gradient-Aware Logit Adjustment Loss for Long-Tailed Classifier

ICASSP 2024accepted

In the real-world setting, data often follows a long-tailed distribution, where head classes contain significantly more training samples than tail classes. Consequently, models trained on such data tend to be biased toward head classes. The medium of this bias is imbalanced gradients, which include…

Cited by 0SourceScholar
2024

LDMVFI: Video Frame Interpolation with Latent Diffusion Models

AAAI 2024technical

Existing works on video frame interpolation (VFI) mostly employ deep neural networks that are trained by minimizing the L1, L2, or deep feature space distance (e.g. VGG loss) between their outputs and ground-truth frames. However, recent works have shown that these metrics are poor indicators of per…

2024

LTGC: Long-tail Recognition via Leveraging LLMs-driven Generated Content

CVPR 2024poster

Long-tail recognition is challenging because it requires the model to learn good representations from tail categories and address imbalances across all categories. In this paper we propose a novel generative and fine-tuning framework LTGC to handle long-tail recognition via leveraging generated cont…

Cited by 16SourcePDFScholar
2024

Loop Structure-Aware Learning for Fully Automated Pulmonary Fissure Completeness Assessment

ICASSP 2024accepted

Pulmonary fissures are anatomical biomarkers used to evaluate the severity of chronic obstructive pulmonary disease. The completeness of the fissures is significantly associated with this disease. This work proposes a new fully automated fissure completeness assessment framework on the basis of deep…

Cited by 0SourceScholar
2024

MTKD: Multi-Teacher Knowledge Distillation for Image Super-Resolution

ECCV 2024poster

"Knowledge distillation (KD) has emerged as a promising technique in deep learning, typically employed to enhance a compact student network through learning from their high-performance but more complex teacher variant. When applied in the context of image super-resolution, most KD approaches are mod…

2024

Semi-supervised Knowledge Transfer Across Multi-omic Single-cell Data

NeurIPS 2024poster

Knowledge transfer between multi-omic single-cell data aims to effectively transfer cell types from scRNA-seq data to unannotated scATAC-seq data. Several approaches aim to reduce the heterogeneity of multi-omic data while maintaining the discriminability of cell types with extensive annotated data.…

Cited by 0SourcePDFScholar
2024

SkillNet-X: A Multilingual Multitask Model with Sparsely Activated Skills

ICASSP 2024accepted

Traditional multitask learning methods typically can only leverage shared knowledge within specific tasks or languages, resulting in a loss of either cross-language or cross-task knowledge. This paper proposes a general multilingual multitask model, named SkillNet-X, which enables a single model to…

Cited by 0SourceScholar
2024

Skip-Timeformer: Skip-Time Interaction Transformer for Long Sequence Time-Series Forecasting

IJCAI 2024poster

Recent studies have raised questions about the suitability of the Transformer architecture for long sequence time-series forecasting. These forecasting models leverage Transformers to capture dependencies between multiple time steps in a time series, with embedding tokens composed of data from indiv…

Cited by 9SourcePDFScholar
2024

VBench: Comprehensive Benchmark Suite for Video Generative Models

CVPR 2024highlight

Video generation has witnessed significant advancements yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal evaluation system should pro…

2023

Can Graph Neural Networks Learn to Solve the MaxSAT Problem? (Student Abstract)

AAAI 2023technical

The paper presents an attempt to bridge the gap between machine learning and symbolic reasoning. We build graph neural networks (GNNs) to predict the solution of the Maximum Satisfiability (MaxSAT) problem, an optimization variant of SAT. Two closely related graph representations are adopted, and we…

2023

Contrastive Self-Supervised Learning for Automated Multi-Modal Dance Performance Assessment

ICASSP 2023accepted

A fundamental challenge of analyzing human motion is to effectively represent human movements both spatially and temporally. We propose a contrastive self-supervised strategy to tackle this challenge. Particularly, we focus on dancing, which involves a high level of physical and intellectual abiliti…

Cited by 0SourceScholar
2023

Fed-CBS: A Heterogeneity-Aware Client Sampling Mechanism for Federated Learning via Class-Imbalance Reduction

ICML 2023poster

Due to the often limited communication bandwidth of edge devices, most existing federated learning (FL) methods randomly select only a subset of devices to participate in training at each communication round. Compared with engaging all the available clients, such a random-selection mechanism could l…

Cited by 58SourcePDFScholar
2023

HiNeRV: Video Compression with Hierarchical Encoding-based Neural Representation

NeurIPS 2023poster

Learning-based video compression is currently a popular research topic, offering the potential to compete with conventional standard video codecs. In this context, Implicit Neural Representations (INRs) have previously been used to represent and compress image and video content, demonstrating relati…

Cited by 55SourcePDFScholar
2023

MDCS: More Diverse Experts with Consistency Self-distillation for Long-tailed Recognition

ICCV 2023poster

Recently, multi-expert methods have led to significant improvements in long-tail recognition (LTR). We summarize two aspects that need further enhancement to contribute to LTR boosting: (1) More diverse experts; (2) Lower model variance. However, the previous methods didn't handle them well. To this…

Cited by 16PDFcodeScholar
2023

MixPro: Data Augmentation with MaskMix and Progressive Attention Labeling for Vision Transformer

ICLR 2023poster

The recently proposed data augmentation TransMix employs attention labels to help visual transformers (ViT) achieve better robustness and performance. However, TransMix is deficient in two aspects: 1) The image cropping method of TransMix may not be suitable for vision transformer. 2) At the early s…

2023

Purifier: Defending Data Inference Attacks via Transforming Confidence Scores

AAAI 2023technical

Neural networks are susceptible to data inference attacks such as the membership inference attack, the adversarial model inversion attack and the attribute inference attack, where the attacker could infer useful information such as the membership, the reconstruction or the sensitive attributes of a…

Cited by 19SourcePDFScholar
2023

Skillnet-NLG: General-Purpose Natural Language Generation with a Sparsely Activated Approach

ICASSP 2023accepted

We present SkillNet-NLG, a sparsely activated approach that handles many natural language generation tasks with one model. Different from traditional dense models that always activate all the parameters, SkillNet-NLG selectively activates relevant parts of the parameters to accomplish a task, where…

Cited by 0SourceScholar
2021

A Simplified Wiener Beamformer Based on Covariance Matrix Modelling

ICASSP 2021accepted

This paper is devoted to the problem of adaptive beamforming with small-spaced microphone arrays. In this context, the Wiener filter is an optimal beamformer in the mean-squared error (MSE) sense. However, it requires good estimates of the covariance matrices of the speech signal of interest and noi…

Cited by 0SourceScholar
2021

Improving Faithfulness in Abstractive Summarization with Contrast Candidate Generation and Selection

NAACL 2021long

Despite significant progress in neural abstractive summarization, recent studies have shown that the current models are prone to generating summaries that are unfaithful to the original context. To address the issue, we study contrast candidate generation and selection as a model-agnostic post-proce…

Cited by 118SourcePDFScholar
2021

Learning Temporal Consistency for Low Light Video Enhancement From Single Images

CVPR 2021poster

Single image low light enhancement is an important task and it has many practical applications. Most existing methods adopt a single image approach. Although their performance is satisfying on a static single image, we found, however, they suffer serious temporal instability when handling low light…

Cited by 165PDFcodeScholar
2021

On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation

ACL 2021long

In the recent advances of natural language processing, the scale of the state-of-the-art models and datasets is usually extensive, which challenges the application of sample-based explanation methods in many aspects, such as explanation interpretability, efficiency, and faithfulness. In this work, f…

Cited by 0SourcePDFScholar
2021

Self-Triggered Based Coordinate Control With Low Communication for Tethered Multi-UAV Collaborative Transportation

RA-L 2021

In this letter, a self-triggered based coordinate control scheme with low communication requirements is investigated for a team of unmanned aerial vehicles (UAVs) collaboratively transporting a suspended payload. In most existing research on collaborative transportation, the limited communication ab

Cited by 43SourceScholar
2021

Speech Emotion Recognition with Multiscale Area Attention and Data Augmentation

ICASSP 2021accepted

In Speech Emotion Recognition (SER), emotional characteristics often appear in diverse forms of energy patterns in spectrograms. Typical attention neural network classifiers of SER are usually optimized on a fixed attention granularity. In this paper, we apply multiscale area attention in a deep con…

Cited by 0SourceScholar
2020

Autonomous Obstacle Avoidance for UAV based on Fusion of Radar and Monocular Camera

IROS 2020poster

UAVs face many challenges in autonomous obstacle avoidance in large outdoor scenarios, specifically the long communication distance from ground stations. The computing power of onboard computers is limited, and the unknown obstacles cannot be accurately detected. In this paper, an autonomous obstacl…

Cited by 45SourceScholar
2020

Distributionally Robust Local Non-parametric Conditional Estimation

NeurIPS 2020poster

Conditional estimation given specific covariate values (i.e., local conditional estimation or functional estimation) is ubiquitously useful with applications in engineering, social and natural sciences. Existing data-driven non-parametric estimators mostly focus on structured homogeneous data (e.g.,…

2020

Distributionally Robust Policy Evaluation and Learning in Offline Contextual Bandits

ICML 2020poster

Policy learning using historical observational data is an important problem that has found widespread applications. However, existing literature rests on the crucial assumption that the future environment where the learned policy will be deployed is the same as the past environment that has generate…

Cited by 69SourcePDFScholar
2020

Unsupervised Instance Segmentation in Microscopy Images via Panoptic Domain Adaptation and Task Re-Weighting

CVPR 2020poster

Unsupervised domain adaptation (UDA) for nuclei instance segmentation is important for digital pathology, as it alleviates the burden of labor-intensive annotation and domain shift across datasets. In this work, we propose a Cycle Consistency Panoptic Domain Adaptive Mask R-CNN (CyC-PDAM) architectu…

Cited by 98PDFcodeScholar
2019

ACFNet: Attentional Class Feature Network for Semantic Segmentation

ICCV 2019poster

Recent works have made great progress in semantic segmentation by exploiting richer context, most of which are designed from a spatial perspective. In contrast to previous works, we present the concept of class center which extracts the global context from a categorical perspective. This class-level…

Cited by 357PDFScholar
2017

Locally-Transferred Fisher Vectors for Texture Classification

ICCV 2017poster

Texture classification has been extensively studied in computer vision. Recent research shows that the combination of Fisher vector (FV) encoding and convolutional neural network (CNN) provides significant improvement in texture classification over the previous feature representation methods. Howeve…

Cited by 75PDFScholar
2017

Preoperative planning for the multi-arm surgical robot using PSO-GP-based performance optimization

ICRA 2017poster

For the robotically-assisted minimally invasive surgery, preoperative planning is essential towards assisting surgeons to prepare the intervention and to decide the best access to the surgical site. Many recent studies in preoperative planning have focused on the pose selection of the robot and the…

Cited by 9SourceScholar
2015

An under-actuated manipulation controller based on Workspace Analysis and Gaussian Processes

IROS 2015poster

The kinematic modelling has been applied to many controllers of under-actuated manipulators. Most of these studies assume that the control process is conducted within the workspace. However, as such a kinematic model cannot describe the situations when the stable grasping is violated in the real env…

Cited by 7SourceScholar
2015

Fusing Subcategory Probabilities for Texture Classification

CVPR 2015poster

Texture, as a fundamental characteristic of objects, has attracted much attention in computer vision research. Performance of texture classification is however still lacking for some challenging cases, largely due to the high intra-class variation and low inter-class distinction. To tackle these iss…

Cited by 22SourcePDFScholar