← Search

Jing wang

133 accepted papers

2026

Breaking Semantic Boundaries: Distribution-Guided Semantic Exploration for Creative Generation

CVPR 2026

Text-to-image (T2I) diffusion models effectively produce semantically aligned images, but their reliance on training distributions constrains their capacity for synthesizing truly novel, out-of-distribution concepts. Existing methods attempt to enhance creativity through semantic exploration, such a

Cited by 0SourceScholar
2026

CADC: Content Adaptive Diffusion-Based Generative Image Compression

CVPR 2026

Diffusion-based generative image compression has demonstrated remarkable potential for achieving realistic reconstruction at ultra-low bitrates. The key to unlocking this potential lies in making the entire compression process content-adaptive, ensuring that the encoder's representation and the deco

Cited by 0SourceScholar
2026

Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation

AAAI 2026technical

Multimodal Emotion Recognition in Conversation (MERC) aims to predict speakers’ emotions by integrating textual, acoustic, and visual cues. Existing approaches either struggle to capture complex cross‑modal interactions or experience gradient conflicts and unstable training when using deeper archite

Cited by 0SourcePDFScholar
2026

DiffStyle3D: Consistent 3D Gaussian Stylization via Attention Optimization

ICML 2026poster

3D style transfer enables the creation of visually expressive 3D content, enriching the visual appearance of 3D scenes and objects. However, existing VGG- and CLIP-based methods struggle to model multi-view consistency within the model itself, while diffusion-based approaches can capture such consis…

Cited by 0SourceScholar
2026

DivControl: Knowledge Diversion for Controllable Image Generation

AAAI 2026technical

Diffusion models have advanced from text-to-image (T2I) to image-to-image (I2I) generation by incorporating structured inputs such as depth maps, enabling fine-grained spatial control. However, existing methods either train separate models for each condition or rely on unified architectures with ent

Cited by 0SourcePDFScholar
2026

DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO

CVPR 2026

Reinforcement learning (RL), particularly GRPO, improves image generation quality significantly by comparing the relative performance of images generated within the same group. However, in the later stages of training, the model tends to produce homogenized outputs, lacking creativity and visual div

Cited by 0SourceScholar
2026

EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion

AAAI 2026technical

Despite the remarkable developments achieved by recent 3D generation works, scaling these methods to geographic extents, such as modeling thousands of square kilometers of Earth’s surface, remains an open challenge. We address this through a dual innovation in data infrastructure and model architect

Cited by 0SourcePDFScholar
2026

FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion Models

CVPR 2026

The training of diffusion models is computationally intensive, making effective pre-training essential. However, real-world deployments often demand models of variable sizes due to diverse memory and computational constraints, posing challenges when corresponding pre-trained versions are unavailable

Cited by 0SourceScholar
2026

Fast Inverse Lithography via GRPO Reinforced Flow Matching

ICML 2026poster

In semiconductor manufacturing, lithography projects circuit layouts onto silicon wafers through an optical mask. As circuit features shrink below the wavelength of light, optical diffraction causes the printed patterns to deviate from their intended layouts. Inverse Lithography Technology (ILT) add…

Cited by 0SourceScholar
2026

Fast and Interpretable Protein Substructure Alignment via Optimal Transport

ICLR 2026poster

Proteins are essential biological macromolecules that execute life functions. Local motifs within protein structures, such as active sites, are the most critical components for linking structure to function and are key to understanding protein evolution and enabling protein engineering. Existing com…

Cited by 0SourceScholar
2026

GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping

CVPR 2026

Recently, GRPO-based reinforcement learning has shown remarkable progress in optimizing flow-matching models, effectively improving their alignment with task-specific rewards. Within these frameworks, the policy update relies on importance-ratio clipping to constrain overconfident positive and negat

Cited by 0SourcecodeScholar
2026

Geometry-based Schrödinger Bridges for Trustworthy Multimodal Fusion

ICML 2026poster

Real-world multimodal systems must be robust against low-quality data, such as sensor noise, incomplete multimodal data and conflicting inputs. However, existing trustworthy fusion methods rely on the model's own prediction confidence to judge data quality. This creates a circular dependency: when a…

Cited by 0SourceScholar
2026

Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization

CVPR 2026

Generating high-fidelity audio that is both semantically meaningful and temporally synchronized with silent videos remains a challenging problem in video-to-audio generation. Existing approaches often fail to capture fine-grained temporal correspondence between visual events and audio dynamics, lead

Cited by 0SourcecodeScholar
2026

IF-VidCap: Can Video Caption Models Follow Instructions?

ICLR 2026poster

Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptiv…

Cited by 0SourcecodeScholar
2026

Inconsistency-aware Multimodal Schrodinger Bridge for Deepfake Localization

CVPR 2026

Audio-visual deepfake localization demands interval-level outputs that serve as temporal evidence. Despite recent progress, symmetric fusion under single-sided or asynchronous forgeries propagates cross-modal noise, degrading high-precision localization. We present IaMSB, an inconsistency-aware mult

Cited by 0SourceScholar
2026

Knowledge Diversion for Efficient Morphology Control and Policy Transfer

ICML 2026poster

Universal morphology control aims to learn a universal policy that generalizes across heterogeneous robot morphologies, with Transformer-based controllers emerging as a dominant choice. However, such architectures incur substantial computational costs, resulting in high deployment overhead, and exis…

Cited by 0SourceScholar
2026

Learngene: Inheritable ‘Genes’ in Intelligent Agents (Abstract Reprint)

AAAI 2026technical

Biological intelligence has driven significant progress in artificial intelligence (AI), but a critical gap remains: biological systems inherit innate abilities from genes, with brains initialized by blueprints refined over 3.5 billion years of evolution, while machines rely heavily on inefficient,

Cited by 0SourcePDFScholar
2026

MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On

CVPR 2026

Virtual try-on aims to fit an in-shop clothing image onto a specific human body. An optimal virtual try-on method should provide diverse and flexible dressing options, accurately reflecting the varied wearing styles encountered in real-life scenarios, tailored to individual preferences and fashion a

Cited by 0SourceScholar
2026

Multimodal Fusion via Self-Consistent Task-Gradient Fields

ICML 2026poster

Multimodal learning aims to preserve as much task-related information as possible from different inputs. However, current fusion designs often distort the feedback loop to feature extractors. Aggressively merging modalities entangles their representations, making the feature extractors fragile to in…

Cited by 0SourceScholar
2026

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

CVPR 2026

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex dynamics. This limitation arises primarily because existing approa

Cited by 0SourceScholar
2026

RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers

AAAI 2026technical

The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource

Cited by 0SourcePDFScholar
2026

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion

AAAI 2026technical

Denoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trained diffusion priors. However, they typically require multi-step iterations in inference, which limits efficiency. To a

Cited by 0SourcePDFScholar
2026

Self-Supervised Weight Templates for Scalable Vision Model Initialization

ICML 2026poster

The increasing scale and complexity of modern model parameters underscore the importance of pre-trained models. However, deployment often demands architectures of varying sizes, exposing limitations of conventional pre-training and fine-tuning. To address this, we propose SWEET, a self-supervised fr…

Cited by 0SourceScholar
2026

U$^3$CF: Unbiased, Unconfounding, and Unified Causal Framework for Multi-Target Domain Adaptation

ICML 2026poster

Multi-target domain adaptation (MTDA) trains a model using a labeled source domain and several unlabeled target domains, aiming to enhance performance across all targets. However, existing methods lack a principled causal formulation and often rely on empirical domain-invariance enforcement, which c…

Cited by 0SourceScholar
2026

When and How to Adapt: Subject Shifts Detection and Prototype-Guided Correction for Online EEG Decoding

IJCAI 2026

Online EEG decoding is pivotal for real-world Brain-Computer Interfaces (BCIs) but confronts significant challenges arising from continuous distribution shifts, including both inter- and intra-subject variations. Existing Unsupervised Continual Domain Adaptation (UCDA) methods typically rely on rigi

Cited by 0Scholar
2025

Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies

NeurIPS 2025poster

Existing imitation learning methods decouple perception and action, which overlooks the causal reciprocity between sensory representations and action execution that humans naturally leverage for adaptive behaviors. To bridge this gap, we introduce Action-Guided Diffusion Policy (DP-AG), a unified re…

Cited by 0SourcecodeScholar
2025

AdaptCMVC: Robust Adaption to Incremental Views in Continual Multi-view Clustering

CVPR 2025poster

Most Multi-view Clustering approaches assume that all views are available for clustering. However, this assumption is often unrealistic as views are incrementally accumulated over time, leading to a need for continual multi-view clustering (CMVC) methods. Current approaches to CMVC leverage late fus…

Cited by 0SourcePDFScholar
2025

An End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion Models

CVPR 2025poster

Existing conditional Denoising Diffusion Probabilistic Models (DDPMs) with a Noise-Conditional Framework (NCF) remain challenging for 3D scene understanding tasks, as the complex geometric details in scenes increase the difficulty of fitting the gradients of the data distribution (the scores) from s…

2025

A³-Net: Calibration-Free Multi-View 3D Hand Reconstruction for Enhanced Musical Instrument Learning

IJCAI 2025

Precise 3D hand posture is essential for learning musical instruments. Reconstructing highly precise 3D hand gestures enables learners to correct and master proper techniques through 3D simulation and Extended Reality. However, exsiting methods typically rely on precisely calibrated multi-camera sys

Cited by 0SourcePDFScholar
2025

COSDA: Counterfactual-based Susceptibility Risk Framework for Open-Set Domain Adaptation

ICML 2025poster

Open-Set Domain Adaptation (OSDA) aims to transfer knowledge from the labeled source domain to the unlabeled target domain that contains unknown categories, thus facing the challenges of domain shift and unknown category recognition. While recent works have demonstrated the potential of causality fo…

Cited by 0SourcePDFScholar
2025

CircuitFusion: Multimodal Circuit Representation Learning for Agile Chip Design

ICLR 2025poster

The rapid advancements of AI rely on the support of integrated circuits (ICs). However, the growing complexity of digital ICs makes the traditional IC design process costly and time-consuming. In recent years, AI-assisted IC design methods have demonstrated great potential, but most methods are task…

2025

CoDTS: Enhancing Sparsely Supervised Collaborative Perception with a Dual Teacher-Student Framework

AAAI 2025technical

Current collaborative perception methods often rely on fully annotated datasets, which can be expensive to obtain in practical situations. To reduce annotation costs, some works adopt sparsely supervised learning techniques and generate pseudo labels for the missing instances. However, these methods…

Cited by 0SourcePDFScholar
2025

CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control

ACL 2025finding

Retrieval-augmented generation (RAG) has emerged as a promising solution for mitigating hallucinations of large language models (LLMs) with retrieved external knowledge. Adaptive RAG enhances this approach by enabling dynamic retrieval during generation, activating retrieval only when the query exce…

2025

DRP: A Decomposition-Reflection-Prediction Framework for Long-Horizon Robot Task Planning using Large Language Models

IROS 2025

Large language models have demonstrated powerful reasoning capabilities, and their integration with robotics has revolutionized human-computer interaction and automated task planning. However, LLMs are unaware of environmental knowledge and possible state changes in the environment during planning,

Cited by 0SourcecodeScholar
2025

DiMa: Understanding the Hardness of Online Matching Problems via Diffusion Models

ICML 2025poster

We explore the potential of \emph{AI-enhanced combinatorial optimization theory}, taking online bipartite matching (OBM) as a case study. In the theoretical study of OBM, the \emph{hardness} corresponds to a performance \emph{upper bound} of a specific online algorithm or any possible online algorit…

Cited by 0SourcePDFScholar
2025

Diffusion Model with Multi-layer Wavelet Transform for Low-Light Image Enhancement

ICASSP 2025accepted

Low-light image enhancement methods based on diffusion models, though effective in improving image quality, often overrely on noise sensitivity and neglect the reconstruction deviations due to the naive up- and down-sampling operations. To address this issue, we propose a novel diffusion model, MWT-…

Cited by 0SourceScholar
2025

DreamGen: Unlocking Generalization in Robot Learning through Video World Models

CoRL 2025poster

In this work, we unlock new capabilities in robot learning from neural trajectories, synthetic robot data generated from video world models. Our proposed recipe is simple, but powerful: we take the most recent state-of-the-art video generative models (world models), adapt them to the target robot em…

Cited by 0SourcecodeScholar
2025

ECO: Evolving Core Knowledge for Efficient Transfer

NeurIPS 2025poster

Knowledge in modern neural networks is often entangled and structurally opaque, making current transfer methods—typically based on reusing entire parameter sets—inefficient and inflexible. Efforts to improve flexibility by reusing partial parameters frequently depend on handcrafted heuristics or rig…

Cited by 0SourceScholar
2025

Enhancing Consistency of Flow-Based Image Editing through Kalman Control

NeurIPS 2025poster

Flow-based generative models have gained popularity for image generation and editing. For instruction-based image editing, it is critical to ensure that modifications are confined to the targeted regions. Yet existing methods often fail to maintain consistency in non-targeted regions between the ori…

Cited by 0SourceScholar
2025

Evaluating and Mitigating Object Hallucination in Large Vision-Language Models: Can They Still See Removed Objects?

NAACL 2025long

Large Vision-Language Models (LVLMs) have a significant issue with object hallucinations, where researchers have noted that LVLMs often mistakenly determine objects as present in images where they do not actually exist. Some recent studies evaluate the occurrence of object hallucinations by asking L…

Cited by 0SourcePDFScholar
2025

FLARE: Robot Learning with Implicit World Modeling

CoRL 2025poster

We introduce **F**uture **LA**tent **R**presentation Alignm**E**nt (**FLARE**), a novel framework that integrates predictive world modeling into robot policy learning. By aligning features from a diffusion transformer with latent embeddings of future observations, **FLARE** enables a diffusion trans…

Cited by 0SourceScholar
2025

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

IJCAI 2025

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text control, equivalently guiding different frame generations wi

Cited by 0SourcePDFScholar
2025

From Indicators to Insights: Diversity-Optimized for Medical Series-Text Decoding via LLMs

NeurIPS 2025poster

Medical time-series analysis differs fundamentally from general ones by requiring specialized domain knowledge to interpret complex signals and clinical context. Large language models (LLMs) hold great promise for augmenting medical time-series analysis by complementing raw series with rich contextu…

Cited by 0SourcecodeScholar
2025

FuncGenFoil: Airfoil Generation and Editing Model in Function Space

NeurIPS 2025poster

Aircraft manufacturing is the jewel in the crown of industry, in which generating high-fidelity airfoil geometries with controllable and editable representations remains a fundamental challenge. Existing deep learning methods, which typically rely on predefined parametric representations (e.g., Bézi…

Cited by 0SourcecodeScholar
2025

Generalizable Hand-Object Modeling from Monocular RGB Images via 3D Gaussians

NeurIPS 2025poster

Recent advances in hand-object interaction modeling have employed implicit representations, such as Signed Distance Functions (SDF) and Neural Radiance Fields (NeRF) to reconstruct hands and objects with arbitrary topology and photo-realistic detail. However, these methods often rely on dense 3D sur…

Cited by 0SourceScholar
2025

Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update

ICML 2025poster

We study the stochastic linear bandits with heavy-tailed noise. Two principled strategies for handling heavy-tailed noise, truncation and median-of-means, have been introduced to heavy-tailed bandits. Nonetheless, these methods rely on specific noise assumptions or bandit structures, limiting their…

Cited by 0SourcePDFScholar
2025

KIND: Knowledge Integration and Diversion for Training Decomposable Models

ICML 2025poster

Pre-trained models have become the preferred backbone due to the increasing complexity of model parameters. However, traditional pre-trained models often face deployment challenges due to their fixed sizes, and are prone to negative transfer when discrepancies arise between training tasks and target…

2025

Label Distribution Learning with Biased Annotations Assisted by Multi-Label Learning

IJCAI 2025

Multi-label learning (MLL) has gained attention for its ability to represent real-world data. Label Distribution Learning (LDL), an extension of MLL to learning from label distributions, faces challenges in collecting accurate label distributions. To address the issue of biased annotations, based on

Cited by 0SourcePDFScholar
2025

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

ICCV 2025poster

Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue to face challenges in maintaining subject consistency due to the lack of fine-grained guidance and inter-frame interact…

Cited by 0SourcePDFScholar
2025

M2PAIR: A High-Quality Acoustic Impulse Response Computation Model

ICASSP 2025accepted

Acoustic Impulse Response (AIR) provides crucial spatial information about the environment, significantly enhancing audio immersion. However, achieving high perceptual quality while computing AIR in real-time for interactive audio-video media (IAVM) presents a challenging problem. This study propose…

Cited by 0SourceScholar
2025

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding

ICCV 2025poster

Accurate driving behavior recognition and reasoning are critical for autonomous driving video understanding. However, existing methods often tend to dig out the shallow causal, fail to address spurious correlations across modalities, and ignore the ego-vehicle level causality modeling. To overcome t…

2025

MonoSG: Monocular 3D Object Detection With Stereo Guidance

RA-L 2025

In the context of autonomous driving, monocular 3D detection is regarded as a fundamental and essential task due to its convenience, speed, and low cost. However, the lack of depth information in monocular images presents significant challenges for predicting object 3D information. Although existing

Cited by 5SourceScholar
2025

MusicMamba: A Dual-Feature Modeling Approach for Generating Chinese Traditional Music with Modal Precision

ICASSP 2025accepted

In recent years, deep learning has advanced the MIDI domain, solidifying music generation as a key application of artificial intelligence. However, most research focuses on Western music, facing challenges in generating Chinese traditional melodies, particularly in capturing modal characteristics an…

Cited by 0SourceScholar
2025

Online Video Understanding: OVBench and VideoChat-Online

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have significantly progressed in offline video understanding. However, applying these models to real-world scenarios, such as autonomous driving and human-computer interaction, presents unique challenges due to the need for real-time processing of continuous…

Cited by 0SourcePDFScholar
2025

Optimal Information Retention for Time-Series Explanations

ICML 2025poster

Explaining deep models for time-series data is crucial for identifying key patterns in sensitive domains, such as healthcare and finance. However, due to the lack of unified optimization criterion, existing explanation methods often suffer from redundancy and incompleteness, where irrelevant pattern…

2025

PT-T2I/V: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Image/Video-Task

ICLR 2025poster

The global self-attention mechanism in diffusion transformers involves redundant computation due to the sparse and redundant nature of visual information, and the attention map of tokens within a spatial window shows significant similarity. To address this redundancy, we propose the Proxy-Tokenized…

2025

PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution

CVPR 2025poster

Pre-trained video generation models hold great potential for generative video super-resolution (VSR). However, adapting them for full-size VSR, as most existing methods do, suffers from unnecessary intensive full-attention computation and fixed output resolution. To overcome these limitations, we ma…

Cited by 0SourcePDFScholar
2025

Prior-aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose Estimation

ICCV 2025poster

3D hand pose estimation plays a critical role in various human-computer interaction tasks. Single-frame 3D hand pose estimation methods have poor temporal smoothness and are easily affected by self-occlusion, which severely impacts their practical applicability. Traditional joint-based sequential po…

Cited by 0SourcePDFScholar
2025

REFED: A Subject Real-time Dynamic Labeled EEG-fNIRS Synchronized Recorded Emotion Dataset

NeurIPS 2025poster

Affective brain-computer interfaces (aBCIs) play a crucial role in personalized human–computer interaction and neurofeedback modulation. To develop practical and effective aBCI paradigms and to investigate the spatial-temporal dynamics of brain activity under emotional inducement, portable electroen…

Cited by 0SourceScholar
2025

RankMatch: A Novel Approach to Semi-Supervised Label Distribution Learning Leveraging Rank Correlation between Labels

NeurIPS 2025poster

Pseudo label based semi-supervised learning (SSL) for single-label and multi-label classification tasks has been extensively studied; however, semi-supervised label distribution learning (SSLDL) remains a largely unexplored area. Existing SSL methods fail in SSLDL because the pseudo-labels they ge…

Cited by 0SourceScholar
2025

Redefining <Creative> in Dictionary: Towards an Enhanced Semantic Understanding of Creative Generation

CVPR 2025poster

Creative remains an inherently abstract concept for both humans and diffusion models. While text-to-image (T2I) diffusion models can easily generate out-of-distribution concepts like "a blue banana", they struggle with generating combinatorial objects such as "a creative mixture that resembles a let…

2025

SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action Recognition

ICCV 2025poster

Open-vocabulary action recognition (OVAR) extends recognition systems to identify unseen action categories. While large-scale vision-language models (VLMs) like CLIP have enabled OVAR in image domains, their adaptation to event data remains underexplored. Event cameras offer high temporal resolution…

2025

SEHAP: Secure and Efficient Handover Authentication Protocol in LEO Satellite Non-Terrestrial Networks

ICASSP 2025accepted

LEO satellite non-terrestrial networks (NTN) utilize satellites in Low Earth Orbit (LEO) to dynamically establish global communication service and own significant promise. The dynamic nature of LEO satellite NTN necessities efficient handover authentication protocols. However existing schemes cannot…

Cited by 0SourceScholar
2025

SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering

AAAI 2025technical

The general capabilities of large language models (LLMs) make them the infrastructure for various AI applications, but updating their inner knowledge requires significant resources. Recent model editing is a promising technique for efficiently updating a small amount of knowledge of LLMs and has att…

2025

Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience

EMNLP 2025

Large language models (LLMs) generate human-aligned content under certain safety constraints. However, the current known technique “jailbreak prompt” can circumvent safety-aligned measures and induce LLMs to output malicious content. Research on Jailbreaking can help identify vulnerabilities in LLMs

Cited by 0SourcePDFScholar
2025

StreamForest: Efficient Online Video Understanding with Persistent Event Memory

NeurIPS 2025spotlight

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains limited due to storage constraints of historical visual features and insufficient real-time spatiotemporal reasoning. To a…

Cited by 0SourceScholar
2025

SynC-LLM: Generation of Large-Scale Synthetic Circuit Code with Hierarchical Language Models

EMNLP 2025

In recent years, AI-assisted integrated circuit (IC) design methods have shown great potential in boosting IC design efficiency. However, this emerging technique is fundamentally limited by the serious scarcity of publicly accessible large-scale circuit design data, which are mostly private IPs owne

2025

Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling

EMNLP 2025

Direct Prompt Injection (DPI) attacks pose a critical security threat to Large Language Models (LLMs) due to their low barrier of execution and high potential damage. To address the impracticality of existing white-box/gray-box methods and the poor transferability of black-box methods, we propose an

Cited by 0SourcePDFScholar
2025

Understanding the Unfairness in Network Quantization

ICML 2025poster

Network quantization, one of the most widely studied model compression methods, effectively quantizes a floating-point model to obtain a fixed-point one with negligible accuracy loss. Although great success was achieved in reducing the model size, it may exacerbate the unfairness in model accuracy…

Cited by 0SourcePDFScholar
2025

Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose Estimation

NeurIPS 2025poster

Multi-view 3D human pose estimation (HPE) leverages complementary information across views to improve accuracy and robustness. Traditional methods rely on camera calibration to establish geometric correspondences, which is sensitive to calibration accuracy and lacks flexibility in dynamic settings.…

Cited by 0SourceScholar
2025

Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate Ratio

EMNLP 2025

Existing research in multi-hop questions has identified two reasoning modes: latent reasoning and factual shortcuts, but has not deeply investigated how these modes differ during inference. This impacts both model generalization ability and downstream reasoning tasks. In this work, we systematically

Cited by 0SourcePDFScholar
2025

Vicinity-Guided Discriminative Latent Diffusion for Privacy-Preserving Domain Adaptation

NeurIPS 2025poster

Recent work on latent diffusion models (LDMs) has focused almost exclusively on generative tasks, leaving their potential for discriminative transfer largely unexplored. We introduce Discriminative Vicinity Diffusion (DVD), a novel LDM-based framework for a more practical variant of source-free doma…

Cited by 0SourcecodeScholar
2025

WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models

CVPR 2025poster

The growing complexity of model parameters underscores the significance of pre-trained models. However, deployment constraints often necessitate models of varying sizes, exposing limitations in the conventional pre-training and fine-tuning paradigm, particularly when target model sizes are incompati…

2025

WISA: World simulator assistant for physics-aware text-to-video generation

NeurIPS 2025spotlight

Recent advances in text-to-video (T2V) generation, exemplified by models such as Sora and Kling, have demonstrated strong potential for constructing world simulators. However, existing T2V models still struggle to understand abstract physical principles and to generate videos that faithfully obey ph…

Cited by 0SourcecodeScholar
2024

AFBench: A Large-scale Benchmark for Airfoil Design

NeurIPS 2024poster

Data-driven generative models have emerged as promising approaches towards achieving efficient mechanical inverse design. However, due to prohibitively high cost in time and money, there is still lack of open-source and large-scale benchmarks in this field. It is mainly the case for airfoil inverse…

2024

Adaptive FSS: A Novel Few-Shot Segmentation Framework via Prototype Enhancement

AAAI 2024technical

The Few-Shot Segmentation (FSS) aims to accomplish the novel class segmentation task with a few annotated images. Current FSS research based on meta-learning focuses on designing a complex interaction mechanism between the query and support feature. However, unlike humans who can rapidly learn new t…

2024

AutoDrop: Training Deep Learning Models with Automatic Learning Rate Drop

UAI 2024poster

Modern deep learning (DL) architectures are trained using variants of the SGD algorithm and typically rely on the user to manually drop the learning rate when the training curve saturates. In this paper, we develop an algorithm, that we call AutoDrop, that realizes the learning rate drop automatical…

2024

Cluster-Learngene: Inheriting Adaptive Clusters for Vision Transformers

NeurIPS 2024poster

In recent years, the merging of vast datasets with powerful computational resources has led to the emergence of large pre-trained models in the field of deep learning. However, the common practices often overgeneralize the applicability of these models, overlooking the task-specific resource constra…

Cited by 1SourcePDFScholar
2024

DMT: Comprehensive Distillation with Multiple Self-Supervised Teachers

ICASSP 2024accepted

Numerous self-supervised learning paradigms, such as contrastive learning and masked image modeling, have been proposed to acquire powerful and general representations from unlabeled data. However, these models are commonly pretrained within their specific framework alone, failing to consider the co…

Cited by 0SourceScholar
2024

Exploiting Multi-Label Correlation in Label Distribution Learning

IJCAI 2024poster

Label Distribution Learning (LDL) is a novel machine learning paradigm that assigns label distribution to each instance. Numerous LDL methods proposed to leverage label correlation in the learning process to solve the exponential-sized output space; among these, many exploited the low-rank structur…

2024

Exploring Active Learning in Meta-Learning: Enhancing Context Set Labeling

ECCV 2024poster

"Most meta-learning methods assume that the (very small) context set used to establish a new task at test time is passively provided. In some settings, however, it is feasible to actively select which points to label; the potential gain from a careful choice is substantial, but the setting requires…

2024

HPipe: Large Language Model Pipeline Parallelism for Long Context on Heterogeneous Cost-effective Devices

NAACL 2024industry

Micro-enterprises and individual developers emerge analysis demands for long sequence with powerful Large Language Models (LLMs). They try to deploy the LLMs at local, but only possess various commodity devices and the unreliable interconnection between devices. Existing parallel techniques do not l…

Cited by 6SourcePDFScholar
2024

LightCodec: A High Fidelity Neural Audio Codec with Low Computation Complexity

ICASSP 2024accepted

The audio codec is one of the core modules in audio communication for real-time transmission. With the development of neural networks, end-to-end audio codecs have emerged and demonstrated effects beyond conventional codecs. However, current neural network-based codecs have the weakness of high comp…

Cited by 0SourceScholar
2024

Neural Rate Control for Learned Video Compression

ICLR 2024poster

The learning-based video compression method has made significant progress in recent years, exhibiting promising compression performance compared with traditional video codecs. However, prior works have primarily focused on advanced compression architectures while neglecting the rate control techniqu…

Cited by 6SourcePDFScholar
2024

Non-Intrusive Speech Quality Assessment with Multi-Task Learning Based on Tensor Network

ICASSP 2024accepted

With the growing significance of non-intrusive speech quality assessment in speech systems, existing methods predominantly rely on neural networks to extract low-order features. Typically, these features undergo a low-dimensional linear transformation, yielding the network’s output. However, the int…

Cited by 0SourceScholar
2024

SURER: Structure-Adaptive Unified Graph Neural Network for Multi-View Clustering

AAAI 2024technical

Deep Multi-view Graph Clustering (DMGC) aims to partition instances into different groups using the graph information extracted from multi-view data. The mainstream framework of DMGC methods applies graph neural networks to embed structure information into the view-specific representations and fuse…

Cited by 9SourcePDFScholar
2023

Electromagnetic Clutch-Based Ankle Exosuit for Assisting Stroke Survivors With Different Body Sizes

RA-L 2023

Exosuits can be effective in aiding stroke rehabilitation. However, single-motor exosuits are <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">challenging</b> to adapt to users with different body sizes. <italic xmlns:mml="http://www.w3.org/1998/Math/Ma

Cited by 5SourceScholar
2023

Exploiting Interactivity and Heterogeneity for Sleep Stage Classification Via Heterogeneous Graph Neural Network

ICASSP 2023accepted

Sleep stage classification based on physiological time-series is essential for sleep quality evaluation and the diagnosis of sleep disorders in clinical practice. Existing machine learning studies have achieved adequate results in sleep stage classification. However, those methods neglect the signif…

Cited by 0SourceScholar
2023

Robust Video Anomaly Detection Framework via Prior Knowledge and Multi-Path Frame Prediction

ICASSP 2023accepted

Video anomaly detection aims to automatically detect abnormal objects or behaviors. Most existing methods tackle the problem by minimizing the reconstruction errors stemming from the lack of anomalous data, which leads to poor interpretability and robustness. Focus on the context-dependent nature of…

Cited by 0SourceScholar
2023

Sample-Adapt Fusion Network for RGB-D Hand Detection in the Wild

ICASSP 2023accepted

RGB and depth modalities provide complementary information, which can be effectively utilized to improve the performance of hand detection in the wild. Most existing fusion-based methods model the channel-wise or spatial-wise cross-modal correlation to exploit the complementary RGB-D information, in…

Cited by 0SourceScholar
2023

Semi-Supervised Sound Event Detection with Pre-Trained Model

ICASSP 2023accepted

Sound event detection (SED) is an interesting but challenging task due to the scarcity of data and diverse sound events in real life. In this paper, we focus on the semi-supervised SED task, and combine pre-trained model from other field to assist in improving the detection effect. Pre-trained model…

Cited by 0SourceScholar
2023

Tight and fast generalization error bound of graph embedding in metric space

ICML 2023poster

Recent studies have experimentally shown that we can achieve in non-Euclidean metric space effective and efficient graph embedding, which aims to obtain the vertices' representations reflecting the graph's structure in the metric space. Specifically, graph embedding in hyperbolic space has experimen…

Cited by 1SourcePDFScholar
2022

Detecting Tampered Scene Text in the Wild

ECCV 2022poster

"Text manipulation technologies cause serious worries in recent years, however, corresponding tampering detection methods have not been well explored. In this paper, we introduce a new task, named Tampered Scene Text Detection (TSTD), to localize text instances and recognize the texture authenticity…

2022

Learning to Socially Navigate in Pedestrian-rich Environments with Interaction Capacity

ICRA 2022poster

Existing navigation policies for autonomous robots tend to focus on collision avoidance while ignoring human-robot interactions in social life. For instance, robots can pass along the corridor safer and easier if pedestrians notice them. Sounds have been considered as an efficient way to attract the…

Cited by 18SourceScholar
2022

Low-Pass Filtering SGD for Recovering Flat Optima in the Deep Learning Optimization Landscape

AISTATS 2022poster

In this paper, we study the sharpness of a deep learning (DL) loss landscape around local minima in order to reveal systematic mechanisms underlying the generalization abilities of DL models. Our analysis is performed across varying network and optimizer hyper-parameters, and involves a rich family…

2022

MOS Predictor for Synthetic Speech with I-Vector Inputs

ICASSP 2022accepted

Based on deep learning technology, non-intrusive methods have received increasing attention for synthetic speech quality assessment since it does not need reference signals. Meanwhile, i-vector has been widely used in paralinguistic speech attribute recognition such as speaker and emotion recognitio…

Cited by 0SourceScholar
2022

Med-DANet: Dynamic Architecture Network for Efficient Medical Volumetric Segmentation

ECCV 2022poster

"For 3D medical image (e.g. CT and MRI) segmentation, the difficulty of segmenting each slice in a clinical case varies greatly. Previous research on volumetric medical image segmentation in a slice-by-slice manner conventionally use the identical 2D deep neural network to segment all the slices of…

2022

Metric-Fair Active Learning

ICML 2022spotlight

Active learning has become a prevalent technique for designing label-efficient algorithms, where the central principle is to only query and fit “informative” labeled instances. It is, however, known that an active learning algorithm may incur unfairness due to such instance selection procedure. In t…

Cited by 9SourcePDFScholar
2022

Modeling Aspect Correlation for Aspect-based Sentiment Analysis via Recurrent Inverse Learning Guidance

COLING 2022main

Aspect-based sentiment analysis (ABSA) aims to distinguish sentiment polarity of every specific aspect in a given sentence. Previous researches have realized the importance of interactive learning with context and aspects. However, these methods are ill-studied to learn complex sentence with multipl…

Cited by 3SourcePDFScholar
2022

Multi-Level Spatial-Temporal Adaptation Network for Motor Imagery Classification

ICASSP 2022accepted

Electroencephalogram (EEG) signals for motor imagery (MI) are easily influenced by the environment and the state of the subject, which exhibit temporal and spatial variance. And this variance is more significant across subjects and sessions, which imposes limitations on the cross-domain MI tasks. To…

Cited by 0SourceScholar
2021

Asymmetric Gained Deep Image Compression With Continuous Rate Adaptation

CVPR 2021poster

With the development of deep learning techniques, the combination of deep learning with image compression has drawn lots of attention. Recently, learned image compression methods had exceeded their classical counterparts in terms of rate-distortion performance. However, continuous rate adaptation re…

Cited by 162PDFScholar
2021

From Two to One: A New Scene Text Recognizer With Visual Language Modeling Network

ICCV 2021poster

In this paper, we abandon the dominant complex language model and rethink the linguistic learning process in the scene text recognition. Different from previous methods considering the visual and linguistic information in two separate structures, we propose a Visual Language Modeling Network (Vision…

Cited by 184PDFcodeScholar
2021

Generalization Bounds for Graph Embedding Using Negative Sampling: Linear vs Hyperbolic

NeurIPS 2021poster

Graph embedding, which represents real-world entities in a mathematical space, has enabled numerous applications such as analyzing natural languages, social networks, biochemical networks, and knowledge bases. It has been experimentally shown that graph embedding in hyperbolic space can represent hi…

Cited by 12SourcePDFScholar
2021

Generalization Error Bound for Hyperbolic Ordinal Embedding

ICML 2021spotlight

Hyperbolic ordinal embedding (HOE) represents entities as points in hyperbolic space so that they agree as well as possible with given constraints in the form of entity $i$ is more similar to entity $j$ than to entity $k$. It has been experimentally shown that HOE can obtain representations of hiera…

Cited by 14SourcePDFScholar
2021

Improving OCR-Based Image Captioning by Incorporating Geometrical Relationship

CVPR 2021poster

OCR-based image captioning aims to automatically describe images based on all the visual entities (both visual objects and scene text) in images. Compared with conventional image captioning, the reasoning of scene text is required for OCR-based image captioning since the generated descriptions often…

Cited by 50PDFcodeScholar
2021

Learning To Filter: Siamese Relation Network for Robust Tracking

CVPR 2021poster

Despite the great success of Siamese-based trackers, their performance under complicated scenarios is still not satisfying, especially when there are distractors. To this end, we propose a novel Siamese relation network, which introduces two efficient modules, i.e. Relation Detector (RD) and Refinem…

Cited by 146PDFcodeScholar
2021

SalientSleepNet: Multimodal Salient Wave Detection Network for Sleep Staging

IJCAI 2021poster

Sleep staging is fundamental for sleep assessment and disease diagnosis. Although previous attempts to classify sleep stages have achieved high classification performance, several challenges remain open: 1) How to effectively extract salient waves in multimodal sleep data; 2) How to capture the mult…

2021

Scene Text Retrieval via Joint Text Detection and Similarity Learning

CVPR 2021poster

Scene text retrieval aims to localize and search all text instances from an image gallery, which are the same or similar with a given query text. Such a task is usually realized by matching a query text to the recognized words, outputted by an end-to-end scene text spotter. In this paper, we address…

Cited by 47PDFcodeScholar
2021

Self-Supervised Learning for Sleep Stage Classification with Predictive and Discriminative Contrastive Coding

ICASSP 2021accepted

The purpose of this paper is to learn efficient representations from raw electroencephalogram (EEG) signals for sleep stage classification via self-supervised learning (SSL). Although supervised methods have gained favorable performance, they heavily rely on manually labeled datasets. Recently, SSL…

Cited by 0SourceScholar
2020

Discovering Latent Class Labels for Multi-Label Learning

IJCAI 2020poster

Existing multi-label learning (MLL) approaches mainly assume all the labels are observed and construct classification models with a fixed set of target labels (known labels). However, in some real applications, multiple latent labels may exist outside this set and hide in the data, especially for la…

Cited by 0SourcePDFScholar
2020

GraphSleepNet: Adaptive Spatial-Temporal Graph Convolutional Networks for Sleep Stage Classification

IJCAI 2020poster

Sleep stage classification is essential for sleep assessment and disease diagnosis. However, how to effectively utilize brain spatial features and transition information among sleep stages continues to be challenging. In particular, owing to the limited knowledge of the human brain, predefining a su…

2020

Knowledge Consistency between Neural Networks and Beyond

ICLR 2020poster

This paper aims to analyze knowledge consistency between pre-trained deep neural networks. We propose a generic definition for knowledge consistency between neural networks at different fuzziness levels. A task-agnostic method is designed to disentangle feature components, which represent the consis…

Cited by 41SourceScholar
2018

Practical Considerations of a BMI Application for Detecting Acute Pain Signals

ICASSP 2018accepted

Brain-machine interfaces (BMIs) have been an important research area in closed-loop neuroscience and neuroengineering. In real-time neuroscience applications, many issues require special consideration, such as trial variability, spike sorting noise or multi-unit activity. For a BMI application of de…

Cited by 0SourceScholar
2016

Walk and Learn: Facial Attribute Representation Learning From Egocentric Video and Contextual Data

CVPR 2016oral

The way people look in terms of facial attributes (ethnicity, hair color, facial hair, etc.) and the clothes or accessories they wear (sunglasses, hat, hoodies, etc.) is highly dependent on geo-location and weather condition, respectively. This work explores, for the first time, the use of this cont…

Cited by 137PDFScholar