← Search

Yichi Zhang

79 accepted papers

2026

D-FUSEr: Diverse Failure, Unified Success via Error-Distribution Shaping in LLM Reasoning

ICML 2026poster

Test-time scaling methods such as majority vote aggregation and iterative refinement (e.g., self-reflection or multi-agent inference) improve reasoning performance by leveraging multiple solution samples. However, their efficacy depends not only on raw performance, but critically on the distribution…

Cited by 0SourceScholar
2026

Exploring the Basin-Like Loss Landscape in Large Language Models

ICLR 2026poster

We discover the emergence of \textit{basins} in the loss landscape of large language models. As model scale increases, LLMs become progressively more resilient to random perturbations in the parameter space, giving rise to expansive stability regions where models exhibit nearly identical performance…

Cited by 0SourcecodeScholar
2026

Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification

ICML 2026poster

As LLM-powered agents have been used for high-stakes decision-making, such as clinical diagnosis, it becomes critical to develop reliable verification of their decisions to facilitate trustworthy deployment. Yet, existing verifiers usually underperform owing to a lack of domain knowledge and limited…

Cited by 0SourceScholar
2026

Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks

ICML 2026poster

Deeper modern architectures are costly to train, making hyperparameter transfer preferable to expensive repeated tuning. Maximal Update Parametrization ($\mu$P) helps explain why many hyperparameters transfer across width. Yet depth scaling is less understood for modern architectures, whose computat…

Cited by 0SourceScholar
2026

MESA: Improving MoE Safety Alignment via Decentralized Expertise

ICML 2026poster

Mixture-of-Experts (MoE) architectures have emerged as a popular paradigm for scaling Large Language Models (LLMs), enabling greater capacity with reduced computational cost by dynamically routing inputs to the most relevant experts based on learned patterns. However, this also introduces a critical…

Cited by 0SourceScholar
2026

Meta-Router: Bridging Gold-standard and Preference-based Evaluations in LLM Routing

ICLR 2026poster

In language tasks requiring extensive human-model interaction, the inference cost of large language models (LLMs) can be substantial. To reduce expenses while preserving the quality of the responses, an LLM router selects among candidate models to balance between the expected response quality and t…

Cited by 0SourcecodeScholar
2026

PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography

AAAI 2026technical

Positron emission tomography (PET) is a cornerstone of modern oncologic and neurologic imaging, distinguished by its unique ability to illuminate dynamic metabolic processes that transcend the anatomical focus of traditional imaging technologies. Radiology reports are essential for clinical decision

Cited by 0SourcePDFScholar
2026

Pixel Motion Diffusion is What We Need for Robot Control

CVPR 2026

We present DAWN (Diffusion is All We Need for robot control), a unified diffusion-based framework for language-conditioned robotic manipulation that bridges high-level motion intent and low-level robot action via structured pixel motion representation. In DAWN, both the high-level and low-level cont

Cited by 0SourcecodeScholar
2026

Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective

ICML 2026poster

Recent studies have shown that semantic watermarks, which embed information into the initial noise of latent diffusion models (LDMs), are vulnerable to black-box forgery attacks. However, existing methods primarily rely on empirical evidence and lack a rigorous theoretical understanding of the condi…

Cited by 0SourceScholar
2026

Robust Training of Neural Networks at Arbitrary Precision and Sparsity

ICLR 2026poster

The discontinuous operations inherent in quantization and sparsification introduce a long-standing obstacle to backpropagation, particularly in ultra-low precision and sparse regimes. While the community has long viewed quantization as unfriendly to gradient descent due to its lack of smoothness, we…

Cited by 0SourceScholar
2026

Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention

ICLR 2026poster

Although Large Reasoning Models (LRMs) have progressed in solving complex problems, their chain-of-thought (CoT) reasoning often contains harmful content that can persist even when the final responses appear safe. We show that this issue still remains in existing methods which overlook the unique si…

Cited by 0SourceScholar
2026

UniHR: Hierarchical Representation Learning for Unified Knowledge Graph Link Prediction

AAAI 2026technical

Real-world knowledge graphs (KGs) contain not only standard triple-based facts, but also more complex, heterogeneous types of facts, such as hyper-relational facts with auxiliary key-value pairs, temporal facts with additional timestamps, and nested facts that imply relationships between facts. Thes

Cited by 0SourcePDFScholar
2026

rMMEA: Robust Multi-Modal Entity Alignment with Missing and Noise Visual Modality

AAAI 2026technical

Recently, multi-modal embedding methods have flourished in entity alignment. As state-of-the-art approaches evolve rapidly, visual modality (i.e., images) missing emerges as a critical challenge. While visual modality typically offers the most informative signals in multi-modal entity alignment (MME

Cited by 0SourcePDFScholar
2025

A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation

ICLR 2025poster

This work tackles the information loss bottleneck of vector-quantization (VQ) autoregressive image generation by introducing a novel model architecture called the 2-Dimensional Autoregression (DnD) Transformer. The DnD-Transformer predicts more codes for an image by introducing a new direction, **mo…

2025

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies

CoRL 2025poster

In this paper, we propose AimBot, a lightweight visual augmentation technique that provides explicit spatial cues to improve visuomotor policy learning in robotic manipulation. AimBot overlays shooting lines and scope reticles onto multi-view RGB images, offering auxiliary visual guidance that encod…

Cited by 0SourcecodeScholar
2025

Atomic Diffusion Models for Small Molecule Structure Elucidation from NMR Spectra

NeurIPS 2025poster

Nuclear Magnetic Resonance (NMR) spectroscopy is a cornerstone technique for determining the structures of small molecules and is especially critical in the discovery of novel natural products and clinical therapeutics. Yet, interpreting NMR spectra remains a time-consuming, manual process requiring…

Cited by 0SourceScholar
2025

Balanced Rate-Distortion Optimization in Learned Image Compression

CVPR 2025highlight

Learned image compression (LIC) using deep learning architectures has seen significant advancements, yet standard rate-distortion (R-D) optimization often encounters imbalanced updates due to diverse gradients of the rate and distortion objectives. This imbalance can lead to suboptimal optimization,…

2025

Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens

NeurIPS 2025poster

We present **Brain Harmony (BrainHarmonix)**, the first multimodal brain foundation model that unifies structural morphology and functional dynamics into compact 1D token representations. The model was pretrained on two of the largest neuroimaging datasets to date, encompassing 64,594 T1-weighted s…

Cited by 0SourceScholar
2025

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space

ACL 2025finding

Large Language Models (LLMs), despite advanced general capabilities, still suffer from numerous safety risks, especially jailbreak attacks that bypass safety protocols. Understanding these vulnerabilities through black-box jailbreak attacks, which better reflect real-world scenarios, offers critical…

2025

CogAtom: From Cognitive Atoms to Olympiad-level Mathematical Reasoning in Large Language Models

EMNLP 2025

Mathematical reasoning poses significant challenges for Large Language Models (LLMs) due to its demand for multi-step reasoning and abstract conceptual integration. While recent test-time scaling techniques rely heavily on high-quality, challenging problems, the scarcity of Olympiad-level math probl

2025

DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios

NeurIPS 2025poster

Despite the remarkable advances of Large Language Models (LLMs) across diverse cognitive tasks, the rapid enhancement of these capabilities also introduces emergent deception behaviors that may induce severe risks in high-stakes deployments. More critically, the characterization of deception across…

Cited by 0SourcecodeScholar
2025

Exploring the Generalizability of Factual Hallucination Mitigation via Enhancing Precise Knowledge Utilization

EMNLP 2025

Large Language Models (LLMs) often struggle to align their responses with objective facts, resulting in the issue of factual hallucinations , which can be difficult to detect and mislead users without relevant knowledge. Although post-training techniques have been employed to mitigate the issue, exi

2025

FIG: Flow with Interpolant Guidance for Linear Inverse Problems

ICLR 2025poster

Diffusion and flow matching models have recently been used to solve various linear inverse problems in image restoration, such as super-resolution and inpainting. Using a pre-trained diffusion or flow-matching model as a prior, most existing methods modify the reverse-time sampling process by incorp…

2025

Have We Designed Generalizable Structural Knowledge Promptings? Systematic Evaluation and Rethinking

ACL 2025long

Large language models (LLMs) have demonstrated exceptional performance in text generation within current NLP research. However, the lack of factual accuracy is still a dark cloud hanging over the LLM skyscraper. Structural knowledge prompting (SKP) is a prominent paradigm to integrate external knowl…

2025

Improve Representation for Imbalanced Regression through Geometric Constraints

CVPR 2025poster

In representation learning, uniformity refers to the uniform feature distribution in the latent space (i.e., unit hypersphere). Previous work has shown that improving uniformity contributes to the learning of under-represented classes. However, most of the previous work focused on classification; th…

2025

K-ON: Stacking Knowledge on the Head Layer of Large Language Model

AAAI 2025technical

Recent advancements in large language models (LLMs) have significantly improved various natural language processing (NLP) tasks. Typically, LLMs are trained to predict the next token, aligning well with many NLP tasks. However, in knowledge graph (KG) scenarios, entities are the fundamental units an…

Cited by 0SourcePDFScholar
2025

Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance

EMNLP 2025

Large vision-language models (LVLMs) have achieved impressive results in vision-language tasks. However, Therefore, we propose LACING, designed to address such bias with Mu ̲ L timodal Du ̲ A l-attention Me ̲ C han ̲ I sm (MDA) a ̲ N d Soft-Image ̲ G uidance (SIG). Specifically, MDA adopts a paralle

Cited by 0SourcePDFScholar
2025

MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

ICLR 2025poster

Multi-modal Large Language Models (MLLMs) have recently showcased superior proficiency in general visual scenarios. However, we identify their mathematical capabilities remain under-explored with three areas to be improved: visual encoding of math diagrams, diagram-language alignment, and chain-of-t…

2025

MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

NAACL 2025long

Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, often assessed through multiple-choice questions (MCQs) that include an image, a question, and several options. However, many benchmarks used for such evaluations suffer from systematic biases. Remar…

2025

Mitigating Overthinking in Large Reasoning Models via Manifold Steering

NeurIPS 2025poster

Recent advances in Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in solving complex tasks such as mathematics and coding. However, these models frequently exhibit a phenomenon known as *overthinking* during inference, characterized by excessive validation loops and redundan…

Cited by 0SourcecodeScholar
2025

Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning

ICLR 2025poster

Learning high-quality multi-modal entity representations is an important goal of multi-modal knowledge graph (MMKG) representation learning, which can en- hance reasoning tasks within the MMKGs, such as MMKG completion (MMKGC). The main challenge is to collaboratively model the structural informatio…

2025

Noise-powered Multi-modal Knowledge Graph Representation Framework

COLING 2025main

The rise of Multi-modal Pre-training highlights the necessity for a unified Multi-Modal Knowledge Graph (MMKG) representation learning framework. Such a framework is essential for embedding structured knowledge into multi-modal Large Language Models effectively, alleviating issues like knowledge mis…

2025

Proactive Assistant Dialogue Generation from Streaming Egocentric Videos

EMNLP 2025

Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactive, proactive assistance based on streaming visual inputs, yet their development is constrained by the costly and labor-i

Cited by 0SourcePDFScholar
2025

RL-Guider: Leveraging Historical Decisions and Feedback for Drug Editing with Large Language Models

ACL 2025finding

Recent success of large language models (LLMs) in diverse domains showcases their potential to revolutionize scientific fields, including drug editing. Traditional drug editing relies on iterative conversations with domain experts, refining the drug until the desired property is achieved. This inter…

2025

STAIR: Improving Safety Alignment with Introspective Reasoning

ICML 2025oral

Ensuring the safety and harmlessness of Large Language Models (LLMs) has become equally critical as their performance in applications. However, existing safety alignment methods typically suffer from safety-performance trade-offs and susceptibility to jailbreak attacks, primarily due to their relian…

2025

SegAnyPET: Universal Promptable Segmentation from Positron Emission Tomography Images

ICCV 2025poster

Positron Emission Tomography (PET) is a powerful molecular imaging tool that plays a crucial role in modern medical diagnostics by visualizing radio-tracer distribution to reveal physiological processes. Accurate organ segmentation from PET images is essential for comprehensive multi-systemic analys…

2025

Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation

AAAI 2025technical

Multi-modal knowledge graph completion (MMKGC) aims to discover unobserved knowledge from given multi-modal knowledge graphs (MMKG), collaboratively leveraging structural information from the triples and multi-modal information of the entities to overcome the inherent incompleteness. Existing MMKGC…

2025

Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors

ACL 2025finding

Intelligent tutoring agents powered by large language models (LLMs) have been increasingly explored to deliver personalized knowledge in areas such as language learning and science education. However, their capabilities in guiding users to solve complex real-world tasks remain underexplored. To addr…

2024

"SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models"

ECCV 2024poster

"We present , a versatile multi-modal large language model (MLLM) with a joint mixing of model weights, visual embeddings and image scales. First, for stronger vision-language alignment, we unfreeze the large language model (LLM) during pre-training, and introduce a weight mix strategy between LLMs…

2024

Another Way to the Top: Exploit Contextual Clustering in Learned Image Coding

AAAI 2024technical

While convolution and self-attention are extensively used in learned image compression (LIC) for transform coding, this paper proposes an alternative called Contextual Clustering based LIC (CLIC) which primarily relies on clustering operations and local attention for correlation characterization and…

Cited by 8SourcePDFScholar
2024

Eliciting Honest Information from Authors Using Sequential Review

AAAI 2024technical

In the setting of conference peer review, the conference aims to accept high-quality papers and reject low-quality papers based on noisy review scores. A recent work proposes the isotonic mechanism, which can elicit the ranking of paper qualities from an author with multiple submissions to help impr…

2024

Exploring the Transferability of Visual Prompting for Multimodal Large Language Models

CVPR 2024highlight

Although Multimodal Large Language Models (MLLMs) have demonstrated promising versatile capabilities their performance is still inferior to specialized models on downstream tasks which makes adaptation necessary to enhance their utility. However fine-tuning methods require independent training for e…

2024

GROUNDHOG: Grounding Large Language Models to Holistic Segmentation

CVPR 2024poster

Most multimodal large language models (MLLMs) learn language-to-object grounding through causal language modeling where grounded objects are captured by bounding boxes as sequences of location tokens. This paradigm lacks pixel-level representations that are important for fine-grained visual understa…

Cited by 47SourcePDFScholar
2024

Gliding over the Pareto Front with Uniform Designs

NeurIPS 2024poster

Multiobjective optimization (MOO) plays a critical role in various real-world domains. A major challenge therein is generating $K$ uniform Pareto-optimal solutions to represent the entire Pareto front. To address this issue, this paper firstly introduces \emph{fill distance} to evaluate the $K$ desi…

Cited by 2SourcePDFScholar
2024

Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering

ACL 2024findings

Deploying large language models (LLMs) to real scenarios for domain-specific question answering (QA) is a key thrust for LLM applications, which poses numerous challenges, especially in ensuring that responses are both accommodating to user requirements and appropriately leveraging domain-specific k…

2024

MKGL: Mastery of a Three-Word Language

NeurIPS 2024spotlight

Large language models (LLMs) have significantly advanced performance across a spectrum of natural language processing (NLP) tasks. Yet, their application to knowledge graphs (KGs), which describe facts in the form of triplets and allow minimal hallucinations, remains an underexplored frontier. In th…

Cited by 1SourcePDFScholar
2024

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

ECCV 2024poster

"The remarkable progress of Multi-modal Large Language Models (MLLMs) has gained unparalleled attention. However, their capabilities in visual math problem-solving remain insufficiently evaluated and understood. We investigate current benchmarks to incorporate excessive visual content within textual…

2024

MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models

NeurIPS 2024poster

Despite the superior capabilities of Multimodal Large Language Models (MLLMs) across diverse tasks, they still face significant trustworthiness challenges. Yet, current literature on the assessment of trustworthy MLLMs remains limited, lacking a holistic evaluation to offer thorough insights into fu…

Cited by 5SourcecodeScholar
2024

PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

ACL 2024findings

We present PCA-Bench, a multimodal decision-making benchmark for evaluating the integrated capabilities of Multimodal Large Language Models (MLLMs). Departing from previous benchmarks focusing on simplistic tasks and individual model capability, PCA-Bench introduces three complex scenarios: autonomo…

2024

PINNacle: A Comprehensive Benchmark of Physics-Informed Neural Networks for Solving PDEs

NeurIPS 2024poster

While significant progress has been made on Physics-Informed Neural Networks (PINNs), a comprehensive comparison of these methods across a wide range of Partial Differential Equations (PDEs) is still lacking. This study introduces PINNacle, a benchmarking tool designed to fill this gap. PINNacle pro…

2024

Rethinking Model Ensemble in Transfer-based Adversarial Attacks

ICLR 2024poster

It is widely recognized that deep learning models lack robustness to adversarial examples. An intriguing property of adversarial examples is that they can transfer across different models, which enables black-box attacks without any knowledge of the victim model. An effective strategy to improve the…

2024

SC-NeuS: Consistent Neural Surface Reconstruction from Sparse and Noisy Views

AAAI 2024technical

The recent neural surface reconstruction approaches using volume rendering have made much progress by achieving impressive surface reconstruction quality, but are still limited to dense and highly accurate posed views. To overcome such drawbacks, this paper pays special attention on the consistent s…

2024

Unleashing the Power of Imbalanced Modality Information for Multi-modal Knowledge Graph Completion

COLING 2024main

Multi-modal knowledge graph completion (MMKGC) aims to predict the missing triples in the multi-modal knowledge graphs by incorporating structural, visual, and textual information of entities into the discriminant models. The information from different modalities will work together to measure the tr…

2023

Binarized Neural Machine Translation

NeurIPS 2023poster

The rapid scaling of language models is motivating research using low-bitwidth quantization. In this work, we propose a novel binarization technique for Transformers applied to machine translation (BMT), the first of its kind. We identify and address the problem of inflated dot-product variance when…

2023

Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework

ACL 2023findings

Despite recent success on various tasks, deep learning techniques still perform poorly on adversarial examples with small perturbations. While optimization-based methods for adversarial attacks are well-explored in the field of computer vision, it is impractical to directly apply them in natural lan…

2023

Can Foundation Models Watch, Talk and Guide You Step by Step to Make a Cake?

EMNLP 2023long findings

Despite tremendous advances in AI, it remains a significant challenge to develop interactive task guidance systems that can offer situated, personalized guidance and assist humans in various tasks. These systems need to have a sophisticated understanding of the user as well as the environment, and…

Cited by 0SourcecodeScholar
2023

Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?

EMNLP 2023long main

Vision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world. However, known as visual illusions, human's perception of reality isn't always faithful to the physical world. This raises a key question: do VLMs have the similar kind of i…

Cited by 0SourcecodeScholar
2023

Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio Representations

ICASSP 2023accepted

This work investigates pretrained audio representations for few shot Sound Event Detection. We specifically address the task of few shot detection of novel acoustic sequences, or sound events with semantically meaningful temporal structure, without assuming access to non-target audio. We develop pro…

Cited by 0SourceScholar
2023

Revisiting the Evaluation of Image Synthesis with GANs

NeurIPS 2023poster

A good metric, which promises a reliable comparison between solutions, is essential for any well-defined task. Unlike most vision tasks that have per-sample ground-truth, image synthesis tasks target generating unseen data and hence are usually evaluated through a distributional distance between one…

2023

Understanding the Robustness of 3D Object Detection With Bird's-Eye-View Representations in Autonomous Driving

CVPR 2023poster

3D object detection is an essential perception task in autonomous driving to understand the environments. The Bird's-Eye-View (BEV) representations have significantly improved the performance of 3D detectors with camera inputs on popular benchmarks. However, there still lacks a systematic understand…

2022

DANLI: Deliberative Agent for Following Natural Language Instructions

EMNLP 2022main

Recent years have seen an increasing amount of work on embodied AI agents that can perform tasks by following human language instructions. However, most of these agents are reactive, meaning that they simply learn and imitate behaviors encountered in the training data. These reactive agents are insu…

2022

Understanding Hyperdimensional Computing for Parallel Single-Pass Learning

NeurIPS 2022accept

Hyperdimensional computing (HDC) is an emerging learning paradigm that computes with high dimensional binary vectors. There is an active line of research on HDC in the community of emerging hardware because of its energy efficiency and ultra-low latency---but HDC suffers from low model accuracy, wit…

2021

BulletTrain: Accelerating Robust Neural Network Training via Boundary Example Mining

NeurIPS 2021poster

Neural network robustness has become a central topic in machine learning in recent years. Most training algorithms that improve the model's robustness to adversarial and common corruptions also introduce a large computational overhead, requiring as many as ten times the number of forward and backwar…

Cited by 21SourcePDFScholar
2021

Drop Redundant, Shrink Irrelevant: Selective Knowledge Injection for Language Pretraining

IJCAI 2021poster

Previous research has demonstrated the power of leveraging prior knowledge to improve the performance of deep models in natural language processing. However, traditional methods neglect the fact that redundant and irrelevant knowledge exists in external knowledge bases. In this study, we launched an…

Cited by 34SourcePDFScholar
2021

Interpretable and Low-Resource Entity Matching via Decoupling Feature Learning from Decision Making

ACL 2021long

Entity Matching (EM) aims at recognizing entity records that denote the same real-world object. Neural EM models learn vector representation of entity descriptions and match entities end-to-end. Though robust, these methods require many annotated resources for training, and lack of interpretability.…

2021

Is Multi-Hop Reasoning Really Explainable? Towards Benchmarking Reasoning Interpretability

EMNLP 2021main

Multi-hop reasoning has been widely studied in recent years to obtain more interpretable link prediction. However, we find in experiments that many paths given by these models are actually unreasonable, while little work has been done on interpretability evaluation for them. In this paper, we propos…

2021

Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-Modal Pretraining

ICCV 2021poster

Nowadays, customer's demands for E-commerce are more diversified, which introduces more complications to the product retrieval industry. Previous methods are either subject to single-modal input or perform supervised image-level product retrieval, thus fail to accommodate real-life scenarios where e…

Cited by 75PDFcodeScholar
2021

Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding

EMNLP 2021finding

Large-scale, pre-trained language models (LMs) have achieved human-level performance on a breadth of language understanding tasks. However, evaluations only based on end task performance shed little light on machines’ true ability in language understanding and reasoning. In this paper, we highlight…

2020

Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision Activations

ICLR 2020poster

We propose precision gating (PG), an end-to-end trainable dynamic dual-precision quantization technique for deep neural networks. PG computes most features in a low precision and only a small proportion of important features in a higher precision to preserve accuracy. The proposed approach is appl…

Cited by 33SourcecodeScholar
2019

Seq2Seq Attentional Siamese Neural Networks for Text-dependent Speaker Verification

ICASSP 2019accepted

In this paper, we present a Sequence-to-Sequence Attentional Siamese Neural Network (Seq2Seq-ASNN) that leverages temporal alignment information for end-to-end speaker verification. In prior works of speaker discriminative neural networks, utterance-level evaluation/enrollment speaker representation…

Cited by 0SourceScholar
2018

Joint Speaker Diarization and Recognition Using Convolutional and Recurrent Neural Networks

ICASSP 2018accepted

Speaker diarization (detecting who-spoke-when using relative identity labels) and speaker recognition (detecting absolute identity labels without timing) are different but related tasks that often need to be completed simultaneously in many scenarios. Traditional methods, however, address them indep…

Cited by 0SourceScholar
2018

Visualization and Interpretation of Siamese Style Convolutional Neural Networks for Sound Search by Vocal Imitation

ICASSP 2018accepted

Designing systems that allow users to search sounds through vocal imitation augments the current text-based search engines and advances human-computer interaction. Previously we proposed a Siamese style convolutional network called IMINET for sound search by vocal imitation, which jointly addresses…

Cited by 0SourceScholar