← Search

Zhuo Chen

156 accepted papers

2026

Balancing Marker and Markerless Modes in Vision-Based Tactile Sensors with a Translucent Skin

ICRA 2026poster

Vision-based tactile sensors (VBTS) face an inherent trade-off in tactile skin design. Opaque ink markers enable accurate force and tangential displacement estimation but occlude geometric features essential for object and texture classification. Conversely, markerless skins preserve surface details…

Cited by 0Scholar
2026

Dens3R: A Foundation Model for 3D Geometry Prediction

ICLR 2026poster

Recent advances in dense 3D reconstruction have led to significant progress, yet achieving accurate unified geometric prediction remains a major challenge. Most existing methods are limited to predicting a single geometry quantity from input images. However, geometric quantities such as depth, surfa…

Cited by 0SourcecodeScholar
2026

Direct Preference Optimization for Speech Autoregressive Diffusion Models

ICASSP 2026poster

Autoregressive diffusion models (ARDMs) have recently been applied to speech generation, achieving state-of-the-art (SOTA) performance in zero-shot text-to-speech. By autoregressively generating continuous speech tokens with next-token diffusion, these models offer a promising alternative to next-to…

Cited by 0SourcePDFScholar
2026

Force-Aware 3D Contact Modeling for Stable Grasp Generation

AAAI 2026technical

Contact-based grasp generation plays a crucial role in various applications. Recent methods typically focus on the geometric structure of objects, producing grasps with diverse hand poses and plausible contact points. However, these approaches often overlook the physical attributes of the grasp, spe

Cited by 0SourcePDFScholar
2026

From Natural Alignment to Conditional Controllability in Multimodal Dialogue

ICLR 2026poster

The recent advancement of Artificial Intelligence Generated Content (AIGC) has led to significant strides in modeling human interaction, particularly in the context of multi-modal dialogue. While current methods impressively generates realistic dialogue in speech and vision modalities, challenges r…

Cited by 0SourcecodeScholar
2026

L-CUBE: Isolating Long-Context Capacity from Knowledge with Controllable Mutual Information Scaling

ICML 2026poster

Evaluating long-context language models on natural language conflates architectural capacity to capture dependencies with semantic knowledge and vocabulary statistics. When models fail at long contexts, we cannot determine whether failures stem from fundamental architectural limitations or insuffici…

Cited by 0SourceScholar
2026

LABO: LLM-Accelerated Bayesian Optimization through Broad Exploration and Selective Experimentation

ICML 2026poster

The high cost and data scarcity in scientific exploration have motivated the use of large language models (LLMs) as knowledge-driven components in Bayesian optimization (BO). However, existing approaches typically embed LLMs directly into the sampling or surrogate modeling pipeline, without fully le…

Cited by 0SourceScholar
2026

MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis

CVPR 2026

Physically-based rendering (PBR) materials are fundamental to photorealistic graphics, yet their creation remains labor-intensive and requires specialized expertise. While generative models have advanced material synthesis, existing methods lack a unified representation bridging natural image appear

Cited by 0SourceScholar
2026

Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation

CVPR 2026

Reinforcement learning (RL) has demonstrated remarkable success in text and image generation, yet its potential in 3D generation remains largely unexplored. Existing attempts typically rely on offline direct preference optimization (DPO) method, which suffers from low training efficiency and limited

Cited by 0SourceScholar
2026

MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts

CVPR 2026

Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks.In 3D visual geometry reconstruction, large-scale training has likewise proven effective for learning versatile representations.However, further scaling of 3

Cited by 0SourcecodeScholar
2026

Multi-Modal Fact Knowledge Generation for Imbalanced Cross-Source Entity Alignment

AAAI 2026technical

Multi-modal imbalanced cross-source entity alignment aims to identify equivalent entity pairs across multi-modal knowledge graphs (MMKGs) that encompass diverse data sources with imbalanced modality, which poses significant challenges due to the non-uniform distribution of information across differe

Cited by 0SourcePDFScholar
2026

Native-Domain Cross-Attention for Camera-LiDAR Extrinsic Calibration Under Large Initial Perturbations

RA-L 2026

Accurate camera–LiDAR fusion relies on precise extrinsic calibration, which fundamentally depends on establishing reliable cross-modal correspondences under potentially large misalignments. Existing learning-based methods typically project LiDAR points into depth maps for feature fusion, which disto

Cited by 0SourceScholar
2026

POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling

CVPR 2026

Face relighting aims to synthesize realistic portraits under novel illumination while preserving identity and geometry. However, progress remains constrained by the limited availability of large-scale, physically consistent illumination data. To address this, we introduce POLAR, a large-scale and ph

Cited by 0SourceScholar
2026

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

ICLR 2026poster

We introduce Part-X-MLLM, a native 3D multimodal large language model that unifies diverse 3D tasks by formulating them as programs in a structured, executable grammar. Given an RGB point cloud and a natural language prompt, our model autoregressively generates a single, coherent token sequence enco…

Cited by 4SourcecodeScholar
2026

PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World

ICML 2026poster

Synthesizing physics-grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI. Existing methods predominantly focus on static geometry, overlooking the functional properties essential for interaction. We propose that interactive asset generation must be rooted in fu…

Cited by 0SourceScholar
2026

QuadGPT: Native Quadrilateral Mesh Generation with Autoregressive Models

ICLR 2026poster

The generation of quadrilateral-dominant meshes is a cornerstone of professional 3D content creation. However, existing generative models generate quad meshes by first generating triangle meshes and then merging triangles into quadrilaterals with some specific rules, which typically produces quad m…

Cited by 0SourceScholar
2026

Repurposing Synthetic Data for Fine-grained Search Agent Supervision

ICLR 2026poster

LLM-based search agents are increasingly trained on entity-centric synthetic data to solve complex, knowledge-intensive tasks. However, prevailing training methods like Group Relative Policy Optimization (GRPO) discard this rich entity information, relying instead on sparse, outcome-based rewards. T…

Cited by 0SourceScholar
2026

Scaling Agents via Continual Pre-training

ICLR 2026poster

Large language models (LLMs) have evolved into agentic systems capable of autonomous tool use and multi-step reasoning for complex problem-solving. However, post-training approaches building upon general-purpose foundation models consistently underperform in agentic tasks, particularly in open-sourc…

Cited by 0SourcecodeScholar
2026

SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action Planning

CVPR 2026

Vision-Language-Action (VLA) models have emerged as a promising paradigm where pretrained Vision-Language Models (VLMs) serve as System 2 for high-level reasoning, connected to action experts as System 1 for low-level motor control.However, current works fail to genuinely leverage VLM capabilities:

Cited by 0SourceScholar
2026

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

ICLR 2026poster

Aligning large generative models with human feedback is a critical challenge. In speech synthesis, this is particularly pronounced due to the lack of a large-scale human preference dataset, which hinders the development of models that truly align with human perception. To address this, we introduce…

Cited by 0SourceScholar
2026

The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy

CVPR 2026

Training-free image editing with large diffusion models has become practical, yet faithfully performing complex non-rigid edits (e.g., pose or shape changes) remains highly challenging. We identify a key underlying cause: attention collapse in existing attention sharing mechanisms, where either posi

Cited by 0SourcecodeScholar
2026

UniHR: Hierarchical Representation Learning for Unified Knowledge Graph Link Prediction

AAAI 2026technical

Real-world knowledge graphs (KGs) contain not only standard triple-based facts, but also more complex, heterogeneous types of facts, such as hyper-relational facts with auxiliary key-value pairs, temporal facts with additional timestamps, and nested facts that imply relationships between facts. Thes

Cited by 0SourcePDFScholar
2026

Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery

ICLR 2026poster

Scientific discovery is increasingly constrained by costly experiments and limited budgets, making efficient optimization essential for AI for science. Bayesian Optimization (BO), while widely adopted for balancing exploration and exploitation, suffers from slow cold-start performance and poor scala…

Cited by 0SourceScholar
2026

ViTacGen: Robotic Pushing with Vision-To-Touch Generation

ICRA 2026poster

Robotic pushing is a fundamental manipulation task that requires tactile feedback to capture subtle contact forces and dynamics between the end-effector and the object. However, real tactile sensors often face hardware limitations and deployment challenges, while vision-only policies struggle with s…

2026

Visual-Tactile Peg-in-Hole Assembly Learning From Peg-Out-of-Hole Disassembly

RA-L 2026

Peg-in-hole (PiH) assembly is a fundamental yet challenging robotic manipulation task. While reinforcement learning (RL) has shown promise in tackling such tasks, it requires extensive exploration. In this paper, we propose a novel visual-tactile skill learning framework for the PiH task that levera

Cited by 0SourceScholar
2026

X-Part: High Fidelity And Structure Coherent Shape Decomposition And Completion

CVPR 2026

Generating 3D shapes at part level is pivotal for downstream applications such as mesh retopology, UV mapping, and 3D printing. However, existing part-based generation methods often lack sufficient controllability and suffer from poor semantically meaningful decomposition. To this end, we introduce

Cited by 0SourcecodeScholar
2026

rMMEA: Robust Multi-Modal Entity Alignment with Missing and Noise Visual Modality

AAAI 2026technical

Recently, multi-modal embedding methods have flourished in entity alignment. As state-of-the-art approaches evolve rapidly, visual modality (i.e., images) missing emerges as a critical challenge. While visual modality typically offers the most informative signals in multi-modal entity alignment (MME

Cited by 0SourcePDFScholar
2025

AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and Segmentation

ICCV 2025poster

The challenge of multimodal semantic segmentation lies in establishing semantically consistent and segmentable multimodal fusion features under conditions of significant visual feature discrepancies. Existing methods commonly construct cross-modal self-attention fusion frameworks or introduce additi…

2025

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment

ACL 2025long

Modern zero-shot text-to-speech (TTS) systems, despite using extensive pre-training, often struggle in challenging scenarios such as tongue twisters, repeated words, code-switching, and cross-lingual synthesis, leading to intelligibility issues. To address these limitations, this paper leverages pre…

2025

AniSDF: Fused-Granularity Neural Surfaces with Anisotropic Encoding for High-Fidelity 3D Reconstruction

ICLR 2025poster

Neural radiance fields have recently revolutionized novel-view synthesis and achieved high-fidelity renderings. However, these methods sacrifice the geometry for the rendering quality, limiting their further applications including relighting and deformation. How to synthesize photo-realistic rende…

Cited by 1SourcePDFScholar
2025

Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference Optimization

NeurIPS 2025poster

We introduce Auto-Connect, a novel approach for automatic rigging that explicitly preserves skeletal connectivity through a connectivity-preserving tokenization scheme. Unlike previous methods that predict bone positions represented as two joints or first predict points before determining connectivi…

Cited by 0SourceScholar
2025

Compact R-X-Y Stage and Dual-Finger Micromanipulator under Inverted Optical Microscope for Microassembly

IROS 2025

Microassembly plays an important role in fabricating complex structures with small basic components in industrial and biomedical fields. Inverted optical microscope could provide high-quality image feedback for microassembly with its continuously improving resolution. However, a compact stage capabl

Cited by 0SourceScholar
2025

Consistency of Physics-Informed Neural Networks for Second-Order Elliptic Equations

NeurIPS 2025poster

The physics-informed neural networks (PINNs) are widely applied in solving differential equations. However, few studies have discussed their consistency. In this paper, we consider the consistency of PINNs when applied to second-order elliptic equations with Dirichlet boundary conditions. We first p…

Cited by 0SourceScholar
2025

Contactless and Economical Chemical Reaction Platform Based on Ultrasonic Field

IROS 2025

Chemical reactions constitute a cornerstone of fundamental scientific inquiry, yet traditional methodologies and platforms are encumbered by excessive reagent and consumable demands. Emerging alternatives, such as microfluidic systems, while innovative, suffer from intricate fabrication processes an

Cited by 0SourceScholar
2025

DO-CoLM: Dynamic 3D Conformation Relationships Capture with Self-Adaptive Ordering Molecular Relational Modeling in Language Models

IJCAI 2025

Molecular Relational Learning (MRL) aims to understand interactions between molecular pairs, playing a critical role in advancing biochemical research. Recently, Large Language Models (LLMs), with their extensive knowledge bases and advanced reasoning capabilities, have emerged as powerful tools for

Cited by 0SourcePDFScholar
2025

Dataset Distillation as Data Compression: A Rate-Utility Perspective

ICCV 2025poster

Driven by the "scale-is-everything" paradigm, modern machine learning increasingly demands ever-larger datasets and models, yielding prohibitive computational and storage requirements. Dataset distillation mitigates this by compressing an original dataset into a small set of synthetic samples, while…

Cited by 0SourcePDFScholar
2025

Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference

EMNLP 2025

Despite the advancements made in Vision Large Language Models (VLLMs), like text Large Language Models (LLMs), they have limitations in addressing questions that require real-time information or are knowledge-intensive. Indiscriminately adopting Retrieval Augmented Generation (RAG) techniques is an

2025

DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation

ICML 2025poster

Several recent studies have attempted to autoregressively generate continuous speech representations without discrete speech tokens by combining diffusion and autoregressive models, yet they often face challenges with excessive computational loads or suboptimal outcomes. In this work, we propose Dif…

Cited by 1SourcePDFScholar
2025

ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering

AAAI 2025technical

The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, existing methods still have some limitations: 1) repetitions, transpositions, and omissions in the output synthesized speec…

2025

Enhanced Rolling Motion of Magnetic Microparticles by Turning Interface Lubrication

IROS 2025

Micro-nano robots must break the symmetry of the flow field to generate net displacement in the low Reynolds number environment. The spherical micro-robots utilize the frictional forces generated through interaction with the surface. We designed a magnetic microroller robot powered by the rotating A

Cited by 0SourceScholar
2025

ExpTalk: Diverse Emotional Expression via Adaptive Disentanglement and Refined Alignment for Speech-Driven 3D Facial Animation

IJCAI 2025

Speech-driven 3D facial animation aims to create lifelike facial expressions that synchronize accurately with speech. Despite significant progress, many existing methods may focus on generating facial animation with a fixed emotional state, neglecting the diverse transformations of facial emotions u

Cited by 0SourcePDFScholar
2025

FreeMesh: Boosting Mesh Generation with Coordinates Merging

ICML 2025poster

The next-coordinate prediction paradigm has emerged as the de facto standard in current auto-regressive mesh generation methods. Despite their effectiveness, there is no efficient measurement for the various tokenizers that serialize meshes into sequences. In this paper, we introduce a new metric P…

Cited by 0SourcePDFScholar
2025

Graph-guided Cross-composition Feature Disentanglement for Compositional Zero-shot Learning

ACL 2025finding

Disentanglement of visual features of primitives (i.e., attributes and objects) has shown exceptional results in Compositional Zero-shot Learning (CZSL). However, due to the feature divergence of an attribute (resp. object) when combined with different objects (resp. attributes), it is challenging t…

2025

Have We Designed Generalizable Structural Knowledge Promptings? Systematic Evaluation and Rethinking

ACL 2025long

Large language models (LLMs) have demonstrated exceptional performance in text generation within current NLP research. However, the lack of factual accuracy is still a dark cloud hanging over the LLM skyscraper. Structural knowledge prompting (SKP) is a prominent paradigm to integrate external knowl…

2025

Infer the Whole from a Glimpse of a Part: Keypoint-Based Knowledge Graph for Vehicle Re-Identification

AAAI 2025technical

Vehicle re-identification aims to match vehicles across non-overlapping camera views. Many existing methods extract features from one specific image, and these methods lack view-invariance when comparing vehicles of different orientations. As a result, discriminative parts obscured by viewpoint chan…

Cited by 0SourcePDFScholar
2025

K-ON: Stacking Knowledge on the Head Layer of Large Language Model

AAAI 2025technical

Recent advancements in large language models (LLMs) have significantly improved various natural language processing (NLP) tasks. Typically, LLMs are trained to predict the next token, aligning well with many NLP tasks. However, in knowledge graph (KG) scenarios, entities are the fundamental units an…

Cited by 0SourcePDFScholar
2025

KBM: Delineating Knowledge Boundary for Adaptive Retrieval in Large Language Models

EMNLP 2025

Large Language Models (LLMs) often struggle with dynamically changing knowledge and handling unknown static information. Retrieval-Augmented Generation (RAG) is employed to tackle these challenges and has a significant impact on improving LLM performance. In fact, we find that not all questions need

2025

L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling

NeurIPS 2025poster

We present a universal theoretical framework for understanding *long-context language modeling* based on a *bipartite* mutual information scaling law that we rigorously verify in natural language. We demonstrate that bipartite mutual information captures multi-token interactions distinct from and sc…

Cited by 0SourcecodeScholar
2025

Language Model Can Listen While Speaking

AAAI 2025technical

Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM), have significantly enhanced speech-based conversational AI. However, these models are limited to turn-based conversation, lacking the ability to interact with humans i…

Cited by 2SourcePDFScholar
2025

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

NeurIPS 2025poster

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through ite…

Cited by 0SourcecodeScholar
2025

Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning

NeurIPS 2025spotlight

Existing pretrained models for 3D mesh generation often suffer from data biases and produce low-quality results, while global reinforcement learning (RL) methods rely on object-level rewards that struggle to capture local structure details. To address these challenges, we present $\textbf{Mesh-RFT}$…

Cited by 0SourceScholar
2025

ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models

NeurIPS 2025poster

Molecular Relational Learning (MRL) aims to understand interactions between molecular pairs, playing a critical role in advancing biochemical research. With the recent development of large language models (LLMs), a growing number of studies have explored the integration of MRL with LLMs and achieved…

Cited by 0SourceScholar
2025

Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation

ICCV 2025poster

Generating sewing patterns in garment design is receiving increasing attention due to its CG-friendly and flexible-editing nature. Previous sewing pattern generation methods have been able to produce exquisite clothing, but struggle to design complex garments with detailed control. To address these…

Cited by 0SourcePDFScholar
2025

Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning

ICLR 2025poster

Learning high-quality multi-modal entity representations is an important goal of multi-modal knowledge graph (MMKG) representation learning, which can en- hance reasoning tasks within the MMKGs, such as MMKG completion (MMKGC). The main challenge is to collaboratively model the structural informatio…

2025

Noise-powered Multi-modal Knowledge Graph Representation Framework

COLING 2025main

The rise of Multi-modal Pre-training highlights the necessity for a unified Multi-Modal Knowledge Graph (MMKG) representation learning framework. Such a framework is essential for embedding structured knowledge into multi-modal Large Language Models effectively, alleviating issues like knowledge mis…

2025

On-Chip Dynamic Mechanical Characterization: from Cells to Nucleus

IROS 2025

Traditional single-cell mechanical characterization techniques (e.g., atomic force microscopy) often face limitations in throughput, require invasive labeling, or fail to replicate physiological microenvironments, impeding their clinical utility for rapid cancer cell analysis. To address these limit

Cited by 0SourceScholar
2025

One-for-More: Continual Diffusion Model for Anomaly Detection

CVPR 2025poster

With the rise of generative models, there is a growing interest in unifying all tasks within a generative framework. Anomaly detection methods also fall into this scope and utilize diffusion models to generate or reconstruct normal samples when given arbitrary anomaly images. However, our study foun…

2025

Scaling Mesh Generation via Compressive Tokenization

CVPR 2025poster

We propose a compressive yet effective mesh tokenization, Blocked and Patchified Tokenization (BPT), facilitating the generation of meshes exceeding 8k faces. BPT compresses mesh sequences by employing block-wise indexing and patch aggregation, reducing their length by approximately 75% compared to…

2025

Semantic-Guided Illumination-Aware Deformable Transformer for RGB-T Object Detection

RA-L 2025

RGB-T object detection in autonomous driving has been researched increasingly in recent years. Nevertheless, several problems limit the performance of RGB-T fusion perception. Initially, although illumination awareness is a mature technology to guide fusion process, the outputs of previous methods l

Cited by 1SourceScholar
2025

Sound-VECaps: Improving Audio Generation with Visually Enhanced Captions

ICASSP 2025accepted

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work…

Cited by 0SourceScholar
2025

Sounding that Object: Interactive Object-Aware Image to Audio Generation

ICML 2025poster

Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an interactive object-aware audio generation model that grounds sound generation in user-selected visual objects within images. Our m…

Cited by 0SourcePDFScholar
2025

Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systems

CVPR 2025highlight

By locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data to cloud servers. However, recent studies have revealed that these intermediate features may not sufficiently preserve p…

2025

Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation

AAAI 2025technical

Multi-modal knowledge graph completion (MMKGC) aims to discover unobserved knowledge from given multi-modal knowledge graphs (MMKG), collaboratively leveraging structural information from the triples and multi-modal information of the entities to overcome the inherent incompleteness. Existing MMKGC…

2025

Towards Reliable Large Audio Language Model

ACL 2025finding

Recent advancements in large audio language models (LALMs) have demonstrated impressive results and promising prospects in universal understanding and reasoning across speech, music, and general sound. However, these models still lack the ability to recognize their knowledge boundaries and refuse to…

2025

TransForce: Transferable Force Prediction for Vision-Based Tactile Sensors with Sequential Image Translation

ICRA 2025

Vision-based tactile sensors (VBTSs) provide highresolution tactile images crucial for robot in-hand manipulation. However, force sensing in VBTSs is underutilized due to the costly and time-intensive process of acquiring paired tactile images and force labels. In this study, we introduce a transfer

Cited by 10SourceScholar
2025

Unifying Within and Across: Intra-Modality Multi-View Fusion and Inter-Modality Alignment for Knowledge Graph Completion

ICASSP 2025accepted

Multi-modal knowledge graph completion (MMKGC) enhances the structural and semantic richness of knowledge graphs by integrating diverse information across modalities. However, existing methods often either overlook the diversity within a single modality or fail to ensure effective cross-modality ali…

Cited by 0SourceScholar
2025

ViTacGen: Robotic Pushing With Vision-to-Touch Generation

RA-L 2025

Robotic pushing is a fundamental manipulation task that requires tactile feedback to capture subtle contact forces and dynamics between the end-effector and the object. However, real tactile sensors often face hardware limitations such as high costs and fragility, and deployment challenges involving

Cited by 2SourcecodeScholar
2025

Which Tasks Should Be Compressed Together? A Causal Discovery Approach for Efficient Multi-Task Representation Compression

ICLR 2025poster

Conventional image compression methods are inadequate for intelligent analysis, as they overemphasize pixel-level precision while neglecting semantic significance and the interaction among multiple tasks. This paper introduces a Taskonomy-Aware Multi-Task Compression framework comprising (1) inter-…

Cited by 0SourcePDFScholar
2024

3D-Aware Face Editing via Warping-Guided Latent Direction Learning

CVPR 2024poster

3D facial editing a longstanding task in computer vision with broad applications is expected to fast and intuitively manipulate any face from arbitrary viewpoints following the user's will. Existing works have limitations in terms of intuitiveness generalization and efficiency. To overcome these cha…

2024

A Unified Image Compression Method for Human Perception and Multiple Vision Tasks

ECCV 2024poster

"Recent advancements in end-to-end image compression demonstrate the potential to surpass traditional codecs regarding rate-distortion performance. However, current methods either prioritize human perceptual quality or solely optimize for one or a few predetermined downstream tasks, neglecting a mor…

Cited by 0SourcePDFScholar
2024

Acoustically Driven Micropipette for Hydrodynamic Manipulation of Mouse Oocytes

ICRA 2024poster

Micromanipulation techniques that can achieve controlled fine operations at the micro scale play an important role in biomedical fields including embryo engineering, gene engineering, drug screening, and cell analysis. However, micromanipulation of biological micro-objects, such as cells and micro t…

Cited by 0SourceScholar
2024

DET: A Dual-Encoding Transformer for Relational Graph Embedding

COLING 2024main

Despite recent successes in natural language processing and computer vision, Transformer faces scalability issues when processing graphs, e.g., computing the full node-to-node attention on knowledge graphs (KGs) with million of entities is still infeasible. The existing methods mitigate this problem…

2024

Deep Domain Adaptation Regression for Force Calibration of Optical Tactile Sensors

IROS 2024

Optical tactile sensors provide robots with rich force information for robot grasping in unstructured environments. The fast and accurate calibration of three-dimensional contact forces holds significance for new sensors and existing tactile sensors which may have incurred damage or aging. However,

Cited by 8SourcecodeScholar
2024

Domain-Agnostic Molecular Generation with Chemical Feedback

ICLR 2024poster

The generation of molecules with desired properties has become increasingly popular, revolutionizing the way scientists design molecular structures and providing valuable support for chemical and drug design. However, despite the potential of language models in molecule generation, they face challen…

2024

Dual Mapping of 2D StyleGAN for 3D-Aware Image Generation and Manipulation (Student Abstract)

AAAI 2024technical

3D-aware GANs successfully solve the problem of 3D-consistency generation and furthermore provide a 3D shape of the generated object. However, the application of the volume renderer disturbs the disentanglement of the latent space, which makes it difficult to manipulate 3D-aware GANs and lowers the…

Cited by 0SourcePDFScholar
2024

Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts

ACL 2024findings

In the era of large language models, applying techniques such as Retrieval Augmented Generation can better address Open-Domain Question-Answering problems. Due to constraints including model sizes and computing resources, the length of context is often limited, and it becomes challenging to empower…

2024

Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering

ACL 2024findings

Deploying large language models (LLMs) to real scenarios for domain-specific question answering (QA) is a key thrust for LLM applications, which poses numerous challenges, especially in ensuring that responses are both accommodating to user requirements and appropriately leveraging domain-specific k…

2024

LLM-based Multi-Level Knowledge Generation for Few-shot Knowledge Graph Completion

IJCAI 2024poster

Knowledge Graphs (KGs) are pivotal in various NLP applications but often grapple with incompleteness, especially due to the long-tail problem where infrequent, unpopular relationships drastically reduce the KG completion performance. In this paper, we focus on Few-shot Knowledge Graph Completion (FK…

Cited by 6SourcePDFScholar
2024

MKGL: Mastery of a Three-Word Language

NeurIPS 2024spotlight

Large language models (LLMs) have significantly advanced performance across a spectrum of natural language processing (NLP) tasks. Yet, their application to knowledge graphs (KGs), which describe facts in the form of triplets and allow minimal hallucinations, remains an underexplored frontier. In th…

Cited by 1SourcePDFScholar
2024

Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models

ICLR 2024poster

Large Language Models (LLMs), with their remarkable task-handling capabilities and innovative outputs, have catalyzed significant advancements across a spectrum of fields. However, their proficiency within specialized domains such as biomolecular studies remains limited. To address this challenge, w…

2024

Multi-times Monte Carlo Rendering for Inter-reflection Reconstruction

NeurIPS 2024poster

Inverse rendering methods have achieved remarkable performance in reconstructing high-fidelity 3D objects with disentangled geometries, materials, and environmental light. However, they still face huge challenges in reflective surface reconstruction. Although recent methods model the light trace to…

Cited by 1SourcePDFScholar
2024

OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step

NeurIPS 2024poster

Despite significant advancements in text generation and reasoning, Large Language Models (LLMs) still face challenges in accurately performing complex arithmetic operations. Language model systems often enable LLMs to generate code for arithmetic operations to achieve accurate calculations. However,…

2024

QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation

NeurIPS 2024poster

We propose **Quan**tum-informed **T**ensor **A**daptation (**QuanTA**), a novel, easy-to-implement, fine-tuning method with no inference overhead for large-scale pre-trained language models. By leveraging quantum-inspired methods derived from quantum circuit structures, QuanTA enables efficient *hig…

2024

Rethinking the Soft Conflict Pseudo Boolean Constraint on MaxSAT Local Search Solvers

IJCAI 2024poster

MaxSAT is an optimization version of the famous NP-complete Satisfiability problem (SAT). Algorithms for MaxSAT mainly include complete solvers and local search incomplete solvers. In many complete solvers, once a better solution is found, a Soft conflict Pseudo Boolean (SPB) constraint will be gene…

2024

Revisit and Outstrip Entity Alignment: A Perspective of Generative Models

ICLR 2024poster

Recent embedding-based methods have achieved great successes in exploiting entity alignment from knowledge graph (KG) embeddings of multiple modalities. In this paper, we study embedding-based entity alignment (EEA) from a perspective of generative models. We show that EEA shares similarities with t…

2024

STViT: Improving Self-Supervised Multi-Camera Depth Estimation with Spatial-Temporal Context and Adversarial Geometry Regularization (Student Abstract)

AAAI 2024technical

Multi-camera depth estimation has recently garnered significant attention due to its substantial practical implications in the realm of autonomous driving. In this paper, we delve into the task of self-supervised multi-camera depth estimation and propose an innovative framework, STViT, featuring sev…

Cited by 1SourcePDFScholar
2024

Self-Improvement Programming for Temporal Knowledge Graph Question Answering

COLING 2024main

Temporal Knowledge Graph Question Answering (TKGQA) aims to answer questions with temporal intent over Temporal Knowledge Graphs (TKGs). The core challenge of this task lies in understanding the complex semantic information regarding multiple types of time constraints (e.g., before, first) in questi…

Cited by 9SourcePDFScholar
2024

Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured Representations

AAAI 2024technical

Large-scale vision-language pre-training has achieved significant performance in multi-modal understanding and generation tasks. However, existing methods often perform poorly on image-text matching tasks that require structured representations, i.e., representations of objects, attributes, and rela…

2024

T-SOT FNT: Streaming Multi-Talker ASR with Text-Only Domain Adaptation Capability

ICASSP 2024accepted

Token-level serialized output training (t-SOT) was recently proposed to address the challenge of streaming multi-talker automatic speech recognition (ASR). T-SOT effectively handles overlapped speech by representing multi-talker transcriptions as a single token stream with ⟨cc⟩ symbols interspersed.…

Cited by 0SourceScholar
2024

TENG: Time-Evolving Natural Gradient for Solving PDEs With Deep Neural Nets Toward Machine Precision

ICML 2024poster

Partial differential equations (PDEs) are instrumental for modeling dynamical systems in science and engineering. The advent of neural networks has initiated a significant shift in tackling these complexities though challenges in accuracy persist, especially for initial value problems. In this paper…

2024

UniMix: Towards Domain Adaptive and Generalizable LiDAR Semantic Segmentation in Adverse Weather

CVPR 2024poster

LiDAR semantic segmentation (LSS) is a critical task in autonomous driving and has achieved promising progress. However prior LSS methods are conventionally investigated and evaluated on datasets within the same domain in clear weather. The robustness of LSS models in unseen scenes and all weather c…

Cited by 46SourcePDFScholar
2024

Unleashing the Power of Imbalanced Modality Information for Multi-modal Knowledge Graph Completion

COLING 2024main

Multi-modal knowledge graph completion (MMKGC) aims to predict the missing triples in the multi-modal knowledge graphs by incorporating structural, visual, and textual information of entities into the discriminant models. The information from different modalities will work together to measure the tr…

2023

ANTN: Bridging Autoregressive Neural Networks and Tensor Networks for Quantum Many-Body Simulation

NeurIPS 2023poster

Quantum many-body physics simulation has important impacts on understanding fundamental science and has applications to quantum materials design and quantum technology. However, due to the exponentially growing size of the Hilbert space with respect to the particle number, a direct simulation is int…

2023

Adaptive Patch Deformation for Textureless-Resilient Multi-View Stereo

CVPR 2023poster

In recent years, deep learning-based approaches have shown great strength in multi-view stereo because of their outstanding ability to extract robust visual features. However, most learning-based methods need to build the cost volume and increase the receptive field enormously to get a satisfactory…

2023

An Adapter Based Multi-Label Pre-Training for Speech Separation and Enhancement

ICASSP 2023accepted

In recent years, self-supervised learning (SSL) has achieved tremendous success in various speech tasks due to its power to extract representations from massive unlabeled data. However, compared with tasks such as speech recognition (ASR), the improvements from SSL representation in speech separatio…

Cited by 0SourceScholar
2023

BEATs: Audio Pre-Training with Acoustic Tokenizers

ICML 2023oral

We introduce a self-supervised learning (SSL) framework BEATs for general audio representation pre-training, where we optimize an acoustic tokenizer and an audio SSL model by iterations. Unlike the previous audio SSL models that employ reconstruction loss for pre-training, our audio SSL model is tra…

2023

DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation Tasks

ICASSP 2023accepted

Self-supervised learning has been successfully applied to various speech recognition and understanding tasks. However, for generative tasks such as speech enhancement and speech separation, most self-supervised speech representations did not show substantial improvements. To deal with this problem,…

Cited by 0SourceScholar
2023

DUET: Cross-Modal Semantic Grounding for Contrastive Zero-Shot Learning

AAAI 2023technical

Zero-shot learning (ZSL) aims to predict unseen classes whose samples have never appeared during training. One of the most effective and widely used semantic information for zero-shot image classification are attributes which are annotations for class-level visual characteristics. However, the curre…

2023

Inverse Reinforcement Learning with Graph Neural Networks for IoT Resource Allocation

ICASSP 2023accepted

The rapid development of Internet of Things (IoT) applications requires efficient computing and communication resource allocation strategies to streamline the existing network operations. These strategies could be formulated as mixed-integer nonlinear programming (MINLP) problems, where the optimal…

Cited by 0SourceScholar
2023

Newton–Cotes Graph Neural Networks: On the Time Evolution of Dynamic Systems

NeurIPS 2023spotlight

Reasoning system dynamics is one of the most important analytical approaches for many scientific studies. With the initial state of a system as input, the recent graph neural networks (GNNs)-based methods are capable of predicting the future state distant in time with high accuracy. Although these m…

2023

Programable On-Chip Fabrication of Magnetic Soft Micro-Robot

IROS 2023poster

In the last decade, researchers have been trying to develop many microrobots that mimic the extraordinary abilities of bionts in complex environments. How to fabricate the biomimetic microrobot with satisfying deformability and complex shapes to realize desired precise motion is the key issue. In th…

Cited by 0SourceScholar
2023

Real-Time Speech Interruption Analysis: from Cloud to Client Deployment

ICASSP 2023accepted

Meetings are an essential form of communication for all types of organizations, and remote collaboration systems have been much more widely used since the COVID-19 pandemic. One major issue with remote meetings is that it is challenging for remote participants to interrupt and speak. We have recentl…

Cited by 0SourceScholar
2023

Self-Supervised Learning with Bi-Label Masked Speech Prediction for Streaming Multi-Talker Speech Recognition

ICASSP 2023accepted

Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However, most of the previous studies focused on offline single-talker applications, with limited investigations in multi-talke…

Cited by 0SourceScholar
2023

Simulating Realistic Speech Overlaps Improves Multi-Talker ASR

ICASSP 2023accepted

Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including over-lapping speech of multiple speakers. Due to the difficulty in acquiring real conversation data with high-quality human transcriptions, a naïve simulation of multi-talker…

Cited by 18SourceScholar
2023

Speech Separation with Large-Scale Self-Supervised Learning

ICASSP 2023accepted

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the exploration of the SSL-based SS by massively scaling up both the pre-training data (more than 300K hours) and fine-tuning…

Cited by 0SourceScholar
2023

Target Sound Extraction with Variable Cross-Modality Clues

ICASSP 2023accepted

Automatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture of sources. It often uses a model conditioned on a fixed form of target sound clues, such as a sound class label, which l…

Cited by 0SourceScholar
2023

Vararray Meets T-Sot: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

ICASSP 2023accepted

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geometry. Our framework, named t-SOT-VA, capitalizes on independently developed two recent technologies; array-geometry-agno…

Cited by 0SourceScholar
2022

All-Neural Beamformer for Continuous Speech Separation

ICASSP 2022accepted

Continuous speech separation (CSS) aims to separate overlapping voices from a continuous influx of conversational audio containing an unknown number of utterances spoken by an unknown number of speakers. A common application scenario is transcribing a meeting conversation recorded by a microphone ar…

Cited by 0SourceScholar
2022

Collaboration of Experts: Achieving 80% Top-1 Accuracy on ImageNet with 100M FLOPs

ICML 2022spotlight

In this paper, we propose a Collaboration of Experts (CoE) framework to assemble the expertise of multiple networks towards a common goal. Each expert is an individual network with expertise on a unique portion of the dataset, contributing to the collective capacity. Given a sample, delegator select…

Cited by 13SourcePDFScholar
2022

Continuous Speech Separation with Recurrent Selective Attention Network

ICASSP 2022accepted

While permutation invariant training (PIT) based continuous speech separation (CSS) significantly improves the conversation transcription accuracy, it often suffers from speech leakages and failures in separation at "hot spot" regions because it has a fixed number of output channels. In this paper,…

Cited by 0SourceScholar
2022

Continuous Streaming Multi-Talker ASR with Dual-Path Transducers

ICASSP 2022accepted

Streaming recognition of multi-talker conversations has so far been evaluated only for 2-speaker single-turn sessions. In this paper, we investigate it for multi-turn meetings containing multiple speakers using the Streaming Unmixing and Recognition Transducer (SURT) model, and show that naively ext…

Cited by 0SourceScholar
2022

Molecular Contrastive Learning with Chemical Element Knowledge Graph

AAAI 2022technical

Molecular representation learning contributes to multiple downstream tasks such as molecular property prediction and drug design. To properly represent molecules, graph contrastive learning is a promising paradigm as it utilizes self-supervision signals and has no requirements for human annotations.…

2022

On-Chip Automatic Trapping and Rotating for Zebrafish Embryo Injection

RA-L 2022

Zebrafish embryo injection is often required in biomedical research using zebrafish. In the injecting operation, trapping and rotating the zebrafish embryo to achieve a proper posture is essential for the high success rate. We proposed an on-chip platform capable of efficient and automatic trapping

Cited by 8SourceScholar
2022

One Model to Enhance Them All: Array Geometry Agnostic Multi-Channel Personalized Speech Enhancement

ICASSP 2022accepted

With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect with friends and families. Single-channel personalized speech enhancement (PSE) methods show promising results compared…

Cited by 0SourceScholar
2022

Personalized speech enhancement: new models and Comprehensive evaluation

ICASSP 2022accepted

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and thus improve the speech quality of online video conferencing systems for various acoustic scenarios. In this work, we pr…

Cited by 0SourceScholar
2022

Structural Triangulation: A Closed-Form Solution to Constrained 3D Human Pose Estimation

ECCV 2022poster

"We propose Structural Triangulation, a closed-form solution for optimal 3D human pose considering multi-view 2D pose estimations, calibrated camera parameters, and bone lengths. To start with, we focus on embedding structural constraints of human body in the process of 2D-to-3D inference using tria…

2022

Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers Using End-to-End Speaker-Attributed ASR

ICASSP 2022accepted

This paper presents Transcribe-to-Diarize, a new approach for neural speaker diarization that uses an end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR). The E2E SA-ASR is a joint model that was recently proposed for speaker counting, multi-talker speech recognition, and speak…

Cited by 0SourceScholar
2022

Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-Training

ICASSP 2022accepted

Self-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years have witnessed great successes in applying self-supervised learning in speech recognition, while limited exploration was attemp…

Cited by 0SourceScholar
2022

VarArray: Array-Geometry-Agnostic Continuous Speech Separation

ICASSP 2022accepted

Continuous speech separation using a microphone array was shown to be promising in dealing with the speech overlap problem in natural conversation transcription. This paper proposes VarArray, an array-geometry-agnostic speech separation neural network model. The proposed model is applicable to any n…

Cited by 47SourceScholar
2021

Don't Shoot Butterfly with Rifles: Multi-Channel Continuous Speech Separation with Early Exit Transformer

ICASSP 2021accepted

With its strong modeling capacity that comes from a multi-head and multi-layer structure, Transformer is a very powerful model for learning a sequential representation and has been successfully applied to speech separation recently. However, multi-channel speech separation sometimes does not necessa…

Cited by 0SourceScholar
2021

Dual-Path Modeling for Long Recording Speech Separation in Meetings

ICASSP 2021accepted

The continuous speech separation (CSS) is a task to separate the speech sources from a long, partially overlapped recording, which involves a varying number of speakers. A straightforward extension of conventional utterance-level speech separation to the CSS task is to segment the long recording wit…

Cited by 0SourceScholar
2021

Knowledge-aware Zero-Shot Learning: Survey and Perspective

IJCAI 2021poster

Zero-shot learning (ZSL) which aims at predicting classes that have never appeared during the training using external knowledge (a.k.a. side information) has been widely investigated. In this paper we present a literature review towards ZSL in the perspective of external knowledge, where we categori…

Cited by 80SourcePDFScholar
2021

Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020

ICASSP 2021accepted

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2020. We will first explain our system design to address issues in handling real multi-talker reco…

Cited by 0SourceScholar
2021

Minimum Bayes Risk Training for End-to-End Speaker-Attributed ASR

ICASSP 2021accepted

Recently, an end-to-end speaker-attributed automatic speech recognition (E2E SA-ASR) model was proposed as a joint model of speaker counting, speech recognition and speaker identification for monaural overlapped speech. In the previous study, the model parameters were trained based on the speaker-at…

Cited by 0SourceScholar
2021

Rethinking The Separation Layers In Speech Separation Networks

ICASSP 2021accepted

Modules in all existing speech separation networks can be categorized into single-input-multi-output (SIMO) modules and single-input-single-output (SISO) modules. SIMO modules generate more outputs than input, and SISO modules keep the numbers of input and output the same. While the majority of sepa…

Cited by 0SourceScholar
2020

A Learning Approach to Cooperative Communication System Design

ICASSP 2020accepted

The cooperative relay network is a type of multi-terminal communication system. We present in this paper a Neural Network (NN)-based autoencoder (AE) approach to optimize its design. This approach implements a classical three-node cooperative system as one AE model, and uses a two-stage scheme to tr…

Cited by 0SourceScholar
2020

Continuous Speech Separation: Dataset and Analysis

ICASSP 2020accepted

This paper describes a dataset and protocols for evaluating continuous speech separation algorithms. Most prior speech separation studies use pre-segmented audio signals, which are typically generated by mixing speech utterances on computers so that they fully overlap. Also, the separation algorithm…

Cited by 0SourceScholar
2020

Dual-Path RNN: Efficient Long Sequence Modeling for Time-Domain Single-Channel Speech Separation

ICASSP 2020accepted

Recent studies in deep learning-based speech separation have proven the superiority of time-domain approaches to conventional time-frequency-based methods. Unlike the time-frequency domain approaches, the time-domain separation systems often receive input sequences consisting of a huge number of tim…

Cited by 0SourceScholar
2020

End-to-end Microphone Permutation and Number Invariant Multi-channel Speech Separation

ICASSP 2020accepted

An important problem in ad-hoc microphone speech separation is how to guarantee the robustness of a system with respect to the locations and numbers of microphones. The former requires the system to be invariant to different indexing of the microphones with the same locations, while the latter requi…

Cited by 0SourceScholar
2020

Improving Deep CNN Networks with Long Temporal Context for Text-Independent Speaker Verification

ICASSP 2020accepted

Deep CNN networks have shown great success in various tasks for text-independent speaker recognition. In this paper, we explore two approaches for modeling long temporal contexts to improve the performance of the ResNet networks. The first approach is simply integrating the utterance-level mean and…

Cited by 0SourceScholar
2020

Mesh-Guided Multi-View Stereo With Pyramid Architecture

CVPR 2020poster

Multi-view stereo (MVS) aims to reconstruct 3D geometry of the target scene by using only information from 2D images. Although much progress has been made, it still suffers from textureless regions. To overcome this difficulty, we propose a mesh-guided MVS method with pyramid architecture, which mak…

Cited by 38PDFcodeScholar
2020

PuppeteerGAN: Arbitrary Portrait Animation With Semantic-Aware Appearance Transformation

CVPR 2020poster

Portrait animation, which aims to animate a still portrait to life using poses extracted from target frames, is an important technique for many real-world entertainment applications. Although recent works have achieved highly realistic results on synthesizing or controlling human head images, the pu…

Cited by 58PDFScholar
2020

Real-Time Task Offloading for Large-Scale Mobile Edge Computing

ICASSP 2020accepted

Mobile-edge computing (MEC) is a promising technology to support computation-intensive and delay-sensitive applications at smart devices by offloading their local tasks to the network edge. In this paper, we propose a novel index based real-time task offloading policy for an asynchronous large-scale…

Cited by 1SourceScholar
2020

Robot-Assisted and Wearable Sensor-Mediated Autonomous Gait Analysis§

ICRA 2020poster

In this paper, we propose an autonomous gait analysis system consisting of a mobile robot and custom-engineered instrumented insoles. The robot is equipped with an on-board RGB-D sensor, the insoles feature inertial sensors and force sensitive resistors. This system is motivated by the need for a ro…

Cited by 22SourceScholar
2019

Low-latency Speaker-independent Continuous Speech Separation

ICASSP 2019accepted

Speaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of which contains no overlapping speech segment. A separated, or cleaned, version of e…

Cited by 0SourceScholar
2019

Single-channel Speech Extraction Using Speaker Inventory and Attention Network

ICASSP 2019accepted

Neural network-based speech separation has received a surge of interest in recent years. Previously proposed methods either are speaker independent or extract a target speaker's voice by using his or her voice snippet. In applications such as home devices or office meeting transcriptions, a possible…

Cited by 76SourceScholar
2018

Developing Far-Field Speaker System Via Teacher-Student Learning

ICASSP 2018accepted

In this study, we develop the keyword spotting (KWS) and acoustic model (AM) components in a far-field speaker system. Specifically, we use teacher-student (T/S) learning to adapt a close-talk well-trained production AM to far-field by using parallel close-talk and simulated far-field data. We also…

Cited by 0SourceScholar
2018

Efficient Integration of Fixed Beamformers and Speech Separation Networks for Multi-Channel Far-Field Speech Separation

ICASSP 2018accepted

Speech separation research has significantly progressed in recent years thanks to the rapid advances in deep learning technology. However the performance of recently proposed single-channel neural network-based speech separation methods is still limited especially in reverberant environments. To pus…

Cited by 0SourceScholar
2018

Image Quality Assessment Based Label Smoothing in Deep Neural Network Learning

ICASSP 2018accepted

For many computer vision problems, deep neural networks are trained and validated based on the assumption that the input images are pristine (i.e., artifact-free). However, digital images are subject to a wide range of distortions in real application scenarios, while the practical issues regarding i…

Cited by 0SourceScholar
2018

Mobile Bayesian Spectrum Learning for Heterogeneous Networks

ICASSP 2018accepted

Spectrum sensing in heterogeneous networks is very challenging as it usually requires a large number of static secondary users (SUs) to capture the global spectrum states. In this paper, we tackle the spectrum sensing in heterogeneous networks from a new perspective. We exploit the mobility of multi…

Cited by 0SourceScholar
2018

Multi-Microphone Neural Speech Separation for Far-Field Multi-Talker Speech Recognition

ICASSP 2018accepted

This paper describes a neural network approach to far-field speech separation using multiple microphones. Our proposed approach is speaker-independent and can learn to implicitly figure out the number of speakers constituting an input speech mixture. This is realized by utilizing the permutation inv…

Cited by 0SourceScholar
2018

Speaker-Invariant Training Via Adversarial Learning

ICASSP 2018accepted

We propose a novel adversarial multi-task learning scheme, aiming at actively curtailing the inter-talker feature variability while maximizing its senone discriminability so as to enhance the performance of a deep neural network (DNN) based ASR system. We call the scheme speaker-invariant training (…

Cited by 0SourceScholar
2017

Deep clustering and conventional networks for music separation: Stronger together

ICASSP 2017accepted

Deep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However, little is known about its effectiveness in other challenging si…

Cited by 0SourceScholar
2016

Deep clustering: Discriminative embeddings for segmentation and separation

ICASSP 2016accepted

We address the problem of "cocktail-party" source separation in a deep learning framework called deep clustering. Previous deep network approaches to separation have shown promising performance in scenarios with a fixed number of sources, each belonging to a distinct signal class, such as speech and…

Cited by 0SourceScholar