← Search

Jian Zhu

36 accepted papers

2026

Boomda: Balanced Multi-objective Optimization for Multimodal Domain Adaptation

AAAI 2026technical

Multimodal learning, while contributing to numerous success stories across various fields, faces the challenge of prohibitively expensive manual annotation. To address the scarcity of annotated data, a popular solution is unsupervised domain adaptation, which has been extensively studied in unimodal

Cited by 0SourcePDFScholar
2026

Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models

ICML 2026poster

Visual prostheses hold great promise for restoring vision in blind individuals. While researchers have successfully utilized M/EEG signals to evoke visual perceptions during the brain decoding stage of visual prostheses, the complementary process of converting images into M/EEG signals in the brain …

Cited by 0SourceScholar
2026

SAGA: Structural Aggregation Guided Alignment with Dynamic View and Neighborhood Order Selection for Multiview Graph Domain Adaptation

ICLR 2026poster

Graph domain adaptation (GDA) transfers knowledge from a labeled source graph to an unlabeled target graph to alleviate label scarcity. In multi-view graphs, the challenge of mitigating domain shift is constrained by structural information across various views. Moreover, within each view, structures…

Cited by 0SourcecodeScholar
2025

Adversarial Alignment with Anchor Dragging Drift (A3D2): Multimodal Domain Adaptation with Partially Shifted Modalities

ACL 2025long

Multimodal learning has celebrated remarkable success across diverse areas, yet faces the challenge of prohibitively expensive data collection and annotation when adapting models to new environments. In this context, domain adaptation has gained growing popularity as a technique for knowledge transf…

2025

Developing multilingual speech synthesis system for Ojibwe, Mi’kmaq, and Maliseet

NAACL 2025short

We present lightweight flow matching multilingual text-to-speech (TTS) systems for Ojibwe, Mi’kmaq, and Maliseet, three Indigenous languages in North America. Our results show that training a multilingual TTS model on three typologically similar languages can improve the performance over monolingual…

2025

Dynamic SRM Curriculum for Trustworthy Multi-modal Classification

ICASSP 2025accepted

Trustworthy multi-modal learning integrates multiple sources of data reliably. However, the current methods still focus on performance improvement by developing deep multi-modal networks. These approaches frequently encounter challenges due to the inherent non-convex nature of deep neural networks a…

Cited by 0SourceScholar
2025

High-Force Electroadhesion Based on Unique Liquid-Solid Dielectrics for UAV Perching

ICRA 2025

Electroadhesion (EA), as an electrostatically driven, controllable adhesion technology, has unique attributes such as low noise, robust adaptability, and energy efficiency. However, its adhesion pressure is still low (0.1~10kPa) which may significantly limit its applications. This paper presents an

Cited by 2SourceScholar
2025

Multimodal Deformation Estimation of Soft Pneumatic Gripper During Operation

IROS 2025

Soft pneumatic robots are gaining significant attention due to their compliance and adaptability in unstructured environments. While emerging dual-chamber soft pneumatic robots can achieve complex 3D deformations beyond conventional single-axis bending, real-time proprioception remains challenging d

Cited by 0SourceScholar
2025

Trusted Mamba Contrastive Network for Multi-View Clustering

ICASSP 2025accepted

Multi-view clustering can partition data samples into their categories by learning a consensus representation in an unsupervised way and has received more and more attention in recent years. However, there is an untrusted fusion problem. The reasons for this problem are as follows: 1) The current me…

Cited by 10SourceScholar
2025

ZIPA: A family of efficient models for multilingual phone recognition

ACL 2025long

We present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. We first curated IPA PACK++, a large-scale multilingual speech corpus with 17,000+ hours of normalized phone transcriptions and a novel evaluation set capturing u…

2024

Adaptive Confidence Multi-View Hashing for Multimedia Retrieval

ICASSP 2024accepted

The multi-view hash method converts heterogeneous data from multiple views into binary hash codes, which is one of the critical technologies in multimedia retrieval. However, the current methods mainly explore the complementarity among multiple views while lacking confidence in learning and fusion.…

Cited by 0SourceScholar
2024

Embedded 3D Printing of Silicone for Soft Actuator with Stiffness Gradient and Programmable Workspace

IROS 2024poster

Soft pneumatic actuators can accomplish various customizable deformation/motion through the distribution of cavities and gradients in stiffness. However, traditional manufacturing methods, say molding, struggle to produce soft actuators with both complex cavities and desirable stiffness distribution…

Cited by 0SourceScholar
2024

Generalize for Future: Slow and Fast Trajectory Learning for CTR Prediction

AAAI 2024technical

Deep neural networks (DNNs) have achieved significant advancements in click-through rate (CTR) prediction by demonstrating strong generalization on training data. However, in real-world scenarios, the assumption of independent and identically distributed (i.i.d.) conditions, which is fundamental to…

Cited by 0SourcePDFScholar
2024

Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh Recovery

NeurIPS 2024poster

Human Mesh Recovery (HMR) is the task of estimating a parameterized 3D human mesh from an image. There is a kind of methods first training a regression model for this problem, then further optimizing the pretrained regression model for any specific sample individually at test time. However, the pret…

2024

The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language

NAACL 2024long

In this project, we demonstrate that phoneme-based models for speech processing can achieve strong crosslinguistic generalizability to unseen languages. We curated the IPAPACK, a massively multilingual speech corpora with phonemic transcriptions, encompassing more than 115 languages from diverse lan…

2023

RWKV: Reinventing RNNs for the Transformer Era

EMNLP 2023long findings

Transformers have revolutionized almost all natural language processing (NLP) tasks but suffer from memory and computational complexity that scales quadratically with sequence length. In contrast, recurrent neural networks (RNNs) exhibit linear scaling in memory and computational requirements but st…

Cited by 0SourceScholar
2022

Bootstrapping meaning through listening: Unsupervised learning of spoken sentence embeddings

EMNLP 2022finding

Inducing semantic representations directly from speech signals is a highly challenging task but has many useful applications in speech mining and spoken language understanding. This study tackles the unsupervised learning of semantic representations for spoken utterances. Through converting speech s…

2022

Modeling of viscoelastic dielectric elastomer actuators based on the sparse identification method

ICRA 2022poster

Dielectric elastomer actuators (DEAs) have been widely employed to drive various soft robots, due to their quiet fast muscle-like behavior. It is significant but challenging to model and control these soft actuators, due to their viscoelastic property, irregular geometry, complex structure, etc. In…

Cited by 5SourceScholar
2022

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

NeurIPS 2022accept

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large lan…

Cited by 214SourcePDFScholar
2021

Idiosyncratic but not Arbitrary: Learning Idiolects in Online Registers Reveals Distinctive yet Consistent Individual Styles

EMNLP 2021main

An individual’s variation in writing style is often a function of both social and personal attributes. While structured social variation has been extensively studied, e.g., gender based variation, far less is known about how to characterize individual styles due to their idiosyncratic nature. We int…

2021

Scalable Discriminative Discrete Hashing For Large-Scale Cross-Modal Retrieval

ICASSP 2021accepted

Cross-modal hashing has received increasing research attentions due to its less storage and efficient retrieval. However, most existing cross-modal hashing methods focus only on exploring multi-modal information, while underestimate the significance of local and Euclidean structure information on th…

Cited by 0SourceScholar
2020

An Earthworm-like Soft Robot with Integration of Single Pneumatic Actuator and Cellular Structures for Peristaltic Motion

IROS 2020poster

Earthworm-like soft robots have been widely studied for various applications, such as medical endoscopy and pipeline inspection. Many actuation modes have been chosen to drive the soft robots, including pneumatic actuators, dielectric elastomeric actuators, and shape memory actuators. Pneumatic actu…

Cited by 15SourceScholar
2019

Deep Reinforcement Learning in Soft Viscoelastic Actuator of Dielectric Elastomer

RA-L 2019

Dielectric elastomer actuators (DEAs) have been widely employed as artificial muscles in soft robots. Due to material viscoelasticity and nonlinear electromechanical coupling, it is challenging to accurately model a viscoelastic DEA, especially when the actuator is of a complex or irregular configur

Cited by 32SourceScholar
2019

Denoising Convolutional Autoencoder Based B-mode Ultrasound Tongue Image Feature Extraction

ICASSP 2019accepted

B-mode ultrasound tongue imaging is widely used in the speech production field. However, efficient interpretation is in a great need for the tongue image sequences. Inspired by the recent success of unsupervised deep learning approach, we explore unsupervised convolutional network architecture for t…

Cited by 0SourceScholar
2019

Predicting Tongue Motion in Unlabeled Ultrasound Videos Using Convolutional Lstm Neural Networks

ICASSP 2019accepted

A challenge in speech production research is to predict future tongue movements based on a short period of past tongue movements. This study tackles speaker-dependent tongue motion prediction problem in unlabeled ultrasound videos with convolutional long short-term memory (ConvLSTM) networks. The mo…

Cited by 0SourceScholar
2018

Modelling and Control of a Novel Soft Crawling Robot Based on a Dielectric Elastomer Actuator

ICRA 2018poster

Soft robots have recently evoked extensive attention due to their abilities to work effectively in unstructured environments. As an actuation technology of soft robots, dielectric elastomers exhibit many intriguing attributes such as large strain and high energy density. This work presents a novel d…

Cited by 34SourceScholar
2018

Topology Optimized Design, Fabrication, and Characterization of a Soft Cable-Driven Gripper

RA-L 2018

Soft-bodied robots, due to their intrinsic compliance, have shown great potential for operating within unstructured environment and interacting with unknown objects. This letter deals with automatic design and fabrication of soft robots. From a structure point of view, we synthesize a soft cable-dri

Cited by 125SourceScholar
2017

A frog-inspired swimming robot based on dielectric elastomer actuators

IROS 2017poster

Frogs are capable of multiple locomotion modes including jumping and swimming, which enables them to adapt to various environmental conditions. This paper demonstrates a frog-inspired robot, which can mimic the swimming motion of a natural frog. The robot is developed based on dielectric elastomer a…

Cited by 71SourceScholar
2017

Networked soft actuators with large deformations

ICRA 2017poster

Soft actuators play an important role in producing motions in soft robots, and dielectric elastomers have shown great promise because of their considerable voltage-induced deformation. In particular, air-filled dielectric elastomer actuators have been well studied, where the air inside provides pres…

Cited by 12SourceScholar