← Search

Dong Yang

34 accepted papers

2026

A Dual-Channel Framework for Blind Perceptual Quality Assessment in Bilateral Teleoperation

ICRA 2026poster

This paper proposes a perceptual no-reference (blind) haptic quality assessment framework for predicting the Quality of Experience (QoE) in teleoperation systems with force feedback. The proposed approach employs a deep neural network that combines semantic and distortion-based channels. The semanti…

Cited by 0Scholar
2026

DA-DFGAS:Differentiable Federated Graph Neural Architecture Search with Distribution-Aware Attentive Aggregation

AAAI 2026technical

Graph Neural Networks (GNNs) have demonstrated superior performance in processing centralized graph-structured data. However, real-world privacy and security concerns hinder data centralization and shareing, leading to severe data isolation (data silos). While Federated Learning (FL) offers a distri

Cited by 0SourcePDFScholar
2026

MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss

AAAI 2026technical

Medical image synthesis is an important topic for both clinical and research applications. Recently, diffusion models have become a leading approach in this area. Despite their strengths, many existing methods struggle with (1) limited generalizability, only working for specific body regions or voxe

Cited by 0SourcePDFScholar
2026

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

ICLR 2026poster

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study the design choices across model architecture and data curati…

Cited by 0SourcecodeScholar
2025

Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging

NeurIPS 2025poster

Recent progress in vision-language modeling for 3D medical imaging has been fueled by large-scale computed tomography (CT) corpora with paired free-text reports, stronger architectures, and powerful pretrained models. This has enabled applications such as automated report generation and text-conditi…

Cited by 0SourcecodeScholar
2025

Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis

NeurIPS 2025poster

We propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike conventional FM modules, which use the coarse representations from the weak generator as conditions, SFM constructs interme…

Cited by 0SourcecodeScholar
2025

VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge

CVPR 2025highlight

Generalist vision language models (VLMs) have made significant strides in computer vision, but they fall short in specialized fields like healthcare, where expert knowledge is essential. Current large multimodal models like Gemini and GPT-4o are insufficient for medical tasks due to their reliance o…

Cited by 5SourcePDFScholar
2025

VISTA3D: A Unified Segmentation Foundation Model For 3D Medical Imaging

CVPR 2025poster

Foundation models for interactive segmentation in 2D natural images and videos have sparked significant interest in building 3D foundation models for medical imaging. However, the domain gaps and clinical use cases for 3D medical imaging require a dedicated model that diverges from existing 2D solut…

2024

FedBPT: Efficient Federated Black-box Prompt Tuning for Large Language Models

ICML 2024poster

Pre-trained language models (PLM) have revolutionized the NLP landscape, achieving stellar performances across diverse tasks. These models, while benefiting from vast training data, often require fine-tuning on specific data to cater to distinct downstream tasks. However, this data adaptation proces…

Cited by 34SourcePDFScholar
2024

HPF-SLAM: An Efficient Visual SLAM System Leveraging Hybrid Point Features

ICRA 2024poster

Visual SLAM is an essential tool in diverse applications such as robot perception and extended reality, where feature-based methods are prevalent due to their accuracy and robustness. However, existing methods employ either hand-crafted or solely learnable point features and are thus limited by the…

Cited by 1SourceScholar
2024

KC-GenRe: A Knowledge-constrained Generative Re-ranking Method Based on Large Language Models for Knowledge Graph Completion

COLING 2024main

The goal of knowledge graph completion (KGC) is to predict missing facts among entities. Previous methods for KGC re-ranking are mostly built on non-generative language models to obtain the probability of each candidate. Recently, generative large language models (LLMs) have shown outstanding perfor…

2024

Prompt Space Optimizing Few-shot Reasoning Success with Large Language Models

NAACL 2024findings

Prompt engineering is an essential technique for enhancing the abilities of large language models (LLMs) by providing explicit and specific instructions. It enables LLMs to excel in various tasks, such as arithmetic reasoning, question answering, summarization, relation extraction, machine translati…

2023

A Canonicalization-Enhanced Known Fact-Aware Framework For Open Knowledge Graph Link Prediction

IJCAI 2023poster

Open knowledge graph (OpenKG) link prediction aims to predict missing factual triples in the form of (head noun phrase, relation phrase, tail noun phrase). Since triples are not canonicalized, previous methods either focus on canonicalizing noun phrases (NPs) to reduce graph sparsity, or utilize tex…

2023

CoMave: Contrastive Pre-training with Multi-scale Masking for Attribute Value Extraction

ACL 2023findings

Attribute Value Extraction (AVE) aims to automatically obtain attribute value pairs from product descriptions to aid e-commerce. Despite the progressive performance of existing approaches in e-commerce platforms, they still suffer from two challenges: 1) difficulty in identifying values at different…

2023

Communication-Efficient Vertical Federated Learning with Limited Overlapping Samples

ICCV 2023poster

Federated learning is a popular collaborative learning approach that enables clients to train a global model without sharing their local data. Vertical federated learning (VFL) deals with scenarios in which the data on clients have different feature spaces but share some overlapping samples. Existin…

Cited by 18PDFScholar
2023

Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-Speech

ICASSP 2023accepted

Pause insertion, also known as phrase break prediction and phrasing, is an essential part of TTS systems because proper pauses with natural duration significantly enhance the rhythm and intelligibility of synthetic speech. However, conventional phrasing models ignore various speakers’ different styl…

Cited by 0SourceScholar
2023

Fair Federated Medical Image Segmentation via Client Contribution Estimation

CVPR 2023poster

How to ensure fairness is an important topic in federated learning (FL). Recent studies have investigated how to reward clients based on their contribution (collaboration fairness), and how to achieve uniformity of performance across clients (performance fairness). Despite achieving progress on eith…

Cited by 63SourcePDFScholar
2023

Haptic Dataset Augmentation with Subjective QoE Labels using Conditional Generative Adversarial Network

IROS 2023poster

This paper proposes a novel Generative Adversarial Network (GAN)-based strategy to augment subjective haptic Quality of Experience (QoE) datasets for bilateral teleoperation with haptic feedback without conducting time-consuming subjective experiments. In our previous work, we proposed a multi-asses…

Cited by 1SourceScholar
2023

Low-Complexity Acoustic Echo Cancellation with Neural Kalman Filtering

ICASSP 2023accepted

The Kalman filter has been adopted in acoustic echo cancellation due to its robustness to double-talk, fast convergence, and good steady-state performance. The performance of Kalman filter is closely related to the estimation accuracy of the state noise covariance and the observation noise covarianc…

Cited by 0SourceScholar
2023

Neural Deformable Models for 3D Bi-Ventricular Heart Shape Reconstruction and Modeling from 2D Sparse Cardiac Magnetic Resonance Imaging

ICCV 2023poster

We propose a novel neural deformable model (NDM) targeting at the reconstruction and modeling of 3D bi-ventricular shape of the heart from 2D sparse cardiac magnetic resonance (CMR) imaging data. We model the bi-ventricular shape using blended deformable superquadrics, which are parameterized by a s…

Cited by 6PDFcodeScholar
2023

SRI-Graph: A Novel Scene-Robot Interaction Graph for Robust Scene Understanding

ICRA 2023poster

We propose a novel scene-robot interaction graph (SRI-Graph) that exploits the known position of a mobile manipulator for robust and accurate scene understanding. Compared to the state-of-the-art scene graph approaches, the proposed SRI-Graph captures not only the relationships between the objects,…

Cited by 5SourceScholar
2022

Auto-FedRL: Federated Hyperparameter Optimization for Multi-Institutional Medical Image Segmentation

ECCV 2022poster

"Federated learning (FL) is a distributed machine learning technique that enables collaborative model training while avoiding explicit data sharing. The inherent privacy-preserving property of FL algorithms makes them especially attractive to the medical field. However, in case of heterogeneous clie…

2022

Closing the Generalization Gap of Cross-Silo Federated Medical Image Segmentation

CVPR 2022poster

Cross-silo federated learning (FL) has attracted much attention in medical imaging analysis with deep learning in recent years as it can resolve the critical issues of insufficient data, data privacy, and training efficiency. However, there can be a generalization gap between the model trained from…

Cited by 84PDFcodeScholar
2022

GammaE: Gamma Embeddings for Logical Queries on Knowledge Graphs

EMNLP 2022main

Embedding knowledge graphs (KGs) for multi-hop logical reasoning is a challenging problem due to massive and complicated structures in many KGs. Recently, many promising works projected entities and queries into a geometric space to efficiently find answers. However, it remains challenging to model…

2022

HyperSegNAS: Bridging One-Shot Neural Architecture Search With 3D Medical Image Segmentation Using HyperNet

CVPR 2022poster

Semantic segmentation of 3D medical images is a challenging task due to the high variability of the shape and pattern of objects (such as organs or tumors). Given the recent success of deep learning in medical image segmentation, Neural Architecture Search (NAS) has been introduced to find high-perf…

Cited by 41PDFScholar
2022

Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image Analysis

CVPR 2022poster

Vision Transformers (ViT)s have shown great performance in self-supervised learning of global and local representations that can be transferred to downstream applications. Inspired by these results, we introduce a novel self-supervised learning framework with tailored proxy tasks for medical image a…

Cited by 796PDFcodeScholar
2022

Skill-CPD: Real-time Skill Refinement for Shared Autonomy in Manipulator Teleoperation

IROS 2022

Advanced wireless communication networks provide lower latency and a higher transmission rate. Although this is an enabler for many new teleoperation applications, the risk of network instability or packet drop is still unavoidable. Real-time manipulator teleoperation requires data transmission with

Cited by 8SourcecodeScholar
2021

DeepTag: An Unsupervised Deep Learning Method for Motion Tracking on Cardiac Tagging Magnetic Resonance Images

CVPR 2021poster

Cardiac tagging magnetic resonance imaging (t-MRI) is the gold standard for regional myocardium deformation and cardiac strain estimation. However, this technique has not been widely used in clinical diagnosis, as a result of the difficulty of motion tracking encountered with t-MRI images. In this p…

Cited by 47PDFcodeScholar
2021

DiNTS: Differentiable Neural Network Topology Search for 3D Medical Image Segmentation

CVPR 2021poster

Recently, neural architecture search(NAS) has been applied to automatically search high-performance networks for medical image segmentation. The NAS search space usually contains a network topology level(controlling connections among cells with different spatial scales) and a cell level(operations w…

Cited by 116PDFcodeScholar
2021

I2UV-HandNet: Image-to-UV Prediction Network for Accurate and High-Fidelity 3D Hand Mesh Modeling

ICCV 2021poster

Reconstructing a high-precision and high-fidelity 3D human hand from a color image plays a central role in replicating a realistic virtual hand in human-computer interaction and virtual reality applications. Current methods are lacking in accuracy and fidelity due to various hand poses and severe oc…

Cited by 73PDFScholar
2021

T-AutoML: Automated Machine Learning for Lesion Segmentation Using Transformers in 3D Medical Imaging

ICCV 2021poster

Lesion segmentation in medical imaging has been an important topic in clinical research. Researchers have proposed various detection and segmentation algorithms to address this task. Recently, deep learning-based approaches have significantly improved the performance over conventional methods. Howev…

Cited by 36PDFScholar
2020

C2FNAS: Coarse-to-Fine Neural Architecture Search for 3D Medical Image Segmentation

CVPR 2020poster

3D convolution neural networks (CNN) have been proved very successful in parsing organs or tumours in 3D medical images, but it remains sophisticated and time-consuming to choose or design proper 3D networks given different task contexts. Recently, Neural Architecture Search (NAS) is proposed to sol…

Cited by 182PDFScholar
2019

An Alarm System for Segmentation Algorithm Based on Shape Model

ICCV 2019accepted

It is usually hard for a learning system to predict correctly on rare events that never occur in the training data, and there is no exception for segmentation algorithms. Meanwhile, manual inspection of each case to locate the failures becomes infeasible due to the trend of large data scale and limi…

Cited by 30SourcePDFScholar