← Search

Ying Chen

64 accepted papers

2026

EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model

ICLR 2026poster

Subject-driven generation is a critical task in creative AI; yet current state-of-the-art methods present a stark trade-off. They either rely on computationally expensive, per-subject fine-tuning, sacrificing efficiency and zero-shot capability, or employ feed-forward architectures built on diffusio…

Cited by 0SourcecodeScholar
2026

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI

AAAI 2026technical

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annot

Cited by 0SourcePDFScholar
2026

Hybrid-Driven Disc-Shaped Autonomous Underwater Vehicle With High Maneuverability and Gliding Capability: Design and Experiments

RA-L 2026

This paper presents the mechatronic design and implementation of a hybrid-driven disc-shaped autonomous underwater vehicle (HD-AUV). The hybrid-driven system integrates a buoyancy adjustment system and propeller thrusters, enabling the HD-AUV to achieve both high maneuverability motion and energy-ef

Cited by 1SourceScholar
2026

HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics Prediction

CVPR 2026

Spatial Transcriptomics (ST) merges the benefits of pathology images and gene expression, linking molecular profiles with tissue structure to analyze spot-level function comprehensively. Predicting gene expression from histology images is a cost-effective alternative to expensive ST technologies. Ho

Cited by 0SourcecodeScholar
2026

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models

CVPR 2026

Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization (DPO) methods rely on multi-sample ranking and task-specific critic models, which is inefficient and often yields ambiguous global supervision. To addr

Cited by 0SourcecodeScholar
2026

Rethinking Diffusion Model-Based Video Super-Resolution: Leveraging Dense Guidance from Aligned Features

CVPR 2026

Diffusion model (DM) based Video Super-Resolution (VSR) approaches achieve impressive perceptual quality. Diffusion model (DM) based Video Super-Resolution (VSR) approaches achieve impressive perceptual quality. However, existing DM-based VSR methods over-prioritize perceptual synthesis while neglec

Cited by 0SourcecodeScholar
2026

S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything Without Supervision

AAAI 2026technical

Recent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This

Cited by 0SourcePDFScholar
2026

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

ICML 2026poster

Medical diagnosis demands models that can process multimodal medical inputs, such as medical images and patient histories, and generate diverse outputs including textual reports and visual content, such as annotations or segmentation masks. Despite this need, existing medical AI models disrupt this …

Cited by 0SourceScholar
2026

Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration

ICLR 2026poster

We present Vivid-VR, a DiT-based generative video restoration method built upon an advanced T2V foundation model, where ControlNet is leveraged to control the generation process, ensuring content consistency. However, conventional fine-tuning of such controllable pipelines frequently suffers from di…

Cited by 0SourcecodeScholar
2026

When Do Graph Foundation Models Transfer? A Data-Centric Theory

ICML 2026poster

Graph foundation models (GFMs) aim to reuse a single backbone across diverse graph domains, yet their transfer is often uneven and can exhibit negative transfer. While most prior work improves transfer through architectural or adaptation choices, we ask a data-centric question: *which properties of …

Cited by 0SourceScholar
2025

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training

CVPR 2025poster

In rapidly evolving field of vision-language models (VLMs), contrastive language-image pre-training (CLIP) has made significant strides, becoming foundation for various downstream tasks. However, relying on one-to-one (image, text) contrastive paradigm to learn alignment from large-scale messy web d…

2025

FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching

ICCV 2025poster

Facial attractiveness prediction (FAP) has long been an important computer vision task, which could be widely applied in live videos with facial retouching. However, previous FAP datasets are either small or closed-source. Moreover, the corresponding FAP models exhibit limited generalization and ada…

2025

Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection

AAAI 2025technical

Large Language Models (LLMs) have revolutionized text generation, making detecting machine-generated text increasingly challenging. Although past methods have achieved good performance on detecting pure machine-generated text, those detectors have poor performance on distinguishing machine-revised t…

2025

ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection

IJCAI 2025

Auditory attention detection (AAD) aims to identify the direction of the attended speaker in multi-speaker environments from brain signals, such as Electroencephalography (EEG) signals. However, existing EEG-based AAD methods overlook the spatio-temporal dependencies of EEG signals, limiting their d

2025

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

IJCAI 2025

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which h

2025

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

ICCV 2025poster

Vision-language pretraining (VLP) enables open-world generalization beyond predefined labels, a critical capability in surgery due to the diversity of procedures, instruments, and patient anatomies. However, applying VLP to ophthalmic surgery presents unique challenges, including limited vision-lang…

2025

PEDE: Enhance Multi-modal Sarcasm Detection in Videos via Prompted Emotion Distributions

ICASSP 2025accepted

Multi-modal sarcasm detection is crucial for understanding human communications. A key aspect of multi-modal sarcasm detection is the analysis of emotion incongruity. However, the advancement of emotion analysis in video is hindered by the scarcity of labeled datasets, which are limited in both scal…

Cited by 0SourceScholar
2025

SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding

CVPR 2025poster

Despite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scal…

2025

Symbolic Representation for Any-to-Any Generative Tasks

CVPR 2025poster

We propose a symbolic generative task description language and a corresponding inference engine that can represent arbitrary multimodal tasks as structured symbolic flows. Unlike conventional generative models, which rely on large-scale training and implicit neural representations to learn cross-mod…

2024

3D Object Detection with VI-SLAM Point Clouds: The Impact of Object and Environment Characteristics on Model Performance

ICRA 2024poster

3D object detection (OD) is a crucial element in scene understanding. However, most existing 3D OD models have been tailored to work with light detection and ranging (LiDAR) and RGB-D point cloud data, leaving their performance on commonly available visual-inertial simultaneous localization and mapp…

Cited by 1SourceScholar
2024

ADHD Diagnosis and Biomarker Detection Based on Multimodal Graph Convolutional Neural Network

ICASSP 2024accepted

In this study, we apply a graph convolutional network (GCN) in attention deficit hyperactivity disorder (ADHD) classification by using multimodal data. Here, multimodal data is integrated to construct a dual graph for leveraging the modality information. Then, a GCN learning model is performed withi…

Cited by 0SourceScholar
2024

Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses

AISTATS 2024poster

Matrix sensing problems exhibit pervasive non-convexity, plaguing optimization with a proliferation of suboptimal spurious solutions. Avoiding convergence to these critical points poses a major challenge. This work provides new theoretical insights that help demystify the intricacies of the non-conv…

2024

Enhancing Quality of Compressed Images by Mitigating Enhancement Bias Towards Compression Domain

CVPR 2024poster

Existing quality enhancement methods for compressed images focus on aligning the enhancement domain with the raw domain to yield realistic images. However these methods exhibit a pervasive enhancement bias towards the compression domain inadvertently regarding it as more realistic than the raw domai…

Cited by 3SourcePDFScholar
2024

Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction

CVPR 2024poster

Whole Slide Image (WSI) classification is often formulated as a Multiple Instance Learning (MIL) problem. Recently Vision-Language Models (VLMs) have demonstrated remarkable performance in WSI classification. However existing methods leverage coarse-grained pathogenetic descriptions for visual repre…

2024

Toward Tiny and High-quality Facial Makeup with Data Amplify Learning

ECCV 2024poster

"Contemporary makeup approaches primarily hinge on unpaired learning paradigms, yet they grapple with the challenges of inaccurate supervision (e.g., face misalignment) and sophisticated facial prompts (including face parsing, and landmark detection). These challenges prohibit low-cost deployment of…

2024

Tuning-Free Image Customization with Image and Text Guidance

ECCV 2024poster

"Despite significant advancements in image customization with diffusion models, current methods still have several limitations: 1) unintended changes in non-target areas when regenerating the entire image; 2) guidance solely by a reference image or text descriptions; and 3) time-consuming fine-tunin…

2024

Unsupervised Continual Anomaly Detection with Contrastively-Learned Prompt

AAAI 2024technical

Unsupervised Anomaly Detection (UAD) with incremental training is crucial in industrial manufacturing, as unpredictable defects make obtaining sufficient labeled data infeasible. However, continual learning methods primarily rely on supervised annotations, while the application in UAD is limited due…

2023

ADHD Classification with Biomarker Identification Using a Triplet Loss Attention Auto-Encoding Network

ICASSP 2023accepted

Deep learning methods have been widely applied in Attention Deficit Hyperactivity Disorder (ADHD) classification in the past decade due to their effective learned features. However, these features are lack of neurobiological meanings and hard to be biomarkers. Here, we proposed an attention auto-enc…

Cited by 0SourceScholar
2023

Adaptive Assignment for Geometry Aware Local Feature Matching

CVPR 2023poster

The detector-free feature matching approaches are currently attracting great attention thanks to their excellent performance. However, these methods still struggle at large-scale and viewpoint variations, due to the geometric inconsistency resulting from the application of the mutual nearest neighbo…

2023

Copyright-Certified Distillation Dataset: Distilling One Million Coins into One Bitcoin with Your Private Key

AAAI 2023technical

The rapid development of neural network dataset distillation in recent years has provided new ideas in many areas such as continuous learning, neural network architecture search and privacy preservation. Dataset distillation is a very effective method to distill large training datasets into small da…

Cited by 2SourcePDFScholar
2023

MD-VQA: Multi-Dimensional Quality Assessment for UGC Live Videos

CVPR 2023poster

User-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing…

2023

TextShield: Beyond Successfully Detecting Adversarial Sentences in text classification

ICLR 2023poster

Adversarial attack serves as a major challenge for neural network models in NLP, which precludes the model's deployment in safety-critical applications. A recent line of work, detection-based defense, aims to distinguish adversarial sentences from benign ones. However, {the core limitation of previo…

Cited by 6SourcePDFScholar
2023

Word-level Prefix/Suffix Sense Detection: A Case Study on Negation Sense with Few-shot Learning

ACL 2023findings

Morphological analysis is an important research issue in the field of natural language processing. In this study, we propose a context-free morphological analysis task, namely word-level prefix/suffix sense detection, which deals with the ambiguity of sense expressed by prefix/suffix. To research th…

2022

AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-Time Image Enhancement

CVPR 2022poster

The 3D Lookup Table (3D LUT) is a highly-efficient tool for real-time image enhancement tasks, which models a non-linear 3D color transform by sparsely sampling it into a discretized 3D lattice. Previous works have made efforts to learn image-adaptive output color values of LUTs for flexible enhance…

Cited by 82PDFcodeScholar
2022

DuQM: A Chinese Dataset of Linguistically Perturbed Natural Questions for Evaluating the Robustness of Question Matching Models

EMNLP 2022main

In this paper, we focus on the robustness evaluation of Chinese Question Matching (QM) models. Most of the previous work on analyzing robustness issues focus on just one or a few types of artificial adversarial examples. Instead, we argue that a comprehensive evaluation should be conducted on natura…

2022

DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search Engine

EMNLP 2022main

In this paper, we present DuReader-retrieval, a large-scale Chinese dataset for passage retrieval. DuReader-retrieval contains more than 90K queries and over 8M unique passages from a commercial search engine. To alleviate the shortcomings of other datasets and ensure the quality of our benchmark, w…

2022

Dynamic Low-Resolution Distillation for Cost-Efficient End-to-End Text Spotting

ECCV 2022poster

"End-to-end text spotting has attached great attention recently due to its benefits on global optimization and high maintainability for real applications. However, the input scale has always been a tough trade-off since recognizing a small text instance usually requires enlarging the whole image, wh…

2022

Guide Local Feature Matching by Overlap Estimation

AAAI 2022technical

Local image feature matching under large appearance, viewpoint, and distance changes is challenging yet important. Conventional methods detect and match tentative local features across the whole images, with heuristic consistency checks to guarantee reliable matches. In this paper, we introduce a no…

2022

KATG: Keyword-Bias-Aware Adversarial Text Generation for Text Classification

AAAI 2022technical

Recent work has shown that current text classification models are vulnerable to small adversarial perturbation to inputs, and adversarial training that re-trains the models with the support of adversarial examples is the most popular way to alleviate the impact of the perturbation. However, current…

Cited by 6SourcePDFScholar
2022

Passive Inverted Ultra-Short Baseline Positioning for a Disc-Shaped Autonomous Underwater Vehicle: Design and Field Experiments

RA-L 2022

Underwater positioning is critical to autonomous underwater vehicles (AUVs) for navigation and geo-referencing. The rapid attenuation of the electromagnetic wave in the underwater environment prevents the use of traditional positioning methods such as the Global Positioning System, whereupon acousti

Cited by 32SourceScholar
2022

SepLUT: Separable Image-Adaptive Lookup Tables for Real-Time Image Enhancement

ECCV 2022poster

"Image-adaptive lookup tables (LUTs) have achieved great success in real-time image enhancement tasks due to their high efficiency for modeling color transforms. However, they embed the complete transform, including the color component-independent and the component-correlated parts, into only a sing…

2021

Development of an Intention-Based Adaptive Neural Cooperative Control Strategy for Upper-Limb Robotic Rehabilitation

RA-L 2021

Robotic rehabilitation therapy has become an important technology to recover the motor ability of disabled individuals. Clinical studies indicate that involving the active intention of patient into rehabilitation training contributes to promoting the performance of therapies. An adaptive neural coop

Cited by 33SourceScholar
2021

KDExplainer: A Task-oriented Attention Model for Explaining Knowledge Distillation

IJCAI 2021poster

Knowledge distillation (KD) has recently emerged as an efficacious scheme for learning compact deep neural networks (DNNs). Despite the promising results achieved, the rationale that interprets the behavior of KD has yet remained largely understudied. In this paper, we introduce a novel task-oriente…

2021

MANGO: A Mask Attention Guided One-Stage Scene Text Spotter

AAAI 2021technical

Recently end-to-end scene text spotting has become a popular research topic due to its advantages of global optimization and high maintainability in real applications. Most methods attempt to develop various region of interest (RoI) operations to concatenate the detection part and the sequence recog…

2021

Refining Language Models with Compositional Explanations

NeurIPS 2021spotlight

Pre-trained language models have been successful on text classification tasks, but are prone to learning spurious correlations from biased datasets, and are thus vulnerable when making inferences in a new domain. Prior work reveals such spurious patterns via post-hoc explanation algorithms which com…

2021

Temporal-Coded Deep Spiking Neural Network with Easy Training and Robust Performance

AAAI 2021technical

Spiking neural network (SNN) is promising but the development has fallen far behind conventional deep neural networks (DNNs) because of difficult training. To resolve the training problem, we analyze the closed-form input-output response of spiking neurons and use the response expression to build ab…

2020

End-to-End Emotion-Cause Pair Extraction with Graph Convolutional Network

COLING 2020main

Emotion-cause pair extraction (ECPE), which aims at simultaneously extracting emotion-cause pairs that express emotions and their corresponding causes in a document, plays a vital role in understanding natural languages. Considering that most emotions usually have few causes mentioned in their conte…

2020

High-Accuracy Classification of Attention Deficit Hyperactivity Disorder with L2, 1-Norm Linear Discriminant Analysis

ICASSP 2020accepted

Attention Deficit Hyperactivity Disorder (ADHD) is a high incidence of neurobehavioral disease in school-age children. Its neurobiological classification is meaningful for clinicians. The existing ADHD classification methods suffer from two problems, i.e., insufficient data and noise disturbance. He…

Cited by 0SourceScholar
2019

Adaptive Control of Aerobatic Quadrotor Maneuvers in the Presence of Propeller-Aerodynamic-Coefficient and Torque-Latency Time-Variations

ICRA 2019poster

We present a study of the dynamics and control of a 28-gram quadrotor during the execution of aerobatic maneuvers in the presence of propeller-aerodynamic-coefficient and torque-latency time-variations. First, through a momentum-theory-based analysis of the flow field surrounding the robot during ae…

Cited by 13SourceScholar
2019

Bee+: A 95-mg Four-Winged Insect-Scale Flying Robot Driven by Twinned Unimorph Actuators

RA-L 2019

We introduce Bee <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">+</sup> , a 95-mg four-winged microrobot with improved controllability and open-loop-response characteristics with respect to those exhibited by state-of-the-art two-winged microrobots wit

Cited by 53SourceScholar
2018

Nonlinear Adaptive Control of Quadrotor Multi-Flipping Maneuvers in the Presence of Time-Varying Torque Latency

IROS 2018poster

The dynamics of quadrotors are affected by time-varying torque latency, which can greatly alter the stability robustness and performance of the closed-loop control schemes employed for flight; this issue is especially relevant during the execution of aerobatic maneuvers such as high-speed multi-flip…

Cited by 14SourceScholar
2016

Generation and real-time implementation of high-speed controlled maneuvers using an autonomous 19-gram quadrotor

ICRA 2016

We present a new experimental method for the generation and real-time implementation of high-speed aerobatic maneuvers, including multiple flips, on a 19-gram autonomous quadrotor. A key element in the proposed approach is the design and experimental tuning of a gain scheduling control strategy in w

Cited by 21SourceScholar
2016

Spread spectrum compressed sensing MRI using chirp radio frequency pulses

ICASSP 2016accepted

Compressed sensing has shown great potential in reducing data acquisition time in magnetic resonance imaging (MRI). Recently, a spread spectrum compressed sensing MRI method modulates an image with a quadratic phase. It performs better than the conventional compressed sensing MRI with variable densi…

Cited by 0SourceScholar