← Search

zhen chen

26 accepted papers

2026

FlowAnyTime: Efficient Fine-tuning with Intra-Inter Frame Distillation for All-Weather Optical Flow Estimation

AAAI 2026technical

Motion estimation in degraded scenes has long been a significant challenge, primarily attributed to substantial scene variations and insufficient training data. Existing approaches typically address this limitation by incorporating additional training strategies or modifying network architectures wi

Cited by 0SourcePDFScholar
2026

Fragile by Design: On the Limits of Adversarial Defenses in Personalized DreamBooth Generation

AAAI 2026technical

Personalized AI applications such as DreamBooth enable the generation of customized content from user images, but also raise significant privacy concerns, particularly the risk of facial identity leakage. Recent defense mechanisms like Anti-DreamBooth attempt to mitigate this risk by injecting adver

Cited by 0SourcePDFScholar
2026

From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

CVPR 2026

There is growing interest in biomedical vision--language models trained on scientific literature. However, most pipelines compress rich multi-panel figures and long captions into coarse figure-level pairs, discarding the fine-grained correspondences clinicians rely on when zooming into local structu

Cited by 0SourceScholar
2026

Obstacle-Aware IBVS Target Tracking Via Feature-Space Projection and Virtual Imaging Guidance with ADP-Shaped Terminal Cost

ICRA 2026poster

We propose a visual-servoing and obstacle-avoidance controller for a wheeled mobile robot (WMR) with a two-axis gimbal camera that operates without mapping, using only vision and lightweight forward sensing. A task-allocation MPC with online terminal-cost iteration is introduced. Specifically, task …

Cited by 0Scholar
2026

ReWeaver: Towards Simulation-Ready and Topology-Accurate Garment Reconstruction

CVPR 2026

High-quality 3D garment reconstruction plays a crucial role in mitigating the sim-to-real gap in applications such as digital avatars, virtual try-on and robotic manipulation. However, existing garment reconstruction methods typically rely on unstructured representations, such as 3D Gaussian Splats,

Cited by 0SourcecodeScholar
2025

Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-Driven Surface Normal-Aware Tracking and Mapping

ICRA 2025

Simultaneous Localization and Mapping (SLAM) is essential for precise surgical interventions and robotic tasks in minimally invasive procedures. While recent advancements in 3D Gaussian Splatting (3DGS) have improved SLAM with high-quality novel view synthesis and fast rendering, these systems strug

Cited by 18SourcecodeScholar
2025

Adversarial Training for Probabilistic Robustness

ICCV 2025poster

Deep learning (DL) has shown transformative potential across industries, yet its sensitivity to adversarial examples (AEs) limits its reliability and broader deployment. Research on DL robustness has developed various techniques, with adversarial training (AT) established as a leading approach to co…

2025

Beyond Logits: Aligning Feature Dynamics for Effective Knowledge Distillation

ACL 2025long

Knowledge distillation (KD) compresses large language models (LLMs), known as teacher models, into lightweight versions called student models, enabling efficient inference and downstream applications. However, prevailing approaches accomplish this by predominantly focusing on matching the final outp…

2025

ICIMG-Net: Inject Context Information to Motion Generation for Optical Flow Estimation

ICASSP 2025accepted

Although the overall performance of existing optical flow estimation methods has improved rapidly, motion discontinuities caused by large displacements and occlusions remain significant challenges for accurate optical flow estimation. To address this issue, we propose a novel Inject Context Informat…

Cited by 0SourceScholar
2025

MotionFlow: Joint Motion Priors and Appearance Enhancement for High-Accuracy Optical Flow Estimation

ICASSP 2025accepted

Although optical flow estimation has improved significantly in recent years, large displacements and occlusions remain challenging for current methods due to motion discontinuities that may hinder accurate feature correspondences in these regions, leading to degraded performance. To address this cha…

Cited by 0SourceScholar
2025

POQD: Performance-Oriented Query Decomposer for Multi-vector retrieval

ICML 2025poster

Although Multi-Vector Retrieval (MVR) has achieved the state of the art on many information retrieval (IR) tasks, its performance highly depends on how to decompose queries into smaller pieces, say phrases or tokens. However, optimizing query decomposition for MVR performance is not end-to-end diffe…

2025

PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt Relocation

ICCV 2025accepted

Visual prompt tuning (VPT), i.e., fine-tuning some lightweight prompt tokens, provides an efficient and effective approach for adapting pre-trained models to various downstream tasks. However, most prior art indiscriminately uses a fixed prompt distribution across different tasks, neglecting the imp…

2025

Polyp-Gen: Realistic and Diverse Polyp Image Generation for Endoscopic Dataset Expansion

ICRA 2025

Automated diagnostic systems (ADS) have shown significant potential in the early detection of polyps during endoscopic examinations, thereby reducing the incidence of colorectal cancer. However, due to high annotation costs and strict privacy concerns, acquiring high-quality endoscopic images poses

Cited by 8SourcecodeScholar
2025

SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline Inference

ICRA 2025

Surgical phase recognition is critical for assisting surgeons in understanding surgical videos. Existing studies focused more on online surgical phase recognition, by leveraging preceding frames to predict the current frame. Despite great progress, they formulated the task as a series of frame-wise

Cited by 4SourcecodeScholar
2025

TANDEM: Bi-Level Data Mixture Optimization with Twin Networks

NeurIPS 2025poster

The capabilities of large language models (LLMs) significantly depend on training data drawn from various domains. Optimizing domain-specific mixture ratios can be modeled as a bi-level optimization problem, which we simplify into a single-level penalized form and solve with twin networks: a proxy m…

Cited by 0SourceScholar
2025

U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation

AAAI 2025technical

U-Net has become a cornerstone in various visual applications such as image segmentation and diffusion probability models. While numerous innovative designs and improvements have been introduced by incorporating transformers or MLPs, the networks are still limited to linearly modeling patterns as we…

2025

UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation

ICLR 2025poster

Pre-training techniques significantly enhance the performance of semantic segmentation tasks with limited training data. However, the efficacy under a large domain gap between pre-training (e.g. RGB) and fine-tuning (e.g. infrared) remains underexplored. In this study, we first benchmark the infrare…

2025

Universal Domain Adaptive Object Detection via Dual Probabilistic Alignment

AAAI 2025technical

Domain Adaptive Object Detection (DAOD) transfers knowledge from a labeled source domain to an unannotated target domain under closed-set assumption. Universal DAOD (UniDAOD) extends DAOD to handle open-set, partial-set, and closed-set domain adaptation. In this paper, we first unveil two issues: do…

2024

ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention Understanding

IROS 2024

Surgical instrument segmentation is crucial in surgical scene understanding, thereby facilitating surgical safety. Existing algorithms directly detected all instruments of predefined categories in the input image, lacking the capability to segment specific instruments according to the surgeon’s inte

Cited by 16SourcecodeScholar
2024

Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAM

NeurIPS 2024poster

As the vision foundation models like the Segment Anything Model (SAM) demonstrate potent universality, they also present challenges in giving ambiguous and uncertain predictions. Significant variations in the model output and granularity can occur with simply subtle changes in the prompt, contradict…

2024

GaussianGrasper: 3D Language Gaussian Splatting for Open-Vocabulary Robotic Grasping

RA-L 2024

Constructing a 3D scene capable of accommodating open-ended language queries, is a pivotal pursuit in the domain of robotics, which facilitates robots in executing object manipulations based on human language directives. To achieve this, some research efforts have been dedicated to the development o

Cited by 102SourcecodeScholar
2024

KGTS: Contrastive Trajectory Similarity Learning over Prompt Knowledge Graph Embedding

AAAI 2024technical

Trajectory similarity computation serves as a fundamental functionality of various spatial information applications. Although existing deep learning similarity computation methods offer better efficiency and accuracy than non-learning solutions, they are still immature in trajectory embedding and su…

Cited by 26SourcePDFScholar
2024

TARP-VP: Towards Evaluation of Transferred Adversarial Robustness and Privacy on Label Mapping Visual Prompting Models

NeurIPS 2024poster

Adversarial robustness and privacy of deep learning (DL) models are two widely studied topics in AI security. Adversarial training (AT) is an effective approach to improve the robustness of DL models against adversarial attacks. However, while models with AT demonstrate enhanced robustness, they be…

Cited by 0SourcePDFScholar
2022

Free Lunch for Cross-Domain Occluded Face Recognition without Source Data

ICASSP 2022accepted

Most recognizing occluded faces methods focus on synthetic-occluded faces for training due to the lack of real-occluded data. However, the performance may suffer from degradation since the synthetic-occluded and real-occluded face images are under different distributions. Hence, it draws our eyes to…

Cited by 0SourceScholar
2022

Improving Continual Relation Extraction through Prototypical Contrastive Learning

COLING 2022main

Continual relation extraction (CRE) aims to extract relations towards the continuous and iterative arrival of new data, of which the major challenge is the catastrophic forgetting of old tasks. In order to alleviate this critical problem for enhanced CRE performance, we propose a novel Continual Rel…

2021

Diagnose Like A Pathologist: Weakly-Supervised Pathologist-Tree Network for Slide-Level Immunohistochemical Scoring

AAAI 2021technical

The immunohistochemistry (IHC) test of biopsy tissue is crucial to develop targeted treatment and evaluate prognosis for cancer patients. The IHC staining slide is usually digitized into the whole-slide image (WSI) with gigapixels for quantitative image analysis. To perform a whole image prediction…

Cited by 44SourcePDFScholar