← Search

Zhenyu Wang

35 accepted papers

2026

IAFMNet: Information-Aware Feature Modulation for Efficient Super-Resolution

CVPR 2026

Single Image Super-Resolution (SISR) aims to reconstruct a high-resolution (HR) image from a low-resolution (LR) input, a task that becomes increasingly challenging under real-world computational constraints. However, most efficient SISR methods adopt lightweight, spatially uniform strategies that a

Cited by 0SourceScholar
2026

MUSE: Multimodal Uncertainty-Based Self-Driven Evolution for Robust Physiological-Signal–Based Driver Fatigue Detection

AAAI 2026technical

Precise detection of driver mental fatigue is critical for reducing traffic accidents and enhancing road safety. Compared with vision-based detection—which is susceptible to illumination and occlusion—multimodal physiological‑signal-based approaches integrate complementary information from diverse

Cited by 0SourcePDFScholar
2026

Tracking through Severe Occlusion via Event-Derived Transient Cues

CVPR 2026

Tracking targets with high-speed and nonlinear motion under occlusion remains challenging due to spatial appearance deprivation and temporal trajectory fragmentation caused by missing visual cues. Existing methods typically either dynamically update templates to maintain appearance similarity or emp

Cited by 0SourceScholar
2025

An Efficient Dialogue Policy Agent with Model-Based Causal Reinforcement Learning

COLING 2025main

Dialogue policy trains an agent to select dialogue actions frequently implemented via deep reinforcement learning (DRL). The model-based reinforcement methods built a world model to generate simulated data to alleviate the sample inefficiency. However, traditional world model methods merely consider…

Cited by 0SourcePDFScholar
2025

Bridging Task Boundaries: Remote Sensing Image-Text Retrieval via Dictionary-Driven Adaptation

ICASSP 2025accepted

Given image (or text), remote sensing image-text retrieval (RSITR) aims to retrieve corresponding text (or image) within diverse remote sensing data. However, due to the complex scenes and compact distribution of targets in remote sensing data, existing methods, particularly those leveraging large m…

Cited by 0SourceScholar
2025

CMIF-VIO: A Novel Cross Modal Interaction Framework for Visual Inertial Odometry

RA-L 2025

Visual Inertial Odometry (VIO) estimates predicted trajectories through self motion. With the popularization of artificial intelligence, deep learning-based VIO methods have shown better performance than traditional geometry-based VIO methods. However, in deep learning methods, how to better achieve

Cited by 5SourceScholar
2025

Geometry and Force-Informed Robotic Assembly with Small Relative Initial Deviations for Circular Electrical Connectors

ICRA 2025

Circular electrical connectors (CECs) have a wide range of applications in scenarios that require reliable connections. However, sockets are often located in narrow scenes with random spatial orientations, complex lighting conditions, and obstructions from cables, making it difficult to accurately l

Cited by 0SourceScholar
2025

Large Language Models in Bioinformatics: A Survey

ACL 2025finding

Large Language Models (LLMs) are revolutionizing bioinformatics, enabling advanced analysis of DNA, RNA, proteins, and single-cell data. This survey provides a systematic review of recent advancements, focusing on genomic sequence modeling, RNA structure prediction, protein function inference, and s…

Cited by 0SourcePDFScholar
2025

Layered Image Vectorization via Semantic Simplification

CVPR 2025poster

This work presents a progressive image vectorization technique that reconstructs the raster image as layer-wise vectors from semantic-aligned macro structures to finer details. Our approach introduces a new image simplification method leveraging the feature-average effect in the Score Distillation S…

Cited by 6SourcePDFScholar
2025

PatternCIR Benchmark and TisCIR: Advancing Zero-Shot Composed Image Retrieval in Remote Sensing

IJCAI 2025

Remote sensing composed image retrieval (RSCIR) is a new vision-language task that takes a composed query of an image and text, aiming to search for a target remote sensing image satisfying two conditions from intricate remote sensing imagery. However, the existing attribute-based benchmark Patternc

Cited by 0SourcePDFScholar
2025

PortLLM: Personalizing Evolving Large Language Models with Training-Free and Portable Model Patches

ICLR 2025poster

As large language models (LLMs) increasingly shape the AI landscape, fine-tuning pretrained models has become more popular than in the pre-LLM era for achieving optimal performance in domain-specific tasks. However, pretrained LLMs such as ChatGPT are periodically evolved (i.e., model parameters are…

Cited by 1SourcePDFScholar
2025

RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

RSS 2025poster

Developing robust and general-purpose manipulation policies is a key goal in robotics. To achieve effective generalization, it is essential to construct comprehensive datasets that encompass a large number of demonstration trajectories and diverse tasks. Unlike vision or language data, which can be…

Cited by 20PDFScholar
2024

BVT-IMA: Binary Vision Transformer with Information-Modified Attention

AAAI 2024technical

As a compression method that can significantly reduce the cost of calculations and memories, model binarization has been extensively studied in convolutional neural networks. However, the recently popular vision transformer models pose new challenges to such a technique, in which the binarized model…

Cited by 1SourcePDFScholar
2024

GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing

NeurIPS 2024spotlight

Despite the success achieved by existing image generation and editing methods, current models still struggle with complex problems including intricate text prompts, and the absence of verification and self-correction mechanisms makes the generated images unreliable. Meanwhile, a single model tends…

Cited by 23SourcePDFScholar
2024

One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection

NeurIPS 2024poster

The current trend in computer vision is to utilize one universal model to address all various tasks. Achieving such a universal model inevitably requires incorporating multi-domain data for joint training to learn across multiple problem scenarios. In point cloud based 3D object detection, however,…

Cited by 2SourcePDFScholar
2024

RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

NeurIPS 2024poster

A fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they still face challenges in two areas: (1) insufficient reasoning ability to tackle…

Cited by 5SourcePDFScholar
2023

A Novel Metamorphic Foot Mechanism With Toe Joints Based on Spring-Loaded Linkages

RA-L 2023

The toe joints play an important role in human walking and running movement patterns. In this letter, we propose a design method for metamorphic foot structures based on spring-loaded linkages to realize metatarsal-toe switching by using the self-recovery and self-stabilization properties of the spr

Cited by 4SourceScholar
2023

BSH-Det3D: Improving 3D Object Detection with BEV Shape Heatmap

IROS 2023poster

The progress of LiDAR-based 3D object detection has significantly enhanced developments in autonomous driving and robotics. However, due to the limitations of LiDAR sensors, object shapes suffer from deterioration in occluded and distant areas, which creates a fundamental challenge to 3D perception.…

Cited by 7SourcecodeScholar
2023

Detecting Everything in the Open World: Towards Universal Object Detection

CVPR 2023poster

In this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We…

2023

Disentangled Training with Adversarial Examples for Robust Small-Footprint Keyword Spotting

ICASSP 2023accepted

A keyword spotting (KWS) engine continuously running on the device is exposed to various speech signals that are usually unseen beforehand. It is a challenging problem to build a small-footprint and high-performing KWS model with robustness under different acoustic environments. In this paper, we ex…

Cited by 0SourceScholar
2023

HQP-MVS:High-Quality Plane Priors Assisted Multi-View Stereo for Low-Textured Areas

ICASSP 2023accepted

The completeness of reconstructed models in low-textured areas in multi-view stereo is still a challenging problem because of the unreliable photometric consistency. Since these areas always exhibit planar properties, many methods explicitly construct planar priors to assist in optimizing depth esti…

Cited by 0SourceScholar
2023

Haptic Dataset Augmentation with Subjective QoE Labels using Conditional Generative Adversarial Network

IROS 2023poster

This paper proposes a novel Generative Adversarial Network (GAN)-based strategy to augment subjective haptic Quality of Experience (QoE) datasets for bilateral teleoperation with haptic feedback without conducting time-consuming subjective experiments. In our previous work, we proposed a multi-asses…

Cited by 1SourceScholar
2023

SoulChat: Improving LLMs' Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations

EMNLP 2023short findings

Large language models (LLMs) have been widely applied in various fields due to their excellent capability for memorizing knowledge and chain of thought (CoT). When these language models are applied in the field of psychological counseling, they often rush to provide universal advice. However, when u…

Cited by 0SourcecodeScholar
2023

Uni3DETR: Unified 3D Detection Transformer

NeurIPS 2023poster

Existing point cloud based 3D detectors are designed for the particular scene, either indoor or outdoor ones. Because of the substantial differences in object distribution and point density within point clouds collected from various environments, coupled with the intricate nature of 3D metrics, ther…

2022

Graph Learning Based Autoencoder for Hyperspectral Band Selection

ICASSP 2022accepted

Hyperspectral band selection aims to identify an optimal sub-set of bands from hyperspectral images (HSIs). Most existing methods explore the relationships between pair-wise pixels in a fixed graph. However, the quality of the initial fixed graph may be influenced by noises and user-defined paramete…

Cited by 0SourceScholar
2022

Hybrid Physical Metric For 6-DoF Grasp Pose Detection

ICRA 2022poster

6-DoF grasp pose detection of multi-grasp and multi-object is a challenge task in the field of intelligent robot. To imitate human reasoning ability for grasping objects, data driven methods are widely studied. With the introduction of large-scale datasets, we discover that a single physical metric…

Cited by 22SourcecodeScholar
2022

Noisy Boundaries: Lemon or Lemonade for Semi-Supervised Instance Segmentation?

CVPR 2022poster

Current instance segmentation methods rely heavily on pixel-level annotated images. The huge cost to obtain such fully-annotated images restricts the dataset scale and limits the performance. In this paper, we formally address semi-supervised instance segmentation, where unlabeled images are employe…

Cited by 40PDFcodeScholar
2022

Rethinking Depth Estimation for Multi-View Stereo: A Unified Representation

CVPR 2022poster

Depth estimation is solved as a regression or classification problem in existing learning-based multi-view stereo methods. Although these two representations have recently demonstrated their excellent performance, they still have apparent shortcomings, e.g., regression methods tend to overfit due to…

Cited by 168PDFcodeScholar
2022

VTC-LFC: Vision Transformer Compression with Low-Frequency Components

NeurIPS 2022accept

Although Vision transformers (ViTs) have recently dominated many vision tasks, deploying ViT models on resource-limited devices remains a challenging problem. To address such a challenge, several methods have been proposed to compress ViTs. Most of them borrow experience in convolutional neural netw…

Cited by 38SourcePDFScholar
2021

Automatic Curriculum Learning With Over-repetition Penalty for Dialogue Policy Learning

AAAI 2021technical

Dialogue policy learning based on reinforcement learning is difficult to be applied to real users to train dialogue agents from scratch because of the high cost. User simulators, which choose random user goals for the dialogue agent to train on, have been considered as an affordable substitute for r…

Cited by 20SourcePDFScholar
2021

Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification

NeurIPS 2021poster

Semi-supervised learning aims to leverage a large amount of unlabeled data for performance boosting. Existing works primarily focus on image classification. In this paper, we delve into semi-supervised learning for object detection, where labeled data are more labor-intensive to collect. Current met…

Cited by 33SourcePDFScholar
2021

Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection

CVPR 2021poster

In this paper, we delve into semi-supervised object detection where unlabeled images are leveraged to break through the upper bound of fully-supervised object detection models. Previous semi-supervised methods based on pseudo labels are severely degenerated by noise and prone to overfit to noisy lab…

Cited by 101PDFScholar
2021

Efficient Dialogue Complementary Policy Learning via Deep Q-network Policy and Episodic Memory Policy

EMNLP 2021main

Deep reinforcement learning has shown great potential in training dialogue policies. However, its favorable performance comes at the cost of many rounds of interaction. Most of the existing dialogue policy methods rely on a single learning system, while the human brain has two specialized learning a…

Cited by 13SourcePDFScholar
2020

A multi-view approach for Mandarin non-native mispronunciation verification

ICASSP 2020accepted

Traditionally, the performance of non-native mispronunciation verification systems relied on effective phone-level labelling of non-native corpora. In this study, a multi-view approach is proposed to incorporate discriminative feature representations which requires less annotation for non-native mis…

Cited by 0SourceScholar