← Search

Lihua ZHang

51 accepted papers

2026

Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control

CVPR 2026

In computational pathology, understanding and generation have evolved along disparate paths: advanced understanding models already exhibit diagnostic-level competence, whereas generative models largely simulate pixels. Progress remains hindered by three coupled factors: the scarcity of large, high-q

Cited by 0SourcecodeScholar
2026

Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models

CVPR 2026

Multimodal biomedical Vision-Language Models (VLMs) exhibit immense potential in the field of Continual Learning (CL). However, they confront a core dilemma: how to preserve fine-grained intra-modality features while bridging the significant domain gap across different modalities. To address this ch

Cited by 0SourcecodeScholar
2026

Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology

ICLR 2026poster

Recent years have witnessed remarkable progress in multimodal learning within computational pathology. Existing models primarily rely on vision and language modalities; however, language alone lacks molecular specificity and offers limited pathological supervision, leading to representational bottle…

Cited by 0SourcecodeScholar
2026

KiRAS: Keyframe Guided Self-Imitation for Robust and Adaptive Skill Learning in Quadruped Robots

ICRA 2026poster

With advances in reinforcement learning and imitation learning, quadruped robots can acquire diverse skills within a single policy by imitating multiple skill-specific datasets. However, the lack of datasets on complex terrains limits the ability of such multi-skill policies to generalize effectivel…

2026

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-turn Dialogue

ICML 2026poster

Multimodal Large Language Models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermined by {hallucination snowballing}: a phenomenon where initial errors amplify across conversational turns, leading to a collapse in coherence. This f…

Cited by 0SourceScholar
2026

MUJICA: Multi-Skill Unified Joint Integration of Control Architecture for Wheeled-Legged Robots

ICRA 2026poster

Wheeled-legged robots hold promise for traversing complex terrains and offer superior mobility compared to legged robots. However, wheeled-legged robots must effectively balance both wheeled driving and legged control. Furthermore, due to noisy proprioceptive sensing and real-world motor constraints…

2026

ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation (VLN) requires agents to accurately perceive complex visual environments and reason over navigation instructions and histories. However, existing methods passively process redundant visual inputs and treat all historical contexts indiscriminately, resulting in ineffici

Cited by 0SourceScholar
2026

RENet: Fault-Tolerant Motion Control for Quadruped Robots Via Redundant Estimator Networks under Visual Collapse

ICRA 2026poster

Vision-based locomotion in outdoor environments presents significant challenges for quadruped robots. Accurate environmental prediction and effective handling of depth sensor noise during real-world deployment remain difficult, severely restricting the outdoor applications of such algorithms. To add…

2026

Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding

CVPR 2026

Document understanding is a long-standing practical task. Vision-Language Models (VLMs) have gradually become a primary approach in this domain, demonstrating effective performance on single-page tasks. However, their effectiveness diminishes when handling long documents. In such scenarios, clues ar

Cited by 0SourceScholar
2026

SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension

AAAI 2026technical

Satire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely visual forms, remains a challenging task for current vision-langu

Cited by 0SourcePDFScholar
2026

SpaCRD: Multimodal Deep Fusion of Histology and Spatial Transcriptomics for Cancer Region Detection

AAAI 2026technical

Accurate detection of cancer tissue regions (CTR) enables deeper analysis of the tumor microenvironment and offers crucial insights into treatment response. Traditional CTR detection methods, which typically rely on the rich cellular morphology in histology images, are susceptible to a high rate of

Cited by 0SourcePDFScholar
2026

UniMGS: Unifying Mesh and 3D Gaussian Splatting with Single-Pass Rasterization and Proxy-Based Deformation

AAAI 2026technical

Joint rendering and deformation of mesh and 3D Gaussian Splatting (3DGS) have significant value as both representations offer complementary advantages for graphics applications. However, due to differences in representation and rendering pipelines, existing studies render meshes and 3DGS separately,

Cited by 0SourcePDFScholar
2025

BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation

AAAI 2025technical

With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all necessary structures. Many current 3D scene generation methods…

2025

CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation

ICASSP 2025accepted

Automatic medical report generation (MRG), which possesses significant research value as it can aid radiologists in clinical diagnosis and report composition, has garnered increasing attention. Despite recent progress, generating accurate reports remains arduous due to the requirement for precise cl…

Cited by 0SourceScholar
2025

Continuous Control of Diverse Skills in Quadruped Robots Without Complete Expert Datasets

ICRA 2025

Learning diverse skills for quadruped robots presents significant challenges, such as mastering complex transitions between different skills and handling tasks of varying difficulty. Existing imitation learning methods, while successful, rely on expensive datasets to reproduce expert behaviors. Insp

Cited by 1SourceScholar
2025

Debiased Multimodal Understanding for Human Language Sequences

AAAI 2025technical

Human multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures…

Cited by 1SourcePDFScholar
2025

Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful Comparators

AAAI 2025technical

Despite their remarkable capabilities, Large Language Models (LLMs) are prone to generate responses that contradict verifiable facts, i.e., unfaithful hallucination content. Existing efforts generally focus on optimizing model parameters or editing semantic representations, which compromise the inte…

2025

MAFD: Fine-Grained Motion Style Transfer with Adaptive Signal Fusion

ICASSP 2025accepted

Motion style transfer allows for the swift switching of different styles within the same motion for virtual avatars, offering significant efficiency gains and enhanced motion diversity compared to traditional motion capture methods. However, many existing methods struggle with controlling fine detai…

Cited by 0SourceScholar
2025

MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation

CVPR 2025poster

Diffusion models have shown excellent performance in text-to-image generation. However, existing methods often suffer from performance bottlenecks when dealing with complex prompts involving multiple objects, characteristics, and relations. Therefore, we propose a Multi-agent Collaboration-based Co…

Cited by 0SourcePDFScholar
2025

MMPF: Multi-Modal Perception Framework for Abnormal Medical Condition Detection

AAAI 2025technical

As the global population ages and the incidence of chronic diseases increases, the demand for early detection of abnormal medical conditions is increasing. Traditional health monitoring methods often require significant resources and specialized personnel, limiting their widespread use. Leveraging a…

Cited by 0SourcePDFScholar
2025

Music-Driven Legged Robots: Synchronized Walking to Rhythmic Beats

ICRA 2025

We address the challenge of effectively controlling the locomotion of legged robots by incorporating precise frequency and phase characteristics, which is often ignored in locomotion policies that do not account for the periodic nature of walking. We propose a hierarchical architecture that integrat

Cited by 0SourcecodeScholar
2025

RENet: Fault-Tolerant Motion Control for Quadruped Robots via Redundant Estimator Networks Under Visual Collapse

RA-L 2025

Vision-based locomotion in outdoor environments presents significant challenges for quadruped robots. Accurate environmental prediction and effective handling of depth sensor noise during real-world deployment remain difficult, severely restricting the outdoor applications of such algorithms. To add

Cited by 4SourcecodeScholar
2025

SAF: Local Shape-aware Face-based Garment Collision Handling via Neural SDFs

ICASSP 2025accepted

Learning-based garment prediction presents an appealing alternative to physics-based methods owing to its high efficiency. However, the predicted garments can exhibit noticeable penetrations into the body. Many collision-handling methods operate in the point domain, which has an inherent limitation…

Cited by 0SourceScholar
2025

V-Fusion: 2D Detection-enhanced Multimodal 3D BEV Object Detection

ICASSP 2025accepted

Integrating information from multiple sensors enhances the performance of autonomous vehicle perception systems. However, current multimodal 3D object detection methods focus on unifying modalities into a bird’s-eye view (BEV) representation, which overlooks the inherent characteristics of camera pe…

Cited by 0SourceScholar
2024

A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing Modalities

AAAI 2024technical

Multimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of modality completeness. However, many inevitable factors in real-world scenarios lead to uncertain missing modalities, which invalidate the fixed multimodal fusion…

Cited by 13SourcePDFScholar
2024

CPR-Coach: Recognizing Composite Error Actions based on Single-class Training

CVPR 2024poster

Fine-grained medical action analysis plays a vital role in improving medical skill training efficiency but it faces the problems of data and algorithm shortage. Cardiopulmonary Resuscitation (CPR) is an essential skill in emergency treatment. Currently the assessment of CPR skills mainly depends on…

2024

Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities

CVPR 2024poster

Multimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However in real-world applications some practical factors cause uncertain modality missingness which drastically degrades the model's…

Cited by 15SourcePDFScholar
2024

De-confounded Data-free Knowledge Distillation for Handling Distribution Shifts

CVPR 2024poster

Data-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However a long-overlooked iss…

Cited by 6SourcePDFScholar
2024

HybridOcc: NeRF Enhanced Transformer-Based Multi-Camera 3D Occupancy Prediction

RA-L 2024

Vision-based 3D semantic scene completion (SSC) describes autonomous driving scenes through 3D volume representations. However, the occlusion of invisible voxels by scene surfaces poses challenges to current SSC methods in hallucinating refined 3D geometry. This letter proposes HybridOcc, a hybrid 3

Cited by 17SourceScholar
2024

Multi-Task Learning of Active Fault-Tolerant Controller for Leg Failures in Quadruped robots

ICRA 2024poster

Electric quadruped robots used in outdoor exploration are susceptible to leg-related electrical or mechanical failures. Unexpected joint power loss and joint locking can immediately pose a falling threat. Typically, controllers lack the capability to actively sense the condition of their own joints…

Cited by 5SourceScholar
2024

Offline Reinforcement Learning with Generative Adversarial Networks and Uncertainty Estimation

ICASSP 2024accepted

In recent years, offline reinforcement learning has attracted considerable attention in artificial intelligence. By generating a static dataset through a behavior policy, it is unable to engage in online interactions with the environment. However, this inevitably leads to states or actions undergoin…

Cited by 0SourceScholar
2024

Offline Reinforcement Learning with Policy Guidance and Uncertainty Estimation

ICASSP 2024accepted

Offline reinforcement learning is an approach for transforming static datasets into powerful decision engines. It cannot interact with the environment online, which leads to distribution shifts. Previous approaches addressed this problem by making the current policy as close as possible to the behav…

Cited by 0SourceScholar
2024

Optimal Auction Design with User Coupons in Advertising Systems

IJCAI 2024poster

Online advertising is a major revenue source for most Internet companies. The advertising opportunities are usually sold to advertisers through auctions that take into account the bids of the advertisers and the click-through rates (CTRs) and the conversion rates (CVRs) of the users. Standard auctio…

Cited by 0SourcePDFScholar
2024

PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications

NeurIPS 2024poster

Developing intelligent pediatric consultation systems offers promising prospects for improving diagnostic efficiency, especially in China, where healthcare resources are scarce. Despite recent advances in Large Language Models (LLMs) for Chinese medicine, their performance is sub-optimal in pediatri…

2024

Robust Emotion Recognition in Context Debiasing

CVPR 2024poster

Context-aware emotion recognition (CAER) has recently boosted the practical applications of affective computing techniques in unconstrained environments. Mainstream CAER methods invariably extract ensemble representations from diverse contexts and subject-centred characteristics to perceive the targ…

Cited by 23SourcePDFScholar
2024

Robust Proximal Adversarial Reinforcement Learning Under Model Mismatch

RA-L 2024

Reinforcement learning (RL) can generate high-performance control policies for complex tasks in simulation through an end-to-end approach. However, the RL policy is not robust to uncertainties caused by modeling mismatch between simulation and real environments, making it difficult to transfer to th

Cited by 3SourceScholar
2024

Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning

NeurIPS 2024poster

Multimodal Sentiment Analysis (MSA) is an important research area that aims to understand and recognize human sentiment through multiple modalities. The complementary information provided by multimodal fusion promotes better sentiment analysis compared to utilizing only a single modality. Neverthele…

Cited by 1SourcePDFScholar
2023

A Perturbation-Based Policy Distillation Framework with Generative Adversarial Nets

ICASSP 2023accepted

We study the problem of imitation learning in automated decision systems, in which a learner is trained to imitate an expert demonstrator. A widely used method is adversarial imitation learning that alternately optimizes a generator (learner) and a discriminator (reward function). However, the discr…

Cited by 0SourceScholar
2023

AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception

ICCV 2023poster

Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIs…

Cited by 54PDFcodeScholar
2023

Context De-Confounded Emotion Recognition

CVPR 2023poster

Context-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representa…

2023

D-CONFORMER: Deformable Sparse Transformer Augmented Convolution for Voxel-Based 3D Object Detection

ICASSP 2023accepted

Although CNN-based and Transformer-based detectors have made impressive improvements in 3D object detection, these two network paradigms suffer from the interference of insufficient receptive field and local detail weakening, which significantly limits the feature extraction performance of the backb…

Cited by 0SourceScholar
2023

How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent Perception

NeurIPS 2023poster

Multi-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay,…

2023

Learning Unbiased Rewards with Mutual Information in Adversarial Imitation Learning

ICASSP 2023accepted

A powerful method for automated decision systems is Adversarial Imitation Learning (AIL). It is based on a generative adversarial framework that alternately optimizes a generator (learner) and a discriminator (reward function). In the popular mind, a high-accuracy discriminator results in informativ…

Cited by 0SourceScholar
2023

Towards Simultaneous Segmentation Of Liver Tumors And Intrahepatic Vessels Via Cross-Attention Mechanism

ICASSP 2023accepted

Accurate visualization of liver tumors and their surrounding blood vessels is essential for noninvasive diagnosis and prognosis prediction of tumors. In medical image segmentation, there is still a lack of in-depth research on the simultaneous segmentation of liver tumors and peritumoral blood vesse…

Cited by 0SourceScholar
2022

CA-SpaceNet: Counterfactual Analysis for 6D Pose Estimation in Space

IROS 2022poster

Reliable and stable 6D pose estimation of un-cooperative space objects plays an essential role in on-orbit servicing and debris removal missions. Considering that the pose estimator is sensitive to background interference, this paper proposes a counterfactual analysis framework named CA-SpaceNet to…

Cited by 21SourcecodeScholar
2022

Emotion Recognition for Multiple Context Awareness

ECCV 2022poster

"Understanding emotion in context is a rising hotspot in the computer vision community. Existing methods lack reliable context semantics to mitigate uncertainty in expressing emotions and fail to model multiple context representations complementarily. To alleviate these issues, we present a context-…

2022

Robust Adversarial Reinforcement Learning with Dissipation Inequation Constraint

AAAI 2022technical

Robust adversarial reinforcement learning is an effective method to train agents to manage uncertain disturbance and modeling errors in real environments. However, for systems that are sensitive to disturbances or those that are difficult to stabilize, it is easier to learn a powerful adversary than…

Cited by 20SourcePDFScholar
2022

Stpointgcn: Spatial Temporal Graph Convolutional Network for Multiple People Recognition Using Millimeter-Wave Radar

ICASSP 2022accepted

Gait recognition is a new biometric technology, which aims to identify people by their walking posture. Compared with fingerprint recognition, face recognition and other technologies, gait recognition usually has the characteristics of long-distance non-contact and difficulty in camouflage. And comp…

Cited by 0SourceScholar