← Search

Zhihong Liu

6 accepted papers

2026

X-MoGe: A Cross-Modal Adaptation Framework with Mixture-of-Experts and Geometry Guidance for Heterogeneous Collaborative Perception

ICML 2026poster

Multi-agent collaborative perception improves perception range and robustness in autonomous driving. However, most existing methods assume homogeneous sensors and perception networks, which is unrealistic in real-world heterogeneous systems. Differences in sensing modalities and independently traine…

Cited by 0SourceScholar
2025

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

ICML 2025poster

We introduce Orthus, a unified multimodal model that excels in generating interleaved images and text from mixed-modality inputs by simultaneously handling discrete text tokens and continuous image features under the \textbf{AR} modeling principle. The continuous treatment of visual signals minimize…

Cited by 8SourcePDFScholar
2024

Image-Based Distributed Predictive Visual Servo Control for Cooperative Tracking of Multiple Fixed-Wing UAVs

RA-L 2024

This letter proposes a novel approach combining the distributed model predictive control (DMPC) and the image-based visual servoing (IBVS) for cooperative tracking problem of multiple fixed-wing Unmanned Aerial Vehicles (UAVs) equipped with pan-tilt cameras. In particular, the target is unknown and

Cited by 8SourceScholar
2021

Multi-Scale Selective Feedback Network with Dual Loss for Real Image Denoising

IJCAI 2021poster

The feedback mechanism in the human visual system extracts high-level semantics from noisy scenes. It then guides low-level noise removal, which has not been fully explored in image denoising networks based on deep learning. The commonly used fully-supervised network optimizes parameters through pai…

Cited by 8SourcePDFScholar
2021

Pseudo 3D Auto-Correlation Network for Real Image Denoising

CVPR 2021poster

The extraction of auto-correlation in images has shown great potential in deep learning networks, such as the self-attention mechanism in the channel domain and the self-similarity mechanism in the spatial domain. However, the realization of the above mechanisms mostly requires complicated module st…

Cited by 35PDFScholar
2019

P3SGD: Patient Privacy Preserving SGD for Regularizing Deep CNNs in Pathological Image Classification

CVPR 2019poster

Recently, deep convolutional neural networks (CNNs) have achieved great success in pathological image classification. However, due to the limited number of labeled pathological images, there are still two challenges to be addressed: (1) overfitting: the performance of a CNN model is undermined by th…

Cited by 57PDFScholar