← Search

Zhenbo Song

14 accepted papers

2026

Neural–Evolutionary Symbolic Regression with Global Constraints: Constraint-Aware Decoding and Reward Shaping

ICML 2026poster

Symbolic regression discovers interpretable mathematical expressions from data and is central to scientific modeling. Recent neural approaches typically linearize expression trees into token sequences for sequential generation, but this representation weakens access to the underlying hierarchy and m…

Cited by 0SourceScholar
2026

SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite Imagery

ICLR 2026poster

Generating multiview-consistent $360^\circ$ ground-level scenes from satellite imagery is a challenging task with broad applications in simulation, autonomous navigation, and digital twin cities. Existing approaches primarily focus on synthesizing individual ground-view panoramas, often relying on a…

Cited by 0SourceScholar
2025

Controllable Satellite-to-Street-View Synthesis with Precise Pose Alignment and Zero-Shot Environmental Control

ICLR 2025poster

Generating street-view images from satellite imagery is a challenging task, particularly in maintaining accurate pose alignment and incorporating diverse environmental conditions. While diffusion models have shown promise in generative tasks, their ability to maintain strict pose alignment throughou…

Cited by 0SourcePDFScholar
2025

Gradient-Based Adversarial Attacks on Deep LiDAR Odometry

ICRA 2025

Adversarial attacks have been recently investigated in LiDAR perception problems for autonomous driving, where a small perturbation of source inputs can result in incorrect predictions. However, most previous studies focus on attacks on single-frame perception modules, lacking explorations of attack

Cited by 2SourceScholar
2025

Lightweight Yet High-Performance Defect Detector for Uav-Based Large-Scale Infrastructure Real-Time Inspection

ICRA 2025

Defect diagnosis in urban infrastructure is crucial for public safety. Traditional manual inspections face significant challenges in terms of accuracy and cost-effectiveness. In this paper, we propose a lightweight and hardware-friendly large-scale infrastructure detector, CUPID, highly suitable for

Cited by 2SourceScholar
2024

Harmonizing Stochasticity and Determinism: Scene-responsive Diverse Human Motion Prediction

NeurIPS 2024poster

Diverse human motion prediction (HMP) is a fundamental application in computer vision that has recently attracted considerable interest. Prior methods primarily focus on the stochastic nature of human motion, while neglecting the specific impact of external environment, leading to the pronounced art…

Cited by 4SourcePDFScholar
2023

EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

ICCV 2023poster

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content. To address this issue, this paper proposes an end-to-end neur…

Cited by 117PDFcodeScholar
2023

GIDP: Learning a Good Initialization and Inducing Descriptor Post-enhancing for Large-scale Place Recognition

ICRA 2023poster

Large-scale place recognition is a fundamental but challenging task, which plays an increasingly important role in autonomous driving and robotics. Existing methods have achieved acceptable good performance, however, most of them are concentrating on designing elaborate global descriptor learning ne…

Cited by 0SourceScholar
2023

Learning Dense Flow Field for Highly-accurate Cross-view Camera Localization

NeurIPS 2023poster

This paper addresses the problem of estimating the 3-DoF camera pose for a ground-level image with respect to a satellite image that encompasses the local surroundings. We propose a novel end-to-end approach that leverages the learning of dense pixel-wise flow fields in pairs of ground and satellite…

Cited by 9SourcePDFScholar
2023

Robust Single Image Reflection Removal Against Adversarial Attacks

CVPR 2023poster

This paper addresses the problem of robust deep single-image reflection removal (SIRR) against adversarial attacks. Current deep learning based SIRR methods have shown significant performance degradation due to unnoticeable distortions and perturbations on input images. For a comprehensive robustnes…

2022

Object Level Depth Reconstruction for Category Level 6D Object Pose Estimation from Monocular RGB Image

ECCV 2022poster

"Recently, RGBD-based category-level 6D object pose estimation has achieved promising improvement in performance, however, the requirement of depth information prohibits broader applications. In order to relieve this problem, this paper proposes a novel approach named Object Level Depth reconstructi…

Cited by 34SourcePDFScholar
2022

SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

AAAI 2022technical

Simultaneous Localization and Mapping (SLAM) and Autonomous Driving are becoming increasingly more important in recent years. Point cloud-based large scale place recognition is the spine of them. While many models have been proposed and have achieved acceptable performance by learning short-range lo…

Cited by 78SourcePDFScholar
2020

Deep Novel View Synthesis from Colored 3D Point Clouds

ECCV 2020poster

We propose a new deep neural network which takes a colored 3D point cloud of a scene, and directly synthesizes a photo-realistic image from an arbitrary viewpoint. Key contributions of this work include a deep point feature extraction module, an image synthesis module, and an image refinement module…

2020

End-to-end Learning for Inter-Vehicle Distance and Relative Velocity Estimation in ADAS with a Monocular Camera

ICRA 2020poster

Inter-vehicle distance and relative velocity estimations are two basic functions for any ADAS (Advanced driver-assistance systems). In this paper, we propose a monocular camera based inter-vehicle distance and relative velocity estimation method based on end-to-end training of a deep neural network.…

Cited by 26SourceScholar