← Search

Ziang Zhang

18 accepted papers

2026

Cross-Modal Semantic Decoupling and Transfer for Text-to-Visible-Infrared Person Re-Identification

ICML 2026poster

Text-to-Image Person Re-Identification (TI-ReID) retrieves visible pedestrian images using text queries. Yet in low-light or nighttime settings, visible images lack sufficient identity details, while infrared images effectively capture pedestrian contours and textures. To enable all-day surveillance…

Cited by 0SourceScholar
2026

SpatialHand: Generative Object Manipulation from 3D Prespective

ICLR 2026poster

We introduce SpatialHand, a novel framework for generative object insertion with precise 3D control. Current generative object manipulation methods primarily operate within the 2D image plane, but often fail to grasp 3D scene complexities, leading to ambiguities in an object's 3D position, orientati…

Cited by 0SourceScholar
2026

SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation

CVPR 2026

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited attention to fine-grained spatial relationships, often producing images that appear

Cited by 0SourcecodeScholar
2025

Data-Efficiently Learn Large Language Model for Universal 3D Scene Perception

NAACL 2025findings

3D scene understanding has gained significant attention due to its wide range of applications. However, existing methods for 3D scene understanding are limited to specific downstream tasks, which hinders their practicality in real-world applications. This paper presents Chat-3D, which combines the 3…

Cited by 0SourcePDFScholar
2025

GenSpace: Benchmarking Spatially-Aware Image Generation

NeurIPS 2025poster

Humans can intuitively compose and arrange scenes in the 3D space for photography. However, can advanced AI image generators plan scenes with similar 3D spatial awareness when creating images from text or image prompts? We present GenSpace, a novel benchmark and evaluation pipeline to comprehensivel…

Cited by 0SourceScholar
2025

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

ICLR 2025poster

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Meanwhile, multimodal representation models have emerged as the foundation for these versatile multimodal understanding and generation pipeline. Models like CLIP, CLAP and ImageBind…

Cited by 11SourcePDFScholar
2025

OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup

ICLR 2025poster

Query-based sound separation (QSS) effectively isolate sound signals that match the content of a given query, enhancing the understanding of audio data. However, most existing QSS methods rely on a single modality for separation, lacking the ability to fully leverage homologous but heterogeneous inf…

2025

Orient Anything V2: Unifying Orientation and Rotation Understanding

NeurIPS 2025spotlight

This work presents Orient Anything V2, an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. Building upon Orient Anything V1, which defines orientation via a single unique front face, V2 extends this capability to handle objects w…

Cited by 0SourceScholar
2025

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models

ICML 2025poster

Orientation is a fundamental attribute of objects, essential for understanding their spatial pose and arrangement. However, practical solutions for estimating the orientation of open-world objects in monocular images remain underexplored. In this work, we introduce Orient Anything, the first foundat…

2025

SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language

CVPR 2025poster

Contrastive Language-Image Pre-training (CLIP) learns robust visual models through language supervision, making it a crucial visual encoding technique for various applications. However, CLIP struggles with comprehending spatial concepts in images, potentially restricting the spatial intelligence of…

2025

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

ICLR 2025poster

Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTo…

2024

Extending Multi-modal Contrastive Representations

NeurIPS 2024poster

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the expensive training costs limit their further development. Ins…

2024

FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion

ICML 2024poster

Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained unified spaces. In this work, we propose FreeBind, an idea that t…

2023

Connecting Multi-modal Contrastive Representations

NeurIPS 2023poster

Multi-modal Contrastive Representation (MCR) learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across various modalities. However, the reliance on massive high-quality data pairs l…

2022

A Robust Reference Path Selection Method for Path Planning Algorithm

RA-L 2022

In this letter, a general robust reference path selection method (RPSM) that can be integrated into current existing motion planning algorithms is proposed to improve the mobile performance of autonomous patrol robots. The proposed RPSM maintains a dynamic array of path candidates that contains newl

Cited by 14SourceScholar
2022

Direction and Trajectory Tracking Control for Nonholonomic Spherical Robot by Combining Sliding Mode Controller and Model Prediction Controller

RA-L 2022

A spherical robot is a nonlinear, nonholonomic, and unstable system which increases the difficulty of the direction and trajectory tracking problem. In this study, we propose a new direction controller Hierarchical Terminal Sliding Mode Controller (HTSMC), an instruction planning controller called M

Cited by 33SourceScholar
2021

Fuzzy PID Controller Based on Yaw Angle Prediction of a Spherical Robot

IROS 2021poster

In this paper, a fuzzy PID controller based on yaw angle prediction is applied to design an attitude controller for a spherical rolling robot. The robot consists of a 2-DOF pendulum located inside a spherical shell with freedom to rotate about the transversal and longitudinal axis. The proposed cont…

Cited by 24SourceScholar