← Search

Bo Zhao

57 accepted papers

2026

Compensating Distribution Drifts in Continual Learning with Pre-trained Vision Transformers

AAAI 2026technical

Recent advances have shown that sequential fine-tuning (SeqFT) of pre-trained vision transformers (ViTs), followed by classifier refinement using approximate distributions of class features, can be an effective strategy for class-incremental learning (CIL). However, this approach is susceptible to d

Cited by 0SourcePDFScholar
2026

Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success

ICML 2026poster

Model merging combines knowledge from separately fine-tuned models, yet success factors remain poorly understood. While recent work treats mergeability as an intrinsic property, we show with an architecture-agnostic framework that it fundamentally depends on both the merging method and the partner t…

Cited by 0SourceScholar
2026

Emergence of Hierarchical Emotion Organization in Large Language Models

ICML 2026poster

As large language models (LLMs) increasingly power conversational agents, understanding how they model users' emotional states is critical for ethical deployment. Inspired by emotion wheels, i.e., a psychological framework that argues emotions organize hierarchically, we analyze probabilistic depend…

Cited by 0SourceScholar
2026

Encode Geometric Diagram as Geo-Graph in Geometry Problem Solving

AAAI 2026technical

Geometry Problem Solving has become a hot topic these years due to its complexity of enabling the machine with geometric abstraction, multi-modal reasoning and mathematical capabilities. Majority of research works place their attention on the fusion of multi-modal data or the synergistic combination

Cited by 0SourcePDFScholar
2026

Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment

CVPR 2026

Vision-Language-Action (VLA) models have emerged as a powerful framework that unifies perception, language, and control, enabling robots to perform diverse tasks through multimodal understanding. However, current VLA models typically contain massive parameters and rely heavily on large-scale robot d

Cited by 0SourcecodeScholar
2026

FLOW: Optimal Transport-Driven Feature Warping for Generalized Remote Physiological Measurement

CVPR 2026

Remote photoplethysmography (rPPG) enables non-contact physiological measurement from facial videos but often suffers from severe performance degradation under domain shifts. Traditional STMap-based methods [??] rely on predefined spatio-temporal representations that offer engineered robustness but

Cited by 0SourceScholar
2026

MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs

ICLR 2026poster

Developing Large Language Models (LLMs) to cooperate and compete effectively within multi-agent systems (MASs) is a critical step towards more advanced intelligence. While reinforcement learning (RL) has proven effective for enhancing reasoning in single-agent tasks, its extension to multi-turn, mul…

Cited by 13SourcecodeScholar
2026

Mask2IV: Interaction-Centric Video Generation via Mask Trajectories

AAAI 2026technical

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training, and affordance reasoning. However, existing methods often s

Cited by 0SourcePDFScholar
2026

PHASE-Net: Physics-Grounded Harmonic Attention System for Efficient Remote Photoplethysmography Measurement

CVPR 2026

Remote photoplethysmography (rPPG) measurement enables non-contact physiological monitoring but suffers from accuracy degradation under head motion and illumination changes. Existing deep learning methods are mostly heuristic and lack theoretical grounding, limiting robustness and interpretability.

Cited by 0SourcecodeScholar
2026

PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing

ICLR 2026poster

Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains highly susceptible to illumination changes, motion artifacts, and limited temporal modeling. Large Language Models (LLMs) excel at capturing long-range dependencies, offering a potential solution but struggl…

Cited by 0SourceScholar
2026

TexEditor: Structure-Preserving Text-Driven texture Editing

ICML 2026poster

Text-guided texture editing aims to modify object appearance while preserving the underlying geometric structure. However, our empirical analysis reveals that even SOTA editing models frequently struggle to maintain structural consistency during texture editing, despite the intended changes being pu…

Cited by 0SourceScholar
2026

Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents

ICML 2026poster

Large language models (LLMs) are increasingly deployed as autonomous agents for multi-turn decision-making tasks. However, current agents typically rely on fixed cognitive patterns: non-thinking models generate immediate responses, while thinking models engage in deep reasoning uniformly. This rigid…

Cited by 0SourceScholar
2026

Unsupervised Camouflaged Object Detection with Dual-Eigenvector Spectral Pseudo-Labeling and Contrastive Refinement

ICML 2026poster

Unsupervised Camouflaged Object Detection (UCOD) aims to identify objects concealed in their surroundings without relying on pixel-level labels. Existing methods rely solely on simple post-processing of DINO high-dimensional features to generate pseudo labels for training. However, these methods suf…

Cited by 0SourceScholar
2025

BOOD: Boundary-based Out-Of-Distribution Data Generation

ICML 2025poster

Harnessing the power of diffusion models to synthesize auxiliary training data based on latent space features has proven effective in enhancing out-of-distribution (OOD) detection performance. However, extracting effective features outside the in-distribution (ID) boundary in latent space remains ch…

Cited by 0SourcePDFScholar
2025

Causal-R: A Causal-Reasoning Geometry Problem Solver for Optimized Solution Exploration

NeurIPS 2025poster

The task of geometry problem solving has been a long-standing focus in the automated mathematics community and draws growing attention due to its complexity for both symbolic and neural models. Although prior studies have explored various effective approaches for enhancing problem solving performanc…

Cited by 0SourceScholar
2025

GIST: Guided Interpretable Large Language Model Strategy Transfer for Multi-Task Reinforcement Learning

ICASSP 2025accepted

Multi-task reinforcement learning (MTRL) presents critical challenges, such as the complexities of task-switching and maintaining knowledge retention across various tasks, especially within control tasks. These challenges frequently result in sub-optimal decision-making and catastrophic forgetting d…

Cited by 0SourceScholar
2025

MLVU: Benchmarking Multi-task Long Video Understanding

CVPR 2025poster

The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in vi…

2025

MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers

ICCV 2025poster

Fully comprehending scientific papers by machines reflects a high level of Artificial General Intelligence, requiring the ability to reason across fragmented and heterogeneous sources of information, presenting a complex and practically significant challenge. While Vision-Language Models (VLMs) have…

2025

MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval

ACL 2025long

Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massi…

2025

MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval

NeurIPS 2025poster

Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on the end-to-end LVU performance, making them inappropriate for…

Cited by 0SourceScholar
2025

Na Vid-4D: Unleashing Spatial Intelligence in Egocentric RGB-D Videos for Vision-and-Language Navigation

ICRA 2025

Understanding and reasoning about the 4D space-time is crucial for Vision-and-Language Navigation (VLN). However, previous works lack in-depth exploration in this aspect, resulting in bottlenecked spatial perception and action precision of VLN agents. In this work, we introduce NaVid-4D, a Vision La

Cited by 4SourceScholar
2025

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

ICCV 2025poster

The use of Multimodal Large Language Models (MLLMs) as an end-to-end solution for Embodied AI and Autonomous Driving has become a prevailing trend. While MLLMs have been extensively studied for visual semantic understanding tasks, their ability to perform precise and quantitative spatial-temporal un…

Cited by 0SourcePDFScholar
2025

SpatialBot: Precise Spatial Understanding with Vision Language Models

ICRA 2025

Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding; however, they still struggle with spatial understanding, which is fundamental to embodied AI. In this paper, we propose SpatialBot, a model designed to enhance spatial understanding by utilizing both RGB an

Cited by 167SourcecodeScholar
2025

Towards Universal Dataset Distillation via Task-Driven Diffusion

CVPR 2025poster

Dataset distillation (DD) condenses key information from large-scale datasets into smaller synthetic datasets, reducing storage and computational costs for training networks. However, recent research has primarily focused on image classification tasks, with limited expansion to detection and segment…

Cited by 0SourcePDFScholar
2025

Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly

CVPR 2025poster

**M**ultimodal **L**arge **L**anguage **M**odels (MLLMs) have displayed remarkable performance in multimodal tasks, particularly in visual comprehension. However, we reveal that MLLMs often generate incorrect answers even when they understand the visual content. To this end, we manually construct a…

2025

UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation

NeurIPS 2025poster

Data augmentation using generative models has emerged as a powerful paradigm for enhancing performance in computer vision tasks. However, most existing augmentation approaches primarily focus on optimizing intrinsic data attributes -- such as fidelity and diversity -- to generate visually high-quali…

Cited by 0SourceScholar
2025

Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

CVPR 2025poster

Long video understanding poses a significant challenge for current Multi-modal Large Language Models (MLLMs). Notably, the MLLMs are constrained by their limited context lengths and the substantial costs while processing long videos. Although several existing methods attempt to reduce visual tokens,…

2024

Fetch and Forge: Efficient Dataset Condensation for Object Detection

NeurIPS 2024poster

Dataset condensation (DC) is an emerging technique capable of creating compact synthetic datasets from large originals while maintaining considerable performance. It is crucial for accelerating network training and reducing data storage requirements. However, current research on DC mainly focuses o…

Cited by 1SourcePDFScholar
2024

Improving Convergence and Generalization Using Parameter Symmetries

ICLR 2024oral

In many neural networks, different values of the parameters may result in the same loss value. Parameter space symmetries are loss-invariant transformations that change the model parameters. Teleportation applies such transformations to accelerate optimization. However, the exact mechanism behind th…

2024

Omni6DPose: A Benchmark and Model for Universal 6D Object Pose Estimation and Tracking

ECCV 2024poster

"6D object pose estimation is crucial in the field of computer vision. However, it suffers from a significant lack of large-scale and diverse datasets, impeding comprehensive model evaluation and curtailing downstream applications. To address these issues, this paper introduces , a substantial bench…

Cited by 12SourcePDFScholar
2024

RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Multi-Modal Large Language Model Learning

RSS 2024poster

We need to trust robots that use often opaque AI methods. They need to explain themselves to us, and we need to trust their explanation. In this regard, explainability plays a critical role in trustworthy autonomous decision-making to foster transparency and acceptance among end users, especially in…

Cited by 83SourcePDFScholar
2024

Real-Fake: Effective Training Data Synthesis Through Distribution Matching

ICLR 2024poster

Synthetic training data has gained prominence in numerous learning tasks and scenarios, offering advantages such as dataset augmentation, generalization evaluation, and privacy preservation. Despite these benefits, the efficiency of synthetic data generated by current methodologies remains inferior…

2024

SegVol: Universal and Interactive Volumetric Medical Image Segmentation

NeurIPS 2024spotlight

Precise image segmentation provides clinical study with instructive information. Despite the remarkable progress achieved in medical image segmentation, there is still an absence of a 3D foundation segmentation model that can segment a wide range of anatomical categories with easy user interaction.…

2024

Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?

NeurIPS 2024poster

How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks…

2024

VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

ACL 2024long

Multi-modal retrieval becomes increasingly popular in practice. However, the existing retrievers are mostly text-oriented, which lack the capability to process visual information. Despite the presence of vision-language models like CLIP, the current methods are severely limited in representing the t…

2023

Accelerating Dataset Distillation via Model Augmentation

CVPR 2023highlight

Dataset Distillation (DD), a newly emerging field, aims at generating much smaller but efficient synthetic training datasets from large ones. Existing DD methods based on gradient matching achieve leading performance; however, they are extremely computationally intensive as they require continuously…

2023

DYffusion: A Dynamics-informed Diffusion Model for Spatiotemporal Forecasting

NeurIPS 2023poster

While diffusion models can successfully generate data and make predictions, they are predominantly designed for static images. We propose an approach for training diffusion models for dynamics forecasting that leverages the temporal dynamics encoded in the data, directly coupling it with the diffusi…

2023

Everyone's Preference Changes Differently: A Weighted Multi-Interest Model For Retrieval

ICML 2023poster

User embeddings (vectorized representations of a user) are essential in recommendation systems. Numerous approaches have been proposed to construct a representation for the user in order to find similar items for retrieval tasks, and they have been proven effective in industrial recommendation syste…

Cited by 9SourcePDFScholar
2023

Symmetries, Flat Minima, and the Conserved Quantities of Gradient Flow

ICLR 2023poster

Empirical studies of the loss landscape of deep networks have revealed that many local minima are connected through low-loss valleys. Yet, little is known about the theoretical origin of such valleys. We present a general framework for finding continuous symmetries in the parameter space, which carv…

2022

CAFE: Learning To Condense Dataset by Aligning Features

CVPR 2022poster

Dataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one. State-of-the-art approaches largely rely on learning the synthetic data by matching the gradients between the real and synthetic data batches. Despite the intu…

Cited by 277PDFcodeScholar
2022

FedInv: Byzantine-Robust Federated Learning by Inversing Local Model Updates

AAAI 2022technical

Federated learning (FL) is a privacy-preserving distributed machine learning paradigm that enables multiple clients to collaboratively train statistical models without disclosing raw training data. However, the inaccessible local training data and uninspectable local training process make FL suscept…

Cited by 61SourcePDFScholar
2022

LIMO: Latent Inceptionism for Targeted Molecule Generation

ICML 2022spotlight

Generation of drug-like molecules with high binding affinity to target proteins remains a difficult and resource-intensive task in drug discovery. Existing approaches primarily employ reinforcement learning, Markov sampling, or deep generative models guided by Gaussian processes, which can be prohib…

2018

Left-Right Comparative Recurrent Model for Stereo Matching

CVPR 2018poster

Leveraging the disparity information from both left and right views is crucial for stereo disparity estimation. Left-right consistency check is an effective way to enhance the disparity estimation by referring to the information from the opposite view. However, the conventional left-right consisten…

Cited by 115SourcePDFScholar
2018

MSplit LBI: Realizing Feature Selection and Dense Estimation Simultaneously in Few-shot and Zero-shot Learning

ICML 2018oral

It is one typical and general topic of learning a good embedding model to efficiently learn the representation coefficients between two spaces/subspaces. To solve this task, $L_{1}$ regularization is widely used for the pursuit of feature selection and avoiding overfitting, and yet the sparse estima…

Cited by 23SourcePDFScholar
2017

Memory-Augmented Attribute Manipulation Networks for Interactive Fashion Search

CVPR 2017poster

We introduce a new fashion search protocol where attribute manipulation is allowed within the interaction between users and search engines, e.g. manipulating the color attribute of the clothing from red to blue. It is particularly useful for image-based search when the query image cannot perfectly m…

Cited by 173PDFScholar