← Search

Fang Zhao

31 accepted papers

2026

DART: Navigating Last-Mile Heterogeneity in Instant Delivery via Distribution-Adaptive Splines

IJCAI 2026

On-demand delivery platforms rely on Travel Time Estimation (TTE) to balance courier earnings and overdue risks. In collaboration with one of China's largest platforms, we address a critical "Fairness Gap" in TTE: current systems fail to capture complex delivery patterns in GNSS-denied environments,

Cited by 0Scholar
2026

HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment

AAAI 2026technical

Contrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-fo

Cited by 0SourcePDFScholar
2026

HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal Grounding

ICLR 2026poster

Temporal grounding in long, untrimmed videos is critical for real-world video understanding, yet it remains a challenging task owing to complex temporal structures and pervasive visual redundancy. Existing methods rely heavily on supervised training with task-specific annotations, which inherently l…

Cited by 0SourceScholar
2026

Matting Anything 2: Towards Video Matting for Anything

ICLR 2026poster

Video matting is a crucial task for many applications, but existing methods face significant limitations. They are often domain-specific, focusing primarily on human portraits, and rely on the mask of first frame that is challenging to acquire for transparent or intricate objects like fire or smoke.…

Cited by 0SourceScholar
2026

MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference

ICLR 2026poster

We present MeanCache, a training-free caching framework for efficient Flow Matching inference. Existing caching methods reduce redundant computation but typically rely on instantaneous velocity information (e.g., feature caching), which often leads to severe trajectory deviations and error accumulat…

Cited by 0SourceScholar
2026

One-to-More: High-Fidelity Training-Free Anomaly Generation with Attention Control

CVPR 2026

Industrial anomaly detection (AD) is characterized by an abundance of normal images but a scarcity of anomalous ones. Although numerous few-shot anomaly synthesis methods have been proposed to augment anomalous data for downstream AD tasks, most existing approaches require time-consuming training an

Cited by 0SourceScholar
2026

ST-DiffPlanner: A Safety-Enhanced Topology-Aware Diffusion Planner for Global Path Planning

ICRA 2026poster

In complex environments, traditional path planning methods rely on manually defined models, requiring tedious adjustments under varying scenarios or constraints. They also suffer from unstable time overhead and exponentially increasing computational costs as environmental complexity grows. Deep lear…

Cited by 0Scholar
2026

TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven Learning

ICML 2026poster

Visual data is typically a prerequisite for training existing video anomaly detection (VAD) methods. However, obtaining sufficient annotated anomaly data for training is challenging and not scalable due to the rarity of anomaly data and the wide variety of abnormal events. In this work, we advocate …

Cited by 0SourceScholar
2025

A Conditional KAN Diffusion Network for Human Activity Recognition with Missing Sensor Signal Series

ICASSP 2025accepted

Human Activity Recognition (HAR) is crucial for applications like urban traffic management and health monitoring but faces challenges in handling complex patterns and missing sensor data. In this work, we propose a conditional Kolmogorov-Arnold network diffusion (CKAD) framework for HAR, which separ…

Cited by 0SourceScholar
2025

CSS: Overcoming Pose and Scene Challenges in Crowd-Sourced 3D Gaussian Splatting

ICASSP 2025accepted

We introduce Crowd-Sourced Splatting (CSS), a novel 3D Gaussian Splatting (3DGS) pipeline designed to overcome the challenges of pose-free scene reconstruction using crowd-sourced imagery. The dream of reconstructing historically significant but inaccessible scenes from collections of photographs ha…

Cited by 0SourceScholar
2025

DGS-SLAM: A Visual Dense SLAM Based on Gaussian Splatting in Dynamic Environments

ICRA 2025

Visual dense SLAM can facilitate pose estimation and map reconstruction for sensor carriers in unknown environments. However, in uncontrolled environments such as offices, shopping malls, and train stations, frequent occurrences of people walking back and forth or temporary movement of objects withi

Cited by 1SourceScholar
2025

Kernel-Aware Graph Prompt Learning for Few-Shot Anomaly Detection

AAAI 2025technical

Few-shot anomaly detection (FSAD) aims to detect unseen anomaly regions with the guidance of very few normal support images from the same class. Existing FSAD methods usually find anomalies by directly designing complex text prompts to align them with visual features under the prevailing large visio…

2025

LOG-SLAM: Large-Scale Outdoor Gaussian SLAM for Dense Mapping and Loop Closure in Kilometer-Scale Scene Reconstruction

IROS 2025

The success of 3D Gaussian splatting in 3D reconstruction has recently led to efforts to integrate it with SLAM systems. However, most existing research has focused on indoor tracking and mapping, while outdoor Gaussian SLAM methods still heavily rely expensive LiDAR sensor. To address these challen

Cited by 0SourceScholar
2025

LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation

NeurIPS 2025spotlight

We present LeMiCa, a training-free and efficient acceleration framework for diffusion-based video generation. While existing caching strategies primarily focus on reducing local heuristic errors, they often overlook the accumulation of global errors, leading to noticeable content degradation between…

Cited by 0SourceScholar
2025

Map-Free Visual Relocalization Enhanced by Instance Knowledge and Depth Knowledge

ICASSP 2025accepted

Map-free visual relocalization computes camera pose using only a query image and a reference image. Therefore, it is hindered by challenges in feature-point matching and the absence of scale information in monocular images. These issues may cause significant rotational and metric errors, leading to…

Cited by 0SourceScholar
2025

SDD-SLAM: Semantic-Driven Dynamic SLAM With Gaussian Splatting

RA-L 2025

Recently, significant advancements have been made in 3D Gaussian Splatting SLAM for dynamic environments. However, most existing methods primarily address active dynamic objects, such as people and vehicles, and fail to account for the impact of passive dynamic objects on localization and mapping. T

Cited by 10SourceScholar
2024

SMORE-SLAM: Semantic Monocular SLAM with Scale Correction and Reverse Loop Utilization in Outdoor Environments

IROS 2024poster

In large-scale outdoor environments, vehicles often encounter situations like retracing their path or turning around, leading to many reverse loop closures where the vehicles traverse previously covered paths from opposite viewpoints. Existing monocular SLAM methods, due to insufficient utilization…

Cited by 0SourceScholar
2024

SynSP: Synergy of Smoothness and Precision in Pose Sequences Refinement

CVPR 2024poster

Predicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily co…

2023

Learning Anchor Transformations for 3D Garment Animation

CVPR 2023poster

This paper proposes an anchor-based deformation model, namely AnchorDEF, to predict 3D garment animation from a body motion sequence. It deforms a garment mesh template by a mixture of rigid transformations with extra nonlinear displacements. A set of anchors around the mesh surface is introduced to…

Cited by 13SourcePDFScholar
2023

Skinned Motion Retargeting With Residual Perception of Motion Semantics & Geometry

CVPR 2023poster

A good motion retargeting cannot be reached without reasonable consideration of source-target differences on both the skeleton and shape geometry levels. In this work, we propose a novel Residual RETargeting network (R2ET) structure, which relies on two neural modification modules, to adjust the sou…

2022

Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation

CVPR 2022poster

Learning semantic segmentation from weakly-labeled (e.g., image tags only) data is challenging since it is hard to infer dense object regions from sparse semantic tags. Despite being broadly studied, most current efforts directly learn from limited semantic annotations carried by individual image or…

Cited by 194PDFcodeScholar
2021

Learning Anchored Unsigned Distance Functions With Gradient Direction Alignment for Single-View Garment Reconstruction

ICCV 2021poster

While single-view 3D reconstruction has made significant progress benefiting from deep shape representations in recent years, garment reconstruction is still not solved well due to open surfaces, diverse topologies and complex geometric details. In this paper, we propose a novel learnable Anchored U…

Cited by 55PDFcodeScholar
2021

Lvio-Fusion: A Self-adaptive Multi-sensor Fusion SLAM Framework Using Actor-critic Method

IROS 2021poster

State estimation with sensors is essential for mobile robots. Due to different performance of sensors in different environments, how to fuse measurements of various sensors is a problem. In this paper, we propose a tightly coupled multi-sensor fusion framework, Lvio-Fusion, which fuses stereo camera…

Cited by 48SourcecodeScholar
2020

Human Parsing Based Texture Transfer from Single Image to 3D Human via Cross-View Consistency

NeurIPS 2020poster

This paper proposes a human parsing based texture transfer model via cross-view consistency learning to generate the texture of 3D human body from a single image. We use the semantic parsing of human body as input for providing both the shape and pose information to reduce the appearance variation…

2020

Region Graph Embedding Network for Zero-Shot Learning

ECCV 2020poster

Most of the existing Zero-Shot Learning (ZSL) approaches learn direct embeddings from global features or image parts (regions) to the semantic space, which, however, fail to capture the appearance relationships between different local regions within a single image. In this paper, to model the relati…

Cited by 195SourcePDFScholar
2020

Unsupervised Domain Adaptation with Noise Resistible Mutual-Training for Person Re-identification

ECCV 2020poster

Unsupervised domain adaptation (UDA) in the task of person re-identification (re-ID) is highly challenging due to large domain divergence and no class overlap between domains. Pseudo-label based self-training is one of the representative techniques to address UDA. However, label noise caused by unsu…

Cited by 227SourcePDFScholar
2018

Towards Pose Invariant Face Recognition in the Wild

CVPR 2018poster

Pose variation is one key challenge in face recognition. As opposed to current techniques for pose invariant face recognition, which either directly extract pose invariant features for recognition, or first normalize profile face images to frontal pose before feature extraction, we argue that it is…

Cited by 300SourcePDFScholar
2018

Weakly Supervised Phrase Localization With Multi-Scale Anchored Transformer Network

CVPR 2018poster

In this paper, we propose a novel weakly supervised model, Multi-scale Anchored Transformer Network (MATN), to accurately localize free-form textual phrases with only image-level supervision. The proposed MATN takes region proposals as localization anchors, and learns a multi-scale correspondence ne…

Cited by 75SourcePDFScholar
2017

Dual-Agent GANs for Photorealistic and Identity Preserving Profile Face Synthesis

NeurIPS 2017poster

Synthesizing realistic profile faces is promising for more efficiently training deep pose-invariant models for large-scale unconstrained face recognition, by populating samples with extreme poses and avoiding tedious annotations. However, learning from synthetic faces may not achieve the desired pe…

2015

Deep Semantic Ranking Based Hashing for Multi-Label Image Retrieval

CVPR 2015poster

With the rapid growth of web images, hashing has received increasing interests in large scale image retrieval. Research efforts have been devoted to learning compact binary codes that preserve semantic similarity based on labels. However, most of these hashing methods are designed to handle simple b…

Cited by 741SourcePDFScholar