← Search

Ping Hu

37 accepted papers

2026

Counterfactual Occlusion-Aware Learning via Visibility Intervention for LiDAR Anomaly Detection

ICML 2026poster

LiDAR point cloud anomaly detection is critical for autonomous system safety, yet most existing methods rely only on visible measurements, overlooking occlusion as a structured consequence of the LiDAR sensing process. We argue that anomalies are characterized not only by what is observed, but also …

Cited by 0SourceScholar
2026

Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language Models

AAAI 2026technical

Vision-language models (VLMs) pre-trained on natural image and language data, such as CLIP, have exhibited significant potential in few-shot image recognition tasks, leading to development of various efficient transfer learning methods. These methods exploit inherent pre-learned knowledge in VLMs an

Cited by 0SourcePDFScholar
2026

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

ICML 2026poster

E-commerce short videos represent a high-revenue segment of the online video industry characterized by a goal-driven format and dense multi-modal signals. Current models often struggle with these videos because existing benchmarks focus primarily on general-purpose tasks and neglect the reasoning of…

Cited by 1SourceScholar
2026

Graph Smoothing for Enhanced Local Geometry Learning in Point Cloud Analysis

AAAI 2026technical

Graph-based methods have proven to be effective in capturing relationships among points for 3D point cloud analysis. However, these methods often suffer from suboptimal graph structures, particularly due to sparse connections at boundary points and noisy connections in junction areas. To address the

Cited by 0SourcePDFScholar
2026

Revisiting Confidence Calibration for Misclassification Detection in VLMs

ICLR 2026poster

Confidence calibration has been widely studied to improve the trustworthiness of predictions in vision-language models (VLMs). However, we theoretically reveal that standard confidence calibration inherently _impairs_ the ability to distinguish between correct and incorrect predictions (i.e., Miscla…

Cited by 0SourceScholar
2026

Structure-to-Intensity Diffusion for Adverse-Weather LiDAR Generation

CVPR 2026

Adverse-weather LiDAR point cloud generation is challenged by complex weather-induced degradations. These degradations affect geometry and reflectance in fundamentally different ways, making joint modeling difficult and ambiguous, especially when diverse real-world training data is limited. To addre

Cited by 0SourceScholar
2025

Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding

AAAI 2025technical

Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance o…

2025

Boundary Probing for Input Privacy Protection When Using LMM Services

ICCV 2025poster

Alongside the rapid development of Large Multimodal Models (LMMs) like GPT-4V, privacy concerns also rise. As LMMs are commonly deployed as cloud services, users are typically required to upload their personal images and videos to the cloud to access these services, raising great concerns about visu…

Cited by 0SourcePDFScholar
2025

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

NeurIPS 2025poster

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particul…

Cited by 0SourceScholar
2025

JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systems

CVPR 2025poster

Unmanned Aerial Vehicles (UAVs) are widely adopted across various fields, yet they raise significant privacy and safety concerns, demanding robust monitoring solutions. Existing anti-UAV methods primarily focus on position tracking but fail to capture UAV behavior and intent. To address this, we int…

Cited by 0SourcePDFScholar
2025

Seeking Proxy Point via Stable Feature Space for Noisy Correspondence Learning

IJCAI 2025

To meet the growing demand for cross-modal training data, directly collecting multimodal data from the Internet has become prevalent. However, such data inevitably suffer from Noisy Correspondence. Previous works focused on recasting soft labels to mitigate noise's negative impact. We explore a nove

2025

Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse Weather

CVPR 2025poster

Existing LiDAR semantic segmentation models often suffer from decreased accuracy when exposed to adverse weather conditions. Recent methods addressing this issue focus on enhancing training data through weather simulation or universal augmentation techniques. However, few works have studied the nega…

Cited by 0SourcePDFScholar
2024

Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters

CVPR 2024poster

Continual learning can empower vision-language models to continuously acquire new knowledge without the need for access to the entire historical dataset. However mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout lifelong learning and (…

2024

Exploring the Role of Node Diversity in Directed Graph Representation Learning

IJCAI 2024poster

Many methods of Directed Graph Neural Networks (DGNNs) are designed to equally treat nodes in the same neighbor set (i.e., out-neighbor set and in-neighbor set) for every node, without considering the node diversity in directed graphs, so they are often unavailable to adaptively acquire suitable inf…

Cited by 3SourcePDFScholar
2024

Harnessing Text-to-Image Diffusion Models for Category-Agnostic Pose Estimation

ECCV 2024oral

"Category-Agnostic Pose Estimation (CAPE) aims to detect keypoints of an arbitrary unseen category in images, based on several provided examples of that category. This is a challenging task, as the limited data of unseen categories makes it difficult for models to generalize effectively. To address…

Cited by 10SourcePDFScholar
2024

Koala: Key Frame-Conditioned Long Video-LLM

CVPR 2024highlight

Long video question answering is a challenging task that involves recognizing short-term activities and reasoning about their fine-grained relationships. State-of-the-art video Large Language Models (vLLMs) hold promise as a viable solution due to their demonstrated emergent capabilities on new task…

Cited by 33SourcePDFScholar
2024

Self-Supervised Heterogeneous Graph Learning: a Homophily and Heterogeneity View

ICLR 2024poster

Self-supervised heterogeneous graph learning has achieved promising results in various real applications, but it still suffers from the following issues: (i) meta-paths can be employed to capture the homophily in the heterogeneous graph, but meta-paths are human-defined, requiring substantial exper…

Cited by 9SourcePDFScholar
2024

Towards Dynamic-Prompting Collaboration for Source-Free Domain Adaptation

IJCAI 2024poster

In domain adaptation, challenges such as data privacy constraints can impede access to source data, catalyzing the development of source-free domain adaptation (SFDA) methods. However, current approaches heavily rely on models trained on source data, posing the risk of overfitting and suboptimal gen…

Cited by 0SourcePDFScholar
2023

Diffusion-based Image Translation with Label Guidance for Domain Adaptive Semantic Segmentation

ICCV 2023poster

Translating images from a source domain to a target domain for learning target models is one of the most common strategies in domain adaptive semantic segmentation (DASS). However, existing methods still struggle to preserve semantically-consistent local details between the original and translated i…

Cited by 32PDFScholar
2023

GAIT: Generating Aesthetic Indoor Tours with Deep Reinforcement Learning

ICCV 2023poster

Placing and orienting a camera to compose aesthetically meaningful shots of a scene is not only a key objective in real-world photography and cinematography but also for virtual content creation. The framing of a camera often significantly contributes to the story telling in movies, games, and mixed…

Cited by 3PDFcodeScholar
2023

Joint Attribute and Model Generalization Learning for Privacy-Preserving Action Recognition

NeurIPS 2023poster

Privacy-Preserving Action Recognition (PPAR) aims to transform raw videos into anonymous ones to prevent privacy leakage while maintaining action clues, which is an increasingly important problem in intelligent vision applications. Despite recent efforts in this task, it is still challenging to deal…

Cited by 4SourcePDFScholar
2023

Token Boosting for Robust Self-Supervised Visual Transformer Pre-Training

CVPR 2023poster

Learning with large-scale unlabeled data has become a powerful tool for pre-training Visual Transformers (VTs). However, prior works tend to overlook that, in real-world scenarios, the input data may be corrupted and unreliable. Pre-training VTs on such corrupted data can be challenging, especially…

Cited by 6SourcePDFScholar
2022

DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations

NeurIPS 2022accept

Solving multi-label recognition (MLR) for images in the low-label regime is a challenging task with many real-world applications. Recent work learns an alignment between textual and visual spaces to compensate for insufficient image labels, but loses accuracy because of the limited amount of availab…

Cited by 141SourcePDFScholar
2022

ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes

CVPR 2022poster

Less than 35% of recyclable waste is being actually recycled in the US, which leads to increased soil and sea pollution and is one of the major concerns of environmental researchers as well as the common public. At the heart of the problem are the inefficiencies of the waste sorting process (separat…

Cited by 67PDFcodeScholar
2021

Real-Time Semantic Segmentation With Fast Attention

RA-L 2021

In deep CNN based models for semantic segmentation, high accuracy relies on rich spatial context (large receptive fields) and fine spatial details (high resolution), both of which incur high computational costs. In this letter, we propose a novel architecture that addresses both challenges and achie

Cited by 143SourcecodeScholar
2020

Temporally Distributed Networks for Fast Video Semantic Segmentation

CVPR 2020poster

We present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of a deep CNN can be approximated by composing features extracted from several shallower sub-networks. Leveraging the inhe…

Cited by 250PDFScholar
2018

Motion-Guided Cascaded Refinement Network for Video Object Segmentation

CVPR 2018poster

Deep CNNs have achieved superior performance in many tasks of computer vision and image understanding. However, it is still difficult to effectively apply deep CNNs to video object segmentation(VOS) since treating video frames as separate and static will lose the information hidden in motion. To tac…

Cited by 131SourcePDFScholar
2017

Global Context-Aware Attention LSTM Networks for 3D Action Recognition

CVPR 2017poster

Long Short-Term Memory (LSTM) networks have shown superior performance in 3D human action recognition due to their power in modeling the dynamics and dependencies in sequential data. Since not all joints are informative for action analysis and the irrelevant joints often bring a lot of noise, we nee…

Cited by 822PDFScholar