← Search

Weidong Cai

35 accepted papers

2026

Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking

ICML 2026poster

Vision–language models (VLMs) now support both direct Instruct and explicit-reasoning Thinking modes, but practitioners lack principled ways to decide when reasoning helps or how much computation to allocate at test time. We investigate whether VLMs encode meta-cognitive signals for adaptive inferen…

Cited by 0SourceScholar
2026

HiFusion: Hierarchical Intra-Spot Alignment and Regional Context Fusion for Spatial Gene Expression Prediction from Histopathology

AAAI 2026technical

Spatial transcriptomics (ST) bridges gene expression and tissue morphology but faces clinical adoption barriers due to technical complexity and prohibitive costs. While computational methods predict gene expression from H&E-stained whole-slide images (WSIs), existing approaches often fail to capture

Cited by 0SourcePDFScholar
2026

NeuroSeg Meets DINOv3: Transferring 2D Self-Supervised Visual Priors to 3D Neuron Segmentation via DINOv3 Initialization

CVPR 2026

2D visual foundation models, such as DINOv3, a self-supervised model trained on large-scale natural images, have demonstrated strong zero-shot generalization, capturing rich global context and fine-grained structural cues. However, an analogous 3D foundation model for downstream volumetric neuroimag

Cited by 0SourcecodeScholar
2026

RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction

ICML 2026poster

Histopathology whole-slide images (WSIs) are routinely acquired in clinical practice and contain rich tissue morphology but lack direct molecular architecture and functional programs defining pathological states, whereas RNA sequencing (RNA-seq) provides genome-wide transcriptional profiles at subst…

Cited by 0SourceScholar
2026

Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance

ICML 2026poster

Autoregressive (AR) models based on next-scale prediction are rapidly emerging as a powerful tool for image generation, but they face a critical weakness: information inconsistencies between patches across timesteps introduced by progressive resolution scaling. These inconsistencies scatter guidance…

Cited by 0SourceScholar
2026

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

AAAI 2026technical

Universal multimodal embedding models are essential in various tasks. Existing approaches typically use in-batch mining to identify hard negatives by measuring the similarity of query-candidate pairs. However, these methods often struggle to capture subtle semantic differences among candidates and l

Cited by 0SourcePDFScholar
2025

CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination

AAAI 2025technical

Contrastive Language-Image Pre-training (CLIP) has achieved excellent performance over a wide range of tasks. However, the effectiveness of CLIP heavily relies on a substantial corpus of pre-training data, resulting in notable consumption of computational resources. Although knowledge distillation h…

Cited by 5SourcePDFScholar
2025

Multimodal Causal Reasoning Benchmark: Challenging Multimodal Large Language Models to Discern Causal Links Across Modalities

ACL 2025finding

Multimodal Large Language Models (MLLMs) have showcased exceptional Chain-of-Thought (CoT) reasoning ability in complex textual inference tasks including causal reasoning. However, will these causalities remain straightforward when crucial hints hide in visual details? If not, what factors might inf…

Cited by 0SourcePDFScholar
2024

Advancements in 3D Lane Detection Using LiDAR Point Clouds: From Data Collection to Model Development

ICRA 2024poster

Advanced Driver-Assistance Systems (ADAS) have successfully integrated learning-based techniques into vehicle perception and decision-making. However, their application in 3D lane detection for effective driving environment perception is hindered by the lack of comprehensive LiDAR datasets. The spar…

Cited by 4SourcecodeScholar
2024

Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights

ECCV 2024poster

"(CIC) evolves traditional image captioning into a more complex domain, necessitating the ability for multimodal reasoning. It aims to generate image captions given specific contextual information. This paper further introduces a novel domain of (). Unlike CIC, which solely relies on broad context,…

2024

Enhancing Advanced Visual Reasoning Ability of Large Language Models

EMNLP 2024main

Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models’ advanced reasoning ability. Traditional Vision-Language models (VLMs) perform well in visual perception tasks while struggling with complex reasoning scenarios. Converse…

Cited by 7SourcePDFScholar
2024

PaintHuman: Towards High-Fidelity Text-to-3D Human Texturing via Denoised Score Distillation

AAAI 2024technical

Recent advances in zero-shot text-to-3D human generation, which employ the human model prior (e.g., SMPL) or Score Distillation Sampling (SDS) with pre-trained text-to-image diffusion models, have been groundbreaking. However, SDS may provide inaccurate gradient directions under the weak diffusion g…

2024

RWKV-CLIP: A Robust Vision-Language Representation Learner

EMNLP 2024main

Contrastive Language-Image Pre-training (CLIP) has significantly improved performance in various vision-language tasks by expanding the dataset with image-text pairs obtained from the web. This paper further explores CLIP from the perspectives of data and model architecture. To mitigate the impact o…

2024

Seeing Unseen: Discover Novel Biomedical Concepts via Geometry-Constrained Probabilistic Modeling

CVPR 2024poster

Machine learning holds tremendous promise for transforming the fundamental practice of scientific discovery by virtue of its data-driven nature. With the ever-increasing stream of research data collection it would be appealing to autonomously explore patterns and insights from observational data for…

Cited by 6SourcePDFScholar
2024

V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation Models

AAAI 2024technical

Building artificial intelligence (AI) systems on top of a set of foundation models (FMs) is becoming a new paradigm in AI research. Their representative and generative abilities learnt from vast amounts of data can be easily adapted and transferred to a wide range of downstream tasks without extra t…

2023

CelebV-Text: A Large-Scale Facial Text-Video Dataset

CVPR 2023poster

Text-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and highly relevant texts. This paper presents CelebV-Text, a large-scale, diverse, a…

2023

PaRot: Patch-Wise Rotation-Invariant Network via Feature Disentanglement and Pose Restoration

AAAI 2023technical

Recent interest in point cloud analysis has led rapid progress in designing deep learning methods for 3D models. However, state-of-the-art models are not robust to rotations, which remains an unknown prior to real applications and harms the model performance. In this work, we introduce a novel Patch…

2023

SQUID: Deep Feature In-Painting for Unsupervised Anomaly Detection

CVPR 2023poster

Radiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. To exploit this structured information, we propose the use of Space-aware Memory Queues for In-painting and Detecting anomalies…

2023

Taxonomy Adaptive Cross-Domain Adaptation in Medical Imaging via Optimization Trajectory Distillation

ICCV 2023poster

The success of automated medical image analysis depends on large-scale and expert-annotated training sets. Unsupervised domain adaptation (UDA) has been raised as a promising approach to alleviate the burden of labeled data collection. However, they generally operate under the closed-set adaptation…

Cited by 15PDFcodeScholar
2022

Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

IJCAI 2022poster

Dense captioning in 3D point clouds is an emerging vision-and-language task involving object-level 3D scene understanding. Apart from coarse semantic class prediction and bounding box regression as in traditional 3D object detection, 3D dense captioning aims at producing a further and finer instance…

2021

Exploiting Edge-Oriented Reasoning for 3D Point-Based Scene Graph Analysis

CVPR 2021poster

Scene understanding is a critical problem in computer vision. In this paper, we propose a 3D point-based scene graph generation (SGGpoint) framework to effectively bridge perception and reasoning to achieve scene understanding via three sequential stages, namely scene graph construction, reasoning,…

Cited by 65PDFScholar
2021

Walk in the Cloud: Learning Curves for Point Clouds Shape Analysis

ICCV 2021poster

Discrete point cloud objects lack sufficient shape descriptors of 3D geometries. In this paper, we present a novel method for aggregating hypothetical curves in point clouds. Sequences of connected points (curves) are initially grouped by taking guided walks in the point clouds, and then subsequentl…

Cited by 367PDFcodeScholar
2020

Unsupervised Instance Segmentation in Microscopy Images via Panoptic Domain Adaptation and Task Re-Weighting

CVPR 2020poster

Unsupervised domain adaptation (UDA) for nuclei instance segmentation is important for digital pathology, as it alleviates the burden of labor-intensive annotation and domain shift across datasets. In this work, we propose a Cycle Consistency Panoptic Domain Adaptive Mask R-CNN (CyC-PDAM) architectu…

Cited by 98PDFcodeScholar
2017

Deep Clustering via Joint Convolutional Autoencoder Embedding and Relative Entropy Minimization

ICCV 2017poster

In this paper, we propose a new clustering model, called DEeP Embedded RegularIzed ClusTering (DEPICT), which efficiently maps data into a discriminative embedding subspace and precisely predicts cluster assignments. DEPICT generally consists of a multinomial logistic regression function stacked on…

Cited by 585PDFcodeScholar
2017

Locally-Transferred Fisher Vectors for Texture Classification

ICCV 2017poster

Texture classification has been extensively studied in computer vision. Recent research shows that the combination of Fisher vector (FV) encoding and convolutional neural network (CNN) provides significant improvement in texture classification over the previous feature representation methods. Howeve…

Cited by 75PDFScholar
2017

Regularized Modal Regression with Applications in Cognitive Impairment Prediction

NeurIPS 2017poster

Linear regression models have been successfully used to function estimation and model selection in high-dimensional data analysis. However, most existing methods are built on least squares with the mean square error (MSE) criterion, which are sensitive to outliers and their performance may be degrad…

Cited by 41SourcePDFScholar
2015

Fusing Subcategory Probabilities for Texture Classification

CVPR 2015poster

Texture, as a fundamental characteristic of objects, has attracted much attention in computer vision research. Performance of texture classification is however still lacking for some challenging cases, largely due to the high intra-class variation and low inter-class distinction. To tackle these iss…

Cited by 22SourcePDFScholar
2015

Robust Saliency Detection via Regularized Random Walks Ranking

CVPR 2015poster

In the field of saliency detection, many graph-based algorithms heavily depend on the accuracy of the pre-processed superpixel segmentation, which leads to significant sacrifice of detail information from the input image. In this paper, we propose a novel bottom-up saliency detection approach that t…

Cited by 279SourcePDFScholar