← Search

Junjie Zhou

21 accepted papers

2026

Advancing Cancer Prognosis with Hierarchical Fusion of Genomic, Proteomic and Pathology Imaging Data from a Systems Biology Perspective

CVPR 2026

To enhance the precision of cancer prognosis, recent research has increasingly focused on multimodal survival methods by integrating genomic data and histology images. However, current approaches overlook the fact that the proteome serves as an intermediate layer bridging genomic alterations and his

Cited by 0SourceScholar
2026

OmniGen2: Towards Instruction-Aligned Multimodal Generation

CVPR 2026

Multimodal generative models can process instructions in various modalities and demonstrate outstanding performance across a wide range of image generation tasks. However, their robustness in complex real-world scenarios remains limited due to insufficient generalized instruction alignment. We intro

Cited by 0SourcecodeScholar
2026

ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement

CVPR 2026

While existing generation and unified models excel at general image generation, they struggle with tasks requiring deep reasoning, planning, and precise data-to-visual mapping abilities beyond general scenarios. To push beyond the existing limitations, we introduce a new and challenging task: creati

Cited by 0SourcecodeScholar
2025

AcZeroTS: Active Learning for Zero-shot Tissue Segmentation in Pathology Images

ICCV 2025poster

Tissue segmentation in pathology images is crucial for computer-aided diagnostics of human cancers. Traditional tissue segmentation models rely heavily on large-scale labeled datasets, where every tissue type must be annotated by experts. However, due to the complexity of tumor micro-environment, co…

Cited by 0SourcePDFScholar
2025

Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval

ACL 2025long

With the popularity of multimodal techniques, it receives growing interests to acquire useful information in visual forms. In this work, we formally define an emerging IR paradigm called Visualized Information Retrieval, or Vis-IR, where multimodal information, such as texts, images, tables and char…

2025

DAMM-Diffusion: Learning Divergence-Aware Multi-Modal Diffusion Model for Nanoparticles Distribution Prediction

CVPR 2025highlight

The prediction of nanoparticles (NPs) distribution is crucial for the diagnosis and treatment of tumors. Recent studies indicate that the heterogeneity of tumor microenvironment (TME) highly affects the distribution of NPs across tumors. Hence, it has become a research hotspot to generate the NPs di…

2025

Design and Performance Study of an Underwater Soft Snake-like Robot

IROS 2025

In this paper, we propose a design of an underwater soft snake-like robot prototype that uses two actuators made of 3D-printed soft materials to build the robot body. Control signals with appropriate displacement phases and different voltages are used to control the water pump to drive the soft actu

Cited by 0SourceScholar
2025

MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image Classification

NeurIPS 2025poster

Prompt learning has emerged as a promising paradigm for adapting pre-trained vision-language models (VLMs) to few-shot whole slide image (WSI) classification by aligning visual features with textual representations, thereby reducing annotation cost and enhancing model generalization. Nevertheless, e…

Cited by 0SourceScholar
2025

MLVU: Benchmarking Multi-task Long Video Understanding

CVPR 2025poster

The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in vi…

2025

MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval

ACL 2025long

Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massi…

2025

MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval

NeurIPS 2025poster

Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on the end-to-end LVU performance, making them inappropriate for…

Cited by 0SourceScholar
2025

OmniGen: Unified Image Generation

CVPR 2025poster

The emergence of Large Language Models (LLMs) has unified language generation tasks and revolutionized human-machine interaction. However, in the realm of image generation, a unified model capable of handling various tasks within a single framework remains largely unexplored. In this work, we introd…

2025

Robust Multimodal Survival Prediction with Conditional Latent Differentiation Variational AutoEncoder

CVPR 2025poster

The integrative analysis of histopathological images and genomic data has received increasing attention for survival prediction of human cancers. However, the existing studies always hold the assumption that full modalities are available. As a matter of fact, the cost for collecting genomic data is…

2025

Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

CVPR 2025poster

Long video understanding poses a significant challenge for current Multi-modal Large Language Models (MLLMs). Notably, the MLLMs are constrained by their limited context lengths and the substantial costs while processing long videos. Although several existing methods attempt to reduce visual tokens,…

2024

State Estimation by Joint Approach With Dynamic Modeling and Observer for Soft Actuator

RA-L 2024

In order to achieve a significant reduction in state estimation error and improved convergence speed, ensuring real-time responsiveness and computational efficiency, this article proposes a joint approach that combines dynamic modeling and observers to achieve accurate nonlinear state estimation of

Cited by 1SourceScholar
2024

Tumor Micro-environment Interactions Guided Graph Learning for Survival Analysis of Human Cancers from Whole-slide Pathological Images

CVPR 2024poster

The recent advance of deep learning technology brings the possibility of assisting the pathologist to predict the patients' survival from whole-slide pathological images (WSIs). However most of the prevalent methods only worked on the sampled patches in specifically or randomly selected tumor areas…

2024

VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

ACL 2024long

Multi-modal retrieval becomes increasingly popular in practice. However, the existing retrievers are mostly text-oriented, which lack the capability to process visual information. Despite the presence of vision-language models like CLIP, the current methods are severely limited in representing the t…

2023

Design, Modeling, and Control of a Legless Squamate Reptiles Inspired Soft Crawling Robot

RA-L 2023

For soft crawling robots moving in two-dimensional space, accurate dynamics models are the basis for control and navigation. However, it is still challenging for nonlinear and low-bandwidth soft robot systems. This article details two dynamics models that can be applied to predict the motion of a so

Cited by 3SourceScholar