← Search

Hong-Yu Zhou

11 accepted papers

2025

MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

ICLR 2025poster

This paper introduces MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities with multigranular annotations for more than 65 diseases. These multigranular annotations encompass both global information, such as modality and o…

2025

No More Sibling Rivalry: Debiasing Human-Object Interaction Detection

ICCV 2025poster

Detection transformers have been applied to human-object interaction (HOI) detection, enhancing the localization and recognition of human-action-object triplets in images. Despite remarkable progress, this study identifies a critical issue--"Toxic Siblings" bias--which hinders the interaction decode…

Cited by 0SourcePDFScholar
2023

Activate and Reject: Towards Safe Domain Generalization under Category Shift

ICCV 2023poster

Albeit the notable performance on in-domain test points, it is non-trivial for deep neural networks to attain satisfactory accuracy when deploying in the open world, where novel domains and object classes often occur. In this paper, we study a practical problem of Domain Generalization under Categor…

Cited by 8PDFScholar
2023

Advancing Radiograph Representation Learning with Masked Record Modeling

ICLR 2023poster

Modern studies in radiograph representation learning (R$^2$L) rely on either self-supervision to encode invariant semantics or associated radiology reports to incorporate medical expertise, while the complementarity between them is barely noticed. To explore this, we formulate the self- and report-c…

2023

DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models

NeurIPS 2023poster

A long-standing goal of AI systems is to perform complex multimodal reasoning like humans. Recently, large language models (LLMs) have made remarkable strides in such multi-step reasoning on the language modality solely by leveraging the chain of thought (CoT) to mimic human thinking. However, the t…

Cited by 100SourcePDFScholar
2023

Protein Representation Learning via Knowledge Enhanced Primary Structure Reasoning

ICLR 2023poster

Protein representation learning has primarily benefited from the remarkable development of language models (LMs). Accordingly, pre-trained protein models also suffer from a problem in LMs: a lack of factual knowledge. The recent solution models the relationships between protein and associated knowle…

Cited by 24SourcePDFScholar
2022

Attribute Surrogates Learning and Spectral Tokens Pooling in Transformers for Few-Shot Learning

CVPR 2022poster

This paper presents new hierarchically cascaded transformers that can improve data efficiency through attribute surrogates learning and spectral tokens pooling. Vision transformers have recently been thought of as a promising alternative to convolutional neural networks for visual recognition. But w…

Cited by 71PDFcodeScholar
2021

Bottom-Up Shift and Reasoning for Referring Image Segmentation

CVPR 2021poster

Referring image segmentation aims to segment the referent that is the corresponding object or stuff referred by a natural language expression in an image. Its main challenge lies in how to effectively and efficiently differentiate between the referent and other objects of the same category as the re…

Cited by 101PDFcodeScholar
2021

Preservational Learning Improves Self-Supervised Medical Image Models by Reconstructing Diverse Contexts

ICCV 2021poster

Preserving maximal information is the basic principle of designing self-supervised learning methodologies. To reach this goal, contrastive learning adopts an implicit way which is contrasting image pairs. However, we believe it is not fully optimal to simply use the contrastive estimation for preser…

Cited by 116PDFcodeScholar
2017

Adaptive Feeding: Achieving Fast and Accurate Detections by Adaptively Combining Object Detectors

ICCV 2017poster

Object detection aims at high speed and accuracy simultaneously. However, fast models are usually less accurate, while accurate models cannot satisfy our need for speed. A fast model can be 10 times faster but 50% less accurate than an accurate model. In this paper, we propose Adaptive Feeding (AF)…

Cited by 37PDFScholar