← Search

Baosheng Yu

27 accepted papers

2026

AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs

ICLR 2026poster

The application of large language models (LLMs) in the medical field has garnered significant attention, yet their reasoning capabilities in more specialized domains like anesthesiology remain underexplored. To bridge this gap, we introduce AnesSuite, the first comprehensive dataset suite specifical…

Cited by 0SourcecodeScholar
2026

ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling

ICLR 2026poster

Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time scaling (TTS). However, while most PRMs exhibit substantial gains in mathematical domains, the scarcity of domain-specific…

Cited by 0SourceScholar
2026

Cross-Sample Augmented Test-Time Adaptation for Personalized Intraoperative Hypotension Prediction

AAAI 2026technical

Intraoperative hypotension (IOH) poses significant surgical risks, but accurate prediction remains challenging due to patient-specific variability. While test-time adaptation (TTA) offers a promising approach for personalized prediction, the rarity of IOH events often leads to unreliable test-time t

Cited by 0SourcePDFScholar
2026

Remodeling Semantic Relationships in Vision-Language Fine-Tuning

AAAI 2026technical

Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus lead

Cited by 0SourcePDFScholar
2025

Bi-Level Optimization for Self-Supervised AI-Generated Face Detection

ICCV 2025poster

AI-generated face detectors trained via supervised learning typically rely on synthesized images from specific generators, limiting their generalization to emerging generative techniques. To overcome this limitation, we introduce a self-supervised method based on bi-level optimization. In the inner…

2025

SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation

NeurIPS 2025poster

Enhancing large language models by simply scaling up datasets has begun to yield diminishing returns, shifting the spotlight to data quality. Monte Carlo Tree Search (MCTS) has emerged as a powerful technique for generating high-quality chain-of-thought data, yet conventional approaches typically re…

Cited by 0SourceScholar
2024

MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models

IJCAI 2024poster

Foundation models have demonstrated significant emergent abilities, holding great promise for enhancing embodied agents' reasoning and planning capacities. However, the absence of a comprehensive benchmark for evaluating embodied agents with multimodal observations in complex environments remains a…

2023

Knowledge-Aware Federated Active Learning with Non-IID Data

ICCV 2023poster

Federated learning enables multiple decentralized clients to learn collaboratively without sharing local data. However, the expensive annotation cost on local clients remains an obstacle in utilizing local data. In this paper, we propose a federated active learning paradigm to efficiently learn a gl…

Cited by 25PDFcodeScholar
2022

BatchFormer: Learning To Explore Sample Relationships for Robust Representation Learning

CVPR 2022poster

Despite the success of deep neural networks, there are still many challenges in deep representation learning due to the data scarcity issues such as data imbalance, unseen distribution, and domain shift. To address the above-mentioned issues, a variety of methods have been devised to explore the sam…

Cited by 101PDFcodeScholar
2022

Contrastive Boundary Learning for Point Cloud Segmentation

CVPR 2022poster

Point cloud segmentation is fundamental in understanding 3D environments. However, current 3D point cloud segmentation methods usually perform poorly on scene boundaries, which degenerates the overall segmentation performance. In this paper, we focus on the segmentation of scene boundaries. Accordin…

Cited by 184PDFcodeScholar
2022

Discovering Human-Object Interaction Concepts via Self-Compositional Learning

ECCV 2022poster

"A comprehensive understanding of human-object interaction (HOI) requires detecting not only a small portion of predefined HOI concepts (or categories) but also other reasonable HOI concepts, while current approaches usually fail to explore a huge portion of unknown HOI concepts (i.e., unknown but r…

2022

Improving Fine-Grained Visual Recognition in Low Data Regimes via Self-Boosting Attention Mechanism

ECCV 2022poster

"The challenge of fine-grained visual recognition often lies in discovering the key discriminative regions. While such regions can be automatically identified from a large-scale labeled dataset, a similar method might become less effective when only a few annotations are available. In low data regim…

2022

Learning Affinity From Attention: End-to-End Weakly-Supervised Semantic Segmentation With Transformers

CVPR 2022poster

Weakly-supervised semantic segmentation (WSSS) with image-level labels is an important and challenging task. Due to the high training efficiency, end-to-end solutions for WSSS have received increasing attention from the community. However, current methods are mainly based on convolutional neural net…

Cited by 262PDFcodeScholar
2022

MeshMAE: Masked Autoencoders for 3D Mesh Data Analysis

ECCV 2022poster

"Recently, self-supervised pre-training has advanced Vision Transformers on various tasks w.r.t. different data modalities, e.g., image and 3D point cloud data. In this paper, we explore this learning paradigm for 3D mesh data analysis based on Transformers. Since applying Transformer architectures…

Cited by 59SourcePDFScholar
2022

Resistance Training Using Prior Bias: Toward Unbiased Scene Graph Generation

AAAI 2022technical

Scene Graph Generation (SGG) aims to build a structured representation of a scene using objects and pairwise relationships, which benefits downstream tasks. However, current SGG methods usually suffer from sub-optimal scene graph generation because of the long-tailed distribution of training data. T…

2021

Affordance Transfer Learning for Human-Object Interaction Detection

CVPR 2021poster

Reasoning the human-object interactions (HOI) is essential for deeper scene understanding, while object affordances (or functionalities) are of great importance for human to discover unseen HOIs with novel objects. Inspired by this, we introduce an affordance transfer learning approach to jointly de…

Cited by 139PDFcodeScholar
2021

Contrastive Graph Poisson Networks: Semi-Supervised Learning with Extremely Limited Labels

NeurIPS 2021poster

Graph Neural Networks (GNNs) have achieved remarkable performance in the task of semi-supervised node classification. However, most existing GNN models require sufficient labeled data for effective network training. Their performance can be seriously degraded when labels are extremely limited. To ad…

Cited by 65SourcePDFScholar
2021

Detecting Human-Object Interaction via Fabricated Compositional Learning

CVPR 2021poster

Human-Object Interaction (HOI) detection, inferring the relationships between human and objects from images/videos, is a fundamental task for high-level scene understanding. However, HOI detection usually suffers from the open long-tailed nature of interactions with objects, while human has extremel…

Cited by 118PDFcodeScholar
2021

Not All Operations Contribute Equally: Hierarchical Operation-Adaptive Predictor for Neural Architecture Search

ICCV 2021poster

Graph-based predictors have recently shown promising results on neural architecture search (NAS). Despite their efficiency, current graph-based predictors treat all operations equally, resulting in biased topological knowledge of cell architectures. Intuitively, not all operations are equally signif…

Cited by 13PDFScholar
2021

SynFace: Face Recognition With Synthetic Data

ICCV 2021poster

With the recent success of deep neural networks, remarkable progress has been achieved on face recognition. However, collecting large-scale real-world training data for face recognition has turned out to be challenging, especially due to the label noise and privacy issues. Meanwhile, existing face r…

Cited by 154PDFcodeScholar
2021

TGRNet: A Table Graph Reconstruction Network for Table Structure Recognition

ICCV 2021poster

A table arranging data in rows and columns is a very effective data structure, which has been widely used in business and scientific research. Considering large-scale tabular data in online and offline documents, automatic table recognition has attracted increasing attention from the document analys…

Cited by 69PDFcodeScholar
2018

Correcting the Triplet Selection Bias for Triplet Loss

ECCV 2018poster

Triplet loss, popular for metric learning, has made a great success in many computer vision tasks, such as fine-grained image classification, image retrieval, and face recognition. Considering that the number of triplets grows cubically with the size of training data, triplet mining is thus indispen…

Cited by 132SourcePDFScholar