← Search

Sheng Zhang

47 accepted papers

2026

CMG3D: Compensation towards Modality Gap for Open-Vocabulary Indoor 3D Object Detection

ICRA 2026poster

Open-vocabulary indoor three-dimensional object detection (OVI3DOD) is used to detect any class of objects in indoor scenes with prompts. Owing to the relatively limited three-dimensional (3D) data, most of the OVI3DOD algorithms perform training with pseudo labels transformed from the openvocabular…

Cited by 0SourceScholar
2026

Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning

CVPR 2026

Effective medical image analysis requires representations that capture both global anatomical structure and fine-grained tissue texture. Current self-supervised approaches exhibit limited capacity to address both requirements simultaneously. Invariance-based methods learn through augmentation consis

Cited by 0SourceScholar
2026

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

CVPR 2026

High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tasks. We investigate strategies for training and data curation to develop a robust multimodal reasoning model in the medic

Cited by 0SourceScholar
2026

Renormalization Group Guided Tensor Network Structure Search

AAAI 2026technical

Tensor network structure search (TN-SS) aims to automatically discover optimal network topologies and rank configurations for efficient tensor decomposition in high-dimensional data representation. Despite recent advances, existing TN-SS methods face significant limitations in computational tractabi

Cited by 0SourcePDFScholar
2026

SIGMA: An Agent-Based Modeling UAV Swarm Simulator for Swarm Intelligence Algorithms (I)

ICRA 2026poster

Swarm intelligence for uncrewed aerial vehicles (UAVs) significantly improves the success rate of executing intricate tasks using “distributed platforms and aggregated effects”. However, the high experimental costs and safety risks constrain its development. This paper introduces SIGMA (Swarm Intell…

Cited by 0Scholar
2026

Towards Pareto-Optimal Tool-Integrated Agents with Pareto Ranking Policy Optimization

ICML 2026spotlight

Recent advances in tool-integrated language agents have significantly improved their ability to solve complex reasoning tasks. However, existing alignment methods predominantly focus on maximizing task accuracy, while overlooking auxiliary objectives such as tool-use efficiency, which are essential …

Cited by 0SourceScholar
2025

BoRe-Depth: Self-Supervised Monocular Depth Estimation with Boundary Refinement for Embedded Systems

IROS 2025

Depth estimation is one of the key technologies for realizing 3D perception in unmanned systems. Monocular depth estimation has been widely researched because of its low-cost advantage, but the existing methods face the challenges of poor depth estimation performance and blurred object boundaries on

Cited by 1SourcecodeScholar
2025

CMG3D: Compensation Towards Modality Gap for Open-Vocabulary Indoor 3D Object Detection

RA-L 2025

For open-vocabulary indoor three-dimensional (3D) object detection (OVI3DOD), there is a gap between the image and the point cloud for indoor scenes, especially on distant objects. However, existing algorithms ignore this problem, which weakens the detection performance. Therefore, we propose Compen

Cited by 0SourceScholar
2025

CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code Completion

EMNLP 2025

Repository-level code completion automatically predicts the unfinished code based on the broader information from the repository. Recent strides in Code Large Language Models (code LLMs) have spurred the development of repository-level code completion methods, yielding promising results. Nevertheles

2025

DANCE: Resource-Efficient Neural Architecture Search with Data-Aware and Continuous Adaptation

IJCAI 2025

Neural Architecture Search (NAS) has emerged as a powerful approach for automating neural network design. However, existing NAS methods face critical limitations in real-world deployments: architectures lack adaptability across scenarios, each deployment context requires costly separate searches, an

2025

FlexiTex: Enhancing Texture Generation via Visual Guidance

AAAI 2025technical

Recent texture generation methods achieve impressive results due to the powerful generative prior they leverage from large-scale text-to-image diffusion models. However, abstract textual prompts are limited in providing global textural or shape information, which results in the texture generation me…

2025

From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning

NAACL 2025long

Motivated by in-context learning (ICL) capabilities of Large Language Models (LLMs), multimodal LLMs with additional visual modality are also exhibited with similar ICL abilities when multiple image-text pairs are provided as demonstrations. However, relatively less work has been done to investigate…

2025

GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning

ICCV 2025poster

Recent advances have shown that video generation models can enhance robot learning by deriving effective robot actions through inverse dynamics. However, these methods heavily depend on the quality of generated data and struggle with fine-grained manipulation due to the lack of environment feedback.…

2025

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

ICLR 2025poster

We introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10 categories of multi-image relations (e.g., multiview, temporal…

2025

Online Optimization of Offloading Video Analytics Tasks to Multiple Edges for Accuracy Maximization

ICASSP 2025accepted

Real-time video analytics (VA) presents challenges due to its computational intensity and latency sensitivity, especially when processed on mobile devices with limited local resources. We propose to offload VA tasks to edge servers with diverse computational capabilities. We present a "detect + trac…

Cited by 0SourceScholar
2025

RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging

EMNLP 2025

We unveil that internal representations in large language models (LLMs) serve as reliable proxies of learned knowledge, and propose **RECALL**, a novel representation-aware model merging framework for continual learning without access to historical data. RECALL computes inter-model similarity from l

2025

RomanTex: Decoupling 3D-aware Rotary Positional Embedded Multi-Attention Network for Texture Synthesis

ICCV 2025poster

Painting textures for existing geometries is a critical yet labor-intensive process in 3D asset generation. Recent advancements in text-to-image (T2I) models have led to significant progress in texture generation. Most existing research approaches this task by first generating images in 2D spaces us…

Cited by 0SourcePDFScholar
2025

SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG

ICASSP 2025accepted

Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrograms of audio stimuli perceived by subjects at the word or letter level using no…

Cited by 0SourceScholar
2025

Volatile MAB-based Configuration Selection for Offloading Video Analytics Tasks to Edges

ICASSP 2025accepted

The demand for video analytics is increasing rapidly. Due to the limited computational and network resources on edge servers, adjusting video configurations such as resolution and frame rate has become an effective strategy to reduce computational and transmission costs. However, this can also compr…

Cited by 0SourceScholar
2024

Code Membership Inference for Detecting Unauthorized Data Use in Code Pre-trained Language Models

EMNLP 2024finding

Code pre-trained language models (CPLMs) have received great attention since they can benefit various tasks that facilitate software development and maintenance. However, CPLMs are trained on massive open-source code, raising concerns about potential data infringement. This paper launches the study…

2024

DocLens: Multi-aspect Fine-grained Medical Text Evaluation

ACL 2024long

Medical text generation aims to assist with administrative work and highlight salient information to support decision-making.To reflect the specific requirements of medical text, in this paper, we propose a set of metrics to evaluate the completeness, conciseness, and attribution of the generated te…

2024

MonoPlane: Exploiting Monocular Geometric Cues for Generalizable 3D Plane Reconstruction

IROS 2024poster

This paper presents a generalizable 3D plane detection and reconstruction framework named MonoPlane. Unlike previous robust estimator-based works (which require multiple images or RGB-D input) and learning-based works (which suffer from domain shift), MonoPlane combines the best of two worlds and es…

Cited by 1SourcecodeScholar
2024

S3A: Towards Realistic Zero-Shot Classification via Self Structural Semantic Alignment

AAAI 2024technical

Large-scale pre-trained Vision Language Models (VLMs) have proven effective for zero-shot classification. Despite the success, most traditional VLMs-based methods are restricted by the assumption of partial source supervision or ideal target vocabularies, which rarely satisfy the open-world scenario…

2024

UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition

ICLR 2024poster

Large language models (LLMs) have demonstrated remarkable generalizability, such as understanding arbitrary entities and relations. Instruction tuning has proven effective for distilling LLMs into more cost-efficient models such as Alpaca and Vicuna. Yet such student models still trail the original…

Cited by 133SourcePDFScholar
2024

mDPO: Conditional Preference Optimization for Multimodal Large Language Models

EMNLP 2024main

Direct preference optimization (DPO) has shown to be an effective method for large language model (LLM) alignment. Recent works have attempted to apply DPO to multimodal scenarios but have found it challenging to achieve consistent improvement. Through a comparative experiment, we identify the uncon…

2023

Benchmarking Diverse-Modal Entity Linking with Generative Models

ACL 2023findings

Entities can be expressed in diverse formats, such as texts, images, or column names and cell values in tables. While existing entity linking (EL) models work well on per modality configuration, such as text-only EL, visual grounding or schema linking, it is more challenging to design a unified mode…

2023

Continual Contrastive Finetuning Improves Low-Resource Relation Extraction

ACL 2023long

Relation extraction (RE), which has relied on structurally annotated corpora for model training, has been particularly challenging in low-resource scenarios and domains. Recent literature has tackled low-resource RE by self-supervised learning, where the solution involves pretraining the entity pair…

Cited by 7SourcePDFScholar
2023

DecAF: Joint Decoding of Answers and Logical Forms for Question Answering over Knowledge Bases

ICLR 2023poster

Question answering over knowledge bases (KBs) aims to answer natural language questions with factual information such as entities and relations in KBs. Previous methods either generate logical forms that can be executed over KBs to obtain final answers or predict answers directly. Empirical results…

2023

Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness

ICLR 2023top-5%

Neural text-to-SQL models have achieved remarkable performance in translating natural language questions into SQL queries. However, recent studies reveal that text-to-SQL models are vulnerable to task-specific perturbations. Previous curated robustness test sets usually focus on individual phenomena…

Cited by 22SourcePDFScholar
2023

Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge

ACL 2023findings

The open-ended Visual Question Answering (VQA) task requires AI models to jointly reason over visual and natural language inputs using world knowledge. Recently, pre-trained Language Models (PLM) such as GPT-3 have been applied to the task and shown to be powerful world knowledge sources. However, t…

Cited by 17SourcePDFScholar
2023

Importance of Synthesizing High-quality Data for Text-to-SQL Parsing

ACL 2023findings

There has been increasing interest in synthesizing data to improve downstream text-to-SQL tasks. In this paper, we examined the existing synthesized datasets and discovered that state-of-the-art text-to-SQL algorithms did not further improve on popular benchmarks when trained with augmented syntheti…

2023

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

NeurIPS 2023spotlight

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leveraging billions of image-text pairs from the public web, but such general-domain vi…

Cited by 828SourcePDFScholar
2023

Optimizing Bi-Encoder for Named Entity Recognition via Contrastive Learning

ICLR 2023poster

We present a bi-encoder framework for named entity recognition (NER), which applies contrastive learning to map candidate text spans and entity types into the same vector representation space. Prior work predominantly approaches NER as sequence labeling or span classification. We instead frame NER a…

2023

PromptCAL: Contrastive Affinity Learning via Auxiliary Prompts for Generalized Novel Category Discovery

CVPR 2023poster

Although existing semi-supervised learning models achieve remarkable success in learning with unannotated in-distribution data, they mostly fail to learn on unlabeled data sampled from novel semantic classes due to their closed-set assumption. In this work, we target a pragmatic but under-explored G…

2022

CLLE: A Benchmark for Continual Language Learning Evaluation in Multilingual Machine Translation

EMNLP 2022finding

Continual Language Learning (CLL) in multilingual translation is inevitable when new languages are required to be translated. Due to the lack of unified and generalized benchmarks, the evaluation of existing methods is greatly influenced by experimental design which usually has a big gap from the in…

2022

FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing

ACL 2022long

We present a benchmark suite of four datasets for evaluating the fairness of pre-trained language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (European Council, USA, Switzerland, and China), five languages (English, German, French, I…

2022

Knowledge-Rich Self-Supervision for Biomedical Entity Linking

EMNLP 2022finding

Entity linking faces significant challenges such as prolific variations and prevalent ambiguities, especially in high-value domains with myriad entities. Standard classification approaches suffer from the annotation bottleneck and cannot effectively handle unseen entities. Zero-shot entity linking h…

Cited by 45SourcePDFScholar
2022

Locally Aggregated Feature Attribution on Natural Language Model Understanding

NAACL 2022long

With the growing popularity of deep-learning models, model understanding becomes more important. Much effort has been devoted to demystify deep neural networks for better explainability. Some feature attribution methods have shown promising results in computer vision, especially the gradient-based m…

2022

Zero-Shot Dependency Parsing with Worst-Case Aware Automated Curriculum Learning

ACL 2022short

Large multilingual pretrained language models such as mBERT and XLM-RoBERTa have been found to be surprisingly effective for cross-lingual transfer of syntactic parsing models Wu and Dredze (2019), but only between related languages. However, source and training languages are rarely related, when pa…

2021

DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement Learning

ICML 2021spotlight

Games are abstractions of the real world, where artificial agents learn to compete and cooperate with other agents. While significant achievements have been made in various perfect- and imperfect-information games, DouDizhu (a.k.a. Fighting the Landlord), a three-player card game, is still unsolved.…

2021

Finite Sample Analysis of Average-Reward TD Learning and $Q$-Learning

NeurIPS 2021poster

The focus of this paper is on sample complexity guarantees of average-reward reinforcement learning algorithms, which are known to be more challenging to study than their discounted-reward counterparts. To the best of our knowledge, we provide the first known finite sample guarantees using both cons…

Cited by 36SourcePDFScholar
2021

Modular Self-Supervision for Document-Level Relation Extraction

EMNLP 2021main

Extracting relations across large text spans has been relatively underexplored in NLP, but it is particularly important for high-value domains such as biomedicine, where obtaining high recall of the latest findings is crucial for practical applications. Compared to conventional information extractio…

Cited by 10SourcePDFScholar
2019

Greedy Orthogonal Pivoting Algorithm for Non-Negative Matrix Factorization

ICML 2019oral

Non-negative matrix factorization is a powerful tool for learning useful representations in the data and has been widely applied in many problems such as data mining and signal processing. Orthogonal NMF, which can improve the locality of decomposition, has drawn considerable interest in solving clu…