← Search

Yawen Zeng

8 accepted papers

2025

COME: Dual Structure-Semantic Learning with Collaborative MoE for Universal Lesion Detection Across Heterogeneous Ultrasound Datasets

ICCV 2025poster

Conventional single-dataset training often fails with new data distributions, especially in ultrasound (US) image analysis due to limited data, acoustic shadows, and speckle noise.Therefore, constructing a universal framework for multi-heterogeneous US datasets is imperative. However, a key challeng…

2025

DataMan: Data Manager for Pre-training Large Language Models

ICLR 2025poster

The performance emergence of large language models (LLMs) driven by data scaling laws makes the selection of pre-training data increasingly important. However, existing methods rely on limited heuristics and human intuition, lacking comprehensive and clear guidelines. To address this, we are inspir…

Cited by 2SourcePDFScholar
2025

QuantAgents: Towards Multi-agent Financial System via Simulated Trading

EMNLP 2025

In this paper, our objective is to develop a multi-agent financial system that incorporates simulated trading , a technique extensively utilized by financial professionals. While current LLM-based agent models demonstrate competitive performance, they still exhibit significant deviations from real-w

2024

Energy-based Automated Model Evaluation

ICLR 2024poster

The conventional evaluation protocols on machine learning models rely heavily on a labeled, i.i.d-assumed testing dataset, which is not often present in real-world applications. The Automated Model Evaluation (AutoEval) shows an alternative to this traditional workflow, by forming a proximal predict…

2024

Multi-Prompts Learning with Cross-Modal Alignment for Attribute-Based Person Re-identification

AAAI 2024technical

The fine-grained attribute descriptions can significantly supplement the valuable semantic information for person image, which is vital to the success of person re-identification (ReID) task. However, current ReID algorithms typically failed to effectively leverage the rich contextual information av…

2022

Distill The Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine Translation

EMNLP 2022main

Past works on multimodal machine translation (MMT) elevate bilingual setup by incorporating additional aligned vision information.However, an image-must requirement of the multimodal dataset largely hinders MMT’s development — namely that it demands an aligned form of [image, source text, target tex…

2021

Multi-Modal Relational Graph for Cross-Modal Video Moment Retrieval

CVPR 2021poster

Given an untrimmed video and a query sentence, cross-modal video moment retrieval aims to rank a video moment from pre-segmented video moment candidates that best matches the query sentence. Pioneering work typically learns the representations of the textual and visual content separately and then ob…

Cited by 86PDFcodeScholar