← Search

Haibin Wang

5 accepted papers

2025

Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models

NeurIPS 2025spotlight

Foundation models demand advanced data processing for their vast, multimodal datasets. However, traditional frameworks struggle with the unique complexities of multimodal data. In response, we present Data-Juicer 2.0, a data processing system backed by 100+ data processing operators spanning text, i…

Cited by 0SourcecodeScholar
2025

Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development

ICML 2025spotlight

The emergence of multimodal large models has advanced artificial intelligence, introducing unprecedented levels of performance and functionality. However, optimizing these models remains challenging due to historically isolated paths of model-centric and data-centric developments, leading to subopti…

Cited by 0SourcePDFScholar
2023

PreNAS: Preferred One-Shot Learning Towards Efficient Neural Architecture Search

ICML 2023poster

The wide application of pre-trained models is driving the trend of once-for-all training in one-shot neural architecture search (NAS). However, training within a huge sample space damages the performance of individual subnets and requires much computation to search for a optimal model. In this paper…

2022

HIE-SQL: History Information Enhanced Network for Context-Dependent Text-to-SQL Semantic Parsing

ACL 2022findings

Recently, context-dependent text-to-SQL semantic parsing which translates natural language into SQL in an interaction process has attracted a lot of attentions. Previous works leverage context dependence information either from interaction history utterances or previous predicted queries but fail in…

Cited by 35SourcePDFScholar
2018

A Deep Neural Network Based Method of Source Localization in a Shallow Water Environment

ICASSP 2018accepted

This paper applies deep neural network (DNN) to source localization in a shallow water environment because of its powerful modeling capability and the little dependence on the prior knowledge of environmental parameters. The classical two-stage scheme is adopted, in which feature extraction and DNN…

Cited by 0SourceScholar