← Search

Hongbin zhou

14 accepted papers

2026

Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets

ICLR 2026poster

Large Language Models (LLMs) have seen remarkable advancements, achieving state-of-the-art results in diverse applications. Fine-tuning, an important step for adapting LLMs to specific downstream tasks, typically involves further training on corresponding datasets. However, a fundamental discrepancy…

Cited by 0SourceScholar
2026

JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization

CVPR 2026

Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed safety filters. Existing jailbreak attacks either rely on proxy-loss optimization instead of the true end-to-end objecti

Cited by 0SourcecodeScholar
2025

Chimera: Improving Generalist Model with Domain-Specific Experts

ICCV 2025poster

Large Multi-modal Models (LMMs), trained on web-scale datasets predominantly composed of natural images, have demonstrated remarkable performance on general tasks. However, these models often exhibit limited specialized capabilities for domain-specific tasks that require extensive domain prior knowl…

Cited by 0SourcePDFScholar
2025

GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training

ICLR 2025poster

Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images…

Cited by 8SourcePDFScholar
2025

LaTeXNet: A Specialized Model for Converting Visual Tables and Equations to LaTeX Code

ICASSP 2025accepted

LaTeX provides precise representation of complex elements (i.e., tables and equations) in scientific documents. However, the automated transcription of visual representations into LaTeX code is challenging and prone to errors. This paper introduces LaTeXNet, a specialized model designed to automate…

Cited by 0SourceScholar
2025

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

CVPR 2025poster

Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems. Despite recent progress, current document parsing methods have not been fairly and comprehensively evaluated due to the nar…

2025

StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

AAAI 2025technical

Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using language model-based or diffusion-based approaches, several challenges remain: 1) cur…

2025

Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling

ACL 2025long

Expressive zero-shot voice conversion (VC) is a critical and challenging task that aims to transform the source timbre into an arbitrary unseen speaker while preserving the original content and expressive qualities. Despite recent progress in zero-shot VC, there remains considerable potential for im…

2024

Promptvc: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

ICASSP 2024accepted

Stylistic voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or reference speech to control the conversion process, which leads to limitations…

Cited by 0SourceScholar
2024

VeloVox: A Low-Cost and Accurate 4D Object Detector with Single-Frame Point Cloud of Livox LiDAR

ICRA 2024poster

Combining motion prediction in LiDAR-based 3D object detection is an effective method for improving overall accuracy, especially the downstream autonomous driving tasks. The recent development of low-cost LiDARs (e.g. Livox LiDAR) enables us to explore such 4D perception systems with a lower budget…

Cited by 1SourcecodeScholar
2024

ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous Driving

NeurIPS 2024poster

Offboard perception aims to automatically generate high-quality 3D labels for autonomous driving (AD) scenes. Existing offboard methods focus on 3D object detection with closed-set taxonomy and fail to match human-level recognition capability on the rapidly evolving perception tasks. Due to heavy re…

2023

DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds

ICCV 2023poster

Existing offboard 3D detectors always follow a modular pipeline design to take advantage of unlimited sequential point clouds. We have found that the full potential of offboard 3D detectors is not explored mainly due to two reasons: (1) the onboard multi-object tracker cannot generate sufficient com…

Cited by 35PDFcodeScholar
2023

Symbolization, Prompt, and Classification: A Framework for Implicit Speaker Identification in Novels

EMNLP 2023long findings

Speaker identification in novel dialogues can be widely applied to various downstream tasks, such as producing multi-speaker audiobooks and converting novels into scripts. However, existing state-of-the-art methods are limited to handling explicit narrative patterns like "Tom said, '...'", unable to…

Cited by 0SourceScholar
2022

Improving Cross-Lingual Speech Synthesis with Triplet Training Scheme

ICASSP 2022accepted

Recent advances in cross-lingual text-to-speech (TTS) made it possible to synthesize speech in a language foreign to a monolingual speaker. However, there is still a large gap between the pronunciation of generated cross-lingual speech and that of native speakers in terms of naturalness and intellig…

Cited by 0SourceScholar