← Search

Jingyu Zhang

28 accepted papers

2026

GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward Model

ICML 2026poster

Aligning Multimodal Large Language Models (MLLMs) with human preferences remains a fundamental challenge. While Generative Reward Models (GRMs) offer a promising reasoning-based alternative to scalar models, they are often hindered by severe position bias and prohibitively high computational overhea…

Cited by 0SourceScholar
2026

Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs

ICLR 2026poster

Multimodal Large Language Models (MLLMs) pose critical safety challenges, as they are susceptible not only to adversarial attacks such as jailbreaking but also to inadvertently generating harmful content for benign users. While internal safety alignment via Supervised Fine-Tuning (SFT) and Reinforce…

Cited by 0SourceScholar
2026

Teach to Reason Safely: Policy-Guided Safety Tuning for MLRMs

ICLR 2026poster

Multimodal Large Reasoning Models (MLRMs) have exhibited remarkable capabilities in complex multimodal tasks. However, our findings reveal a critical trade-off: reasoning-based models are more prone to generating harmful content, leading to degradation in safety performance. This paper presents a la…

Cited by 0SourceScholar
2026

The Alignment Waltz: Jointly Training Agents to Collaborate for Safety

ICLR 2026poster

Harnessing the power of LLMs requires a delicate dance between being helpful and harmless, leading to two critical challenges: vulnerability to adversarial attacks that elicit unsafe content, and a tendency for overrefusal on benign but sensitive prompts. Current approaches often navigate this dance…

Cited by 0SourceScholar
2025

Active Modeling and Compensation Control of Yoshimura Manipulator Using Koopman Operator

IROS 2025

The integration of origami structures into soft robotics has enriched the adaptability and functionality of the soft robots. Our research group has developed a cable-driven origami robot attached to an arc frame, which enables its deployment in an MR bore and manipulation of medical tools. However,

Cited by 0SourceScholar
2025

Certified Mitigation of Worst-Case LLM Copyright Infringement

EMNLP 2025

The exposure of large language models (LLMs) to copyrighted material during pre-training raises concerns about unintentional copyright infringement post deployment. This has driven the development of “copyright takedown” methods—post-training approaches aimed at preventing models from generating con

2025

Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements

ICLR 2025poster

The current paradigm for safety alignment of large language models (LLMs) follows a _one-size-fits-all_ approach: the model refuses to interact with any content deemed unsafe by the model provider. This approach lacks flexibility in the face of varying social norms across cultures and regions. In ad…

Cited by 0SourcePDFScholar
2025

Core: Robust Factual Precision with Informative Sub-Claim Identification

ACL 2025finding

Hallucinations pose a challenge to the application of large language models (LLMs) thereby motivating the development of metrics to evaluate factual precision. We observe that popular metrics using the Decompose-Then-Verify framework, such as FActScore, can be manipulated by adding obvious or repeti…

2025

DSRC: Learning Density-Insensitive and Semantic-Aware Collaborative Representation Against Corruptions

AAAI 2025technical

As a potential application of Vehicle-to-Everything (V2X) communication, multi-agent collaborative perception has achieved significant success in 3D object detection. While these methods have demonstrated impressive results on standard benchmarks, the robustness of such approaches in the face of com…

2025

Deformation Configuration Estimation for Soft Continuum Robot Utilizing Seq2Seq Learning

RA-L 2025

Inspired by biological tentacles, soft continuum robots exhibit the potential for navigating through narrow spaces and operating in complex environments, offering extensive application possibilities. However, owing to their inherent compliance, soft continuum robots may undergo unpredictable deforma

Cited by 1SourceScholar
2025

Jailbreak Distillation: Renewable Safety Benchmarking

EMNLP 2025

Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a novel benchmark construction framework that “distills” jailbreak attacks into high-quality and easily-updatable safety ben

Cited by 0SourcePDFScholar
2025

Planning and Compliant Control for Laparoscopic Ultrasound Scanning System

RA-L 2025

Laparoscopic ultrasound (LUS) serves as a critical technology in intraoperative surgeries, particularly for guiding complex procedures in liver diseases. However, the development of robotic systems for LUS examination remains hindered by challenges such as high costs and the absence of force feedbac

Cited by 0SourceScholar
2025

RATIONALYST: Pre-training Process-Supervision for Improving Reasoning

ACL 2025long

The reasoning steps generated by LLMs might be incomplete, as they mimic logical leaps common in everyday communication found in their pre-training data: underlying rationales are frequently left implicit (unstated). To address this challenge, we introduce RATIONALYST, a model for process-supervisio…

2025

SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses

AAAI 2025technical

Can LLMs consistently improve their previous outputs for better results? For this to be true, LLMs would need to be better at discriminating among previously-generated alternatives, than generating initial responses. We explore the validity of this hypothesis in practice. We first formulate a unifie…

2025

TurkingBench: A Challenge Benchmark for Web Agents

NAACL 2025long

Can advanced multi-modal models effectively tackle complex web-based tasks? Such tasks are often found on crowdsourcing platforms, where crowdworkers engage in challenging micro-tasks within web-based environments.Building on this idea, we present TurkingBench, a benchmark consisting of tasks presen…

2025

Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data

NAACL 2025long

To trust the fluent generations of large language models (LLMs), humans must be able to _verify_ their correctness against trusted, external sources. Recent efforts, such as providing citations via retrieved documents or post-hoc provenance, enhance verifiability but provide no guarantees on their c…

2024

DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation

NeurIPS 2024poster

Non-autoregressive Transformers (NATs) are recently applied in direct speech-to-speech translation systems, which convert speech across different languages without intermediate text data. Although NATs generate high-quality outputs and offer faster inference than autoregressive models, they tend to…

2024

ERMVP: Communication-Efficient and Collaboration-Robust Multi-Vehicle Perception in Challenging Environments

CVPR 2024poster

Collaborative perception enhances perception performance by enabling autonomous vehicles to exchange complementary information. Despite its potential to revolutionize the mobile industry challenges in various environments such as communication bandwidth limitations localization errors and informatio…

2024

Learning the Inverse Kinematics of Magnetic Continuum Robot for Teleoperated Navigation

IROS 2024

Magnetic continuum robots are subject to external magnetic fields and deformed remotely, simplifying the robot’s transmission mechanism and providing it with significant potential for miniaturization and operational flexibility. However, modeling magnetic field distribution generated by permanent ma

Cited by 2SourceScholar
2024

SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation

NAACL 2024long

Existing watermarked generation algorithms employ token-level designs and therefore, are vulnerable to paraphrase attacks. To address this issue, we introduce watermarking on the semantic representation of sentences. We propose SemStamp, a robust sentence-level semantic watermarking algorithm that u…

2024

Soft Hybrid Actuated Hierarchical Bronchoscope Robot for Deep Lung Examination

RA-L 2024

Lungdiseases are becoming one of the world's most serious health issues. Soft bronchoscope robots can achieve safe and controllable lung navigation, which will be crucial for the future early examination of lung diseases. However, due to the single driving method, the large size, and insufficient fl

Cited by 10SourceScholar
2024

The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts

ACL 2024findings

As the influence of large language models (LLMs) spans across global communities, their safety challenges in multilingual settings become paramount for alignment research. This paper examines the variations in safety challenges faced by LLMs across different languages and discusses approaches to all…

Cited by 55SourcePDFScholar
2024

k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text

ACL 2024findings

Recent watermarked generation algorithms inject detectable signatures during language generation to facilitate post-hoc detection. While token-level watermarks are vulnerable to paraphrase attacks, SemStamp (Hou et al., 2023) applies watermark on the semantic representation of sentences and demonstr…

2023

An Efficient Multi-solution Solver for the Inverse Kinematics of 3-Section Constant-Curvature Robots

RSS 2023poster

Piecewise constant curvature is a popular kinematics framework for continuum robots. Computing the model parameters from the desired end pose, known as the inverse kinematics problem, is fundamental in manipulation, tracking and planning tasks. In this paper, we propose an efficient multi-solution s…

Cited by 7SourcePDFScholar
2023

Geo-Seq2seq: Twitter User Geolocation on Noisy Data through Sequence to Sequence Learning

ACL 2023findings

Location information can support social media analyses by providing geographic context. Some of the most accurate and popular Twitter geolocation systems rely on rule-based methods that examine the user-provided profile location, which fail to handle informal or noisy location names. We propose Geo-…

Cited by 4SourcePDFScholar
2023

On the Blind Spots of Model-Based Evaluation Metrics for Text Generation

ACL 2023long

In this work, we explore a useful but often neglected methodology for robustness analysis of text generation evaluation metrics: stress tests with synthetic data. Basically, we design and synthesize a wide range of potential errors and check whether they result in a commensurate drop in the metric s…

2023

On the Zero-Shot Generalization of Machine-Generated Text Detectors

EMNLP 2023short findings

The rampant proliferation of large language models, fluent enough to generate text indistinguishable from human-written language, gives unprecedented importance to the detection of machine-generated text. This work is motivated by an important research question: How will the detectors of machine-gen…

Cited by 0SourceScholar
2023

Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception

ICCV 2023poster

Multi-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this em…

Cited by 68PDFcodeScholar