← Search

Junlin Wu

14 accepted papers

2026

DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning

ICLR 2026poster

Real-world large language model deployments (e.g., conversational AI systems, code generation assistants) naturally generate abundant implicit user dissatisfaction (DSAT) signals, as users iterate toward better answers through refinements, corrections, and expressed preferences, while explicit satis…

Cited by 0SourceScholar
2026

Learning Vision-Based Neural Network Controllers with Semi-Probabilistic Safety Guarantees

AAAI 2026technical

Ensuring safety in autonomous systems with vision-based control remains a critical challenge due to the high dimensionality of image inputs and the fact that the relationship between true system state and its visual manifestation is unknown. Existing methods for learning-based control in such setti

Cited by 0SourcePDFScholar
2026

SurgSync: Time-Synchronized Multi-Modal Data Collection Framework and Dataset for Surgical Robotics

ICRA 2026poster

Most existing robotic surgery systems adopt a human-in-the-loop paradigm, often with the surgeon directly teleoperating the robotic system. Adding intelligence to these robots would enable higher-level control, such as supervised autonomy or even full autonomy. However, artificial intelligence (AI) …

2025

Melody Structure Transfer Network: Generating Music with Separable Self-Attention

ICASSP 2025accepted

Most existing symbolic music generation methods focus on generating short pieces, typically less than 8 bars and occasionally up to 32 bars. Generating long music sequences requires effective representation of coherent musical structures. Vanilla self-attention face challenges in capturing subtle lo…

Cited by 0SourceScholar
2025

Towards Green VAE: A Light Pixel-weighting Technique to Enhance Variational AutoEncoder

ICASSP 2025accepted

Variational autoencoders (VAEs) has been a popular generative model for its effectiveness, mathematical foundation, and its impact to other approaches in deep generative learning. For its relatively light-weights and easiness for training, compared with Generative Adversarial Networks (GANs) or othe…

Cited by 0SourceScholar
2024

Axioms for AI Alignment from Human Feedback

NeurIPS 2024spotlight

In the context of reinforcement learning from human feedback (RLHF), the reward function is generally derived from maximum likelihood estimation of a random utility model based on pairwise comparisons made by humans. The problem of learning a reward function is one of preference aggregation that, we…

Cited by 17SourcePDFScholar
2024

RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models

ACL 2024long

Reinforcement Learning with Human Feedback (RLHF) is a methodology designed to align Large Language Models (LLMs) with human preferences, playing an important role in LLMs alignment. Despite its advantages, RLHF relies on human annotators to rank the text, which can introduce potential security vuln…

2024

Verified Safe Reinforcement Learning for Neural Network Dynamic Models

NeurIPS 2024poster

Learning reliably safe autonomous control is one of the core problems in trustworthy autonomy. However, training a controller that can be formally verified to be safe remains a major challenge. We introduce a novel approach for learning verified safe control policies in nonlinear neural dynamical sy…

2023

Exact Verification of ReLU Neural Control Barrier Functions

NeurIPS 2023poster

Control Barrier Functions (CBFs) are a popular approach for safe control of nonlinear systems. In CBF-based control, the desired safety properties of the system are mapped to nonnegativity of a CBF, and the control input is chosen to ensure that the CBF remains nonnegative for all time. Recently, ma…

2023

Neural Lyapunov Control for Discrete-Time Systems

NeurIPS 2023poster

While ensuring stability for linear systems is well understood, it remains a major challenge for nonlinear systems. A general approach in such cases is to compute a combination of a Lyapunov function and an associated control policy. However, finding Lyapunov functions for general nonlinear systems…

2022

Manipulating Elections by Changing Voter Perceptions

IJCAI 2022poster

The integrity of elections is central to democratic systems. However, a myriad of malicious actors aspire to influence election outcomes for financial or political benefit. A common means to such ends is by manipulating perceptions of the voting public about select candidates, for example, through m…

Cited by 6SourcePDFScholar
2022

Robust Deep Reinforcement Learning through Bootstrapped Opportunistic Curriculum

ICML 2022spotlight

Despite considerable advances in deep reinforcement learning, it has been shown to be highly vulnerable to adversarial perturbations to state observations. Recent efforts that have attempted to improve adversarial robustness of reinforcement learning can nevertheless tolerate only very small perturb…

2018

Robust Camera Pose Estimation via Consensus on Ray Bundle and Vector Field

IROS 2018poster

Estimating the camera pose requires point correspondences. However, in practice, correspondences are inevitably corrupted by outliers, which affects the pose estimation. We propose a general and accurate outlier removal strategy for robust camera pose estimation. The proposed strategy can detect out…

Cited by 7SourceScholar
2017

Combining points and lines for camera pose estimation and optimization in monocular visual odometry

IROS 2017poster

In this paper, we propose a unified model for camera pose estimation and a novel strategy for pose optimization by combining points and lines in monocular visual odometry. Our proposed unified model treats point and line features equivalently, which is applicable for all the minimal cases requiring…

Cited by 23SourceScholar