← Search

Haoran Li

109 accepted papers

2026

4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models

CVPR 2026

World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct realistic, dynamic, and physically consistent 3D/4D worlds from images, videos, or text. These models not only need to prod

Cited by 0SourcecodeScholar
2026

ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior

CVPR 2026

Multimodal Large Language Models (MLLMs) are increasingly vulnerable to multimodal Indirect Prompt Injection (IPI) attacks, which embed malicious instructions in images, videos, or audio to hijack model behavior. Existing defenses, designed primarily for text-only LLMs, are unsuitable for countering

Cited by 0SourcecodeScholar
2026

BioFormer: Rethinking Cross-Subject Generalization via Spectral Structural Alignment in Biomedical Time-Series

ICML 2026poster

Cross-subject generalization in biomedical time-series (BTS) refers to training on data from some subjects and testing on unseen subjects. The key challenge is to suppress subject-specific variability in BTS representations. Most existing methods implicitly suppress the variability through model bui…

Cited by 0SourceScholar
2026

CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment

ICRA 2026poster

The spatial information inherent in 3D point clouds is crucial for robotic manipulation. However, existing 3D pre-training methods face a fundamental trade-off: Masked Autoencoding (MAE) excels at capturing spatial-geometric features but lacks semantics, whereas contrastive learning, while able to d…

2026

DiffuDepGrasp: Diffusion-Based Depth Noise Modeling Empowers Sim-To-Real Robotic Grasping

ICRA 2026poster

Accurate spatial-geometric perception remains fundamental to robotic grasping, yet physical artifacts in real depth maps like voids and noise establish a significant sim-to-real gap that critically impedes policy transfer. Training-time strategies like procedural noise injection or learned mappings …

Cited by 0Scholar
2026

FedHPro: Federated Hyper-Prototype Learning via Gradient Matching

ICML 2026poster

Federated Learning (FL) enables collaborative training of distributed clients while protecting privacy. To enhance generalization capability in FL, prototype-based FL is in the spotlight, since shared global prototypes offer semantic anchors for aligning client-specific local prototypes. However, ex…

Cited by 0SourceScholar
2026

Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease Diagnosis

AAAI 2026technical

Multimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention m

Cited by 0SourcePDFScholar
2026

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

ICASSP 2026poster

Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instructions. To address this, we adopt a reinforcement learning (RL) based post-training strategy for MLLMs in multi-image grou…

Cited by 0SourcePDFScholar
2026

KANFIS: A Neuro-Symbolic Framework for Interpretable and Uncertainty-Aware Learning

ICML 2026poster

Adaptive Neuro-Fuzzy Inference System (ANFIS) was designed to combine the learning capabilities of neural network with the reasoning transparency of fuzzy logic. However, conventional ANFIS architectures suffer from structural complexity, where the product-based inference mechanism causes an exponen…

Cited by 0SourceScholar
2026

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

RSS 2026poster

Recent robot foundation models largely rely on large-scale behavior cloning, which imitates expert actions but discards transferable dynamics knowledge embedded in heterogeneous embodied data. While the Unified World Model (UWM) formulation has the potential to leverage such diverse data, existing i…

Cited by 0SourceScholar
2026

Maximizing mutual information between prompt and response improves LLM performance with no additional data

ICML 2026poster

While post-training has successfully improved large language models across a variety of domains from open-ended text generation to mathematics, these gains heavily rely on human-labeled data or external verifiers. Existing data has already been exploited and new high-quality data is expensive to col…

Cited by 0SourceScholar
2026

Memory-Statistics Tradeoff in Continual Learning with Structural Regularization

ICLR 2026poster

We study the statistical performance of a continual learning problem with two linear regression tasks in a well-specified random design setting. We consider a structural regularization algorithm that incorporates a generalized $\ell_2$-regularization tailored to the Hessian of the previous task for…

Cited by 0SourceScholar
2026

On the Tension Between Optimality and Adversarial Robustness in Policy Optimization

ICLR 2026poster

Achieving optimality and adversarial robustness in deep reinforcement learning has long been regarded as conflicting goals. Nonetheless, recent theoretical insights presented in CAR suggest a potential alignment, raising the important question of how to realize this in practice. This paper first ide…

Cited by 0SourceScholar
2026

Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic Verification

ICML 2026poster

Large Language Models (LLMs) show remarkable capabilities, yet their stochastic next-token prediction creates logical inconsistencies and reward hacking that formal symbolic systems avoid. To bridge this gap, we introduce a formal logic verification-guided framework that dynamically interleaves form…

Cited by 0SourceScholar
2026

Reinforcement Learning Control Outperforms Iterative Learning in Exoskeleton-Assisted Gait Training

ICRA 2026poster

Learning-based controllers are increasingly adopted in lower-extremity powered exoskeletons, yet their advantages over traditional adaptive approaches remain underexplored. We compared two adaptive assist-as-needed (AAN) controllers for gait training with an ankle exoskeleton: a reinforcement learni…

Cited by 0Scholar
2026

Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learning

CVPR 2026

Zero-shot unsupervised reinforcement learning (URL) offers a promising direction for building generalist agents capable of generalizing to unseen tasks without additional supervision. Among existing approaches, successor representations (SR) have emerged as a prominent paradigm due to their effectiv

Cited by 0SourcecodeScholar
2026

Semi-Supervised Synthetic Data Generation with Fine-Grained Relevance Control for Short Video Search Relevance Modeling

AAAI 2026technical

Synthetic data is widely adopted in embedding models to ensure diversity in training data distributions across dimensions such as difficulty, length, and language. However, existing prompt-based synthesis methods struggle to capture domain-specific data distributions, particularly in data-scarce dom

Cited by 0SourcePDFScholar
2026

Shear-Based Grasp Control for Multi-Fingered Underactuated Tactile Robotic Hands

ICRA 2026poster

This paper presents a shear-based control scheme for grasping and manipulating delicate objects with a Pisa/IIT anthropomorphic SoftHand equipped with soft biomimetic tactile sensors on all five fingertips. These `microTac' tactile sensors are miniature versions of the TacTip vision-based tactile se…

2026

SoftHand Model-W: A 3D-Printed, Anthropomorphic, Underactuated Robot Hand with Integrated Wrist and Carpal Tunnel

ICRA 2026poster

This paper presents the SoftHand Model-W: a 3D-printed, underactuated, anthropomorphic robot hand based on the Pisa/IIT SoftHand, with an integrated antagonistic tendon mechanism and 2 degree-of-freedom tendon-driven wrist. These four degrees-of-acuation provide active flexion and extension to the f…

2026

StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models

AAAI 2026technical

Human writers often begin their stories with an overarching mental scene, where they envision the interactions between characters and their environment. Inspired by this creative process, we propose a novel approach to long-form story generation, termed hybrid bottom-up long-form story generation, u

Cited by 0SourcePDFScholar
2026

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning

RSS 2026poster

Pretrained on large-scale and diverse datasets, VLA models demonstrate strong generalization and adaptability as general-purpose robotic policies. However, Supervised Fine-Tuning (SFT), which serves as the primary mechanism for adapting VLAs to downstream domains, requires substantial amounts of tas…

Cited by 0SourceScholar
2026

Towards Zero-Shot Diabetic Retinopathy Grading: Learning Generalized Knowledge via Prompt-Driven Matching and Emulating

AAAI 2026technical

As one of the primary causes of visual impairment, Diabetic Retinopathy (DR) requires accurate and robust grading to facilitate timely diagnosis and intervention. Different from conventional DR grading methods that utilize single-view images, recent clinical studies have revealed that multi-view fun

Cited by 0SourcePDFScholar
2026

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models

RSS 2026poster

Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-trained on large and diverse datasets, they typically rely on embodiment-specific fine-tuning to achieve strong performance in downstream tasks. This…

Cited by 0SourceScholar
2025

A-PSRO: A Unified Strategy Learning Method with Advantage Metric for Normal-form Games

ICML 2025poster

Solving the Nash equilibrium in normal-form games with large-scale strategy spaces presents significant challenges. Open-ended learning frameworks, such as PSRO and its variants, have emerged as effective solutions. However, these methods often lack an efficient metric for evaluating strategy improv…

Cited by 0SourcePDFScholar
2025

AFL: A Single-Round Analytic Approach for Federated Learning with Pre-trained Models

CVPR 2025poster

In this paper, we introduce analytic federated learning (AFL), a new training paradigm that brings analytical (i.e., closed-form) solutions to the federated learning (FL) with pre-trained models. Our AFL draws inspiration from analytic learning---a gradient-free technique that trains neural networks…

Cited by 2SourcePDFScholar
2025

Advancing Collaborative Debates with Role Differentiation through Multi-Agent Reinforcement Learning

ACL 2025long

Multi-agent collaborative tasks exhibit exceptional capabilities in natural language applications and generation. By prompting agents to assign clear roles, it is possible to facilitate cooperation and achieve complementary capabilities among LLMs. A common strategy involves adopting a relatively ge…

Cited by 0SourcePDFScholar
2025

Advancing Object-Goal Navigation through LLM-enhanced Object Affinities Transfer

IROS 2025

Object-goal navigation requires mobile robots to efficiently locate targets with visual and spatial information, yet existing methods struggle with generalization in unseen environments. Heuristic approaches with naive metrics fail in complex layouts, while graph-based and learning-based methods suf

Cited by 7SourceScholar
2025

ArenaSim: A High-Performance Simulation Platform for Multi-Robot Self-Play Learning

RA-L 2025

In this letter, we introduce <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">ArenaSim</i>, a novel simulation platform designed for realistic and efficient self-play learning in multi-robot cooperative-competitive games. Compared to previous simulati

Cited by 0SourceScholar
2025

Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods

EMNLP 2025

With the development of technology, large language models (LLMs) have dominated the downstream natural language processing (NLP) tasks. However, because of the LLMs’ instruction-following abilities and inability to distinguish the instructions in the data content, such as web pages from search engin

2025

Can Indirect Prompt Injection Attacks Be Detected and Removed?

ACL 2025long

Prompt injection attacks manipulate large language models (LLMs) by misleading them to deviate from the original input instructions and execute maliciously injected instructions, because of their instruction-following capabilities and inability to distinguish between the original input instructions…

2025

ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

RSS 2025poster

Vision-Language-Action (VLA) models have shown substantial potential in real-world robotic manipulation. However, fine-tuning these models through supervised learning struggles to achieve robust performance due to limited, inconsistent demonstrations, especially in contact-rich environments. In this…

Cited by 6PDFcodeScholar
2025

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning

EMNLP 2025

While Large Language Models (LLMs) exhibit remarkable capabilities, they also introduce significant safety and privacy risks. Current mitigation strategies often fail to preserve contextual reasoning capabilities in risky scenarios. Instead, they rely heavily on sensitive pattern matching to protect

Cited by 0SourcePDFScholar
2025

Defense Against Prompt Injection Attack by Leveraging Attack Techniques

ACL 2025long

With the advancement of technology, large language models (LLMs) have achieved remarkable performance across various natural language processing (NLP) tasks, powering LLM-integrated applications like Microsoft Copilot. However, as LLMs continue to evolve, new vulnerabilities, especially prompt injec…

Cited by 0SourcePDFScholar
2025

Distributed Invariant Kalman Filter for Object-Level Multi-Robot Pose SLAM

ICRA 2025

Cooperative localization and target tracking are essential for multi-robot systems to implement high-level tasks. To this end, we propose a distributed invariant Kalman filter (KF) based on covariance intersection (CI) for effective multi-robot pose estimation. The paper utilizes the object-level me

Cited by 0SourcecodeScholar
2025

Educational SoftHand-A: Building an Anthropomorphic Hand with Soft Synergies using LEGO® MINDSTORMS®

IROS 2025

This paper introduces an anthropomorphic robot hand built entirely using LEGO MINDSTORMS: the Educational SoftHand-A, a tendon-driven, highly-underactuated robot hand based on the Pisa/IIT SoftHand and related hands. To be suitable for an educational context, the design is constrained to use only st

Cited by 0SourceScholar
2025

Essentia: Boosting Artifact Removal from EEG through Semantic Guidance Utilizing Diffusion Model

ICASSP 2025accepted

Electroencephalography (EEG) is a time-series signal containing semantic information that can be used to determine human brain activities. Artifacts within EEG data can interfere with the intrinsic distribution of this semantic information, so removing artifacts is crucial for improving EEG analysis…

Cited by 0SourceScholar
2025

FedDifRC: Unlocking the Potential of Text-to-Image Diffusion Models in Heterogeneous Federated Learning

ICCV 2025poster

Federated learning aims at training models collaboratively across participants while protecting privacy. However, one major challenge for this paradigm is the data heterogeneity issue, where biased data preferences across multiple clients, harming the model's convergence and performance. In this pap…

2025

FetchBot: Learning Generalizable Object Fetching in Cluttered Scenes via Zero-Shot Sim2Real

CoRL 2025oral

Generalizable object fetching in cluttered scenes remains a fundamental and application-critical challenge in embodied AI. Closely packed objects cause inevitable occlusions, making safe action generation particularly difficult. Under such partial observability, effective policies must not only gene…

Cited by 0SourceScholar
2025

GAPartManip: A Large-Scale Part-Centric Dataset for Material-Agnostic Articulated Object Manipulation

ICRA 2025

Effectively manipulating articulated objects in household scenarios is a crucial step toward achieving general embodied artificial intelligence. Mainstream research in 3D vision has primarily focused on manipulation through depth perception and pose detection. However, in real-world environments, th

Cited by 8SourcecodeScholar
2025

HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models

NeurIPS 2025poster

Vision-Language Models (VLMs) have made significant progress in multimodal tasks. However, their performance often deteriorates in long-context scenarios, particularly long videos. While Rotary Position Embedding (RoPE) has been widely adopted for length generalization in Large Language Models (LLMs…

Cited by 0SourceScholar
2025

InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding

EMNLP 2025

Grounding large language models (LLMs) in external knowledge sources is a promising method for faithful prediction. While existing grounding approaches work well for simple queries, many real-world information needs require synthesizing multiple pieces of evidence. We introduce “integrative groundin

2025

MCIP: Protecting MCP Safety via Model Contextual Integrity Protocol

EMNLP 2025

As Model Context Protocol (MCP) introduces an easy-to-use ecosystem for users and developers, it also brings underexplored safety risks. Its decentralized architecture, which separates clients and servers, poses unique challenges for systematic safety analysis. This paper proposes a novel framework

Cited by 0SourcePDFScholar
2025

MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models

ICLR 2025poster

Artificial Intelligence (AI) has demonstrated significant potential in healthcare, particularly in disease diagnosis and treatment planning. Recent progress in Medical Large Vision-Language Models (Med-LVLMs) has opened up new possibilities for interactive diagnostic tools. However, these models oft…

2025

On the Role of Entity and Event Level Conceptualization in Generalizable Reasoning: A Survey of Tasks, Methods, Applications, and Future Directions

EMNLP 2025

Conceptualization, a fundamental element of human cognition, plays a pivotal role in human generalizable reasoning.Generally speaking, it refers to the process of sequentially abstracting specific instances into higher-level concepts and then forming abstract knowledge that can be applied in unfamil

Cited by 0SourcePDFScholar
2025

Prescribed-Time Safe Pursuit Control with Dynamic Obstacle and Occlusion Avoidance

IROS 2025

Performing target tracking and surveillance in dynamic obstacle environments requires maintaining continuous visual focus on the target while ensuring collision avoidance. This paper presents a safety-critical tracking control method that ensures dynamic obstacles remain outside the camera’s line of

Cited by 0SourceScholar
2025

PrivaCI-Bench: Evaluating Privacy with Contextual Integrity and Legal Compliance

ACL 2025long

Recent advancements in generative large language models (LLMs) have enabled wider applicability, accessibility, and flexibility. However, their reliability and trustworthiness are still in doubt, especially for concerns regarding individuals’ data privacy. Great efforts have been made on privacy by…

2025

Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory

NAACL 2025long

Privacy research has attracted wide attention as individuals worry that their private data can be easily leaked during interactions with smart devices, social platforms, and AI applications. Existing works mostly consider privacy attacks and defenses on various sub-fields. Within each field, various…

2025

PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration

ACL 2025long

The widespread usage of online Large Language Models (LLMs) inference services has raised significant privacy concerns about the potential exposure of private information in user inputs. Existing privacy protection methods for LLMs suffer from either insufficient privacy protection with performance…

Cited by 0SourcePDFScholar
2025

Purity Law for Neural Routing Problem Solvers with Enhanced Generalizability

NeurIPS 2025poster

Achieving generalization in neural approaches across different scales and distributions remains a significant challenge for routing problems. A key obstacle is that neural networks often fail to learn robust principles for identifying universal patterns and deriving optimal solutions from diverse in…

Cited by 0SourceScholar
2025

Reinforcement Learning Assist-As-Needed Control Promotes Recovery of Walking Speed Following Ankle Weight Perturbations

IROS 2025

Self-selected walking speed is a key outcome for exercise-based rehabilitation programs following lower-extremity trauma. This work introduces a novel reinforcement learning-based assist-as-needed (RL-AAN) controller for ankle exoskeletons, aimed at gait speed training. Built on an actor–critic arch

Cited by 0SourceScholar
2025

RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis

EMNLP 2025

The success of large language models (LLMs) has attracted many individuals to fine-tune them for domain-specific tasks by uploading their data. However, in sensitive areas like healthcare and finance, privacy concerns often arise. One promising solution is to generate synthetic data with Differentia

2025

Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models

AAAI 2025technical

With rapid advances, generative large language models (LLMs) dominate various Natural Language Processing (NLP) tasks from understanding to reasoning. Yet, language models' inherent vulnerabilities may be exacerbated due to increased accessibility and unrestricted model training on massive data. A m…

2025

So Far Yet So Near: Time Series Data Augmentation with Exploring non-Semantic Boundaries based on Reinforcement Learning

ICASSP 2025accepted

Data augmentation effectively expands feature distribution in time series classification, enhancing downstream task performance. However, existing techniques often fail to maintain semantic consistency between augmented and original time series data, causing label noise and thereby degrading downstr…

Cited by 0SourceScholar
2025

TANDEM: Bi-Level Data Mixture Optimization with Twin Networks

NeurIPS 2025poster

The capabilities of large language models (LLMs) significantly depend on training data drawn from various domains. Optimizing domain-specific mixture ratios can be modeled as a bi-level optimization problem, which we simplify into a single-level penalized form and solve with twin networks: a proxy m…

Cited by 0SourceScholar
2025

The Primacy of Magnitude in Low-Rank Adaptation

NeurIPS 2025spotlight

Low-Rank Adaptation (LoRA) offers a parameter-efficient paradigm for tuning large models. While recent spectral initialization methods improve convergence and performance over the naive “Noise \& Zeros” scheme, their extra computational and storage overhead undermines efficiency. In this paper, we e…

Cited by 0SourceScholar
2025

TopicAttack: An Indirect Prompt Injection Attack via Topic Transition

EMNLP 2025

Large language models (LLMs) have shown remarkable performance across a range of NLP tasks. However, their strong instruction-following capabilities and inability to distinguish instructions from data content make them vulnerable to indirect prompt injection attacks. In such attacks, instructions wi

Cited by 0SourcePDFScholar
2025

Unsupervised Zero-Shot Reinforcement Learning via Dual-Value Forward-Backward Representation

ICLR 2025poster

Online unsupervised reinforcement learning (URL) can discover diverse skills via reward-free pre-training and exhibits impressive downstream task adaptation abilities through further fine-tuning. However, online URL methods face challenges in achieving zero-shot generalization, i.e., directly applyi…

Cited by 0SourcePDFScholar
2025

Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations

NeurIPS 2025poster

Humans can efficiently extract knowledge and learn skills from the videos within only a few trials and errors. However, it poses a big challenge to replicate this learning process for autonomous agents, due to the complexity of visual input, the absence of action or reward signals, and the limitatio…

Cited by 0SourceScholar
2025

Vox-UDA: Voxel-wise Unsupervised Domain Adaptation for Cryo-Electron Subtomogram Segmentation with Denoised Pseudo-Labeling

AAAI 2025technical

Cryo-Electron Tomography (cryo-ET) is a 3D imaging technology that facilitates the study of macromolecular structures at near-atomic resolution. Recent volumetric segmentation approaches on cryo-ET images have drawn widespread interest in the biological sector. However, existing methods heavily rely…

2024

3D-GOI: 3D GAN Omni-Inversion for Multifaceted and Multi-object Editing

ECCV 2024poster

"The current GAN inversion methods typically can only edit the appearance and shape of a single object and background while overlooking spatial information. In this work, we propose a 3D editing framework, to enable multifaceted editing of affine information (scale, translation, and rotation) on mul…

2024

AnyRotate: Gravity-Invariant In-Hand Object Rotation with Sim-to-Real Touch

CoRL 2024poster

Human hands are capable of in-hand manipulation in the presence of different hand motions. For a robot hand, harnessing rich tactile information to achieve this level of dexterity still remains a significant challenge. In this paper, we present AnyRotate, a system for gravity-invariant multi-axis in…

Cited by 19SourceScholar
2024

BioTacTip: A Soft Biomimetic Optical Tactile Sensor for Efficient 3D Contact Localization and 3D Force Estimation

RA-L 2024

In this study, we introduce a new soft biomimetic optical tactile sensor based on mimicking the interlocking structure of the epidermal-dermal boundary: the BioTacTip. The primary sensing unit comprises a sharp white tip surrounded by four black cover tips that when subjected to an external force em

Cited by 23SourceScholar
2024

CNFA: Conditional Normalizing Flow for Query-Limited Attack

ICASSP 2024accepted

Traditional black-box attack methods rely on sufficient feedback from the victim model through a large number of queries until the attack is successful. This may not be acceptable in real applications, since the deployed system may be equipped with certain defense mechanisms and only return the fina…

Cited by 0SourceScholar
2024

Design and Control of a Transformable Multi-Mode Mobile Robot

RA-L 2024

Conventional mobile robots typically include a single locomotion mode and require additional arms to transport objects. To address the challenges of traversing in diverse environments and transporting objects, a novel transformable multi-mode Mecanum-wheeled mobile robot is proposed in this paper. O

Cited by 7SourceScholar
2024

DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling

ECCV 2024poster

"Text-to-3D scene generation holds immense potential for the gaming, film, and architecture sectors. Despite significant progress, existing methods struggle with maintaining high quality, consistency, and editing flexibility. In this paper, we propose , a 3D Gaussian-based novel text-to-3D scene gen…

2024

Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization

NeurIPS 2024poster

With high-dimensional state spaces, visual reinforcement learning (RL) faces significant challenges in exploitation and exploration, resulting in low sample efficiency and training stability. As a time-efficient diffusion model, although consistency models have been validated in online state-based R…

Cited by 4SourcePDFScholar
2024

GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory

EMNLP 2024main

Privacy issues arise prominently during the inappropriate transmission of information between entities. Existing research primarily studies privacy by exploring various privacy attacks, defenses, and evaluations within narrowly predefined patterns, while neglecting that privacy is not an isolated, c…

2024

LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing

EMNLP 2024main

Claim: This work is not advocating the use of LLMs for paper (meta-)reviewing. Instead, wepresent a comparative analysis to identify and distinguish LLM activities from human activities. Two research goals: i) Enable better recognition of instances when someone implicitly uses LLMs for reviewing act…

2024

Metric from Human: Zero-shot Monocular Metric Depth Estimation via Test-time Adaptation

NeurIPS 2024poster

Monocular depth estimation (MDE) is fundamental for deriving 3D scene structures from 2D images. While state-of-the-art monocular relative depth estimation (MRDE) excels in estimating relative depths for in-the-wild images, current monocular metric depth estimation (MMDE) approaches still face chall…

Cited by 4SourcePDFScholar
2024

NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding

EMNLP 2024finding

Large Language Models (LLMs) have sparked substantial interest and debate concerning their potential emergence of Theory of Mind (ToM) ability. Theory of mind evaluations currently focuses on testing models using machine-generated data or game settings prone to shortcuts and spurious correlations, w…

2024

PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models

ACL 2024long

The rapid development of language models (LMs) brings unprecedented accessibility and usage for both models and users. On the one hand, powerful LMs achieve state-of-the-art performance over numerous downstream NLP tasks. On the other hand, more and more attention is paid to unrestricted model acces…

2024

Quad Bayer Joint Demosaicing and Denoising Based on Dual Encoder Network with Joint Residual Learning

AAAI 2024technical

The recent imaging technology Quad Bayer CFA brings better imaging PSNR and higher visual quality compared to traditional Bayer CFA, but also serious challenges for demosaicing and denoising during the ISP pipeline. In this paper, we propose a novel dual encoder network, namely DRNet, to achieve joi…

Cited by 12SourcePDFScholar
2024

RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models

EMNLP 2024main

The recent emergence of Medical Large Vision Language Models (Med-LVLMs) has enhanced medical diagnosis. However, current Med-LVLMs frequently encounter factual issues, often generating responses that do not align with established medical facts. Retrieval-Augmented Generation (RAG), which utilizes e…

2024

TacShade: A New 3D-printed Soft Optical Tactile Sensor Based on Light, Shadow and Greyscale for Shape Reconstruction

ICRA 2024poster

In this paper, we present the TacShade: a newly designed 3D-printed soft optical tactile sensor. The sensor is developed for shape reconstruction under the inspiration of sketch drawing that uses the density of sketch lines to draw light and shadow, resulting in the creation of a 3D-view effect. Tac…

Cited by 1SourceScholar
2024

Towards Optimal Adversarial Robust Q-learning with Bellman Infinity-error

ICML 2024oral

Establishing robust policies is essential to counter attacks or disturbances affecting deep reinforcement learning (DRL) agents. Recent studies explore state-adversarial robustness and suggest the potential lack of an optimal robust policy (ORP), posing challenges in setting strict robustness constr…

2024

ViTacTip: Design and Verification of a Novel Biomimetic Physical Vision-Tactile Fusion Sensor

ICRA 2024poster

Tactile sensing is significant for robotics since it can obtain physical contact information during manipulation. To capture multimodal contact information within a compact framework, we designed a novel sensor called ViTacTip, which seamlessly integrates both tactile and visual perception capabilit…

Cited by 15SourceScholar
2023

Bi-Touch: Bimanual Tactile Manipulation With Sim-to-Real Deep Reinforcement Learning

RA-L 2023

Bimanual manipulation with tactile feedback will be key to human-level robot dexterity. However, this topic is less explored than single-arm settings, partly due to the availability of suitable hardware along with the complexity of designing effective controllers for tasks with relatively large stat

Cited by 49SourceScholar
2023

Multi-step Jailbreaking Privacy Attacks on ChatGPT

EMNLP 2023long findings

With the rapid progress of large language models (LLMs), many downstream NLP tasks can be well solved given appropriate prompts. Though model developers and researchers work hard on dialog safety to avoid generating harmful content from LLMs, it is still challenging to steer AI-generated content (AI…

Cited by 0SourcecodeScholar
2023

Pretrained Transformers for Seizure Detection

ICASSP 2023accepted

Epilepsy is a neurological disorder characterized by seizures that can disrupt a patient’s quality of life. EEG has been used to detect underlying neural activity for diagnosis and treatment. However, standard methods of seizure detection are time-consuming and require manual detection by a trained…

Cited by 0SourceScholar
2023

Priori Anchor Labels Supervised Scalable Multi-View Bipartite Graph Clustering

AAAI 2023technical

Although multi-view clustering (MVC) has achieved remarkable performance by integrating the complementary information of views, it is inefficient when facing scalable data. Proverbially, anchor strategy can mitigate such a challenge a certain extent. However, the unsupervised dynamic strategy usuall…

Cited by 15SourcePDFScholar
2023

Sentence Embedding Leaks More Information than You Expect: Generative Embedding Inversion Attack to Recover the Whole Sentence

ACL 2023findings

Sentence-level representations are beneficial for various natural language processing tasks. It is commonly believed that vector representations can capture rich linguistic properties. Currently, large language models (LMs) achieve state-of-the-art performance on sentence embedding. However, some re…

2023

Sub-network Discovery and Soft-masking for Continual Learning of Mixed Tasks

EMNLP 2023long findings

Continual learning (CL) has two main objectives: preventing catastrophic forgetting (CF) and encouraging knowledge transfer (KT). The existing literature mainly focused on overcoming CF. Some work has also been done on KT when the tasks are similar. To our knowledge, only one method has been propose…

Cited by 0SourcecodeScholar
2023

Tactile-Driven Gentle Grasping for Human-Robot Collaborative Tasks

ICRA 2023poster

This paper presents a control scheme for force sensitive, gentle grasping with a Pisa/IIT anthropomorphic SoftHand equipped with a miniaturised version of the TacTip optical tactile sensor on all five fingertips. The tactile sensors provide high-resolution information about a grasp and how the finge…

Cited by 10SourceScholar
2023

Tuna: Instruction Tuning using Feedback from Large Language Models

EMNLP 2023long findings

Instruction tuning of open-source large language models (LLMs) like LLaMA, using direct outputs from more powerful LLMs such as Instruct-GPT and GPT-4, has proven to be a cost-effective way to align model behaviors with human preferences. However, the instruction-tuned model has only seen one respon…

Cited by 0SourcecodeScholar
2023

Visual-Tactile Robot Grasping Based on Human Skill Learning From Demonstrations Using a Wearable Parallel Hand Exoskeleton

RA-L 2023

The soft fingers and strategic grasping skills enable the human hands to grasp objects in a stable manner. This letter is to model human grasping skills and transfer the learned skills to robots to improve grasping quality and success rate. First, we designed a wearable tool-like parallel hand exosk

Cited by 24SourceScholar
2022

AnswerSumm: A Manually-Curated Dataset and Pipeline for Answer Summarization

NAACL 2022long

Community Question Answering (CQA) fora such as Stack Overflow and Yahoo! Answers contain a rich resource of answers to a wide range of community-based questions. Each question thread can receive a large number of answers with different perspectives. One goal of answer summarization is to produce a…

2022

BRL/Pisa/IIT SoftHand: A Low-Cost, 3D-Printed, Underactuated, Tendon-Driven Hand With Soft and Adaptive Synergies

RA-L 2022

This letter introduces the BRL/Pisa/IIT (BPI) SoftHand: a single actuator-driven, low-cost, 3D-printed, tendon-driven, underactuated robot hand that can be used to perform a range of grasping tasks. Based on the adaptive synergies of the Pisa/IIT SoftHand, we design a new joint system and tendon rou

Cited by 37SourceScholar
2022

CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning

NAACL 2022long

Factual inconsistencies in generated summaries severely limit the practical applications of abstractive dialogue summarization. Although significant progress has been achieved by using pre-trained neural language models, substantial amounts of hallucinated content are found during the human evaluati…

Cited by 71SourcePDFScholar
2022

Investigating Crowdsourcing Protocols for Evaluating the Factual Consistency of Summaries

NAACL 2022long

Current pre-trained models applied for summarization are prone to factual inconsistencies that misrepresent the source text. Evaluating the factual consistency of summaries is thus necessary to develop better models. However, the human evaluation setup for evaluating factual consistency has not been…

Cited by 20SourcePDFScholar
2022

JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and Summarization

EMNLP 2022main

The popularity of multimodal dialogue has stimulated the need for a new generation of dialogue agents with multimodal interactivity.When users communicate with customer service, they may express their requirements by means of text, images, or even videos. Visual information usually acts as discrimin…

2022

PRINCE: Prefix-Masked Decoding for Knowledge Enhanced Sequence-to-Sequence Pre-Training

EMNLP 2022main

Pre-trained Language Models (PLMs) have shown effectiveness in various Natural Language Processing (NLP) tasks. Denoising autoencoder is one of the most successful pre-training frameworks, learning to recompose the original text given a noise-corrupted one. The existing studies mainly focus on injec…

2022

STRUDEL: Structured Dialogue Summarization for Dialogue Comprehension

EMNLP 2022main

Abstractive dialogue summarization has long been viewed as an important standalone task in natural language processing, but no previous work has explored the possibility of whether abstractive dialogue summarization can also be used as a means to boost an NLP system’s performance on other important…

Cited by 2SourcePDFScholar
2022

You Don’t Know My Favorite Color: Preventing Dialogue Representations from Revealing Speakers’ Private Personas

NAACL 2022long

Social chatbots, also known as chit-chat chatbots, evolve rapidly with large pretrained language models. Despite the huge progress, privacy concerns have arisen recently: training data of large language models can be extracted via model inversion attacks. On the other hand, the datasets used for tra…

2021

ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument Mining

ACL 2021long

While online conversations can cover a vast amount of information in many different formats, abstractive text summarization has primarily focused on modeling solely news articles. This research gap is due, in part, to the lack of standardized datasets for summarizing online discussions. To address t…

2021

Improving Zero and Few-Shot Abstractive Summarization with Intermediate Fine-tuning and Data Augmentation

NAACL 2021long

Models pretrained with self-supervised objectives on large text corpora achieve state-of-the-art performance on English text summarization tasks. However, these models are typically fine-tuned on hundreds of thousands of data points, an infeasible requirement when applying summarization to new, nich…

Cited by 117SourcePDFScholar
2021

K-PLUG: Knowledge-injected Pre-trained Language Model for Natural Language Understanding and Generation in E-Commerce

EMNLP 2021finding

Existing pre-trained language models (PLMs) have demonstrated the effectiveness of self-supervised learning for a broad range of natural language processing (NLP) tasks. However, most of them are not explicitly aware of domain-specific knowledge, which is essential for downstream tasks in many domai…

2021

Learn to Copy from the Copying History: Correlational Copy Network for Abstractive Summarization

EMNLP 2021main

The copying mechanism has had considerable success in abstractive summarization, facilitating models to directly copy words from the input text to the output summary. Existing works mostly employ encoder-decoder attention, which applies copying at each time step independently of the former ones. How…

2021

Syntax-augmented Multilingual BERT for Cross-lingual Transfer

ACL 2021long

In recent years, we have seen a colossal effort in pre-training multilingual text encoders using large-scale corpora in many languages to facilitate cross-lingual transfer learning. However, due to typological differences across languages, the cross-lingual transfer is challenging. Nevertheless, lan…

2020

AutoGAN-Distiller: Searching to Compress Generative Adversarial Networks

ICML 2020poster

The compression of Generative Adversarial Networks (GANs) has lately drawn attention, due to the increasing demand for deploying GANs into mobile devices for numerous applications such as image translation, enhancement and editing. However, compared to the substantial efforts to compressing other de…

2020

Multimodal Sentence Summarization via Multimodal Selective Encoding

COLING 2020main

This paper studies the problem of generating a summary for a given sentence-image pair. Existing multimodal sequence-to-sequence approaches mainly focus on enhancing the decoder by visual signals, while ignoring that the image can improve the ability of the encoder to identify highlights of a news e…

Cited by 41SourcePDFScholar
2020

On the Faithfulness for E-commerce Product Summarization

COLING 2020main

In this work, we present a model to generate e-commerce product summaries. The consistency between the generated summary and the product attributes is an essential criterion for the ecommerce product summarization task. To enhance the consistency, first, we encode the product attribute table to guid…