← Search

Wei Xu

140 accepted papers

2026

Bridging the Semantic Gap: Leveraging LLMs for Hierarchical Interest Evolution in Sequential Recommendation

IJCAI 2026

Accurate user behavior modeling is fundamental to the prediction of click-through rates (CTR) in industrial recommendation systems and online advertising. Traditional discriminative models, which rely on isolated ID features, struggle to capture the evolving nature of user intents across multiple ch

Cited by 0Scholar
2026

Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?

ICLR 2026poster

Vision-language models (VLMs) have demonstrated strong performance in image geolocation, a capability further sharpened by frontier multimodal large reasoning models (MLRMs). This poses a significant privacy risk, as these widely accessible models can be exploited to infer sensitive locations from c…

Cited by 0SourceScholar
2026

Flipping the Dialogue: Training and Evaluating User Language Models

ICLR 2026poster

Conversations with LMs involve two participants: a human user leading the conversation, and an LM assistant responding to the user's request. To satisfy this specific role, LMs are post-trained to be helpful assistants - optimized to produce exhaustive and well-structured responses, often free of am…

Cited by 0SourceScholar
2026

GDP: Enhancing End-To-End Autonomous Driving with Goal-Driven Planner

ICRA 2026poster

End-to-end (E2E) autonomous driving has emerged as a promising paradigm with the pervasive power of model architectures and the availability of large-scale driving datasets. Despite tremendous efforts in recent research, most E2E driving frameworks rely on rather general driving commands, such as "G…

Cited by 0Scholar
2026

Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking

ICML 2026oral

Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We introduce **JAILBREAK FOUNDRY (JBF)**, a system that addresses this gap via a m…

Cited by 0SourceScholar
2026

Kronos: A Foundation Model for the Language of Financial Markets

AAAI 2026technical

The success of large-scale pre-training paradigm, exemplified by Large Language Models (LLMs), has inspired the development of Time Series Foundation Models (TSFMs). However, their application to financial candlestick (K-line) data remains limited, often underperforming non-pre-trained architectures

Cited by 0SourcePDFScholar
2026

Learning to Route Languages for Multilingual Preference Optimization

ICML 2026poster

Large language models (LLMs) are trained on heterogeneous multilingual corpora, yet existing preference optimization methods often implicitly restrict each training question to a single response language or rely on a fixed dominant language for supervision. We propose language-routed preference opti…

Cited by 0SourceScholar
2026

UniVBench: Towards Unified Evaluation for Video Foundation Models

CVPR 2026

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation benchmarks remain fragmented and limited in scope, as they each

Cited by 0SourcecodeScholar
2025

A Generalized Bisimulation Metric of State Similarity between Markov Decision Processes: From Theoretical Propositions to Applications

NeurIPS 2025poster

The bisimulation metric (BSM) is a powerful tool for computing state similarities within a Markov decision process (MDP), revealing that states closer in BSM have more similar optimal value functions. While BSM has been successfully utilized in reinforcement learning (RL) for tasks like state repres…

Cited by 0SourceScholar
2025

Automated Red Teaming for Text-to-Image Models through Feedback-Guided Prompt Iteration with Vision-Language Models

ICCV 2025poster

Text-to-image models have achieved remarkable progress in generating high-quality images from textual prompts, yet their potential for misuse like generating unsafe content remains a critical concern. Existing safety mechanisms, such as filtering and fine-tuning, remain insufficient in preventing vu…

2025

Blurry-Edges: Photon-Limited Depth Estimation from Defocused Boundaries

CVPR 2025poster

Extracting depth information from photon-limited, defocused images is challenging because depth from defocus (DfD) relies on accurate estimation of defocus blur, which is fundamentally sensitive to image noise. We present a novel approach to robustly measure object depths from photon-limited images…

Cited by 0SourcePDFScholar
2025

CARE: Multilingual Human Preference Learning for Cultural Awareness

EMNLP 2025

Language Models (LMs) are typically tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied. In this paper, we systematically analyze how native human cultural preferences can be incorpora

2025

CROSSNEWS: A Cross-Genre Authorship Verification and Attribution Benchmark

AAAI 2025technical

Authorship models have historically generalized poorly to new domains because of the wide distribution of author-identifying signals across domains. In particular, the effects of topic and genre are highly domain-dependent and impact authorship analysis performance greatly. This paper addresses the…

2025

Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?

ACL 2025finding

Jailbreak attacks have been observed to largely fail against recent reasoning models enhanced by Chain-of-Thought (CoT) reasoning. However, the underlying mechanism remains underexplored, and relying solely on reasoning capacity may raise security concerns. In this paper, we try to answer the questi…

Cited by 0SourcePDFScholar
2025

Generalization of Humanoid Arm Manipulation Based on Keypoints Synergetic Guidance

RA-L 2025

In humanoid robot manipulation imitation learning, arm and tool synergies are required to accomplish tasks. However, the existence of arm and tool shape variations within the demonstrators and between the demonstrator-robot impacts the generalization performance. This paper models the arm and tool a

Cited by 0SourceScholar
2025

Generating CAD Code with Vision-Language Models for 3D Designs

ICLR 2025poster

Generative AI has transformed the fields of Design and Manufacturing by providing efficient and automated methods for generating and modifying 3D objects. One approach involves using Large Language Models (LLMs) to generate Computer- Aided Design (CAD) scripting code, which can then be executed to r…

Cited by 4SourcePDFScholar
2025

Hierarchical Reinforcement Learning for Articulated Tool Manipulation with Multifingered Hand

IROS 2025

Manipulating articulated tools, such as tweezers or scissors, has rarely been explored in previous research. Unlike rigid tools, articulated tools change their shape dynamically, creating unique challenges for dexterous robotic hands. In this work, we present a hierarchical, goal-conditioned reinfor

Cited by 2SourceScholar
2025

How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation

EMNLP 2025

As Large Language Models (LLMs) are widely deployed in diverse scenarios, the extent to which they could tacitly spread misinformation emerges as a critical safety concern. Current research primarily evaluates LLMs on explicit false statements, overlooking how misinformation often manifests subtly a

2025

Learning Multi-Stage Pick-and-Place With a Legged Mobile Manipulator

RA-L 2025

Quadruped-based mobile manipulation presents significant challenges in robotics due to the diversity of required skills, the extended task horizon, and partial observability. After presenting a multi-stage pick-and-place task as a succinct yet sufficiently rich setup that captures key desiderata for

Cited by 4SourcecodeScholar
2025

Map2Traj: Street Map Piloted Zero-shot Trajectory Generation Method for Wireless Network Optimization

IJCAI 2025

In modern wireless networks, user mobility modeling plays a pivotal role in learning-based network optimization, particularly in tasks such as user association and resource allocation. Traditional random mobility models, e.g., random waypoint and Gauss Markov model, often fail to accurately capture

Cited by 0SourcePDFScholar
2025

NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding

NeurIPS 2025poster

Underwater exploration offers critical insights into our planet and attracts increasing attention for its broader applications in resource exploration, national security, etc. We study the underwater scene understanding methods, which aim to achieve automated underwater exploration. The underwater s…

Cited by 0SourcecodeScholar
2025

Nuclear Deployed!: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents

ACL 2025finding

Large language models (LLMs) are evolving into autonomous decision-makers, raising concerns about catastrophic risks in high-stakes scenarios, particularly in Chemical, Biological, Radiological and Nuclear (CBRN) domains. Based on the insight that such risks can originate from trade-offs between the…

2025

On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena

NAACL 2025long

Language Models (LMs) have been shown to exhibit a strong preference towards entities associated with Western culture when operating in non-Western languages. In this paper, we aim to uncover the origins of entity-related cultural biases in LMs by analyzing several contributing factors, including th…

2025

Opportunistic Collaborative Planning with Large Vision Model Guided Control and Joint Query-Service Optimization

IROS 2025

Navigating autonomous vehicles in open scenarios is a challenge due to the difficulties in handling unseen objects. Existing solutions either rely on small models that struggle with generalization or large models that are resource-intensive. While collaboration between the two offers a promising sol

Cited by 1SourceScholar
2025

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning

EMNLP 2025

Instruction tuning has underscored the significant potential of large language models (LLMs) in producing more human controllable and effective outputs in various domains. In this work, we focus on the data selection problem for task-specific instruction tuning of LLMs. Prevailing methods primarily

2025

SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkage

ACL 2025finding

Large language models (LLMs) have made significant advancements across various tasks, but their safety alignment remains a major concern. Exploring jailbreak prompts can expose LLMs’ vulnerabilities and guide efforts to secure them. Existing methods primarily design sophisticated instructions for th…

2025

SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?

EMNLP 2025

Large language models (LLMs) are increasingly used in interactive applications, and human evaluation remains the gold standard for assessing their performance in multi-turn conversations. Since human studies are costly, time-consuming, and hard to reproduce, recent work explores using LLMs to simula

Cited by 0SourcePDFScholar
2025

The Impact of Visual Information in Chinese Characters: Evaluating Large Models’ Ability to Recognize and Utilize Radicals

NAACL 2025long

The glyphic writing system of Chinese incorporates information-rich visual features in each character, such as radicals that provide hints about meaning or pronunciation. However, there has been no investigation into whether contemporary Large Language Models (LLMs) and Vision-Language Models (VLMs)…

2025

WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

ICLR 2025poster

Large language models (LLMs) have shown remarkable potential as autonomous agents, particularly in web-based tasks. However, existing LLM web agents face significant limitations: high-performing agents rely on expensive proprietary LLM APIs, while open LLMs lack the necessary decision-making capabi…

2025

iDPA: Instance Decoupled Prompt Attention for Incremental Medical Object Detection

ICML 2025poster

Existing prompt-based approaches have demonstrated impressive performance in continual learning, leveraging pre-trained large-scale models for classification tasks; however, the tight coupling between foreground-background information and the coupled attention between prompts and image-text tokens p…

Cited by 0SourcePDFScholar
2024

A Unified Framework for 3D Scene Understanding

NeurIPS 2024poster

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are typically tailored to a specific task, limiting their underst…

2024

Can Large Language Models Mine Interpretable Financial Factors More Effectively? A Neural-Symbolic Factor Mining Agent Model

ACL 2024findings

Finding interpretable factors for stock returns is the most vital issue in the empirical asset pricing domain. As data-driven methods, existing factor mining models can be categorized into symbol-based and neural-based models. Symbol-based models are interpretable but inefficient, while neural-based…

Cited by 1SourcePDFScholar
2024

ChatHF: Collecting Rich Human Feedback from Real-time Conversations

EMNLP 2024system demonstrations

We introduce ChatHF, an interactive annotation framework for chatbot evaluation, which integrates configurable annotation within a chat interface. ChatHF can be flexibly configured to accommodate various chatbot evaluation tasks, for example detecting offensive content, identifying incorrect or misl…

Cited by 0SourcePDFScholar
2024

Continuous Field Reconstruction from Sparse Observations with Implicit Neural Networks

ICLR 2024poster

Reliably reconstructing physical fields from sparse sensor data is a challenge that frequenty arises in many scientific domains. In practice, the process generating the data is often not known to sufficient accuracy. Therefore, there is a growing interest in the deep neural network route to the prob…

2024

Course-Correction: Safety Alignment Using Synthetic Preferences

EMNLP 2024industry

The risk of harmful contents generated by large language models (LLMs) becomes a critical concern. This paper systematically evaluates and enhances LLMs’ capability to perform course-correction, , the model can steer away from generating harmful content autonomously. First, we introduce the C2-Eval…

2024

Dynamic Adapter Meets Prompt Tuning: Parameter-Efficient Transfer Learning for Point Cloud Analysis

CVPR 2024poster

Point cloud analysis has achieved outstanding performance by transferring point cloud pre-trained models. However existing methods for model adaptation usually update all model parameters i.e. full fine-tuning paradigm which is inefficient as it relies on high computational costs (e.g. training GPU…

2024

FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence

ACL 2024long

Plain language summarization with LLMs can be useful for improving textual accessibility of technical content. But how factual are these summaries in a high-stakes domain like medicine? This paper presents FactPICO, a factuality benchmark for plain language summarization of medical texts describing…

2024

FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning

ICML 2024poster

In this work, we investigate how to leverage pre-trained visual-language models (VLM) for online Reinforcement Learning (RL). In particular, we focus on sparse reward tasks with pre-defined textual task descriptions. We first identify the problem of reward misalignment when applying VLM as a reward…

2024

Granular Privacy Control for Geolocation with Vision Language Models

EMNLP 2024main

Vision Language Models (VLMs) are rapidly advancing in their capability to answer information-seeking questions. As these models are widely deployed in consumer applications, they could lead to new privacy risks due to emergent abilities to identify people in photos, geolocate images, etc. As we dem…

2024

Having Beer after Prayer? Measuring Cultural Bias in Large Language Models

ACL 2024long

As the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial. Despite advancements in multilingual capabilities, models are not designed with appropriate cultural nuances. In this paper, we show that multilingual and Arabic monolin…

2024

InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification

ACL 2024long

Text simplification aims to make technical texts more accessible to laypeople but often results in deletion of information and vagueness. This work proposes InfoLossQA, a framework to characterize and recover simplification-induced information loss in form of question-and-answer (QA) pairs. Building…

2024

Integrating Sensing, Communication, and Computation in the Sky

ICASSP 2024accepted

Unmanned Aerial Vehicle (UAV)-mounted edge devices are particularly advantageous for federated edge learning (FEEL) due to their flexibility and mobility in efficient data collection. In UAV-assisted FEEL, sensing, computation, and communication are coupled and compete for limited onboard resources,…

Cited by 0SourceScholar
2024

Kill Two Birds with One Stone: Rethinking Data Augmentation for Deep Long-tailed Learning

ICLR 2024poster

Real-world tasks are universally associated with training samples that exhibit a long-tailed class distribution, and traditional deep learning models are not suitable for fitting this distribution, thus resulting in a biased trained model. To surmount this dilemma, massive deep long-tailed learning…

Cited by 13SourcePDFScholar
2024

Knowledge Conflicts for LLMs: A Survey

EMNLP 2024main

This survey provides an in-depth analysis of knowledge conflicts for large language models (LLMs), highlighting the complex challenges they encounter when blending contextual and parametric knowledge. Our focus is on three categories of knowledge conflicts: context-memory, inter-context, and intra-m…

2024

MFCalib: Single-shot and Automatic Extrinsic Calibration for LiDAR and Camera in Targetless Environments Based on Multi-Feature Edge

IROS 2024poster

This paper presents MFCalib, an innovative extrinsic calibration technique for LiDAR and RGB camera that operates automatically in targetless environments with a single data capture. At the heart of this method is using a rich set of edge information, significantly enhancing calibration accuracy and…

Cited by 3SourcecodeScholar
2024

Make Bricks with a Little Straw: Large-Scale Spatio-Temporal Graph Learning with Restricted GPU-Memory Capacity

IJCAI 2024poster

Traffic prediction plays a key role in various smart city applications, which can help traffic managers make traffic plans in advance, assist online ride-hailing companies in deploying vehicles reasonably, and provide early warning of congestion for safety authorities. While increasingly complex mod…

Cited by 2SourcePDFScholar
2024

Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding

ACL 2024long

Language style is often used by writers to convey their intentions, identities, and mastery of language. In this paper, we show that current large language models struggle to capture some language styles without fine-tuning. To address this challenge, we investigate whether LLMs can be meta-trained…

2024

NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms

ACL 2024long

The performance of Large Language Models (LLMs) degrades from the temporal drift between data used for model training and newer text seen during inference. One understudied avenue of language change causing data drift is the emergence of neologisms – new word forms – over time. We create a diverse r…

2024

PointMamba: A Simple State Space Model for Point Cloud Analysis

NeurIPS 2024poster

Transformers have become one of the foundational architectures in point cloud analysis tasks due to their excellent global modeling ability. However, the attention mechanism has quadratic complexity, making the design of a linear complexity method with global modeling appealing. In this paper, we pr…

2024

ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment

EMNLP 2024main

We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This paper introduces ReadMe++, a multilingual multi-domain data…

2024

Reducing Privacy Risks in Online Self-Disclosures with Language Models

ACL 2024long

Self-disclosure, while being common and rewarding in social media interaction, also poses privacy risks. In this paper, we take the initiative to protect the user-side privacy associated with online self-disclosure through detection and abstraction. We develop a taxonomy of 19 self-disclosure catego…

Cited by 19SourcePDFScholar
2024

Robot Policy Learning with Temporal Optimal Transport Reward

NeurIPS 2024poster

Reward specification is one of the most tricky problems in Reinforcement Learning, which usually requires tedious hand engineering in practice. One promising approach to tackle this challenge is to adopt existing expert video demonstrations for policy learning. Some recent work investigates how to l…

2024

Seamless Virtual Reality With Integrated Synchronizer and Synthesizer for Autonomous Driving

RA-L 2024

Virtual reality (VR) is a promising data engine for autonomous driving (AD). However, data fidelity in this paradigm is often degraded by VR inconsistency, for which the existing VR approaches become ineffective, as they ignore the inter-dependency between low-level VR synchronizer designs (i.e., da

Cited by 8SourceScholar
2024

Semi-Autonomous Grasping Control of Prosthetic Hand and Wrist Based on Motion Prior Field

RA-L 2024

Grasping multiple affordance parts and from arbitrary directions for complex shaped objects still remains a challenging problem for prosthetic hand with wrist. We propose a semi-autonomous control method that uses only an integrated in-hand camera to predict the final grasping part on an object as t

Cited by 7SourceScholar
2024

Solving Motion Planning Tasks with a Scalable Generative Model

ECCV 2024poster

"As autonomous driving systems being deployed to millions of vehicles, there is a pressing need of improving the system’s scalability, safety and reducing the engineering cost. A realistic, scalable, and practical simulator of the driving world is highly desired. In this paper, we present an efficie…

2024

The Earth is Flat because...: Investigating LLMs’ Belief towards Misinformation via Persuasive Conversation

ACL 2024long

Large language models (LLMs) encapsulate vast amounts of knowledge but still remain vulnerable to external misinformation. Existing research mainly studied this susceptibility behavior in a single-turn setting. However, belief can change during a multi-turn conversation, especially a persuasive one.…

Cited by 61SourcePDFScholar
2024

Two Fists, One Heart: Multi-Objective Optimization Based Strategy Fusion for Long-tailed Learning

ICML 2024poster

Real-world data generally follows a long-tailed distribution, which makes traditional high-performance training strategies unable to show their usual effects. Various insights have been proposed to alleviate this challenging distribution. However, some observations indicate that models trained on lo…

Cited by 4SourcePDFScholar
2024

VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAE

ICLR 2024poster

Unsupervised video object learning seeks to decompose video scenes into structural object representations without any supervision from depth, optical flow, or segmentation. We present VONet, an innovative approach that is inspired by MONet. While utilizing a U-Net architecture, VONet employs an effi…

2024

Variable Admittance Control Using Velocity-Curvature Patterns to Enhance Physical Human-Robot Interaction

RA-L 2024

This letter introduces a variable admittance control approach aimed at enhancing intuitive human-robot interaction by considering both direct and indirect human intentions. The magnitude of force serves as a representation of direct intentions, delineating preferences for rapid or precise motions. D

Cited by 4SourceScholar
2024

Walking in Others’ Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias

EMNLP 2024main

The common toxicity and societal bias in contents generated by large language models (LLMs) necessitate strategies to reduce harm. Present solutions often demand white-box access to the model or substantial training, which is impractical for cutting-edge commercial LLMs. Moreover, prevailing prompti…

Cited by 10SourcePDFScholar
2023

A Computational Interface to Translate Strategic Intent from Unstructured Language in a Low-Data Setting

EMNLP 2023long findings

Many real-world tasks involve a mixed-initiative setup, wherein humans and AI systems collaboratively perform a task. While significant work has been conducted towards enabling humans to specify, through language, exactly how an agent should complete a task (i.e., low-level specification), prior wor…

Cited by 0SourcecodeScholar
2023

Color Guided Depth Map Super-Resolution with Nonlocla Autoregres-Sive Modeling

ICASSP 2023accepted

Depth map captured by 3D cameras usually suffers from low resolution and insufficient quality, which limits its applications in real world. Thus, it is an essential task to develop efficient and effective techniques to handle various depth degradations. In this paper, we propose a color guided depth…

Cited by 0SourceScholar
2023

CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model

CVPR 2023poster

Supervised crowd counting relies heavily on costly manual labeling, which is difficult and expensive, especially in dense scenes. To alleviate the problem, we propose a novel unsupervised framework for crowd counting, named CrowdCLIP. The core idea is built on two observations: 1) the recent contras…

2023

Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSA

EMNLP 2023long main

Large language models (e.g., GPT-4) are uniquely capable of producing highly rated text simplification, yet current human evaluation methods fail to provide a clear understanding of systems' specific strengths and weaknesses. To address this limitation, we introduce SALSA, an edit-based human annota…

Cited by 0SourceScholar
2023

Efficient Multi-Task and Transfer Reinforcement Learning With Parameter-Compositional Framework

RA-L 2023

In this work, we investigate the potential of improving multi-task training and also leveraging it for transferring in the reinforcement learning setting. We identify several challenges towards this goal and propose a transferring approach with a parameter-compositional formulation. We investigate w

Cited by 13SourceScholar
2023

Frustratingly Easy Label Projection for Cross-lingual Transfer

ACL 2023findings

Translating training data into many languages has emerged as a practical solution for improving cross-lingual transfer. For tasks that involve span-level annotations, such as information extraction or question answering, an additional label projection step is required to map annotated spans onto the…

2023

G2CNN: Geometric Prior Based GCNN for Single-View 3D Reconstruction with Loop Subdivision

ICASSP 2023accepted

Single-view 3D reconstruction is a fundamental operation in computer vision. Although significant progress has been made by learning-based approaches, it remains a challenge that the reconstructed mesh is usually coarse since the geometric prior is ignored. In this paper, we propose a geometric prio…

Cited by 0SourceScholar
2023

Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments

ACL 2023long

We present a human-in-the-loop evaluation framework for fact-checking novel misinformation claims and identifying social media messages that support them. Our approach extracts check-worthy claims, which are aggregated and ranked for review. Stance classifiers are then used to identify tweets suppor…

2023

Integrated Sensing and Full-Duplex Communication: Joint Transceiver Beamforming and Power Allocation

ICASSP 2023accepted

In this paper, we investigate the beamforming design for an integrated sensing and communication (ISAC) system involved full-duplex (FD) communications. Specifically, an FD ISAC base station (BS) performs target detection and communicates with multiple downlink users and uplink users reusing the sam…

Cited by 0SourceScholar
2023

LENS: A Learnable Evaluation Metric for Text Simplification

ACL 2023long

Training learnable metrics using modern language models has recently emerged as a promising method for the automatic evaluation of machine translation. However, existing human evaluation datasets for text simplification have limited annotations that are based on unitary or outdated models, making th…

2023

Multilingual Simplification of Medical Texts

EMNLP 2023long main

Automated text simplification aims to produce simple versions of complex texts. This task is especially useful in the medical domain, where the latest medical findings are typically communicated via complex and technical articles. This creates barriers for laypeople seeking access to up-to-date med…

Cited by 0SourcecodeScholar
2023

Revisiting non-English Text Simplification: A Unified Multilingual Benchmark

ACL 2023long

Recent advancements in high-quality, large-scale English resources have pushed the frontier of English Automatic Text Simplification (ATS) research. However, less work has been done on multilingual text simplification due to the lack of a diverse evaluation benchmark that covers complex-simple sente…

2023

Super-Resolution Information Enhancement for Crowd Counting

ICASSP 2023accepted

Crowd counting is a challenging task due to the heavy occlusions, scales, and density variations. Existing methods handle these challenges effectively while ignoring low-resolution (LR) circumstances. The LR circumstances weaken the counting performance deeply for two crucial reasons: 1) limited det…

Cited by 0SourceScholar
2023

Swarm-LIO: Decentralized Swarm LiDAR-inertial Odometry

ICRA 2023poster

Accurate self and relative state estimation are the critical preconditions for completing swarm tasks, e.g., collaborative autonomous exploration, target tracking, search and rescue. This paper proposes Swarm-LIO: a fully decentralized state estimation method for aerial swarm systems, in which each…

Cited by 38SourceScholar
2023

Swashplateless-Elevon Actuation for a Dual-Rotor Tail-Sitter VTOL UAV

IROS 2023poster

In this paper, we propose a novel swashplateless-elevon actuation (SEA) for dual-rotor tail-sitter vertical takeoff and landing (VTOL) unmanned aerial vehicles (UAVs). In contrast to the conventional elevon actuation (CEA) which controls both pitch and yaw using elevons, the SEA adopts swash-platele…

Cited by 5SourceScholar
2023

Teaching the Pre-trained Model to Generate Simple Texts for Text Simplification

ACL 2023findings

Randomly masking text spans in ordinary texts in the pre-training stage hardly allows models to acquire the ability to generate simple texts. It can hurt the performance of pre-trained models on text simplification tasks. In this paper, we propose a new continued pre-training strategy to teach the p…

2023

Transferable Adversarial Attack for Both Vision Transformers and Convolutional Networks via Momentum Integrated Gradients

ICCV 2023poster

Visual Transformers (ViTs) and Convolutional Neural Networks (CNNs) are the two primary backbone structures extensively used in various vision tasks. Generating transferable adversarial examples for ViTs is difficult due to ViTs' superior robustness, while transferring adversarial examples across Vi…

Cited by 41PDFScholar
2022

Distributionally Robust $Q$-Learning

ICML 2022spotlight

Reinforcement learning (RL) has demonstrated remarkable achievements in simulated environments. However, carrying this success to real environments requires the important attribute of robustness, which the existing RL algorithms often lack as they assume that the future deployment environment is the…

Cited by 64SourcePDFScholar
2022

Efficient and Probabilistic Adaptive Voxel Mapping for Accurate Online LiDAR Odometry

RA-L 2022

This letter proposes an efficient and probabilistic adaptive voxel mapping method for LiDAR odometry. The map is a collection of voxels; each contains one plane feature that enables the probabilistic representation of the environment and accurate registration of a new LiDAR scan. We further analyze

Cited by 152SourcecodeScholar
2022

Extracting a Knowledge Base of COVID-19 Events from Social Media

COLING 2022main

We present a manually annotated corpus of 10,000 tweets containing public reports of five COVID-19 events, including positive and negative tests, deaths, denied access to testing, claimed cures and preventions. We designed slot-filling questions for each event type and annotated a total of 28 fine-g…

2022

FAST-LIVO: Fast and Tightly-coupled Sparse-Direct LiDAR-Inertial-Visual Odometry

IROS 2022poster

To achieve accurate and robust pose estimation in Simultaneous Localization and Mapping (SLAM) task, multisensor fusion is proven to be an effective solution and thus provides great potential in robotic applications. This paper proposes FAST-LIVO, a fast LiDAR-Inertial-Visual Odometry system, which…

Cited by 169SourcecodeScholar
2022

Generative Planning for Temporally Coordinated Exploration in Reinforcement Learning

ICLR 2022spotlight

Standard model-free reinforcement learning algorithms optimize a policy that generates the action to be taken in the current time step in order to maximize expected future return. While flexible, it faces difficulties arising from the inefficient exploration due to its single step nature. In this wo…

2022

Multi-Level Spatial-Temporal Adaptation Network for Motor Imagery Classification

ICASSP 2022accepted

Electroencephalogram (EEG) signals for motor imagery (MI) are easily influenced by the environment and the state of the subject, which exhibit temporal and spatial variance. And this variance is more significant across subjects and sessions, which imposes limitations on the cross-domain MI tasks. To…

Cited by 0SourceScholar
2022

Musicyolo: A Sight-Singing Onset/Offset Detection Framework Based on Object Detection Instead of Spectrum Frames

ICASSP 2022accepted

In this paper, we propose MusicYOLO based on object detection to detect the onset and offset in singing for the first time. The onset of the vocal is not as stable and clear as that of musical instruments, which makes the frame-based onset/offset detection methods often not work well. Compared with…

Cited by 0SourceScholar
2022

PaCo: Parameter-Compositional Multi-task Reinforcement Learning

NeurIPS 2022accept

The purpose of multi-task reinforcement learning (MTRL) is to train a single policy that can be applied to a set of different tasks. Sharing parameters allows us to take advantage of the similarities among tasks. However, the gaps between contents and difficulties of different tasks bring us challen…

2022

Society of Agents: Regret Bounds of Concurrent Thompson Sampling

NeurIPS 2022accept

We consider the concurrent reinforcement learning problem where $n$ agents simultaneously learn to make decisions in the same environment by sharing experience with each other. Existing works in this emerging area have empirically demonstrated that Thompson sampling (TS) based algorithms provide a…

Cited by 5SourcePDFScholar
2022

Stanceosaurus: Classifying Stance Towards Multicultural Misinformation

EMNLP 2022main

We present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi and Arabic annotated with stance towards 250 misinformation claims. As far as we are aware, it is the largest corpus annotated with stance towards misinformation claims. The claims in Stanceosaurus originate from 15 fact-check…

Cited by 18SourcePDFScholar
2022

arXivEdits: Understanding the Human Revision Process in Scientific Writing

EMNLP 2022main

Scientific publications are the primary means to communicate research discoveries, where the writing quality is of crucial importance. However, prior work studying the human editing process in this domain mainly focused on the abstract or introduction sections, resulting in an incomplete picture. In…

2021

An Empowerment-based Solution to Robotic Manipulation Tasks with Sparse Rewards

RSS 2021poster

In order to provide adaptive and user-friendly solutions to robotic manipulation; it is important that the agent can learn to accomplish tasks even if they are only provided with very sparse instruction signals. To address the issues reinforcement learning algorithms face when task rewards are spars…

2021

BiSECT: Learning to Split and Rephrase Sentences with Bitexts

EMNLP 2021main

An important task in NLP applications such as sentence simplification is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary. We introduce a novel dataset and a new model for this ‘split and rephrase’ task. Our BiSECT training data consists of 1…

2021

Controllable Text Simplification with Explicit Paraphrasing

NAACL 2021long

Text Simplification improves the readability of sentences through several rewriting transformations, such as lexical paraphrasing, deletion, and splitting. Current simplification systems are predominantly sequence-to-sequence models that are trained end-to-end to perform all these operations simulta…

2021

Generative Particle Variational Inference via Estimation of Functional Gradients

ICML 2021spotlight

Recently, particle-based variational inference (ParVI) methods have gained interest because they can avoid arbitrary parametric assumptions that are common in variational inference. However, many ParVI approaches do not allow arbitrary sampling from the posterior, and the few that do allow such samp…

Cited by 1SourcePDFScholar
2021

R $2$ LIVE: A Robust, Real-Time, LiDAR-Inertial-Visual Tightly-Coupled State Estimator and Mapping

RA-L 2021

In this letter, we propose a robust, real-time tightly-coupled multi-sensor fusion framework, which fuses measurements from LiDAR, inertial sensor, and visual camera to achieve robust and accurate state estimation. Our proposed framework is composed of two parts: the filter-based odometry and factor

Cited by 124SourcecodeScholar
2021

WIKIBIAS: Detecting Multi-Span Subjective Biases in Language

EMNLP 2021finding

Biases continue to be prevalent in modern text and media, especially subjective bias – a special type of bias that introduces improper attitudes or presents a statement with the presupposition of truth. To tackle the problem of detecting and further mitigating subjective bias, we introduce a manuall…

2020

Multi-hop Reading Comprehension across Documents with Path-based Graph Convolutional Network

IJCAI 2020poster

Multi-hop reading comprehension across multiple documents attracts much attentions recently. In this paper, we propose a novel approach to tackle this multi-hop reading comprehension problem. Inspired by the human reasoning processing, we introduce a path-based graph with reasoning paths which extra…

Cited by 0SourcePDFScholar
2019

CAMEL: A Weakly Supervised Learning Framework for Histopathology Image Segmentation

ICCV 2019accepted

Histopathology image analysis plays a critical role in cancer diagnosis and treatment. To automatically segment the cancerous regions, fully supervised segmentation algorithms require labor-intensive and time-consuming labeling at the pixel level. In this research, we propose CAMEL, a weakly supervi…

2019

UnOS: Unified Unsupervised Optical-Flow and Stereo-Depth Estimation by Watching Videos

CVPR 2019poster

In this paper, we propose UnOS, an unified system for unsupervised optical flow and stereo depth estimation using convolutional neural network (CNN) by taking advantages of their inherent geometrical consistency based on the rigid-scene assumption. UnOS significantly outperforms other state-of-the-a…

Cited by 194PDFScholar
2018

DeLS-3D: Deep Localization and Segmentation With a 3D Semantic Map

CVPR 2018poster

For applications such as augmented reality, autonomous driving, self-localization/camera pose estimation and scene parsing are crucial technologies. In this paper, we propose a unified framework to tackle these two problems simultaneously. The uniqueness of our design is a sensor fusion scheme which…

2018

Guided Feature Transformation (GFT): A Neural Language Grounding Module for Embodied Agents

CoRL 2018

Recently there has been a rising interest in training agents, embodied in virtual environments, to perform language-directed tasks by deep reinforcement learning. In this paper, we propose a simple but effective neural language grounding module for embodied agents that can be trained end to end from

2018

Interactive Grounded Language Acquisition and Generalization in a 2D World

ICLR 2018poster

We build a virtual agent for learning language in a 2D maze-like world. The agent sees images of the surrounding environment, listens to a virtual teacher, and takes actions to receive rewards. It interactively learns the teacher’s language from scratch based on two language use cases: sentence-dire…

2018

LEGO: Learning Edge With Geometry All at Once by Watching Videos

CVPR 2018poster

Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network is attracting significant attention. In this paper, we introduce a “3D as-smooth-as-possible (3D-ASAP)” prior inside the pipeline, which enables joint estimation of edges and 3D scene, yiel…

Cited by 215SourcePDFScholar
2018

Occlusion Aware Unsupervised Learning of Optical Flow

CVPR 2018poster

It has been recently shown that a convolutional neural network can learn optical flow estimation with unsuper- vised learning. However, the performance of the unsuper- vised methods still has a relatively large gap compared to its supervised counterpart. Occlusion and large motion are some of the ma…

Cited by 376SourcePDFScholar
2016

Attention to Scale: Scale-Aware Semantic Image Segmentation

CVPR 2016poster

Incorporating multi-scale features in fully convolutional neural networks (FCNs) has been a key element to achieving state-of-the-art performance on semantic image segmentation. One common way to extract multi-scale features is to feed multiple resized input images to a shared deep network and then…

Cited by 1726PDFScholar
2016

CNN-RNN: A Unified Framework for Multi-Label Image Classification

CVPR 2016oral

While deep convolutional neural networks (CNNs) have shown a great success in single-label image classification, it is important to note that most real world images contain multiple labels, which could correspond to different objects, scenes, actions and attributes in an image. Traditional approache…

Cited by 1717PDFScholar
2016

Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks

CVPR 2016oral

We present an approach that exploits hierarchical Recurrent Neural Networks (RNNs) to tackle the video captioning problem, i.e., generating one or multiple sentences to describe a realistic video. Our hierarchical framework contains a sentence generator and a paragraph generator. The sentence genera…

Cited by 742PDFScholar
2015

Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question

NeurIPS 2015poster

In this paper, we present the mQA model, which is able to answer questions about the content of an image. The answer can be a sentence, a phrase or a single word. Our model contains four components: a Long Short-Term Memory (LSTM) to extract the question representation, a Convolutional Neural Networ…

Cited by 692SourcePDFScholar
2015

Look and Think Twice: Capturing Top-Down Visual Attention With Feedback Convolutional Neural Networks

ICCV 2015poster

While feedforward deep convolutional neural networks (CNNs) have been a great success in computer vision, it is important to remember that the human visual contex contains generally more feedback connections than foward connections. In this paper, we will briefly introduce the background of feedback…

Cited by 530PDFcodeScholar