← Search

Wei Tang

73 accepted papers

2026

Agile Collision Avoidance for Deformable-Tethered Multi-Robot Systems Via Zone-Aware Hierarchical Learning and VLM-Guided Control

ICRA 2026poster

Navigating Linked Multi-Component Robotic Systems (L-MCRS)---robot pairs tethered by passive flexible hoses---through dynamic pedestrian environments is fundamentally harder than rigid multi-robot coordination, as the uncontrollable hose creates a variable-geometry collision footprint spanning 118 p…

Cited by 0Scholar
2026

Artemis: Structured Visual Reasoning for Perception Policy Learning

ICML 2026poster

Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural language. Empirical observations indicate that such purely linguistic intermediate reasoning often reduces performance on perception tasks. We argue that the c…

Cited by 0SourceScholar
2026

Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs

ICLR 2026poster

Diagrams represent a form of visual language that encodes abstract concepts and relationships through structured symbols and their spatial arrangements. Unlike natural images, they are inherently symbolic, and entirely artificial. They thus pose unique challenges for Multimodal Large Language Model…

Cited by 0SourceScholar
2026

RARE: Learn to RAnk and REtrieve for Monocular 3D Object Detection

CVPR 2026

Monocular 3D object detection from a single RGB image remains challenging due to two fundamental challenges: the ill-posed nature of 3D localization, where multiple plausible configurations can correspond to the same 2D observation, and unreliable confidence estimation that fails to reflect true loc

Cited by 0SourcecodeScholar
2026

SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning

ICML 2026poster

Reinforcement learning with verifiable rewards has shown notable effectiveness in enhancing large language models (LLMs) reasoning performance, especially in mathematics tasks. However, such improvements often come with reduced outcome diversity, where the model concentrates probability mass on a na…

Cited by 0SourceScholar
2025

"I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities

ICASSP 2025accepted

Spoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, annotating their Spoken NER data is costly. In this paper, we demonstrate that existing Spoken NER systems perform poorly w…

Cited by 0SourceScholar
2025

An Exceptional Dataset For Rare Pancreatic Tumor Segmentation

ICASSP 2025accepted

Pancreatic NEuroendocrine Tumors (pNETs) are very rare endocrine neoplasms that account for less than 5% of all pancreatic malignancies, with an incidence of only 1–1.5 cases per 100,000. Early detection of pNETs is critical for improving patient survival, but the rarity of pNETs makes segmenting th…

Cited by 0SourceScholar
2025

Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning

ICASSP 2025accepted

Existing video captioning methods merely provide shallow or simplistic representations of object behaviors, resulting in superficial and ambiguous descriptions. However, object behavior is dynamic and complex. To comprehensively capture the essence of object behavior, we propose a dynamic action sem…

Cited by 0SourceScholar
2025

Detection and Geographic Localization of Natural Objects in the Wild: A Case Study on Palms

IJCAI 2025

Palms are ecologically and economically indicators of tropical forest health, biodiversity, and human impact that support local economies and global forest product supply chains. While palm detection in plantations is well-studied, efforts to map naturally occurring palms in dense forests remain lim

2025

Disentangling Language and Culture for Evaluating Multilingual Large Language Models

ACL 2025long

This paper introduces a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of LLMs. By decomposing the evaluation along the dimensions of linguistic medium and cultural context, this framework enables a nuanced analysis of LLMs’ ability to process questions within both…

Cited by 0SourcePDFScholar
2025

DoCIA: An Online Document-Level Context Incorporation Agent for Speech Translation

ACL 2025finding

Document-level context is crucial for handling discourse challenges in text-to-text document-level machine translation (MT). Despite the increased discourse challenges introduced by noise from automatic speech recognition (ASR), the integration of document-level context in speech translation (ST) re…

2025

FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation

CVPR 2025poster

Robotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational efficiency due to the iterative denoising process during in…

Cited by 0SourcePDFScholar
2025

GAP-RL: Grasps as Points for RL Towards Dynamic Object Grasping

RA-L 2025

Dynamic grasping of moving objects in complex, continuous motion scenarios remains challenging. Reinforcement Learning (RL) has been applied in various robotic manipulation tasks, benefiting from its closed-loop property. However, existing RL-based methods do not fully explore the potential for enha

Cited by 7SourceScholar
2025

Investigating Numerical Translation with Large Language Models

ICASSP 2025accepted

The inaccurate translation of numbers can lead to significant security issues, ranging from financial setbacks to medical inaccuracies. While large language models (LLMs) have made significant advancements in machine translation, their capacity for translating numbers has not been thoroughly explore…

Cited by 0SourceScholar
2025

Large Language Model Should Understand Pinyin for Chinese ASR Error Correction

ICASSP 2025accepted

Large language models (LLMs) can enhance automatic speech recognition (ASR) systems through generative error correction (GEC). In this paper, we propose Pinyin-enhanced GEC (PY-GEC), which leverages Pinyin—the phonetic representation of Mandarin Chinese—as supplementary information to improve Chines…

Cited by 0SourceScholar
2025

Learning Partonomic 3D Reconstruction from Image Collections

CVPR 2025poster

Reconstructing the 3D shape of an object from a single-view image is a fundamental task in computer vision. Recent advances in differentiable rendering have enabled 3D reconstruction from image collections using only 2D annotations. However, these methods mainly focus on whole-object reconstruction…

2025

PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic Manipulation

CVPR 2025poster

Robotic manipulation based on visual observations and natural language instructions is a long-standing challenge in robotics. Yet prevailing approaches model action distribution by adopting explicit or implicit representations, which often struggle to achieve a trade-off between accuracy and efficie…

Cited by 0SourcePDFScholar
2025

Region-Centric 6-Dof Grasp Detection: A Data-Efficient Solution for Cluttered Scenes

IROS 2025

Robotic grasping, serving as the cornerstone of robot manipulation, is fundamental for embodied intelligence. Manipulation in challenging scenarios demands grasp detection algorithms with higher efficiency and generalizability. However, for general 6-Dof grasp detection, most data-driven methods dir

Cited by 0SourceScholar
2025

Rethinking Smoothness for Fast and Adaptable Entity Alignment Decoding

NAACL 2025findings

Entity alignment (EA) is crucial for integrating multi-source knowledge graphs (KGs), aiming to identify equivalent entities across different graphs. However, most existing EA decoding methods rely on both entity and relation embeddings, limiting their generalizability and efficiency, especially in…

2025

RevPv8: Reverse Information for Road Space and Lane Line Segmentation in Highway Surveillance Scene

ICASSP 2025accepted

Road space and lane line segmentation are vital tasks in intelligent traffic surveillance, yet they receive less attention compared to similar tasks in autonomous driving. In highway surveillance, segmenting occluded road areas and lane lines is crucial for comprehensive perception, though it adds c…

Cited by 0SourceScholar
2025

SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World Models

NeurIPS 2025poster

World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions due to disrupted spatial structure, inefficient decoding, and…

Cited by 0SourceScholar
2025

The Rise of Parameter Specialization for Knowledge Storage in Large Language Models

NeurIPS 2025poster

Over time, a growing wave of large language models from various series has been introduced to the community. Researchers are striving to maximize the performance of language models with constrained parameter sizes. However, from a microscopic perspective, there has been limited research on how to be…

Cited by 0SourceScholar
2025

Towards Precise Embodied Dialogue Localization via Causality Guided Diffusion

CVPR 2025poster

Embodied localization based on vision and natural language dialogues presents a persistent challenge in embodied intelligence. Existing methods often approach this task as an image translation problem, leveraging encoder-decoder architectures to predict heatmaps. However, these methods frequently ex…

Cited by 0SourcePDFScholar
2024

A + B: A General Generator-Reader Framework for Optimizing LLMs to Unleash Synergy Potential

ACL 2024findings

Retrieval-Augmented Generation (RAG) is an effective solution to supplement necessary knowledge to large language models (LLMs). Targeting its bottleneck of retriever performance, “generate-then-read” pipeline is proposed to replace the retrieval stage with generation from the LLM itself. Although p…

2024

Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language Models

NeurIPS 2024poster

Large language models (LLMs) have achieved impressive performance across various natural language benchmarks, prompting a continual need to curate more difficult datasets for larger LLMs, which is costly and time-consuming. In this paper, we propose to automate dataset updating and provide systemati…

2024

CFD-enabled Approach for Optimizing CPG Control Network for Underwater Soft Robotic Fish

IROS 2024poster

Central Pattern Generators (CPG) nonlinear oscillation network is being increasingly used in the control of multi-joint collaborative robots. The motion attitude of robots can be effectively adjusted by tuning parameters of the CPG neural network. However, the mapping from CPG parameters to motion a…

Cited by 0SourceScholar
2024

Decentralized Communication-Maintained Coordination for Multi-Robot Exploration: Achieving Connectivity and Adaptability

IROS 2024poster

The realm of multi-robot autonomous exploration tasks underscores the critical role of communication in coordinating group activities. This paper introduces an innovative decentralized multi-robot exploration algorithm, meticulously crafted to ensure unbroken communication within robotic groups, a c…

Cited by 1SourceScholar
2024

Exploiting Conjugate Label Information for Multi-Instance Partial-Label Learning

IJCAI 2024poster

Multi-instance partial-label learning (MIPL) addresses scenarios where each training sample is represented as a multi-instance bag associated with a candidate label set containing one true label and several false positives. Existing MIPL algorithms have primarily focused on mapping multi-instance ba…

2024

Intrinsic Robustness of Prophet Inequality to Strategic Reward Signaling

NeurIPS 2024poster

Prophet inequality concerns a basic optimal stopping problem and states that simple threshold stopping policies --- i.e., accepting the first reward larger than a certain threshold --- can achieve tight $\frac{1}{2}$-approximation to the optimal prophet value. Motivated by its economic applications,…

Cited by 0SourcePDFScholar
2024

Invisible Backdoor Attack against 3D Point Cloud Classifier in Graph Spectral Domain

AAAI 2024technical

3D point cloud has been wildly used in security crucial domains, such as self-driving and 3D face recognition. Backdoor attack is a serious threat that usually destroy Deep Neural Networks (DNN) in the training stage. Though a few 3D backdoor attacks are designed to achieve guaranteed attack efficie…

2024

LLMs-as-Instructors: Learning from Errors Toward Automating Model Improvement

EMNLP 2024finding

This paper introduces the innovative “LLMs-as-Instructors” framework, which leverages the advanced Large Language Models (LLMs) to autonomously enhance the training of smaller target models. Inspired by the theory of “Learning from Errors”, this framework employs an instructor LLM to meticulously an…

Cited by 11SourcePDFScholar
2024

Learning Anomalies with Normality Prior for Unsupervised Video Anomaly Detection

ECCV 2024poster

"Unsupervised video anomaly detection (UVAD) aims to detect abnormal events in videos without any annotations. It remains challenging because anomalies are rare, diverse, and usually not well-defined. Existing UVAD methods are purely data-driven and perform unsupervised learning by identifying vario…

2024

Multi-Instance Partial-Label Learning with Margin Adjustment

NeurIPS 2024poster

Multi-instance partial-label learning (MIPL) is an emerging learning framework where each training sample is represented as a multi-instance bag associated with a candidate label set. Existing MIPL algorithms often overlook the margins for attention scores and predicted probabilities, leading to sub…

2024

Performative Prediction with Bandit Feedback: Learning through Reparameterization

ICML 2024poster

Performative prediction, as introduced by Perdomo et al., is a framework for studying social prediction in which the data distribution itself changes in response to the deployment of a model. Existing work in this field usually hinges on three assumptions that are easily violated in practice: that t…

2024

Portable Planner for Enhancing Ground Robots Exploration Performance in Unstructured Environments

RA-L 2024

In this letter, we present a novel portable strategy for the autonomous exploration of highly unstructured three-dimensional environments using ground robots. The proposed planner leverages elevation mapping to estimate traversability, enabling efficient environment mapping while conserving computat

Cited by 5SourceScholar
2024

QRMeM: Unleash the Length Limitation through Question then Reflection Memory Mechanism

EMNLP 2024finding

While LLMs have made notable advancements in natural language processing, they continue to struggle with processing extensive text. Memory mechanisms offer a flexible solution for managing long contexts, utilizing techniques such as compression, summarization, and structuring to facilitate nuanced a…

2024

Region-aware Grasp Framework with Normalized Grasp Space for Efficient 6-DoF Grasping

CoRL 2024poster

A series of region-based methods succeed in extracting regional features and enhancing grasp detection quality. However, faced with a cluttered scene with potential collision, the definition of the grasp-relevant region stays inconsistent. In this paper, we propose Normalized Grasp Space (NGS) from…

Cited by 1SourceScholar
2024

Towards Generalizable Multi-Object Tracking

CVPR 2024poster

Multi-Object Tracking (MOT) encompasses various tracking scenarios each characterized by unique traits. Effective trackers should demonstrate a high degree of generalizability across diverse scenarios. However existing trackers struggle to accommodate all aspects or necessitate hypothesis and experi…

2023

Deeply Coupled Cross-Modal Prompt Learning

ACL 2023findings

Recent advancements in multimodal foundation models (e.g., CLIP) have excelled in zero-shot generalization. Prompt tuning involved in the knowledge transfer from foundation models to downstream tasks has gained significant attention recently. Existing prompt-tuning methods in cross-modal learning, h…

2023

Disambiguated Attention Embedding for Multi-Instance Partial-Label Learning

NeurIPS 2023poster

In many real-world tasks, the concerned objects can be represented as a multi-instance bag associated with a candidate label set, which consists of one ground-truth label and several false positive labels. Multi-instance partial-label learning (MIPL) is a learning paradigm to deal with such tasks an…

Cited by 13SourcePDFScholar
2023

Efficient Heatmap-Guided 6-Dof Grasp Detection in Cluttered Scenes

RA-L 2023

Fast and robust object grasping in clutter is a crucial component of robotics. Most current works resort to the whole observed point cloud for 6-Dof grasp generation, ignoring the guidance information excavated from global semantics, thus limiting high-quality grasp generation and real-time performa

Cited by 55SourcecodeScholar
2023

Encoding Human Behavior in Information Design through Deep Learning

NeurIPS 2023poster

We initiate the study of $\textit{behavioral information design}$ through deep learning. In information design, a $\textit{sender}$ aims to persuade a $\textit{receiver}$ to take certain actions by strategically revealing information. We address scenarios in which the receiver might exhibit differen…

Cited by 4SourcePDFScholar
2023

Learning Sparse Alignments via Optimal Transport for Cross-Domain Fake News Detection

ICASSP 2023accepted

Fake news causes cognitive misperception among the audience and spreads panic to the public. It is crucial to detect fake news and prevent its spread early. Previous methods focus on excavating distinguishable features from news contents in a single domain with deep models, which are difficult to ge…

Cited by 0SourceScholar
2023

Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action Localization

ICCV 2023poster

Semi-Supervised Temporal Action Localization (SS-TAL) aims to improve the generalization ability of action detectors with large-scale unlabeled videos. Albeit the recent advancement, one of the major challenges still remains: noisy pseudo labels hinder efficient learning on abundant unlabeled videos…

Cited by 10PDFcodeScholar
2023

MotionTrack: Learning Robust Short-Term and Long-Term Motions for Multi-Object Tracking

CVPR 2023poster

The main challenge of Multi-Object Tracking (MOT) lies in maintaining a continuous trajectory for each target. Existing methods often learn reliable motion patterns to match the same target between adjacent frames and discriminative appearance features to re-identify the lost targets after a long pe…

2023

Multi-Stream Representation Learning for Pedestrian Trajectory Prediction

AAAI 2023technical

Forecasting the future trajectory of pedestrians is an important task in computer vision with a range of applications, from security cameras to autonomous driving. It is very challenging because pedestrians not only move individually across time but also interact spatially, and the spatial and tempo…

2022

Learning Disentangled Classification and Localization Representations for Temporal Action Localization

AAAI 2022technical

A common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that thi…

Cited by 20SourcePDFScholar
2022

Learning To Refactor Action and Co-Occurrence Features for Temporal Action Localization

CVPR 2022poster

The main challenge of Temporal Action Localization is to retrieve subtle human actions from various co-occurring ingredients, e.g., context and background, in an untrimmed video. While prior approaches have achieved substantial progress through devising advanced action detectors, they still suffer f…

Cited by 57PDFScholar
2022

Path Planning of Multi-Robot Systems With Boolean Specifications Based on Simulated Annealing

RA-L 2022

In this letter, we address the path planning of multi-robot systems (i.e., a team of identical mobile robots) with a global high-level specification that is given as a Boolean formula over some regions of the environment. The task is composed of logical requirements on the trajectories and the final

Cited by 36SourceScholar
2022

UniRel: Unified Representation and Interaction for Joint Relational Triple Extraction

EMNLP 2022main

Relational triple extraction is challenging for its difficulty in capturing rich correlations between entities and relations. Existing works suffer from 1) heterogeneous representations of entities and relations, and 2) heterogeneous modeling of entity-entity interactions and entity-relation interac…

2021

ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action Localization

AAAI 2021technical

The object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foregroun…

Cited by 86SourcePDFScholar
2021

An Optical Spatial Localization System for Tracking Unmanned Aerial Vehicles Using a Single Dynamic Vision Sensor

IROS 2021poster

This paper reports a novel optical localization method, including both the hardware design and algorithm design, to track mobile Unmanned Aerial Vehicles (UAVs). The method relies on a circle-shaped blinking LED marker installed on the UAV and uses a single Dynamic Vision Sensing (DVS) camera to sen…

Cited by 15SourceScholar
2021

Enriching Local and Global Contexts for Temporal Action Localization

ICCV 2021poster

Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by…

Cited by 148PDFcodeScholar
2021

Meta Pairwise Relationship Distillation for Unsupervised Person Re-Identification

ICCV 2021poster

Unsupervised person re-identification (Re-ID) remains challenging due to the lack of ground-truth labels. Existing methods often rely on estimated pseudo labels via iterative clustering and classification, and they are unfortunately highly susceptible to performance penalties incurred by the inaccur…

Cited by 51PDFcodeScholar
2021

Unlimited Neighborhood Interaction for Heterogeneous Trajectory Prediction

ICCV 2021poster

Understanding complex social interactions among agents is a key challenge for trajectory prediction. Most existing methods consider the interactions between pairwise traffic agents or in a local area, while the nature of interactions is unlimited, involving an uncertain number of agents and non-loca…

Cited by 35PDFcodeScholar
2021

Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and Context

AAAI 2021technical

Weakly-supervised Temporal Action Localization (WS-TAL) methods learn to localize temporal starts and ends of action instances in a video under only video-level supervision. Existing WS-TAL methods rely on deep features learned for action recognition. However, due to the mismatch between classificat…

Cited by 32SourcePDFScholar
2020

A Comprehensive Study of Weight Sharing in Graph Networks for 3D Human Pose Estimation

ECCV 2020poster

Graph convolutional networks (GCNs) have been applied to 3D human pose estimation (HPE) from 2D body joint detections and have shown encouraging performance. One limitation of the vanilla graph convolution is that it models the relationships between neighboring nodes via a shared weight matrix. This…

Cited by 180SourcePDFScholar
2020

Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization

ECCV 2020poster

Weakly-supervised Temporal Action Localization (W-TAL) aims to classify and localize all action instances in an untrimmed video under only video-level supervision. However, without frame-level annotations, it is challenging for W-TAL methods to identify false positive action proposals and generate a…

2017

Efficient Online Local Metric Adaptation via Negative Samples for Person Re-Identification

ICCV 2017poster

Many existing person re-identification (PRID) methods typically attempt to train a faithful global metric offline to cover the enormous visual appearance variations, so as to directly use it online on various probes for identity matching. However, their need for a huge set of positive training pairs…

Cited by 98PDFScholar