← Search

Tao Wang

139 accepted papers

2026

Bayesian Decomposition and Semantic Completion for Few-shot Semantic Segmentation

CVPR 2026

Few-shot Semantic Segmentation (FSS) aims to segment objects of novel categories given only a handful of labeled examples. However, existing methods often rely on complex category-specific modeling, resulting in high computational cost and limited generalization under low-data regimes. To address th

Cited by 0SourceScholar
2026

CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution

AAAI 2026technical

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual cues often conflict, requiring models to perform structured

Cited by 0SourcePDFScholar
2026

D-FUSEr: Diverse Failure, Unified Success via Error-Distribution Shaping in LLM Reasoning

ICML 2026poster

Test-time scaling methods such as majority vote aggregation and iterative refinement (e.g., self-reflection or multi-agent inference) improve reasoning performance by leveraging multiple solution samples. However, their efficacy depends not only on raw performance, but critically on the distribution…

Cited by 0SourceScholar
2026

DIPP: A Diffusion-Based Potential Planner for Synergistic Navigation and Mapping

ICRA 2026poster

Object-Goal Navigation (ObjectNav) requires an embodied agent to search for and reach a target object category in previously unseen environments using only onboard egocentric observations, which is a fundamental capability for long-horizon autonomous robots. Current Object-Goal Navigation methods ty…

Cited by 0Scholar
2026

DiMA: Distinguishing Resident and Tourist Preferences via Multi-Modal LLM Alignment for Out-of-Town Cross-Domain Recommendation

AAAI 2026technical

Out-of-Town (OOT) recommendation aims to provide personalized suggestions for users in unfamiliar cities. However, OOT recommendation faces two fundamental challenges: the difficulty of reasoning across modalities, as preference signals in disparate formats such as images and text are hard to compar

Cited by 0SourcePDFScholar
2026

Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design

CVPR 2026

Following the success of Group Relative Policy Optimization (GRPO) in foundation LLMs, an increasing number of works have sought to adapt GRPO to Visual Large Language Models (VLLMs) for visual perception tasks (e.g., detection and segmentation). However, much of this line of research rests on a lon

Cited by 0SourcecodeScholar
2026

GS^2: Graph-based Spatial Distribution Optimization for Compact 3D Gaussian Splatting

CVPR 2026

3D Gaussian Splatting (3DGS) has demonstrated breakthrough performance in novel view synthesis and real-time rendering. Nevertheless, its practicality is constrained by the high memory cost due to a huge number of Gaussian points. Many pruning-based 3DGS variants have been proposed for memory saving

Cited by 0SourcecodeScholar
2026

Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-Experts

AAAI 2026technical

Automated scoring of written constructed responses typically relies on separate models per task, straining computational resources, storage, and maintenance in real-world education settings. We propose UniMoE-Guided, a knowledge-distilled multi-task Mixture-of-Experts (MoE) approach that transfers e

Cited by 0SourcePDFScholar
2026

Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge Editing

AAAI 2026technical

Multimodal Knowledge Editing (MKE) extends traditional knowledge editing to settings involving both textual and visual modalities. However, existing MKE benchmarks primarily assess final answer correctness, neglecting the quality of intermediate reasoning and robustness to visually rephrased inputs.

Cited by 0SourcePDFScholar
2026

LLM-Orchestrated Diagnose–Plan–Treat for Mixed-Degradation CT Reconstruction

IJCAI 2026

Clinical Computed Tomography (CT) reconstruction often faces mixed degradations, where quantum noise, streak artifacts, and geometric distortions co-occur with various compositions and severities. Recently, all-in-one frameworks have outperformed traditional single-task models through degradation-sp

Cited by 0Scholar
2026

MCOO-SLAM: A Multi-Camera Omnidirectional Object SLAM System

RA-L 2026

Object-level SLAM offers structured and semantically meaningful environment representations, making it more interpretable and suitable for high-level robotic tasks. However, most existing approaches rely on RGB-D sensors or monocular views, which suffer from narrow fields of view, occlusion sensitiv

Cited by 2SourceScholar
2026

MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science

ICLR 2026oral

We introduce MedAgentGym, a scalable and interactive training environment designed to enhance coding-based biomedical reasoning capabilities in large language model (LLM) agents. MedAgentGym comprises 72,413 task instances across 129 categories derived from 12 authentic real-world biomedical scenari…

Cited by 0SourcecodeScholar
2026

RN-D: Discretized Categorical Actors with Regularized Networks for On-Policy Reinforcement Learning

ICML 2026poster

On-policy deep reinforcement learning remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradients are noisy and policy updates must be conservative. In this paper, w…

Cited by 0SourceScholar
2026

Statistical Early Stopping for Reasoning Models

ICML 2026poster

While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries. We introduce statistically principled early stopping methods that monitor uncertainty s…

Cited by 0SourceScholar
2026

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

ICML 2026poster

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction…

Cited by 0SourceScholar
2025

A Crab-Inspired Soft Gripper with Single-Finger Dexterous Grasping Capabilities

IROS 2025

Soft grippers conform to the shape and surface properties of the objects to be grasped, effectively avoiding damage to soft and fragile items. Despite the variety of existing soft gripper designs, their structures lack sufficient flexibility for effectively grasping slender objects or operating in n

Cited by 0SourceScholar
2025

A Hubness Perspective on Representation Learning for Graph-Based Multi-View Clustering

CVPR 2025poster

Recent graph-based multi-view clustering (GMVC) methods typically encode view features into high-dimensional spaces and construct graphs based on distance similarity. However, the high dimensionality of the embeddings often leads to the hubness problem, where a few points repeatedly appear in the ne…

2025

ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation

NeurIPS 2025poster

While humans effortlessly draw visual objects and shapes by adaptively allocating attention based on their complexity, existing multimodal large language models (MLLMs) remain constrained by rigid token representations. Bridging this gap, we propose ALTo, an adaptive length tokenizer for autoregress…

Cited by 0SourcecodeScholar
2025

Association-Focused Path Aggregation for Graph Fraud Detection

NeurIPS 2025poster

Fraudulent activities have caused substantial negative social impacts and are exhibiting emerging characteristics such as intelligence and industrialization, posing challenges of high-order interactions, intricate dependencies, and the sparse yet concealed nature of fraudulent entities. Existing gra…

Cited by 0SourcecodeScholar
2025

Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction

IJCAI 2025

Multimodal Information Extraction (MIE) has gained attention for extracting structured information from multimedia sources. Traditional methods tackle MIE tasks separately, missing opportunities to share knowledge across tasks. Recent approaches unify these tasks into a generation problem using inst

2025

DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech

ICASSP 2025accepted

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the…

Cited by 0SourceScholar
2025

Enhancing the Flexibility of a Quadruped Robot with a 2-DOF Active Spine Using Nonlinear Model Predictive Control

IROS 2025

For quadrupeds, a flexible spine allows them to traverse space and make quick turns. From the perspective of mechanical design in quadruped robots, an active spine with 2 degrees of freedom (2-DOF) can achieve dynamic posture adjustment similar to biological organisms which allows for pitch and yaw

Cited by 0SourceScholar
2025

Foundations of Top-$k$ Decoding for Language Models

NeurIPS 2025poster

Top-$k$ decoding is a widely used method for sampling from LLMs: at each token, only the largest $k$ next-token-probabilities are kept, and the next token is sampled after re-normalizing them to sum to unity. Top-$k$ and other sampling methods are motivated by the intuition that true next-token dist…

Cited by 0SourceScholar
2025

GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization

ACL 2025long

Recent advances in large language models have highlighted the critical need for precise control over model outputs through predefined constraints. While existing methods attempt to achieve this through either direct instruction-response synthesis or preferential response optimization, they often str…

2025

Gaussian Herding across Pens: An Optimal Transport Perspective on Global Gaussian Reduction for 3DGS

NeurIPS 2025spotlight

3D Gaussian Splatting (3DGS) has emerged as a powerful technique for radiance field rendering, but it typically requires millions of redundant Gaussian primitives, overwhelming memory and rendering budgets. Existing compaction approaches address this by pruning Gaussians based on heuristic importanc…

Cited by 0SourceScholar
2025

HeRo: A State Machine-Based, Fault-Tolerant Framework for Heterogeneous Multi-Robot Collaboration

ICRA 2025

Heterogeneous robots can work together to accomplish a variety of complex tasks and have shown great potential in many fields. There are many efforts to make robot task orchestration more efficient. However, current methods still have some limitations, including the lack of a high-level abstraction

Cited by 0SourceScholar
2025

HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

ICCV 2025poster

The remarkable performance of large multimodal models (LMMs) has attracted significant interest from the image segmentation community.To align with the next-token-prediction paradigm, current LMM-driven segmentation methods either use object boundary points to represent masks or introduce special se…

2025

IMM-MOT: A Novel 3D Multi-object Tracking Framework with Interacting Multiple Model Filter

IROS 2025

3D Multi-Object Tracking (MOT) provides the trajectories of surrounding objects, assisting robots or vehicles in smarter path planning and obstacle avoidance. Existing 3D MOT methods based on the Tracking-by-Detection framework typically use a single motion model to track an object throughout its en

Cited by 1SourcecodeScholar
2025

LLFA: Fusing Global Illumination and Local Priors for Low-Light Face Image Enhancement with Adaptor

ICASSP 2025accepted

Low-light image enhancement problem has been widely studied. However, most existing methods do not perform well on low-light face images due to no specific facial characteristic considerations. We first create large-scale low-light face datasets with synthesized and real-world images to address the…

Cited by 0SourceScholar
2025

LLM-GAN: Constructing Generative Adversarial Network Through Large Language Models for Explainable Fake News Detection

ICASSP 2025accepted

Explainable fake news detection predicts the authenticity of news items with annotated explanations. Today, Large Language Models (LLMs) are known for their powerful natural language understanding and explanation generation abilities. However, using LLMs for explainable fake news detection remains t…

Cited by 0SourceScholar
2025

MOERL: When Mixture-of-Experts Meet Reinforcement Learning for Adverse Weather Image Restoration

ICCV 2025poster

Adverse weather conditions, such as rain, snow, and haze, introduce complex degradations that present substantial challenges for effective image restoration. Existing all-in-one models often rely on fixed network structures, limiting their ability to adapt to the varying characteristics of different…

Cited by 0SourcePDFScholar
2025

MVINS: A Magnetism&vision Aided Inertial Navigation System for Autonomous Underwater Vehicles

RA-L 2025

We present a robust underwater navigation system that integrates magnetic, visual, and inertial measurements from commercial off-the-shelf sensors. Visual Inertial Navigation Systems (VINS) face challenges when used for Autonomous Underwater Vehicle (AUV) localization in perceptually degraded enviro

Cited by 3SourceScholar
2025

MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR Diffusion

ICCV 2025poster

Physically-based rendering (PBR) has become a cornerstone in modern computer graphics, enabling realistic material representation and lighting interactions in 3D scenes. In this paper, we present MaterialMVP, a novel end-to-end model for generating PBR textures from 3D meshes and image prompts, addr…

Cited by 0SourcePDFScholar
2025

Open-Det: An Efficient Learning Framework for Open-Ended Detection

ICML 2025poster

Open-Ended object Detection (OED) is a novel and challenging task that detects objects and generates their category names in a free-form manner, without requiring additional vocabularies during inference. However, the existing OED models, such as GenerateU, require large-scale datasets for training,…

2025

QCRD: Quality-guided Contrastive Rationale Distillation for Large Language Models

EMNLP 2025

The deployment of large language models (LLMs) faces considerable challenges concerning resource constraints and inference efficiency. Recent research has increasingly focused on smaller, task-specific models enhanced by distilling knowledge from LLMs. However, prior studies have often overlooked th

Cited by 0SourcePDFScholar
2025

SALS: Sparse Attention in Latent Space for KV Cache Compression

NeurIPS 2025poster

Large Language Models (LLMs) capable of handling extended contexts are in high demand, yet their inference remains challenging due to substantial Key-Value (KV) cache size and high memory bandwidth requirements. Previous research has demonstrated that KV cache exhibits low-rank characteristics withi…

Cited by 0SourceScholar
2025

Seamless Transition Control in Spring-Legged Quadrotors: A Hybrid Dynamics Perspective with Guaranteed Feasibility

IROS 2025

Legged aerial-terrestrial robots have garnered significant research attention in recent years due to their enhanced environmental adaptability through combined aerial and terrestrial locomotion. However, existing passive spring-legged aerial robots exhibit limited motion versatility, demonstrating s

Cited by 0SourceScholar
2025

Segmentation-Guided Sparse Transformer for Under-Display Camera Image Restoration

ICASSP 2025accepted

Under-display Camera is an emerging technology for full-screen display with a camera under the display. However, the current implementation of UDC causes serious image degradation. Incident light required for camera imaging undergoes attenuation and diffraction when passing through the display. Curr…

Cited by 0SourceScholar
2025

StickMotion: Generating 3D Human Motions by Drawing a Stickman

CVPR 2025poster

Text-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text inputs. This paper introduces StickMotion, an efficient diffusion-based network designed for multi-condition scenarios, wh…

2025

UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model

NAACL 2025findings

Significant advancements has recently been achieved in the field of multi-modal large language models (MLLMs), demonstrating their remarkable capabilities in understanding and reasoning across diverse tasks. However, these models are often trained for specific tasks and rely on task-specific input-o…

2025

WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification

ICASSP 2025accepted

Recent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before compression and extract them from reconstructed speech for verification, but face limitations such as separate training pr…

Cited by 0SourceScholar
2024

An Online Automatic Calibration Method for Infrastructure-Based LiDAR-Camera via Cross-modal Object Matching

IROS 2024poster

In indoor environments where the Global Navigation Satellite System (GNSS) isn’t available, the infrastructure-based LiDAR-camera joint array can provide high-precision localization for mobile robots, such as Autonomous Valet Parking (AVP). The primary challenge in employing the infrastructure-based…

Cited by 0SourceScholar
2024

Controlled Decoding from Language Models

ICML 2024poster

KL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective and propose a modular solver for it, called *controlled decoding (CD)*. CD exerts control through a separate *prefix scor…

Cited by 86SourcePDFScholar
2024

Cut, Bond and Play: Volume-Preserved Reprogrammable Soft Pneumatic Actuators

RA-L 2024

Reprogrammable design enables soft actuators to change their performances after fabrication and obtain new deformation modes and functions. The reprogrammable design of soft pneumatic actuators is relatively difficult due to the intrinsic material being chemically inactive. This work proposes a new

Cited by 5SourceScholar
2024

DFA-GNN: Forward Learning of Graph Neural Networks by Direct Feedback Alignment

NeurIPS 2024poster

Graph neural networks (GNNs) are recognized for their strong performance across various applications, with the backpropagation (BP) algorithm playing a central role in the development of most GNN models. However, despite its effectiveness, BP has limitations that challenge its biological plausibilit…

Cited by 1SourcePDFScholar
2024

Development of Negative-Pressure Artificial Muscles With Fiber Constraints and Pre-Stretched Soft Skin

RA-L 2024

Negative-pressure artificial muscles based on internal support and flexible skin provide a way to develop high-performance artificial muscles. However, the random and disordered skin wrinkles formed during the contraction may lead to uncertainty in the actuation behavior, and the flexible but inexte

Cited by 3SourceScholar
2024

Fewer-Token Neural Speech Codec with Time-Invariant Codes

ICASSP 2024accepted

Language model based text-to-speech (TTS) models, like VALL-E, have gained attention for their outstanding in-context learning capability in zero-shot scenarios. Neural speech codec is a critical component of these models, which can convert speech into discrete token representations. However, excess…

Cited by 0SourceScholar
2024

Generated and Pseudo Content guided Prototype Refinement for Few-shot Point Cloud Segmentation

NeurIPS 2024spotlight

Few-shot 3D point cloud semantic segmentation aims to segment query point clouds with only a few annotated support point clouds. Existing prototype-based methods learn prototypes from the 3D support set to guide the segmentation of query point clouds. However, they encounter the challenge of low pro…

Cited by 1SourcePDFScholar
2024

GroundingGPT: Language Enhanced Multi-modal Grounding Model

ACL 2024long

Multi-modal large language models (MLLMs) have demonstrated remarkable performance across various tasks. However, these models often prioritize capturing global information and overlook the importance of perceiving local information. This limitation hinders their ability to effectively understand fi…

2024

Learning Speech Representation from Contrastive Token-Acoustic Pretraining

ICASSP 2024accepted

For fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate representations extracted from speech should serve as a "bridge" between text and acoustic information, containing info…

Cited by 0SourceScholar
2024

Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding

ICASSP 2024accepted

Recently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations and using two sequence-to-sequence tasks to decouple TTS. However, existing methods suffer from three problems: the high-…

Cited by 0SourceScholar
2024

OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models

ECCV 2024poster

"Personalization is an important topic in text-to-image generation, especially the challenging multi-concept personalization. Current multi-concept methods are struggling with identity preservation, occlusion, and the harmony between foreground and background. In this work, we propose OMG, an occlus…

2024

PANORAMIA: Privacy Auditing of Machine Learning Models without Retraining

NeurIPS 2024poster

We present PANORAMIA, a privacy leakage measurement framework for machine learning models that relies on membership inference attacks using generated data as non-members. By relying on generated non-member data, PANORAMIA eliminates the common dependency of privacy measurement tools on in-distributi…

2024

Rethinking the Representation in Federated Unsupervised Learning with Non-IID Data

CVPR 2024poster

Federated learning achieves effective performance in modeling decentralized data. In practice client data are not well-labeled which makes it potential for federated unsupervised learning (FUSL) with non-IID data. However the performance of existing FUSL methods suffers from insufficient representat…

Cited by 17SourcePDFScholar
2024

SynSP: Synergy of Smoothness and Precision in Pose Sequences Refinement

CVPR 2024poster

Predicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily co…

2024

Trend-Aware Supervision: On Learning Invariance for Semi-supervised Facial Action Unit Intensity Estimation

AAAI 2024technical

With the increasing need for facial behavior analysis, semi-supervised AU intensity estimation using only keyframe annotations has emerged as a practical and effective solution to relieve the burden of annotation. However, the lack of annotations makes the spurious correlation problem caused by AU c…

Cited by 0SourcePDFScholar
2024

Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression

ACL 2024long

Key-value (KV) caching is an important technique to accelerate the inference of large language models (LLMs), but incurs significant memory overhead. To compress the size of KV cache, existing methods often compromise precision or require extra data for calibration, limiting their practicality in LL…

2024

VSFormer: Visual-Spatial Fusion Transformer for Correspondence Pruning

AAAI 2024technical

Correspondence pruning aims to find correct matches (inliers) from an initial set of putative correspondences, which is a fundamental task for many applications. The process of finding is challenging, given the varying inlier ratios between scenes/image pairs due to significant visual differences. H…

2024

Visual-Inertial-Wheel Odometry With Wheel-Aided Maximum-a-Posteriori Initialization for Ground Robots

RA-L 2024

In recent years, Visual-Inertial Odometry (VIO) has demonstrated remarkable results using low-cost and complementary sensors. However, these methods often encounter initialization failure and suffer reduced robustness or low trajectory accuracy under challenging scenarios. In this letter, we propose

Cited by 8SourceScholar
2024

Zero-Shot Aerial Object Detection with Visual Description Regularization

AAAI 2024technical

Existing object detection models are mainly trained on large-scale labeled datasets. However, annotating data for novel aerial object classes is expensive since it is time-consuming and may require expert knowledge. Thus, it is desirable to study label-efficient object detection methods on aerial im…

2023

BLEURT Has Universal Translations: An Analysis of Automatic Metrics by Minimum Risk Training

ACL 2023long

Automatic metrics play a crucial role in machine translation. Despite the widespread use of n-gram-based metrics, there has been a recent surge in the development of pre-trained model-based metrics that focus on measuring sentence semantics. However, these neural metrics, while achieving higher corr…

2023

DecomFormer: Decompose Self-Attention Via Fourier Transform for VHR Aerial Image Scene Classification

ICASSP 2023accepted

Very high-resolution (VHR) aerial image scene classification is an essential task for aerial image understanding. Although transformer-based models have demonstrated strong ability in natural image classification, transformer-based methods on VHR aerial image tasks are still lack of concern because…

Cited by 0SourceScholar
2023

Graph Propagation Transformer for Graph Representation Learning

IJCAI 2023poster

This paper presents a novel transformer architecture for graph representation learning. The core insight of our method is to fully consider the information propagation among nodes and edges in a graph when building the attention module in the transformer blocks. Specifically, we propose a new attent…

2023

Improving Speech Translation by Fusing Speech and Text

EMNLP 2023long findings

In speech translation, leveraging multimodal data to improve model performance and address limitations of individual modalities has shown significant effectiveness. In this paper, we harness the complementary strengths of speech and text to improve speech translation. However, speech and text are di…

Cited by 0SourcecodeScholar
2023

Punctuation-level Attack: Single-shot and Single Punctuation Can Fool Text Models

NeurIPS 2023poster

The adversarial attacks have attracted increasing attention in various fields including natural language processing. The current textual attacking models primarily focus on fooling models by adding character-/word-/sentence-level perturbations, ignoring their influence on human perception. In this p…

Cited by 3SourcePDFScholar
2023

Stay In The Middle: A Semi-Supervised Model for CT Metal Artifact Reduction

ICASSP 2023accepted

Metal artifacts degrade CT image’s quality. Recently, some deep learning-based metal artifact reduction (MAR) methods have been developed. Supervised MAR methods don’t perform well in clinical due to the domain gap between simulated and clinical data. Although this problem can be avoided in an unsup…

Cited by 0SourceScholar
2023

Ultra-High-Definition Low-Light Image Enhancement: A Benchmark and Transformer-Based Method

AAAI 2023technical

As the quality of optical sensors improves, there is a need for processing large-scale images. In particular, the ability of devices to capture ultra-high definition (UHD) images and video places new demands on the image processing pipeline. In this paper, we consider the task of low-light image enh…

2022

A Novel Framework Based on Medical Concept Driven Attention for Explainable Medical Code Prediction via External Knowledge

ACL 2022findings

Medical code prediction from clinical notes aims at automatically associating medical codes with the clinical notes. Rare code problem, the medical codes with low occurrences, is prominent in medical code prediction. Recent studies employ deep neural networks and the external knowledge to tackle it.…

2022

ADD 2022: the first Audio Deep Synthesis Detection Challenge

ICASSP 2022accepted

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three t…

Cited by 0SourceScholar
2022

BézierPalm: A Free Lunch for Palmprint Recognition

ECCV 2022poster

"Palmprints are private and stable information for biometric recognition. In the deep learning era, the development of palmprint recognition is limited by the lack of sufficient training data. In this paper, by observing that palmar creases are the key information to deep-learning-based palmprint re…

Cited by 21SourcePDFScholar
2022

Causal Intervention for Subject-Deconfounded Facial Action Unit Recognition

AAAI 2022technical

Subject-invariant facial action unit (AU) recognition remains challenging for the reason that the data distribution varies among subjects. In this paper, we propose a causal inference framework for subject-invariant facial action unit recognition. To illustrate the causal effect existing in AU recog…

Cited by 29SourcePDFScholar
2022

Context-Aware Mask Prediction Network for End-to-End Text-Based Speech Editing

ICASSP 2022accepted

The text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major drawback of current systems is that edited speech often sounds unnatural and it is not obvious how to synthesize records…

Cited by 0SourceScholar
2022

Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action Recognition

AAAI 2022technical

In recent years, self-supervised representation learning for skeleton-based action recognition has been developed with the advance of contrastive learning methods. The existing contrastive learning methods use normal augmentations to construct similar positive samples, which limits the ability to ex…

2022

Design and Modelling of Multi-DOF Manipulator Driven by Hysteresis-Attenuated Pneumatic Artificial Muscles

RA-L 2022

In this article, a multi-DOF manipulator driven by novel hysteresis-attenuated pneumatic artificial muscles (PAMs) is designed, and the multi-DOF manipulator is composed of three manipulator units assembled in cascade, and each manipulator is driven by hysteresis-attenuated PAMs. Compared with the c

Cited by 6SourceScholar
2022

Discrete Listwise Personalized Ranking for Fast Top-N Recommendation with Implicit Feedback

IJCAI 2022poster

We address the efficiency problem of personalized ranking from implicit feedback by hashing users and items with binary codes, so that top-N recommendation can be fast executed in a Hamming space by bit operations. However, current hashing methods for top-N recommendation fail to align their learnin…

2022

FedInv: Byzantine-Robust Federated Learning by Inversing Local Model Updates

AAAI 2022technical

Federated learning (FL) is a privacy-preserving distributed machine learning paradigm that enables multiple clients to collaboratively train statistical models without disclosing raw training data. However, the inaccessible local training data and uninspectable local training process make FL suscept…

Cited by 61SourcePDFScholar
2022

GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

ICML 2022spotlight

Scaling language models with more data, compute and parameters has driven significant progress in natural language processing. For example, thanks to scaling, GPT-3 was able to achieve strong results on in-context learning tasks. However, training these large dense models requires significant amount…

Cited by 765SourcePDFScholar
2022

ITA: Image-Text Alignments for Multi-Modal Named Entity Recognition

NAACL 2022long

Recently, Multi-modal Named Entity Recognition (MNER) has attracted a lot of attention. Most of the work utilizes image information through region-level visual representations obtained from a pretrained object detector and relies on an attention mechanism to model the interactions between image and…

2022

On Mitigating Hard Clusters for Face Clustering

ECCV 2022poster

"Face clustering is a promising way to scale up face recognition systems using large-scale unlabeled face images. It remains challenging to identify small or sparse face image clusters that we call hard clusters, which is caused by the heterogeneity, i.e., high variations in size and sparsity, of th…

2022

PDD-Net: A Precise Defect Detection Network Based on Point Set Representation

ICASSP 2022accepted

Defect detection has been widely studied in computer vision and used in industrial production. However, most existing methods for defect detection mainly suffer three drawbacks: i) Low-contrast problem between defects and background. ii) Large scale changes in defects size. iii) Extreme imbalance pr…

Cited by 0SourceScholar
2022

Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on Transformer

AAAI 2022technical

Occluded person re-identification is a challenging task as human body parts could be occluded by some obstacles (e.g. trees, cars, and pedestrians) in certain scenes. Some existing pose-guided methods solve this problem by aligning body parts according to graph matching, but these graph-based method…

2022

PoseTriplet: Co-Evolving 3D Human Pose Estimation, Imitation, and Hallucination Under Self-Supervision

CVPR 2022oral

Existing self-supervised 3D human pose estimation schemes have largely relied on weak supervisions like consistency loss to guide the learning, which, inevitably, leads to inferior results in real-world scenarios with unseen poses. In this paper, we propose a novel self-supervised approach that allo…

Cited by 57PDFcodeScholar
2022

Powerful Graph Convolutional Networks with Adaptive Propagation Mechanism for Homophily and Heterophily

AAAI 2022technical

Graph Convolutional Networks (GCNs) have been widely applied in various fields due to their significant power on processing graph-structured data. Typical GCN and its variants work under a homophily assumption (i.e., nodes with same class are prone to connect to each other), while ignoring the heter…

Cited by 128SourcePDFScholar
2022

Towards Real-World HDRTV Reconstruction: A Data Synthesis-Based Approach

ECCV 2022poster

"Existing deep learning based HDRTV reconstruction methods assume one kind of tone mapping operators (TMOs) as the degradation procedure to synthesize SDRTV-HDRTV pairs for supervised training. In this paper, we argue that, although traditional TMOs exploit efficient dynamic range compression priors…

2022

Uncertainty-Guided Pixel Contrastive Learning for Semi-Supervised Medical Image Segmentation

IJCAI 2022poster

Recently, contrastive learning has shown great potential in medical image segmentation. Due to the lack of expert annotations, however, it is challenging to apply contrastive learning in semi-supervised scenes. To solve this problem, we propose a novel uncertainty-guided pixel contrastive learning m…

2021

A Unified Encoding of Structures in Transition Systems

EMNLP 2021main

Transition systems usually contain various dynamic structures (e.g., stacks, buffers). An ideal transition-based model should encode these structures completely and efficiently. Previous works relying on templates or neural network structures either only encode partial structure information or suffe…

2021

Autocorrect in the Process of Translation — Multi-task Learning Improves Dialogue Machine Translation

NAACL 2021industry

Automatic translation of dialogue texts is a much needed demand in many real life scenarios. However, the currently existing neural machine translation delivers unsatisfying results. In this paper, we conduct a deep analysis of a dialogue corpus and summarize three major issues on dialogue translati…

2021

Automated Concatenation of Embeddings for Structured Prediction

ACL 2021long

Pretrained contextualized embeddings are powerful word representations for structured prediction tasks. Recent work found that better word representations can be obtained by concatenating different types of embeddings. However, the selection of embeddings to form the best concatenated representation…

2021

Bi-Level Style and Prosody Decoupling Modeling for Personalized End-to-End Speech Synthesis

ICASSP 2021accepted

End-to-end framework can generate high-quality and high-similarity speech in the personalized speech synthesis task. However, the generalization of out-of-domain texts is still a challenging task. Limited target data leads to unacceptable errors and poor prosody and similarity performance of the syn…

Cited by 0SourceScholar
2021

Cross-Modal Representation Learning for Lightweight and Accurate Facial Action Unit Detection

RA-L 2021

In this letter, we focus on designing an effective method for lightweight and accurate facial action unit (AU) detection, which is essential for emotional communication in most human-robot interaction scenarios. AU detection is a delicate and challenging task because the subtle fleeting appearance c

Cited by 8SourceScholar
2021

Deep Reinforcement Learning for Multi-contact Motion Planning of Hexapod Robots

IJCAI 2021poster

Legged locomotion in a complex environment requires careful planning of the footholds of legged robots. In this paper, a novel Deep Reinforcement Learning (DRL) method is proposed to implement multi-contact motion planning for hexapod robots moving on uneven plum-blossom piles. First, the motion of…

Cited by 15SourcePDFScholar
2021

Direct Multi-view Multi-person 3D Pose Estimation

NeurIPS 2021poster

We present Multi-view Pose transformer (MvP) for estimating multi-person 3D poses from multi-view images. Instead of estimating 3D joint locations from costly volumetric representation or reconstructing the per-person 3D pose from multiple detected 2D poses as in previous methods, MvP directly regre…

2021

End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural Networks

ICCV 2021poster

Video instance segmentation is a challenging task that extends image instance segmentation to the video domain. Existing methods either rely only on single-frame information for the detection and segmentation subproblems or handle tracking as a separate post-processing step, which limit their capabi…

Cited by 38PDFcodeScholar
2021

Improving Named Entity Recognition by External Context Retrieving and Cooperative Learning

ACL 2021long

Recent advances in Named Entity Recognition (NER) show that document-level contexts can significantly improve model performance. In many application scenarios, however, such contexts are not available. In this paper, we propose to find external contexts of a sentence by retrieving and selecting a se…

2021

Learning to Navigate in a VUCA Environment: Hierarchical Multi-expert Approach

IROS 2021poster

Despite decades of efforts, robot navigation in a real scenario with volatility, uncertainty, complexity, and ambiguity (VUCA for short), remains a challenging topic. Inspired by the central nervous system (CNS), we propose a hierarchical multi-expert learning framework for autonomous navigation in…

Cited by 9SourceScholar
2021

MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity Representations

EMNLP 2021main

Entity retrieval, which aims at disambiguating mentions to canonical entities from massive KBs, is essential for many tasks in natural language processing. Recent progress in entity retrieval shows that the dual-encoder structure is a powerful and efficient framework to nominate candidates if entiti…

2021

Multi-Scale Separable Network for Ultra-High-Definition Video Deblurring

ICCV 2021poster

Although recent research has witnessed a significant progress on the video deblurring task, these methods struggle to reconcile inference efficiency and visual quality simultaneously, especially on ultra-high-definition (UHD) videos (e.g., 4K resolution). To address the problem, we propose a novel d…

Cited by 36PDFcodeScholar
2021

Multi-View Cross-Lingual Structured Prediction with Minimum Supervision

ACL 2021long

In structured prediction problems, cross-lingual transfer learning is an efficient way to train quality models for low-resource languages, and further improvement can be obtained by learning from multiple source languages. However, not all source models are created equal and some may hurt performanc…

Cited by 7SourcePDFScholar
2021

PnP-DETR: Towards Efficient Visual Analysis With Transformers

ICCV 2021poster

Recently, DETR pioneered the solution of vision tasks with transformers, it directly translates the image feature map into the object detection result. Though effective, translating the full feature map can be costly due to redundant computation on some area like the background. In this work, we enc…

Cited by 115PDFcodeScholar
2021

Prosody and Voice Factorization for Few-Shot Speaker Adaptation in the Challenge M2voc 2021

ICASSP 2021accepted

The paper describes the CASIA speech synthesis system entry for challenge M2VoC 2021. The low similarity and naturalness of synthesized speech remains a challenging problem for speaker adaptation with few resources. Since the end-to-end acoustic model is too complex to interpret, overfitting will oc…

Cited by 0SourceScholar
2021

Real-Time Image Enhancer via Learnable Spatial-Aware 3D Lookup Tables

ICCV 2021poster

Recently, deep learning-based image enhancement algorithms achieved state-of-the-art (SOTA) performance on several publicly available datasets. However, most existing methods fail to meet practical requirements either for visual perception or for computation efficiency, especially for high-resolutio…

Cited by 98PDFScholar
2021

Risk Minimization for Zero-shot Sequence Labeling

ACL 2021long

Zero-shot sequence labeling aims to build a sequence labeler without human-annotated datasets. One straightforward approach is utilizing existing systems (source models) to generate pseudo-labeled datasets and train a target sequence labeler accordingly. However, due to the gap between the source an…

Cited by 3SourcePDFScholar
2021

Secoco: Self-Correcting Encoding for Neural Machine Translation

EMNLP 2021finding

This paper presents Self-correcting Encoding (Secoco), a framework that effectively deals with noisy input for robust neural machine translation by introducing self-correcting predictors. Different from previous robust approaches, Secoco enables NMT to explicitly correct noisy inputs and delete spec…

2021

Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor

ACL 2021long

Knowledge distillation is a critical technique to transfer knowledge between models, typically from a large model (the teacher) to a more fine-grained one (the student). The objective function of knowledge distillation is typically the cross-entropy between the teacher and the student’s output distr…

2021

Tokens-to-Token ViT: Training Vision Transformers From Scratch on ImageNet

ICCV 2021poster

Transformers, which are popular for language modeling, have been explored for solving vision tasks recently, e.g., the Vision Transformer (ViT) for image classification. The ViT model splits each image into a sequence of tokens with fixed length and then applies multiple Transformer layers to model…

Cited by 2573PDFcodeScholar
2021

Ultra-High-Definition Image Dehazing via Multi-Guided Bilateral Learning

CVPR 2021poster

During the last couple of years, convolutional neural networks (CNNs) have achieved significant success in the single image dehazing task. Unfortunately, most existing deep dehazing models have high computational complexity, which hinders their application to high-resolution images, especially for U…

Cited by 248PDFcodeScholar
2021

Ultra-High-Definition Image HDR Reconstruction via Collaborative Bilateral Learning

ICCV 2021poster

Existing single image high dynamic range (HDR) reconstruction attempt to expand the range of luminance. They are not effective to generate plausible textures and colors in the reconstructed results, especially for high-density pixels in ultra-high-definition (UHD) images.To address these problems, w…

Cited by 35PDFScholar
2021

Word Reordering for Zero-shot Cross-lingual Structured Prediction

EMNLP 2021main

Adapting word order from one language to another is a key problem in cross-lingual structured prediction. Current sentence encoders (e.g., RNN, Transformer with position embeddings) are usually word order sensitive. Even with uniform word form representations (MUSE, mBERT), word order discrepancies…

2020

Central Similarity Quantization for Efficient Image and Video Retrieval

CVPR 2020poster

Existing data-dependent hashing methods usually learn hash functions from pairwise or triplet data relationships, which only capture the data similarity locally, and often suffer from low learning efficiency and low collision rate. In this work, we propose a new global similarity metric, termed as c…

Cited by 400PDFcodeScholar
2020

Focusing on Attention: Prosody Transfer and Adaptative Optimization Strategy for Multi-Speaker End-to-End Speech Synthesis

ICASSP 2020accepted

End-to-end speech synthesis can generate high-quality synthetic speech and achieve high similarity scores with low-resource adaptation data. However, the generalization of out-domain texts is still a challenging task. The limited adaptation data leads to unacceptable errors and the poor prosody perf…

Cited by 0SourceScholar
2020

Overcoming Classifier Imbalance for Long-Tail Object Detection With Balanced Group Softmax

CVPR 2020oral

Solving long-tail large vocabulary object detection with deep learning based models is a challenging and demanding task, which is however under-explored. In this work, we provide the first systematic analysis on the underperformance of state-of-the-art models in front of long-tail distribution. We f…

Cited by 351PDFcodeScholar
2020

Revisiting Knowledge Distillation via Label Smoothing Regularization

CVPR 2020oral

Knowledge Distillation (KD) aims to distill the knowledge of a cumbersome teacher model into a lightweight student model. Its success is generally attributed to the privileged information on similarities among categories provided by the teacher model, and in this sense, only strong teacher models ar…

Cited by 719PDFcodeScholar
2020

The Devil is in Classification: A Simple Framework for Long-tail Instance Segmentation

ECCV 2020poster

Most existing object instance detection and segmentation models only work well on fairly balanced benchmarks where per-category training sample numbers are comparable, such as COCO. They tend to suffer performance drop on realistic datasets that are usually long-tailed. This work aims to study and a…

2019

Eagle Shoal: A new designed modular tactile sensing dexterous hand for domestic service robots

ICRA 2019poster

This paper introduces a new designed modular tactile sensing dexterous hand for domestic service robots. This fully-actuated hand consists of 1 palm and 3 fingers, with embedded tactile sensors, motors and control boards. The palm and each finger have 2 degrees of freedom (DOFs). The modular design…

Cited by 22SourceScholar
2018

A Fluid-Filled Tubular Dielectric Elastomer Variable Stiffness Structure Inspired by the Hydrostatic Skeleton Principle

ICRA 2018poster

This work presents a novel variable stiffness structure consisting of a fiber-constrained dielectric elastomer tube filled with insulating oil. The tensile stiffness of the structure can be adjusted by voltages and its initial value can be customized according to the initial pre-stretch of the mater…

Cited by 3SourceScholar
2018

A Fluid-Filled Tubular Dielectric Elastomer Variable Stiffness Structure Inspired by the Hydrostatic Skeleton Principle *Research supported by the National Natural Science Foundation of China (No.51675413)

ICRA 2018

This work presents a novel variable stiffness structure consisting of a fiber-constrained dielectric elastomer tube filled with insulating oil. The tensile stiffness of the structure can be adjusted by voltages and its initial value can be customized according to the initial pre-stretch of the mater

Cited by 15SourceScholar
2018

Constrained Confidence Matching for Planar Object Tracking

ICRA 2018poster

Tracking planar objects has a wide range of applications in robotics. Conventional template tracking algorithms, however, often fail to observe fast object motion or drift significantly after a period of time, due to drastic object appearance change. To address such challenges, we propose a novel co…

Cited by 7SourceScholar
2017

Design and control of an inchworm-inspired soft robot with omega-arching locomotion

ICRA 2017poster

This paper presents an inchworm inspired soft robot composed of the soft body, the front foot as well as the back foot. Compared to the traditional inchworm-type robot consisting of rigid components, the driven mode for the soft robot is more simple. The soft robot inspired by the inchworm has highe…

Cited by 46SourceScholar