← Search

Xiang Zhang

102 accepted papers

2026

DexCtrl: Sim-To-Real Dexterity with Adaptive Controller Learning

ICRA 2026poster

Dexterous manipulation has advanced rapidly, with policies now capable of performing complex, contact-rich tasks in simulation. However, transferring these policies from simulation to real world remains a significant challenge. A key obstacle is the mismatch in low-level controller dynamics, where s…

Cited by 0Scholar
2026

Event-Guided Super-Resolving Blurry Image via Asymmetric Integral Driven Consistency

AAAI 2026technical

Super-Resolution from a Blurry low-resolution image (SRB) constitutes a severely ill-posed inverse problem. Current learning-based SRB approaches primarily rely on synthetic, well-labeled paired datasets to regularize solution spaces, yet they exhibit limited generalizability in practical applicatio

Cited by 0SourcePDFScholar
2026

Guardians of the Hair: Rescuing Soft Boundaries in Depth, Stereo, and Novel Views

CVPR 2026

Soft boundaries, like thin hairs, are commonly observed in natural and computer-generated imagery, but they remain challenging for 3D vision due to the ambiguous mixing of foreground and background cues. This paper introduces Guardians of the Hair (HairGuard), a framework designed to recover fine-gr

Cited by 0SourceScholar
2026

How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use

ICLR 2026poster

As Large Language Models (LLMs) are increasingly applied in high-stakes domains, their ability to reason strategically under uncertainty becomes critical. Poker provides a rigorous testbed, requiring not only strong actions but also principled, game-theoretic reasoning. In this paper, we conduct a s…

Cited by 0SourceScholar
2026

PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction

CVPR 2026

We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed distance fields and post-hoc layout optimization, PixARMesh jointly predicts object layout and geometry within a unifie

Cited by 0SourcecodeScholar
2026

Repurposing Foundation Model for Generalizable Medical Time Series Classification

ICLR 2026poster

Medical time series (MedTS) classification suffers from poor generalizability in real-world deployment due to inter- and intra-dataset heterogeneity, such as varying numbers of channels, signal lengths, task definitions, and patient characteristics. % implicit patient characteristics, variable chann…

Cited by 0SourcecodeScholar
2026

Robust Localization for Autonomous Vehicles in Highway Scenes

ICRA 2026poster

Localization for autonomous vehicles on highways remains under-explored compared to urban roads, and state-of-the-art methods for urban scenes degrade when directly applied to highways. We identify key challenges including environment change under information homogeneity, heavy occlusion, degraded G…

2026

Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI

ICLR 2026oral

Understanding how the brain encodes visual information is a central challenge in neuroscience and machine learning. A promising approach is to reconstruct visual stimuli—essentially images—from functional Magnetic Resonance Imaging (fMRI) signals. This involves two stages: transforming fMRI signals…

Cited by 0SourceScholar
2026

UrbanFeel:A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective

ICLR 2026poster

Urban development impacts over half of the global population, making human-centered understanding of its structural and perceptual changes essential for smart city planning. While Multimodal Large Language Models (MLLMs) have shown remarkable capabilities across various domains, existing benchmarks…

Cited by 0SourcecodeScholar
2026

When to Think, When to Speak: Learning Disclosure Policies for Large Language Model Reasoning

ICML 2026poster

Standard Chain-of-Thought (CoT) reasoning trades reliability for responsiveness: in a single user-visible token stream, more deliberation delays meaningful output, imposing a ``silence tax.'' We introduce \emph{Side-by-Side (SxS) Interleaved Reasoning}, a training framework that makes \emph{disclosu…

Cited by 0SourceScholar
2026

Wi-CBR: Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior Recognition

AAAI 2026technical

The challenge in WiFi-based cross-domain Behavior Recognition lies in the significant interference of domain-specific signals on gesture variation. However, previous methods alleviate this interference by mapping the phase from multiple domains into a common feature space. If the Doppler Frequency S

Cited by 0SourcePDFScholar
2026

You Only Need One Stage: Novel-View Synthesis from a Single Blind Face Image

AAAI 2026technical

We propose a novel one-stage method, NVB-Face, for generating consistent Novel-View images directly from a single Blind Face image. Existing approaches to novel-view synthesis for objects or faces typically require a high-resolution RGB image as input. When dealing with degraded images, the conventi

Cited by 0SourcePDFScholar
2025

Accelerating Inverse Kinematic Solutions for a Cable-Driven Soft Robotic Manipulator via Physics-Informed Neural Network

IROS 2025

Cable-driven soft manipulators, with inherent compliance and hyper-redundancy, offer significant advantages in unstructured environments but present formidable challenges in modeling of inverse kinematics due to nonlinear deformations and underactuation. In this paper, building on a modified forward

Cited by 0SourceScholar
2025

Bidirectional Representations Augmented Autoregressive Biological Sequence Generation: Application in De Novo Peptide Sequencing

NeurIPS 2025poster

Autoregressive (AR) models, common in sequence generation, are limited in many biological tasks like de novo peptide sequencing and protein modeling by their unidirectional nature, failing to capture crucial global bidirectional token dependencies. Non-Autoregressive (NAR) models offer holistic, bid…

Cited by 0SourcecodeScholar
2025

Capsizing-Guided Trajectory Optimization for Autonomous Navigation with Rough Terrain

IROS 2025

It is a challenging task for ground robots to autonomously navigate in harsh environments due to the presence of non-trivial obstacles and uneven terrain. This requires trajectory planning that balances safety and efficiency. The primary challenge is to generate a feasible trajectory that prevents r

Cited by 0SourceScholar
2025

Credible and Detailed 3D Face Reconstruction in Large Pose

ICASSP 2025accepted

The existing monocular methods face huge challenges in reconstructing credible details of non-visible areas in large pose images. Due to the fact that facial details are lost in non-visible areas of large pose images, existing methods lose basis when reconstructing details, resulting in unreliable r…

Cited by 0SourceScholar
2025

Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing

ICML 2025poster

Peptide sequencing—the process of identifying amino acid sequences from mass spectrometry data—is a fundamental task in proteomics. Non-Autoregressive Transformers (NATs) have proven highly effective for this task, outperforming traditional methods. Unlike autoregressive models, which generate token…

2025

DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion

ICCV 2025poster

We propose DepR, a depth-guided single-view scene reconstruction framework that integrates instance-level diffusion within a compositional paradigm. Instead of reconstructing the entire scene holistically, DepR generates individual objects and subsequently composes them into a coherent 3D layout. Un…

Cited by 0SourcePDFScholar
2025

DualEqui: A Dual-Space Hierarchical Equivariant Network for Large Biomolecules

NeurIPS 2025poster

Geometric graph neural networks (GNNs) that respect E(3) symmetries have achieved strong performance on small molecule modeling, but they face scalability and expressiveness challenges when applied to large biomolecules such as RNA and proteins. These systems require models that can simultaneously c…

Cited by 0SourceScholar
2025

Efficient Personalized Adaptation for Physiological Signal Foundation Model

ICML 2025poster

Time series analysis is crucial across various fields like energy, environment, transportation, finance and health. Deep learning has significantly advanced this field, particularly, the Time Series Foundation Model (TSFM) excels in multiple domains due to extensive pre-training. In this work, we fo…

Cited by 0SourcePDFScholar
2025

Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors

ICLR 2025poster

The advent of large language models (LLMs) has revolutionized the field of text generation, producing outputs that closely mimic human-like writing. Although academic and industrial institutions have developed detectors to prevent the malicious usage of LLM-generated texts, other research has doubt…

Cited by 1SourcePDFScholar
2025

In-Pipe Navigation Development Environment and a Smooth Path Planning Method on Pipeline Surface

ICRA 2025

Autonomous in-pipe inspection robots can automatically navigate through complex pipeline networks and detect potential risks from corrosion and defects, demonstrating great potential for replacing costly manual inspections. However, there is no publicly available simulation environment where researc

Cited by 6SourceScholar
2025

Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task Learning

AAAI 2025technical

Generative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often…

Cited by 0SourcePDFScholar
2025

L3A: Label-Augmented Analytic Adaptation for Multi-Label Class Incremental Learning

ICML 2025poster

Class-incremental learning (CIL) enables models to learn new classes continually without forgetting previously acquired knowledge. Multi-label CIL (MLCIL) extends CIL to a real-world scenario where each sample may belong to multiple classes, introducing several challenges: label absence, which leads…

2025

LLM-Driven Hierarchical Planning: Long-horizon Task Allocation for Multi-Robot Systems in Cross-Regional Environments

IROS 2025

Long-horizon composite task planning for multi-robot systems in cross-regional complex scenarios faces dual challenges: spatial-semantic comprehension of natural language described tasks and collaborative optimization of subtask al-location. To address these challenges, this paper proposes a progres

Cited by 1SourceScholar
2025

Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers

ICCV 2025poster

We present Lay-Your-Scene (shorthand LayouSyn), a novel text-to-layout generation pipeline for natural scenes. Prior scene layout generation methods are either closed-vocabulary or use proprietary large language models for open-vocabulary generation, limiting their modeling capabilities and broader…

2025

MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention

CoRL 2025poster

Aligning robot behavior with human preferences is crucial for deploying embodied AI agents in human-centered environments. A promising solution is interactive imitation learning from human intervention, where a human expert observes the policy's execution and provides interventions as feedback. Howe…

Cited by 0SourceScholar
2025

NesTools: A Dataset for Evaluating Nested Tool Learning Abilities of Large Language Models

COLING 2025main

Large language models (LLMs) combined with tool learning have gained impressive results in real-world applications. During tool learning, LLMs may call multiple tools in nested orders, where the latter tool call may take the former response as its input parameters. However, current research on the n…

2025

OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps

NeurIPS 2025poster

Despite steady progress in layout-to-image generation, current methods still struggle with layouts containing significant overlap between bounding boxes. We identify two primary challenges: (1) large overlapping regions and (2) overlapping instances with minimal semantic distinction. Through both qu…

Cited by 0SourcecodeScholar
2025

Retrieval is Not Enough: Enhancing RAG through Test-Time Critique and Optimization

NeurIPS 2025poster

Retrieval-augmented generation (RAG) has become a widely adopted paradigm for enabling knowledge-grounded large language models (LLMs). However, standard RAG pipelines often fail to ensure that model reasoning remains consistent with the evidence retrieved, leading to factual inconsistencies or unsu…

Cited by 0SourcecodeScholar
2025

STAR: A Benchmark for Astronomical Star Fields Super-Resolution

NeurIPS 2025spotlight

Super-resolution (SR) advances astronomical imaging by enabling cost-effective high-resolution capture, crucial for detecting faraway celestial objects and precise structural analysis. However, existing datasets for astronomical SR (ASR) exhibit three critical limitations: flux inconsistency, object…

Cited by 0SourcecodeScholar
2025

Stephanie: Step-by-Step Dialogues for Mimicking Human Interactions in Social Conversations

NAACL 2025findings

In the rapidly evolving field of natural language processing, dialogue systems primarily employ a single-step dialogue paradigm. Although this paradigm is commonly adopted, it lacks the depth and fluidity of human interactions and does not appear natural. We introduce a novel **Step**-by-Step Dialog…

Cited by 2SourcePDFScholar
2025

Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling

ACL 2025long

Expressive zero-shot voice conversion (VC) is a critical and challenging task that aims to transform the source timbre into an arbitrary unseen speaker while preserving the original content and expressive qualities. Despite recent progress in zero-shot VC, there remains considerable potential for im…

2025

Transtreaming: Adaptive Delay-aware Transformer for Real-time Streaming Perception

AAAI 2025technical

Real-time object detection is critical for the decision-making process for many real-world applications, such as collision avoidance and path planning in autonomous driving. This work presents an innovative real-time streaming perception method, Transtreaming, which addresses the challenge of real-t…

2025

UGM2N: An Unsupervised and Generalizable Mesh Movement Network via M-Uniform Loss

NeurIPS 2025poster

Partial differential equations (PDEs) form the mathematical foundation for modeling physical systems in science and engineering, where numerical solutions demand rigorous accuracy-efficiency tradeoffs. Mesh movement techniques address this challenge by dynamically relocating mesh nodes to rapidly-va…

Cited by 0SourceScholar
2025

Universal Biological Sequence Reranking for Improved De Novo Peptide Sequencing

ICML 2025poster

De novo peptide sequencing is a critical task in proteomics. However, the performance of current deep learning-based methods is limited by the inherent complexity of mass spectrometry data and the heterogeneous distribution of noise signals, leading to data-specific biases. We present RankNovo, the…

2025

VertexRegen: Mesh Generation with Continuous Level of Detail

ICCV 2025poster

We introduce VertexRegen, a novel mesh generation framework that enables generation at a continuous level of detail. Existing autoregressive methods generate meshes in a partial-to-complete manner and thus intermediate steps of generation represent incomplete structures. VertexRegen takes inspiratio…

Cited by 0SourcePDFScholar
2025

Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMs

ACL 2025long

Despite the remarkable successes of Large Language Models (LLMs), the underlying Transformer architecture has inherent limitations in handling complex reasoning tasks. Chain-of-Thought (CoT) prompting has emerged as a practical workaround, but most CoT-based methods rely on a single generic prompt l…

2025

YOLO-Count: Differentiable Object Counting for Text-to-Image Generation

ICCV 2025poster

We propose YOLO-Count, a differentiable open-vocabulary object counting model that tackles both general counting challenges and enables precise quantity control for text-to-image (T2I) generation. A core contribution is the 'cardinality' map, a novel regression target that accounts for variations in…

Cited by 0SourcePDFScholar
2024

Bayesian Diffusion Models for 3D Shape Reconstruction

CVPR 2024poster

We present Bayesian Diffusion Models (BDM) a prediction algorithm that performs effective Bayesian inference by tightly coupling the top-down (prior) information with the bottom-up (data-driven) procedure via joint diffusion processes. We demonstrate the application of BDM on the 3D shape reconstruc…

2024

BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation

NeurIPS 2024poster

By training over large-scale datasets, zero-shot monocular depth estimation (MDE) methods show robust performance in the wild but often suffer from insufficient detail. Although recent diffusion-based MDE approaches exhibit a superior ability to extract details, they struggle in geometrically comple…

Cited by 7SourcePDFScholar
2024

Bridging the Sim-to-Real Gap with Dynamic Compliance Tuning for Industrial Insertion

ICRA 2024poster

Contact-rich manipulation tasks often exhibit a large sim-to-real gap. For instance, industrial assembly tasks frequently involve tight insertions where the clearance is less than 0.1 mm and can even be negative when dealing with a deformable receptacle. This narrow clearance leads to complex contac…

Cited by 10SourcecodeScholar
2024

Contact-Rich SE(3)-Equivariant Robot Manipulation Task Learning via Geometric Impedance Control

RA-L 2024

This letter presents a differential geometric control approach that leverages <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">SE(3)</i> group invariance and equivariance to increase transferability in learning robot manipulation tasks that involve in

Cited by 23SourcecodeScholar
2024

ContraNovo: A Contrastive Learning Approach to Enhance De Novo Peptide Sequencing

AAAI 2024technical

De novo peptide sequencing from mass spectrometry (MS) data is a critical task in proteomics research. Traditional de novo algorithms have encountered a bottleneck in accuracy due to the inherent complexity of proteomics data. While deep learning-based methods have shown progress, they reduce the pr…

2024

EiffHDR: An Efficient Network for Multi-Exposure High Dynamic Range Imaging

ICASSP 2024accepted

While recent progress in Multi-exposure HDR imaging is promising, the growing complexity of state-of-the-art (SOTA) methods poses challenges for their analysis and comparison. In this paper, we analyze the motivations and approaches behind previous SOTA works and introduce EiffHDR, an efficient Mult…

Cited by 0SourceScholar
2024

Harnessing with Twisting: Single-Arm Deformable Linear Object Manipulation for Industrial Harnessing Task

IROS 2024poster

Wire-harnessing tasks pose great challenges to be automated by the robot due to the complex dynamics and unpredictable behavior of the deformable wire. Traditional methods, often reliant on dual-robot arms or tactile sensing, face limitations in adaptability, cost, and scalability. This paper introd…

Cited by 0SourceScholar
2024

History-Aware Planning for Risk-free Autonomous Navigation on Unknown Uneven Terrain

ICRA 2024poster

It is challenging for the mobile robot to achieve autonomous and mapless navigation in the unknown environment with uneven terrain. In this study, we present a layered and systematic pipeline. At the local level, we maintain a tree structure that is dynamically extended with the navigation. This str…

Cited by 1SourcecodeScholar
2024

In-Hand Following of Deformable Linear Objects Using Dexterous Fingers with Tactile Sensing

IROS 2024

Most research on deformable linear object (DLO) manipulation assumes rigid grasping. However, beyond rigid grasping and re-grasping, in-hand following is also an essential skill that humans use to dexterously manipulate DLOs, which requires continuously changing the grasp point by in-hand sliding wh

Cited by 13SourceScholar
2024

Medformer: A Multi-Granularity Patching Transformer for Medical Time-Series Classification

NeurIPS 2024poster

Medical time series (MedTS) data, such as Electroencephalography (EEG) and Electrocardiography (ECG), play a crucial role in healthcare, such as diagnosing brain and heart diseases. Existing methods for MedTS classification primarily rely on handcrafted biomarkers extraction and CNN-based models, wi…

2024

Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning

CoRL 2024poster

The increasing complexity of tasks in robotics demands efficient strategies for multitask and continual learning. Traditional models typically rely on a universal policy for all tasks, facing challenges such as high computational costs and catastrophic forgetting when learning new tasks. To address…

Cited by 17SourceScholar
2024

Universal Prompt Optimizer for Safe Text-to-Image Generation

NAACL 2024long

Text-to-Image (T2I) models have shown great performance in generating images based on textual prompts. However, these models are vulnerable to unsafe input to generate unsafe content like sexual, harassment and illegal-activity images. Existing studies based on image checker, model fine-tuning and e…

2024

Unsupervised Learning of Neural Semantic Mappings with the Hungarian Algorithm for Compositional Semantics

ICASSP 2024accepted

Neural semantic parsing maps natural languages (NL) to equivalent formal semantics which are compositional and deduce the sentence meanings by composing smaller parts. To learn a well-defined semantics, semantic parsers must recognize small parts, which are semantic mappings between NL and semantic…

Cited by 0SourceScholar
2023

Certifiably Robust Graph Contrastive Learning

NeurIPS 2023poster

Graph Contrastive Learning (GCL) has emerged as a popular unsupervised graph representation learning method. However, it has been shown that GCL is vulnerable to adversarial attacks on both the graph structure and node attributes. Although empirical approaches have been proposed to enhance the robus…

2023

Contrast Everything: A Hierarchical Contrastive Framework for Medical Time-Series

NeurIPS 2023poster

Contrastive representation learning is crucial in medical time series analysis as it alleviates dependency on labor-intensive, domain-specific, and scarce expert annotations. However, existing contrastive learning methods primarily focus on one single data level, which fails to fully exploit the int…

2023

Don’t Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMs

EMNLP 2023long main

Large language models (LLMs) have demonstrated exceptional natural language understanding abilities, and have excelled in a variety of natural language processing (NLP) tasks. Despite the fact that most LLMs are trained predominantly on English, multiple studies have demonstrated their capabilities…

Cited by 0SourceScholar
2023

Efficient Sim-to-real Transfer of Contact-Rich Manipulation Skills with Online Admittance Residual Learning

CoRL 2023poster

Learning contact-rich manipulation skills is essential. Such skills require the robots to interact with the environment with feasible manipulation trajectories and suitable compliance control parameters to enable safe and stable contact. However, learning these skills is challenging due to data inef…

Cited by 24SourceScholar
2023

FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction

ACL 2023long

The recent advent of self-supervised pre-training techniques has led to a surge in the use of multimodal learning in form document understanding. However, existing approaches that extend the mask language modeling to other modalities require careful multi-task tuning, complex reconstruction target d…

2023

GC-Flow: A Graph-Based Flow Network for Effective Clustering

ICML 2023poster

Graph convolutional networks (GCNs) are *discriminative models* that directly model the class posterior $p(y|\mathbf{x})$ for semi-supervised classification of graph data. While being effective, as a representation learning approach, the node representations extracted from a GCN often miss useful in…

2023

Generalizing Event-Based Motion Deblurring in Real-World Scenarios

ICCV 2023poster

Event-based motion deblurring has shown promising results by exploiting low-latency events. However, current approaches are limited in their practical usage, as they assume the same spatial resolution of inputs and specific blurriness distributions. This work addresses these limitations and aims to…

Cited by 29PDFcodeScholar
2023

Interpret ESG Rating’s Impact on the Industrial Chain Using Graph Neural Networks

IJCAI 2023poster

We conduct a quantitative analysis of the development of the industry chain from the environmental, social, and governance (ESG) perspective, which is an overall measure of sustainability. Factors that may impact the performance of the industrial chain have been studied in the literature, such as g…

Cited by 9SourcePDFScholar
2023

Knowledge-Spreader: Learning Semi-Supervised Facial Action Dynamics by Consistifying Knowledge Granularity

ICCV 2023poster

Recent studies on dynamic facial action unit (AU) detection have extensively relied on dense annotations. However, manual annotations are difficult, time-consuming, and costly. The canonical semi-supervised learning (SSL) methods ignore the consistency, extensibility, and adaptability of structural…

Cited by 10PDFScholar
2023

Pic2Word: Mapping Pictures to Words for Zero-Shot Composed Image Retrieval

CVPR 2023poster

In Composed Image Retrieval (CIR), a user combines a query image with text to describe their intended target. Existing methods rely on supervised learning of CIR models using labeled triplets consisting of the query image, text specification, and the target image. Labeling such triplets is expensive…

2023

Prefix Conditioning Unifies Language and Label Supervision

CVPR 2023poster

Pretraining visual models on web-scale image-caption datasets has recently emerged as a powerful alternative to traditional pretraining on image classification data. Image-caption datasets are more "open-domain", containing broader scene types and vocabulary words, and result in models that have str…

Cited by 15SourcePDFScholar
2023

ReactioNet: Learning High-Order Facial Behavior from Universal Stimulus-Reaction by Dyadic Relation Reasoning

ICCV 2023poster

Diverse visual stimuli can evoke various human affective states, which are usually manifested in an individual's muscular actions and facial expressions. In lab-controlled emotion datasets, such a critical component (i.e., stimulus) was commonly designed in a limited way, making researchers incapabl…

Cited by 6PDFScholar
2023

Surface-Sampling Based Objective Quality Assessment Metrics for Meshes

ICASSP 2023accepted

In this paper, we prove that it is feasible to perform mesh quality assessment by sampling it into point cloud. We propose a general and efficient surface-sampling based framework that can deal with various types and levels of distortions with less complexity. In this method, the original and distor…

Cited by 0SourceScholar
2023

Time Series Contrastive Learning with Information-Aware Augmentations

AAAI 2023technical

Various contrastive learning approaches have been proposed in recent years and achieve significant empirical success. While effective and prevalent, contrastive learning has been less explored for time series data. A key component of contrastive learning is to select appropriate augmentations imposi…

2023

Weakly-Supervised Text-Driven Contrastive Learning for Facial Behavior Understanding

ICCV 2023poster

Contrastive learning has shown promising potential for learning robust representations by utilizing unlabeled data. However, constructing effective positive-negative pairs for contrastive learning on facial behavior datasets remains challenging. This is because such pairs inevitably encode the s…

Cited by 13PDFScholar
2022

A Character-Level Length-Control Algorithm for Non-Autoregressive Sentence Summarization

NeurIPS 2022accept

Sentence summarization aims at compressing a long sentence into a short one that keeps the main gist, and has extensive real-world applications such as headline generation. In previous work, researchers have developed various approaches to improve the ROUGE score, which is the main evaluation metric…

2022

AutoGCL: Automated Graph Contrastive Learning via Learnable View Generators

AAAI 2022technical

Contrastive learning has been widely applied to graph representation learning, where the view generators play a vital role in generating effective contrastive samples. Most of the existing contrastive learning methods employ pre-defined view generation methods, e.g., node drop or edge perturbation,…

2022

Class Guided Channel Weighting Network for Fine-Grained Semantic Segmentation

AAAI 2022technical

Deep learning has achieved promising performance on semantic segmentation, but few works focus on semantic segmentation at the fine-grained level. Fine-grained semantic segmentation requires recognizing and distinguishing hundreds of sub-categories. Due to the high similarity of different sub-catego…

Cited by 2SourcePDFScholar
2022

EGCN: An Ensemble-based Learning Framework for Exploring Effective Skeleton-based Rehabilitation Exercise Assessment

IJCAI 2022poster

Recently, some skeleton-based physical therapy systems have been attempted to automatically evaluate the correctness or quality of an exercise performed by rehabilitation subjects. However, in terms of algorithms and evaluation criteria, the task remains not fully explored regarding making full use…

2022

Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language Models

NeurIPS 2022accept

Traditional knowledge distillation (KD) methods manually design student architectures to compress large models given pre-specified computational cost. This requires several trials to find viable students, and repeating the process with change in computational budget. We use Neural Architecture Searc…

2022

Graph-Guided Network for Irregularly Sampled Multivariate Time Series

ICLR 2022poster

In many domains, including healthcare, biology, and climate science, time series are irregularly sampled with varying time intervals between successive readouts and different subsets of variables (sensors) observed at different time points. Here, we introduce RAINDROP, a graph neural network that em…

2022

Improving HowNet-Based Chinese Word Sense Disambiguation with Translations

EMNLP 2022finding

Word sense disambiguation (WSD) is the task of identifying the intended sense of a word in context. While prior work on unsupervised WSD has leveraged lexical knowledge bases, such as WordNet and BabelNet, these resources have proven to be less effective for Chinese. Instead, the most widely used le…

Cited by 10SourcePDFScholar
2022

Learning Insertion Primitives with Discrete-Continuous Hybrid Action Space for Robotic Assembly Tasks

ICRA 2022poster

This paper introduces a discrete-continuous action space to learn insertion primitives for robotic assembly tasks. Primitives are sequences of elementary actions with certain exit conditions, such as “pushing down the peg until contact”. Since the primitive is an abstraction of robot control command…

Cited by 51SourceScholar
2022

Offline-Online Learning of Deformation Model for Cable Manipulation With Graph Neural Networks

RA-L 2022

Manipulating deformable linear objects by robots has a wide range of applications, e.g., manufacturing and medical surgery. To complete such tasks, an accurate dynamics model for predicting the deformation is critical for robust control. In this letter, we deal with this challenge by proposing a hyb

Cited by 69SourceScholar
2022

Self-Supervised Contrastive Pre-Training For Time Series via Time-Frequency Consistency

NeurIPS 2022accept

Pre-training on time series poses a unique challenge due to the potential mismatch between pre-training and target domains, such as shifts in temporal dynamics, fast-evolving trends, and long-range and short-cyclic effects, which can lead to poor downstream performance. While domain adaptation metho…

2021

Controlling Neural Networks with Rule Representations

NeurIPS 2021poster

We propose a novel training method that integrates rules into deep learning, in a way the strengths of the rules are controllable at inference. Deep Neural Networks with Controllable Rule Representations (DeepCTRL) incorporates a rule encoder into the model coupled with a rule-based objective, enabl…

Cited by 50SourcePDFScholar
2021

Event-Based Synthetic Aperture Imaging With a Hybrid Network

CVPR 2021poster

Synthetic aperture imaging (SAI) is able to achieve the see through effect by blurring out the off-focus foreground occlusions and reconstructing the in-focus occluded targets from multi-view images. However, very dense occlusions and extreme lighting conditions may bring significant disturbances to…

Cited by 39PDFcodeScholar
2021

InfoGCL: Information-Aware Graph Contrastive Learning

NeurIPS 2021poster

Various graph contrastive learning models have been proposed to improve the performance of tasks on graph datasets in recent years. While effective and prevalent, these models are usually carefully customized. In particular, despite all recent work create two contrastive views, they differ in a vari…

Cited by 235SourcePDFScholar
2021

Learning Variable Impedance Control via Inverse Reinforcement Learning for Force-Related Tasks

RA-L 2021

Many manipulation tasks require robots to interact with unknown environments. In such applications, the ability to adapt the impedance according to different task phases and environment constraints is crucial for safety and performance. Although many approaches based on deep reinforcement learning (

Cited by 111SourceScholar
2021

Online Learning of Unknown Dynamics for Model-Based Controllers in Legged Locomotion

RA-L 2021

The performance of a model-based controller can severely suffer when its model inaccurately represents the real world dynamics. We propose to learn a time-varying, locally linear residual model along the robot's current trajectory, to compensate for the prediction errors of the controller's model. S

Cited by 65SourceScholar
2021

Transformer-Style Relational Reasoning with Dynamic Memory Updating for Temporal Network Modeling

AAAI 2021technical

Network modeling aims to learn the latent representations of nodes such that the representations preserve both network structures and node attribute information. This problem is fundamental due to its prevalence in numerous domains. However, existing approaches either target the static networks or s…

Cited by 24SourcePDFScholar
2020

Parameterized Explainer for Graph Neural Network

NeurIPS 2020poster

Despite recent progress in Graph Neural Networks (GNNs), explaining predictions made by GNNs remains a challenging open problem. The leading method mainly addresses the local explanations (i.e., important subgraph structure and node features) to interpret why a GNN model makes the prediction for a s…

2018

A Novel Learnable Dictionary Encoding Layer for End-to-End Language Identification

ICASSP 2018accepted

A novel learnable dictionary encoding layer is proposed in this paper for end-to-end language identification. It is inline with the conventional GMM i-vector approach both theoretically and practically. We imitate the mechanism of traditional GMM training and Supervector encoding procedure on the to…

Cited by 79SourceScholar
2015

An under-actuated manipulation controller based on Workspace Analysis and Gaussian Processes

IROS 2015poster

The kinematic modelling has been applied to many controllers of under-actuated manipulators. Most of these studies assume that the control process is conducted within the workspace. However, as such a kinematic model cannot describe the situations when the stable grasping is violated in the real env…

Cited by 7SourceScholar