← Search

Hao Wu

112 accepted papers

2026

Active Surface-Driven Reconfigurable Gripper: Robust Grasping and Sequential Manipulation of Thin Objects

RSS 2026poster

Robotic grippers face substantial challenges in grasping and manipulating thin objects. Most existing grippers rely on highly precise approach and grasp motions, which limits robustness and reduces applicability. This paper explores thin-object grasping using books as a representative example. Here,…

Cited by 0SourceScholar
2026

AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation

ICML 2026poster

Vision-Language Navigation (VLN) requires agents to follow natural language instructions by grounding them in sequential visual observations over long horizons. Explicit reasoning could enhance temporal consistency and perception–action alignment, but reasoning at fixed steps often leads to suboptim…

Cited by 0SourceScholar
2026

Benchmarking LLMs’ Mathematical Reasoning with Unseen Random Variables Questions

AAAI 2026technical

Recent studies have raised significant concerns regarding the reliability of current mathematical benchmarks, highlighting key limitations such as simplistic design and potential data contamination that undermine evaluation accuracy. Consequently, developing a reliable benchmark that effectively eva

Cited by 0SourcePDFScholar
2026

DeepSenseMoE: Harnessing Power of Time Series Foundation Models for Few-Shot Human Activity Recognition

AAAI 2026technical

Recent advances in Time Series Foundation Models (TSFMs) have fundamentally revolutionized general time series analysis across domains like finance, retail, weather, and power. However, how to unlock the hidden capacity of general-purpose TSFMs for wearable activity recognition still remains largely

Cited by 1SourcePDFScholar
2026

Diverse Human Driving Vehicle Simulation in Background Traffic for Autonomous Driving Tests

AAAI 2026technical

Realistic background traffic is critical to the simulation platforms for autonomous driving (AD) testing. Given that most vehicles in reality are driven by human beings, introducing human driving (HD) vehicles to the background traffic is necessary to be able to discover more problems of the tested

Cited by 0SourcePDFScholar
2026

HiDivDrop: Vision Token Reduction in MLLMs via Late Injection and Differentiable Top-K

ICLR 2026poster

The computational cost of Multimodal Large Language Models (MLLMs), driven by the quadratic complexity of processing vision tokens, remains a significant barrier to their widespread adoption. While progressive vision token pruning is a promising solution, we find that its full potential has been unr…

Cited by 0SourcecodeScholar
2026

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

CVPR 2026

With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still require task-specific fine-tuning and large-scale audiovisual datasets, resulting in high computational costs that hinder accessibility for resource-c

Cited by 0SourcecodeScholar
2026

InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning

ICRA 2026poster

Leveraging pretrained Vision-Language Models (VLMs) to map language instruction and visual observations to raw low-level actions, Vision-Language-Action models (VLAs) hold great promise for achieving general-purpose robotic systems. Despite their advancements, existing VLAs tend to spuriously correl…

2026

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a)…

Cited by 0SourcecodeScholar
2026

NeuralOM: Neural Ocean Model for Subseasonal-to-Seasonal Simulation

AAAI 2026technical

Long-term, high-fidelity simulation of slow-changing physical systems, such as the ocean and climate, presents a fundamental challenge in scientific computing. Traditional autoregressive machine learning models often fail in these tasks as minor errors accumulate and lead to rapid forecast degradati

Cited by 0SourcePDFScholar
2026

PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting

ICML 2026poster

Coupled spatiotemporal forecasting is important for predicting the future evolution of multiple interacting dynamical systems, such as in climate models. However, existing methods are severely constrained by the persistent bottleneck of compounding errors. In coupled systems, errors from each subsys…

Cited by 0SourceScholar
2026

RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

AAAI 2026technical

We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research

Cited by 0SourcePDFScholar
2026

Relaxation-Aware Multimodal Sensing of Soft Gripper Driven by Structure-Perception-Learning

RSS 2026poster

Achieving stable, sustained grasping with soft robotic hands remains a fundamental challenge. Compliance enables safe and adaptive contact, yet the intrinsic viscoelasticity of soft polymers leads to stress relaxation and a continuous decay of grasping force during holding. Inspired by human graspin…

Cited by 0SourceScholar
2026

Role-Level Inductive Bias for Cross-Task Generalization in Multi-Agent Reinforcement Learning

ICML 2026poster

Achieving cross-task generalization remains a critical challenge in Multi-Agent Reinforcement Learning (MARL), fundamentally relying on effective inductive biases. However, existing entity-level biases often overlook collaborative patterns, whereas task-level biases lack sufficient coverage for nove…

Cited by 0SourceScholar
2026

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs

ICML 2026poster

Multimodal Large Language Models (MLLMs) possess intrinsic reasoning and world-knowledge capabilities, yet adapting them for dense retrieval remains challenging. Existing approaches typically rely on invasive parameter updates, such as full fine-tuning and LoRA, which risk disrupting the pre-trained…

Cited by 0SourceScholar
2026

State Proficiency-Based Adaptive Fine-Tuning for Offline-to-Online Reinforcement Learning

AAAI 2026technical

In offline-to-online (O2O) reinforcement learning, achieving efficient performance improvement while maintaining training stability remains a critical challenge for effective fine-tuning. Existing O2O methods usually focus on the balance between policy improvement and policy constraint during online

Cited by 0SourcePDFScholar
2026

S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm Detection

AAAI 2026technical

Multimodal sarcasm detection (MSD) aims to identify sarcasm polarity from diverse modalities (i.e., image–text pairs), a task that has received increasing attention. While significant progress has been made, existing approaches still face two major issues: lack of explainability and weak generalizab

Cited by 0SourcePDFScholar
2026

Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models

CVPR 2026

Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities in Chain-of-Thought (CoT) reasoning. However, existing LVLM reasoning paradigms only begin reasoning after the entire video becomes available, introducing unnecessary latency and diminishing attention to early visual cues

Cited by 0SourceScholar
2026

UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking

CVPR 2026

One-stream Transformer-based trackers achieve advanced performance in visual object tracking suffer from significant computational overhead that hinders real-time deployment. While token pruning offers a path to efficiency, a critical limitation persists: no existing work performs pruning jointly ac

Cited by 0SourcecodeScholar
2026

Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning

ICML 2026poster

We present $\textit{Video-in-the-Loop}$ (ViTL), a two-stage long-video QA framework that preserves a fixed token budget by first $\textit{localizing}$ question-relevant interval(s) with a low-fps skim and then $\textit{answering}$ via span-aware reallocation of visual tokens at higher effective fram…

Cited by 3SourceScholar
2025

A Bio-Inspired Sand-Rolling Robot: Effect of Body Shape on Sand Rolling Performance

ICRA 2025

The capability of effectively moving on complex terrains such as sand and gravel can empower our robots to robustly operate in outdoor environments, and assist with critical tasks such as environment monitoring, search-and-rescue, and supply delivery. Inspired by the Mount Lyell salamander's ability

Cited by 2SourceScholar
2025

Adaptive Fault-Tolerant Control of Wheeled Mobile Robots With Multiple Actuator Faults and Saturation

RA-L 2025

The actuator fault problems under saturation bring significant challenges to the stable and accurate tracking of wheeled mobile robots (WMRs) in industrial applications. This letter proposes a novel adaptive fault-tolerant control (FTC) method for WMR systems simultaneously considering uncertain mul

Cited by 7SourceScholar
2025

Breaking the Discretization Barrier of Continuous Physics Simulation Learning

NeurIPS 2025poster

The modeling of complicated time-evolving physical dynamics from partial observations is a long-standing challenge. Particularly, observations can be sparsely distributed in a seemingly random or unstructured manner, making it difficult to capture highly nonlinear features in a variety of scientific…

Cited by 0SourcecodeScholar
2025

CellVerse: Do Large Language Models Really Understand Cell Biology?

NeurIPS 2025poster

Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis ta…

Cited by 0SourcecodeScholar
2025

CoDe: Communication Delay-Tolerant Multi-Agent Collaboration via Dual Alignment of Intent and Timeliness

AAAI 2025technical

Communication has been widely employed to enhance multi-agent collaboration. Previous research has typically assumed delay-free communication, a strong assumption that is challenging to meet in practice. However, real-world agents suffer from channel delays, receiving messages sent at different time…

Cited by 0SourcePDFScholar
2025

Consistent Sampling and Simulation: Molecular Dynamics with Energy-Based Diffusion Models

NeurIPS 2025poster

In recent years, diffusion models trained on equilibrium molecular distributions have proven effective for sampling biomolecules. Beyond direct sampling, the score of such a model can also be used to derive the forces that act on molecular systems. However, while classical diffusion sampling usually…

Cited by 0SourcecodeScholar
2025

Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting

ICCV 2025poster

Spatiotemporal forecasting tasks, such as traffic flow, combustion dynamics, and weather forecasting, often require complex models that suffer from low training efficiency and high memory consumption. This paper proposes a lightweight framework, Spectral Decoupled Knowledge Distillation, which trans…

2025

From General Relation Patterns to Task-Specific Decision-Making in Continual Multi-Agent Coordination

IJCAI 2025

Continual Multi-Agent Reinforcement Learning (Co-MARL) requires agents to address catastrophic forgetting issues while learning new coordination policies with the dynamics team. In this paper, we delve into the core of Co-MARL, namely Relation Patterns, which refer to agents’ general understanding o

Cited by 0SourcePDFScholar
2025

FunEditor: Achieving Complex Image Edits via Function Aggregation with Diffusion Models

AAAI 2025technical

Diffusion models have demonstrated outstanding performance in generative tasks, making them ideal candidates for image editing. Recent studies highlight their ability to apply desired edits effectively by following textual instructions, yet with two key challenges remaining. First, these models stru…

Cited by 0SourcePDFScholar
2025

Improved Approximations for Hard Graph Problems using Predictions

ICML 2025poster

We design improved approximation algorithms for NP-hard graph problems by incorporating predictions (e.g., learned from past data). Our prediction model builds upon and extends the $\varepsilon$-prediction framework by Cohen-Addad, d'Orsi, Gupta, Lee, and Panigrahi (NeurIPS 2024). We consider an ed…

Cited by 0SourcePDFScholar
2025

Improving Knowledge Base Question Answering via Retrieval Enhancement and Stepwise Reasoning

ICASSP 2025accepted

The large-scale knowledge base question-answering (KBQA) has become increasingly vital across various fields. In the era of large language models (LLMs), leveraging knowledge base retrieval combined with large models for knowledge reasoning has become the mainstream approach for KBQA. However, this…

Cited by 0SourceScholar
2025

Learning Graph Quantized Tokenizers

ICLR 2025poster

Transformers serve as the backbone architectures of Foundational Models, where domain-specific tokenizers allow them to adapt to various domains. Graph Transformers (GTs) have recently emerged as leading models in geometric deep learning, outperforming Graph Neural Networks (GNNs) in various graph l…

2025

Learning-Augmented Frequent Directions

ICLR 2025spotlight

An influential paper of Hsu et al. (ICLR'19) introduced the study of learning-augmented streaming algorithms in the context of frequency estimation. A fundamental problem in the streaming literature, the goal of frequency estimation is to approximate the number of occurrences of items appearing in a…

Cited by 0SourcePDFScholar
2025

MagicCity: Geometry-Aware 3D City Generation from Satellite Imagery with Multi-View Consistency

ICCV 2025poster

Directly generating 3D cities from satellite imagery opens up new possibilities for gaming and mapping services. However, this task remains challenging due to the limited information in satellite views, making it difficult for existing methods to achieve both photorealistic textures and geometric ac…

2025

MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds

EMNLP 2025

The rapid advancement of large language models has intensified public concerns about the potential misuse. Therefore, it is important to build trustworthy AI-generated text detection systems. Existing methods neglect stylistic modeling and mostly rely on static thresholds, which greatly limits the d

2025

One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models

ICCV 2025poster

Vision-Language Pre-training (VLP) models have exhibited unprecedented capability in many applications by taking full advantage of the learned multimodal alignment. However, previous studies have shown they are vulnerable to maliciously crafted adversarial samples. Despite recent success, these atta…

2025

OneForecast: A Universal Framework for Global and Regional Weather Forecasting

ICML 2025poster

Accurate weather forecasts are important for disaster prevention, agricultural planning, etc. Traditional numerical weather prediction (NWP) methods offer physically interpretable high-accuracy predictions but are computationally expensive and fail to fully leverage rapidly growing historical data.…

2025

Open-CK: A Large Multi-Physics Fields Coupling benchmarks in Combustion Kinetics

ICLR 2025poster

In this paper, we use the Fire Dynamics Simulator (FDS) combined with the {\fontfamily{lmtt}\selectfont \textit{supercomputer}} support to create a \textbf{C}ombustion \textbf{K}inetics (CK) dataset for machine learning and scientific research. This dataset captures the development of fires in indus…

2025

Pixel2Feature Attack (P2FA): Rethinking the Perturbed Space to Enhance Adversarial Transferability

ICML 2025poster

Adversarial examples have been shown to deceive Deep Neural Networks (DNNs), raising widespread concerns about this security threat. More seriously, as different DNN models share critical features, feature-level attacks can generate transferable adversarial examples, thereby deceiving black-box mode…

Cited by 0SourcePDFScholar
2025

PriFold: Biological Priors Improve RNA Secondary Structure Predictions

AAAI 2025technical

Predicting RNA secondary structures is crucial for understanding RNA function, designing RNA-based therapeutics, and studying molecular interactions within cells. Existing deep-learning-based methods for RNA secondary structure prediction have mainly focused on local structural properties, often ove…

2025

SCA3D: Enhancing Cross-Modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation

ICRA 2025

The cross-modal 3D retrieval task aims to achieve mutual matching between text descriptions and 3D shapes. This has the potential to enhance the interaction between natural language and the 3D environment, especially within the realms of robotics and embodied artificial intelligence (AI) application

Cited by 8SourcecodeScholar
2025

Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning

NeurIPS 2025poster

Scientific discoveries increasingly rely on complex multimodal reasoning based on information-intensive scientific data and domain-specific expertise. Empowered by expert-level scientific benchmarks, scientific Multimodal Large Language Models (MLLMs) hold the potential to significantly enhance this…

Cited by 0SourceScholar
2025

StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition

ICCV 2025poster

With the rise of real-world human-AI interaction applications, such as AI assistants, the need for Streaming Video Dialogue is critical. To address this need, we introduce StreamMind, a video LLM framework that achieves ultra-FPS streaming video processing (100 fps on a single A100) and enables proa…

Cited by 0SourcePDFScholar
2025

Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach

EMNLP 2025

The automatic control of mobile devices is essential for efficiently performing complex tasks that involve multiple sequential steps. However, these tasks pose significant challenges due to the limited environmental information available at each step, primarily through visual observations. As a resu

2025

VERO: Verification and Zero-Shot Feedback Acquisition for Few-Shot Multimodal Aspect-Level Sentiment Classification

AAAI 2025technical

Deep learning approaches for multimodal aspect-level sentiment classification (MALSC) often require extensive data, which is costly and time-consuming to obtain. To mitigate this, current methods typically fine-tune small-scale pretrained models like BERT and BART with few-shot examples. While these…

2024

Adaptive Abrupt Disturbance Rejection Tracking Control for Wheeled Mobile Robots

RA-L 2024

Uncertain disturbances increase the difficulty of robust tracking control for wheeled mobile robots (WMRs) in industrial scenarios, especially when exhibiting abrupt changes. This letter proposes an adaptive abrupt disturbance-rejection sliding mode controller (SMC). To address the increased variabi

Cited by 13SourceScholar
2024

CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks

ECCV 2024poster

"Transferable targeted adversarial attacks aim to mislead models into outputting adversary-specified predictions in black-box scenarios. Recent studies have introduced single-target attacks that train a generator for each target class to generate highly transferable perturbations, resulting in subst…

2024

Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion Model

NeurIPS 2024poster

Spatio-temporal (ST) prediction has garnered a De facto attention in earth sciences, such as meteorological prediction, human mobility perception. However, the scarcity of data coupled with the high expenses involved in sensor deployment results in notable data imbalances. Furthermore, models that a…

Cited by 2SourcePDFScholar
2024

DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement

CVPR 2024poster

We present Dive Into the Boundaries (DIBS) a novel pretraining framework for dense video captioning (DVC) that elaborates on improving the quality of the generated event captions and their associated pseudo event boundaries from unlabeled videos. By leveraging the capabilities of diverse large langu…

Cited by 10SourcePDFScholar
2024

Divide-and-Conquer Predictive Coding: a structured Bayesian inference algorithm

NeurIPS 2024poster

Unexpected stimuli induce "error" or "surprise" signals in the brain. The theory of predictive coding promises to explain these observations in terms of Bayesian inference by suggesting that the cortex implements variational inference in a probabilistic graphical model. However, when applied to mach…

Cited by 1SourcePDFScholar
2024

Earthfarsser: Versatile Spatio-Temporal Dynamical Systems Modeling in One Model

AAAI 2024technical

Efficiently modeling spatio-temporal (ST) physical processes and observations presents a challenging problem for the deep learning community. Many recent studies have concentrated on meticulously reconciling various advantages, leading to designed models that are neither simple nor practical. To add…

2024

Faster Differentially Private Top-$k$ Selection: A Joint Exponential Mechanism with Pruning

NeurIPS 2024poster

We study the differentially private top-$k$ selection problem, aiming to identify a sequence of $k$ items with approximately the highest scores from $d$ items. Recent work by Gillenwater et al. (2022) employs a direct sampling approach from the vast collection of $O(d^k)$ possible length-$k$ sequenc…

Cited by 0SourcePDFScholar
2024

LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic Images

CVPR 2024highlight

Lithic Use-Wear Analysis (LUWA) using microscopic images is an underexplored vision-for-science research area. It seeks to distinguish the worked material which is critical for understanding archaeological artifacts material interactions tool functionalities and dental records. However this challeng…

Cited by 4SourcePDFScholar
2024

MuSR: Multi-Scale 3D Scenes Reconstruction based on Monocular Video

ICASSP 2024accepted

Three-dimensional (3D) scene reconstruction, particularly from monocular videos, is a significant challenge in large-scale scenarios due to difficulty handling varying object sizes and high computational resource needs. This paper introduces MuSR, a novel multi-scale reconstruction method addressing…

Cited by 0SourceScholar
2024

NuwaDynamics: Discovering and Updating in Causal Spatio-Temporal Modeling

ICLR 2024spotlight

Spatio-temporal (ST) prediction plays a pivotal role in earth sciences, such as meteorological prediction, urban computing. Adequate high-quality data, coupled with deep models capable of inference, are both indispensable and prerequisite for achieving meaningful results. However, the sparsity of da…

Cited by 12SourcePDFScholar
2024

PURE: Prompt Evolution with Graph ODE for Out-of-distribution Fluid Dynamics Modeling

NeurIPS 2024poster

This work studies the problem of out-of-distribution fluid dynamics modeling. Previous works usually design effective neural operators to learn from mesh-based data structures. However, in real-world applications, they would suffer from distribution shifts from the variance of system parameters and…

Cited by 5SourcePDFScholar
2024

PathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Pathology

ECCV 2024oral

"The emergence of Large Multimodal Models (LMMs) has unlocked remarkable potential in AI, particularly in pathology. However, the lack of specialized, high-quality benchmark impeded their development and precise evaluation. To address this, we introduce PathMMU, the largest and highest-quality exper…

Cited by 9SourcePDFScholar
2024

Prometheus: Out-of-distribution Fluid Dynamics Modeling with Disentangled Graph ODE

ICML 2024poster

Fluid dynamics modeling has received extensive attention in the machine learning community. Although numerous graph neural network (GNN) approaches have been proposed for this problem, the problem of out-of-distribution (OOD) generalization remains underexplored. In this work, we propose a new large…

Cited by 10SourcePDFScholar
2024

Revisiting Graph-Based Fraud Detection in Sight of Heterophily and Spectrum

AAAI 2024technical

Graph-based fraud detection (GFD) can be regarded as a challenging semi-supervised node binary classification task. In recent years, Graph Neural Networks (GNN) have been widely applied to GFD, characterizing the anomalous possibility of a node by aggregating neighbor information. However, fraud gra…

2024

Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling

NeurIPS 2024poster

In current deep learning tasks, Adam-style optimizers—such as Adam, Adagrad, RMSprop, Adafactor, and Lion—have been widely used as alternatives to SGD-style optimizers. These optimizers typically update model parameters using the sign of gradients, resulting in more stable convergence curves. The l…

Cited by 6SourcePDFScholar
2024

UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation

AAAI 2024technical

Multi-object tracking (MOT) in video sequences remains a challenging task, especially in scenarios with significant camera movements. This is because targets can drift considerably on the image plane, leading to erroneous tracking outcomes. Addressing such challenges typically requires supplementary…

2024

VCR-Graphormer: A Mini-batch Graph Transformer via Virtual Connections

ICLR 2024poster

Graph transformer has been proven as an effective graph learning method for its adoption of attention mechanism that is capable of capturing expressive representations from complex topological and feature information of graphs. Graph transformer conventionally performs dense attention (or global att…

2024

VIDAR: Data Quality Improvement for Monocular 3D Reconstruction through In-situ Visual Interaction

ICRA 2024poster

3D reconstruction based on monocular videos has attracted wide attention, and existing reconstruction methods usually work in a reconstruction-after-scanning manner. However, these methods suffer from insufficient data collection problems due to the lack of effective guidance for users during the sc…

Cited by 2SourceScholar
2023

A Closer Look at Few-shot Classification Again

ICML 2023poster

Few-shot classification consists of a training phase where a model is learned on a relatively large dataset and an adaptation phase where the learned model is adapted to previously-unseen tasks with limited labeled samples. In this paper, we empirically prove that the training algorithm and the adap…

2023

A Simple Framework for Text-Supervised Semantic Segmentation

CVPR 2023poster

Text-supervised semantic segmentation is a novel research topic that allows semantic segments to emerge with image-text contrasting. However, pioneering methods could be subject to specifically designed network architectures. This paper shows that a vanilla contrastive language-image pre-training (C…

2023

An Application of Quantum Mechanics to Attention Methods in Computer Vision

ICASSP 2023accepted

This work proposes the quantum-state-based mapping (QSM) for machine learning. QSM uses wave functions that describe microscopic particle systems as mappings. By QSM, original inputs or features extracted by neural networks are processed as quantum states to train wave function parameters. QSM has a…

Cited by 0SourceScholar
2023

Do We Really Need Complicated Model Architectures For Temporal Networks?

ICLR 2023top-5%

Recurrent neural network (RNN) and self-attention mechanism (SAM) are the de facto methods to extract spatial-temporal information for temporal graph learning. Interestingly, we found that although both RNN and SAM could lead to a good performance, in practice neither of them is always necessary. In…

Cited by 156SourcePDFScholar
2023

Enhancing Robustness and Imperceptibility of Blind Watermarking with Improved Message Processor

ICASSP 2023accepted

The current state-of-the-art(SOTA) blind watermark embedding method MBRS based on deep learning is less robust to Crop, and additional diffusion layers need to be added for optimization. However, the diffusion layer will make the model less robust to noise other than Crop. Therefore, MBRS which need…

Cited by 0SourceScholar
2023

Exploring Universal Singing Speech Language Identification Using Self-Supervised Learning Based Front-End Features

ICASSP 2023accepted

Despite the great performance of language identification (LID), there is a lack of large-scale singing LID databases to support the research of singing language identification (SLID). This paper proposed a over 3200 hours dataset used for singing language identification, called Slingua. As the basel…

Cited by 0SourceScholar
2023

IDEA: An Invariant Perspective for Efficient Domain Adaptive Image Retrieval

NeurIPS 2023poster

In this paper, we investigate the problem of unsupervised domain adaptive hashing, which leverage knowledge from a label-rich source domain to expedite learning to hash on a label-scarce target domain. Although numerous existing approaches attempt to incorporate transfer learning techniques into dee…

Cited by 6SourcePDFScholar
2023

SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

ICML 2023poster

Large language models (LLMs) show excellent performance but are compute- and memory-intensive. Quantization can reduce memory and accelerate inference. However, existing methods cannot maintain accuracy and hardware efficiency at the same time. We propose SmoothQuant, a training-free, accuracy-prese…

2022

A Meta-framework for Spatiotemporal Quantity Extraction from Text

ACL 2022long

News events are often associated with quantities (e.g., the number of COVID-19 patients or the number of arrests in a protest), and it is often important to extract their type, time, and location from unstructured text in order to analyze these quantity events. This paper thus formulates the NLP pro…

Cited by 13SourcePDFScholar
2022

CODE-MVP: Learning to Represent Source Code from Multiple Views with Contrastive Pre-Training

NAACL 2022findings

Recent years have witnessed increasing interest in code representation learning, which aims to represent the semantics of source code into distributed vectors. Currently, various works have been proposed to represent the complex semantics of source code from different views, including plain text, Ab…

2022

Compilable Neural Code Generation with Compiler Feedback

ACL 2022findings

Automatically generating compilable programs with (or without) natural language descriptions has always been a touchstone problem for computational linguistics and automated software engineering. Existing deep-learning approaches model code generation as text generation, either constrained by gramma…

Cited by 73SourcePDFScholar
2022

Contrastive Vision-Language Pre-training with Limited Resources

ECCV 2022poster

"Pioneering dual-encoder pre-training works (e.g., CLIP and ALIGN) have revealed the potential of aligning multi-modal representations with contrastive learning. However, these works require a tremendous amount of data and computational resources (e.g., billion-level web data and hundreds of GPUs),…

2022

Exploiting Unlabeled Data for Target-Oriented Opinion Words Extraction

COLING 2022main

Target-oriented Opinion Words Extraction (TOWE) is a fine-grained sentiment analysis task that aims to extract the corresponding opinion words of a given opinion target from the sentence. Recently, deep learning approaches have made remarkable progress on this task. Nevertheless, the TOWE task still…

2022

Fixed and Sliding FBG Sensors-Based Triaxial Tip Force Sensing for Cable-Driven Continuum Robots

ICRA 2022poster

Tip force sensing for cable-driven continuum robots are vital to provide the force information for safe and reliable human-robot interaction. However, traditional triaxial force sensors usually have a complicated structure occupying its inner lumen, without enough space for additional instrumental t…

Cited by 6SourceScholar
2022

Grasping State Analysis of Soft Manipulator Based on Flexible Tactile Sensor Array

IROS 2022poster

Although the grasping state analysis is vital in the study of manipulators, the grasping state analysis of soft manipulators as an independent research topic is not much so far. This paper proposes a novel pneumatic soft manipulator with a flexible tactile sensor array (SM-FTSA). The flexible tactil…

Cited by 2SourceScholar
2022

L-CoDe:Language-Based Colorization Using Color-Object Decoupled Conditions

AAAI 2022technical

Colorizing a grayscale image is inherently an ill-posed problem with multi-modal uncertainty. Language-based colorization offers a natural way of interaction to reduce such uncertainty via a user-provided caption. However, the color-object coupling and mismatch issues make the mapping from word to c…

Cited by 41SourcePDFScholar
2021

Boosting Offline Reinforcement Learning with Residual Generative Modeling

IJCAI 2021poster

Offline reinforcement learning (RL) tries to learn the near-optimal policy with recorded offline experience without online exploration.Current offline RL research includes: 1) generative modeling, i.e., approximating a policy using fixed data; and 2) learning the state-action value function. While m…

Cited by 15SourcePDFScholar
2021

Conjugate Energy-Based Models

ICML 2021spotlight

In this paper, we propose conjugate energy-based models (CEBMs), a new class of energy-based models that define a joint density over data and latent variables. The joint density of a CEBM decomposes into an intractable distribution over data and a tractable posterior over latent variables. CEBMs hav…

Cited by 7SourcePDFScholar
2021

Continuous Cnn For Nonuniform Time Series

ICASSP 2021accepted

CNN for time series data implicitly assumes that the data are uniformly sampled, whereas many event-based and multi-modal data are nonuniform or have heterogeneous sampling rates. Directly applying regular CNN to nonuniform time series is ungrounded, because it is unable to recognize and extract com…

Cited by 0SourceScholar
2021

FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling

NeurIPS 2021poster

The recently proposed FixMatch achieved state-of-the-art results on most semi-supervised learning (SSL) benchmarks. However, like other modern SSL algorithms, FixMatch uses a pre-defined constant threshold for all classes to select unlabeled data that contribute to the training, thus failing to cons…

2021

Learning proposals for probabilistic programs with inference combinators

UAI 2021poster

We develop operators for construction of proposals in probabilistic programs, which we refer to as inference combinators. Inference combinators define a grammar over importance samplers that compose primitive operations such as application of a transition kernel and importance resampling. Proposals…

2021

Learning the Best Pooling Strategy for Visual Semantic Embedding

CVPR 2021poster

Visual Semantic Embedding (VSE) is a dominant approach for vision-language retrieval, which aims at learning a deep embedding space such that visual data are embedded close to their semantic text labels or descriptions. Recent VSE models use complex methods to better contextualize and aggregate mult…

Cited by 296PDFcodeScholar
2021

Self-Supervised Learning for Sleep Stage Classification with Predictive and Discriminative Contrastive Coding

ICASSP 2021accepted

The purpose of this paper is to learn efficient representations from raw electroencephalogram (EEG) signals for sleep stage classification via self-supervised learning (SSL). Although supervised methods have gained favorable performance, they heavily rely on manually labeled datasets. Recently, SSL…

Cited by 0SourceScholar
2021

Training Spiking Neural Networks with Accumulated Spiking Flow

AAAI 2021technical

The fast development of neuromorphic hardwares promotes Spiking Neural Networks (SNNs) to a thrilling research avenue. Current SNNs, though much efficient, are less effective compared with leading Artificial Neural Networks (ANNs) especially in supervised learning tasks. Recent efforts further demon…

2020

A Variational Approach for Learning from Positive and Unlabeled Data

NeurIPS 2020poster

Learning binary classifiers only from positive and unlabeled (PU) data is an important and challenging task in many real-world applications, including web text classification, disease gene identification and fraud detection, where negative samples are difficult to verify experimentally. Most recent PU l…

2020

Amortized Population Gibbs Samplers with Neural Sufficient Statistics

ICML 2020poster

We develop amortized population Gibbs (APG) samplers, a class of scalable methods that frame structured variational inference as adaptive importance sampling. APG samplers construct high-dimensional proposals by iterating over updates to lower-dimensional blocks of variables. We train each condition…

Cited by 7SourcePDFScholar
2020

Structured Multi-Hashing for Model Compression

CVPR 2020poster

Despite the success of deep neural networks (DNNs), state-of-the-art models are too large to deploy on low-resource devices or common server configurations in which multiple models are held in memory. Model compression methods address this limitation by reducing the memory footprint, latency, or ene…

Cited by 19PDFScholar
2019

Accurate Vehicle Detection Using Multi-camera Data Fusion and Machine Learning

ICASSP 2019accepted

Computer-vision methods have been extensively used in intelligent transportation systems for vehicle detection. However, the detection of severely occluded or partially observed vehicles due to the limited camera fields of view remains a challenge. This paper presents a multi-camera vehicle detectio…

Cited by 0SourceScholar
2019

Structured Disentangled Representations

AISTATS 2019poster

Deep latent-variable models learn representations of high-dimensional data in an unsupervised manner. A number of recent efforts have focused on learning representations that disentangle statistically independent axes of variation by introducing modifications to the standard objective function. Thes…

2019

Unified Visual-Semantic Embeddings: Bridging Vision and Language With Structured Meaning Representations

CVPR 2019oral

We propose the Unified Visual-Semantic Embeddings (Unified VSE) for learning a joint space of visual representation and textual semantics. The model unifies the embeddings of concepts at different levels: objects, attributes, relations, and full scenes. We view the sentential semantics as a combinat…

Cited by 221PDFcodeScholar
2018

Mixed Precision Training

ICLR 2018poster

Increasing the size of a neural network typically improves accuracy but also increases the memory and compute requirements for training the model. We introduce methodology for training deep neural networks using half-precision floating point numbers, without losing model accuracy or having to modify…

Cited by 2212SourcePDFScholar
2018

MorphNet: Fast & Simple Resource-Constrained Structure Learning of Deep Networks

CVPR 2018poster

We present MorphNet, an approach to automate the design of neural network structures. MorphNet iteratively shrinks and expands a network, shrinking via a resource-weighted sparsifying regularizer on activations and expanding via a uniform multiplicative factor on all layers. In contrast to previou…

Cited by 432SourcePDFScholar