← Search

Tao Huang

59 accepted papers

2026

AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios

AAAI 2026technical

Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current RMOT research remains mostly confined to ground-level scenarios, which constrain

Cited by 0SourcePDFScholar
2026

Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation

CVPR 2026

Vision-Language-Action (VLA) models have shown great performance in robotic manipulation by mapping visual observations and language instructions directly to actions. However, they remain brittle under distribution shifts: when test scenarios change, VLAs often reproduce memorized trajectories inste

Cited by 0SourcecodeScholar
2026

Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification

ICLR 2026poster

Large vision-language models (LVLMs) have shown substantial advances in multimodal understanding and generation. However, when presented with incompetent or adversarial inputs, they frequently produce unreliable or even harmful contents, such as fact hallucinations or dangerous instructions. This mi…

Cited by 0SourcecodeScholar
2026

Distilling Cross-Modal Knowledge via Feature Disentanglement

AAAI 2026technical

Knowledge distillation (KD) has proven highly effective for compressing large models and enhancing the performance of smaller ones. However, its effectiveness diminishes in cross-modal scenarios, such as vision-to-language distillation, where inconsistencies in representation across modalities lead

Cited by 0SourcePDFScholar
2026

Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control

CVPR 2026

Humans intuitively move to sound, but current humanoid robots lack expressive improvisational capabilities, confined to predefined motions or sparse commands. Generating motion from audio and then retargeting it to robots relies on explicit motion reconstruction, leading to cascaded errors, high lat

Cited by 0SourceScholar
2026

Enhanced Privacy Leakage from Noise-Perturbed Gradients via Gradient-Guided Conditional Diffusion Models

AAAI 2026technical

Federated learning synchronizes models through gradient transmission and aggregation. However, these gradients pose significant privacy risks, as sensitive training data is embedded within them. Existing gradient inversion attacks suffer from significantly degraded reconstruction performance when gr

Cited by 0SourcePDFScholar
2026

From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance

ICLR 2026poster

Natural language offers a natural interface for humanoid robots, but existing text-to-motion pipelines remain cumbersome and unreliable. They typically decode human motion, retarget it to robot morphology, and then track it with a physics-based controller. However, this multi-stage process is prone…

Cited by 0SourceScholar
2026

IAFMNet: Information-Aware Feature Modulation for Efficient Super-Resolution

CVPR 2026

Single Image Super-Resolution (SISR) aims to reconstruct a high-resolution (HR) image from a low-resolution (LR) input, a task that becomes increasingly challenging under real-world computational constraints. However, most efficient SISR methods adopt lightweight, spatially uniform strategies that a

Cited by 0SourceScholar
2026

Kronecker Generative Networks: A General Neural Architecture for Parameter-Efficient Learning Across Classification Tasks

ICML 2026poster

Modern neural networks derive much of their effectiveness from rich connectivity patterns. Yet, existing architectures often fix the topology at either the sparse or dense extremes, thereby limiting structural flexibility and analysis. We propose Kronecker Generative Networks (KGNs), an algebraic fr…

Cited by 0SourceScholar
2026

Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions

CVPR 2026

Video data is more cost-effective than motion capture data for learning 3D character controllers, yet using it to generate realistic and physically plausible motions remains challenging. Previous approaches typically rely on off-the-shelf motion reconstruction techniques to extract 3D kinematic traj

Cited by 0SourcecodeScholar
2026

MS-Occ: Multi-Stage LiDAR-Camera Fusion for 3D Semantic Occupancy Prediction

RA-L 2026

Accurate 3D semantic occupancy perception is essential for autonomous driving in complex environments with diverse and irregular objects. While vision-centric methods suffer from geometric inaccuracies, LiDAR-based approaches often lack rich semantic information. To address these limitations, MS-Occ

Cited by 2SourceScholar
2026

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

ICML 2026poster

Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often suffers from a substantial visual domain gap and limited environmental diversity, resulting in weak real-world generaliza…

Cited by 0SourceScholar
2025

Adversarial Training and Cross-modal Feature Fusion in Multimodal Sentiment Analysis

ICASSP 2025accepted

Multimodal sentiment analysis recognizes emotions through text, audio, and visual modalities, but data incompleteness is a major challenge. Existing methods often focus on specific types of deficiencies and perform poorly when multiple types of noise are present simultaneously. To address this issue…

Cited by 0SourceScholar
2025

BeamDojo: Learning Agile Humanoid Locomotion on Sparse Footholds

RSS 2025poster

Traversing risky terrains with sparse footholds poses a significant challenge for humanoid robots, requiring precise foot placements and stable locomotion. Existing approaches designed for quadrupedal robots often fail to generalize to humanoid robots due to differences in foot geometry and unstable…

Cited by 6PDFScholar
2025

Bridging Task Boundaries: Remote Sensing Image-Text Retrieval via Dictionary-Driven Adaptation

ICASSP 2025accepted

Given image (or text), remote sensing image-text retrieval (RSITR) aims to retrieve corresponding text (or image) within diverse remote sensing data. However, due to the complex scenes and compact distribution of targets in remote sensing data, existing methods, particularly those leveraging large m…

Cited by 0SourceScholar
2025

Efficient Multi-Robot Task and Path Planning in Large-Scale Cluttered Environments

RA-L 2025

As the potential of multi-robot systems continues to be explored and validated across various real-world applications, such as package delivery, search and rescue, and autonomous exploration, the need to improve the efficiency and quality of task and path planning has become increasingly urgent, par

Cited by 4SourceScholar
2025

EfficientVMamba: Atrous Selective Scan for Light Weight Visual Mamba

AAAI 2025technical

Prior efforts in light-weight model development mainly centered on CNN and Transformer-based designs yet faced persistent challenges. CNNs adept at local feature extraction compromise resolution while Transformers offer global reach but escalate computational demands O(N^2). This ongoing trade-off b…

2025

GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scene

ICCV 2025poster

Generalizable active mapping in complex unknown environments remains a critical challenge for mobile robots. Existing methods, constrained by limited training data and conservative exploration strategies, struggle to generalize across scenes with diverse layouts and complex connectivity. To enable s…

Cited by 0SourcePDFScholar
2025

Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring Classes

CVPR 2025poster

Recent approaches, such as data augmentation, adversarial training, and transfer learning, have shown potential in addressing the issue of performance degradation caused by distributional shifts. However, they typically demand careful design in terms of data or models and lack awareness of the impac…

Cited by 0SourcePDFScholar
2025

Learning Humanoid Locomotion with Perceptive Internal Model

ICRA 2025

In contrast to quadruped robots that can navigate diverse terrains using a “blind” policy, humanoid robots require accurate perception for stable locomotion due to their high degrees of freedom and inherently unstable morphology. However, incorporating perceptual signals often introduces additional

Cited by 84SourceScholar
2025

Learning Humanoid Standing-up Control across Diverse Postures

RSS 2025poster

Standing-up control is crucial for humanoid robots, with the potential for integration into current locomotion and loco-manipulation systems. Existing approaches are either limited to simulations that neglect hardware constraints or rely on predefined ground-specific motion trajectories, failing to…

Cited by 6PDFScholar
2025

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders

CVPR 2025poster

Visual encoders are fundamental components in vision-language models (VLMs), each showcasing unique strengths derived from various pre-trained visual foundation models. To leverage the various capabilities of these encoders, recent studies incorporate multiple encoders within a single VLM, leading t…

2025

Parameterized Blur Kernel Prior Learning for Local Motion Deblurring

CVPR 2025poster

Unlike global motion blur, Local Motion Deblurring (LMD) presents a more complex challenge, as it requires precise restoration of blurry regions while preserving the sharpness of the background. Existing LMD methods rely on manually annotated blur masks and often overlook the blur kernel's character…

Cited by 0SourcePDFScholar
2025

PatternCIR Benchmark and TisCIR: Advancing Zero-Shot Composed Image Retrieval in Remote Sensing

IJCAI 2025

Remote sensing composed image retrieval (RSCIR) is a new vision-language task that takes a composed query of an image and text, aiming to search for a target remote sensing image satisfying two conditions from intricate remote sensing imagery. However, the existing attribute-based benchmark Patternc

Cited by 0SourcePDFScholar
2025

Problem Solving-Oriented Programming Knowledge Tracing from Behavior to Thought

ICASSP 2025accepted

Programming knowledge tracing (programming KT) aims to analyze the dynamic programming states in solving problems based on historical behaviors and predict future performance. In programming, a student’s thought process can lead to multiple solutions for the same problem. However, current programmin…

Cited by 0SourceScholar
2025

Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets

ICLR 2025poster

The pre-training of visual representations has enhanced the efficiency of robot learning. Due to the lack of large-scale in-domain robotic datasets, prior works utilize in-the-wild human videos to pre-train robotic visual representation. Despite their promising results, representations from human vi…

2025

SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures

ICCV 2025poster

While the Segment Anything Model (SAM) transforms interactive segmentation with zero-shot abilities, its inherent vulnerabilities present a single-point risk, potentially leading to the failure of downstream applications. Proactively evaluating these transferable vulnerabilities is thus imperative.…

2025

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

ICML 2025poster

In vision-language models (VLMs), visual tokens usually consume a significant amount of computational overhead, despite their sparser information density compared to text tokens. To address this, most existing methods learn a network to prune redundant visual tokens and require additional training d…

2025

Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models

NeurIPS 2025poster

In this work, we address the task of table image to LaTeX code generation, with the goal of automating the reconstruction of high-quality, publication-ready tables from visual inputs. A central challenge of this task lies in accurately handling complex tables—those with large sizes, deeply nested st…

Cited by 0SourceScholar
2025

VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching

NeurIPS 2025poster

Vision-Language-Action (VLA) models have demonstrated strong multi-modal reasoning capabilities, enabling direct action generation from visual perception and language instructions in an end-to-end manner. However, their substantial computational cost poses a challenge for real-time robotic control,…

Cited by 0SourcecodeScholar
2024

A Framework for Real-time Generation of Multi-directional Traversability Maps in Unstructured Environments

ICRA 2024poster

In complex unstructured environments, accurate terrain traversability analysis is a fundamental requirement for the successful execution of any movements of ground robots, especially given that terrain traversability often exhibits anisotropy. However, the difficulty in obtaining multi-directional t…

Cited by 1SourceScholar
2024

BE-SLAM: BEV-Enhanced Dynamic Semantic SLAM with Static Object Reconstruction

IROS 2024poster

The quality of a robot’s environmental perception determines whether it can achieve more intelligent applications, such as semantic interaction with humans. SLAM, on the other hand, is one of the crucial capabilities for a robot to perceive its environment. However, when only a monocular image is pr…

Cited by 0SourceScholar
2024

FreeKD: Knowledge Distillation via Semantic Frequency Prompt

CVPR 2024poster

Knowledge distillation (KD) has been applied to various tasks successfully and mainstream methods typically boost the student model via spatial imitation losses. However the consecutive downsamplings induced in the spatial domain of teacher model is a type of corruption hindering the student from an…

2024

On the Federated Learning Framework for Cooperative Perception

RA-L 2024

Cooperative perception (CP) is essential to enhance the efficiency and safety of future transportation systems, requiring extensive data sharing among vehicles on the road, which raises significant privacy concerns. Federated learning offers a promising solution by enabling data privacy-preserving c

Cited by 10SourceScholar
2024

Strong Transferable Adversarial Attacks via Ensembled Asymptotically Normal Distribution Learning

CVPR 2024highlight

Strong adversarial examples are crucial for evaluating and enhancing the robustness of deep neural networks. However the performance of popular attacks is usually sensitive for instance to minor image transformations stemming from limited information -- typically only one input example a handful of…

2024

Unveiling the Tapestry of Consistency in Large Vision-Language Models

NeurIPS 2024poster

Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information. However, when faced with prompts in different sizes of solution spaces, LVLMs fail to always give consistent answers regarding the same knowl…

2023

Demonstration-Guided Reinforcement Learning with Efficient Exploration for Task Automation of Surgical Robot

ICRA 2023poster

Task automation of surgical robot has the potentials to improve surgical efficiency. Recent reinforcement learning (RL) based approaches provide scalable solutions to surgical automation, but typically require extensive data collection to solve a task if no prior knowledge is given. This issue is kn…

Cited by 29SourcecodeScholar
2023

Human-in-the-Loop Embodied Intelligence With Interactive Simulation Environment for Surgical Robot Learning

RA-L 2023

Surgical robot automation has attracted increasing research interest over the past decade, expecting its potential to benefit surgeons, nurses and patients. Recently, the learning paradigm of embodied intelligence has demonstrated promising ability to learn good control policies for various complex

Cited by 55SourcecodeScholar
2023

Knowledge Diffusion for Distillation

NeurIPS 2023poster

The representation gap between teacher and student is an emerging topic in knowledge distillation (KD). To reduce the gap and improve the performance, current methods often resort to complicated training schemes, loss functions, and feature alignments, which are task-specific and feature-specific. I…

2023

Low-Light Image Enhancement with Multi-Stage Residue Quantization and Brightness-Aware Attention

ICCV 2023poster

Low-light image enhancement (LLIE) aims to recover illumination and improve the visibility of low-light images. Conventional LLIE methods often produce poor results because they neglect the effect of noise interference. Deep learning-based LLIE methods focus on learning a mapping function between lo…

Cited by 26PDFcodeScholar
2023

Masked Distillation with Receptive Tokens

ICLR 2023poster

Distilling from the feature maps can be fairly effective for dense prediction tasks since both the feature discriminability and localization information can be well transferred. However, not every pixel contributes equally to the performance, and a good student should learn from what really matters…

2023

Value-Informed Skill Chaining for Policy Learning of Long-Horizon Tasks with Surgical Robot

IROS 2023poster

Reinforcement learning is still struggling with solving long-horizon surgical robot tasks which involve multiple steps over an extended duration of time due to the policy exploration challenge. Recent methods try to tackle this problem by skill chaining, in which the long-horizon task is decomposed…

Cited by 7SourcecodeScholar
2022

Data Agnostic Filter Gating For Efficient Deep Networks

ICASSP 2022accepted

Filter pruning is essential for deploying a well-trained CNN model on edge computation devices with a target computation budget (e.g., FLOPs). Current filter pruning methods mainly focus on leveraging feature maps to analyze the importance of filters, and prune those with less impact on the value of…

Cited by 0SourceScholar
2022

DyRep: Bootstrapping Training With Dynamic Re-Parameterization

CVPR 2022poster

Structural re-parameterization (Rep) methods achieve noticeable improvements on simple VGG-style networks. Despite the prevalence, current Rep methods simply re-parameterize all operations into an augmented network, including those that rarely contribute to the model's performance. As such, the pric…

Cited by 42PDFcodeScholar
2022

GreedyNASv2: Greedier Search With a Greedy Path Filter

CVPR 2022poster

Training a good supernet in one-shot NAS methods is difficult since the search space is usually considerably huge (e.g., 13^ 21 ). In order to enhance the supernet's evaluation ability, one greedy strategy is to sample good paths, and let the supernet lean towards the good ones and ease its evaluati…

Cited by 22PDFScholar
2021

Deep Gaussian Scale Mixture Prior for Spectral Compressive Imaging

CVPR 2021poster

In coded aperture snapshot spectral imaging (CASSI) system, the real-world hyperspectral image (HSI) can be reconstructed from the captured compressive image in a snapshot. Model-based HSI reconstruction methods employed hand-crafted priors to solve the reconstruction problem, but most of which achi…

Cited by 185PDFScholar
2021

Improving Privacy Guarantee and Efficiency of Latent Dirichlet Allocation Model Training Under Differential Privacy

EMNLP 2021finding

Latent Dirichlet allocation (LDA), a widely used topic model, is often employed as a fundamental tool for text analysis in various applications. However, the training process of the LDA model typically requires massive text corpus data. On one hand, such massive data may expose private information i…

Cited by 5SourcePDFScholar
2021

Locally Free Weight Sharing for Network Width Search

ICLR 2021spotlight

Searching for network width is an effective way to slim deep neural networks with hardware budgets. With this aim, a one-shot supernet is usually leveraged as a performance evaluator to rank the performance \wrt~different width. Nevertheless, current methods mainly follow a manually fixed weight sha…

Cited by 45SourcePDFScholar
2021

Neighbor2Neighbor: Self-Supervised Denoising From Single Noisy Images

CVPR 2021poster

In the last few years, image denoising has benefited a lot from the fast development of neural networks. However, the requirement of large amounts of noisy-clean image pairs for supervision limits the wide use of these models. Although there have been a few attempts in training an image denoising mo…

Cited by 451PDFcodeScholar
2021

Prioritized Architecture Sampling With Monto-Carlo Tree Search

CVPR 2021poster

One-shot neural architecture search (NAS) methods significantly reduce the search cost by considering the whole search space as one network, which only needs to be trained once. However, current methods select each operation independently without considering previous layers. Besides, the historical…

Cited by 66PDFcodeScholar
2020

GreedyNAS: Towards Fast One-Shot NAS With Greedy Supernet

CVPR 2020poster

Training a supernet matters for one-shot neural architecture search (NAS) methods since it serves as a basic performance estimator for different architectures (paths). Current methods mainly hold the assumption that a supernet should give a reasonable ranking over all paths. They thus treat all path…

Cited by 188PDFScholar