← Search

Tao Chen

130 accepted papers

2026

A Real-Time 6-DoF Posture Estimation Method for High-Speed 6-Axis Industrial Manipulator Control Using a 2D Laser Profiler

ICRA 2026poster

6-DoF posture estimation is a critical technique in robotics. However, a significant gap exists between the two primary approaches—camera-based methods and laser tracker systems—in terms of cost and performance. To bridge this gap, this work proposes a method to calculate the 6-DoF posture of a mani…

Cited by 0SourceScholar
2026

Beyond Frequency: Scoring-Driven Debiasing for Object Detection via Blueprint-Prompted Image Synthesis

ICLR 2026poster

This paper presents a generation-based debiasing framework for object detection. Prior debiasing methods are often limited by the representation diversity of samples, while naive generative augmentation often preserves the biases it aims to solve. Moreover, our analysis reveals that simply generatin…

Cited by 0SourcecodeScholar
2026

Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale

ICML 2026poster

Large Language Models (LLMs) have rapidly advanced, with Gemini-3-Pro setting a new performance milestone. In this work, we explore collective intelligence as an alternative to monolithic scaling, and demonstrate that open-source LLMs' collaboration can surpass Gemini-3-Pro. We first revisit LLM rou…

Cited by 0SourceScholar
2026

Beyond Quadratic: Linear-Time Change Detection with RWKV

AAAI 2026technical

Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependencies at a prohibitive computational cost. This paper introduces ChangeRWKV, a new architecture that reconciles this confli

Cited by 0SourcePDFScholar
2026

E²LoRA: Efficient and Effective Low-Rank Adaptation with Entropy-Guided Adaptive Sharing

ICLR 2026poster

As large pre-trained models rapidly scale, Parameter-Efficient Fine-Tuning (PEFT) through methods like Low-Rank Adaptation (LoRA) becomes increasingly crucial. While LoRA has emerged as a cornerstone of PEFT, excelling at preserving performance with minimal additional parameters, exploring paramete…

Cited by 0SourceScholar
2026

FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision–Language Models

ICML 2026poster

Efficiently enhancing the reasoning capabilities of Vision-Language Models (VLMs) by merging them with Large Reasoning Models (LRMs) has emerged as a promising direction. However, existing methods typically operate at a coarse-grained layer level, which often leads to a trade-off between injecting r…

Cited by 0SourceScholar
2026

Fair-FedMOE: Group-Fair One-Shot Federated Learning via Prototype-Guided Experts for Medical Imaging Analysis

ICML 2026poster

Group fairness can ensure equitable performance across different demographic subgroups for medical image analysis. However, the current fine-tuned foundation models (FMs) exhibit significant subgroup disparity. One-shot federated learning (OFL) can potentially mitigate this by leveraging cross-insti…

Cited by 0SourceScholar
2026

FedUP: Uncertainty-Aware Personalized Federated Learning via Probabilistic Prototypes

IJCAI 2026

Prototype-based federated learning enables efficient knowledge sharing by exchanging class prototypes rather than full model parameters. However, heterogeneous client data and limited local samples increase prototype estimation variance, making many client prototypes unreliable. Existing methods usu

Cited by 0Scholar
2026

GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy

ICML 2026poster

Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained credit assignment paradigm, where all tokens within the same response receive the identical reward. In this paper, we prop…

Cited by 0SourceScholar
2026

Gradient Intrinsic Dimensionality Alignment:Narrowing The Gap Between Low-Rank Adaptation and Full Fine-Tuning

ICLR 2026poster

Parameter-Efficient Fine-Tuning (PEFT) techniques, such as Low-Rank Adaptation (LoRA) and its variants, have emerged as critical tools for adapting large pretrained models under limited computational resources. However, a notable performance gap persists between these LoRA methods and Full Fine-Tuni…

Cited by 0SourceScholar
2026

HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?

ICML 2026poster

Recently, the physics reasoning capabilities of (M)LLMs have attracted growing attention. However, existing physics benchmarks suffer from two major gaps: they neither provide systematic and up-to-date coverage of physics Olympiads, nor enable direct performance comparison with humans. To bridge the…

Cited by 0SourceScholar
2026

Learning from Human Gaze: Human-like Robot Social Navigation in Dense Crowds

AAAI 2026technical

Robot navigation in dense crowds requires understanding social cues that humans naturally use, yet existing methods struggle with real-world complexity. We investigate two questions: (1) Where do pedestrians look when navigating crowds? and (2) Can eye tracking improve robot navigation? To answer, w

Cited by 0SourcePDFScholar
2026

MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMs

ICML 2026poster

Logical reasoning is a fundamental aspect of human intelligence and an essential capability for multimodal large language models (MLLMs). Despite the significant advancement in multimodal reasoning, existing benchmarks fail to comprehensively evaluate their reasoning abilities due to the lack of exp…

Cited by 0SourceScholar
2026

Mitigating Low-Quality Reasoning in MLLMs: Self-Driven Refined Multimodal CoT with Selective Thinking and Step-wise Visual Enhancement

AAAI 2026technical

Current Multimodal Chain-of-Thought (MCoT) methods suffer from low-quality multimodal reasoning, characterized by overthinking on simple queries and inefficient utilization of visual information, resulting in vast inefficient and ineffective computations. In this paper, we discover that Multimodal L

Cited by 0SourcePDFScholar
2026

MotionGPT3: Human Motion as a Second Modality

ICLR 2026poster

With the rapid progress of large language models (LLMs), multimodal frameworks that unify understanding and generation have become promising, yet they face increasing complexity as the number of modalities and tasks grows. We observe that motion quantization introduces approximation errors that cap…

Cited by 0SourceScholar
2026

PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation

CVPR 2026

Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract image-text alignment cues from cost volumes through a serial structure of spatial and class aggregations, leading to knowle

Cited by 0SourcecodeScholar
2026

PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic Segmentation

CVPR 2026

Training-free open-vocabulary semantic segmentation (OVSS) promises rapid adaptation to new label sets without retraining. Yet, many methods rely on heavy post-processing or handle text and vision in isolation, leaving cross-modal geometry underutilized. Others introduce auxiliary vision backbones o

Cited by 0SourcecodeScholar
2026

RegionE: Adaptive Region-Aware Generation for Efficient Image Editing

ICLR 2026poster

Recently, instruction-based image editing (IIE) has received widespread attention. In practice, IIE often modifies only specific regions of an image, while the remaining areas largely remain unchanged. Although these two types of regions differ significantly in generation difficulty and computationa…

Cited by 0SourceScholar
2026

Revisiting Learning with Noisy Labels: Active Forgetting and Noise Suppression

CVPR 2026

Learning with noisy labels (LNL) has received growing attention, with most prior work following the paradigm of clean-sample reliance (e.g., sample selection). However, this reliance also imposes intrinsic limitations, as overfitting to even a few noisy samples is inevitable, creating a major bottle

Cited by 0SourcecodeScholar
2026

Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach

CVPR 2026

Multimodal large language models suffer from substantial inference overhead since multimodal KV Cache grows proportionally with the visual input length. Existing multimodal KV cache compression methods mostly rely on attention score to reduce cache size, which makes them are incompatible with establ

Cited by 0SourceScholar
2026

SCI-Verifier: Scientific Verifier with Thinking

ICLR 2026poster

As large language models (LLMs) are increasingly applied to scientific reasoning, the complexity of answer formats and the diversity of equivalent expressions make answer verification a critical yet challenging task. Existing verification studies in scientific domains suffer from two major limitatio…

Cited by 0SourcecodeScholar
2026

SIGMA-PPG: Statistical-prior Informed Generative Masking Architecture for PPG Foundation Model

ICML 2026poster

Current foundation model for photoplethysmography (PPG) signals is challenged by the intrinsic redundancy and noise of the signal. Standard masked modeling often yields trivial solutions while contrastive methods lack morphological precision. To address these limitations, we propose a Statistical-pr…

Cited by 0SourceScholar
2026

Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism

CVPR 2026

Long video understanding is a key challenge that plagues the advancement of Multimodal Large language Models (MLLMs). In this paper, we study this problem from the perspective of visual memory mechanism, and proposed a novel and training-free approach, termed Flexible Memory (FlexMem). In principle,

Cited by 0SourcecodeScholar
2026

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

AAAI 2026technical

While Diffusion Transformers (DiTs) have achieved breakthroughs in video generation, this long sequence generation task remains constrained by the quadratic complexity of attention mechanisms, resulting in significant inference latency. Through detailed analysis of attention maps in Video Diffusion

Cited by 0SourcePDFScholar
2026

Tailoring the Training: Difficulty-Aware Learning Strategy Allocation for Large Language Models

ICML 2026poster

Although reinforcement learning (RL) enhances the reasoning capabilities of large language models (LLMs), it is primarily learned from the model's self-generated distribution, limiting its ability to acquire reasoning skills beyond its initial knowledge. To overcome this, we propose a Difficulty-Awa…

Cited by 0SourceScholar
2026

Toward Practical Equilibrium Propagation: Brain-inspired Recurrent Neural Network with Feedback Regulation and Residual Connections

ICLR 2026poster

Brain-like intelligent systems need brain-like learning methods. Equilibrium Propagation (EP) is a biologically plausible learning framework with strong potential for brain-inspired computing hardware. However, existing implementations of EP suffer from instability and prohibitively high computation…

Cited by 0SourceScholar
2025

All-in-One: Transferring Vision Foundation Models into Stereo Matching

AAAI 2025technical

As a fundamental vision task, stereo matching has made remarkable progress. While recent iterative optimization-based methods have achieved promising performance, their feature extraction capabilities still have room for improvement. Inspired by the ability of vision foundation models (VFMs) to ext…

Cited by 1SourcePDFScholar
2025

BadRefSR: Backdoor Attacks Against Reference-based Image Super Resolution

ICASSP 2025accepted

Reference-based image super-resolution (RefSR) represents a promising advancement in super-resolution (SR). In contrast to single-image super-resolution (SISR), RefSR leverages an additional reference image to help recover high-frequency details, yet its vulnerability to backdoor attacks has not bee…

Cited by 0SourceScholar
2025

Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models

EMNLP 2025

Large language models (LLMs) have shown remarkable capabilities in general domains, but their application to multi-omics biology remains underexplored. To address this gap, we introduce Biology-Instructions, the first large-scale instruction-tuning dataset for multi-omics biological sequences, inclu

2025

Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression

NeurIPS 2025poster

With the rise of the fine-tuned–pretrained paradigm, storing numerous fine-tuned models for multi-tasking creates significant storage overhead. Delta compression alleviates this by storing only the pretrained model and the highly compressed delta weights (the differences between fine-tuned and pretr…

Cited by 0SourcecodeScholar
2025

Chimera: Improving Generalist Model with Domain-Specific Experts

ICCV 2025poster

Large Multi-modal Models (LMMs), trained on web-scale datasets predominantly composed of natural images, have demonstrated remarkable performance on general tasks. However, these models often exhibit limited specialized capabilities for domain-specific tasks that require extensive domain prior knowl…

Cited by 0SourcePDFScholar
2025

Consistency-aware Self-Training for Iterative-based Stereo Matching

CVPR 2025poster

Iterative-based methods have become mainstream in stereo matching due to their high performance. However, these methods heavily rely on labeled data and face challenges with unlabeled real-world data. To this end, we propose a consistency-aware self-training framework for iterative-based stereo matc…

Cited by 0SourcePDFScholar
2025

DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models

CVPR 2025poster

Upcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into MoE layers. However, these models still suffer from significant parameter inefficiency due to the introduction of multipl…

Cited by 1SourcePDFScholar
2025

Distributed Cooperative Target Tracking and Active Sensing of Dual-AUV Based on Flank Array Sonar Detection

IROS 2025

When tracking underwater target, autonomous underwater vehicles (AUVs) need to estimate the target state based on the information detected by sensors and plan their own tracking paths accordingly to achieve active sensing of the target. When the sensor equipped on the AUV is a flank array sonar, the

Cited by 0SourceScholar
2025

Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback

ACL 2025long

The scientific research paradigm is undergoing a profound transformation owing to the development of Artificial Intelligence (AI). Recent works demonstrate that various AI-assisted research methods can largely improve research efficiency by improving data analysis, accelerating computation, and fost…

2025

ETA-IK: Execution-Time-Aware Inverse Kinematics for Dual-Arm Systems

IROS 2025

This paper presents ETA-IK, a novel Execution-Time-Aware Inverse Kinematics method tailored for dual-arm robotic systems. The primary goal is to optimize motion execution time by leveraging the redundancy of the entire system, specifically in tasks where only the relative pose of the robots is const

Cited by 0SourceScholar
2025

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have shown impressive video content understanding capabilities but struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding ability of existing MLLMs, we introduce FAVOR-Bench, which comprises 1,776 videos from both…

Cited by 0SourceScholar
2025

GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training

ICLR 2025poster

Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images…

Cited by 8SourcePDFScholar
2025

HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-View Reconstruction

ICLR 2025poster

Reconstructing 3D scenes from multiple viewpoints is a fundamental task in stereo vision. Recently, advances in generalizable 3D Gaussian Splatting have enabled high-quality novel view synthesis for unseen scenes from sparse input views by feed-forward predicting per-pixel Gaussian parameters withou…

2025

ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large-Scale Multitask Dataset

ICML 2025poster

Time-series data are critical in diverse applications, such as industrial monitoring, medical diagnostics, and climate research. However, effectively integrating these high-dimensional temporal signals with natural language for dynamic, interactive tasks remains a significant challenge. To address t…

2025

Once-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model Variants

CVPR 2025poster

Vision-language model (VLM) is one of the most important models for multi-modal tasks. Real industrial applications often meet the challenge of adapting VLMs to different scenarios, such as varying hardware platforms or performance requirements. Traditional methods involve training or fine-tuning to…

2025

One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild

ICASSP 2025accepted

Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and re…

Cited by 4SourceScholar
2025

PAMN: Multi-phase Correlation Modeling for Contrast-Enhanced 3D Medical Image Retrieval

EMNLP 2025

Contrast-enhanced 3D Medical imaging (e.g., CT, MRI) leverages phase sequences to uncover temporal dynamics vital for diagnosing tumors, lesions, and vascular issues. However, current retrieval models primarily focus on spatial features, neglecting phase-specific progression detailed in clinical rep

Cited by 0SourcePDFScholar
2025

PRING: Rethinking Protein-Protein Interaction Prediction from Pairs to Graphs

NeurIPS 2025poster

Deep learning-based computational methods have achieved promising results in predicting protein-protein interactions (PPIs). However, existing benchmarks predominantly focus on isolated pairwise evaluations, overlooking a model's capability to reconstruct biologically meaningful PPI networks, which…

Cited by 0SourcecodeScholar
2025

PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding

NeurIPS 2025poster

While Large Language Models (LLMs) demonstrate strong performance across domains, their long-context capabilities are limited by transient neural activations causing information decay and unstructured feed-forward network (FFN) weights leading to semantic fragmentation. Inspired by the brain’s worki…

Cited by 0SourceScholar
2025

Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning

CVPR 2025poster

Model quantization reduces the bit-width of weights and activations, improving memory efficiency and inference speed in diffusion models. However, achieving 4-bit quantization remains challenging. Existing methods, primarily based on integer quantization and post-training quantization fine-tuning, s…

Cited by 0SourcePDFScholar
2025

Research on Automated Microassembly Technology for ICF Target Core Microdevices Based on Teleoperation

RA-L 2025

This letter presents a teleoperation-based solution to address the challenges of achieving high precision, flexibility, and efficiency in complex microdevice assembly. We developed an automated microassembly system that integrates a decision tree for predicting operator intention with a leader-follo

Cited by 0SourceScholar
2025

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning

ICCV 2025poster

We propose SC-Captioner, a reinforcement learning framework that enables the self-correcting capability of image caption models. Our crucial technique lies in the design of the reward function to incentivize accurate caption corrections. Specifically, the predicted and reference captions are decompo…

2025

Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection

CVPR 2025poster

The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text pairs, enabling strong zero-shot classification and retrieval capabilities on various domains. However, CLIP's training remains computationally intensive…

2025

Vegetable Peeling: A Case Study in Constrained Dexterous Manipulation

ICRA 2025

Recent studies have made significant progress in addressing dexterous manipulation problems, particularly in inhand object reorientation. However, there are few existing works that explore the potential utilization of developed dexterous manipulation controllers for downstream tasks. In this study,

Cited by 17SourcecodeScholar
2024

$\textit{Bifr\"ost}$: 3D-Aware Image Compositing with Language Instructions

NeurIPS 2024poster

This paper introduces $\textit{Bifröst}$, a novel 3D-aware framework that is built upon diffusion models to perform instruction-based image composition. Previous methods concentrate on image compositing at the 2D level, which fall short in handling complex spatial relationships ($\textit{e.g.}$, occ…

Cited by 0SourcePDFScholar
2024

3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection

NeurIPS 2024poster

Transformer-based architectures have been proven successful in detecting 3D objects from point clouds. However, the quadratic complexity of the attention mechanism struggles to encode rich information as point cloud resolution increases. Recently, state space models (SSM) such as Mamba have gained g…

Cited by 0SourcePDFScholar
2024

Adaptive Integration of Partial Label Learning and Negative Learning for Enhanced Noisy Label Learning

AAAI 2024technical

There has been significant attention devoted to the effectiveness of various domains, such as semi-supervised learning, contrastive learning, and meta-learning, in enhancing the performance of methods for noisy label learning (NLL) tasks. However, most existing methods still depend on prior assumpti…

2024

Boosting Residual Networks with Group Knowledge

AAAI 2024technical

Recent research understands the residual networks from a new perspective of the implicit ensemble model. From this view, previous methods such as stochastic depth and stimulative training have further improved the performance of the residual network by sampling and training of its subnets. However,…

2024

DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism

ECCV 2024poster

"Explicit Caption Editing (ECE) — refining reference image captions through a sequence of explicit edit operations (, KEEP, DETELE) — has raised significant attention due to its explainable and human-like nature. After training with carefully designed reference and ground-truth caption pairs, state-…

Cited by 3SourcePDFScholar
2024

EMR-Merging: Tuning-Free High-Performance Model Merging

NeurIPS 2024spotlight

The success of pretrain-finetune paradigm brings about the release of numerous model weights. In this case, merging models finetuned on different tasks to enable a single model with multi-task capabilities is gaining increasing attention for its practicability. Existing model merging methods usually…

2024

Enhanced KPI Anomaly Detection: An Unsupervised Hybrid Model with Dynamic Threshold

ICASSP 2024accepted

Anomaly detection based on key performance indicator (KPI) is an important topic in the field of intelligent operation and maintenance. The problem of insufficient annotated samples is widespread in the industrial Internet, and it severely impairs the performance of data-driven anomaly detection. Pr…

Cited by 0SourceScholar
2024

FLDM-VTON: Faithful Latent Diffusion Model for Virtual Try-on

IJCAI 2024poster

Despite their impressive generative performance, latent diffusion model-based virtual try-on (VTON) methods lack faithfulness to crucial details of the clothes, such as style, pattern, and text. To alleviate these issues caused by the diffusion stochastic nature and latent supervision, we propose a…

Cited by 6SourcePDFScholar
2024

FNP: Fourier Neural Processes for Arbitrary-Resolution Data Assimilation

NeurIPS 2024poster

Data assimilation is a vital component in modern global medium-range weather forecasting systems to obtain the best estimation of the atmospheric state by combining the short-term forecast and observations. Recently, AI-based data assimilation approaches have attracted increasing attention for their…

2024

Foster Adaptivity and Balance in Learning with Noisy Labels

ECCV 2024poster

"Label noise is ubiquitous in real-world scenarios, posing a practical challenge to supervised models due to its effect in hurting the generalization performance of deep neural networks. Existing methods primarily employ the sample selection paradigm and usually rely on dataset-dependent prior knowl…

Cited by 6SourcePDFScholar
2024

LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding Reasoning and Planning

CVPR 2024poster

Recent progress in Large Multimodal Models (LMM) has opened up great possibilities for various applications in the field of human-machine interactions. However developing LMMs that can comprehend reason and plan in complex and diverse 3D environments remains a challenging topic especially considerin…

2024

Learning Multimodal Behaviors from Scratch with Diffusion Policy Gradient

NeurIPS 2024poster

Deep reinforcement learning (RL) algorithms typically parameterize the policy as a deep network that outputs either a deterministic action or a stochastic one modeled as a Gaussian distribution, hence restricting learning to a single behavioral mode. Meanwhile, diffusion models emerged as a powerful…

2024

Lifelong Robot Learning with Human Assisted Language Planners

ICRA 2024poster

Large Language Models (LLMs) have been shown to act like planners that can decompose high-level instructions into a sequence of executable instructions. However, current LLM-based planners are only able to operate with a fixed set of skills. We overcome this critical limitation and present a method…

Cited by 19SourceScholar
2024

MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer

CVPR 2024poster

Vision-Language Transformers (VLTs) have shown great success recently but are meanwhile accompanied by heavy computation costs where a major reason can be attributed to the large number of visual and language tokens. Existing token pruning research for compressing VLTs mainly follows a single-modali…

2024

MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization

COLING 2024main

Given the long textual product information and the product image, Multi-modal Product Summarization (MPS) aims to increase customers’ desire to purchase by highlighting product characteristics with a short textual summary. Existing MPS methods can produce promising results. Nevertheless, they still…

2024

MeshXL: Neural Coordinate Field for Generative 3D Foundation Models

NeurIPS 2024poster

The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately,…

2024

Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression

CVPR 2024poster

Recent Vision Transformer Compression (VTC) works mainly follow a two-stage scheme where the importance score of each model unit is first evaluated or preset in each submodule followed by the sparsity score evaluation according to the target sparsity constraint. Such a separate evaluation process in…

2024

PM-INR: Prior-Rich Multi-Modal Implicit Large-Scale Scene Neural Representation

AAAI 2024technical

Recent advancements in implicit neural representations have contributed to high-fidelity surface reconstruction and photorealistic novel view synthesis. However, with the expansion of the scene scale, such as block or city level, existing methods will encounter challenges because traditional samplin…

Cited by 2SourcePDFScholar
2024

ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation

ICLR 2024poster

Domain shifts such as sensor type changes and geographical situation variations are prevalent in Autonomous Driving (AD), which poses a challenge since AD model relying on the previous domain knowledge can be hardly directly deployed to a new domain without additional costs. In this paper, we provid…

2024

Reconciling Reality through Simulation: A Real-To-Sim-to-Real Approach for Robust Manipulation

RSS 2024poster

Imitation learning methods need significant human supervision to learn policies robust to changes in object poses, physical disturbances, and visual distractors. Reinforcement learning, on the other hand, can explore the environment autonomously to learn robust behaviors but may require impractical…

Cited by 55SourcePDFScholar
2024

S2HPruner: Soft-to-Hard Distillation Bridges the Discretization Gap in Pruning

NeurIPS 2024poster

Recently, differentiable mask pruning methods optimize the continuous relaxation architecture (soft network) as the proxy of the pruned discrete network (hard network) for superior sub-architecture search. However, due to the agnostic impact of the discretization process, the hard network struggles…

Cited by 0SourcePDFScholar
2024

Spear: Evaluate the Adversarial Robustness of Compressed Neural Models

IJCAI 2024poster

As Artificial Intelligence evolves, the neural models vulnerable to adversarial attacks may produce fatal results in critical applications. This paper mainly discusses the robustness of the compressed neural models facing adversarial attacks. A few studies discuss the interaction between model compr…

2024

The Joint Grid-Free DOA and Polarization Estimation Algorithm based on Atomic Norm Minimization

ICASSP 2024accepted

To address the issue of estimation accuracy degradation caused by off-grid in compressed sensing-based direction of arrival (DOA) estimation algorithms for polarized sensitive arrays, this paper proposes a joint estimation algorithm for two-dimensional DOA and polarization parameters based on the at…

Cited by 0SourceScholar
2024

Through the Real World Haze Scenes: Navigating the Synthetic-to-Real Gap in Challenging Image Dehazing

ICRA 2024poster

Dehazing real-world hazy images is challenging due to the complexity of natural haze, varying haze conditions, details preservation, and the risk of overexposure. Existing methods excel in synthetic hazy scenarios but struggle in the real world because they don’t use all available features. Classica…

Cited by 1SourceScholar
2024

Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy

NeurIPS 2024poster

Diffusion models have recently achieved great success in the synthesis of high-quality images and videos. However, the existing denoising techniques in diffusion models are commonly based on step-by-step noise predictions, which suffers from high computation cost, resulting in a prohibitive latency…

2024

VideoMAC: Video Masked Autoencoders Meet ConvNets

CVPR 2024poster

Recently the advancement of self-supervised learning techniques like masked autoencoders (MAE) has greatly influenced visual representation learning for images and videos. Nevertheless it is worth noting that the predominant approaches in existing masked image / video modeling rely excessively on re…

2023

A Large-Scale Outdoor Multi-Modal Dataset and Benchmark for Novel View Synthesis and Implicit Scene Reconstruction

ICCV 2023poster

Neural Radiance Fields (NeRF) has achieved impressive results in single object scene reconstruction and novel view synthesis, as demonstrated on many single modality and single object focused indoor scene datasets like DTU, BMVS, and NeRF Synthetic. However, the study of NeRF on large-scale outdoor…

Cited by 29PDFScholar
2023

AD-PT: Autonomous Driving Pre-Training with Large-scale Point Cloud Dataset

NeurIPS 2023poster

It is a long-term vision for Autonomous Driving (AD) community that the perception models can learn from a large-scale point cloud dataset, to obtain unified representations that can achieve promising results on different tasks or benchmarks. Previous works mainly focus on the self-supervised pre-tr…

2023

Adversarial Amendment is the Only Force Capable of Transforming an Enemy into a Friend

IJCAI 2023poster

Adversarial attack is commonly regarded as a huge threat to neural networks because of misleading behavior. This paper presents an opposite perspective: adversarial attacks can be harnessed to improve neural models if amended correctly. Unlike traditional adversarial defense or adversarial training…

2023

Bi3D: Bi-Domain Active Learning for Cross-Domain 3D Object Detection

CVPR 2023poster

Unsupervised Domain Adaptation (UDA) technique has been explored in 3D cross-domain tasks recently. Though preliminary progress has been made, the performance gap between the UDA-based 3D model and the supervised one trained with fully annotated target domain is still large. This motivates us to con…

2023

Boost Transformer-based Language Models with GPU-Friendly Sparsity and Quantization

ACL 2023findings

Along with the performance improvement in NLP domain, the sizes of transformer-based language models (TLM) are also dramatically increased. Some prior works intend to compress TLM models into more compact forms, but do not fully consider the hardware characters may not support the efficient executio…

2023

Boost Vision Transformer With GPU-Friendly Sparsity and Quantization

CVPR 2023poster

The transformer extends its success from the language to the vision domain. Because of the numerous stacked self-attention and cross-attention blocks in the transformer, which involve many high-dimensional tensor multiplication operations, the acceleration deployment of vision transformer on GPU har…

2023

Breadcrumbs to the Goal: Goal-Conditioned Exploration from Human-in-the-Loop Feedback

NeurIPS 2023poster

Exploration and reward specification are fundamental and intertwined challenges for reinforcement learning. Solving sequential decision making tasks with a non-trivial element of exploration requires either specifying carefully designed reward functions or relying on indiscriminate, novelty seeking…

2023

ConceptFusion: Open-set multimodal 3D mapping

RSS 2023poster

Building 3D maps of the environment is central to robot navigation, planning, and interaction with objects in a scene. Most existing approaches that integrate semantic concepts with 3D maps largely remain confined to the closed-set setting: they can only reason about a finite set of concepts, pre-de…

2023

DexDeform: Dexterous Deformable Object Manipulation with Human Demonstrations and Differentiable Physics

ICLR 2023poster

In this work, we aim to learn dexterous manipulation of deformable objects using multi-fingered hands. Reinforcement learning approaches for dexterous rigid object manipulation would struggle in this setting due to the complexity of physics interaction with deformable objects. At the same time, prev…

Cited by 22SourcePDFScholar
2023

End-to-End 3D Dense Captioning With Vote2Cap-DETR

CVPR 2023poster

3D dense captioning aims to generate multiple captions localized with their associated object regions. Existing methods follow a sophisticated "detect-then-describe" pipeline equipped with numerous hand-crafted components. However, these hand-crafted components would yield suboptimal performance giv…

2023

Executing Your Commands via Motion Diffusion in Latent Space

CVPR 2023poster

We study a challenging task, conditional human motion generation, which produces plausible human motion sequences according to various conditional inputs, such as action classes or textual descriptors. Since human motions are highly diverse and have a property of quite different distribution from co…

2023

Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation

NeurIPS 2023poster

We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone to producing inconsistent results with the conditions becaus…

2023

MotionGPT: Human Motion as a Foreign Language

NeurIPS 2023poster

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multimodal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to human language, often perc…

2023

PDF: Point Diffusion Implicit Function for Large-scale Scene Neural Representation

NeurIPS 2023poster

Recent advances in implicit neural representations have achieved impressive results by sampling and fusing individual points along sampling rays in the sampling space. However, due to the explosively growing sampling space, finely representing and synthesizing detailed textures remains a challenge f…

Cited by 5SourcePDFScholar
2023

Parallel $Q$-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel Simulation

ICML 2023poster

Reinforcement learning is time-consuming for complex tasks due to the need for large amounts of training data. Recent advances in GPU-based simulation, such as Isaac Gym, have sped up data collection thousands of times on a commodity GPU. Most prior works have used on-policy methods like PPO due to…

2023

Reasoning Makes Good Annotators : An Automatic Task-specific Rules Distilling Framework for Low-resource Relation Extraction

EMNLP 2023long findings

Relation extraction is often challenged by insufficient labeled data. Previous methods exploit knowledge from unlabeled data by generating pseudo labels in a self-training pipeline, which suffers a gradual drift problem. Logic rules, a transferable and explainable form of expert knowledge, have achi…

Cited by 0SourceScholar
2023

Robust Geometry-Preserving Depth Estimation Using Differentiable Rendering

ICCV 2023poster

In this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate for mix-dataset training, enhancing generalization across dive…

Cited by 6PDFScholar
2023

TactoFind: A Tactile Only System for Object Retrieval

ICRA 2023poster

We study the problem of object retrieval in scenarios where visual sensing is absent, object shapes are unknown beforehand and objects can move freely, like grabbing objects out of a drawer. Successful solutions require localizing free objects, identifying specific object instances, and then graspin…

Cited by 14SourceScholar
2023

Uni3D: A Unified Baseline for Multi-Dataset 3D Object Detection

CVPR 2023poster

Current 3D object detection models follow a single dataset-specific training and testing paradigm, which often faces a serious detection accuracy drop when they are directly deployed in another dataset. In this paper, we study the task of training a unified 3D detector from multiple datasets. We obs…

2022

Colorization for In Situ Marine Plankton Images

ECCV 2022poster

"Underwater imaging with red-NIR light illumination can avoid phototropic aggregation-induced observational deviation of marine plankton abundance under white light illumination, but this will lead to the loss of critical color information in the collected grayscale images, which is non-preferable t…

Cited by 2SourcePDFScholar
2022

Coordinates Are NOT Lonely - Codebook Prior Helps Implicit Neural 3D representations

NeurIPS 2022accept

Implicit neural 3D representation has achieved impressive results in surface or scene reconstruction and novel view synthesis, which typically uses the coordinate-based multi-layer perceptrons (MLPs) to learn a continuous scene representation. However, existing approaches, such as Neural Radiance Fi…

2022

ED2LM: Encoder-Decoder to Language Model for Faster Document Re-ranking Inference

ACL 2022findings

State-of-the-art neural models typically encode document-query pairs using cross-attention for re-ranking. To this end, models generally utilize an encoder-only (like BERT) paradigm or an encoder-decoder (like T5) approach. These paradigms, however, are not without flaws, i.e., running the model on…

Cited by 15SourcePDFScholar
2022

Efficient Tactile Simulation with Differentiability for Robotic Manipulation

CoRL 2022poster

Efficient simulation of tactile sensors can unlock new opportunities for learning tactile-based manipulation policies in simulation and then transferring the learned policy to real systems, but fast and reliable simulators for dense tactile normal and shear force fields are still under-explored. We…

Cited by 45SourceScholar
2022

Fast and Constrained Absent Keyphrase Generation by Prompt-Based Learning

AAAI 2022technical

Generating absent keyphrases, which do not appear in the input document, is challenging in the keyphrase prediction task. Most previous works treat the problem as an autoregressive sequence-to-sequence generation task, which demonstrates promising results for generating grammatically correct and flu…

2022

HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object Handovers

ICRA 2022poster

We introduce a new simulation benchmark “Han-doverSim” for human-to-robot object handovers. To simulate the giver's motion, we leverage a recent motion capture dataset of hand grasping of objects. We create training and evaluation environments for the receiver with standardized protocols and metrics…

Cited by 29SourcecodeScholar
2022

Pre-Trained Language Models for Interactive Decision-Making

NeurIPS 2022accept

Language model (LM) pre-training is useful in many language processing tasks. But can pre-trained LMs be further leveraged for more general machine learning problems? We propose an approach for using LMs to scaffold learning and generalization in general sequential decision-making problems. In this…

Cited by 229SourcePDFScholar
2022

Stimulative Training of Residual Networks: A Social Psychology Perspective of Loafing

NeurIPS 2022accept

Residual networks have shown great success and become indispensable in today’s deep models. In this work, we aim to re-investigate the training process of residual networks from a novel social psychology perspective of loafing, and further propose a new training strategy to strengthen the performanc…

2022

TopFormer: Token Pyramid Transformer for Mobile Semantic Segmentation

CVPR 2022poster

Although vision transformers (ViTs) have achieved great success in computer vision, the heavy computational cost hampers their applications to dense prediction tasks such as semantic segmentation on mobile devices. In this paper, we present a mobile-friendly architecture named Token Pyramid Vision T…

Cited by 304PDFcodeScholar
2022

Watermark Vaccine: Adversarial Attacks to Prevent Watermark Removal

ECCV 2022poster

"As a common security tool, visible watermarking has been widely applied to protect copyrights of digital images. However, recent works have shown that visible watermarks can be removed by DNNs without damaging their host images. Such watermark-removal techniques pose a great threat to the ownership…

2022

b-DARTS: Beta-Decay Regularization for Differentiable Architecture Search

CVPR 2022oral

Neural Architecture Search (NAS) has attracted increasingly more attention in recent years because of its capability to design deep neural network automatically. Among them, differential NAS approaches such as DARTS, have gained popularity for the search efficiency. However, they suffer from two mai…

Cited by 148PDFcodeScholar
2021

An End-to-End Differentiable Framework for Contact-Aware Robot Design

RSS 2021poster

The current dominant paradigm for robotic manipulation involves two separate stages: manipulator design and control. Because the robot's morphology and how it can be controlled are intimately linked; joint optimization of design and control can significantly improve performance. Existing methods for…

2021

CIL: Contrastive Instance Learning Framework for Distantly Supervised Relation Extraction

ACL 2021long

The journey of reducing noise from distant supervision (DS) generated training data has been started since the DS was first introduced into the relation extraction (RE) task. For the past decade, researchers apply the multi-instance learning (MIL) framework to find the most reliable feature from a b…

2021

Concept-Based Label Embedding via Dynamic Routing for Hierarchical Text Classification

ACL 2021long

Hierarchical Text Classification (HTC) is a challenging task that categorizes a textual description within a taxonomic hierarchy. Most of the existing methods focus on modeling the text. Recently, researchers attempt to model the class representations with some resources (e.g., external dictionaries…

2021

Deep Symmetric Network for Underexposed Image Enhancement With Recurrent Attentional Learning

ICCV 2021poster

Underexposed image enhancement is of importance in many research domains. In this paper, we take this problem as image feature transformation between the underexposed image and its paired enhanced version, and we propose a deep symmetric network for the issue. Our symmetric network adapts invertible…

Cited by 70PDFScholar
2021

EADNet: Efficient Asymmetric Dilated Network For Semantic Segmentation

ICASSP 2021accepted

Due to real-time image semantic segmentation needs on power constrained edge devices, there has been an increasing desire to design lightweight semantic segmentation neural network, to simultaneously reduce computational cost and increase inference speed. In this paper, we propose an efficient asymm…

Cited by 0SourceScholar
2021

Empower Distantly Supervised Relation Extraction with Collaborative Adversarial Training

AAAI 2021technical

With recent advances in distantly supervised (DS) relation extraction (RE), considerable attention is attracted to leverage multi-instance learning (MIL) to distill high-quality supervision from the noisy DS. Here, we go beyond label noise and identify the key bottleneck of DS-MIL to be its low data…

2021

Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation

CVPR 2021poster

Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of…

Cited by 248PDFcodeScholar
2019

End-to-end Change Detection Using a Symmetric Fully Convolutional Network for Landslide Mapping

ICASSP 2019accepted

In this paper, we propose a novel approach based on a symmetric fully convolutional network within pyramid pooling (FCN-PP) for landslide mapping (LM). The proposed approach has three advantages. Firstly, this approach is automatic and insensitive to noise because multivariate morphological reconstr…

Cited by 0SourceScholar
2019

Sim-to-(Multi)-Real: Transfer of Low-Level Robust Control Policies to Multiple Quadrotors

IROS 2019poster

Quadrotor stabilizing controllers often require careful, model-specific tuning for safe operation. We use reinforcement learning to train policies in simulation that transfer remarkably well to multiple different physical quadrotors. Our policies are low-level, i.e., we map the rotorcrafts' state di…

Cited by 145SourceScholar
2018

Hardware Conditioned Policies for Multi-Robot Transfer Learning

NeurIPS 2018poster

Deep reinforcement learning could be used to learn dexterous robotic policies but it is challenging to transfer them to new robots with vastly different hardware properties. It is also prohibitively expensive to learn a new policy from scratch for each robot hardware due to the high sample complexit…