← Search

Chang Liu

244 accepted papers

2026

AINav: Large Language Model-Based Adaptive Interactive Navigation

ICRA 2026poster

Robotic navigation in complex environments remains a critical research challenge. Traditional navigation focuses on optimal trajectory generation within free space, struggling in environments lacking viable paths to the goal, such as disaster zones or cluttered warehouses. To address this gap, we pr…

2026

Adaptive-Learngene: Continual Expansion and Task-Aware Selection of Learngenes for Dynamic Environments

AAAI 2026technical

Pre-trained Vision Transformer (ViT) models have achieved impressive performance across various computer vision tasks. However, most existing pre-trained models are built on fixed datasets and lack the flexibility to incorporate new pre-training data. When additional data becomes available, previous

Cited by 0SourcePDFScholar
2026

Anchor Frame Bridging for Coherent First-Last Frame Video Generation

ICLR 2026poster

First-last frame video generation has recently gained significant attention. It enables coherent motion generation between specified first and last frames. However, this approach suffers from semantic degradation in intermediate frames, causing scene distortion and subject deformation that undermine…

Cited by 0SourceScholar
2026

AuthSig: Safeguarding Scanned Signatures Against Unauthorized Reuse in Paperless Workflows

AAAI 2026technical

With the deepening trend of paperless workflows, signatures as a means of identity authentication are gradually shifting from traditional ink-on-paper to electronic formats. Despite the availability of dynamic pressure-sensitive and PKI-based digital signatures, static scanned signatures remain prev

Cited by 0SourcePDFScholar
2026

BiHiTo: Biomolecular Hierarchy-inspired Tokenization

AAAI 2026technical

Three-dimensional atomic arrangements of biomolecules are key to demystifying biological functions. The rapid expansion of accessible structural data, driven by advances in AI for science, highlights the critical challenge of efficiently modeling large-scale biomolecular structures, which are high-d

Cited by 0SourcePDFScholar
2026

BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration

ICLR 2026poster

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in subject-consistent video generation due to an inherent difficulty in pa…

Cited by 0SourceScholar
2026

Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics Assessment

AAAI 2026technical

The aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature—spanning visual perception, cognition, and emotion—poses fundamental challenges. Although aesthetic descriptions offer a viable representati

Cited by 2SourcePDFScholar
2026

ChaosNexus: A Foundation Model for ODE-based Chaotic System Forecasting with Hierarchical Multi-scale Awareness

ICML 2026poster

Foundation models have shown great promise in achieving zero-shot or few-shot forecasting for ODE-based chaotic systems via large-scale pretraining. However, existing architectures often fail to capture the multi-scale temporal structures and distinct spectral characteristics of chaotic dynamics. To…

Cited by 0SourceScholar
2026

CoT-VLNBench: A Benchmark for Visual Chain-of-Thought Reasoning in Vision-Language-Navigation Robots

AAAI 2026technical

Recent advances in vision language models (VLMs) have demonstrated remarkable potential in embodied navigation tasks. However, existing robot-centric datasets primarily focus on traditional 3D tasks such as perception and prediction, lacking adequate support for vision-language tasks. Vision-languag

Cited by 0SourcePDFScholar
2026

Comp-Attn: Present-and-Align Attention for Compositional Video Genneration

ICML 2026poster

In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main challenges are twofold: 1) Subject presence, where not all subjects can be presented in the video; 2) Inter-subject relat…

Cited by 0SourceScholar
2026

DISSR: DISENTANGLING SPEECH REPRESENTATION FOR DEGRADATION-PRIOR GUIDED CROSS-DOMAIN SPEECH RESTORATION

ICASSP 2026poster

Previous speech restoration (SR) primarily focuses on single-task speech restoration (SSR), which cannot address general speech restoration problems. Training specific SSR models for different distortions is time-consuming and lacks generality. In addition, most studies ignore the problem of model g…

Cited by 0SourcePDFScholar
2026

DRAMA: Next-Gen Dynamic Orchestration for Resilient Multi-Agent Ecosystems in Flux

CVPR 2026

Embodied Multi-Agent Systems have proven highly effective in addressing complex tasks through coordinated collaboration among heterogeneous agents. However, real-world environments and task specifications are inherently dynamic, exhibiting frequent changes, uncertainty, and variability. Despite thes

Cited by 0SourceScholar
2026

Decoupling Vision and Language: Codebook Anchored Visual Adaptation

CVPR 2026

Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders often underperform in domain-specific visual tasks such as medical image diagnosis or fine-grained classification, where representation errors can cascad

Cited by 0SourceScholar
2026

Diagnosing and Improving Diffusion Models by Estimating Optimal Loss Value

ICLR 2026poster

Diffusion models have achieved remarkable success in generative modeling. Despite more stable training, the loss of diffusion models is not indicative of absolute data-fitting quality, since its optimal value is typically not zero but unknown, leading to the confusion between large optimal loss and…

Cited by 0SourceScholar
2026

DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO

CVPR 2026

Reinforcement learning (RL), particularly GRPO, improves image generation quality significantly by comparing the relative performance of images generated within the same group. However, in the later stages of training, the model tends to produce homogenized outputs, lacking creativity and visual div

Cited by 0SourceScholar
2026

EmoDiffTalk: Emotion-aware Diffusion for Editable 3D Gaussian Talking Head

CVPR 2026

Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-modal control. This paper introduces a new editable 3D Gaussian talking head, i.e.

Cited by 0SourcecodeScholar
2026

FlexProtein: Joint Sequence and Structure Pretraining for Protein Modeling

ICLR 2026poster

Protein foundation models have advanced rapidly, with most approaches falling into two dominant paradigms. Sequence-only language models (e.g., ESM-2) capture sequence semantics at scale but lack structural grounding. MSA-based predictors (e.g., AlphaFold 2/3) achieve accurate folding by exploiting…

Cited by 0SourceScholar
2026

Focus, Align, and Sustain: Counteracting Gradient Dilution in Incremental Object Detection

ICML 2026poster

Adapting Detection Transformers to Incremental Object Detection (IOD) poses a systemic challenge, as set-based optimization is inherently destabilized by sequential learning. In this work, we identify Gradient Dilution as the root cause of performance degradation, wherein optimization signals requir…

Cited by 0SourceScholar
2026

FreeText: Training-Free Text Rendering via Attention Localization and Spectral Glyph Injection

ICML 2026poster

Large-scale text-to-image (T2I) diffusion models excel at open-domain synthesis but still struggle with precise text rendering, especially for multi-line layouts, dense typography, and long-tailed scripts such as Chinese. Prior solutions typically necessitate costly retraining or impose rigid extern…

Cited by 0SourceScholar
2026

GROUP RELATIVE POLICY OPTIMIZATION FOR TEXT-TO-SPEECH WITH LARGE LANGUAGE MODELS

ICASSP 2026oral

This paper proposes a GRPO-based approach to enhance the performance of large language model (LLM)-based text-to-speech (TTS) models by deriving rewards from an off-the-shelf automatic speech recognition (ASR) model. Compared to previous reinforcement learning methods for LLM-based TTS, our method r…

Cited by 0SourcePDFScholar
2026

GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception

ICLR 2026poster

The bird’s-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive multi-task perception. However, the discrete grid representation of BEV leads to significant detail loss and limits feature alignment…

Cited by 0SourceScholar
2026

Inheriting Generalizable Knowledge from LLMs to Diverse Vertical Tasks

ICLR 2026poster

Large language models (LLMs) have demonstrated remarkable generalization across diverse tasks, suggesting the existence of task-agnostic, generalizable knowledge encoded within them. However, how to systematically extract and evaluate this knowledge remains unexplored. In this work, we innovatively…

Cited by 0SourcecodeScholar
2026

LEMD: Latent Environment Extrapolation and Message Disentanglement for Dynamic Graph Under Distribution Shift

IJCAI 2026

Dynamic graph neural networks (DyGNNs) are widely used to model evolving interactions, but may fail under data distribution shift. Due to limited and unreliable interventions and insufficient disentanglement, the existing dynamic graph domain generalization approaches lead to suboptimal results. We

Cited by 0Scholar
2026

M2oE: Modular Mixture of Experts for Multi-Morphology Reinforcement Learning of Modular Robots

ICRA 2026poster

Modular robots offer a promising solution for building versatile and adaptable robotic systems. For instance, space exploration robots can be designed to reconfigure to meet diverse task demands across varying environments. However, training such systems by Reinforcement Learning (RL) remains challe…

Cited by 0codeScholar
2026

Monocular Normal Estimation via Shading Sequence Estimation

ICLR 2026oral

Monocular normal estimation aims to estimate normal map from a single RGB image of an object under arbitrary lighting. Existing methods rely on deep models to directly predict normal maps. However, they often suffer from 3D misalignment: while the estimated normal maps may appear to have an overall…

Cited by 0SourcecodeScholar
2026

ProAR: Probabilistic Autoregressive Modeling for Molecular Dynamics

AAAI 2026technical

Understanding the structural dynamics of biomolecules is crucial for uncovering biological functions. As molecular dynamics (MD) simulation data becomes more available, deep generative models have been developed to synthesize realistic MD trajectories. However, existing methods produce fixed-length

Cited by 0SourcePDFScholar
2026

Risk-Aware and Scalable Hierarchical Motion Planning for Large-Scale Robotic Swarms Via CVaR-Constrained MPC (I)

ICRA 2026poster

Motion planning for large-scale robotic swarms presents significant challenges in terms of scalability and safety assurance in cluttered environments. To address these issues, this manuscript proposes a Closed-loop hierarchical Risk-aware swarm mOtion planner using Conditional ValuE at Risk (C-ROVER…

Cited by 0Scholar
2026

S2D-Align: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation

AAAI 2026technical

Radiology Report Generation (RRG) aims to automatically generate diagnostic reports from radiology images. To achieve this, existing methods have leveraged the powerful cross-modal generation capabilities of Multimodal Large Language Models (MLLMs), primarily focusing on optimizing cross-modal align

Cited by 0SourcePDFScholar
2026

SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema

AAAI 2026technical

Although Large language Model (LLM)-powered information extraction (IE) systems have shown impressive capabilities, current fine-tuning paradigms face two major limitations: high training costs and difficulties in aligning with LLM preferences. To address these issues, we propose a novel universal

Cited by 0SourcePDFScholar
2026

SIGMA: An Agent-Based Modeling UAV Swarm Simulator for Swarm Intelligence Algorithms (I)

ICRA 2026poster

Swarm intelligence for uncrewed aerial vehicles (UAVs) significantly improves the success rate of executing intricate tasks using “distributed platforms and aggregated effects”. However, the high experimental costs and safety risks constrain its development. This paper introduces SIGMA (Swarm Intell…

Cited by 0Scholar
2026

Semantic and Terrain-Aware Trajectory Optimization for Uniform Coverage in Obstacle-Laden Environments

ICRA 2026poster

Achieving efficient and uniform coverage in obstacle-laden unknown environments is essential for au- tonomous robots in cleaning, inspection and agricultural op- erations. Unlike most existing approaches that prioritize path length and time optimality, we propose the SHIFT planner framework, which i…

Cited by 0Scholar
2026

StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References

CVPR 2026

Despite the advancements in diffusion-based image style transfer, existing methods are commonly limited by 1) semantic gap: the style reference could miss proper content semantics, causing uncontrollable stylization; 2) reliance on extra constraints (e.g., semantic masks) restricting applicability;

Cited by 0SourcecodeScholar
2026

SurgSync: Time-Synchronized Multi-Modal Data Collection Framework and Dataset for Surgical Robotics

ICRA 2026poster

Most existing robotic surgery systems adopt a human-in-the-loop paradigm, often with the surgeon directly teleoperating the robotic system. Adding intelligence to these robots would enable higher-level control, such as supervised autonomy or even full autonomy. However, artificial intelligence (AI) …

2026

The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path

ICML 2026poster

Large Language Models (LLMs) fundamentally suffer from representation collapse, a bottleneck that severely degrades performance in long contexts. We identify that existing approaches risk drifting into one of two pathological extremes: Homogenization Collapse (e.g., attention sinks causing rank defi…

Cited by 0SourceScholar
2026

Threat2Traffic: Multi-Agent Environment Synthesis for Malware Traffic Generation from Threat Intelligence

ICML 2026poster

Data-driven cybersecurity research is fundamentally constrained by the scarcity of labeled datasets, yet acquiring authentic, large-scale malware traffic remains bottlenecked by obsolescent public datasets, unscalable manual construction, and inflexible sandboxes that fail to satisfy the sample-spec…

Cited by 0SourcecodeScholar
2026

TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning

ICLR 2026poster

Temporal search aims to identify a minimal set of relevant frames from tens of thousands based on a given query, serving as a foundation for accurate long-form video understanding. Many existing works attempt to progressively narrow the search space. However, these approaches typically rely on a han…

Cited by 0SourcecodeScholar
2026

UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking

ICLR 2026poster

Recent advancements in LLM-based information-seeking agents have achieved record-breaking performance on established benchmarks. However, these agents remain heavily reliant on search-engine-indexed knowledge, leaving a critical blind spot: Unindexed Information Seeking (UIS). This paper identifies…

Cited by 0SourcecodeScholar
2026

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions

ICML 2026poster

Layer pruning efficiently reduces Large Language Model (LLM) computational costs but often triggers sudden performance collapse. Existing representation-based analyses struggle to explain this mechanism. We propose studying pruning through decision representation. Focusing on multiple-choice tasks, …

Cited by 0SourceScholar
2026

UniDef: Universal Defense Against Unauthorized Image Manipulation

CVPR 2026

Image protection against unauthorized diffusion-based editing has achieved encouraging progress. However, existing methods face two critical limitations: (1) They only disturb the denoising direction at local step, resulting in generated images still retaining original or edited semantics. (2) Their

Cited by 0SourceScholar
2026

Unifying Language-Action Understanding and Generation for Autonomous Driving

CVPR 2026

Vision-Language-Action (VLA) models are emerging as a promising paradigm for end-to-end autonomous driving, valued for their potential to leverage world knowledge and reason about complex driving scenes. However, existing methods suffer from two critical limitations: a persistent misalignment betwee

Cited by 0SourcecodeScholar
2026

WaveFormer: Frequency-Time Decoupled Vision Modeling with Wave Equation

AAAI 2026technical

Vision modeling has advanced rapidly with Transformers, whose attention mechanisms capture visual dependencies but lack a principled account of how semantic information propagates spatially. We revisit this problem from a wave-based perspective: feature maps are treated as spatial signals whose evol

Cited by 0SourcePDFScholar
2026

Wavelet-Driven 3D Anomaly Detection under Pose-Agnostic and Sparse-View

CVPR 2026

Pose-agnostic anomaly detection (PAD) achieves strong performance in localizing anomalies from arbitrary viewpoints when trained on densely sampled normal data. However, under sparse-view conditions, existing methods face two key challenges: (1) sparse observations lead to overfitting and geometric

Cited by 0SourceScholar
2026

Zero-Shot Exocentric Viewpoint-Robust Imitation Learning (VIL): Bridging Handheld Gripper and Exocentric Views

ICRA 2026poster

Recent advances in robot learning have motivated integrated pipelines that combine hardware for data collection with imitation learning algorithms. Existing data collection methods like leader–follower, VR/AR, and exoskeletons rely on costly hardware and exhibit limited scalability, while imitation …

Cited by 0codeScholar
2025

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training

CVPR 2025poster

In rapidly evolving field of vision-language models (VLMs), contrastive language-image pre-training (CLIP) has made significant strides, becoming foundation for various downstream tasks. However, relying on one-to-one (image, text) contrastive paradigm to learn alignment from large-scale messy web d…

2025

Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance Segmentation

AAAI 2025technical

Temporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video data, previous methods leverage the image-text pretraining model for recognizing object instances by separately aligning…

Cited by 0SourcePDFScholar
2025

Assessing Robustness of Multi-Modal Large Language Models in Image Classification through Hierarchical WordNet-Based Evaluation

ICASSP 2025accepted

The advancement of multi-modal large language models (MLLMs) has significantly enhanced their capability to process and understand diverse data types, integrating text, images, and other modalities. Despite their impressive performance, evaluating the robustness of these models remains challenging d…

Cited by 0SourceScholar
2025

Beyond Circuit Connections: A Non-Message Passing Graph Transformer Approach for Quantum Error Mitigation

ICLR 2025poster

Despite the progress in quantum computing, one major bottleneck against the practical utility is its susceptibility to noise, which frequently occurs in current quantum systems. Existing quantum error mitigation (QEM) methods either lack generality to noise and circuit types or fail to capture the g…

Cited by 2SourcePDFScholar
2025

BoRe-Depth: Self-Supervised Monocular Depth Estimation with Boundary Refinement for Embedded Systems

IROS 2025

Depth estimation is one of the key technologies for realizing 3D perception in unmanned systems. Monocular depth estimation has been widely researched because of its low-cost advantage, but the existing methods face the challenges of poor depth estimation performance and blurred object boundaries on

Cited by 1SourcecodeScholar
2025

Conditional-Balanced Adversarial Delta Tuning for Cross-Domain Implicit Discourse Relation Recognition

ICASSP 2025accepted

Implicit discourse relation recognition (IDRR) is faced with a domain dilemma. Recent studies have achieved breakthroughs in standard datasets, while they are not appropriate in domains with insufficient data, such as bio-medicine. In this paper, we treat this problem as a cross-domain IDRR task, wh…

Cited by 0SourceScholar
2025

Context-Aware Multi-Scale Polyp Segmentation Network

ICASSP 2025accepted

Colonoscopy is the gold standard for detecting colorectal lesions and is critical for early screening and prevention of colorectal cancer. However, accurate polyp segmentation remains a challenging task due to the diverse morphology, varying sizes and indistinct boundaries of polyps. To address thes…

Cited by 0SourceScholar
2025

Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation

ICML 2025poster

Efficient access to high-quality information is vital for online platforms. To promote more useful information, users not only create new content but also evaluate existing content, often through helpfulness voting. Although aggregated votes help service providers rank their user content, these vote…

Cited by 0SourcePDFScholar
2025

CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation

ICASSP 2025accepted

Voice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However, the ground truth of the converted speech does not exist in a non-parallel VC scenario, which induces the train-inference…

Cited by 0SourceScholar
2025

DCA: Dividing and Conquering Amnesia in Incremental Object Detection

AAAI 2025technical

Incremental object detection (IOD) aims to cultivate an object detector that can continuously localize and recognize novel classes while preserving its performance on previous classes. Existing methods achieve certain success by improving knowledge distillation and exemplar replay for transformer-ba…

2025

DCMKC: A Dual Consistency Matching Approach for Multi-hop Question Answering in LLMs

EMNLP 2025

Reasoning based on chains of thought (CoTs) enables large language models (LLMs) to solve problems by thinking step by step and becomes the mainstream solution for Question-Answering (QA) tasks. Knowledge graph (KG)-enhanced CoT technology helps correct factual errors or predict reasoning direction.

2025

De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks

ICML 2025poster

The rapid advancement of speech generation models has heightened privacy and security concerns related to voice cloning (VC). Recent studies have investigated disrupting unauthorized voice cloning by introducing adversarial perturbations. However, determined attackers can mitigate these protective p…

2025

DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning

EMNLP 2025

Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies

Cited by 0SourcePDFScholar
2025

DigitalLLaVA: Incorporating Digital Cognition Capability for Physical World Comprehension in Multimodal LLMs

AAAI 2025technical

Multimodal Large Language Models (MLLMs) have shown remarkable cognitive capabilities in various cross-modal tasks.However, existing MLLMs struggle with tasks that require physical digital cognition, such as accurately reading an electric meter or pressure gauge. This limitation significantly reduce…

Cited by 0SourcePDFScholar
2025

Do LLMs Know and Understand Domain Conceptual Knowledge?

EMNLP 2025

This paper focuses on the task of generating concept sememe trees to study whether Large Language Models (LLMs) can understand and generate domain conceptual knowledge. Concept sememe tree is a hierarchical structure that represents lexical meaning by combining sememes and their relationships.To thi

2025

E2Former: An Efficient and Equivariant Transformer with Linear-Scaling Tensor Products

NeurIPS 2025spotlight

Equivariant Graph Neural Networks (EGNNs) have demonstrated significant success in modeling microscale systems, including those in chemistry, biology and materials science. However, EGNNs face substantial computational challenges due to the high cost of constructing edge features via spherical tenso…

Cited by 0SourceScholar
2025

Efficient ANN-SNN Conversion with Error Compensation Learning

ICML 2025poster

Artificial neural networks (ANNs) have demonstrated outstanding performance in numerous tasks, but deployment in resource-constrained environments remains a challenge due to their high computational and memory requirements. Spiking neural networks (SNNs) operate through discrete spike events and off…

Cited by 0SourcePDFScholar
2025

Efficient and Scalable Density Functional Theory Hamiltonian Prediction through Adaptive Sparsity

ICML 2025poster

Hamiltonian matrix prediction is pivotal in computational chemistry, serving as the foundation for determining a wide range of molecular properties. While SE(3) equivariant graph neural networks have achieved remarkable success in this domain, their substantial computational cost—driven by high-orde…

2025

Enhancing the Scalability and Applicability of Kohn-Sham Hamiltonians for Molecular Systems

ICLR 2025spotlight

Density Functional Theory (DFT) is a pivotal method within quantum chemistry and materials science, with its core involving the construction and solution of the Kohn-Sham Hamiltonian. Despite its importance, the application of DFT is frequently limited by the substantial computational resources requ…

Cited by 0SourcePDFScholar
2025

Exploiting Diffusion Prior for Real-World Image Dehazing with Unpaired Training

AAAI 2025technical

Unpaired training has been verified as one of the most effective paradigms for real scene dehazing by learning from unpaired real-world hazy and clear images. Although numerous studies have been proposed, current methods demonstrate limited generalization for various real scenes due to limited featu…

2025

FedWMSAM: Fast and Flat Federated Learning via Weighted Momentum and Sharpness-Aware Minimization

NeurIPS 2025poster

In federated learning (FL), models must \emph{converge quickly} under tight communication budgets while \emph{generalizing} across non-IID client distributions. These twin requirements have naturally led to two widely used techniques: client/server \emph{momentum} to accelerate progress, and \emph{s…

Cited by 0SourcecodeScholar
2025

GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray Diffusion

ICCV 2025poster

Accurate surface reconstruction from unposed images is crucial for efficient 3D object or scene creation. However, it remains challenging, particularly for the joint camera pose estimation. Previous approaches have achieved impressive pose-free surface reconstruction results in dense-view settings,…

2025

Haptic-ACT: Bridging Human Intuition with Compliant Robotic Manipulation via Immersive VR

IROS 2025

Robotic manipulation is essential for the widespread adoption of robots in industrial and home settings and has long been a focus within the robotics community. Advances in artificial intelligence have introduced promising learning-based methods to address this challenge, with imitation learning eme

Cited by 8SourceScholar
2025

How Does Topology Bias Distort Message Passing in Graph Recommender? A Dirichlet Energy Perspective

NeurIPS 2025poster

Graph-based recommender systems have achieved remarkable effectiveness by modeling high-order interactions between users and items. However, such approaches are significantly undermined by popularity bias, which distorts the interaction graph’s structure—referred to as topology bias. This leads to o…

Cited by 0SourcecodeScholar
2025

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly

ICCV 2025poster

Numerous synthesized videos from generative models, especially human-centric ones that simulate realistic human actions, pose significant threats to human information security and authenticity. While progress has been made in binary forgery video detection, the lack of fine-grained understanding of…

Cited by 0SourcePDFScholar
2025

Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning

AAAI 2025technical

Reinforcement learning (RL) often encounters delayed and sparse feedback in real-world applications, even with only episodic rewards. Previous approaches have made some progress in reward redistribution for credit assignment but still face challenges, including training difficulties due to redundan…

2025

Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition

CVPR 2025poster

Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic elements. In scenarios with degraded visual quality, lingui…

2025

Long-term Intracortical Neural activity and Kinematics (LINK): An intracortical neural dataset for chronic brain-machine interfaces, neuroscience, and machine learning

NeurIPS 2025poster

Intracortical brain-machine interfaces (iBMIs) have enabled movement and speech in people living with paralysis by using neural data to decode behaviors in real-time. However, intracortical neural recordings exhibit significant instabilities over time, which poses problems for iBMIs, neuroscience, a…

Cited by 0SourcecodeScholar
2025

One Filters All: A Generalist Filter For State Estimation

NeurIPS 2025poster

Estimating hidden states in dynamical systems, also known as optimal filtering, is a long-standing problem in various fields of science and engineering. In this paper, we introduce a general filtering framework, $\textbf{LLM-Filter}$, which leverages large language models (LLMs) for state estimation…

Cited by 0SourceScholar
2025

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

ICML 2025poster

We introduce Orthus, a unified multimodal model that excels in generating interleaved images and text from mixed-modality inputs by simultaneously handling discrete text tokens and continuous image features under the \textbf{AR} modeling principle. The continuous treatment of visual signals minimize…

Cited by 8SourcePDFScholar
2025

PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution

CVPR 2025poster

Pre-trained video generation models hold great potential for generative video super-resolution (VSR). However, adapting them for full-size VSR, as most existing methods do, suffers from unnecessary intensive full-attention computation and fixed output resolution. To overcome these limitations, we ma…

Cited by 0SourcePDFScholar
2025

Query-Based and Unnoticeable Graph Injection Attack from Neighborhood Perspective

IJCAI 2025

The robustness of Graph Neural Networks (GNNs) has become an increasingly important topic due to their expanding range of applications. Various attack methods have been proposed to explore the vulnerabilities of GNNs, ranging from Graph Modification Attacks (GMA) to the more practical and flexible G

2025

Robust Heterogeneous Graph Classification for Molecular Property Prediction with Information Bottleneck

AAAI 2025technical

Heterogeneous Graph Neural Networks (HGNNs) have achieved state-of-the-art performance in classifying molecular graphs, capitalizing on their ability to capture rich semantics. However, HGNNs for molecule property prediction exhibit significant susceptibility to adversarial attacks—a challenge that…

Cited by 0SourcePDFScholar
2025

Robust State Estimation for Legged Robots With Dual Beta Kalman Filter

RA-L 2025

Existing state estimation algorithms for legged robots that rely on proprioceptive sensors often overlook foot slippage and leg deformation in the physical world, leading to large estimation errors. To address this limitation, we propose a comprehensive measurement model that accounts for both foot

Cited by 6SourceScholar
2025

ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment

ICCV 2025poster

Perpetual 3D scene generation aims to produce long-range and coherent 3D view sequences, which is applicable for long-term video synthesis and 3D scene reconstruction. Existing methods follow a "navigate-and-imagine" fashion and rely on outpainting for successive view expansion. However, the generat…

Cited by 0SourcePDFScholar
2025

Specifying What You Know or Not for Multi-Label Class-Incremental Learning

AAAI 2025technical

Existing class incremental learning is mainly designed for single-label classification task, which is ill-equipped for multi-label scenarios due to the inherent contradiction of learning objectives for samples with incomplete labels. We argue that the main challenge to overcome this contradiction in…

2025

Spherical Scissor-Like Reconfigurable Palm Design in Robotic Hands: Insights from Human Hand Functionality

IROS 2025

The human palm demonstrates spatial reconfigurability during the gripping process and forms a spherical grasping envelope. Based on these observations, this study designs a reconfigurable spherical palm that incorporates a spatial scissor mechanism, which only requires a single actuator to reshape t

Cited by 0SourceScholar
2025

Symmetry and Fusion Data Augmentation for Semi-Supervised Medical Segmentation

ICASSP 2025accepted

In semi-supervised medical image segmentation, appropriately merging labeled and unlabeled data before network training instead of using them separately can effectively reduce knowledge loss, mitigate distribution discrepancies and promote efficient knowledge transfer to unlabeled data. However, exi…

Cited by 0SourceScholar
2025

TAD-E2E: A Large-scale End-to-end Autonomous Driving Dataset

ICCV 2025poster

End-to-end autonomous driving technology has recently become a focal point of research and application in autonomous driving. State-of-the-art (SOTA) methods are often trained and evaluated on the NuScenes dataset. However, the NuScenes dataset, introduced in 2019 for 3D perception tasks, faces seve…

Cited by 0SourcePDFScholar
2025

Temporal-aware Query Routing for Real-time Video Instance Segmentation

ICCV 2025poster

With the rise of applications such as embodied intelligence, developing high real-time online video instance segmentation (VIS) has become increasingly important. However, through time profiling of the components in advanced online VIS architecture (i.e., transformer-based architecture), we find tha…

Cited by 0SourcePDFScholar
2025

Toward Robust Early Detection of Alzheimer's Disease via an Integrated Multimodal Learning Approach

ICASSP 2025accepted

Alzheimer’s Disease (AD) is a complex neurodegenerative disorder marked by memory loss, executive dysfunction, and personality changes. Early diagnosis is challenging due to subtle symptoms and varied presentations, often leading to misdiagnosis with traditional unimodal diagnostic methods due to th…

Cited by 0SourceScholar
2025

Tune-Your-Style: Intensity-tunable 3D Style Transfer with Gaussian Splatting

ICCV 2025poster

3D style transfer refers to the artistic stylization of 3D assets based on reference style images. Recently, 3DGS-based stylization methods have drawn considerable attention, primarily due to their markedly enhanced training and rendering speeds. However, a vital challenge for 3D style transfer is t…

2025

Wavelet-Driven Masked Image Modeling: A Path to Efficient Visual Representation

AAAI 2025technical

Masked Image Modeling (MIM) has garnered significant attention in self-supervised learning, thanks to its impressive capacity to learn scalable visual representations tailored for downstream tasks. However, images inherently contain abundant redundant information, leading the pixel-based MIM reconst…

Cited by 0SourcePDFScholar
2025

When Schrodinger Bridge Meets Real-World Image Dehazing with Unpaired Training

ICCV 2025poster

Recent advancements in unpaired dehazing, particularly those using GANs, show promising performance in processing real-world hazy images. However, these methods tend to face limitations due to the generator's limited transport mapping capability, which hinders the full exploitation of their effectiv…

Cited by 0SourcePDFScholar
2025

iSegMan: Interactive Segment-and-Manipulate 3D Gaussians

CVPR 2025poster

The efficient rendering and explicit nature of 3DGS promote the advancement of 3D scene manipulation.However, existing methods typically encounter challenges in controlling the manipulation region and are unable to furnish the user with interactive feedback, which inevitably leads to unexpected resu…

Cited by 0SourcePDFScholar
2024

3D Point Cloud Semantic Segmentation Based on Diffusion Model

ICASSP 2024accepted

Point cloud segmentation plays a crucial role in extracting unique attributes and separating various objects, thereby enabling semantic comprehension and analysis. In this paper, we introduce a novel point cloud segmentation approach based on Diffusion Probabilistic Network (DDPM). The proposed mode…

Cited by 0SourceScholar
2024

ACM-MILP: Adaptive Constraint Modification via Grouping and Selection for Hardness-Preserving MILP Instance Generation

ICML 2024spotlight

Data plays a pivotal role in the development of both classic and learning-based methods for Mixed-Integer Linear Programming (MILP). However, the scarcity of data in real-world applications underscores the necessity for MILP instance generation methods. Currently, these methods primarily rely on ite…

2024

ADAM: Dense Retrieval Distillation with Adaptive Dark Examples

ACL 2024findings

To improve the performance of the dual-encoder retriever, one effective approach is knowledge distillation from the cross-encoder ranker. Existing works prepare training instances by pairing each query with one positive and a batch of negatives. However, most hard negatives mined by advanced dense r…

Cited by 5SourcePDFScholar
2024

ASPIRe: An Informative Trajectory Planner with Mutual Information Approximation for Target Search and Tracking

ICRA 2024poster

This paper proposes an informative trajectory planning approach, namely, adaptive particle filter tree with sigma point-based mutual information reward approximation (ASPIRe), for mobile target search and tracking (SAT) in cluttered environments with limited sensing field of view. We develop a novel…

Cited by 4SourceScholar
2024

Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise

ICRA 2024poster

3D articulated objects are inherently challenging for manipulation due to the varied geometries and intricate functionalities associated with articulated objects. Point-level affordance, which predicts the per-point actionable score and thus proposes the best point to interact with, has demonstrated…

Cited by 16SourceScholar
2024

Aspect-based Sentiment Analysis with Context Denoising

NAACL 2024findings

Given a sentence and a particular aspect term, aspect-based sentiment analysis (ABSA) aims to predict the sentiment polarity towards this aspect term, which provides fine-grained analysis on sentiment understanding and it has attracted much attention in recent years. In order to achieve a good perfo…

2024

Assemblage: Automatic Binary Dataset Construction for Machine Learning

NeurIPS 2024poster

Binary code is pervasive, and binary analysis is a key task in reverse engineering, malware classification, and vulnerability discovery. Unfortunately, while there exist large corpuses of malicious binaries, obtaining high-quality corpuses of benign binaries for modern systems has proven challenging…

2024

Boosting LLM Agents with Recursive Contemplation for Effective Deception Handling

ACL 2024findings

Recent advances in large language models (LLMs) have led to significant success in using LLMs as agents. Nevertheless, a common assumption that LLMs always process honest information neglects the widespread deceptive or misleading content in human and AI-generated material. This oversight might expo…

2024

Bootstrapping Large Language Models for Radiology Report Generation

AAAI 2024technical

Radiology report generation (RRG) aims to automatically generate a free-text description from a specific clinical radiograph, e.g., chest X-Ray images. Existing approaches tend to perform RRG with specific models trained on the public yet limited data from scratch, where they often lead to inferior…

2024

Continuous Relational Diffusion Driven Topic Model with Multi-grained Text for Microblog

COLING 2024main

Topic model is a statistical model that leverages unsupervised learning to mine hidden topics in document collections. The data sparsity and colloquialism of social texts make it difficult to accurately mine the topics. Traditional methods assume that there are only 0/1-state relationships between t…

Cited by 0SourcePDFScholar
2024

DOZE: A Dataset for Open-Vocabulary Zero-Shot Object Navigation in Dynamic Environments

RA-L 2024

Zero-Shot Object Navigation (ZSON) requires agents to autonomously locate and approach unseen objects in unfamiliar environments and has emerged as a particularly challenging task within the domain of Embodied AI. Existing datasets for developing ZSON algorithms lack consideration of dynamic obstacl

Cited by 8SourceScholar
2024

FaceChain-SuDe: Building Derived Class to Inherit Category Attributes for One-shot Subject-Driven Generation

CVPR 2024poster

Recently subject-driven generation has garnered significant interest due to its ability to personalize text-to-image generation. Typical works focus on learning the new subject's private attributes. However an important fact has not been taken seriously that a subject is not an isolated new concept…

2024

FreestyleRet: Retrieving Images from Style-Diversified Queries

ECCV 2024poster

"Image Retrieval aims to retrieve corresponding images based on a given query. In application scenarios, users intend to express their retrieval intent through various query styles. However, current retrieval tasks predominantly focus on text-query retrieval exploration, leading to limited retrieval…

2024

Global and Local Hierarchical Prompt Tuning Framework for Multi-level Implicit Discourse Relation Recognition

COLING 2024main

Multi-level implicit discourse relation recognition (MIDRR) is a challenging task to recognize the hierarchical discourse relations between the arguments with the absence of connectives. Recent methods tend to incorporate the static hierarchical structure containing all senses (defined as global hie…

Cited by 1SourcePDFScholar
2024

GraCo: Granularity-Controllable Interactive Segmentation

CVPR 2024highlight

Interactive Segmentation (IS) segments specific objects or parts in the image according to user input. Current IS pipelines fall into two categories: single-granularity output and multi-granularity output. The latter aims to alleviate the spatial ambiguity present in the former. However the multi-gr…

2024

Infusing Self-Consistency into Density Functional Theory Hamiltonian Prediction via Deep Equilibrium Models

NeurIPS 2024poster

In this study, we introduce a unified neural network architecture, the Deep Equilibrium Density Functional Theory Hamiltonian (DEQH) model, which incorporates Deep Equilibrium Models (DEQs) for predicting Density Functional Theory (DFT) Hamiltonians. The DEQH model inherently captures the self-consi…

2024

Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models

CVPR 2024poster

Generative models have recently exhibited exceptional capabilities in text-to-image generation but still struggle to generate image sequences coherently. In this work we focus on a novel yet challenging task of generating a coherent image sequence based on a given storyline denoted as open-ended vis…

2024

Is Function Similarity Over-Engineered? Building a Benchmark

NeurIPS 2024poster

Binary analysis is a core component of many critical security tasks, including reverse engineering, malware analysis, and vulnerability detection. Manual analysis is often time-consuming, but identifying commonly-used or previously-seen functions can reduce the time it takes to understand a new file…

2024

L2P-MIP: Learning to Presolve for Mixed Integer Programming

ICLR 2024poster

Modern solvers for solving mixed integer programming (MIP) often rely on the branch-and-bound (B&B) algorithm which could be of high time complexity, and presolving techniques are well designed to simplify the instance as pre-processing before B&B. However, such presolvers in existing literature or…

Cited by 6SourcePDFScholar
2024

LLM-Empowered State Representation for Reinforcement Learning

ICML 2024poster

Conventional state representations in reinforcement learning often omit critical task-related details, presenting a significant challenge for value networks in establishing accurate mappings from states to task rewards. Traditional methods typically depend on extensive sample learning to enrich stat…

2024

Learning Spatially Collaged Fourier Bases for Implicit Neural Representation

AAAI 2024technical

Existing approaches to Implicit Neural Representation (INR) can be interpreted as a global scene representation via a linear combination of Fourier bases of different frequencies. However, such universal basis functions can limit the representation capability in local regions where a specific compon…

Cited by 5SourcePDFScholar
2024

Learning from the Web: Language Drives Weakly-Supervised Incremental Learning for Semantic Segmentation

ECCV 2024poster

"Current weakly-supervised incremental learning for semantic segmentation (WILSS) approaches only consider replacing pixel-level annotations with image-level labels, while the training images are still from well-designed datasets. In this work, we argue that widely available web images can also be c…

2024

Local and Global Feature Adaptive Adjustment Network for Remote Sensing Image Scene Classification

ICASSP 2024accepted

Convolutional neural network (CNN)-based methods have been extensively used for remote sensing scene classification (RSSC) and have obtained remarkable classification results. However, its limitations in extracting global features have hindered further improvement. Transformers can directly capture…

Cited by 0SourceScholar
2024

MHGRL: An Effective Representation Learning Model for Electronic Health Records

COLING 2024main

Electronic health records (EHRs) serve as a digital repository storing comprehensive medical information about patients. Representation learning for EHRs plays a crucial role in healthcare applications. In this paper, we propose a Multimodal Heterogeneous Graph-enhanced Representation Learning, deno…

2024

Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning

ICML 2024poster

Although pre-trained models such as Contrastive Language-Image Pre-Training (CLIP) show impressive generalization results, their robustness is still limited under Out-of-Distribution (OOD) scenarios. Instead of undesirably leveraging human annotation as commonly done, it is possible to leverage the…

Cited by 15SourcePDFScholar
2024

MatchTime: Towards Automatic Soccer Game Commentary Generation

EMNLP 2024main

Soccer is a globally popular sport with a vast audience, in this paper, we consider constructing an automatic soccer game commentary model to improve the audiences’ viewing experience. In general, we make the following contributions: *First*, observing the prevalent video-text misalignment in existi…

2024

MuST: Robust Image Watermarking for Multi-Source Tracing

AAAI 2024technical

In recent years, with the popularity of social media applications, massive digital images are available online, which brings great convenience to image recreation. However, the use of unauthorized image materials in multi-source composite images is still inadequately regulated, which may cause signi…

2024

MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models

NeurIPS 2024poster

Despite the superior capabilities of Multimodal Large Language Models (MLLMs) across diverse tasks, they still face significant trustworthiness challenges. Yet, current literature on the assessment of trustworthy MLLMs remains limited, lacking a holistic evaluation to offer thorough insights into fu…

Cited by 5SourcecodeScholar
2024

ParCo: Part-Coordinating Text-to-Motion Synthesis

ECCV 2024poster

"We study a challenging task: text-to-motion synthesis, aiming to generate motions that align with textual descriptions and exhibit coordinated movements. Currently, the part-based methods introduce part partition into the motion synthesis process to achieve finer-grained generation. However, these…

2024

Parallel Vertex Diffusion for Unified Visual Grounding

AAAI 2024technical

Unified visual grounding (UVG) capitalizes on a wealth of task-related knowledge across various grounding tasks via one-shot training, which curtails retraining costs and task-specific architecture design efforts. Vertex generation-based UVG methods achieve this versatility by unified modeling objec…

Cited by 27SourcePDFScholar
2024

Physical Consistency Bridges Heterogeneous Data in Molecular Multi-Task Learning

NeurIPS 2024poster

In recent years, machine learning has demonstrated impressive capability in handling molecular science tasks. To support various molecular properties at scale, machine learning models are trained in the multi-task learning paradigm. Nevertheless, data of different molecular properties are often not…

Cited by 1SourcePDFScholar
2024

Quantum Interference Model for Semantic Biases of Glosses in Word Sense Disambiguation

AAAI 2024technical

Word Sense Disambiguation (WSD) aims to determine the meaning of the target word according to the given context. Currently, a single representation enhanced by glosses from different dictionaries or languages is used to characterize each word sense. By analyzing the similarity between glosses of the…

Cited by 7SourcePDFScholar
2024

RGBGrasp: Image-Based Object Grasping by Capturing Multiple Views During Robot arm Movement With Neural Radiance Fields

RA-L 2024

Robotic research encounters a significant hurdle when it comes to the intricate task of grasping objects that come in various shapes, materials, and textures. Unlike many prior investigations that heavily leaned on specialized point-cloud cameras or abundant RGB visual data to gather 3D insights for

Cited by 25SourceScholar
2024

Referring Image Editing: Object-level Image Editing via Referring Expressions

CVPR 2024poster

Significant advancements have been made in image editing with the recent advance of the Diffusion model. However most of the current methods primarily focus on global or subject-level modifications and often face limitations when it comes to editing specific objects when there are other objects coex…

Cited by 14SourcePDFScholar
2024

Representation Degeneration Problem in Prompt-based Models for Natural Language Understanding

COLING 2024main

Prompt-based fine-tuning (PF), by aligning with the training objective of pre-trained language models (PLMs), has shown improved performance on many few-shot natural language understanding (NLU) benchmarks. However, the word embedding space of PLMs exhibits anisotropy, which is called the representa…

2024

Risk-Aware Non-Myopic Motion Planner for Large-Scale Robotic Swarm Using CVaR Constraints

IROS 2024poster

Swarm robotics has garnered significant attention due to its ability to accomplish elaborate and synchronized tasks. Existing methodologies for motion planning of swarm robotic systems mainly encounter difficulties in scalability and safety guarantee. To address these limitations, we propose a Risk-…

Cited by 1SourceScholar
2024

Self-Consistency Training for Density-Functional-Theory Hamiltonian Prediction

ICML 2024poster

Predicting the mean-field Hamiltonian matrix in density functional theory is a fundamental formulation to leverage machine learning for solving molecular science problems. Yet, its applicability is limited by insufficient labeled data for training. In this work, we highlight that Hamiltonian predict…

Cited by 5SourcePDFScholar
2024

Similarity Knowledge Distillation with Calibrated Mask

ICASSP 2024accepted

In this paper, we propose a novel and efficient method for knowledge distillation, which is structurally simple and requires negligible computation overhead. Our method includes three modules. The first module is the calibrated mask, which avoids the teacher model’s incorrect representation to distu…

Cited by 0SourceScholar
2024

SwarmPRM: Probabilistic Roadmap Motion Planning for Large-Scale Swarm Robotic Systems

IROS 2024poster

Large-scale swarm robotic systems consisting of numerous cooperative agents show considerable promise for performing autonomous tasks across various sectors. Nonetheless, traditional motion planning approaches often face a trade-off between scalability and solution quality due to the exponential gro…

Cited by 1SourceScholar
2024

Towards General Loop Invariant Generation: A Benchmark of Programs with Memory Manipulation

NeurIPS 2024poster

Program verification is vital for ensuring software reliability, especially in the context of increasingly complex systems. Loop invariants, remaining true before and after each iteration of loops, are crucial for this verification process. Traditional provers and machine learning based methods for…

Cited by 2SourcePDFScholar
2024

Transformer-Empowered Multi-Modal Item Embedding for Enhanced Image Search in E-commerce

AAAI 2024technical

Over the past decade, significant advances have been made in the field of image search for e-commerce applications. Traditional image-to-image retrieval models, which focus solely on image details such as texture, tend to overlook useful semantic information contained within the images. As a result,…

Cited by 1SourcePDFScholar
2024

UNIDEAL: Curriculum Knowledge Distillation Federated Learning

ICASSP 2024accepted

Federated Learning (FL) has emerged as a promising approach to enable collaborative learning among multiple clients while preserving data privacy. However, cross-domain FL tasks, where clients possess data from different domains or distributions, remain a challenging problem due to the inherent hete…

Cited by 0SourceScholar
2024

VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model

ICML 2024poster

In the realm of household robotics, the Zero-Shot Object Navigation (ZSON) task empowers agents to adeptly traverse unfamiliar environments and locate objects from novel categories without prior explicit training. This paper introduces VoroNav, a novel semantic exploration framework that proposes th…

2023

ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

CVPR 2023highlight

Recently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship amo…

Cited by 50SourcePDFScholar
2023

Ambiguity-Resistant Semi-Supervised Learning for Dense Object Detection

CVPR 2023poster

With basic Semi-Supervised Object Detection (SSOD) techniques, one-stage detectors generally obtain limited promotions compared with two-stage clusters. We experimentally find that the root lies in two kinds of ambiguities: (1) Selection ambiguity that selected pseudo labels are less accurate, since…

2023

Attend, Select and Eliminate: Accelerating Multi-turn Response Selection with Dual-attention-based Content Elimination

ACL 2023findings

Although the incorporation of pre-trained language models (PLMs) significantly pushes the research frontier of multi-turn response selection, it brings a new issue of heavy computation costs. To alleviate this problem and make the PLM-based response selection model both effective and efficient, we p…

Cited by 1SourcePDFScholar
2023

AutoStegaFont: Synthesizing Vector Fonts for Hiding Information in Documents

AAAI 2023technical

Hiding information in text documents has been a hot topic recently, with the most typical schemes of utilizing fonts. By constructing several fonts with similar appearances, information can be effectively represented and embedded in documents. However, due to the unstructured characteristic, font ve…

Cited by 3SourcePDFScholar
2023

CORE: Cooperative Training of Retriever-Reranker for Effective Dialogue Response Selection

ACL 2023long

Establishing retrieval-based dialogue systems that can select appropriate responses from the pre-built index has gained increasing attention. Recent common practice is to construct a two-stage pipeline with a fast retriever (e.g., bi-encoder) for first-stage recall followed by a smart response reran…

Cited by 5SourcePDFScholar
2023

Complementary Attention for Multi-Agent Reinforcement Learning

ICML 2023poster

In cooperative multi-agent reinforcement learning, centralized training with decentralized execution (CTDE) shows great promise for a trade-off between independent Q-learning and joint action learning. However, vanilla CTDE methods assumed a fixed number of agents could hardly adapt to real-world sc…

Cited by 10SourcePDFScholar
2023

Context-Aware Transformer for 3D Point Cloud Automatic Annotation

AAAI 2023technical

3D automatic annotation has received increased attention since manually annotating 3D point clouds is laborious. However, existing methods are usually complicated, e.g., pipelined training for 3D foreground/background segmentation, cylindrical object proposals, and point completion. Furthermore, the…

Cited by 5SourcePDFScholar
2023

DeAR: A Deep-Learning-Based Audio Re-recording Resilient Watermarking

AAAI 2023technical

Audio watermarking is widely used for leaking source tracing. The robustness of the watermark determines the traceability of the algorithm. With the development of digital technology, audio re-recording (AR) has become an efficient and covert means to steal secrets. AR process could drastically dest…

Cited by 43SourcePDFScholar
2023

DiffusionRet: Generative Text-Video Retrieval with Diffusion Model

ICCV 2023poster

Existing text-video retrieval solutions are, in essence, discriminant models focused on maximizing the conditional likelihood, i.e., p(candidates|query). While straightforward, this de facto paradigm overlooks the underlying data distribution p(query), which makes it challenging to identify out-of-d…

Cited by 72PDFcodeScholar
2023

Discover and Align Taxonomic Context Priors for Open-world Semi-Supervised Learning

NeurIPS 2023poster

Open-world Semi-Supervised Learning (OSSL) is a realistic and challenging task, aiming to classify unlabeled samples from both seen and novel classes using partially labeled samples from the seen classes. Previous works typically explore the relationship of samples as priors on the pre-defined sing…

2023

Discovering Informative and Robust Positives for Video Domain Adaptation

ICLR 2023poster

Unsupervised domain adaptation for video recognition is challenging where the domain shift includes both spatial variations and temporal dynamics. Previous works have focused on exploring contrastive learning for cross-domain alignment. However, limited variations in intra-domain positives, false cr…

Cited by 9SourcePDFScholar
2023

Fuzzy Positive Learning for Semi-Supervised Semantic Segmentation

CVPR 2023poster

Semi-supervised learning (SSL) essentially pursues class boundary exploration with less dependence on human annotations. Although typical attempts focus on ameliorating the inevitable error-prone pseudo-labeling, we think differently and resort to exhausting informative semantics from multiple proba…

2023

Gradually Excavating External Knowledge for Implicit Complex Question Answering

EMNLP 2023long findings

Recently, large language models (LLMs) have gained much attention for the emergence of human-comparable capabilities and huge potential. However, for open-domain implicit question-answering problems, LLMs may not be the ultimate solution due to the reasons of: 1) uncovered or out-of-date domain know…

Cited by 0SourceScholar
2023

ILSGAN: Independent Layer Synthesis for Unsupervised Foreground-Background Segmentation

AAAI 2023technical

Unsupervised foreground-background segmentation aims at extracting salient objects from cluttered backgrounds, where Generative Adversarial Network (GAN) approaches, especially layered GANs, show great promise. However, without human annotations, they are typically prone to produce foreground and ba…

2023

LaPE: Layer-adaptive Position Embedding for Vision Transformers with Independent Layer Normalization

ICCV 2023poster

Position information is critical for Vision Transformers (VTs) due to the permutation-invariance of self-attention operations. A typical way to introduce position information is adding the absolute Position Embedding (PE) to patch embedding before entering VTs. However, this approach operates the sa…

Cited by 10PDFcodeScholar
2023

Length-Adaptive Distillation: Customizing Small Language Model for Dynamic Token Pruning

EMNLP 2023long findings

Pre-trained language models greatly improve the performance of various tasks but at a cost of high computation overhead. To facilitate practical applications, there are mainly two lines of research to accelerate model inference: model compression and dynamic computation (e.g., dynamic token pruning)…

Cited by 0SourceScholar
2023

MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

ICCV 2023poster

Video object segmentation (VOS) aims at segmenting a particular object throughout the entire video clip sequence. The state-of-the-art VOS methods have achieved excellent performance (e.g., 90+% J&F) on existing datasets. However, since the target objects in these existing datasets are usually relat…

Cited by 148PDFcodeScholar
2023

MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

ICCV 2023poster

This paper strives for motion expressions guided video segmentation, which focuses on segmenting objects in video content based on a sentence describing the motion of the objects. Existing referring video object datasets typically focus on salient objects and use language expressions that contain ex…

Cited by 110PDFcodeScholar
2023

More than Classification: A Unified Framework for Event Temporal Relation Extraction

ACL 2023long

Event temporal relation extraction (ETRE) is usually formulated as a multi-label classification task, where each type of relation is simply treated as a one-hot label. This formulation ignores the meaning of relations and wipes out their intrinsic dependency. After examining the relation definitions…

2023

Multi-granularity Interaction Simulation for Unsupervised Interactive Segmentation

ICCV 2023poster

Interactive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and medical image analysis. Typically, massive and expansive pixel-level annotations are spent to train deep models by object-o…

Cited by 10PDFScholar
2023

Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation

CVPR 2023poster

Weakly supervised semantic segmentation is typically inspired by class activation maps, which serve as pseudo masks with class-discriminative regions highlighted. Although tremendous efforts have been made to recall precise and complete locations for each class, existing methods still commonly suffe…

2023

Out-of-Distributed Semantic Pruning for Robust Semi-Supervised Learning

CVPR 2023poster

Recent advances in robust semi-supervised learning (SSL) typical filters out-of-distribution (OOD) information at the sample level. We argue that an overlooked problem of robust SSL is its corrupted information on semantic level, practically limiting the development of the field. In this paper, we t…

2023

Revocable Deep Reinforcement Learning with Affinity Regularization for Outlier-Robust Graph Matching

ICLR 2023poster

Graph matching (GM) has been a building block in various areas including computer vision and pattern recognition. Despite recent impressive progress, existing deep GM methods often have obvious difficulty in handling outliers, which are ubiquitous in practice. We propose a deep reinforcement learnin…

Cited by 11SourcePDFScholar
2023

TG-VQA: Ternary Game of Video Question Answering

IJCAI 2023poster

Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions, i.e., annotations or priors, current contrastive learning-based VideoQA methods remains challenging to perform fine-gr…

Cited by 17SourcePDFScholar
2023

THMA: Tencent HD Map AI System for Creating HD Map Annotations

AAAI 2023technical

Nowadays, autonomous vehicle technology is becoming more and more mature. Critical to progress and safety, high-definition (HD) maps, a type of centimeter-level map collected using a laser sensor, provide accurate descriptions of the surrounding environment. The key challenge of HD map production is…

Cited by 13SourcePDFScholar
2023

Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment

IJCAI 2023poster

Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they fail to leverage the heterogeneous concepts in data. In this…

2023

TopoSeg: Topology-Aware Nuclear Instance Segmentation

ICCV 2023poster

Nuclear instance segmentation has been critical for pathology image analysis in medical science, e.g., cancer diagnosis. Current methods typically adopt pixel-wise optimization for nuclei boundary exploration, where rich structural information could be lost for subsequent quantitative morphology ass…

Cited by 25PDFcodeScholar
2023

Towards Effective Adversarial Textured 3D Meshes on Physical Face Recognition

CVPR 2023highlight

Face recognition is a prevailing authentication solution in numerous biometric applications. Physical adversarial attacks, as an important surrogate, can identify the weaknesses of face recognition systems and evaluate their robustness before deployed. However, most existing physical attacks are eit…

2023

Towards Real-World Burst Image Super-Resolution: Benchmark and Method

ICCV 2023poster

Despite substantial advances, single-image super-resolution (SISR) is always in a dilemma to reconstruct high-quality images with limited information from one input image, especially in realistic scenarios. In this paper, we establish a large-scale real-world burst super-resolution dataset, i.e., Re…

Cited by 34PDFcodeScholar
2023

UATVR: Uncertainty-Adaptive Text-Video Retrieval

ICCV 2023poster

With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer text-video pairs to the same embedding space and craft cross-mod…

Cited by 64PDFcodeScholar
2023

Video-Text As Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning

CVPR 2023highlight

Contrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined video-text pairs. To clarify this coarse-grained global interaction and move a step further, we have to encounter challe…

2023

WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image Segmentation

IJCAI 2023poster

The top-down and bottom-up methods are two mainstreams of referring segmentation, while both methods have their own intrinsic weaknesses. Top-down methods are chiefly disturbed by Polar Negative (PN) errors owing to the lack of fine-grained cross-modal alignment. Bottom-up methods are mainly perturb…

Cited by 4SourcePDFScholar
2022

Deep Neural Network Fusion via Graph Matching with Applications to Model Ensemble and Federated Learning

ICML 2022spotlight

Model fusion without accessing training data in machine learning has attracted increasing interest due to the practical resource-saving and data privacy issues. During the training process, the neural weights of each model can be randomly permuted, and we have to align the channels of each layer bef…

2022

Distilling Representations from GAN Generator via Squeeze and Span

NeurIPS 2022accept

In recent years, generative adversarial networks (GANs) have been an actively studied topic and shown to successfully produce high-quality realistic images in various domains. The controllable synthesis ability of GAN generators suggests that they maintain informative, disentangled, and explainable…

2022

Experimental Validation of the Usage of Kinematic Singularities to Produce Periodic High-Powered Motion

ICRA 2022poster

This paper reports on preliminary experimental results of recently proposed mechanism kinematics for a legged robot. The proposed kinematics creates a mapping from a series-elastic actuator to a foot motion that includes a pair of singularities within a fully rotatable kinematic circuit. Such a circ…

Cited by 0SourceScholar
2022

How to Represent Context Better? An Empirical Study on Context Modeling for Multi-turn Response Selection

EMNLP 2022finding

Building retrieval-based dialogue models that can predict appropriate responses based on the understanding of multi-turn context messages is a challenging problem. Early models usually concatenate all utterances or independently encode each dialogue turn, which may lead to an inadequate understandin…

Cited by 4SourcePDFScholar
2022

Learning To Learn Across Diverse Data Biases in Deep Face Recognition

CVPR 2022poster

Convolutional Neural Networks have achieved remarkable success in face recognition, in part due to the abundant availability of data. However, the data used for training CNNs is often imbalanced. Prior works largely focus on the long-tailed nature of face datasets in data volume per identity, or foc…

Cited by 26PDFScholar
2022

Multi-Granularity Structural Knowledge Distillation for Language Model Compression

ACL 2022long

Transferring the knowledge to a small model through distillation has raised great interest in recent years. Prevailing methods transfer the knowledge derived from mono-granularity language units (e.g., token-level or sample-level), which is not enough to represent the rich semantics of a text and ma…

2022

Multimodal Transformer for Automatic 3D Annotation and Object Detection

ECCV 2022poster

"Despite a growing number of datasets being collected for training 3D object detection models, significant human effort is still required to annotate 3D boxes on LiDAR scans. To automate the annotation and facilitate the production of various customized datasets, we propose an end-to-end multimodal…

2022

PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior

ICLR 2022poster

Denoising diffusion probabilistic models have been recently proposed to generate high-quality samples by estimating the gradient of the data density. The framework assumes the prior noise as a standard Gaussian distribution, whereas the corresponding data distribution may be more complicated than th…

2022

ProphetChat: Enhancing Dialogue Generation with Simulation of Future Conversation

ACL 2022long

Typical generative dialogue models utilize the dialogue history to generate the response. However, since one dialogue utterance can often be appropriately answered by multiple distinct responses, generating a desired response solely based on the historical information is not easy. Intuitively, if th…

2022

Reciprocal Learning of Knowledge Retriever and Response Ranker for Knowledge-Grounded Conversations

COLING 2022main

Grounding dialogue agents with knowledge documents has sparked increased attention in both academia and industry. Recently, a growing body of work is trying to build retrieval-based knowledge-grounded dialogue systems. While promising, these approaches require collecting pairs of dialogue context an…

Cited by 4SourcePDFScholar
2022

Rethinking Task-Specific Knowledge Distillation: Contextualized Corpus as Better Textbook

EMNLP 2022main

Knowledge distillation has been proven effective when customizing small language models for specific tasks. Here, a corpus as ‘textbook’ plays an indispensable role, only through which the teacher can teach the student. Prevailing methods adopt a two-stage distillation paradigm: general distillation…

Cited by 9SourcePDFScholar
2022

SMASH: Improving SMAll Language Models’ Few-SHot Ability with Prompt-Based Distillation

EMNLP 2022finding

Large-scale language models coupled with prompts have shown remarkable performance on few-shot learning. However, through systematic experiments, we find that the few-shot performance of small language models is poor, and using prompts on them brings fewer improvements than on larger ones. In this p…

2022

Test-time Fourier Style Calibration for Domain Generalization

IJCAI 2022poster

The topic of generalizing machine learning models learned on a collection of source domains to unknown target domains is challenging. While many domain generalization (DG) methods have achieved promising results, they primarily rely on the source domains at train-time without manipulating the target…

2021

Arrhythmia Classification with Heartbeat-Aware Transformer

ICASSP 2021accepted

Electrocardiography (ECG) is a conventional method in arrhythmia diagnosis. In this paper, we proposed a novel neural network model which treats typical heartbeat classification task as ‘Translation’ problem. By introducing Transformer structure into model, and adding heartbeat-aware attention mecha…

Cited by 0SourceScholar
2021

Beyond Bounding-Box: Convex-Hull Feature Adaptation for Oriented and Densely Packed Object Detection

CVPR 2021poster

Detecting oriented and densely packed objects remains challenging for spatial feature aliasing caused by the intersection of reception fields between objects. In this paper, we propose a convex-hull feature adaptation (CFA) approach for configuring convolutional features in accordance with oriented…

Cited by 222PDFcodeScholar
2021

Beyond Max-Margin: Class Margin Equilibrium for Few-Shot Object Detection

CVPR 2021poster

Few-shot object detection has made encouraging progress by reconstructing novel class objects using the feature representation learned upon a set of base classes. However, an implicit contradiction about reconstruction and classification is unfortunately ignored. On the one hand, to precisely recons…

Cited by 215PDFcodeScholar
2021

Computational Design and Fabrication of Corrugated Mechanisms from Behavioral Specifications

ICRA 2021poster

Orthogonally assembled double-layered corrugated (OADLC) mechanisms are a class of foldable structures that harness origami-inspired methods to enhance the structural stiffness of resulting devices; these mechanisms have extensive applications due to their lightweight, compact nature as well as thei…

Cited by 5SourceScholar
2021

ECACL: A Holistic Framework for Semi-Supervised Domain Adaptation

ICCV 2021poster

This paper studies Semi-Supervised Domain Adaptation (SSDA), a practical yet under-investigated research topic that aims to learn a model of good performance using unlabeled samples and a few labeled samples in the target domain, with the help of labeled samples from a source domain. Several SSDA me…

Cited by 76PDFcodeScholar