← Search

Xu Yang

119 accepted papers

2026

Adaptive-Learngene: Continual Expansion and Task-Aware Selection of Learngenes for Dynamic Environments

AAAI 2026technical

Pre-trained Vision Transformer (ViT) models have achieved impressive performance across various computer vision tasks. However, most existing pre-trained models are built on fixed datasets and lack the flexibility to incorporate new pre-training data. When additional data becomes available, previous

Cited by 0SourcePDFScholar
2026

Agent as Student: Learning From Informative Cues for Active Open-Vocabulary Recognition

RA-L 2026

Active recognition, a fundamental task in embodied vision, aims to improve recognition performance by dynamically adjusting viewpoints and poses to mitigate the negative impacts of occlusion and blind spots. Although existing active recognition methods possess basic viewpoint adaptation capabilities

Cited by 0SourceScholar
2026

Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning

ICML 2026spotlight

As a foundational task in human-centric cross-modal intelligence, motion-language retrieval aims to bridge the semantic gap between natural language and human motion, enabling intuitive motion analysis, yet existing approaches predominantly focus on aligning entire motion sequences with global textu…

Cited by 0SourceScholar
2026

Breaking Semantic Boundaries: Distribution-Guided Semantic Exploration for Creative Generation

CVPR 2026

Text-to-image (T2I) diffusion models effectively produce semantically aligned images, but their reliance on training distributions constrains their capacity for synthesizing truly novel, out-of-distribution concepts. Existing methods attempt to enhance creativity through semantic exploration, such a

Cited by 0SourceScholar
2026

CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs

ICLR 2026poster

Chain-of-Thought (CoT) prompting has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs). However, existing implementations, such as in-context learning and fine-tuning, remain costly and inefficient. To improve CoT reasoning at a lower cost, and in…

Cited by 0SourceScholar
2026

Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous Distillation

CVPR 2026

Cloud-Edge Continual Test-Time Adaptation (CTTA)--with edge devices processing real-time data and the cloud offering strong computing power--is a critical paradigm for models that adapt to dynamic data distributions in real-world scenarios. However, most existing frameworks assume architectural homo

Cited by 0SourceScholar
2026

Efficient and Effective In-context Demonstration Selection with Coreset

AAAI 2026technical

In-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. However, the effectiveness of this approach is heavily reliant on the selection of demonstrations, a process that is NP-hard.

Cited by 0SourcePDFScholar
2026

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

AAAI 2026technical

CLIP (Contrastive Language-Image Pre-training) has attracted widespread attention for its multimodal generalizable knowledge, which is significant for downstream tasks. However, the computational overhead of a large number of parameters and large-scale pre-training poses challenges of pre-training a

Cited by 0SourcePDFScholar
2026

FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents

ICML 2026poster

Fine-tuning large language models for vertical domains remains a labor-intensive and expensive process, requiring domain experts to curate data, configure training, and iteratively diagnose model behavior. Despite growing interest in autonomous machine learning, no prior work has tackled end-to-end …

Cited by 0SourceScholar
2026

Flatter Tokens are More Valuable for Speculative Draft Model Training

ICLR 2026poster

Speculative Decoding (SD) is a key technique for accelerating Large Language Model (LLM) inference, but it typically requires training a draft model on a large dataset. We approach this problem from a data-centric perspective, finding that not all training samples contribute equally to the SD accept…

Cited by 0SourcecodeScholar
2026

GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent System

ICLR 2026poster

In the realm of multi-agent systems, the challenge of partial observability is a critical barrier to effective coordination and decision-making. Existing approaches, such as belief state estimation and inter-agent communication, often fall short. Belief-based methods are limited by their focus on pa…

Cited by 0SourceScholar
2026

GraphIC: A Graph-Based In-Context Example Retrieval Model for Multi-Step Reasoning

AAAI 2026technical

In-context learning (ICL) enhances large language models (LLMs) by incorporating demonstration examples, yet its effectiveness heavily depends on the quality of selected examples. Current methods typically use text embeddings to measure semantic similarity, which often introduces bias in multi-step

Cited by 0SourcePDFScholar
2026

Inheriting Generalizable Knowledge from LLMs to Diverse Vertical Tasks

ICLR 2026poster

Large language models (LLMs) have demonstrated remarkable generalization across diverse tasks, suggesting the existence of task-agnostic, generalizable knowledge encoded within them. However, how to systematically extract and evaluate this knowledge remains unexplored. In this work, we innovatively…

Cited by 0SourcecodeScholar
2026

Learngene: Inheritable ‘Genes’ in Intelligent Agents (Abstract Reprint)

AAAI 2026technical

Biological intelligence has driven significant progress in artificial intelligence (AI), but a critical gap remains: biological systems inherit innate abilities from genes, with brains initialized by blueprints refined over 3.5 billion years of evolution, while machines rely heavily on inefficient,

Cited by 0SourcePDFScholar
2026

Learning Attribute–Affordance Hierarchies in Hyperbolic Space for Open-Vocabulary 3D Object Affordance Grounding

ICML 2026poster

This paper pays attention to open-vocabulary 3D object affordance grounding (OVAG), which aims to localize affordance regions on 3D objects by leveraging interaction images or textual instructions. Most existing methods treat interaction images as sources of external affordance knowledge and align t…

Cited by 0SourceScholar
2026

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

ICML 2026poster

Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment in real-world applications. However, existing robustness benchmarks typically focus on hallucination or misleading textual inputs, while largely overlooking the e…

Cited by 0SourceScholar
2026

Memoria-Bench: A Comprehensive Benchmark for Evaluating Memory in Long-Horizon Autonomous Agents

ICML 2026poster

Memory is a core capability of autonomous agents, yet existing benchmarks evaluate it primarily in constrained settings such as short dialogues or synthetic tasks, failing to reflect realistic agent deployments. We present \textbf{Memoria-Bench}, a benchmark for evaluating agent memory grounded in c…

Cited by 0SourceScholar
2026

Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy

CVPR 2026

The development of affective multimodal language models (MLMs) has long been constrained by a gap between low-level perception and high-level interaction, leading to fragmented affective capabilities and limited generalization. To bridge this gap, we propose a cognitively inspired three-level hierar

Cited by 0SourceScholar
2026

OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset Exploration

ICLR 2026poster

Preference-based reinforcement learning (PbRL) can help avoid sophisticated reward designs and align better with human intentions, showing great promise in various real-world applications. However, obtaining human feedback for preferences can be expensive and time-consuming, which forms a strong bar…

Cited by 0SourceScholar
2026

On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification

ICLR 2026poster

In this work, we present a simple yet theoretically motivated improvement to Supervised Fine-Tuning (SFT) for the Large Language Model (LLM), addressing its limited generalization compared to reinforcement learning (RL). Through mathematical analysis, we reveal that standard SFT gradients implicitly…

Cited by 0SourcecodeScholar
2026

Raise One and Infer Three: Toward Reasoning- and Memory-Augmented Diffusion Policy Generalization

IJCAI 2026

Diffusion policy has shown impressive performance in robotic manipulation tasks while struggling with out-of-distribution shifts and limited demonstrations. Recent advances primarily focus on improving geometric or perceptual representations for diffusion policy. However, these approaches rely heavi

Cited by 0Scholar
2026

Recover to Predict: Progressive Retrospective Learning for Variable-Length Trajectory Prediction

CVPR 2026

Trajectory prediction is critical for autonomous driving, enabling safe and efficient planning in dense, dynamic traffic. Most existing methods optimize prediction accuracy under fixed-length observations. However, real-world driving often yields variable-length, incomplete observations, posing a ch

Cited by 0SourcecodeScholar
2026

Rethinking LLM Ensembling from the Perspective of Mixture Models

ICML 2026spotlight

Model ensembling is a well-established technique for improving the performance of machine learning models. Conventionally, this involves averaging the output distributions of multiple models and selecting the most probable label. This idea has been naturally extended to large language models (LLMs),…

Cited by 0SourceScholar
2026

Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys

AAAI 2026technical

The KV cache in self-attention has emerged as a major bottleneck in long-context and large-batch inference for LLMs. Existing approaches often treat sparsity prediction and compression as separate modules—relying on auxiliary index structures to select relevant tokens, and on complex quantization sc

Cited by 0SourcePDFScholar
2026

S²Flow: Towards Fast and Authentic Training-Free High-Resolution Video Generation

AAAI 2026technical

Rectified flow models have shown strong potential in high-fidelity video generation, yet extending them to high-resolution remains challenging due to the high cost of full attention and error accumulation in the ODE-solving process. In this paper, we propose S^2Flow, a training-free framework that e

Cited by 0SourcePDFScholar
2026

TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV Dramas

CVPR 2026

Video highlight detection aims to identify the most engaging segments in long-form videos, supporting content editing and recommendation, especially for movies and TV dramas. However, existing methods are ill-suited to cinematic content due to its narrative complexity, while the scarcity of annotate

Cited by 0SourceScholar
2026

TinyVPR: Distilling Correct and Confusing Knowledge for Lightweight Visual Place Recognition

ICRA 2026poster

Visual Place Recognition (VPR) is a key technology in autonomous driving, robotics, and augmented reality, requiring efficient and robust localization in large-scale environments. However, most existing methods rely on heavy deep models that are computationally expensive and difficult to deploy on e…

Cited by 0Scholar
2026

Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training

ICML 2026poster

Supervised fine-tuning (SFT) is computationally efficient but often yields inferior generalization compared to reinforcement learning (RL). This gap is primarily driven by RL’s use of on-policy data. We propose a framework to bridge this chasm by enabling On-Policy SFT. We first present ***Distribut…

Cited by 0SourceScholar
2026

Trajectory-Stabilized Inference for Diffusion-Based Video Inpainting

ICML 2026poster

Video inpainting aims to restore missing regions while preserving spatial and temporal coherence. Diffusion-based methods achieve strong per-frame reconstruction, but their sampling implicitly generates temporally coupled latent trajectories whose long-horizon stability is not explicitly modeled, le…

Cited by 0SourceScholar
2026

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions

ICML 2026poster

Layer pruning efficiently reduces Large Language Model (LLM) computational costs but often triggers sudden performance collapse. Existing representation-based analyses struggle to explain this mechanism. We propose studying pruning through decision representation. Focusing on multiple-choice tasks, …

Cited by 0SourceScholar
2026

Unifying Precise Keyframes and Semantic Control via Multi-level Diffusion

CVPR 2026

Text-conditioned human motion in-betweening leverages keyframes for spatio-temporal control, with text providing high-level semantic guidance for the transitions. However, existing methods are unable to establish a coherent alignment between textual semantics and the spatio-temporal constraints prov

Cited by 0SourceScholar
2026

d$^2$Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching

ICLR 2026poster

Diffusion-based large language models (dLLMs), despite their promising performance, still suffer from inferior inference efficiency. This is because dLLMs rely on bidirectional attention and cannot directly benefit from the standard key-value (KV) cache as autoregressive models (ARMs) do. To tackle…

Cited by 0SourcecodeScholar
2025

Adversarial Graph Fusion for Incomplete Multi-view Semi-supervised Learning with Tensorial Imputation

NeurIPS 2025poster

View missing remains a significant challenge in graph-based multi-view semi-supervised learning, hindering their real-world applications. To address this issue, traditional methods introduce a missing indicator matrix and focus on mining partial structure among existing samples in each view for labe…

Cited by 0SourcecodeScholar
2025

Data Selection Matters: Towards Robust Instruction Tuning of Large Multimodal Models

NeurIPS 2025poster

Selecting a compact subset of visual instruction–following data has emerged as an effective way to align large multimodal models with human intentions while avoiding the high cost of full-dataset training. Yet we observe that both full-data training and existing state-of-the-art data selection metho…

Cited by 1SourcecodeScholar
2025

Democratizing High-Fidelity Co-Speech Gesture Video Generation

ICCV 2025poster

Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presents challenges due to the significant one-to-many mapping between audio and visual content, further complicated by the sc…

2025

Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

CVPR 2025poster

Hallucinations in Large Vision-Language Models (LVLMs) significantly undermine their reliability, motivating researchers to explore the causes of hallucination. However, most studies primarily focus on the language aspect rather than the visual. In this paper, we address how LVLMs process visual inf…

2025

Dual-Space Semantic Synergy Distillation for Continual Learning of Unlabeled Streams

NeurIPS 2025poster

Continual learning from unlabeled data streams while effectively combating catastrophic forgetting poses an intractable challenge. Traditional methods predominantly rely on visual clustering techniques to generate pseudo labels, which are frequently plagued by problems such as noise and suboptimal q…

Cited by 0SourceScholar
2025

Energy-Efficient Omnidirectional Locomotion for Wheeled Quadrupeds via Predictive Energy-Aware Nominal Gait Selection

IROS 2025

Wheeled-legged robots combine the efficiency of wheels with the versatility of legs, but face significant energy optimization challenges when navigating diverse environments. In this work, we present a hierarchical control framework that integrates predictive power modeling with residual reinforceme

Cited by 0SourceScholar
2025

Fast Large Language Model Collaborative Decoding via Speculation

ICML 2025poster

Large Language Model (LLM) collaborative decoding techniques improve output quality by combining the outputs of multiple models at each generation step, but they incur high computational costs. In this paper, we introduce **Collaborative decoding via Speculation (CoS)**, a novel framework that accel…

2025

FlowPrune: Accelerating Attention Flow Calculation by Pruning Flow Network

NeurIPS 2025poster

The Transformer architecture serves as the foundation of modern AI systems, powering recent advances in Large Language Models (LLMs) and Large Multimodal Models (LMMs). Central to these models, attention mechanisms capture contextual dependencies via token interactions. Beyond inference, attention h…

Cited by 0SourceScholar
2025

Foreground-aware Prototypical Network for Prohibited Item Detection from X-ray Scans

ICASSP 2025accepted

Automatic inspection of X-ray scans is a critical component of modern safety protocols. It plays an indispensable role in detecting concealed weapons, explosives, and other prohibited items that could pose a threat to public safety. Current surveillance systems perform poorly without human intervent…

Cited by 0SourceScholar
2025

Inheriting Generalized Learngene for Efficient Knowledge Transfer across Multiple Tasks

AAAI 2025technical

In practical applications, it is often necessary to transfer knowledge from large pretrained models to small ones with various architectures for tackling different tasks. The Learngene framework, proposed recently, firstly extracts one compact module termed as learngene from a large well-trained mod…

Cited by 0SourcePDFScholar
2025

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

NeurIPS 2025poster

Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, We introduce KRIS-B…

Cited by 0SourceScholar
2025

Learngene Tells You How to Customize: Task-Aware Parameter Initialization at Flexible Scales

ICML 2025poster

Appropriate parameter initialization strategies are essential for reducing the high computational costs of training large pretrained models in various task scenarios. Graph HyperNetwork (GHN), a parameter initialization method, has recently demonstrated strong performance in initializing models. How…

Cited by 0SourcePDFScholar
2025

Mimic In-Context Learning for Multimodal Tasks

CVPR 2025poster

Recently, In-context Learning (ICL) has become a significant inference paradigm in Large Multimodal Models (LMMs), utilizing a few in-context demonstrations (ICDs) to prompt LMMs for new tasks. However, the synergistic effects in multimodal data increase the sensitivity of ICL performance to the con…

2025

Number it: Temporal Grounding Videos like Flipping Manga

CVPR 2025poster

Video Large Language Models (Vid-LLMs) have made remarkable advancements in comprehending video content for QA dialogue. However, they struggle to extend this visual understanding to tasks requiring precise temporal localization, known as Video Temporal Grounding (VTG). To address this, we introduce…

2025

Q-MiniSAM2: A Quantization-based Benchmark for Resource-Efficient Video Segmentation

IJCAI 2025

Segment Anything Model 2 (SAM2) is a new-generation, high-precision model for image and video segmentation, offering extensive application prospects across numerous computer vision fields. However, as a large-scale model, its huge memory demands and expansive computing costs pose challenges for prac

Cited by 0SourcePDFScholar
2025

R&D-Agent-Quant: A Multi-Agent Framework for Data-Centric Factors and Model Joint Optimization

NeurIPS 2025poster

Financial markets pose fundamental challenges for asset return prediction due to their high dimensionality, non-stationarity, and persistent volatility. Despite advances in large language models and multi-agent systems, current quantitative research pipelines suffer from limited automation, weak int…

Cited by 0SourcecodeScholar
2025

RAPID Hand: Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platfrom for Embodied Intelligence

NeurIPS 2025poster

This paper addresses the scarcity of low-cost but high-dexterity platforms for collecting real-world multi-fingered robot manipulation data towards generalist robot autonomy. To achieve it, we propose the RAPID Hand, a co-optimized hardware and software platform where the compact 20-DoF hand, robus…

Cited by 0SourceScholar
2025

Redefining <Creative> in Dictionary: Towards an Enhanced Semantic Understanding of Creative Generation

CVPR 2025poster

Creative remains an inherently abstract concept for both humans and diffusion models. While text-to-image (T2I) diffusion models can easily generate out-of-distribution concepts like "a blue banana", they struggle with generating combinatorial objects such as "a creative mixture that resembles a let…

2025

SSCM: Self-Supervised Critical Model for Reducing Hallucinations in Chinese Financial Text Generation

ICASSP 2025accepted

Large Language Models (LLMs) show strong performance in natural language processing tasks, but their application in the financial domain is limited. Current methods rely on large datasets and manual prompt engineering, resulting in high data demands, long inference times, and frequent hallucinations…

Cited by 0SourceScholar
2025

Tackling Long-Tailed Data Challenges in Spiking Neural Networks via Heterogeneous Knowledge Distillation

IJCAI 2025

Spiking Neural Networks (SNNs), inspired by the behavior of biological neurons, have gained significant research interest for resource-constrained edge devices and neuromorphic hardware due to their use of binary spike signals for inter-unit communication with low power consumption. However, the abs

Cited by 0SourcePDFScholar
2025

Three-Dimensional Trajectory Prediction with 3DMoTraj Dataset

ICML 2025poster

With the growing interest in embodied and spatial intelligence, accurately predicting trajectories in 3D environments has become increasingly critical. However, no datasets have been explicitly designed to study 3D trajectory prediction. To this end, we contribute a 3D motion trajectory (3DMoTraj) d…

2025

Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient

AAAI 2025technical

Text-to-image diffusion models have achieved remarkable success in generating photorealistic images. However, the inclusion of sensitive information during pre-training poses significant risks. Machine Unlearning (MU) offers a promising solution to eliminate sensitive concepts from these models. Des…

2025

Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark

AAAI 2025technical

The demand for producing short-form videos for sharing on social media platforms has experienced significant growth in recent times. Despite notable advancements in the fields of video summarization and highlight detection, which can create partially usable short films from raw videos, these approac…

2024

A Versatile Framework for Continual Test-Time Domain Adaptation: Balancing Discriminability and Generalizability

CVPR 2024poster

Continual test-time domain adaptation (CTTA) aims to adapt the source pre-trained model to a continually changing target domain without additional data acquisition or labeling costs. This issue necessitates an initial performance enhancement within the present domain without labels while concurrentl…

Cited by 3SourcePDFScholar
2024

BPQP: A Differentiable Convex Optimization Framework for Efficient End-to-End Learning

NeurIPS 2024spotlight

Data-driven decision-making processes increasingly utilize end-to-end learnable deep neural networks to render final decisions. Sometimes, the output of the forward functions in certain layers is determined by the solutions to mathematical optimization problems, leading to the emergence of different…

Cited by 1SourcePDFScholar
2024

Building Variable-Sized Models via Learngene Pool

AAAI 2024technical

Recently, Stitchable Neural Networks (SN-Net) is proposed to stitch some pre-trained networks for quickly building numerous networks with different complexity and performance trade-offs. In this way, the burdens of designing or training the variable-sized networks, which can be used in application s…

2024

C3P-VoxelMap: Compact, Cumulative and Coalescible Probabilistic Voxel Mapping

IROS 2024poster

This work presents a compact, cumulative, and coalescible probabilistic voxel mapping method to enhance performance, accuracy, and memory efficiency in LiDAR odometry. Probabilistic voxel mapping requires storing past point clouds and re-iterating them to update the uncertainty at every iteration, w…

Cited by 0SourcecodeScholar
2024

Cluster-Learngene: Inheriting Adaptive Clusters for Vision Transformers

NeurIPS 2024poster

In recent years, the merging of vast datasets with powerful computational resources has led to the emergence of large pre-trained models in the field of deep learning. However, the common practices often overgeneralize the applicability of these models, overlooking the task-specific resource constra…

Cited by 1SourcePDFScholar
2024

Exploiting Intrinsic Multilateral Logical Rules for Weakly Supervised Natural Language Video Localization

ACL 2024long

Weakly supervised natural language video localization (WS-NLVL) aims to retrieve the moment corresponding to a language query in a video with only video-language pairs utilized during training. Despite great success, existing WS-NLVL methods seldomly consider the complex temporal relations enclosing…

2024

Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models

IJCAI 2024poster

In practice, we usually need to build variable-sized models adapting for diverse resource constraints in different application scenarios, where weight initialization is an important step prior to training. The Learngene framework, introduced recently, firstly learns one compact part termed as learng…

2024

How to Configure Good In-Context Sequence for Visual Question Answering

CVPR 2024poster

Inspired by the success of Large Language Models in dealing with new tasks via In-Context Learning (ICL) in NLP researchers have also developed Large Vision-Language Models (LVLMs) with ICL capabilities. However when implementing ICL using these LVLMs researchers usually resort to the simplest way l…

2024

Initializing Variable-sized Vision Transformers from Learngene with Learnable Transformation

NeurIPS 2024poster

In practical scenarios, it is necessary to build variable-sized models to accommodate diverse resource constraints, where weight initialization serves as a crucial step preceding training. The recently introduced Learngene framework firstly learns one compact module, termed learngene, from a large w…

Cited by 4SourcePDFScholar
2024

LIVE: Learnable In-Context Vector for Visual Question Answering

NeurIPS 2024poster

As language models continue to scale, Large Language Models (LLMs) have exhibited emerging capabilities in In-Context Learning (ICL), enabling them to solve language tasks by prefixing a few in-context demonstrations (ICDs) as context. Inspired by these advancements, researchers have extended these…

2024

Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models

NeurIPS 2024poster

As Archimedes famously said, ``Give me a lever long enough and a fulcrum on which to place it, and I shall move the world'', in this study, we propose to use a tiny Language Model (LM), \eg, a Transformer with 67M parameters, to lever much larger Vision-Language Models (LVLMs) with 9B parameters. Sp…

Cited by 8SourcePDFScholar
2024

Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale Models

NeurIPS 2024poster

Vision Transformers (ViTs) are widely used in a variety of applications, while they usually have a fixed architecture that may not match the varying computational resources of different deployment environments. Thus, it is necessary to adapt ViT architectures to devices with diverse computational ov…

Cited by 3SourcePDFScholar
2024

Long-Tail Class Incremental Learning via Independent Sub-prototype Construction

CVPR 2024poster

Long-tail class incremental learning (LT-CIL) is designed to perpetually acquire novel knowledge from an imbalanced and perpetually evolving data stream while ensuring the retention of previously acquired knowledge. The existing method only re-balances data distribution and ignores exploring the pot…

Cited by 5SourcePDFScholar
2024

MemoNav: Working Memory Model for Visual Navigation

CVPR 2024highlight

Image-goal navigation is a challenging task that requires an agent to navigate to a goal indicated by an image in unfamiliar environments. Existing methods utilizing diverse scene memories suffer from inefficient exploration since they use all historical observations for decision-making without cons…

2024

Mixture of Adversarial LoRAs: Boosting Robust Generalization in Meta-Tuning

NeurIPS 2024poster

This paper introduces AMT, an \textbf{A}dversarial \textbf{M}eta-\textbf{T}uning methodology, to boost the robust generalization of pre-trained models in the out-of-domain (OOD) few-shot learning. To address the challenge of transferring knowledge from source domains to unseen target domains, we con…

2024

Navigating Continual Test-time Adaptation with Symbiosis Knowledge

IJCAI 2024poster

Continual test-time domain adaptation seeks to adapt the source pre-trained model to a continually changing target domain without incurring additional data acquisition or labeling costs. Unfortunately, existing mainstream methods may result in a detrimental cycle. This is attributed to noisy pseudo-…

Cited by 0SourcePDFScholar
2024

Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-On

CVPR 2024poster

Image-based virtual try-on is an increasingly important task for online shopping. It aims to synthesize images of a specific person wearing a specified garment. Diffusion model-based approaches have recently become popular as they are excellent at image synthesis tasks. However these approaches usua…

2024

Transformer as Linear Expansion of Learngene

AAAI 2024technical

We propose expanding the shared Transformer module to produce and initialize Transformers of varying depths, enabling adaptation to diverse resource constraints. Drawing an analogy to genetic expansibility, we term such module as learngene. To identify the expansion mechanism, we delve into the rela…

2024

Unveiling the Unknown: Unleashing the Power of Unknown to Known in Open-Set Source-Free Domain Adaptation

CVPR 2024poster

Open-Set Source-Free Domain Adaptation aims to transfer knowledge in realistic scenarios where the target domain has additional unknown classes compared to the limited-access source domain. Due to the absence of information on unknown classes existing methods mainly transfer knowledge of known class…

2024

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception

ICML 2024poster

This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the "Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision,…

2023

Exploring Diverse In-Context Configurations for Image Captioning

NeurIPS 2023poster

After discovering that Language Models (LMs) can be good in-context few-shot learners, numerous strategies have been proposed to optimize in-context sequence configurations. Recently, researchers in Vision-Language (VL) domains also develop their few-shot learners, while they only use the simplest w…

2023

Learning Trajectory-Word Alignments for Video-Language Tasks

ICCV 2023poster

In a video, an object usually appears as the trajectory, i.e., it spans over a few spatial but longer temporal patches, that contains abundant spatiotemporal contexts. However, modern Video-Language BERTs (VDL-BERTs) neglect this trajectory characteristic that they usually follow image-language BERT…

Cited by 6PDFScholar
2023

Transforming Visual Scene Graphs to Image Captions

ACL 2023long

We propose to TransForm Scene Graphs into more descriptive Captions (TFSGC). In TFSGC, we apply multi-head attention (MHA) to design the Graph Neural Network (GNN) for embedding scene graphs. After embedding, different graph embeddings contain diverse specific knowledge for generating the words with…

2023

Unseen Object Instance Segmentation with Fully Test-time RGB-D Embeddings Adaptation

ICRA 2023poster

Segmenting unseen objects is a crucial ability for the robot since it may encounter new environments during the operation. Recently, a popular solution is leveraging RGB-D features of large-scale synthetic data and directly applying the model to unseen real-world scenarios. However, the domain shift…

Cited by 11SourceScholar
2022

Attention-guided Contrastive Hashing for Long-tailed Image Retrieval

IJCAI 2022poster

Image hashing is to represent an image using a binary code for efficient storage and accurate retrieval. Recently, deep hashing methods have shown great improvements on ideally balanced datasets, however, long-tailed data is more common due to rare samples or data collection costs in the real world.…

2022

EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding Matching

CVPR 2022poster

Current metrics for video captioning are mostly based on the text-level comparison between reference and candidate captions. However, they have some insuperable drawbacks, e.g., they cannot handle videos without references, and they may result in biased evaluation due to the one-to-many nature of vi…

Cited by 43PDFcodeScholar
2022

Image Translation Based Synthetic Data Generation for Industrial Object Detection and Pose Estimation

RA-L 2022

Deep learning-based methods have shown excellent potential on object detection and pose estimation with vast amounts of training data to achieve good performance. Obtaining enough comprehensive manual labeling training data is a time-consuming and error-prone task in industrial scenes, and most curr

Cited by 20SourceScholar
2022

Learning Universal Adversarial Perturbation by Adversarial Example

AAAI 2022technical

Deep learning models have shown to be susceptible to universal adversarial perturbation (UAP), which has aroused wide concerns in the community. Compared with the conventional adversarial attacks that generate adversarial samples at the instance level, UAP can fool the target model for different ins…

2022

Not Just Selection, but Exploration: Online Class-Incremental Continual Learning via Dual View Consistency

CVPR 2022poster

Online class-incremental continual learning aims to learn new classes continually from a never-ending and single-pass data stream, while not forgetting the learned knowledge of old classes. Existing replay-based methods have shown promising performance by storing a subset of old class data. Unfortun…

Cited by 104PDFcodeScholar
2022

Show, Deconfound and Tell: Image Captioning With Causal Inference

CVPR 2022poster

The transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce t…

Cited by 67PDFcodeScholar
2022

Towards End-to-End Image Compression and Analysis with Transformers

AAAI 2022technical

We propose an end-to-end image compression and analysis model with Transformers, targeting to the cloud-based image classification application. Instead of placing an existing Transformer-based image classification model directly after an image codec, we aim to redesign the Vision Transformer (ViT) m…

2021

Auto-Parsing Network for Image Captioning and Visual Question Answering

ICCV 2021poster

We propose an Auto-Parsing Network (APN) to discover and exploit the input data's hidden tree structures for improving the effectiveness of the Transformer-based vision-language systems. Specifically, we impose a Probabilistic Graphical Model (PGM) parameterized by the attention operations on each s…

Cited by 44PDFScholar
2021

Graph Debiased Contrastive Learning with Joint Representation Clustering

IJCAI 2021poster

By contrasting positive-negative counterparts, graph contrastive learning has become a prominent technique for unsupervised graph representation learning. However, existing methods fail to consider the class information and will introduce false-negative samples in the random negative sampling, causi…

Cited by 179SourcePDFScholar
2021

SelfSAGCN: Self-Supervised Semantic Alignment for Graph Convolution Network

CVPR 2021poster

Graph convolution networks (GCNs) are a powerful deep learning approach and have been successfully applied to representation learning on graphs in a variety of real-world applications. Despite their success, two fundamental weaknesses of GCNs limit their ability to represent graph-structured data: p…

Cited by 42PDFcodeScholar
2020

Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning

ECCV 2020poster

Change Captioning is a task that aims to describe the difference between images with natural language. Most existing methods treat this problem as a difference judgment without the existence of distractors such as viewpoint changes. However, in practice, viewpoint changes happen often and can overwh…

Cited by 54SourcePDFScholar
2020

Learning Progressive Joint Propagation for Human Motion Prediction

ECCV 2020poster

Despite the great progress in human motion prediction, it remains a challenging task due to the complicated structural dynamics of human behaviors. In this paper, we address this problem in three aspects. First, to capture the long-range spatial correlations and temporal dependencies, we apply a tra…

Cited by 197SourcePDFScholar
2020

Lifelong Zero-Shot Learning

IJCAI 2020poster

Zero-Shot Learning (ZSL) handles the problem that some testing classes never appear in training set. Existing ZSL methods are designed for learning from a fixed training set, which do not have the ability to capture and accumulate the knowledge of multiple training sets, causing them infeasible to m…

Cited by 0SourcePDFScholar
2019

Unpaired Image Captioning via Scene Graph Alignments

ICCV 2019poster

Most of current image captioning models heavily rely on paired image-caption datasets. However, getting large scale image-caption paired data is labor-intensive and time-consuming. In this paper, we present a scene graph-based approach for unpaired image captioning. Our framework comprises an image…

Cited by 209PDFScholar
2019

Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian Detection

ICCV 2019poster

Multispectral pedestrian detection has shown great advantages under poor illumination conditions, since the thermal modality provides complementary information for the color image. However, real multispectral data suffers from the position shift problem, i.e. the color-thermal image pairs are not st…

Cited by 241PDFcodeScholar
2018

Shuffle-Then-Assemble: Learning Object-Agnostic Visual Relationship Features

ECCV 2018poster

Due to fact that it is prohibitively expensive to completely annotate visual relationships, ie, the (obj1, rel, obj2) triplets, relationship models are inevitably biased to object classes of limited pairwise patterns, leading to poor generalization to rare or unseen object combinations. Therefore, w…