← Search

Rui Zhao

148 accepted papers

2026

Adaptive-Smooth LiDAR-Camera Knowledge Distillation with Heterogeneous Fusion for Multi-View 3D Object Detection

AAAI 2026technical

Multi-view 3D object detection has garnered increasing attention, particularly due to its success in autonomous driving systems. Although multi-view systems possess rich semantic information, their spatial-geometric reasoning capabilities remain limited. Recent studies employ simulated point cloud g

Cited by 0SourcePDFScholar
2026

Cooperative-Competitive Team Play of Real-World Craft Robots

ICRA 2026poster

Multi-agent deep Reinforcement Learning (RL) has made significant progress in developing intelligent game-playing agents in recent years. However, the efficient training of collective robots using multi-agent RL and the transfer of learned policies to real-world applications remain open research que…

2026

Direct 3D-Aware Object Insertion via Decomposed Visual Proxies

ICML 2026poster

Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion-based methods achieve high visual quality but formulate insertion as a simple 2D inpainting task, providing no explicit control over the object’s 3D pose and limiting thei…

Cited by 0SourceScholar
2026

GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models

ICLR 2026poster

Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for Autonomous Driving, Embodied AI and General Artificial Intelligence. Existing spatial-temporal benchmarks mainly focus on egocentric perspective reasoning with images/video…

Cited by 0SourcecodeScholar
2026

Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D Retrieval

AAAI 2026technical

With the daily influx of 3D data on the internet, text-3D retrieval has gained increasing attention. However, current methods face two major challenges: Hierarchy Representation Collapse (HRC) and Redundancy-Induced Saliency Dilution (RISD). HRC compresses abstract-to-specific and whole-to-part hier

Cited by 0SourcePDFScholar
2026

Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning

ICML 2026poster

Existing multimodal reasoning approaches predominantly follow two paradigms: converting visual inputs into text prior to reasoning, or performing end-to-end reasoning within a unified vision–language representation space. Despite their empirical progress, both paradigms suffer from fundamental struc…

Cited by 0SourceScholar
2026

PLaST: Towards Paralinguistic-aware Speech Translation

AAAI 2026technical

Speech translation (ST) aims to translate speech from a source language into text in the target language. Naturally, speech signals contain paralinguistic cues beyond linguistic content, which could influence or even alter the interpretation of a lexically identical sentence, thereby yielding distin

Cited by 0SourcePDFScholar
2026

Spike Stream Memory Transfer for Dynamic Scene Reconstruction

AAAI 2026technical

As a retina-inspired sensor with ultra-high temporal resolution, spike camera can continuously capture dynamic scenes with high-speed motion. It is a key task to restore clear images from spike streams. The quantization effects in spike readout bring degradation to the visual quality of restored ima

Cited by 0SourcePDFScholar
2026

SpikeStereoNet: A Brain-Inspired Framework for Stereo Depth Estimation from Spike Streams

ICLR 2026poster

Conventional frame-based cameras often struggle with stereo depth estimation in rapidly changing scenes. In contrast, bio-inspired spike cameras emit asynchronous events at microsecond-level resolution, providing an alternative sensing modality. However, existing methods lack specialized stereo algo…

Cited by 0SourcecodeScholar
2025

Can an Individual Manipulate the Collective Decisions of Multi-Agents?

EMNLP 2025

Individual Large Language Models (LLMs) have demonstrated significant capabilities across various domains, such as healthcare and law. Recent studies also show that coordinated multi-agent systems exhibit enhanced decision-making and reasoning abilities through collaboration. However, due to the vul

Cited by 0SourcePDFScholar
2025

CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model

NeurIPS 2025poster

Autonomous driving represents a prominent application of artificial intelligence. Recent approaches have shifted from focusing solely on common scenarios to addressing complex, long-tail situations such as subtle human behaviors, traffic accidents, and non-compliant driving patterns. Given the demon…

Cited by 0SourceScholar
2025

DiffE2E: Rethinking End-to-End Driving with a Hybrid Diffusion-Regression-Classification Policy

NeurIPS 2025poster

End-to-end learning has emerged as a transformative paradigm for autonomous driving. However, the inherently multimodal nature of driving behaviors remains a fundamental challenge to robust deployment. We propose DiffE2E, a diffusion-based end-to-end autonomous driving framework. The architecture fi…

Cited by 0SourceScholar
2025

DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles

CVPR 2025poster

Adapting generative models to specific domains presents an effective solution for satisfying specialized requirements. However, adapting to some complex domains remains challenging, especially when these domains require substantial paired data to capture the targeted distributions. Since unpaired da…

2025

Dynamic Feature Fusion for Sign Language Translation Using HyperNetworks

NAACL 2025findings

This paper presents an efficient dual-stream early fusion method for sign language translation. Inspired by the brain’s ability to process color, shape, and motion simultaneously, the method explores complex dependencies between RGB and keypoint streams, improving speed and efficiency. A key challen…

2025

Enhancing Extractive Question Answering in Multiparty Dialogues with Logical Inference Memory Network

COLING 2025main

Multiparty dialogue question answering (QA) in machine reading comprehension (MRC) is a challenging task due to its complex information flow interactions and logical QA inference. Existing models typically handle such QA tasks by decoupling dialogue information at both speaker and utterance levels.…

Cited by 0SourcePDFScholar
2025

Graph Pooling via Dropping Task-Irrelevant Nodes

ICASSP 2025accepted

Graph neural networks (GNNs) face scalability challenges. While recent approaches have adopted pooling strategies inspired by convolutional neural networks (CNNs) to reduce graph size and improve efficiency, these methods often focus on local information and are optimized for single graph-level task…

Cited by 0SourceScholar
2025

ISP2HRNet: Learning to Reconstruct High Resolution Image from Irregularly Sampled Pixels via Hierarchical Gradient Learning

ICCV 2025poster

While image signals are typically defined on a regular 2D grid, there are scenarios where they are only available at irregular positions. In such cases, reconstructing a complete image on regular grid is essential. This paper introduces ISP2HRNet, an end-to-end network designed to reconstruct high r…

2025

KITS: Inductive Spatio-Temporal Kriging with Increment Training Strategy

AAAI 2025technical

Sensors are commonly deployed to perceive the environment. However, due to the high cost, sensors are usually sparsely deployed. Kriging is the tailored task to infer the unobserved nodes (without sensors) using the observed nodes (with sensors). The essence of kriging task is transferability. Recen…

2025

PUMA: Empowering Unified MLLM with Multi-granular Visual Generation

ICCV 2025poster

Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models for visual content generation. However, existing approaches face a trade-off between generatio…

2025

Position: Current Model Licensing Practices are Dragging Us into a Quagmire of Legal Noncompliance

ICML 2025oral

The Machine Learning (ML) community has wit- nessed explosive growth, with millions of ML models being published on the Web. Reusing ML model components has been prevalent nowadays. Developers are often required to choose a license to publish and govern the use of their models. Popular options inclu…

Cited by 0SourcePDFScholar
2025

Re-Aligning Language to Visual Objects with an Agentic Workflow

ICLR 2025poster

Language-based object detection (LOD) aims to align visual objects with language expressions. A large amount of paired data is utilized to improve LOD model generalizations. During the training process, recent studies leverage vision-language models (VLMs) to automatically generate human-like expres…

Cited by 0SourcePDFScholar
2025

RemDet: Rethinking Efficient Model Design for UAV Object Detection

AAAI 2025technical

Object detection in Unmanned Aerial Vehicle (UAV) images has emerged as a focal area of research, which presents two significant challenges: i) objects are typically small and dense within vast images; ii) computational resource constraints render most models unsuitable for real-time deployment. Cur…

2025

Representation Purification for End-to-End Speech Translation

COLING 2025main

Speech-to-text translation (ST) is a cross-modal task that involves converting spoken language into text in a different language. Previous research primarily focused on enhancing speech translation by facilitating knowledge transfer from machine translation, exploring various methods to bridge the g…

2025

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

IROS 2025

Recent advancements in imitation learning have shown promising results in robotic manipulation, driven by the availability of high-quality training data. To improve data collection efficiency, some approaches focus on developing specialized teleoperation devices for robot control, while others direc

Cited by 2SourcecodeScholar
2025

SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action Recognition

ICCV 2025poster

Open-vocabulary action recognition (OVAR) extends recognition systems to identify unseen action categories. While large-scale vision-language models (VLMs) like CLIP have enabled OVAR in image domains, their adaptation to event data remains underexplored. Event cameras offer high temporal resolution…

2025

STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization

NeurIPS 2025poster

Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning from human feedback. However, due to the high cost of obtaining feedback, PbRL typically relies on a limited set of preference-labeled samples. This data scarcity introduces two key inefficiencies: (1) the…

Cited by 0SourceScholar
2025

Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning

RA-L 2025

End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an crucial part of Embodied AI. Despite successes in applying Multimodal Large Language Models (MLLMs) for high-level traffic scene semantic understanding, it remains challenging to effectively tra

Cited by 23SourceScholar
2025

TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality Alignment

AAAI 2025technical

Multivariate time series forecasting (MTSF) aims to learn temporal dynamics among variables to forecast future time series. Existing statistical and deep learning-based methods suffer from limited learnable parameters and small-scale training data. Recently, large language models (LLMs) combining ti…

2025

Towards Cross-Modality Modeling for Time Series Analytics: A Survey in the LLM Era

IJCAI 2025

The proliferation of edge devices has generated an unprecedented volume of time series data across different domains, motivating a variety of well-customized methods. Recently, Large Language Models (LLMs) have emerged as a new paradigm for time series analytics by leveraging the shared sequential n

2025

Tree-KG: An Expandable Knowledge Graph Construction Framework for Knowledge-intensive Domains

ACL 2025long

In knowledge-intensive domains like scientific research, effective decisions rely on organizing and retrieving intricate data. Knowledge graphs (KGs) help by structuring entities, relations, and contextual dependencies, but building KGs in such domains is challenging due to inherent complexity, manu…

2025

UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback

NeurIPS 2025poster

Relighting is a crucial task with both practical demand and artistic value, and recent diffusion models have shown strong potential by enabling rich and controllable lighting effects. However, as they are typically optimized in semantic latent space, where proximity does not guarantee physical corre…

Cited by 0SourcecodeScholar
2025

Unlocking the Power of SAM 2 for Few-Shot Segmentation

ICML 2025poster

Few-Shot Segmentation (FSS) aims to learn class-agnostic segmentation on few classes to segment arbitrary classes, but at the risk of overfitting. To address this, some methods use the well-learned knowledge of foundation models (e.g., SAM) to simplify the learning process. Recently, SAM 2 has exten…

2024

Balancing Speciality and Versatility: a Coarse to Fine Framework for Supervised Fine-tuning Large Language Model

ACL 2024findings

Aligned Large Language Models (LLMs) showcase remarkable versatility, capable of handling diverse real-world tasks. Meanwhile, aligned LLMs are also expected to exhibit speciality, excelling in specific applications. However, fine-tuning with extra data, a common practice to gain speciality, often l…

2024

Boosting Spike Camera Image Reconstruction from a Perspective of Dealing with Spike Fluctuations

CVPR 2024poster

As a bio-inspired vision sensor with ultra-high speed spike cameras exhibit great potential in recording dynamic scenes with high-speed motion or drastic light changes. Different from traditional cameras each pixel in spike cameras records the arrival of photons continuously by firing binary spikes…

2024

CLEAR: Can Language Models Really Understand Causal Graphs?

EMNLP 2024finding

Causal reasoning is a cornerstone of how humans interpret the world. To model and reason about causality, causal graphs offer a concise yet effective solution. Given the impressive advancements in language models, a crucial question arises: can they really understand causal graphs? To this end, we p…

2024

CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models

EMNLP 2024main

Large Language Models (LLMs) excel in diverse tasks but often underperform in specialized fields due to limited domain-specific or proprietary corpus. Continual pre-training (CPT) enhances LLM capabilities by imbuing new domain-specific or proprietary knowledge while replaying general corpus to prev…

Cited by 1SourcePDFScholar
2024

Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal Alignment

AAAI 2024technical

Sign language translation (SLT) aims to convert continuous sign language videos into textual sentences. As a typical multi-modal task, there exists an inherent modality gap between sign language videos and spoken language text, which makes the cross-modal alignment between visual and textual modalit…

2024

Deep Reinforcement Learning-Based Large-Scale Robot Exploration

RA-L 2024

In this work, we propose a deep reinforcement learning (DRL) based reactive planner to solve large-scale Lidar-based autonomous robot exploration problems in 2D action space. Our DRL-based planner allows the agent to reactively plan its exploration path by making implicit predictions about unknown a

Cited by 43SourcecodeScholar
2024

Drag Anything: Motion Control for Anything using Entity Representation

ECCV 2024poster

"We introduce , which utilizes a entity representation to achieve motion control for any object in controllable video generation. Comparison to existing motion control methods, offers several advantages. Firstly, trajectory-based is more user-friendly for interaction, when acquiring other guidance s…

2024

DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing

CVPR 2024poster

Despite recent progress in diffusion-based video editing existing methods are limited to short-length videos due to the contradiction between long-range consistency and frame-wise editing. Prior attempts to address this challenge by introducing video-2D representations encounter significant difficul…

2024

Eliminating Feature Ambiguity for Few-Shot Segmentation

ECCV 2024poster

"Recent advancements in few-shot segmentation (FSS) have exploited pixel-by-pixel matching between query and support features, typically based on cross attention, which selectively activate query foreground (FG) features that correspond to the same-class support FG features. However, due to the larg…

2024

Estimating Noisy Class Posterior with Part-level Labels for Noisy Label Learning

CVPR 2024poster

In noisy label learning estimating noisy class posteriors plays a fundamental role for developing consistent classifiers as it forms the basis for estimating clean class posteriors and the transition matrix. Existing methods typically learn noisy class posteriors by training a classification model w…

2024

EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models

NeurIPS 2024poster

Recent advancements in generation models have showcased remarkable capabilities in generating fantastic content. However, most of them are trained on proprietary high-quality data, and some models withhold their parameters and only provide accessible application programming interfaces (APIs), limiti…

2024

Focus-Then-Decide: Segmentation-Assisted Reinforcement Learning

AAAI 2024technical

Visual Reinforcement Learning (RL) is a promising approach to achieve human-like intelligence. However, it currently faces challenges in learning efficiently within noisy environments. In contrast, humans can quickly identify task-relevant objects in distraction-filled surroundings by applying previ…

2024

Gradient-based Visual Explanation for Transformer-based CLIP

ICML 2024poster

Significant progress has been achieved on the improvement and downstream usages of the Contrastive Language-Image Pre-training (CLIP) vision-language model, while less attention is paid to the interpretation of CLIP. We propose a Gradient-based visual Explanation method for CLIP (Grad-ECLIP), which…

2024

Hybrid Mamba for Few-Shot Segmentation

NeurIPS 2024poster

Many few-shot segmentation (FSS) methods use cross attention to fuse support foreground (FG) into query features, regardless of the quadratic complexity. A recent advance Mamba can also well capture intra-sequence dependencies, yet the complexity is only linear. Hence, we aim to devise a cross (atte…

2024

Instruct-ReID: A Multi-purpose Person Re-identification Task with Instructions

CVPR 2024poster

Human intelligence can retrieve any person according to both visual and language descriptions. However the current computer vision community studies specific person re-identification (ReID) tasks in different scenarios separately which limits the applications in the real world. This paper strives to…

2024

InstructDET: Diversifying Referring Object Detection with Generalized Instructions

ICLR 2024poster

We propose InstructDET, a data-centric method for referring object detection (ROD) that localizes target objects based on user instructions. While deriving from referring expressions (REC), the instructions we leverage are greatly diversified to encompass common user intentions related to object det…

2024

Learning Highly Dynamic Behaviors for Quadrupedal Robots

ICRA 2024poster

Learning highly dynamic behaviors for robots has been a longstanding challenge. Traditional approaches have demonstrated robust locomotion, but the exhibited behaviors lack diversity and agility. They employ approximate models, which lead to compromises in performance. Data-driven approaches have be…

Cited by 5SourceScholar
2024

MotionDirector: Motion Customization of Text-to-Video Diffusion Models

ECCV 2024oral

"Large-scale pre-trained diffusion models have exhibited remarkable capabilities in diverse video generations. Given a set of video clips of the same motion concept, the task of Motion Customization is to adapt existing text-to-video diffusion models to generate videos with this motion. Adaptation m…

2024

Non-Neighbors Also Matter to Kriging: A New Contrastive-Prototypical Learning

AISTATS 2024poster

Kriging aims to estimate the attributes of unseen geo-locations from observations in the spatial vicinity or physical connections. Existing works assume that neighbors’ information offers the basis for estimating the unobserved target while ignoring non-neighbors. However, neighbors could also be qu…

2024

Optical Flow for Spike Camera with Hierarchical Spatial-Temporal Spike Fusion

AAAI 2024technical

As an emerging neuromorphic camera with an asynchronous working mechanism, spike camera shows good potential for high-speed vision tasks. Each pixel in spike camera accumulates photons persistently and fires a spike whenever the accumulation exceeds a threshold. Such high-frequency fine-granularity…

2024

Revisiting Domain-Adaptive Object Detection in Adverse Weather by the Generation and Composition of High-Quality Pseudo-Labels

ECCV 2024poster

"Due to data collection challenges, the mean-teacher learning paradigm has emerged as a critical approach for cross-domain object detection, especially in adverse weather conditions. Despite significant progress, existing methods are still plagued by low-quality pseudo-labels in degraded images. Thi…

2024

Reward Difference Optimization For Sample Reweighting In Offline RLHF

EMNLP 2024finding

With the wide deployment of Large Language Models (LLMs), aligning LLMs with human values becomes increasingly important. Although Reinforcement Learning with Human Feedback (RLHF) proves effective, it is complicated and highly resource-intensive. As such, offline RLHF has been introduced as an alte…

2024

SDA: Simple Discrete Augmentation for Contrastive Sentence Representation Learning

COLING 2024main

Contrastive learning has recently achieved compelling performance in unsupervised sentence representation. As an essential element, data augmentation protocols, however, have not been well explored. The pioneering work SimCSE resorting to a simple dropout mechanism (viewed as continuous augmentation…

2024

Self-Supervised Representation Learning from Arbitrary Scenarios

CVPR 2024poster

Current self-supervised methods can primarily be categorized into contrastive learning and masked image modeling. Extensive studies have demonstrated that combining these two approaches can achieve state-of-the-art performance. However these methods essentially reinforce the global consistency of co…

Cited by 1SourcePDFScholar
2024

Sequential Asynchronous Action Coordination in Multi-Agent Systems: A Stackelberg Decision Transformer Approach

ICML 2024poster

Asynchronous action coordination presents a pervasive challenge in Multi-Agent Systems (MAS), which can be represented as a Stackelberg game (SG). However, the scalability of existing Multi-Agent Reinforcement Learning (MARL) methods based on SG is severely restricted by network architectures or env…

Cited by 5SourcePDFScholar
2024

Signer Diversity-driven Data Augmentation for Signer-Independent Sign Language Translation

NAACL 2024findings

The primary objective of sign language translation (SLT) is to transform sign language videos into natural sentences.A crucial challenge in this field is developing signer-independent SLT systems which requires models to generalize effectively to signers not encountered during training.This challeng…

Cited by 2SourcePDFScholar
2024

Sparse Global Matching for Video Frame Interpolation with Large Motion

CVPR 2024poster

Large motion poses a critical challenge in Video Frame Interpolation (VFI) task. Existing methods are often constrained by limited receptive fields resulting in sub-optimal performance when handling scenarios with large motion. In this paper we introduce a new pipeline for VFI which can effectively…

Cited by 14SourcePDFScholar
2024

T-SOT FNT: Streaming Multi-Talker ASR with Text-Only Domain Adaptation Capability

ICASSP 2024accepted

Token-level serialized output training (t-SOT) was recently proposed to address the challenge of streaming multi-talker automatic speech recognition (ASR). T-SOT effectively handles overlapped speech by representing multi-talker transcriptions as a single token stream with ⟨cc⟩ symbols interspersed.…

Cited by 0SourceScholar
2024

TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Industry Systems

EMNLP 2024industry

Large Language Models (LLMs) have demonstrated proficiency in addressing tasks that necessitate a combination of task planning and the usage of external tools, such as weather and calculator APIs. However, real-world industrial systems present prevalent challenges in task planning and tool usage: nu…

2024

VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point Correspondence

CVPR 2024poster

Current diffusion-based video editing primarily focuses on structure-preserved editing by utilizing various dense correspondences to ensure temporal consistency and motion alignment. However these approaches are often ineffective when the target edit involves a shape change. To embark on video editi…

Cited by 37SourcePDFScholar
2024

X-Adapter: Adding Universal Compatibility of Plugins for Upgraded Diffusion Model

CVPR 2024poster

We introduce X-Adapter a universal upgrader to enable the pretrained plug-and-play modules (e.g. ControlNet LoRA) to work directly with the upgraded text-to-image diffusion model (e.g. SDXL) without further retraining. We achieve this goal by training an additional network to control the frozen upgr…

2024

X-Light: Cross-City Traffic Signal Control Using Transformer on Transformer as Meta Multi-Agent Reinforcement Learner

IJCAI 2024poster

The effectiveness of traffic light control has been significantly improved by current reinforcement learning-based approaches via better cooperation among multiple traffic lights. However, a persisting issue remains: how to obtain a multi-agent traffic signal control algorithm with remarkable transf…

2023

Balancing Logit Variation for Long-Tailed Semantic Segmentation

CVPR 2023poster

Semantic segmentation usually suffers from a long tail data distribution. Due to the imbalanced number of samples across categories, the features of those tail classes may get squeezed into a narrow area in the feature space. Towards a balanced feature distribution, we introduce category-wise variat…

2023

CORA: Adapting CLIP for Open-Vocabulary Detection With Region Prompting and Anchor Pre-Matching

CVPR 2023poster

Open-vocabulary detection (OVD) is an object detection task aiming at detecting objects from novel categories beyond the base categories on which the detector is trained. Recent OVD methods rely on large-scale visual-language pre-trained models, such as CLIP, for recognizing novel objects. We identi…

2023

Contextual Image Masking Modeling via Synergized Contrasting without View Augmentation for Faster and Better Visual Pretraining

ICLR 2023poster

We propose a new contextual masking image modeling (MIM) approach called contrasting-aided contextual MIM (ccMIM), under the MIM paradigm for visual pretraining. Specifically, we adopt importance sampling to select the masked patches with richer semantic information for reconstruction, instead of ra…

Cited by 22SourcePDFScholar
2023

Cycle-consistent Masked AutoEncoder for Unsupervised Domain Generalization

ICLR 2023poster

Self-supervised learning methods undergo undesirable performance drops when there exists a significant domain gap between training and testing scenarios. Therefore, unsupervised domain generalization (UDG) is proposed to tackle the problem, which requires the model to be trained on several different…

Cited by 7SourcePDFScholar
2023

DatasetDM: Synthesizing Data with Perception Annotations Using Diffusion Models

NeurIPS 2023poster

Current deep networks are very data-hungry and benefit from training on large-scale datasets, which are often time-consuming to collect and annotate. By contrast, synthetic data can be generated infinitely using generative models such as DALL-E and diffusion models, with minimal effort and cost. In…

2023

Deeply Coupled Cross-Modal Prompt Learning

ACL 2023findings

Recent advancements in multimodal foundation models (e.g., CLIP) have excelled in zero-shot generalization. Prompt tuning involved in the knowledge transfer from foundation models to downstream tasks has gained significant attention recently. Existing prompt-tuning methods in cross-modal learning, h…

2023

Described Object Detection: Liberating Object Detection with Flexible Expressions

NeurIPS 2023poster

Detecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper, we advance them to a more practical setting called *Described Object Detection* (DOD) by expanding category names to fle…

2023

Exploring Stochastic Autoregressive Image Modeling for Visual Representation

AAAI 2023technical

Autoregressive language modeling (ALM) has been successfully used in self-supervised pre-training in Natural language processing (NLP). However, this paradigm has not achieved comparable results with other self-supervised approaches in computer vision (e.g., contrastive learning, masked image modeli…

2023

Fast and Accurate Factorized Neural Transducer for Text Adaption of End-to-End Speech Recognition Models

ICASSP 2023accepted

Neural transducer is now the most popular end-to-end model for speech recognition, due to its naturally streaming ability. However, it is challenging to adapt it with text-only data. Factorized neural transducer (FNT) model was proposed to mitigate this problem. The improved adaptation ability of FN…

Cited by 0SourceScholar
2023

Human Preference Score: Better Aligning Text-to-Image Models with Human Preference

ICCV 2023poster

Recent years have witnessed a rapid growth of deep generative models, with text-to-image models gaining significant attention from the public. However, existing models often generate images that do not align well with human preferences, such as awkward combinations of limbs and facial expressions. T…

Cited by 122PDFcodeScholar
2023

HumanBench: Towards General Human-Centric Perception With Projector Assisted Pretraining

CVPR 2023poster

Human-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this…

2023

Learning to Super-resolve Dynamic Scenes for Neuromorphic Spike Camera

AAAI 2023technical

Spike camera is a kind of neuromorphic sensor that uses a novel ``integrate-and-fire'' mechanism to generate a continuous spike stream to record the dynamic light intensity at extremely high temporal resolution. However, as a trade-off for high temporal resolution, its spatial resolution is limited,…

Cited by 12SourcePDFScholar
2023

LongFNT: Long-Form Speech Recognition with Factorized Neural Transducer

ICASSP 2023accepted

Traditional automatic speech recognition (ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is more practical in real scenarios. Simply attending longer transcription history for a vanilla neural transducer model shows…

Cited by 0SourceScholar
2023

Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination

AAAI 2023technical

We study the problem of training a Reinforcement Learning (RL) agent that is collaborative with humans without using human data. Although such agents can be obtained through self-play training, they can suffer significantly from the distributional shift when paired with unencountered partners, such…

2023

MeGraph: Capturing Long-Range Interactions by Alternating Local and Hierarchical Aggregation on Multi-Scaled Graph Hierarchy

NeurIPS 2023poster

Graph neural networks, which typically exchange information between local neighbors, often struggle to capture long-range interactions (LRIs) within the graph. Building a graph hierarchy via graph pooling methods is a promising approach to address this challenge; however, hierarchical information pr…

2023

Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models

NeurIPS 2023poster

Public large-scale text-to-image diffusion models, such as Stable Diffusion, have gained significant attention from the community. These models can be easily customized for new concepts using low-rank adaptations (LoRAs). However, the utilization of multiple-concept LoRAs to jointly support multipl…

Cited by 178SourcePDFScholar
2023

PUnifiedNER: A Prompting-Based Unified NER System for Diverse Datasets

AAAI 2023technical

Much of named entity recognition (NER) research focuses on developing dataset-specific models based on data from the domain of interest, and a limited set of related entity types. This is frustrating as each new dataset requires a new model to be trained and stored. In this work, we present a ``vers…

2023

Patch-Level Contrasting without Patch Correspondence for Accurate and Dense Contrastive Representation Learning

ICLR 2023poster

We propose ADCLR: \underline{A}ccurate and \underline{D}ense \underline{C}ontrastive \underline{R}epresentation \underline{L}earning, a novel self-supervised learning framework for learning accurate and dense vision representation. To extract spatial-sensitive information, ADCLR introduces query pat…

Cited by 19SourcePDFScholar
2023

SparseMAE: Sparse Training Meets Masked Autoencoders

ICCV 2023poster

Masked Autoencoders (MAE) and its variants have proven to be effective for pretraining large-scale Vision Transformers (ViTs). However, small-scale models do not benefit from the pretraining mechanisms due to limited capacity. Sparse training is a method of transferring representations from large mo…

Cited by 5PDFcodeScholar
2023

Trust Your Partner's Friends: Hierarchical Cross-Modal Contrastive Pre-Training for Video-Text Retrieval

ICASSP 2023accepted

Video-text retrieval has greatly benefited from the massive web video in recent years, while the performance is still limited to the weak supervision from the uncurated data. In this work, we propose to leverage the well-represented information of each original modality and exploit complementary inf…

Cited by 0SourceScholar
2023

Uni6Dv2: Noise Elimination for 6D Pose Estimation

AISTATS 2023poster

Uni6D is the first 6D pose estimation approach to employ a unified backbone network to extract features from both RGB and depth images. We discover that the principal reasons of Uni6D performance limitations are Instance-Outside and Instance-Inside noise. Uni6D’s simple pipeline design inherently in…

Cited by 14SourcePDFScholar
2023

UniHCP: A Unified Model for Human-Centric Perceptions

CVPR 2023poster

Human-centric perceptions (e.g., pose estimation, human parsing, pedestrian detection, person re-identification, etc.) play a key role in industrial applications of visual models. While specific human-centric tasks have their own relevant semantic aspect to focus on, they also share the same underly…

2023

Unsupervised Optical Flow Estimation with Dynamic Timing Representation for Spike Camera

NeurIPS 2023poster

Efficiently selecting an appropriate spike stream data length to extract precise information is the key to the spike vision tasks. To address this issue, we propose a dynamic timing representation for spike streams. Based on multi-layers architecture, it applies dilated convolutions on temporal dime…

2023

What Makes Pre-trained Language Models Better Zero-shot Learners?

ACL 2023long

Current methods for prompt learning in zero-shot scenarios widely rely on a development set with sufficient human-annotated data to select the best-performing prompt template a posteriori. This is not ideal because in a real-world zero-shot scenario of practical relevance, no labelled data is availa…

2023

Zero-Shot Text-to-Parameter Translation for Game Character Auto-Creation

CVPR 2023poster

Recent popular Role-Playing Games (RPGs) saw the great success of character auto-creation systems. The bone-driven face model controlled by continuous parameters (like the position of bones) and discrete parameters (like the hairstyles) makes it possible for users to personalize and customize in-gam…

2022

Align Representations With Base: A New Approach to Self-Supervised Learning

CVPR 2022poster

Existing symmetric contrastive learning methods suffer from collapses (complete and dimensional) or quadratic complexity of objectives. Departure from these methods which maximize mutual information of two generated views, along either instance or feature dimension, the proposed paradigm introduces…

Cited by 30PDFScholar
2022

Counterfactual Intervention Feature Transfer for Visible-Infrared Person Re-identification

ECCV 2022poster

"Graph-based models have achieved great success in person re-identification tasks recently, which compute the graph topology structure (affinities) among different people first and then pass the information across them to achieve stronger features. But we find existing graph-based methods in the vis…

Cited by 51SourcePDFScholar
2022

Domain Invariant Masked Autoencoders for Self-Supervised Learning from Multi-Domains

ECCV 2022poster

"Generalizing learned representations across significantly different visual domains is a fundamental yet crucial ability of the human visual system. While recent self-supervised learning methods have achieved good performances with evaluation set on the same domain as the training set, they will hav…

Cited by 19SourcePDFScholar
2022

Feature Erasing and Diffusion Network for Occluded Person Re-Identification

CVPR 2022poster

Occluded person re-identification (ReID) aims at matching occluded person images to holistic ones across different camera views. Target Pedestrians (TP) are often disturbed by Non-Pedestrian Occlusions (NPO) and Non-Target Pedestrians (NTP). Previous methods mainly focus on increasing the model's ro…

Cited by 179PDFcodeScholar
2022

Learning Optical Flow from Continuous Spike Streams

NeurIPS 2022accept

Spike camera is an emerging bio-inspired vision sensor with ultra-high temporal resolution. It records scenes by accumulating photons and outputting continuous binary spike streams. Optical flow is a key task for spike cameras and their applications. A previous attempt has been made for spike-based…

2022

Learning from Future: A Novel Self-Training Framework for Semantic Segmentation

NeurIPS 2022accept

Self-training has shown great potential in semi-supervised learning. Its core idea is to use the model learned on labeled data to generate pseudo-labels for unlabeled samples, and in turn teach itself. To obtain valid supervision, active attempts typically employ a momentum teacher for pseudo-label…

2022

Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks

NeurIPS 2022accept

Visual tasks vary a lot in their output formats and concerned contents, therefore it is hard to process them with an identical structure. One main obstacle lies in the high-dimensional outputs in object-level visual tasks. In this paper, we propose an object-centric vision framework, Obj2Seq. Obj2Se…

2022

RECCraft System: Towards Reliable and Efficient Collective Robotic Construction

IROS 2022poster

This research presents a novel Collective Robotic Construction (CRC) system named RECCraft. The RECCraft hardware system is composed of the mobile manipulation vehicles, the cubic blocks, and the folding ramp blocks. Solid connection and easy removal of the blocks are achieved by an electropermanent…

Cited by 4SourceScholar
2022

Relative Contrastive Loss for Unsupervised Representation Learning

ECCV 2022poster

"Defining positive and negative samples is critical for learning visual variations of the semantic classes in an unsupervised manner. Previous methods either construct positive sample pairs as different data augmentations on the same image (i.e., single-instance-positive) or estimate a class prototy…

Cited by 3SourcePDFScholar
2022

Revisiting the Transferability of Supervised Pretraining: An MLP Perspective

CVPR 2022poster

The pretrain-finetune paradigm is a classical pipeline in visual learning. Recent progress on unsupervised pretraining methods shows superior transfer performance to their supervised counterparts. This paper revisits this phenomenon and sheds new light on understanding the transferability gap betwee…

Cited by 72PDFScholar
2022

Scale-Aware Spatio-Temporal Relation Learning for Video Anomaly Detection

ECCV 2022poster

"Recent progress in video anomaly detection (VAD) has shown that feature discrimination is the key to effectively distinguishing anomalies from normal events. We observe that many anomalous events occur in limited local regions, and the severe background noise increases the difficulty of feature lea…

2022

Semi-Supervised Semantic Segmentation Using Unreliable Pseudo-Labels

CVPR 2022poster

The crux of semi-supervised semantic segmentation is to assign pseudo-labels to the pixels of unlabeled images. A common practice is to select the highly confident predictions as the pseudo ground-truth, but it leads to a problem that most pixels may be left unused due to their unreliability. We arg…

Cited by 495PDFcodeScholar
2022

Spatio-Temporal Recurrent Networks for Event-Based Optical Flow Estimation

AAAI 2022technical

Event camera has offered promising alternative for visual perception, especially in high speed and high dynamic range scenes. Recently, many deep learning methods have shown great success in providing model-free solutions to many event-based problems, such as optical flow estimation. However, existi…

2022

Uni6D: A Unified CNN Framework Without Projection Breakdown for 6D Pose Estimation

CVPR 2022oral

As RGB-D sensors become more affordable, using RGB-D images to obtain high-accuracy 6D pose estimation results becomes a better option. State-of-the-art approaches typically use different backbones to extract features for RGB and depth images. They use a 2D CNN for RGB images and a per-pixel point c…

Cited by 53PDFScholar
2022

UniVIP: A Unified Framework for Self-Supervised Visual Pre-Training

CVPR 2022poster

Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the scene and instances, as well as the semantic difference of…

Cited by 41PDFScholar
2022

Unifying Visual Contrastive Learning for Object Recognition from a Graph Perspective

ECCV 2022poster

"Recent contrastive based unsupervised object recognition methods leverage a Siamese architecture, which has two branches composed of a backbone, a projector layer, and an optional predictor layer in each branch. To learn the parameters of the backbone, existing methods have a similar projector laye…

Cited by 8SourcePDFScholar
2022

Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector Learning

NeurIPS 2022accept

Unsupervised pretraining methods for object detection aim to learn object discrimination and localization ability from large amounts of images. Typically, recent works design pretext tasks that supervise the detector to predict the defined object priors. They normally leverage heuristic methods to p…

Cited by 5SourcePDFScholar
2022

Zero-CL: Instance and Feature decorrelation for negative-free symmetric contrastive learning

ICLR 2022poster

For self-supervised contrastive learning, models can easily collapse and generate trivial constant solutions. The issue has been mitigated by recent improvement on objective design, which however often requires square complexity either for the size of instances ($\mathcal{O}(N^{2})$) or feature dime…

Cited by 48SourcePDFScholar
2021

An Automated Framework for Supporting Data-Governance Rule Compliance in Decentralized MIMO Contexts

IJCAI 2021poster

We propose Dr.Aid, a logic-based AI framework for automated compliance checking of data governance rules over data-flow graphs. The rules are modelled using a formal language based on situation calculus and are suitable for decentralized contexts with multi-input-multi-output (MIMO) processes. Dr.Ai…

2021

MST: Masked Self-Supervised Transformer for Visual Representation

NeurIPS 2021poster

Transformer has been widely used for self-supervised pre-training in Natural Language Processing (NLP) and achieved great success. However, it has not been fully explored in visual self-supervised learning. Meanwhile, previous methods only consider the high-level feature and learning representation…

Cited by 180SourcePDFScholar
2021

Progressive Correspondence Pruning by Consensus Learning

ICCV 2021poster

Correspondence pruning aims to correctly remove false matches (outliers) from an initial set of putative correspondences. The selection is challenging since putative matches are typically extremely unbalanced, largely dominated by outliers, and the random distribution of such outliers further compli…

Cited by 91PDFScholar
2021

Structure-Enhanced Attentive Learning For Spine Segmentation From Ultrasound Volume Projection Images

ICASSP 2021accepted

Automatic spine segmentation, based on ultrasound volume projection imaging (VPI), is of great value in clinical applications to diagnose scoliosis in teenagers. In this paper, we propose a novel framework to improve the segmentation accuracy on spine images via structure-enhanced attentive learning…

Cited by 0SourceScholar
2020

Adaptation of RNN Transducer with Text-To-Speech Technology for Keyword Spotting

ICASSP 2020accepted

With the advent of recurrent neural network transducer (RNN-T) model, the performance of keyword spotting (KWS) systems has greatly improved. However, the KWS systems, employed for wake-word detection, still rely on the availability of keyword specific training data for achieving reasonable performa…

Cited by 0SourceScholar
2020

COCAS: A Large-Scale Clothes Changing Person Dataset for Re-Identification

CVPR 2020poster

Recent years have witnessed great progress in person re-identification (re-id). Several academic benchmarks such as Market1501, CUHK03 and DukeMTMC play important roles to promote the re-id research. To our best knowledge, all the existing benchmarks assume the same person will have the same clothes…

Cited by 110PDFScholar
2020

Exploring Pre-Training with Alignments for RNN Transducer Based End-to-End Speech Recognition

ICASSP 2020accepted

Recently, the recurrent neural network transducer (RNN-T) architecture has become an emerging trend in end-to-end automatic speech recognition research due to its advantages of being capable for online streaming speech recognition. However, RNN-T training is made difficult by the huge memory require…

Cited by 0SourceScholar
2020

High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model

ICASSP 2020accepted

While the community keeps promoting end-to-end models over conventional hybrid models, which usually are long short-term memory (LSTM) models trained with a cross entropy criterion followed by a sequence discriminative training criterion, we argue that such conventional hybrid models can still be si…

Cited by 0SourceScholar
2020

Learning to Cluster Faces via Confidence and Connectivity Estimation

CVPR 2020poster

Face clustering is an essential tool for exploiting the unlabeled face data, and has a wide range of applications including face annotation and retrieval. Recent works show that supervised clustering can result in noticeable performance gain. However, they usually involve heuristic steps and require…

Cited by 116PDFcodeScholar
2020

RBF-Softmax: Learning Deep Representative Prototypes with Radial Basis Function Softmax

ECCV 2020poster

Deep neural networks have achieved remarkable successes in learning feature representations for visual classification. However, deep features learned by the softmax cross-entropy loss generally show excessive intra-class variations. We argue that, because the traditional softmax losses aim to optimi…

2020

Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-ID

NeurIPS 2020poster

Domain adaptive object re-ID aims to transfer the learned knowledge from the labeled source domain to the unlabeled target domain to tackle the open-class re-identification problems. Although state-of-the-art pseudo-label-based methods have achieved great success, they did not make full use of all v…

2020

Self-supervising Fine-grained Region Similarities for Large-scale Image Localization

ECCV 2020poster

The task of large-scale retrieval-based image localization is to estimate the geographical location of a query image by recognizing its nearest reference images from a city-scale dataset. However, the general public benchmarks only provide noisy GPS labels associated with the training images, which…

2019

AdaCos: Adaptively Scaling Cosine Logits for Effectively Learning Deep Face Representations

CVPR 2019oral

The cosine-based softmax losses and their variants achieve great success in deep learning based face recognition. However, hyperparameter settings in these losses have significant influences on the optimization path as well as the final recognition performance. Manually tuning those hyperparameters…

Cited by 312PDFScholar
2019

P2SGrad: Refined Gradients for Optimizing Deep Face Models

CVPR 2019poster

Cosine-based softmax losses significantly improve the performance of deep face recognition networks. However, these losses always include sensitive hyper-parameters which can make training process unstable, and it is very tricky to set suitable hyper parameters for a specific dataset. This paper add…

Cited by 48PDFScholar
2018

A Hierarchical Generative Model for Eye Image Synthesis and Eye Gaze Estimation

CVPR 2018poster

In this work, we introduce a Hierarchical Generative Model (HGM) to enable realistic forward eye image synthe- sis, as well as effective backward eye gaze estimation. The proposed HGM consists of a hierarchical generative shape model (HGSM), and a conditional bidirectional generative adversarial net…

2018

Attention-Aware Compositional Network for Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is to identify pedestrians observed from different camera views based on visual appearance. It is a challenging task due to large pose variations, complex background clutters and severe occlusions. Recently, human pose estimation by predicting joint locations was larg…

Cited by 565SourcePDFScholar
2018

Bilateral Ordinal Relevance Multi-Instance Regression for Facial Action Unit Intensity Estimation

CVPR 2018poster

Automatic intensity estimation of facial action units (AUs) is challenging in two aspects. First, capturing subtle changes of facial appearance is quiet difficult. Second, the annotation of AU intensity is scarce and expensive. Intensity annotation requires strong domain knowledge thus only experts…

Cited by 53SourcePDFScholar
2018

Developing Far-Field Speaker System Via Teacher-Student Learning

ICASSP 2018accepted

In this study, we develop the keyword spotting (KWS) and acoustic model (AM) components in a far-field speaker system. Specifically, we use teacher-student (T/S) learning to adapt a close-talk well-trained production AM to far-field by using parallel close-talk and simulated far-field data. We also…

Cited by 0SourceScholar
2018

QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension

ICLR 2018poster

Current end-to-end machine reading and question answering (Q\&A) models are primarily based on recurrent neural networks (RNNs) with attention. Despite their success, these models are often slow for both training and inference due to the sequential nature of RNNs. We propose a new Q\&A architecture…

Cited by 1263SourceScholar
2016

Recurrent support vector machines for speech recognition

ICASSP 2016accepted

Recurrent Neural Networks (RNNs) using Long-Short Term Memory (LSTM) architecture have demonstrated the state-of-the-art performances on speech recognition. Most of deep RNNs use the softmax activation function in the last layer for classification. This paper illustrates small but consistent advanta…

Cited by 0SourceScholar