← Search

YI WANG

176 accepted papers

2026

A Survey of Artificial Intelligence in Endoscopic Surgery Workflow: From Perception to Surgical Support

IJCAI 2026

Endoscopic surgery demands continuous real-time visual decision-making under severe constraints, including a limited field of view, motion blur, and dynamically deforming anatomy. These factors impose substantial cognitive load on surgeons and motivate the integration of artificial intelligence (AI)

Cited by 0Scholar
2026

Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep

CVPR 2026

Diffusion-based video editing has emerged as an important paradigm for high-quality and flexible content generation. However, despite their generality and strong modeling capacity, Diffusion Transformers (DiT) remain computationally expensive due to the iterative denoising process, posing challenges

Cited by 0SourcecodeScholar
2026

Balancing the Experts: Unlocking LoRA-MoE for GRPO via Mechanism-Aware Rewards

ICLR 2026poster

Parameter-efficient Mixture-of-Experts (MoE) architectures, such as LoRA-MoE, enable strong and generalizable fine-tuning. However, a critical problem arises when fine-tuning these architectures with advanced reinforcement learning algorithms such as Group Relative Policy Optimization (GRPO). Tradit…

Cited by 0SourceScholar
2026

Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors

CVPR 2026

Self-supervised contrastive learning has emerged as a powerful paradigm for skeleton-based action recognition by enforcing consistency in the embedding space. However, existing methods rely on binary contrastive objectives that overlook the intrinsic continuity of human motion, resulting in fragment

Cited by 0SourceScholar
2026

BiPA: Bilevel Prompt Adaptation for Underwater Instance Segmentation

CVPR 2026

Underwater instance segmentation is essential for fine-grained scene understanding. However, underwater imagery exhibits a strong domain gap from in-air vision due to severe degradation (e.g., turbidity). Consequently, despite its general segmentation ability, SAM degrades sharply underwater. In thi

Cited by 0SourcecodeScholar
2026

Bridging Inter-View and Client Heterogeneity: Federated Multi-View Clustering Under Non-IID Data

IJCAI 2026

Federated multi-view clustering (FedMVC) has been widely used to discover latent structures in distributed multi-view data, but most methods assume independent and identically distributed (IID) data. In practice, non-IID distributions with partial and imbalanced categories cause clients to learn bia

Cited by 0Scholar
2026

CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation

AAAI 2026technical

While previous multimodal slow-thinking methods have demonstrated remarkable success in single-image understanding scenarios, their effectiveness becomes fundamentally constrained when extended to more complex multi-image comprehension tasks. This limitation stems from their predominant reliance on

Cited by 0SourcePDFScholar
2026

ConstructAI: From Real-Time Safety Insight to Skill Growth in Deployed Construction AI Systems

AAAI 2026technical

Ensuring safety in power grid construction remains a critical yet challenging task, as existing monitoring approaches often lack scalability, timeliness, and adaptability to diverse on-site conditions. To address these limitations, we present ConstructAI, a deployed AI-driven safety management syste

Cited by 0SourcePDFScholar
2026

Cross-Timestep: 3D Diffusion Model with Trans-temporal Memory LSTM and Adaptive Priori Decoding Strategy for Medical Segmentation

ICLR 2026poster

Diffusion models have recently demonstrated significant robustness in medical image segmentation, effectively accommodating variations across different imaging styles. However, their applications remain limited due to: (i) current successes being primarily confined to 2D segmentation tasks—we observ…

Cited by 0SourceScholar
2026

Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection

AAAI 2026technical

Existing core-set selection methods predominantly rely on heuristic scoring signals such as training dynamics or model uncertainty, lacking explicit modeling of data likelihood. This omission may hinder the constructed subset from capturing subtle yet critical distributional structures that underpin

Cited by 0SourcePDFScholar
2026

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding

ICML 2026poster

A precise and comprehensive understanding of human-environment interactions in egocentric vision is essential for next-generation intelligent agents, such as assistive robotics. While existing multimodal large language models (MLLMs) support unified reasoning from scene-level analysis to instance-sp…

Cited by 0SourceScholar
2026

Efficient Time Series Clustering from Multiscale Reservoir Dynamics with Granular-Ball Anchoring Graph Optimization

IJCAI 2026

Time-series clustering remains challenging due to the inherent trade-off between clustering effectiveness and computational efficiency. Similarity-based methods often suffer from quadratic complexity caused by pairwise distance computations, while deep learning–based approaches typically rely on cos

Cited by 0Scholar
2026

EigenScore: OOD Detection using Posterior Covariance in Diffusion Models

ICLR 2026poster

Out-of-distribution (OOD) detection is critical for the safe deployment of machine learning systems in safety-sensitive domains. Diffusion models have recently emerged as powerful generative models, capable of capturing complex data distributions through iterative denoising. Building on this progres…

Cited by 6SourceScholar
2026

ExpVid: A Benchmark for Experiment Video Understanding & Reasoning

ICLR 2026poster

Multimodal Large Language Models (MLLMs) hold promise for accelerating scientific discovery by interpreting complex experimental procedures. However, their true capabilities are poorly understood, as existing benchmarks neglect the fine-grained and long-horizon nature of authentic laboratory work, e…

Cited by 0SourcecodeScholar
2026

FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free Memory

AAAI 2026technical

Text-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation resu

Cited by 0SourcePDFScholar
2026

GRL-SNAM: Geometric Reinforcement Learning with Differential Hamiltonians for Navigation and Mapping in Unknown Environments

ICLR 2026poster

We present GRL-SNAM, a geometric reinforcement learning framework for Simultaneous Navigation and Mapping in unknown environments. GRL-SNAM differs from traditional SLAM and other reinforcement learning methods by relying exclusively on local sensory observations without constructing a global map. O…

Cited by 0SourcecodeScholar
2026

HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding

CVPR 2026

Humans commonly reason about object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. Inspired by this principle, we advocate for a novel framework that leverages emerging multimodal large language models (M

Cited by 0SourceScholar
2026

HO-SFL: Hybrid-Order Split Federated Learning with Backprop-Free Clients and Dimension-Free Aggregation

ICML 2026poster

Fine-tuning large models on edge devices is severely hindered by the memory-intensive backpropagation (BP) in standard frameworks like federated learning and split learning. While substituting BP with zeroth-order optimization can significantly reduce memory footprints, it typically suffers from pro…

Cited by 0SourceScholar
2026

HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

CVPR 2026

Recent methods have made notable progress in the visual quality of hand-object interaction video synthesis. However, most approaches rely on 2D control signals that lack spatial expressiveness and limit the utilization of synthetic 3D conditional data. To address these limitations, we propose HVG-3D

Cited by 0SourceScholar
2026

InterCoser: Interactive 3D Character Creation with Disentangled Fine-Grained Features

AAAI 2026technical

This paper aims to interactively generate and edit disentangled 3D characters based on precise user instructions. Existing methods generate and edit 3D characters via rough and simple editing guidance and entangled representations, making it difficult to achieve precise and comprehensive control ove

Cited by 0SourcePDFScholar
2026

Interaction-aware Representation Modeling With Co-Occurrence Consistency for Egocentric Hand-Object Parsing

ICLR 2026poster

A fine-grained understanding of egocentric human-environment interactions is crucial for developing next-generation embodied agents. One fundamental challenge in this area involves accurately parsing hands and active objects. While transformer-based architectures have demonstrated considerable poten…

Cited by 0SourcecodeScholar
2026

InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models

ICLR 2026poster

Recent benchmarks and datasets have been proposed to improve spatial reasoning in vision-language models (VLMs), yet existing open resources remain limited in scale, visual diversity, and instruction expressiveness. In this work, we introduce InternSpatial, the largest open-source dataset for spatia…

Cited by 0SourceScholar
2026

InternVideo-Next: Towards World-Understanding Video Models

CVPR 2026

Large-scale video-text pretraining achieves strong performance but depends on noisy, synthetic captions with limited semantic coverage, often overlooking implicit world knowledge such as object motion, 3D geometry, and physical cues. In contrast, masked video modeling (MVM) directly exploits spatiot

Cited by 0SourcecodeScholar
2026

MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement Learning

CVPR 2026

Offline Multi-Agent Reinforcement Learning (MARL) is critical for coordinating multiple agents in costly and unsafe environments, yet existing methods struggle with high sensitivity to reward functions and weak generalization to new goals, limiting its practical impact. Inspired by single-agent Offl

Cited by 0SourceScholar
2026

On the Power of Statistics in Class-Incremental Learning with Pretrained Models

ICML 2026poster

Recent class-incremental learning (CIL) methods built on large pre-trained vision models have shown that strong performance can be retained even under strict data access constraints. This raises a fundamental question: which properties of pre-trained representations make such recovery possible in th…

Cited by 0SourceScholar
2026

Open-Vocabulary Spatio-Temporal Scene Graph for Robot Perception and Teleoperation Planning

ICRA 2026poster

Teleoperation via natural-language reduces operator workload and enhances safety in high-risk or remote settings. However, in dynamic remote scenes, transmission latency during bidirectional communication creates gaps between remote perceived states and operator intent, leading to command misunderst…

2026

P2S: Probabilistic Process Supervision for General-Domain Reasoning Question Answering

AAAI 2026technical

While reinforcement learning with verifiable rewards (RLVR) has advanced LLM reasoning in structured domains like mathematics and programming, its application to general-domain reasoning tasks remains challenging due to the absence of verifiable reward signals. To this end, methods like Reinforcemen

Cited by 0SourcePDFScholar
2026

Permutation Equivariant Framelet-based Hypergraph Neural Networks

AAAI 2026technical

Hypergraphs provide a natural and expressive framework for modeling high-order relationships, enabling the representation of group-wise interactions beyond pairwise connections. While hypergraph neural networks (HNNs) have shown promise for learning on such structures, existing models often rely on

Cited by 0SourcePDFScholar
2026

REFO: Reinforced Evolutionary Faithfulness Optimization for Large Language Models

AAAI 2026technical

Despite its success in enriching LLMs with external knowledge, RAG remains plagued by faithfulness hallucinations, where generated text contradicts the retrieved source information. Previous research on faithfulness hallucination in LLMs is frequently hindered by prohibitive manual annotation costs

Cited by 0SourcePDFScholar
2026

RIVER: Real-time Video Interaction Benchmark

ICLR 2026poster

The rapid advancement of multimodal large language models has demonstrated impressive capabilities, yet nearly all operate in an offline paradigm, hindering real-time interactivity. Addressing this gap, we introduce the Real-tIme Video intERaction Bench (RIVER Bench), designed for evaluating online…

Cited by 0SourcecodeScholar
2026

Rethinking Token Reduction for Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) excel in visual understanding and reasoning, but the excessive visual tokens lead to high inference costs. Although recent token reduction methods mitigate this issue, they mainly target single-turn Visual Question Answering (VQA), leaving the more practical mult

Cited by 0SourcecodeScholar
2026

Richer Representations for Neural Algorithmic Reasoning via Auxiliary Reconstruction

AAAI 2026technical

Neural algorithmic reasoning has recently emerged as a popular research direction. It aims to train neural networks to mimic the step-by-step behavior of classical rule-based algorithms. More specifically, the execution of such algorithms can be abstracted as a sequence of states, where each state

Cited by 0SourcePDFScholar
2026

SSTODE: Ocean-Atmosphere Physics-Informed Neural ODEs for Sea Surface Temperature Prediction

AAAI 2026technical

Sea Surface Temperature (SST) is crucial for understanding upper-ocean thermal dynamics and ocean-atmosphere interactions, which have profound economic and social impacts. While data-driven models show promise in SST prediction, their black-box nature often limits interpretability and overlooks key

Cited by 0SourcePDFScholar
2026

Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation

AAAI 2026technical

Textile pattern generation (TPG) aims to synthesize fine-grained textile pattern images based on given clothing images. Although previous studies have not explicitly investigated TPG, existing image-to-image models appear to be natural candidates for this task. However, when applied directly, these

Cited by 0SourcePDFScholar
2026

Think Then Rewrite: Reasoning Enhanced Query Rewriting for Domain Specific Retrieval

AAAI 2026technical

Query rewriting is a crucial task for improving retrieval, especially in professional domains such as law and medicine, where user queries are often underspecified and ambiguous. While large language models (LLMs) offer strong understanding and generation capabilities, existing LLM-based approaches

Cited by 0SourcePDFScholar
2026

Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding

CVPR 2026

Long video understanding is essential for human-like intelligence, enabling coherent perception and reasoning over extended temporal contexts. While the emerging thinking-with-frames paradigm--which alternates between global temporal reasoning and local frame examination--has advanced the reasoning

Cited by 0SourceScholar
2026

Towards an Early Warning System for Ocean Heat Extremes Through AI-Ocean Dynamics Synergy

IJCAI 2026

Ocean heat extremes, including marine heatwaves and the El Ni\~no–Southern Oscillation (ENSO), exert profound impacts on marine ecosystems and socio-economic stability. Establishing robust early warning systems is critical for proactive risk management; however, conventional predictive models often

Cited by 0Scholar
2026

TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning

AAAI 2026technical

Tourism and travel planning increasingly rely on digital assistance, yet existing multimodal AI systems often lack specialized knowledge and contextual understanding of urban environments. We present TraveLLaMA, a specialized multimodal language model designed for comprehensive travel assistance. Ou

Cited by 0SourcePDFScholar
2026

UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation

ICLR 2026poster

Tokenizer is a crucial component for both visual understanding and generation. To advance toward the ultimate goal of universal modeling, recent research has focused on developing a unified tokenizer. However, existing tokenizers face a significant performance trade-off between understanding and gen…

Cited by 0SourceScholar
2026

Urban Socio-Semantic Segmentation with Vision-Language Reasoning

ICLR 2026poster

As hubs of human activity, urban surfaces consist of a wealth of semantic entities. Segmenting these various entities from satellite imagery is crucial for a range of downstream applications. Current advanced segmentation models can reliably segment entities defined by physical attributes (e.g., bui…

Cited by 0SourcecodeScholar
2026

VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning

AAAI 2026technical

Recent advances in video understanding have been driven by MLLMs. But these MLLMs are good at analyzing short videos, while suffering from difficulties in understanding videos with a longer context. To address this difficulty, several agent paradigms have recently been proposed, using MLLMs as agen

Cited by 0SourcePDFScholar
2026

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

ICLR 2026poster

Long-context video modeling is critical for multimodal large language models (MLLMs), enabling them to process movies, online video streams, and so on. Despite its advances, handling long videos remains challenging due to the difficulty in efficiently understanding the extremely long video context.…

Cited by 0SourcecodeScholar
2026

VideoSeeker: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

ICML 2026poster

Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting their ability to identify sparse yet critical evidence amid extensive redundancy. We introduce VideoSeeker, a novel framework that supports iterative disc…

Cited by 13SourceScholar
2025

2.5D Top-K Ranked Multiple Instance Learning to Classify NSCLC PD-L1 Status on CT Images

ICASSP 2025accepted

Classifying the status of NSCLC PD-L1 on chest CT is a cost-effective and non-invasive method. The existing multiple instance learning (MIL) methods are not effective for this task, due to the lack of an efficient feature encoder for 3D instances and ignoring the importance of representative instanc…

Cited by 0SourceScholar
2025

A Deep Learning-Driven Autonomous System for Retinal Vein Cannulation: Validation Using a Chicken Embryo Model

IROS 2025

Retinal vein cannulation (RVC) is a minimally invasive microsurgical procedure for treating retinal vein occlusion (RVO), a leading cause of vision impairment. However, the small size and fragility of retinal veins, coupled with the need for high-precision, tremor-free needle manipulation, create si

Cited by 0SourceScholar
2025

A Non-isotropic Time Series Diffusion Model with Moving Average Transitions

ICML 2025poster

Diffusion models, known for their generative ability, have recently been adapted to time series analysis. Most pioneering works rely on the standard isotropic diffusion, treating each time step and the entire frequency spectrum identically. However, it may not be suitable for time series, which ofte…

Cited by 0SourcePDFScholar
2025

A Robust Quality Evaluator for Panoramic Videos

ICASSP 2025accepted

Most of the existing methods to evaluate the quality of panoramic content mainly focus on studying the quality evaluation of static panoramic images, rather than the more widely used dynamic panoramic videos. Also, the few panoramic video quality metrics that are available have obvious weaknesses in…

Cited by 0SourceScholar
2025

ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction

CVPR 2025poster

Egocentric interaction perception is one of the essential branches in investigating human-environment interaction, which lays the basis for developing next-generation intelligent systems. However, existing egocentric interaction understanding methods cannot yield coherent textual and pixel-level res…

Cited by 0SourcePDFScholar
2025

Adaptive Learning of High-Value Regions for Semi-Supervised Medical Image Segmentation

ICCV 2025poster

Existing semi-supervised learning methods typically mitigate the impact of unreliable predictions by suppressing low-confidence regions. However, these methods fail to explore which regions hold higher learning value and how to design adaptive learning strategies for these regions. To address these…

2025

All Roads Lead to Rome: Exploring Edge Distribution Shifts for Heterophilic Graph Learning

IJCAI 2025

Heterophilic graph neural networks (GNNs) have gained prominence for their ability to learn effective representations in graphs with diverse, attribute-aware relationships. While existing methods leverage attribute inference during message passing to improve performance, they often struggle with cha

Cited by 0SourcePDFScholar
2025

Asymptotically Optimal Sampling-Based Motion Planning Through Anytime Incremental Lazy Bidirectional Heuristic Search

ICRA 2025

This paper introduces Bidirectional Lazy Informed Trees (BLIT*), the first algorithm to incorporate anytime incremental lazy bidirectional heuristic search (Bi-HS) into batch-wise sampling-based motion planning (Bw-SBMP). BLIT* operates on batches of informed states (states that can potentially impr

Cited by 1SourceScholar
2025

Bidirectional Search while Ensuring Meet-In-The-Middle via Effective and Efficient-to-Compute Termination Conditions

IJCAI 2025

In bidirectional heuristic search, the meeting-in-the-middle property (MMP) and the theory of must-expand pairs (MEP) have driven significant recent developments in search efficiency. However, these methodologies typically terminate the search based on minimal priority metrics in the forward and bac

2025

Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel

ICLR 2025poster

Creating high-quality data for training robust language-instructed agents is a long-lasting challenge in embodied AI. In this paper, we introduce a Self-Refining Data Flywheel (SRDF) that generates high-quality and large-scale navigational instruction-trajectory pairs by iteratively refining the dat…

2025

CSD: Weather forecasting with graph neural network based on cross-scale diffusivity

ICASSP 2025accepted

Automated weather stations play a pivotal role in fine-grained weather forecasting, due to their cost-effectiveness and global deployment potential. Data-driven methods, particularly deep learning techniques, have emerged as potent tools for precise weather forecasting. However, capturing meaningful…

Cited by 0SourceScholar
2025

CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding

NeurIPS 2025oral

Coral reefs are vital yet vulnerable ecosystems that require continuous monitoring to support conservation. While coral reef images provide essential information in coral monitoring, interpreting such images remains challenging due to the need for domain expertise. Visual Question Answering (VQA), p…

Cited by 0SourceScholar
2025

DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing

ACL 2025finding

Large Language Models (LLMs) are widely applied in decision making, but their deployment is threatened by jailbreak attacks, where adversarial users manipulate model behavior to bypass safety measures. Existing defense mechanisms, such as safety fine-tuning and model editing, either require extensiv…

2025

Deep Hypergraph Neural Networks with Tight Framelets

AAAI 2025technical

Hypergraphs provide a flexible framework for modeling high-order (complex) interactions among multiple entities, extending beyond traditional pairwise correlations in graph structures. However, deep hypergraph neural networks (HGNNs) often face the challenge of oversmoothing with increasing depth, s…

Cited by 1SourcePDFScholar
2025

DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex Degradations

ICCV 2025poster

Diffusion models have demonstrated exceptional capabilities in image restoration, yet their application to video super-resolution (VSR) faces significant challenges in balancing fidelity with temporal consistency. Our evaluation reveals a critical gap: existing approaches consistently fail on severe…

Cited by 0SourcePDFScholar
2025

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

ICCV 2025poster

In video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input. While linear projectors are widely employed for encapsulation, they introduce semantic indistinctness and temporal inc…

2025

Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models

CVPR 2025poster

Concept Bottleneck Models (CBMs) try to make the decision-making process transparent by exploring an intermediate concept space between the input image and the output prediction. Existing CBMs just learn coarse-grained relations between the whole image and the concepts, less considering local image…

Cited by 1SourcePDFScholar
2025

Efficient Knowledge Transfer in Federated Recommendation for Joint Venture Ecosystem

NeurIPS 2025spotlight

The current Federated Recommendation System (FedRS) focuses on personalized recommendation services and assumes clients are personalized IoT devices (e.g., Mobile phones). In this paper, we deeply dive into new but practical FedRS applications within the joint venture ecosystem. Subsidiaries engage…

Cited by 0SourceScholar
2025

Enhancing Privacy in Multimodal Federated Learning with Information Theory

NeurIPS 2025poster

Multimodal federated learning (MMFL) has gained increasing popularity due to its ability to leverage the correlation between various modalities, meanwhile preserving data privacy for different clients. However, recent studies show that correlation between modalities increase the vulnerability of fed…

Cited by 0SourceScholar
2025

Enhancing Vision-Language Models with Morphological and Taxonomic Knowledge: Towards Coral Recognition for Ocean Health

AAAI 2025technical

Coral reefs play a crucial role in marine ecosystems, offering a nutrient-rich environment and safe shelter for numerous marine species. Automated coral image recognition aids in monitoring ocean health at a scale without experts' manual effort. Recently, large vision-language models like CLIP have…

Cited by 0SourcePDFScholar
2025

FAPEX: Fractional Amplitude-Phase Expressor for Robust Cross-Subject Seizure Prediction

NeurIPS 2025spotlight

Precise, generalizable subject-agnostic seizure prediction (SASP) remains a fundamental challenge due to the intrinsic complexity and significant spectral variability of electrophysiologial signals across individuals and recording modalities. We propose \model{FAPEX}, a novel architecture that intro…

Cited by 0SourceScholar
2025

FedSe: Group-Based Sequential Training Strategies for Mitigating Label Skew in Federated Learning

ICASSP 2025accepted

Federated Learning (FL) has emerged as a promising approach for distributed machine learning, enabling clients to collaboratively train models without sharing their data. However, existing FL methods continue to face challenges when dealing with non-IID data, particularly under conditions of extreme…

Cited by 0SourceScholar
2025

FlashGS: Efficient 3D Gaussian Splatting for Large-scale and High-resolution Rendering

CVPR 2025poster

Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated significant potential over traditional rendering techniques, attracting widespread attention from both industry and academia. However, real-time rendering with 3DGS remains a challenging problem, particularly in large-scale, high-reso…

2025

GraphVCM: Virtual Center Mixing with Distance-Aware Regulation for Class Imbalanced Node Classification

ICASSP 2025accepted

Class imbalance is a prevalent issue in real-world graph-structure data, such as social and citation networks, posing significant challenges for Graph Neural Networks (GNNs). Existing solutions often focus on balancing class distributions via oversampling techniques, which may lead to overfitting an…

Cited by 0SourceScholar
2025

HyperNear: Unnoticeable Node Injection Attacks on Hypergraph Neural Networks

ICML 2025poster

With the growing adoption of Hypergraph Neural Networks (HNNs) to model higher-order relationships in complex data, concerns about their security and robustness have become increasingly important. However, current security research often overlooks the unique structural characteristics of hypergraph…

2025

INDOORWORLD : Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment

EMNLP 2025

Virtual environments are essential to AI agent research. Existing environments for LLM agent research typically focus on either physical task solving or social simulation, with the former oversimplifying agent individuality and social dynamics, and the latter lacking physical grounding of social beh

Cited by 0SourcePDFScholar
2025

Influence-Guided Diffusion for Dataset Distillation

ICLR 2025poster

Dataset distillation aims to streamline the training process by creating a compact yet effective dataset for a much larger original dataset. However, existing methods often struggle with distilling large, high-resolution datasets due to prohibitive resource costs and limited performance, primarily s…

2025

MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

AAAI 2025technical

Auto-regressive models have made significant progress in the realm of text-to-image synthesis, yet devising an appropriate model architecture and training strategy to achieve a satisfactory level remains an important avenue of exploration. In this work, we introduce MARS, a novel framework for T2I g…

2025

ML-GOOD: Towards Multi-Label Graph Out-Of-Distribution Detection

AAAI 2025technical

The out-of-distribution (OOD) detection on graph-structured data is crucial for deploying graph neural networks securely in open-world scenarios. However, existing methods have overlooked the prevalent scenario of multi-label classification in real-world applications. In this work, we investigate th…

2025

MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding

AAAI 2025technical

We introduce MM-Mixing, a multi-modal mixing alignment framework for 3D understanding. MM-Mixing applies mixing-based methods to multi-modal data, preserving and optimizing cross-modal connections while enhancing diversity and improving alignment across modalities. Our proposed two-stage training pi…

Cited by 0SourcePDFScholar
2025

Make Unseen Clear: Occluder Removal for Complete 3D Pedestrian Detection

ICASSP 2025accepted

In autonomous driving, the ability to detect pedestrians accurately is crucial for safety. Some detectors, however, often struggle with occlusions, where pedestrians partially hidden behind objects appear incomplete and are harder to be identified accurately. To alleviate this issue, we introduce Cl…

Cited by 0SourceScholar
2025

Make Your Training Flexible: Towards Deployment-Efficient Video Models

ICCV 2025poster

Current video training methods rely on fixed spatiotemporal sampling grids to extract a predetermined number of tokens, limiting adaptability to diverse computational budgets and resulting in suboptimal accuracy-computation trade-offs. This rigidity constrains high-performance models trained in reso…

2025

Multi-Level Speaker Representation for Target Speaker Extraction

ICASSP 2025accepted

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of speakers may suffer from confusion of speaker identity. In th…

Cited by 0SourceScholar
2025

OccProphet: Pushing the Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with an Observer-Forecaster-Refiner Framework

ICLR 2025poster

Predicting variations in complex traffic environments is crucial for the safety of autonomous driving. Recent advancements in occupancy forecasting have enabled forecasting future 3D occupied status in driving environments by observing historical 2D images. However, high computational demands make o…

2025

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

ICLR 2025spotlight

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains th…

2025

PatchST: A Patch Spatial-Temporal Network for Large-Scale Traffic Forecasting

ICASSP 2025accepted

Traffic prediction plays a critical role in mitigating congestion, optimizing traffic flow, and enhancing urban mobility. Accurate predictions contribute directly to improved safety, reduced travel times, and a more efficient transportation system. For example, when a surge in traffic is anticipated…

Cited by 0SourceScholar
2025

Seg-VAR:Image Segmentation with Visual Autoregressive Modeling

NeurIPS 2025poster

While visual autoregressive modeling (VAR) strategies have shed light on image generation with the autoregressive models, their potential for segmentation, a task that requires precise low-level spatial perception, remains unexplored. Inspired by the multi-scale modeling of classic Mask2Former-based…

Cited by 0SourceScholar
2025

Semantic Representation Attack against Aligned Large Language Models

NeurIPS 2025poster

Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that induce LLMs to generate harmful content. Current methods typically target exact affirmative responses, suffering from lim…

Cited by 0SourceScholar
2025

StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models

NeurIPS 2025poster

The advancement of diffusion models has enhanced the realism of AI-generated content but also raised concerns about misuse, necessitating robust copyright protection and tampering localization. Although recent methods have made progress toward unified solutions, their reliance on post hoc processing…

Cited by 0SourceScholar
2025

StreamForest: Efficient Online Video Understanding with Persistent Event Memory

NeurIPS 2025spotlight

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains limited due to storage constraints of historical visual features and insufficient real-time spatiotemporal reasoning. To a…

Cited by 0SourceScholar
2025

TIDE-Net: A Physics-Based Graph Model for Predicting Tropical Cyclone Impacts on Estuarine Systems

ICASSP 2025accepted

Estuarine systems, located at the interface of land, ocean, and atmosphere, are vital to global ecosystems and economies due to their rich exchanges among multiple environments. Tropical cyclones cause significant fluctuations in salinity and other environmental parameters within estuaries, impactin…

Cited by 0SourceScholar
2025

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment

CVPR 2025poster

Current multimodal large language models (MLLMs) struggle with fine-grained or precise understanding of visuals although they give comprehensive perception and reasoning in a spectrum of vision applications. Recent studies either develop tool-using or unify specific visual tasks into the autoregress…

2025

TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging for MLLMs. This paper proposes TimeSuite, a collection of new designs to adapt the existing short-form video MLLMs for lon…

Cited by 10SourcePDFScholar
2025

Towards a Unified Copernicus Foundation Model for Earth Vision

ICCV 2025poster

Advances in Earth observation (EO) foundation models have unlocked the potential of big satellite data to learn generic representations from space, benefiting a wide range of downstream applications crucial to our planet. However, most existing efforts remain limited to fixed spectral sensors, focus…

2025

V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation

NeurIPS 2025poster

Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the scarcity of real data prevent event-based training datasets f…

Cited by 0SourcecodeScholar
2025

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

ICCV 2025poster

We present VRBench, the first long narrative video benchmark crafted for evaluating large models' multi-step reasoning capabilities, addressing limitations in existing evaluations that overlook temporal reasoning and procedural validity. It comprises 960 long videos (with an average duration of 1.6…

Cited by 0SourcePDFScholar
2025

ViLLa: Video Reasoning Segmentation with Large Language Model

ICCV 2025poster

Recent efforts in video reasoning segmentation (VRS) integrate large language models (LLMs) with perception models to localize and track objects via textual instructions, achieving barely satisfactory results in simple scenarios. However, they struggled to discriminate and deduce the objects from us…

2025

VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception

NeurIPS 2025poster

Inducing reasoning in multimodal large language models (MLLMs) is critical for achieving human-level perception and understanding. Existing methods mainly leverage LLM reasoning to analyze parsed visuals, often limited by static perception stages. This paper introduces Visual Test-Time Scaling (VTTS…

Cited by 0SourceScholar
2025

What We Miss Matters: Learning from the Overlooked in Point Cloud Transformers

NeurIPS 2025poster

Point Cloud Transformers have become a cornerstone in 3D representation for their ability to model long-range dependencies via self-attention. However, these models tend to overemphasize salient regions while neglecting other informative regions, which limits feature diversity and compromises robust…

Cited by 0SourceScholar
2025

When Evolution Strategy Meets Language Models Tuning

COLING 2025main

Supervised Fine-tuning has been pivotal in training autoregressive language models, yet it introduces exposure bias. To mitigate this, Post Fine-tuning, including on-policy and off-policy methods, has emerged as a solution to enhance models further. However, each has its limitations regarding perfor…

2025

When Hypergraph Meets Heterophily: New Benchmark Datasets and Baseline

AAAI 2025technical

Hypergraph neural networks (HNNs) have shown promise in handling tasks characterized by high-order correlations, achieving notable success across various applications. However, there has been limited focus on heterophilic hypergraph learning (HHL), in contrast to the increasing attention given to gr…

Cited by 1SourcePDFScholar
2025

XCotton: Advancing AI-Enabled Hardware/Software Integrated System for Foreign Fiber Cleaning

AAAI 2025technical

Cotton is a critical agricultural product and industrial raw material, playing a key role in the national economies and people's living conditions, particularly in developing countries. However, cotton picking and processing often result in the contamination with various foreign fibers, such as hair…

Cited by 0SourcePDFScholar
2024

A Novel Miniature Flexible Instrument With Unfolding and Decoupling Design for Endoscopic Surgery

RA-L 2024

Nowadays, gastrointestinal cancer has widely impacted people's health worldwide due to its high mortality rate. Early treatment of gastrointestinal cancer by endoscopic procedure can greatly increase survival rates of patients. Nevertheless, current flexible endoscopic instruments lack of degree of

Cited by 7SourceScholar
2024

A Novel Robot Platform With Decoupled Stiffness Control for Endoscopic Surgery

RA-L 2024

Endoscopic robot has garnered significant attention for its ability to offer auxiliary traction and precise maneuverability. However, the low stiffness of its insertion tube makes it susceptible to deformation, which poses great challenges for precise control. Existing variable stiffness technologie

Cited by 3SourceScholar
2024

CDA-MBPO: Corrected Data Aggregation for Model-Based Policy Optimization

ICASSP 2024accepted

Model-based reinforcement learning has shown promise in sample efficiency but suffers from errors accumulated during multi-step model sampling. To tackle this issue, we propose corrected data aggregation for model-based policy optimization. This approach involves aligning simulated trajectories with…

Cited by 0SourceScholar
2024

ChatMusician: Understanding and Generating Music Intrinsically with LLM

ACL 2024findings

While LLMs demonstrate impressive capabilities in musical knowledge, we find that music reasoning is still an unsolved task.We introduce ChatMusician, an open-source large language model (LLM) that integrates intrinsic musical abilities. It is based on continual pre-training and finetuning LLaMA2 on…

2024

DCV2I: A Practical Approach for Supporting Geographers’ Visual Interpretation in Dune Segmentation with Deep Vision Models

AAAI 2024technical

Visual interpretation is extremely important in human geography as the primary technique for geographers to use photograph data in identifying, classifying, and quantifying geographic and topological objects or regions. However, it is also time-consuming and requires overwhelming manual effort from…

Cited by 3SourcePDFScholar
2024

Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?

NeurIPS 2024poster

Video understanding relies on accurate action detection for temporal analysis. However, existing mainstream methods have limitations in real-world applications due to their offline and closed-set evaluation approaches, as well as their dependence on manual annotations. To address these challenges an…

2024

Explaining Time Series via Contrastive and Locally Sparse Perturbations

ICLR 2024poster

Explaining multivariate time series is a compound challenge, as it requires identifying important locations in the time series and matching complex temporal patterns. Although previous saliency-based methods addressed the challenges, their perturbation may not alleviate the distribution shift issue,…

2024

F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental Learning

NeurIPS 2024poster

Online Class Incremental Learning (OCIL) aims to train models incrementally, where data arrive in mini-batches, and previous data are not accessible. A major challenge in OCIL is Catastrophic Forgetting, i.e., the loss of previously learned knowledge. Among existing baselines, replay-based methods s…

2024

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

ICLR 2024spotlight

This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. InternVid contains over 7 million videos lasting nearly 760K hours, yielding 234M video clips accomp…

2024

Learning to Maximize Mutual Information for Chain-of-Thought Distillation

ACL 2024findings

Knowledge distillation, the technique of transferring knowledge from large, complex models to smaller ones, marks a pivotal step towards efficient AI deployment. Distilling Step-by-Step (DSS), a novel method utilizing chain-of-thought (CoT) distillation, has demonstrated promise by imbuing smaller m…

2024

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

CVPR 2024highlight

With the rapid development of Multi-modal Large Language Models (MLLMs) a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However most benchmarks predominantly assess spatial understanding in the static image tasks while overlooking t…

2024

OW3Det: Toward Open-World 3D Object Detection for Autonomous Driving

IROS 2024poster

Despite their success in LIDAR object detection, modern detectors are vulnerable to uncommon instances and corner cases (e.g., a runaway tire) since they are closed-set and static. Networks under the closed-set setup only predict labels of seen classes, while static models suffer from catastrophic f…

Cited by 0SourceScholar
2024

PointPatchMix: Point Cloud Mixing with Patch Scoring

AAAI 2024technical

Data augmentation is an effective regularization strategy for mitigating overfitting in deep neural networks, and it plays a crucial role in 3D vision tasks, where the point cloud data is relatively limited. While mixing-based augmentation has shown promise for point clouds, previous methods mix poi…

Cited by 11SourcePDFScholar
2024

Purpose Enhanced Reasoning through Iterative Prompting: Uncover Latent Robustness of ChatGPT on Code Comprehension

IJCAI 2024poster

Code comments are crucial for gaining in-depth insights to facilitate code comprehension. The key to obtaining these insights lies in precisely summarizing the main purpose of the code. Recent approaches on code comment generation lie in prompting large language models (LLMs) such as ChatGPT, instea…

2024

SyncVIS: Synchronized Video Instance Segmentation

NeurIPS 2024poster

Recent DETR-based methods have advanced the development of Video Instance Segmentation (VIS) through transformers' efficiency and capability in modeling spatial and temporal information. Despite harvesting remarkable progress, existing works follow asynchronous designs, which model video sequences v…

2024

Task-oriented Time Series Imputation Evaluation via Generalized Representers

NeurIPS 2024poster

Time series analysis is widely used in many fields such as power energy, economics, and transportation, including different tasks such as forecasting, anomaly detection, classification, etc. Missing values are widely observed in these tasks, and often leading to unpredictable negative effects on exi…

2024

Terrestrial Locomotion of PogoX: From Hardware Design to Energy Shaping and Step-to-step Dynamics Based Control

ICRA 2024poster

We present a novel controller design on a robotic locomotor that combines an aerial vehicle with a spring-loaded leg. The main motivation is to enable the terrestrial locomotion capability on aerial vehicles so that they can carry heavy loads: heavy enough that flying is no longer possible, e.g., wh…

Cited by 9SourceScholar
2024

Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene Completion

NeurIPS 2024poster

Semantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accuratel…

Cited by 1SourcePDFScholar
2024

Weakly-Supervised Crowd Counting with Token Attention and Fusion: A Simple and Effective Baseline

ICASSP 2024accepted

Conventional crowd counting methods exploit a large number of point annotations to train regression-based neural networks for density map estimation. However, laborious point annotations of human heads (strong supervision) are required in training. This paper presents a simple and effective crowd co…

Cited by 0SourceScholar
2023

Bitstream-Corrupted JPEG Images Are Restorable: Two-Stage Compensation and Alignment Framework for Image Restoration

CVPR 2023poster

In this paper, we study a real-world JPEG image restoration problem with bit errors on the encrypted bitstream. The bit errors bring unpredictable color casts and block shifts on decoded image contents, which cannot be trivially resolved by existing image restoration methods mainly relying on pre-de…

2023

Bitstream-Corrupted Video Recovery: A Novel Benchmark Dataset and Method

NeurIPS 2023poster

The past decade has witnessed great strides in video recovery by specialist technologies, like video inpainting, completion, and error concealment. However, they typically simulate the missing content by manual-designed error masks, thus failing to fill in the realistic video loss in video communica…

2023

Boosting Accuracy and Robustness of Student Models via Adaptive Adversarial Distillation

CVPR 2023poster

Distilled student models in teacher-student architectures are widely considered for computational-effective deployment in real-time applications and edge devices. However, there is a higher risk of student models to encounter adversarial attacks at the edge. Popular enhancing schemes such as adversa…

2023

Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition

AAAI 2023technical

The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with different sca…

2023

Enhancing NeRF akin to Enhancing LLMs: Generalizable NeRF Transformer with Mixture-of-View-Experts

ICCV 2023poster

Cross-scene generalizable NeRF models, which can directly synthesize novel views of unseen scenes, have become a new spotlight of the NeRF field. Several existing attempts rely on increasingly end-to-end "neuralized" architectures, i.e., replacing scene representation and/or rendering modules with p…

Cited by 21PDFcodeScholar
2023

Exploiting Prompt Learning with Pre-Trained Language Models for Alzheimer's Disease Detection

ICASSP 2023accepted

Early diagnosis of Alzheimer’s disease (AD) is crucial in facilitating preventive care and to delay further progression. Speech based automatic AD screening systems provide a non-intrusive and more scalable alternative to other clinical screening techniques. Textual embedding features produced by pr…

Cited by 0SourceScholar
2023

Exploring Self-Supervised Pre-Trained ASR Models for Dysarthric and Elderly Speech Recognition

ICASSP 2023accepted

Automatic recognition of disordered and elderly speech remains a highly challenging task to date due to the difficulty in collecting such data in large quantities. This paper explores a series of approaches to integrate domain adapted Self-Supervised Learning (SSL) pre-trained models into TDNN and C…

Cited by 0SourceScholar
2023

Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision

CVPR 2023poster

In this paper, we consider the problem of open-vocabulary semantic segmentation (OVS), which aims to segment objects of arbitrary classes instead of pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor,…

2023

NeRFLix: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-Viewpoint MiXer

CVPR 2023poster

Neural radiance fields(NeRF) show great success in novel-view synthesis. However, in real-world scenes, recovering high-quality details from the source images is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation…

2023

NeuralLift-360: Lifting an In-the-Wild 2D Photo to a 3D Object With 360deg Views

CVPR 2023highlight

Virtual reality and augmented reality (XR) bring increasing demand for 3D content generation. However, creating high-quality 3D content requires tedious work from a human expert. In this work, we study the challenging task of lifting a single image to a 3D object and, for the first time, demonstrate…

2023

OPRADI: Applying Security Game to Fight Drive under the Influence in Real-World

AAAI 2023technical

Driving under the influence (DUI) is one of the main causes of traffic accidents, often leading to severe life and property losses. Setting up sobriety checkpoints on certain roads is the most commonly used practice to identify DUI-drivers in many countries worldwide. However, setting up checkpoints…

Cited by 0SourcePDFScholar
2023

Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection

CVPR 2023poster

Salient object detection (SOD) aims to mimic the human visual system (HVS) and cognition mechanisms to identify and segment salient objects. However, due to the complexity of these mechanisms, current methods are not perfect. Accuracy and robustness need to be further improved, particularly in compl…

2023

SADI: A Self-Adaptive Decomposed Interpretable Framework for Electric Load Forecasting Under Extreme Events

ICASSP 2023accepted

Accurate prediction of electric load is crucial in power grid planning and management. In this paper, we solve the electric load forecasting problem under extreme events such as scorching heats. One challenge for accurate forecasting is the lack of training samples under extreme conditions. Also loa…

Cited by 0SourceScholar
2023

SSL4EO-L: Datasets and Foundation Models for Landsat Imagery

NeurIPS 2023poster

The Landsat program is the longest-running Earth observation program in history, with 50+ years of data acquisition by 8 satellites. The multispectral imagery captured by sensors onboard these satellites is critical for a wide range of scientific fields. Despite the increasing popularity of deep lea…

2023

Scaling Data Generation in Vision-and-Language Navigation

ICCV 2023oral

Recent research in language-guided visual navigation has demonstrated a significant demand for the diversity of traversable environments and the quantity of supervision for training generalizable agents. To tackle the common data scarcity issue in existing vision-and-language navigation datasets, we…

Cited by 80PDFcodeScholar
2023

ScatterFormer: Locally-Invariant Scattering Transformer for Patient-Independent Multispectral Detection of Epileptiform Discharges

AAAI 2023technical

Patient-independent detection of epileptic activities based on visual spectral representation of continuous EEG (cEEG) has been widely used for diagnosing epilepsy. However, precise detection remains a considerable challenge due to subtle variabilities across subjects, channels and time points. Thus…

2023

TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation

NeurIPS 2023poster

Training on large-scale datasets can boost the performance of video instance segmentation while the annotated datasets for VIS are hard to scale up due to the high labor cost. What we possess are numerous isolated filed-specific datasets, thus, it is appealing to jointly train models across the aggr…

2023

Transplayer: Timbre Style Transfer with Flexible Timbre Control

ICASSP 2023accepted

Music timbre style transfer aims at replacing the instrument timbre in a solo recording with another instrument, while preserving the musical content. Existing GAN-based methods can only achieve timbre style transfer between two given timbres. Inspired by the practice in voice conversion, we propose…

Cited by 0SourceScholar
2023

UniFormerV2: Unlocking the Potential of Image ViTs for Video Understanding

ICCV 2023poster

The prolific performances of Vision Transformers (ViTs) in image tasks have prompted research into adapting the image ViTs for video tasks. However, the substantial gap between image and video impedes the spatiotemporal learning of these image-pretrained models. Though video-specialized models like…

Cited by 58PDFcodeScholar
2023

Unmasked Teacher: Towards Training-Efficient Video Foundation Models

ICCV 2023oral

Video Foundation Models (VFMs) have received limited exploration due to high computational costs and data scarcity. Previous VFMs rely on Image Foundation Models (IFMs), which face challenges in transferring to the video domain. Although VideoMAE has trained a robust ViT from limited data, its low-l…

Cited by 189PDFcodeScholar
2023

VideoMAE V2: Scaling Video Masked Autoencoders With Dual Masking

CVPR 2023poster

Scale is the primary factor for building a powerful foundation model that could well generalize to a variety of downstream tasks. However, it is still challenging to train video foundation models with billions of parameters. This paper shows that video masked autoencoder (VideoMAE) is a scalable and…

2022

Audio-Visual Grounding Referring Expression for Robotic Manipulation

ICRA 2022poster

Referring expressions are commonly used when referring to a specific target in people's daily dialogue. In this paper, we develop a novel task of audio-visual grounding referring expression for robotic manipulation. The robot leverages both the audio and visual information to understand the referrin…

Cited by 19SourceScholar
2022

Design and Modeling of a Compound Twisted and Coiled Actuator Based on Spandex Fibers and an SMA Skeleton

RA-L 2022

Twisted and Coiled Actuators (TCAs) are a class of new artificial muscles for flexible actuations. However, conventional TCAs based on nylon fibers commonly require high driving temperature, which limits their applications. Although the TCAs based on spandex fibers can produce high strain under low

Cited by 11SourceScholar
2022

Diversity Features Enhanced Prototypical Network for Few-shot Intent Detection

IJCAI 2022poster

Few-shot Intent Detection (FSID) is a challenging task in dialogue systems due to the scarcity of available annotated utterances. Although existing few-shot learning approaches have made remarkable progress, they fall short in adapting to the Generalized Few-shot Intent Detection (GFSID) task where…

Cited by 11SourcePDFScholar
2022

Learning Friction Model for Magnet-Actuated Tethered Capsule Robot

ICRA 2022poster

The potential diagnostic applications of magnet-actuated capsules have been greatly increased in recent years. For most of these potential applications, accurate position control of the capsule have been highly demanding. However, the friction between the robot and the environment as well as the dra…

Cited by 3SourceScholar
2022

Nonlinear ICA Using Volume-Preserving Transformations

ICLR 2022poster

Nonlinear ICA is a fundamental problem in machine learning, aiming to identify the underlying independent components (sources) from data which is assumed to be a nonlinear function (mixing function) of these sources. Recent works prove that if the sources have some particular structures (e.g. tempor…

Cited by 23SourcePDFScholar
2022

PalGAN: Image Colorization with Palette Generative Adversarial Networks

ECCV 2022poster

"Multimodal ambiguity and color bleeding remain challenging in colorization. To tackle these problems, we propose a new GAN-based colorization approach PalGAN, integrated with palette estimation and chromatic attention. To circumvent the multimodality issue, we present a new colorization formulation…

2022

Point Cloud Domain Adaptation via Masked Local 3D Structure Prediction

ECCV 2022poster

"The superiority of deep learning based point cloud representations relies on large-scale labeled datasets, while the annotation of point clouds is notoriously expensive. One of the most effective solutions is to transfer the knowledge from existing labeled source data to unlabeled target data. Howe…

2022

SDNET: Lightweight Facial Expression Recognition For Sample Disequilibrium

ICASSP 2022accepted

Facial expression recognition (FER) based on the convolutional neural network (CNN) in the wild have numerous challenges. For instance, the complexity of the network model makes FER tasks difficult to deploy on portable devices. Some approaches design lightweight networks to reduce the model size, w…

Cited by 0SourceScholar
2021

Adversarial Defence by Diversified Simultaneous Training of Deep Ensembles

AAAI 2021technical

Learning-based classifiers are susceptible to adversarial examples. Existing defence methods are mostly devised on individual classifiers. Recent studies showed that it is viable to increase adversarial robustness by promoting diversity over an ensemble of models. In this paper, we propose adversari…

2021

Efficient Folded Attention for Medical Image Reconstruction and Segmentation

AAAI 2021technical

Recently, 3D medical image reconstruction (MIR) and segmentation (MIS) based on deep neural networks have been developed with promising results, and attention mechanism has been further designed for performance enhancement. However, the large size of 3D volume images poses a great computational chal…

2021

Multi-Scale Aligned Distillation for Low-Resolution Detection

CVPR 2021poster

In instance-level detection tasks (e.g., object detection), reducing input resolution is an easy option to improve runtime efficiency. However, this option severely hurts the detection performance. This paper focuses on boosting the performance of a low-resolution model, by distilling knowledge from…

Cited by 80PDFcodeScholar
2021

Temporal Rain Decomposition with Spatial Structure Guidance for Video Deraining

ICASSP 2021accepted

Recently, removing rain streaks from videos has drawn wide concerns in vision and multimedia communities. But existing works ignore the depicts of image inherent structure and rain location to cause details loss, and their adopted manners of exploiting temporal information are still insufficient. In…

Cited by 0SourceScholar
2020

Augmentation Data Synthesis Via Gans: Boosting Latent Fingerprint Reconstruction

ICASSP 2020accepted

Latent fingerprint reconstruction is a vital preprocessing step for its identification. This task is very challenging due to not only existing complicated degradation patterns but also its scarcity of paired training data. To address these challenges, we propose a novel generative adversarial networ…

Cited by 0SourceScholar
2020

Intra Frame Rate Control for Versatile Video Coding with Quadratic Rate-Distortion Modelling

ICASSP 2020accepted

With numerous coding tools adopted in the forthcoming Versatile Video Coding (VVC) standard, much less work has been dedicated to study the corresponding Rate-Distortion (R-D) characteristics. This paper proposes a new quadratic R-D model for Versatile Video Coding. In particular, based on the propo…

Cited by 0SourceScholar
2020

On a videoing control system based on object detection and tracking

IROS 2020poster

In this paper, we propose a camera control system towards occasionally videoing preassigned objects. Based on the technique of real-time visual detection and tracking, using the Kalman filter and re-identification (ReID), we propose continuous composition of lens, based on the atomic rules of shots,…

Cited by 1SourceScholar
2020

RANet: Region Attention Network for Semantic Segmentation

NeurIPS 2020poster

Recent semantic segmentation methods model the relationship between pixels to construct the contextual representations. In this paper, we introduce the \emph{Region Attention Network} (RANet), a novel attention network for modeling the relationship between object regions. RANet divides the image int…

2019

Semantic Super-resolution for Extremely Low-resolution Vehicle License Plate

ICASSP 2019accepted

Vehicle license plate (VLP) super-resolution (SR) is of great demand in intelligent traffic systems. Super-Resolution for extremely low-resolution VLP remains challenging and the state-of-the-art SR methods hardly provide satisfying results for low-resolution (LR) VLPs. In this study, from a new per…

Cited by 0SourceScholar
2018

A Reliable Video Storage Architecture in Hybrid SLC/MLC Nand Flash

ICASSP 2018accepted

In this paper, we propose a reliable video storage architecture in hybrid SLC/MLC storage systems. In this architecture, the video stream is reconstructed as the key cluster and the non-key cluster according to the importance of video restoration. The key cluster is stored in SLC blocks to ensure th…

Cited by 0SourceScholar
2018

Generalized Robust Bayesian Committee Machine for Large-scale Gaussian Process Regression

ICML 2018oral

In order to scale standard Gaussian process (GP) regression to large-scale datasets, aggregation models employ factorized training process and then combine predictions from distributed experts. The state-of-the-art aggregation models, however, either provide inconsistent predictions or require time-…

2018

Image Inpainting via Generative Multi-column Convolutional Neural Networks

NeurIPS 2018poster

In this paper, we propose a generative multi-column network for image inpainting. This network synthesizes different image components in a parallel manner within one stage. To better characterize global structures, we design a confidence-driven reconstruction loss while an implicit diversified MRF r…

2018

Inverse Atmoshperic Scattering Modeling with Convolutional Neural Networks for Single Image Dehazing

ICASSP 2018accepted

Single image dehazing is an ill-posed problem. Most existing works use the atmospheric scattering model (ASM) [1] and some natural priors to dehazing. Recently, DehazeNet [2] was developed using deep learning approach achieves the state-of-the-art results on many test hazy images, which motivates us…

Cited by 0SourceScholar
2017

Incremental Kernel Null Space Discriminant Analysis for Novelty Detection

CVPR 2017poster

Novelty detection, which aims to determine whether a given data belongs to any category of training data or not, is considered to be an important and challenging problem in areas of Pattern Recognition, Machine Learning, etc. Recently, kernel null space method (KNDA) was reported to have state-of-th…

Cited by 61PDFScholar
2016

Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin

ICML 2016poster

We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of s…

2015

A parametric modeling approach for wireless capsule endoscopy hazy image restoration

ICASSP 2015accepted

Wireless capsule endoscopy (WCE) is an innovative solution for gastrointestinal disease detection. The image quality of WCE is not satisfactory for medical applications since some of them are dark or hazy. For the purpose of improving WCE image quality, we take a new way to establish a parametric im…

Cited by 0SourceScholar