← Search

Yan Wang

247 accepted papers

2026

ARMOR: Adaptive Curriculum Meta-Learning for Noise-Robust RAG Reasoning

IJCAI 2026

Retrieval-Augmented Generation (RAG) systems have demonstrated remarkable effectiveness in mitigating hallucinations by incorporating external knowledge. However, the retrieval process inevitably introduces noise, posing significant challenges to RAG robustness. Fundamentally, noise robustness is a

Cited by 0Scholar
2026

AdaHC: Accelerating Multi-Token Prediction with Adaptive Head Chunking with Pipeline Parallelism

ICML 2026poster

Multi-token prediction (MTP) architecture is widely adopted in LLMs. MTP blocks can be appended to the tail of model to predict additional tokens. However, when training with pipeline parallel, MTP leads to more pipeline bubbles and deteriorates the pipeline efficiency. Based on in-depth analysis of…

Cited by 0SourceScholar
2026

Benchmarking and Enhancing VLM for Compressed Image Understanding

ICML 2026poster

With the rapid development of Vision-Language Models (VLMs) and the growing demand for their applications, efficient compression of the image inputs has become increasingly important. Existing VLMs predominantly digest and understand high-bitrate compressed images, while their ability to interpret l…

Cited by 0SourceScholar
2026

C-GNN-PRUNE: A Unified Graph-Based Framework for Structure-Aware Pruning of Mixture-of-Experts Models

AAAI 2026technical

The Mixture-of-Experts (MoE) architecture has emerged as a promising paradigm for scaling large language models (LLMs) by activating only a sparse subset of experts per input. However, its massive parameter size remains a major obstacle to efficient deployment. Existing pruning methods often ignore

Cited by 0SourcePDFScholar
2026

Cognition-Inspired Dual-Stream Semantic Enhancement for Vision-Based Dynamic Emotion Modeling

ICRA 2026poster

The human brain constructs emotional percepts not by processing facial expressions in isolation, but through a dynamic, hierarchical integration of sensory input with semantic and contextual knowledge. However, existing vision-based dynamic emotion modeling approaches often neglect emotion perceptio…

2026

Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced Memory

AAAI 2026technical

Few-shot multimodal industrial anomaly detection is a critical yet underexplored task, offering the ability to quickly adapt to complex industrial scenarios. In few-shot settings, insufficient training samples often fail to cover the diverse patterns present in test samples. This challenge can be mi

Cited by 0SourcePDFScholar
2026

Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning

CVPR 2026

Recent reasoning-augmented Vision-Language-Action (VLA) models have improved the interpretability of end-to-end autonomous driving by generating intermediate reasoning traces. Yet these models primarily describe what they perceive and intend to do, rarely questioning whether their planned actions ar

Cited by 0SourceScholar
2026

Diagnostic-Guided Dynamic Profile Optimization for LLM-based User Simulators in Sequential Recommendation

AAAI 2026technical

Recent advances in large language models (LLMs) have enabled realistic user simulators for developing and evaluating recommender systems (RSs). However, existing LLM-based simulators for RSs face two major limitations: (1) static and single-step prompt-based inference that leads to inaccurate and in

Cited by 0SourcePDFScholar
2026

Discontinuous Galerkin Neural Operator for Pathology Defocus Deblurring

ICML 2026poster

Defocus deblurring in pathological microscopy remains challenging due to the spatially varying and locally discontinuous nature of optical blur induced by a position-dependent integral imaging process. Existing deep learning methods, constrained by shift-invariance assumptions and limited interpreta…

Cited by 0SourceScholar
2026

DualFete: Revisiting Teacher-Student Interactions from a Feedback Perspective for Semi-supervised Medical Image Segmentation

AAAI 2026technical

The teacher-student paradigm has emerged as a canonical framework in semi-supervised learning. When applied to medical image segmentation, the paradigm faces challenges due to inherent image ambiguities, making it particularly vulnerable to erroneous supervision. Crucially, the student

Cited by 0SourcePDFScholar
2026

EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human Animation

AAAI 2026technical

Recent work on human animation usually incorporates large-scale video models, thereby achieving more vivid performance. However, the practical use of such methods is hindered by the slow inference speed and high computational demands. Moreover, traditional work typically employs separate models for

Cited by 0SourcePDFScholar
2026

Efficient Multi-Camera Tokenization with Triplanes for End-To-End Driving

ICRA 2026poster

Autoregressive Transformers are increasingly being deployed as end-to-end robot and autonomous vehicle (AV) policy architectures, owing to their scalability and potential to leverage internet-scale pretraining for generalization. Accordingly, tokenizing sensor data efficiently is paramount to ensuri…

2026

Energy Waveify and Redistribution for Test-Time Adaptation: A Control System Perspective

CVPR 2026

This work tackles a key challenge in test-time energy adaptation: prohibitive time overhead arising from recent state-of-the-art test-time adaptation (TTA) methods, which are built on energy models relying on iterative Monte Carlo or Langevin dynamics sampling with multiple stochastic updates per te

Cited by 0SourcecodeScholar
2026

GO-PRE:Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction

ICML 2026poster

Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing methods rely on surrogate signals—such as parameter uncertainty or geometric heuristics—which are often misaligned with the ultimate goal: the fidelity o…

Cited by 0SourceScholar
2026

GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning

CVPR 2026

Zero-shot 3D Anomaly Detection (ZS3DAD) is an emerging task that aims to detect anomalies in a target dataset without any target training data, which is particularly important in scenarios constrained by sample scarcity and data privacy concerns. While current methods adapt CLIP by projecting 3D poi

Cited by 0SourcecodeScholar
2026

GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting

AAAI 2026technical

Implicit neural representations (INRs) have achieved remarkable success in image representation and compression, but they require substantial training time and memory. Meanwhile, recent 2D Gaussian Splatting (GS) methods (\textit{e.g.}, GaussianImage) offer promising alternatives through efficient p

Cited by 0SourcePDFScholar
2026

HySeg: Learning Generative Priors for Structure-Aware Remote Sensing Segmentation

CVPR 2026

High-resolution remote sensing imagery exhibits complex spatial regularities where topology, continuity, and region adjacency govern semantic organization. However, existing remote sensing image semantic segmentation (RSISS) networks, being predominantly discriminative, estimate strong posteriors fr

Cited by 0SourcecodeScholar
2026

Hyperbolic Relational Prompts for Intersectional Fairness in Medical VLMs

CVPR 2026

Ensuring fairness in medical vision-language models (VLMs) is essential for equitable healthcare, yet existing models amplify biases across demographic subgroups such as race and gender. Traditional fairness mitigation approaches relying on broad distribution alignment, fall short in addressing thes

Cited by 0SourceScholar
2026

JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials

ICML 2026poster

Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy MD simulations, and their models exhibit scaling-law trends similar to large language models. However, the lack of scala…

Cited by 0SourceScholar
2026

Latent Chain-of-Thought World Modeling for End-to-End Autonomous Driving

CVPR 2026

Recent Vision-Language-Action (VLA) models for autonomous driving explore inference-time reasoning as a way to improve driving performance and safety in challenging scenarios. Most prior work uses natural language to express chain-of-thought (CoT) reasoning before producing driving actions. However,

Cited by 0SourceScholar
2026

Lifting Unlabeled Internet-level Data for 3D Scene Understanding

CVPR 2026

Annotated 3D scene data is scarce and expensive to acquire, while abundant unlabeled videos are readily available on the internet. In this paper, we demonstrate that carefully designed data engines can leverage web-curated, unlabeled videos to automatically generate training data, to facilitate end-

Cited by 0SourcecodeScholar
2026

Low-Latency Neural LiDAR Compression with 2D Context Models

ICLR 2026poster

Context modeling is fundamental to LiDAR point cloud compression. Existing methods rely on computationally intensive 3D contexts, such as voxel and octree, which struggle to balance the compression efficiency and coding speed. In this work, we propose a neural LiDAR compressor based on 2D context mo…

Cited by 0SourcecodeScholar
2026

MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement Learning

CVPR 2026

Recent advances in multimodal reward modeling have been largely driven by a paradigm shift from discriminative to generative approaches. Building on this progress, recent studies have further employed reinforcement learning with verifiable rewards (RLVR) to enhance multimodal reward models (MRMs). D

Cited by 0SourcecodeScholar
2026

Memento: Toward an All-Day Proactive Assistant for Ultra-Long Streaming Video

ICLR 2026poster

Multimodal large language models have demonstrated impressive capabilities in visual-language understanding, particularly in offline video tasks. More recently, the emergence of online video modeling has introduced early forms of active interaction. However, existing models, typically limited to ten…

Cited by 0SourceScholar
2026

OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text

ICLR 2026poster

Composed video retrieval presents a complex challenge: retrieving a target video based on a source video and a textual modification instruction. This task demands fine-grained reasoning over multimodal transformations. However, existing benchmarks predominantly focus on vision–text alignment, largel…

Cited by 0SourceScholar
2026

Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression

CVPR 2026

Multi-view image compression (MIC) aims to achieve high compression efficiency by exploiting inter-image correlations, playing a crucial role in 3D applications. As a subfield of MIC, distributed multi-view image compression (DMIC) offers performance comparable to MIC while eliminating the need for

Cited by 0SourceScholar
2026

Physics-Aware Accelerated Unrolling Model for Sparse-View CT Reconstruction

AAAI 2026technical

Deep unrolling models (DUMs) have shown great poten-tial in sparse-view CT reconstruction by combining itera-tive optimization and deep learning. However, most DUMsinsufficiently account for physical degradation from sparse-view imaging, leading to slow convergence and persistentartifacts. To addres

Cited by 0SourcePDFScholar
2026

PromptEmo: Learning Emotion with Bilateral Textual Prompts in Multi-Domain Open-set Scenarios

AAAI 2026technical

Facial Expression Recognition (FER) is crucial to human-computer interaction. Existing cross-domain FER (CD-FER) methods mainly focus on single-source closed-set scenarios, transferring knowledge from a single source domain to a target domain with identical class sets. However, CD-FER faces two real

Cited by 0SourcePDFScholar
2026

Reference Recommendation Based Membership Inference Attack Against Hybrid-Based Recommender Systems

AAAI 2026technical

Recommender systems have been widely deployed across various domains such as e-commerce and social media, and intelligently suggest items like products and potential friends to users based on their preferences and interaction history, which are often privacy-sensitive. Recent studies have revealed t

Cited by 0SourcePDFScholar
2026

SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Names

CVPR 2026

Open-vocabulary object detection (OVD) aims to detect objects described by arbitrary text, but most existing methods operate at a coarse category level and struggle with fine-grained, attribute-sensitive queries. We address this from both model and data perspectives. We propose a Semantic-Retrieval-

Cited by 0SourceScholar
2026

Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning

AAAI 2026technical

The growing exploration of Large Language Models (LLM) and Vision-Language Models (VLM) has opened avenues for enhancing the effectiveness of reinforcement learning (RL). However, existing LLM-based RL methods often focus on the guidance of control policy and encounter the challenge of limited repre

Cited by 0SourcePDFScholar
2026

SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation

ICLR 2026poster

The Distribution Matching Distillation (DMD) has been successfully applied to text-to-image diffusion models such as Stable Diffusion (SD) 1.5. However, vanilla DMD suffers from convergence difficulties on large-scale flow-based text-to-image models, such as SD 3.5 and FLUX. In this paper, we first…

Cited by 0SourcecodeScholar
2026

SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries

AAAI 2026technical

Semantic occupancy has emerged as a powerful representation in world models for its ability to capture rich spatial semantics. However, most existing occupancy world models rely on static and fixed embeddings or grids, which inherently limit the flexibility of perception. Moreover, their ``in-place

Cited by 0SourcePDFScholar
2026

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation

CVPR 2026

Diffusion-based video motion customization facilitates the acquisition of human motion representations from a few video samples, while achieving arbitrary subjects transfer through precise textual conditioning. Existing approaches often rely on semantic-level alignment, expecting the model to learn

Cited by 0SourceScholar
2026

THE END OF MANUAL DECODING: TOWARDS TRULY END-TO-END LANGUAGE MODELS

ICLR 2026poster

The "end-to-end" label for LLMs is a misnomer. In practice, they depend on a non-differentiable decoding process that requires laborious, hand-tuning of hyperparameters like temperature and top-p. This paper introduces AutoDeco, a novel architecture that enables truly "end-to-end'' generation by lea…

Cited by 0SourcecodeScholar
2026

TR-DQ: Time-Rotation Diffusion Quantization

AAAI 2026technical

Diffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impa

Cited by 0SourcePDFScholar
2026

The Pensieve Paradigm: Stateful Language Models with Learned Memory Management

ICLR 2026poster

In the world of Harry Potter, when Dumbledore's mind is overburdened, he extracts memories into a Pensieve to be revisited later. In the world of AI, while we possess the Pensieve—mature databases and retrieval systems, our models inexplicably lack the "wand" to operate it. They remain like a Dumble…

Cited by 0SourceScholar
2026

UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement

CVPR 2026

Image quality assessment (IQA) and image restoration are fundamental problems in low-level vision. Although IQA and restoration are closely connected conceptually, most existing work treats them in isolation. Recent advances in unified multimodal understanding-generation models demonstrate promising

Cited by 0SourcecodeScholar
2026

VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation

CVPR 2026

The rapid evolution of generative AI, including such models as Sora, has intensified the threat of video misinformation. A critical challenge in detecting these AI-generated video misinformation lies in a fundamental disconnect between existing datasets and practical deception tactics. Current datas

Cited by 0SourceScholar
2026

Vector Quantization using Gaussian Variational Autoencoder

ICML 2026poster

Vector-quantized variational autoencoders (VQ-VAEs) are discrete autoencoders that compress images into discrete tokens. However, they are difficult to train due to discretization. In this paper, we propose a simple yet effective technique dubbed __Gaussian Quant (GQ)__, which first trains a Gaussia…

Cited by 0SourceScholar
2026

Your One-Stop Solution for AI-Generated Video Detection

CVPR 2026

Recent advances in generative modeling can create remarkably realistic synthetic videos, making it increasingly difficult for humans to distinguish them from real ones and necessitating reliable detection methods. However, two key limitations hinder the development of this field.**From the dataset p

Cited by 0SourcecodeScholar
2025

Addressing Mark Imbalance in Integration-free Marked Temporal Point Processes

NeurIPS 2025poster

Marked Temporal Point Process (MTPP) has been well studied to model the event distribution in marked event streams, which can be used to predict the mark and arrival time of the next event. However, existing studies overlook that the distribution of event marks is highly imbalanced in many real-worl…

Cited by 0SourcecodeScholar
2025

Advancing Stain Transfer for Multi-Biomarkers: A Human Annotation-Free Method Based on Auxiliary Task Supervision

IJCAI 2025

Histopathological examination primarily relies on hematoxylin and eosin (H&E) and immunohistochemical (IHC) staining. Though IHC provides more crucial molecular information for diagnosis, it is more costly than H&E staining. Stain transfer technology seeks to efficiently generate virtual IHC images

2025

AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models

ACL 2025long

Evaluating the alignment capabilities of large Vision-Language Models (VLMs) is essential for determining their effectiveness as helpful assistants. However, existing benchmarks primarily focus on basic abilities using nonverbal methods, such as yes-no and multiple-choice questions. In this paper, w…

2025

Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models

ACL 2025long

Large language models have shown remarkable performance across a wide range of language tasks, owing to their exceptional capabilities in context modeling. The most commonly used method of context modeling is full self-attention, as seen in standard decoder-only Transformers. Although powerful, this…

Cited by 0SourcePDFScholar
2025

CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression

AAAI 2025technical

Existing learning-based stereo image codec adopt sophisticated transformation with simple entropy models derived from single image codecs to encode latent representations. However, those entropy models struggle to effectively capture the spatial-disparity characteristics inherent in stereo images, w…

2025

Component-Aware Unsupervised Logical Anomaly Generation for Industrial Anomaly Detection

ICRA 2025

Anomaly detection is critical in industrial manufacturing for ensuring product quality and improving efficiency in automated processes. The scarcity of anomalous samples limits traditional detection methods, making anomaly generation essential for expanding the data repository. However, recent gener

Cited by 2SourceScholar
2025

CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object Query

ICRA 2025

Cooperative perception enhances the individual perception capabilities of autonomous vehicles (AVs) by providing a comprehensive view of the environment. However, balancing perception performance and transmission costs remains a significant challenge. Current approaches that transmit regionlevel fea

Cited by 9SourceScholar
2025

D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition.

CVPR 2025poster

The current advancements in Dynamic Facial Expression Recognition (DFER) methods mainly focus on better capturing the spatial and temporal features of facial expressions. However, DFER datasets contain a substantial amount of noisy samples, and few have addressed the issue of handling this noise. We…

Cited by 0SourcePDFScholar
2025

DreamDrive: Generative 4D Scene Modeling from Street View Images

ICRA 2025

Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based methods create 3D scenes from driving logs and synthesize geometry-consistent driving videos through neural rendering, b

Cited by 24SourceScholar
2025

Efficient Multi-Camera Tokenization With Triplanes for End-to-End Driving

RA-L 2025

Autoregressive Transformers are increasingly being deployed as end-to-end robot and autonomous vehicle (AV) policy architectures, owing to their scalability and potential to leverage internet-scale pretraining for generalization. Accordingly, tokenizing sensor data <italic xmlns:mml="http://www.w3.o

Cited by 5SourceScholar
2025

Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection

ACL 2025finding

Large Language Models (LLMs) have been shown to exhibit various biases and stereotypes in their generated content. While extensive research has investigated biases in LLMs, prior work has predominantly focused on explicit bias, with minimal attention to implicit bias and the relation between these t…

Cited by 0SourcePDFScholar
2025

Extrapolated Urban View Synthesis Benchmark

ICCV 2025poster

Photorealistic simulators are essential for the training and evaluation of vision-centric autonomous vehicles (AVs). At their core is Novel View Synthesis (NVS), a crucial capability that generates diverse unseen viewpoints to accommodate the broad and continuous pose distribution of AVs. Recent adv…

2025

F-Adapter: Frequency-Adaptive Parameter-Efficient Fine-Tuning in Scientific Machine Learning

NeurIPS 2025poster

Parameter-efficient fine-tuning (PEFT) powerful pre-trained models for complex downstream tasks has proven effective in vision and language processing, yet this paradigm remains unexplored in scientific machine learning, where the objective is to model complex physical systems. We conduct the first…

Cited by 0SourceScholar
2025

GapMatch: Bridging Instance and Model Perturbations for Enhanced Semi-Supervised Medical Image Segmentation

AAAI 2025technical

Medical image segmentation provides detailed understanding and aids in diagnosis, treatment planning, and monitoring of diseases. Due to the high cost of acquiring labeled data in the field of medical image analysis, semi-supervised segmentation methods have garnered increasing attention. Benefiting…

Cited by 0SourcePDFScholar
2025

IROAM: Improving Roadside Monocular 3D Object Detection Learning from Autonomous Vehicle Data Domain

ICRA 2025

In autonomous driving, The perception capabilities of the ego-vehicle can be improved with roadside sensors, which can provide a holistic view of the environment. However, existing monocular detection methods designed for vehicle cameras are not suitable for roadside cameras due to viewpoint domain

Cited by 0SourceScholar
2025

Integrating Task-Specific and Universal Adapters for Pre-Trained Model-based Class-Incremental Learning

ICCV 2025poster

Class-Incremental Learning (CIL) requires a learning system to continually learn new classes without forgetting. Existing pre-trained model-based CIL methods often freeze the pre-trained network and adapt to incremental tasks using additional lightweight modules such as adapters. However, incorrect…

2025

Jack of All Trades, Master of None: PMP-Guided Adaptive Multi-Teacher Distillation with Meta-Learning

ICASSP 2025accepted

To enhance the robustness and accuracy of the small model, existing approaches combine adversarial training with knowledge distillation, introducing a comprehensive single-teacher model to improve the performance of the student model (small model). However, due to the limited knowledge of a teacher…

Cited by 0SourceScholar
2025

KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models

ICASSP 2025accepted

The bottleneck associated with the key-value(KV) cache presents a significant challenge during the inference processes of large language models. While depth pruning accelerates inference, it requires extensive recovery training, which can take up to two weeks. On the other hand, width pruning retain…

Cited by 0SourceScholar
2025

LEANCODE: Understanding Models Better for Code Simplification of Pre-trained Large Language Models

ACL 2025long

Large Language Models for code often entail significant computational complexity, which grows significantly with the length of the input code sequence. We propose LeanCode for code simplification to reduce training and prediction time, leveraging code contexts in utilizing attention scores to repres…

Cited by 0SourcePDFScholar
2025

LLM4RSR: Large Language Models as Data Correctors for Robust Sequential Recommendation

AAAI 2025technical

Sequential Recommenders (SRs) are trained to predict the next item as the target given its preceding items as the input, assuming every input-target pair is matched and is reliable for training. However, users can be induced by external distractions to click on items inconsistent with their true pre…

2025

Language-Image Models with 3D Understanding

ICLR 2025poster

Multi-modal large language models (MLLMs) have shown incredible capabilities in a variety of 2D vision and language tasks. We extend MLLMs’ perceptual capabilities to ground and reason about images in 3-dimensional space. To that end, we first develop a large-scale pretraining dataset for 2D and 3D…

Cited by 15SourcePDFScholar
2025

LoRATEE: A Secure and Efficient Inference Framework for Multi-Tenant LoRA LLMs Based on TEE

ICASSP 2025accepted

Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning approach that adaptes pre-trained Large Language Models (LLMs) to multi-tenant tasks by generating a variety of LoRA adapters. However, this approach faces significant security challenges and is particularly susceptible to malicious ser…

Cited by 0SourceScholar
2025

MDN: Mamba-Driven Dualstream Network For Medical Hyperspectral Image Segmentation

ICASSP 2025accepted

Medical Hyperspectral Imaging (MHSI) offers potential for computational pathology and precision medicine. However, existing CNN and Transformer struggle to balance segmentation accuracy and speed due to high spatial-spectral dimensionality. In this study, we leverage Mamba’s global context modeling…

Cited by 2SourceScholar
2025

MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes

ICCV 2025poster

4D Gaussian Splatting (4DGS) has recently emerged as a promising technique for capturing complex dynamic 3D scenes with high fidelity. It utilizes a 4D Gaussian representation and a GPU-friendly rasterizer, enabling rapid rendering speeds. Despite its advantages, 4DGS faces significant challenges, n…

2025

MamV2XCalib: V2X-based Target-less Infrastructure Camera Calibration with State Space Model

ICCV 2025poster

As cooperative systems that leverage roadside cameras to assist autonomous vehicle perception become increasingly widespread, large-scale precise calibration of infrastructure cameras has become a critical issue. Traditional manual calibration methods are often time-consuming, labor-intensive, and m…

2025

MambaIC: State Space Models for High-Performance Learned Image Compression

CVPR 2025poster

A high-performance image compression algorithm is crucial for real-time information transmission across numerous fields. Despite rapid progress in image compression, computational inefficiency and poor redundancy modeling still pose significant bottlenecks, limiting practical applications. Inspired…

2025

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding

CVPR 2025poster

Open-vocabulary 3D scene understanding is pivotal for enhancing physical intelligence, as it enables embodied agents to interpret and interact dynamically within real-world environments. This paper introduces MPEC, a novel Masked Point-Entity Contrastive learning method for open-vocabulary 3D semant…

Cited by 2SourcePDFScholar
2025

Medusa: A Multi-Scale High-order Contrastive Dual-Diffusion Approach for Multi-View Clustering

CVPR 2025poster

Deep multi-view clustering methods utilize information from multiple views to achieve enhanced clustering results and have gained increasing popularity in recent years. Most existing methods typically focus on either inter-view or intra-view relationships, aiming to align information across views or…

Cited by 0SourcePDFScholar
2025

MonoMixer: Marrying Convolution and Vision Transformer for Efficient Self-Supervised Monocular Depth Estimation

IJCAI 2025

Self-supervised monocular depth estimation that does not require hard-to-source depth labels for training has been widely studied in recent years. Due to its significant and growing needs, many lightweight but effective architectures have been designed for edge devices. Convolutional Neural Networks

Cited by 0SourcePDFScholar
2025

Multi-Layer Knowledge Distillation for Continual Semantic Segmentation

ICASSP 2025accepted

Recently, knowledge distillation has shown promising results for continual semantic segmentation (CSS). Despite their success, several key issues have not been well addressed in existing studies: 1) Difficulty to balance new and old classes, challenged by catastrophic forgetting. 2) Complex informat…

Cited by 0SourceScholar
2025

Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering

AAAI 2025technical

Retrieval-based multi-image question answering (QA) task involves retrieving multiple question-related images and synthesizing these images to generate an answer. Conventional "retrieve-then-answer" pipelines often suffer from cascading errors because the training objective of QA fails to optimize t…

Cited by 0SourcePDFScholar
2025

OUS: Bridging Scene Context and Facial Features to Overcome the Rigid Cognitive Problem

AAAI 2025technical

Dynamic Facial Expression Recognition (DFER) is crucial for affective computing but often overlooks the impact of scene context. We have identified a significant issue in current DFER tasks: human annotators typically integrate emotions from various angles, including environmental cues and body lang…

2025

PICD: Versatile Perceptual Image Compression with Diffusion Rendering

CVPR 2025poster

Recently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perc…

Cited by 0SourcePDFScholar
2025

Physical-aware Neural Radiance Fields for Efficient Exposure Correction

AAAI 2025technical

Neural Radiance Fields (NeRF) has achieved remarkable success in synthesizing impressive novel views. However, existing methods usually fail to handle scenes with adverse lighting conditions caused by external time variations and different camera settings, leading to poor visual quality. To address…

Cited by 0SourcePDFScholar
2025

Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance

EMNLP 2025

Despite Greece’s pivotal role in the global economy, large language models (LLMs) remain underexplored for Greek financial context due to the linguistic complexity of Greek and the scarcity of domain-specific datasets. While multilingual financial NLP has revealed large performance gaps across langu

2025

ProteinBench: A Holistic Evaluation of Protein Foundation Models

ICLR 2025poster

Recent years have witnessed a surge in the development of protein foundation models, significantly improving performance in protein prediction and generative tasks ranging from 3D structure prediction and protein design to conformational dynamics. However, the capabilities and limitations associated…

Cited by 6SourcePDFScholar
2025

Renderworld: World Model with Self-Supervised 3D Label

ICRA 2025

End-to-end autonomous driving with vision-only is not only more cost-effective compared to LiDAR-vision fusion but also more reliable than traditional methods. To achieve a economical and robust purely visual autonomous driving system, we propose RenderWorld, a vision-only end-to-end autonomous driv

Cited by 47SourceScholar
2025

Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior

ICLR 2025poster

Recent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Previous analyses suggest that DPS accomplishes posterior sampling by approximating…

2025

STORM: Spatio-TempOral Reconstruction Model For Large-Scale Outdoor Scenes

ICLR 2025poster

We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often rely on per-scene optimization, dense observations across space and time, and strong motion supervision, resulting in le…

2025

Scene Graph and Dependency Grammar Enhanced Remote Sensing Change Caption Network (SGD-RSCCN)

COLING 2025main

With the continuous advancement of remote sensing technology, it is easier to obtain high-resolution, multi-temporal and multi-spectral images. The images carry rich information of ground objects. However, how to effectively extract useful information from the complex image data and convert it into…

Cited by 0SourcePDFScholar
2025

Self-Prompting Driven SAM2 for 3D Medical Image Segmentation

ICASSP 2025accepted

The latest advancement in large foundational model, SAM2, has demonstrated significant potential in 3D medical image segmentation due to their capability to effectively segment video streams. However, its application in medical image segmentation presents challenges, requiring extensive training on…

Cited by 0SourceScholar
2025

Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving

ICLR 2025poster

Understanding world dynamics is crucial for planning in autonomous driving. Recent methods attempt to achieve this by learning a 3D occupancy world model that forecasts future surrounding scenes based on current observation. However, 3D occupancy labels are still required to produce promising result…

2025

Simultaneous Modeling of Protein Conformation and Dynamics via Autoregression

NeurIPS 2025poster

Understanding protein dynamics is critical for elucidating their biological functions. The increasing availability of molecular dynamics (MD) data enables the training of deep generative models to efficiently explore the conformational space of proteins. However, existing approaches either fail to…

Cited by 0SourceScholar
2025

Towards Comprehensive and Prerequisite-Free Explainer for Graph Neural Networks

IJCAI 2025

To enhance the reliability and credibility of graph neural networks (GNNs) and improve the transparency of their decision logic, a new field of explainability of GNNs (XGNN) has emerged. However, two major limitations severely degrade the performance and hinder the generalizability of existing XGNN

2025

Transformer Based Multi-view Learning for Integrating Static and Dynamic Complementarity of Brain Function

ICASSP 2025accepted

Dynamic temporal information and static connectivity information derived from functional magnetic resonance imaging (fMRI) can assist in the diagnosis of neurological disorders. However, existing disease diagnosis methods primarily rely on information from a single view, neglecting the advantages of…

Cited by 0SourceScholar
2025

Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis

CVPR 2025poster

Existing 3D vision-language (3D-VL) benchmarks fall short in evaluating 3D-VL models, creating a "mist" that obscures rigorous insights into model capabilities and 3D-VL tasks. This mist persists due to three key limitations. First, flawed test data, like ambiguous referential text in the grounding…

2025

VN-GT: Optimizing Virtual Network Deployment via Game Theory

ICASSP 2025accepted

The static and homogeneous nature of traditional networks presents a significant challenge for our defense efforts. These characteristics enable an experienced attacker to quickly determine our network topology and gather detailed information about the internal hosts through systematic scanning tech…

Cited by 0SourceScholar
2024

A Comparative Study of Explicit and Implicit Gender Biases in Large Language Models via Self-evaluation

COLING 2024main

While extensive work has examined the explicit and implicit biases in large language models (LLMs), little research explores the relation between these two types of biases. This paper presents a comparative study of the explicit and implicit biases in LLMs grounded in social psychology. Social psych…

2024

A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis

AAAI 2024technical

Well-designed prompts have demonstrated the potential to guide text-to-image models in generating amazing images. Although existing prompt engineering methods can provide high-level guidance, it is challenging for novice users to achieve the desired results by manually entering prompts due to a disc…

2024

AdaRevD: Adaptive Patch Exiting Reversible Decoder Pushes the Limit of Image Deblurring

CVPR 2024poster

Despite the recent progress in enhancing the efficacy of image deblurring the limited decoding capability constrains the upper limit of State-Of-The-Art (SOTA) methods. This paper proposes a pioneering work Adaptive Patch Exiting Reversible Decoder (AdaRevD) to explore their insufficient decoding ca…

2024

Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution

ECCV 2024poster

"Pre-trained diffusion models utilized for image generation encapsulate a substantial reservoir of a priori knowledge pertaining to intricate textures. Harnessing the potential of leveraging this a priori knowledge in the context of image super-resolution presents a compelling avenue. Nonetheless, p…

Cited by 3SourcePDFScholar
2024

An Embodied Generalist Agent in 3D World

ICML 2024poster

Leveraging massive knowledge from large language models (LLMs), recent machine learning models show notable successes in general-purpose task solving in diverse domains such as computer vision and robotics. However, several significant challenges remain: (i) most of these models rely on 2D images ye…

2024

Augmenting Lane Perception and Topology Understanding with Standard Definition Navigation Maps

ICRA 2024poster

Autonomous driving has traditionally relied heavily on costly and labor-intensive High Definition (HD) maps, hindering scalability. In contrast, Standard Definition (SD) maps are more affordable and have worldwide coverage, offering a scalable alternative. In this work, we systematically explore the…

Cited by 34SourcecodeScholar
2024

Automatic Captioning based on Visible and Infrared Images

ICRA 2024poster

In this paper, we tackle the task of image captioning with the complementarity of visible light images and infrared images. To address this problem, we propose an RGBIR image fusion captioning model, which can take full advantage of visible light images and infrared images under different conditions…

Cited by 1SourceScholar
2024

Bandwidth-Efficient Inference for Nerual Image Compression

ICASSP 2024accepted

With neural networks growing deeper and feature maps growing larger, limited communication bandwidth with external memory (or DRAM) and power constraints become a bottle-neck in implementing network inference on mobile and edge devices. In this paper, we propose an end-to-end differentiable bandwidt…

Cited by 0SourceScholar
2024

Boosting Neural Representations for Videos with a Conditional Decoder

CVPR 2024highlight

Implicit neural representations (INRs) have emerged as a promising approach for video storage and processing showing remarkable versatility across various video tasks. However existing methods often fail to fully leverage their representation capabilities primarily due to inadequate alignment of int…

2024

Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-modal Language Models

CVPR 2024highlight

While Multi-modal Language Models (MLMs) demon strate impressive multimodal ability they still struggle on providing factual and precise responses for tasks like vi sual question answering (VQA). In this paper we address this challenge from the perspective of contextual informa tion. We propose Caus…

Cited by 5SourcePDFScholar
2024

Clean & Compact: Efficient Data-Free Backdoor Defense with Model Compactness

ECCV 2024poster

"Deep neural networks (DNNs) have been widely deployed in real-world, mission-critical applications, necessitating effective approaches to protect deep learning models against malicious attacks. Motivated by the high stealthiness and potential harm of backdoor attacks, a series of backdoor defense m…

Cited by 2SourcePDFScholar
2024

CogAgent: A Visual Language Model for GUI Agents

CVPR 2024highlight

People are spending an enormous amount of time on digital devices through graphical user interfaces (GUIs) e.g. computer or smartphone screens. Large language models (LLMs) such as ChatGPT can assist people in tasks like writing emails but struggle to understand and interact with GUIs thus limiting…

2024

CogVLM: Visual Expert for Pretrained Language Models

NeurIPS 2024poster

We introduce CogVLM, a powerful open-source visual language foundation model. Different from the popular \emph{shallow alignment} method which maps image features into the input space of language model, CogVLM bridges the gap between the frozen pretrained language model and image encoder by a traina…

2024

Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning

AAAI 2024technical

Open-World Few-Shot Learning (OFSL) is a crucial research field dedicated to accurately identifying target samples in scenarios where data is limited and labels are unreliable. This research holds significant practical implications and is highly relevant to real-world applications. Recently, the adv…

2024

Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities

CVPR 2024poster

Multimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However in real-world applications some practical factors cause uncertain modality missingness which drastically degrades the model's…

Cited by 15SourcePDFScholar
2024

DCL-Net: Dual Contrastive Learning Network for Semi-Supervised Multi-Organ Segmentation

ICASSP 2024accepted

Semi-supervised learning (SSL) is a sound measure to relieve the strict demand of abundant annotated datasets, especially for challenging multi-organ segmentation (MoS). However, most existing SSL methods predict pixels in a single image independently, ignoring the relations among images and categor…

Cited by 0SourceScholar
2024

DeIL: Direct-and-Inverse CLIP for Open-World Few-Shot Learning

CVPR 2024poster

Open-World Few-Shot Learning (OFSL) is a critical field of research concentrating on the precise identification of target samples in environments with scarce data and unreliable labels thus possessing substantial practical significance. Recently the evolution of foundation models like CLIP has revea…

2024

DiffuBox: Refining 3D Object Detection with Point Diffusion

NeurIPS 2024poster

Ensuring robust 3D object detection and localization is crucial for many applications in robotics and autonomous driving. Recent models, however, face difficulties in maintaining high performance when applied to domains with differing sensor setups or geographic locations, often resulting in poor lo…

2024

ECM-OPCC: Efficient Context Model for Octree-Based Point Cloud Compression

ICASSP 2024accepted

Recently, deep learning methods have shown promising results in point cloud compression. However, previous octree-based approaches either lack sufficient context or have high decoding complexity (e.g. > 900s). To address this problem, we propose a sufficient yet efficient context model and design an…

Cited by 0SourceScholar
2024

EMIFF: Enhanced Multi-scale Image Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object Detection

ICRA 2024poster

In autonomous driving, cooperative perception makes use of multi-view cameras from both vehicles and infrastructure, providing a global vantage point with rich semantic context of road conditions beyond a single vehicle viewpoint. Currently, two major challenges persist in vehicle-infrastructure coo…

Cited by 7SourcecodeScholar
2024

EasyTPP: Towards Open Benchmarking Temporal Point Processes

ICLR 2024poster

Continuous-time event sequences play a vital role in real-world domains such as healthcare, finance, online shopping, social networks, and so on. To model such data, temporal point processes (TPPs) have emerged as the most natural and competitive models, making a significant impact in both academic…

2024

Emotion Recognition in Conversation via Dynamic Personality

COLING 2024main

Emotion recognition in conversation (ERC) is a field that aims to classify the emotion of each utterance within conversational contexts. This presents significant challenges, particularly in handling emotional ambiguity across various speakers and contextual factors. Existing ERC approaches have pri…

Cited by 4SourcePDFScholar
2024

Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

EMNLP 2024main

Modern large language models (LLMs) like ChatGPT have shown remarkable performance on general language tasks but still struggle on complex reasoning tasks, which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies. Along this direction, one representat…

2024

Enhancing Adaptability: Hierarchical Frontier-Based Path Planning for Navigation in Challenging Environments

RA-L 2024

Current UAV path planning methods exhibit efficient performance in navigating environments with small obstacles, such as indoor areas and outdoor forests. However, they often encounter challenges when dealing with environments characterized by large obstacles, such as expansive walls and towering st

Cited by 8SourceScholar
2024

Enhancing Event Sequence Modeling with Contrastive Relational Inference

ICASSP 2024accepted

Neural temporal point processes(TPPs) have shown promise for modeling continuous-time event sequences. However, capturing the interactions between events is challenging yet critical for performing inference tasks like forecasting on event sequence data. Existing TPP models have focused on parameteri…

Cited by 0SourceScholar
2024

FD-UAD: Unsupervised Anomaly Detection Platform Based on Defect Autonomous Imaging and Enhancement

IJCAI 2024poster

In industrial quality control, detecting defects is essential. However, manual checks and machine vision encounter challenges in complex conditions, as defects vary among products made of different materials and shapes. We create FD-UAD, Unsupervised Anomaly Detection Platform Based on Defect Autono…

Cited by 0SourcePDFScholar
2024

Fine-Tuning Large Language Model Based Explainable Recommendation with Explainable Quality Reward

AAAI 2024technical

Large language model-based explainable recommendation (LLM-based ER) systems can provide remarkable human-like explanations and have widely received attention from researchers. However, the original LLM-based ER systems face three low-quality problems in their generated explanations, i.e., lack of p…

2024

GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting

ECCV 2024poster

"Implicit neural representations (INRs) recently achieved great success in image representation and compression, offering high visual quality and fast rendering speeds with 10-1000 FPS, assuming sufficient GPU resources are available. However, this requirement often hinders their use on low-end devi…

2024

GeneFormer: Learned Gene Compression using Transformer-Based Context Modeling

ICASSP 2024accepted

The development of gene sequencing technology sparks an explosive growth of gene data. Thus, the storage of gene data has become an important issue. Recently, researchers begin to investigate deep learning-based gene data compression, which outperforms general traditional methods. In this paper, we…

Cited by 0SourceScholar
2024

Idempotence and Perceptual Image Compression

ICLR 2024spotlight

Idempotence is the stability of image codec to re-compression. At the first glance, it is unrelated to perceptual image compression. However, we find that theoretically: 1) Conditional generative model-based perceptual codec satisfies idempotence; 2) Unconditional generative model with idempotence c…

2024

Image2Points: A 3D Point-Based Context Clusters GAN for High-Quality Pet Image Reconstruction

ICASSP 2024accepted

To obtain high-quality Positron emission tomography (PET) images while minimizing radiation exposure, numerous methods have been proposed to reconstruct standard-dose PET (SPET) images from the corresponding low-dose PET (LPET) images. However, these methods heavily rely on voxel-based representatio…

Cited by 0SourceScholar
2024

LCGen: Mining in Low-Certainty Generation for View-consistent Text-to-3D

NeurIPS 2024poster

The Janus Problem is a common issue in SDS-based text-to-3D methods. Due to view encoding approach and 2D diffusion prior guidance, the 3D representation model tends to learn content with higher certainty from each perspective, leading to view inconsistency. In this work, we first model and analyze…

Cited by 0SourcePDFScholar
2024

LLMRG: Improving Recommendations through Large Language Model Reasoning Graphs

AAAI 2024technical

Recommendation systems aim to provide users with relevant suggestions, but often lack interpretability and fail to capture higher-level semantic relationships between user behaviors and profiles. In this paper, we propose a novel approach that leverages large language models (LLMs) to construct pers…

Cited by 16SourcePDFScholar
2024

MLPER: Multi-Level Prompts for Adaptively Enhancing Vision-Language Emotion Recognition

IROS 2024poster

In the field of robotics, vision-based Emotion Recognition (ER) has achieved significant progress, but it still faces the challenge of poor generalization ability under unconstrained conditions (e.g., occlusions and pose variations). In this work, we propose MLPER model, which introduces Vision-Lang…

Cited by 1SourceScholar
2024

MR-ULINS: A Tightly-Coupled UWB-LiDAR-Inertial Estimator With Multi-Epoch Outlier Rejection

RA-L 2024

The LiDAR-inertial odometry (LIO) and the ultra-wideband (UWB) have been integrated to achieve driftless positioning in global navigation satellite system (GNSS)-denied environments. However, the UWB may be affected by systematic range errors (such as the clock drift and the antenna phase center off

Cited by 5SourceScholar
2024

PARA-Drive: Parallelized Architecture for Real-time Autonomous Driving

CVPR 2024poster

Recent works have proposed end-to-end autonomous vehicle (AV) architectures comprised of differentiable modules achieving state-of-the-art driving performance. While they provide advantages over the traditional perception-prediction-planning pipeline (e.g. removing information bottlenecks between co…

Cited by 37SourcePDFScholar
2024

PFCF-Net: A Network Based on Progressive Feature Interaction and Cross-Scale Feature Fusion for Remote Sensing Change Detection

ICASSP 2024accepted

There exist some challenges in accurately capturing temporal change information and efficiently aggregating multi-level information in the field of remote sensing change detection. In order to expand the detection’s receptive field and fully fuse complementary information across different hierarchic…

Cited by 0SourceScholar
2024

Partial Label Learning with a Partner

AAAI 2024technical

In partial label learning (PLL), each instance is associated with a set of candidate labels among which only one is ground-truth. The majority of the existing works focuses on constructing robust classifiers to estimate the labeling confidence of candidate labels in order to identify the correct one…

Cited by 5SourcePDFScholar
2024

Pixel-level Semantic Correspondence through Layout-aware Representation Learning and Multi-scale Matching Integration

CVPR 2024poster

Establishing precise semantic correspondence across object instances in different images is a fundamental and challenging task in computer vision. In this task difficulty arises often due to three challenges: confusing regions with similar appearance inconsistent object scale and indistinguishable n…

2024

Probability-Polarized Optimal Transport for Unsupervised Domain Adaptation

AAAI 2024technical

Optimal transport (OT) is an important methodology to measure distribution discrepancy, which has achieved promising performance in artificial intelligence applications, e.g., unsupervised domain adaptation. However, from the view of transportation, there are still limitations: 1) the local discrimi…

Cited by 4SourcePDFScholar
2024

RepAn: Enhanced Annealing through Re-parameterization

CVPR 2024poster

The simulated annealing algorithm aims to improve model convergence through multiple restarts of training. However existing annealing algorithms overlook the correlation between different cycles neglecting the potential for incremental learning. We contend that a fixed network structure prevents the…

2024

SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding

NeurIPS 2024poster

Accurately identifying and organizing textual content is crucial for the automation of document processing in the field of form understanding. Existing datasets, such as FUNSD and XFUND, support entity classification and relationship prediction tasks but are typically limited to local and entity-lev…

2024

SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

ECCV 2024poster

"3D vision-language (3dvl) grounding, which aims to align language with 3D physical environments, stands as a cornerstone in developing embodied agents. In comparison to recent advancements in the 2D domain, grounding language in 3D scenes faces two significant challenges: (i) the scarcity of paired…

Cited by 71SourcePDFScholar
2024

Set Prediction Guided by Semantic Concepts for Diverse Video Captioning

AAAI 2024technical

Diverse video captioning aims to generate a set of sentences to describe the given video in various aspects. Mainstream methods are trained with independent pairs of a video and a caption from its ground-truth set without exploiting the intra-set relationship, resulting in low diversity of generated…

Cited by 3SourcePDFScholar
2024

Shadow-Free Membership Inference Attacks: Recommender Systems Are More Vulnerable Than You Thought

IJCAI 2024poster

Recommender systems have been successfully applied in many applications. Nonetheless, recent studies demonstrate that recommender systems are vulnerable to membership inference attacks (MIAs), leading to the leakage of users’ membership privacy. However, existing MIAs relying on shadow training suff…

2024

Task-Aware Encoder Control for Deep Video Compression

CVPR 2024poster

Prior research on deep video compression (DVC) for machine tasks typically necessitates training a unique codec for each specific task mandating a dedicated decoder per task. In contrast traditional video codecs employ a flexible encoder controller enabling the adaptation of a single codec to differ…

Cited by 7SourcePDFScholar
2024

Training an Open-Vocabulary Monocular 3D Detection Model without 3D Data

NeurIPS 2024poster

Open-vocabulary 3D object detection has recently attracted considerable attention due to its broad applications in autonomous driving and robotics, which aims to effectively recognize novel classes in previously unseen domains. However, existing point cloud-based open-vocabulary 3D detection models…

Cited by 4SourcePDFScholar
2024

Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding

CVPR 2024poster

The Segment Anything Model (SAM) has garnered significant attention for its versatile segmentation abilities and intuitive prompt-based interface. However its application in medical imaging presents challenges requiring either substantial training costs and extensive medical datasets for full model…

2024

Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models

EMNLP 2024finding

Model fusing has always been an important topic, especially in an era where large language models (LLM) and multi-modal language models (MLM) with different architectures, parameter sizes and training pipelines, are being created all the time. In this work, we propose a post-hoc framework, aiming at…

2024

Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot Learning

CVPR 2024poster

Generative Zero-shot learning (ZSL) learns a generator to synthesize visual samples for unseen classes which is an effective way to advance ZSL. However existing generative methods rely on the conditions of Gaussian noise and the predefined semantic prototype which limit the generator only optimized…

Cited by 19SourcePDFScholar
2023

A Shared-Control Dexterous Robotic System for Assisting Transoral Mandibular Fracture Reduction: Development and Cadaver Study

IROS 2023poster

The rigid and straight nature of conventional surgical drills and screwdrivers makes it difficult to access the posterior mandible for fracture reduction without the creation of facial incisions. To assist transoral mandibular fracture reduction in hard-to-reach areas, we propose a shared-control de…

Cited by 0SourceScholar
2023

AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception

ICCV 2023poster

Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIs…

Cited by 54PDFcodeScholar
2023

Automatic Network Pruning via Hilbert-Schmidt Independence Criterion Lasso under Information Bottleneck Principle

ICCV 2023poster

Most existing neural network pruning methods hand-crafted their importance criteria and structures to prune. This constructs heavy and unintended dependencies on heuristics and expert experience for both the objective and the parameters of the pruning approach. In this paper, we try to solve this pr…

Cited by 17PDFcodeScholar
2023

Bidirectional Copy-Paste for Semi-Supervised Medical Image Segmentation

CVPR 2023poster

In semi-supervised medical image segmentation, there exist empirical mismatch problems between labeled and unlabeled data distribution. The knowledge learned from the labeled data may be largely discarded if treating labeled and unlabeled data separately or training labeled and unlabeled data in an…

2023

Bit Allocation using Optimization

ICML 2023poster

In this paper, we consider the problem of bit allocation in Neural Video Compression (NVC). First, we reveal a fundamental relationship between bit allocation in NVC and Semi-Amortized Variational Inference (SAVI). Specifically, we show that SAVI with GoP (Group-of-Picture)-level likelihood is equiv…

2023

Calibration-Free BEV Representation for Infrastructure Perception

IROS 2023poster

Effective BEV object detection on infrastructure can greatly improve traffic scene understanding and vehicle-to-infrastructure (V2I) cooperative perception. However, cameras installed on infrastructure have various postures, and previous BEV detection methods rely on accurate calibration, which is d…

Cited by 22SourceScholar
2023

Efficient Decision-based Black-box Patch Attacks on Video Recognition

ICCV 2023poster

Although Deep Neural Networks (DNNs) have demonstrated excellent performance, they are vulnerable to adversarial patches that introduce perceptible and localized perturbations to the input. Generating adversarial patches on images has received much attention, while adversarial patches on videos have…

Cited by 23PDFScholar
2023

Exploiting Contextual Objects and Relations for 3D Visual Grounding

NeurIPS 2023poster

3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information…

2023

FedVMR: A New Federated Learning Method for Video Moment Retrieval

ICASSP 2023accepted

Despite the great success achieved, existing video moment retrieval (VMR) methods are developed under the assumption that data are centralizedly stored. However, in real-world applications, due to the inherent nature of data generation and privacy concerns, data are often distributed on different si…

Cited by 0SourceScholar
2023

Generalizing Math Word Problem Solvers via Solution Diversification

AAAI 2023technical

Current math word problem (MWP) solvers are usually Seq2Seq models trained by the (one-problem; one-solution) pairs, each of which is made of a problem description and a solution showing reasoning flow to get the correct answer. However, one MWP problem naturally has multiple solution equations. Th…

2023

Idempotent Learned Image Compression with Right-Inverse

NeurIPS 2023poster

We consider the problem of idempotent learned image compression (LIC). The idempotence of codec refers to the stability of codec to re-compression. To achieve idempotence, previous codecs adopt invertible transforms such as DCT and normalizing flow. In this paper, we first identify that invertibilit…

Cited by 4SourcePDFScholar
2023

Intriguing Findings of Frequency Selection for Image Deblurring

AAAI 2023technical

Blur was naturally analyzed in the frequency domain, by estimating the latent sharp image and the blur kernel given a blurry image. Recent progress on image deblurring always designs end-to-end architectures and aims at learning the difference between blurry and sharp image pairs from pixel-level, w…

2023

Jump Over Block (JOB): An Efficient Line-of-Sight Checker for Grid/Voxel Maps With Sparse Obstacles

RA-L 2023

Line-Of-Sight (LOS) check plays a crucial role in collision avoidance and time comsuming, particularly in scenarios involving large-scale maps with sparse obstacles, as it necessitates a grid-by-grid state check. Specifically, LOS check consumes more than half of the computational time in any-angle

Cited by 1SourceScholar
2023

LION: Label Disambiguation for Semi-supervised Facial Expression Recognition with Progressive Negative Learning

IJCAI 2023poster

Semi-supervised deep facial expression recognition (SS-DFER) has recently attracted rising research interest due to its more practical setting of abundant unlabeled data. However, there are two main problems unconsidered in current SS-DFER methods: 1) label ambiguity, i.e., given labels mismatch wit…

2023

Large Language Models Meet Harry Potter: A Dataset for Aligning Dialogue Agents with Characters

EMNLP 2023long findings

In recent years, Dialogue-style Large Language Models (LLMs) such as ChatGPT and GPT4 have demonstrated immense potential in constructing open-domain dialogue agents. However, aligning these agents with specific characters or individuals remains a considerable challenge due to the complexities of ch…

Cited by 0SourceScholar
2023

MagicNet: Semi-Supervised Multi-Organ Segmentation via Magic-Cube Partition and Recovery

CVPR 2023poster

We propose a novel teacher-student model for semi-supervised multi-organ segmentation. In the teacher-student model, data augmentation is usually adopted on unlabeled data to regularize the consistent training between teacher and student. We start from a key perspective that fixed relative locations…

2023

MedNgage: A Dataset for Understanding Engagement in Patient-Nurse Conversations

ACL 2023findings

Patients who effectively manage their symptoms often demonstrate higher levels of engagement in conversations and interventions with healthcare practitioners. This engagement is multifaceted, encompassing cognitive and social dimensions. Consequently, it is crucial for AI systems to understand the e…

2023

Meta Architecture for Point Cloud Analysis

CVPR 2023poster

Recent advances in 3D point cloud analysis bring a diverse set of network architectures to the field. However, the lack of a unified framework to interpret those networks makes any systematic comparison, contrast, or analysis challenging, and practically limits healthy development of the field. In t…

2023

OMPQ: Orthogonal Mixed Precision Quantization

AAAI 2023technical

To bridge the ever-increasing gap between deep neural networks' complexity and hardware capability, network quantization has attracted more and more research attention. The latest trend of mixed precision quantization takes advantage of hardware's multiple bit-width arithmetic operations to unleash…

2023

Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified Perspective

AAAI 2023technical

Positive Unlabeled (PU) learning, which has a wide range of applications, is becoming increasingly prevalent. However, it suffers from problems such as data imbalance, selection bias, and prior agnostic in real scenarios. Existing studies focus on addressing part of these problems, which fail to pro…

Cited by 4SourcePDFScholar
2023

Privacy-Preserving Adversarial Facial Features

CVPR 2023poster

Face recognition service providers protect face privacy by extracting compact and discriminative facial features (representations) from images, and storing the facial features for real-time recognition. However, such features can still be exploited to recover the appearance of the original face by b…

Cited by 22SourcePDFScholar
2023

Prompt Makes mask Language Models Better Adversarial Attackers

ICASSP 2023accepted

Generating high-quality synonymous perturbations is a core challenge for textual adversarial tasks. However, candidates generated from the masked language model often contain many words that are antonyms or irrelevant to the original words, which limit the perturbation space and affect the attack’s…

Cited by 0SourceScholar
2023

Prompt-augmented Temporal Point Process for Streaming Event Sequence

NeurIPS 2023poster

Neural Temporal Point Processes (TPPs) are the prevalent paradigm for modeling continuous-time event sequences, such as user activities on the web and financial transactions. In real world applications, the event data typically comes in a streaming manner, where the distribution of the patterns may…

Cited by 26SourcePDFScholar
2023

Real-Time Reinforcement Learning for Vision-Based Robotics Utilizing Local and Remote Computers

ICRA 2023poster

Real-time learning is crucial for robotic agents adapting to ever-changing, non-stationary environments. A common setup for a robotic agent is to have two different computers simultaneously: a resource-limited local computer tethered to the robot and a powerful remote computer connected wirelessly.…

Cited by 14SourcecodeScholar
2023

Rethinking Safe Semi-supervised Learning: Transferring the Open-set Problem to A Close-set One

ICCV 2023poster

Conventional semi-supervised learning (SSL) lies in the close-set assumption that the labeled and unlabeled sets contain data with the same seen classes, called in-distribution (ID) data. In contrast, safe SSL investigates a more challenging open-set problem where unlabeled set may involve some out-…

Cited by 12PDFScholar
2023

SSDA3D: Semi-supervised Domain Adaptation for 3D Object Detection from Point Cloud

AAAI 2023technical

LiDAR-based 3D object detection is an indispensable task in advanced autonomous driving systems. Though impressive detection results have been achieved by superior 3D detectors, they suffer from significant performance degeneration when facing unseen domains, such as different LiDAR configurations,…

2023

Stability of Random Forests and Coverage of Random-Forest Prediction Intervals

NeurIPS 2023poster

We establish stability of random forests under the mild condition that the squared response ($Y^2$) does not have a heavy tail. In particular, our analysis holds for the practical version of random forests that is implemented in popular packages like \texttt{randomForest} in \texttt{R}. Empirical re…

Cited by 13SourcePDFScholar
2023

Theoretically Guaranteed Bidirectional Data Rectification for Robust Sequential Recommendation

NeurIPS 2023poster

Sequential recommender systems (SRSs) are typically trained to predict the next item as the target given its preceding (and succeeding) items as the input. Such a paradigm assumes that every input-target pair is reliable for training. However, users can be induced to click on items that are inconsis…

Cited by 4SourcePDFScholar
2023

Waymax: An Accelerated, Data-Driven Simulator for Large-Scale Autonomous Driving Research

NeurIPS 2023poster

Simulation is an essential tool to develop and benchmark autonomous vehicle planning software in a safe and cost-effective manner. However, realistic simulation requires accurate modeling of multi-agent interactive behaviors to be trustworthy, behaviors which can be highly nuanced and complex. To ad…

Cited by 116SourcePDFScholar
2023

mVIL-Fusion: Monocular Visual-Inertial-LiDAR Simultaneous Localization and Mapping in Challenging Environments

RA-L 2023

We propose mVIL-Fusion, a three-level multisensor fusion system that is able to achieve robust state estimation and globally consistent mapping in perceptually degraded environments. First, LiDAR depth-assisted visual-inertial odometry (VIO) with LiDAR odometry (LO) synchronous prediction and distor

Cited by 20SourcecodeScholar
2022

A Contrastive Framework for Neural Text Generation

NeurIPS 2022accept

Text generation is of great importance to many natural language processing applications. However, maximization-based decoding methods (e.g., beam search) of neural language models often lead to degenerate solutions---the generated text is unnatural and contains undesirable repetitions. Existing appr…

2022

A Novel Single-Arm Stapling Robot for Oral and Maxillofacial Surgery - Design and Verification

RA-L 2022

Because of the limited space, the suture of the oral and maxillofacial surgery is a challenging task, which requires oral surgeons to master excellent techniques. This letter presents a single-arm stapling robot for oral and maxillofacial surgery using magnesium alloy staples, as well the stapling s

Cited by 7SourceScholar
2022

AttExplainer: Explain Transformer via Attention by Reinforcement Learning

IJCAI 2022poster

Transformer and its variants, built based on attention mechanisms, have recently achieved remarkable performance in many NLP tasks. Most existing works on Transformer explanation tend to reveal and utilize the attention matrix with human subjective intuitions in a qualitative manner. However, the hu…

2022

Black-Box Dissector: Towards Erasing-Based Hard-Label Model Stealing Attack

ECCV 2022poster

"Previous studies have verified that the functionality of black-box models can be stolen with full probability outputs. However, under the more practical hard-label setting, we observe that existing methods suffer from catastrophic performance degradation. We argue this is due to the lack of rich in…

2022

ContrastMask: Contrastive Learning To Segment Every Thing

CVPR 2022poster

Partially-supervised instance segmentation is a task which requests segmenting objects from novel categories via learning on limited base categories with annotated masks thus eliminating demands of heavy annotation burden. The key to addressing this task is to build an effective class-agnostic mask…

Cited by 52PDFcodeScholar
2022

ELIC: Efficient Learned Image Compression With Unevenly Grouped Space-Channel Contextual Adaptive Coding

CVPR 2022oral

Recently, learned image compression techniques have achieved remarkable performance, even surpassing the best manually designed lossy image coders. They are promising to be large-scale adopted. For the sake of practicality, a thorough investigation of the architecture design of learned image compres…

Cited by 356PDFcodeScholar
2022

Exploiting Playbacks in Unsupervised Domain Adaptation for 3D Object Detection in Self-Driving Cars

ICRA 2022poster

Self-driving cars must detect other traffic participants like vehicles and pedestrians in 3D in order to plan safe routes and avoid collisions. State-of-the-art 3D object detectors, based on deep learning, have shown promising accuracy but are prone to over-fit domain idiosyncrasies, making them fai…

Cited by 25SourceScholar
2022

FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos

CVPR 2022poster

Current benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the "Ha…

Cited by 110PDFcodeScholar
2022

Fixed Neural Network Steganography: Train the images, not the network

ICLR 2022poster

Recent attempts at image steganography make use of advances in deep learning to train an encoder-decoder network pair to hide and retrieve secret messages in images. These methods are able to hide large amounts of data, but they also incur high decoding error rates (around 20%). In this paper, we pr…

2022

Ithaca365: Dataset and Driving Perception Under Repeated and Challenging Weather Conditions

CVPR 2022poster

Advances in perception for self-driving cars have accelerated in recent years due to the availability of large-scale datasets, typically collected at specific locations and under nice weather conditions. Yet, to achieve the high safety requirement, these perceptual systems must operate robustly unde…

Cited by 53PDFScholar
2022

LightPose: A Lightweight and Efficient Model with Transformer for Human Pose Estimation

ICASSP 2022accepted

The prediction of keypoints by generating high-resolution heatmaps has become a popular solution in human pose estimation. While this kind of method requires up-sampling or deconvolution operations, which would bring a great challenge to the acceleration of model inference. If performing keypoint pr…

Cited by 0SourceScholar
2022

Local-Global Feature Aggregation for Light Field Image Super-Resolution

ICASSP 2022accepted

Deep convolutional neural networks (CNNs) have been widely explored in light field (LF) image super-resolution (SR) to achieve remarkable progress. However, most of the existing CNNs-based methods ignore the similarity of local neighbor views in the 4D LF data. Besides, due to the limitations of CNN…

Cited by 0SourceScholar
2022

Multi-Sample Training for Neural Image Compression

NeurIPS 2022accept

This paper considers the problem of lossy neural image compression (NIC). Current state-of-the-art (SOTA) methods adopt uniform posterior to approximate quantization noise, and single-sample pathwise estimator to approximate the gradient of evidence lower bound (ELBO). In this paper, we propose to t…

Cited by 5SourcePDFScholar
2022

Neural Surface Reconstruction of Dynamic Scenes with Monocular RGB-D Camera

NeurIPS 2022accept

We propose Neural-DynamicReconstruction (NDR), a template-free method to recover high-fidelity geometry and motions of a dynamic scene from a monocular RGB-D camera. In NDR, we adopt the neural implicit function for surface representation and rendering such that the captured color and depth can be f…