← Search

hao tang

153 accepted papers

2026

Cross Domain Test Time Scaling: Scale Knowledge and Reasoning on Cross Domains

IJCAI 2026

Test-time scaling (TTS) has demonstrated remarkable potential in enhancing the reasoning capabilities of Large Language Models (LLMs) and Large Vision-Language Models (LVLMs). However, its application has primarily been limited to domains such as mathematics and programming, owing to their reasoning

Cited by 0Scholar
2026

Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models

AAAI 2026technical

Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-lab

Cited by 0SourcePDFScholar
2026

DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledge

CVPR 2026

Open-set fine-grained retrieval (OSFR) is a challenging task where models must generalize to unseen subcategories. Existing methods often fail this, as they embed category-specific semantics from closed-set training labels. Recently, diffusion transformers (DiT) have shown promise by encoding attrib

Cited by 0SourceScholar
2026

FlowPET: Physics-Informed Symplectic Flow Matching for Low-Count PET Reconstruction

ICML 2026poster

Low-count Positron Emission Tomography (PET) reconstruction is severely hindered by the dissipative nature of prevailing generative models, where the inherent phase-space contraction leads to the numerical extinction (``wash-out'') of weak but diagnostically critical lesion signals. To overcome this…

Cited by 0SourceScholar
2026

FourierPET: Deep Fourier-based Unrolled Network for Low-count PET Reconstruction

AAAI 2026technical

Low-count positron emission tomography (PET) reconstruction is a challenging inverse problem due to severe degradations arising from Poisson noise, photon scarcity, and attenuation correction errors. Existing deep learning methods typically address these in the spatial domain with an undifferentiate

Cited by 0SourcePDFScholar
2026

Gram2Token: Enabling Run-time GPU-Native Grammar-Constrained Decoding for LLMs

ICML 2026poster

Grammar-constrained decoding is essential for enabling large language models (LLMs) to efficiently generate structured outputs in applications, such as JSON objects for parameter passing. Existing approaches typically execute grammar constraint masking on the CPU, while LLM inference is performed on…

Cited by 0SourceScholar
2026

Hallucination Begins Where Saliency Drops

ICLR 2026oral

Recent studies have investigated attention dynamics in large vision language models (LVLMs), yet existing methods remain limited in reliably distinguishing hallucinated from correct outputs — primarily because they rely solely on forward-pass attention, ignoring gradient-based signals that reveal ho…

Cited by 0SourcecodeScholar
2026

ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation

AAAI 2026technical

Enabling multi-task adaptation in pre-trained Low-Rank Adaptation (LoRA) models is crucial for enhancing their generalization capabilities. Most existing pre-trained LoRA fusion methods decompose weight matrices, sharing similar parameters, while fusion divergent ones. However, this paradigm inevit

Cited by 0SourcePDFScholar
2026

IMAGGarment+: Efficient Attribute-Wise Diffusion for Garment Generation

AAAI 2026technical

Diffusion models have advanced fine-grained garment generation, yet balancing controllability, efficiency, and texture fidelity remains challenging. Adapter-based methods often yield incoherent details, while full fine-tuning is computationally expensive and prone to overwriting pretrained priors. T

Cited by 0SourcePDFScholar
2026

MIRNet: Integrating Constrained Graph-Based Reasoning with Pre-training for Diagnostic Medical Imaging

AAAI 2026technical

Automated interpretation of medical images demands robust modeling of complex visual-semantic relationships while addressing annotation scarcity, label imbalance, and clinical plausibility constraints. We introduce MIRNet (Medical Image Reasoner Network), a novel framework that integrates self-super

Cited by 0SourcePDFScholar
2026

MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling

ICLR 2026poster

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human–environment interactions. This often leads to unrealistic or physically implaus…

Cited by 0SourceScholar
2026

MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing

CVPR 2026

3D morphing remains challenging due to the difficulty of generating semantically consistent and temporally smooth deformations, especially across categories. We present MorphAny3D, a training-free framework that leverages Structured Latent (SLAT) representations for high-quality 3D morphing. Our key

Cited by 0SourcecodeScholar
2026

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models

CVPR 2026

Multi-modal large language models (MLLMs) have rapidly advanced in visual tasks, yet their spatial understanding remains limited to single images, leaving them ill-suited for physical-world applications that require multi-frame reasoning. In this paper, we propose a framework to equip MLLMs with mul

Cited by 0SourcecodeScholar
2026

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

CVPR 2026

Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties, yet dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient modality-specific modeling, and challenges in reliability. I

Cited by 0SourceScholar
2026

OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis

CVPR 2026

Panoramic dental radiographs require fine-grained spatial reasoning, bilateral symmetry understanding, and multi-step diagnostic verification, yet existing vision-language models operate under a static single-pass paradigm that limits their clinical reliability. In this paper, we introduce OralGPT-P

Cited by 0SourcecodeScholar
2026

PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation

CVPR 2026

Hand-object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains fragmented across three disjoint tracks: (1) pose-only synthesis that predicts MANO trajectories without producing pixels; (

Cited by 0SourcecodeScholar
2026

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

CVPR 2026

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak at spatial grounding and understanding, and VLA systems bui

Cited by 0SourceScholar
2026

Precision-Induced Miscalibration: Understanding and Correcting Confidence Distortion in Quantized Neural Networks

ICML 2026poster

Low-precision arithmetic is pervasive in neural network training and deployment, yet its effect on prediction \textit{confidence}, not just accuracy, remains unexamined. We show that the softmax function amplifies logit-space quantization errors in an input-dependent manner: confidence distortion sc…

Cited by 0SourceScholar
2026

SpikeStereoNet: A Brain-Inspired Framework for Stereo Depth Estimation from Spike Streams

ICLR 2026poster

Conventional frame-based cameras often struggle with stereo depth estimation in rapidly changing scenes. In contrast, bio-inspired spike cameras emit asynchronous events at microsecond-level resolution, providing an alternative sensing modality. However, existing methods lack specialized stereo algo…

Cited by 0SourcecodeScholar
2026

StereoAdapter: Adapting Stereo Depth Estimation to Underwater Scenes

ICRA 2026poster

Underwater stereo depth estimation provides accurate 3D geometry for robotics tasks such as navigation, inspection, and mapping, offering metric depth from low-cost passive cameras while avoiding the scale ambiguity of monocular methods. However, existing approaches face two critical challenges: (i)…

2026

TR-DQ: Time-Rotation Diffusion Quantization

AAAI 2026technical

Diffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impa

Cited by 0SourcePDFScholar
2026

UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing

ICLR 2026poster

In this paper, we propose UniLIP, a unified framework that adapts CLIP for multimodal understanding, generation and editing. Although CLIP excels at understanding, it lacks reconstruction abilities required to be a unified visual encoder. However, previous CLIP-based unified methods fail to balance…

Cited by 0SourcecodeScholar
2026

VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery

ICLR 2026poster

Vision-Language Models (VLMs) have achieved significant progress in multimodal understanding tasks, demonstrating strong capabilities particularly in general tasks such as image captioning and visual reasoning. However, when dealing with specialized cultural heritage domains like 3D vase artifacts,…

Cited by 0SourcecodeScholar
2025

3DS-VLA: A 3D Spatial-Aware Vision Language Action Model for Robust Multi-Task Manipulation

CoRL 2025poster

Recently, 2D vision-language-action (VLA) models have made significant strides in multi-task manipulation. However, these models struggle to reason about 3D spatial relationships from 2D image inputs. Although an increasing number of 3D approaches explicitly integrate 3D information, they encounter…

Cited by 0SourceScholar
2025

A Training-free Synthetic Data Selection Method for Semantic Segmentation

AAAI 2025technical

Training semantic segmenter with synthetic data has been attracting great attention due to its easy accessibility and huge quantities. Most previous methods focused on producing large-scale synthetic image-annotation samples and then training the segmenter with all of them. However, such a solution…

2025

ARNet: Self-Supervised FG-SBIR with Unified Sample Feature Alignment and Multi-Scale Token Recycling

AAAI 2025technical

Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) aims to minimize the distance between sketches and corresponding images in the embedding space. However, scalability is hindered by the growing complexity of solutions, mainly due to the abstract nature of fine-grained sketches. In this paper, we p…

2025

Boosting Adversarial Transferability with Spatial Adversarial Alignment

NeurIPS 2025poster

Deep neural networks are vulnerable to adversarial examples that exhibit transferability across various models. Numerous approaches are proposed to enhance the transferability of adversarial examples, including advanced optimization, data augmentation, and model modifications. However, these methods…

Cited by 0SourceScholar
2025

CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting

IROS 2025

Vehicle-to-everything (V2X) communication plays a crucial role in autonomous driving, enabling cooperation between vehicles and infrastructure. While simulation has significantly contributed to various autonomous driving tasks, its potential for data generation and augmentation in V2X scenarios rema

Cited by 4SourcecodeScholar
2025

Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention

ICML 2025poster

In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera control into the generation process, but their results are oft…

Cited by 8SourcePDFScholar
2025

Combining Induction and Transduction for Abstract Reasoning

ICLR 2025poster

When learning an input-output mapping from very few examples, is it better to first infer a latent function that explains the examples, or is it better to directly predict new test outputs, e.g. using a neural network? We study this question on ARC by training neural models for \emph{induction} (inf…

2025

Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning

IJCAI 2025

Few-shot learning (FSL) addresses the challenge of classifying novel classes with limited training samples. While some methods leverage semantic knowledge from smaller-scale models to mitigate data scarcity, these approaches often introduce noise and bias due to the data’s inherent simplicity. In th

Cited by 0SourcePDFScholar
2025

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

ICCV 2025poster

In recent years, the introduction of Multi-modal Large Language Models (MLLMs) into video understanding tasks has become increasingly prevalent. However, how to effectively integrate temporal information remains a critical research focus. Traditional approaches treat spatial and temporal information…

Cited by 0SourcePDFScholar
2025

Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment

NeurIPS 2025poster

Preference alignment in diffusion models has primarily focused on benign human preferences (e.g., aesthetic). In this paper, we propose a novel perspective: framing unrestricted adversarial example generation as a problem of aligning with adversary preferences. Unlike benign alignment, adversarial a…

Cited by 0SourceScholar
2025

FairSMOE: Mitigating Multi-Attribute Fairness Problem with Sparse Mixture-of-Experts

IJCAI 2025

Real‐world datasets usually contain multiple attributes, making it essential to ensure fairness across all of them simultaneously. However, different attributes may vary in difficulty, and no existing approaches have effectively addressed this issue. Consequently, an attribute‐adaptive strategy is n

Cited by 0SourcePDFScholar
2025

Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass

CVPR 2025poster

Multi-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives. Current leading methods such as DUSt3R employ a fundamentally pairwise approach, processing images in pairs and necessit…

2025

From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks

NAACL 2025long

Large Vision Language Models (LVLMs) achieve great performance on visual-language reasoning tasks, however, the black-box nature of LVLMs hinders in-depth research on the reasoning mechanism. As all images need to be converted into image tokens to fit the input format of large language models (LLMs)…

2025

HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Models

CVPR 2025poster

We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (e.g. text, objects, partia…

Cited by 1SourcePDFScholar
2025

LLM-Guided Probabilistic Program Induction for POMDP Model Estimation

CoRL 2025poster

Partially Observable Markov Decision Processes (POMDPs) model decision making under uncertainty. While there are many approaches to approximately solving POMDPs, we aim to address the problem of learning such models. In particular, we are interested in a subclass of POMDPs wherein the components of…

Cited by 0SourceScholar
2025

MambaIC: State Space Models for High-Performance Learned Image Compression

CVPR 2025poster

A high-performance image compression algorithm is crucial for real-time information transmission across numerous fields. Despite rapid progress in image compression, computational inefficiency and poor redundancy modeling still pose significant bottlenecks, limiting practical applications. Inspired…

2025

MaskSAM: Auto-prompt SAM with Mask Classification for Volumetric Medical Image Segmentation

ICCV 2025poster

The Segment Anything Model (SAM), a prompt-driven foundation model for natural image segmentation, has demonstrated impressive zero-shot performance. However, SAM is not directly applicable to medical image segmentation due to its inability to predict semantic labels, reliance on additional prompts,…

2025

MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task Automation

NAACL 2025system demonstrations

Existing Multimodal Large Language Model (MLLM)-based agents face significant challenges in handling complex GUI (Graphical User Interface) interactions on devices. These challenges arise from the dynamic and structured nature of GUI environments, which integrate text, images, and spatial relationsh…

2025

Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition

AAAI 2025technical

Given the critical role of birds in ecosystems, Fine-Grained Bird Recognition (FGBR) has gained increasing attention, particularly in distinguishing birds within similar subcategories. Although Vision Transformer (ViT)-based methods often outperform Convolutional Neural Network (CNN)-based methods i…

Cited by 0SourcePDFScholar
2025

OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection

IJCAI 2025

Out-of-distribution (OOD) detection is crucial for ensuring the reliability and safety of machine learning models in real-world applications. While zero-shot OOD detection, which requires no training on in-distribution (ID) data, has become feasible with the emergence of vision-language models like

Cited by 0SourcePDFScholar
2025

OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB

IROS 2025

To address the challenge of short-term object pose tracking in dynamic environments with monocular RGB input, we introduce a large-scale synthetic dataset Omni-Pose6D, crafted to mirror the diversity of real-world conditions. We additionally present a benchmarking framework for a comprehensive compa

Cited by 1SourceScholar
2025

PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model

CVPR 2025poster

As interest grows in world models that predict future states from current observations and actions, accurately modeling part-level dynamics has become increasingly relevant for various applications. Existing approaches, such as Puppet-Master, rely on fine-tuning large-scale pre-trained video diffusi…

Cited by 0SourcePDFScholar
2025

PoE-World: Compositional World Modeling with Products of Programmatic Experts

NeurIPS 2025spotlight

Learning how the world works is central to building AI agents that can adapt to complex environments. Traditional world models based on deep-learning demand vast amounts of training data, and do not flexibly update their knowledge from sparse observations. Recent advances in program synthesis usin…

Cited by 0SourcecodeScholar
2025

RoRA: Efficient Fine-Tuning of LLM with Reliability Optimization for Rank Adaptation

ICASSP 2025accepted

Fine-tuning helps large language models (LLM) recover degraded information and enhance task performance. Although Low-Rank Adaptation (LoRA) is widely used and effective for fine-tuning, we have observed that its scaling factor can limit or even reduce performance as the rank size increases. To addr…

Cited by 0SourceScholar
2025

RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness

NeurIPS 2025spotlight

Fine-tuning pre-trained models with custom data leads to numerous expert models on specific tasks. Merging models into one universal model to empower multi-task ability refraining from data leakage has gained popularity. With the expansion in data and model size, parameter-efficient tuning becomes t…

Cited by 0SourceScholar
2025

Similarity Memory Prior is All You Need for Medical Image Segmentation

ICCV 2025poster

In recent years, it has been found that "grandmother cells" in the primary visual cortex (V1) of macaques can directly recognize visual input with complex shapes. This inspires us to examine the value of these cells in promoting the research of medical image segmentation. In this paper, we design a…

2025

Stable-Hair: Real-World Hair Transfer via Diffusion Model

AAAI 2025technical

Current hair transfer methods struggle to handle diverse and intricate hairstyles, limiting their applicability in real-world scenarios. In this paper, we propose a novel diffusion-based hair transfer framework, named Stable-Hair, which robustly transfers a wide range of real-world hairstyles to use…

2025

Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment

AAAI 2025technical

Structured pruning for large language models (LLMs) has garnered significant academic interest due to its ability to efficiently compress and accelerate LLMs by eliminating redundant weight groups at a coarse-grained granularity. Current structured pruning methods for LLMs typically depend on a sing…

2025

Toward Zero-Shot Learning for Visual Dehazing of Urological Surgical Robots

ICRA 2025

Robot-assisted surgery has profoundly influenced current forms of minimally invasive surgery. However, in transurethral urological surgical robots, they need to work in a liquid environment. This causes vaporization of the liquid when shearing and heating is performed, resulting in bubble atomizatio

Cited by 1SourcecodeScholar
2025

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis

NeurIPS 2025poster

Recent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. In particular, panoramic X-rays, a widely used imaging modality in oral radiolog…

Cited by 0SourceScholar
2025

UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface

NeurIPS 2025spotlight

Generalist models have achieved remarkable success in both language and vision-language tasks, showcasing the potential of unified modeling. However, effectively integrating fine-grained perception tasks like detection and segmentation into these models remains a significant challenge. This is prima…

Cited by 0SourcecodeScholar
2025

VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning

ICLR 2025spotlight

Broadly intelligent agents should form task-specific abstractions that selectively expose the essential elements of a task, while abstracting away the complexity of the raw sensorimotor space. In this work, we present Neuro-Symbolic Predicates, a first-order abstraction language that combines the st…

Cited by 3SourcePDFScholar
2024

3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance

ECCV 2024poster

"In this paper, we propose 3DSS-VLG, a weakly supervised approach for 3D Semantic Segmentation with 2D Vision-Language Guidance, an alternative approach that a 3D model predicts dense-embedding for each point which is co-embedded with both the aligned image and text spaces from the 2D vision-languag…

2024

3x2: 3D Object Part Segmentation by 2D Semantic Correspondences

ECCV 2024poster

"3D object part segmentation is essential in computer vision applications. While substantial progress has been made in 2D object part segmentation, the 3D counterpart has received less attention, in part due to the scarcity of annotated 3D datasets, which are expensive to collect. In this work, we p…

2024

ADen: Adaptive Density Representations for Sparse-view Camera Pose Estimation

ECCV 2024oral

"Recovering camera poses from a set of images is a foundational task in 3D computer vision, which powers key applications such as 3D scene/object reconstructions. Classic methods often depend on feature correspondence, such as keypoints, which require the input images to have large overlap and small…

Cited by 1SourcePDFScholar
2024

Code Repair with LLMs gives an Exploration-Exploitation Tradeoff

NeurIPS 2024poster

Iteratively improving and repairing source code with large language models (LLMs), known as refinement, has emerged as a popular way of generating programs that would be too complex to construct in one shot. Given a bank of test cases, together with a candidate program, an LLM can improve that progr…

Cited by 6SourcePDFScholar
2024

Delving into Multimodal Prompting for Fine-Grained Visual Classification

AAAI 2024technical

Fine-grained visual classification (FGVC) involves categorizing fine subdivisions within a broader category, which poses challenges due to subtle inter-class discrepancies and large intra-class variations. However, prevailing approaches primarily focus on uni-modal visual concepts. Recent advancemen…

Cited by 28SourcePDFScholar
2024

Efficient Joint Rectification of Photometric and Geometric Distortions in Document Images

ICASSP 2024accepted

Document images captured with cameras often exhibit photometric and geometric distortions. Here, we propose a novel learning-based approach for efficient joint rectification of document images. Inspired by the strong correlation between visual shadows and physical deformations, we design a shared en…

Cited by 0SourceScholar
2024

Efficient-3Dim: Learning a Generalizable Single-image Novel-view Synthesizer in One Day

ICLR 2024poster

The task of novel view synthesis aims to generate unseen perspectives of an object or scene from a limited set of input images. Nevertheless, synthesizing novel views from a single image remains a significant challenge. Previous approaches tackle this problem by adopting mesh prediction, multi-plane…

Cited by 0SourcePDFScholar
2024

G2P-DDM: Generating Sign Pose Sequence from Gloss Sequence with Discrete Diffusion Model

AAAI 2024technical

The Sign Language Production (SLP) project aims to automatically translate spoken languages into sign sequences. Our approach focuses on the transformation of sign gloss sequences into their corresponding sign pose sequences (G2P). In this paper, we present a novel solution for this task by converti…

2024

GiT: Towards Generalist Vision Transformer through Universal Language Interface

ECCV 2024oral

"This paper proposes a simple, yet effective framework, called , simultaneously applicable for various vision tasks only with a vanilla ViT. Motivated by the universality of the Multi-layer Transformer architecture (e.g., GPT) widely used in large language models (LLMs), we seek to broaden its scope…

2024

HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point Cloud

CVPR 2024highlight

Extracting keypoint locations from input hand frames known as 3D hand pose estimation is a critical task in various human-computer interaction applications. Essentially the 3D hand pose estimation can be regarded as a 3D point subset generative problem conditioned on input frames. Thanks to the rece…

2024

ICON: Incremental CONfidence for Joint Pose and Radiance Field Optimization

CVPR 2024poster

Neural Radiance Fields (NeRF) exhibit remarkable performance for Novel View Synthesis (NVS) given a set of 2D images. However NeRF training requires accurate camera pose for each input view typically obtained by Structure-from-Motion (SfM) pipelines. Recent works have attempted to relax this constra…

Cited by 2SourcePDFScholar
2024

Learning with Unreliability: Fast Few-shot Voxel Radiance Fields with Relative Geometric Consistency

CVPR 2024poster

We propose a voxel-based optimization framework ReVoRF for few-shot radiance fields that strategically addresses the unreliability in pseudo novel view synthesis. Our method pivots on the insight that relative depth relationships within neighboring regions are more reliable than the absolute color v…

2024

Revisiting Adversarial Patches for Designing Camera-Agnostic Attacks against Person Detection

NeurIPS 2024poster

Physical adversarial attacks can deceive deep neural networks (DNNs), leading to erroneous predictions in real-world scenarios. To uncover potential security risks, attacking the safety-critical task of person detection has garnered significant attention. However, we observe that existing attack met…

Cited by 1SourcePDFScholar
2024

SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis

ECCV 2024poster

"Semantic image synthesis (SIS) shows good promises for sensor simulation. However, current best practices in this field, based on GANs, have not yet reached the desired level of quality. As latent diffusion models make significant strides in image generation, we are prompted to evaluate ControlNet,…

Cited by 3SourcePDFScholar
2024

SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation

CVPR 2024poster

Recent advancements in subject-driven image generation have led to zero-shot generation yet precise selection and focus on crucial subject representations remain challenging. Addressing this we introduce the SSR-Encoder a novel architecture designed for selectively capturing any subject from single…

2024

StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion

ECCV 2024poster

"Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trained text-to-image model into an auto-regressive manner. Although these models have shown notable progress, there are st…

2024

Token Transformation Matters: Towards Faithful Post-hoc Explanation for Vision Transformer

CVPR 2024poster

While Transformers have rapidly gained popularity in various computer vision applications post-hoc explanations of their internal mechanisms remain largely unexplored. Vision Transformers extract visual information by representing image regions as transformed tokens and integrating them via attentio…

Cited by 9SourcePDFScholar
2024

Versatile Navigation Under Partial Observability via Value-guided Diffusion Policy

CVPR 2024poster

Route planning for navigation under partial observability plays a crucial role in modern robotics and autonomous driving. Existing route planning approaches can be categorized into two main classes: traditional autoregressive and diffusion-based methods. The former often fails due to its myopic natu…

Cited by 2SourcePDFScholar
2024

WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment

NeurIPS 2024poster

We give a model-based agent that builds a Python program representing its knowledge of the world based on its interactions with the environment. The world model tries to explain its interactions, while also being optimistic about what reward it can achieve. We define this optimism as a logical const…

Cited by 36SourcePDFScholar
2023

Analyzing Acoustic Word Embeddings from Pre-Trained Self-Supervised Speech Models

ICASSP 2023accepted

Given the strong results of self-supervised models on various tasks, there have been surprisingly few studies exploring self-supervised representations for acoustic word embeddings (AWE), fixed-dimensional vectors representing variable-length spoken word segments. In this work, we study several pre-…

Cited by 0SourceScholar
2023

Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution

ICASSP 2023accepted

Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved by conditioning the network of noise predictor with low-resol…

Cited by 0SourceScholar
2023

DE-net: Dynamic Text-Guided Image Editing Adversarial Networks

AAAI 2023technical

Text-guided image editing models have shown remarkable results. However, there remain two problems. First, they employ fixed manipulation modules for various editing requirements (e.g., color changing, texture changing, content adding and removing), which results in over-editing or insufficient edit…

2023

Data Level Lottery Ticket Hypothesis for Vision Transformers

IJCAI 2023poster

The conventional lottery ticket hypothesis (LTH) claims that there exists a sparse subnetwork within a dense neural network and a proper random initialization method, called the winning ticket, such that it can be trained from scratch to almost as good as the dense counterpart. Meanwhile, the resear…

2023

DeepMAD: Mathematical Architecture Design for Deep Convolutional Neural Network

CVPR 2023poster

The rapid advances in Vision Transformer (ViT) refresh the state-of-the-art performances in various vision tasks, overshadowing the conventional CNN-based models. This ignites a few recent striking-back research in the CNN world showing that pure CNN models can achieve as good performance as ViT mod…

2023

Does Graph Distillation See Like Vision Dataset Counterpart?

NeurIPS 2023poster

Training on large-scale graphs has achieved remarkable results in graph representation learning, but its cost and storage have attracted increasing concerns. Existing graph condensation methods primarily focus on optimizing the feature matrices of condensed graphs while overlooking the impact of the…

Cited by 42SourcePDFScholar
2023

Edge Guided GANs with Contrastive Learning for Semantic Image Synthesis

ICLR 2023poster

We propose a novel \underline{e}dge guided \underline{g}enerative \underline{a}dversarial \underline{n}etwork with \underline{c}ontrastive learning (ECGAN) for the challenging semantic image synthesis task. Although considerable improvement has been achieved, the quality of synthesized images is far…

2023

EgoTracks: A Long-term Egocentric Visual Object Tracking Dataset

NeurIPS 2023poster

Visual object tracking is a key component to many egocentric vision problems. However, the full spectrum of challenges of egocentric tracking faced by an embodied AI is underrepresented in many existing datasets; these tend to focus on relatively short, third-person videos. Egocentric video has seve…

2023

GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis

CVPR 2023poster

Synthesizing high-fidelity complex images from text is challenging. Based on large pretraining, the autoregressive and diffusion models can synthesize photo-realistic images. Although these large models have shown notable progress, there remain three flaws. 1) These models require tremendous trainin…

2023

Graph Transformer GANs for Graph-Constrained House Generation

CVPR 2023poster

We present a novel graph Transformer generative adversarial network (GTGAN) to learn effective graph node relations in an end-to-end fashion for the challenging graph-constrained house generation task. The proposed graph-Transformer-based generator includes a novel graph Transformer encoder that com…

Cited by 31SourcePDFScholar
2023

HOTCOLD Block: Fooling Thermal Infrared Detectors with a Novel Wearable Design

AAAI 2023technical

Adversarial attacks on thermal infrared imaging expose the risk of related applications. Estimating the security of these systems is essential for safely deploying them in the real world. In many cases, realizing the attacks in the physical space requires elaborate special perturbations. These solut…

2023

HotBEV: Hardware-oriented Transformer-based Multi-View 3D Detector for BEV Perception

NeurIPS 2023poster

The bird's-eye-view (BEV) perception plays a critical role in autonomous driving systems, involving the accurate and efficient detection and tracking of objects from a top-down perspective. To achieve real-time decision-making in self-driving scenarios, low-latency computation is essential. While re…

Cited by 5SourcePDFScholar
2023

LART: Neural Correspondence Learning with Latent Regularization Transformer for 3D Motion Transfer

NeurIPS 2023poster

3D motion transfer aims at transferring the motion from a dynamic input sequence to a static 3D object and outputs an identical motion of the target with high-fidelity and realistic visual effects. In this work, we propose a novel 3D Transformer framework called LART for 3D motion transfer. With car…

2023

Learning Concordant Attention via Target-aware Alignment for Visible-Infrared Person Re-identification

ICCV 2023poster

Owing to the large distribution gap between the heterogeneous data in Visible-Infrared Person Re-identification (VI Re-ID), we point out that existing paradigms often suffer from the inter-modal semantic misalignment issue and thus fail to align and compare local details properly. In this paper, we…

Cited by 37PDFScholar
2023

Learning Zero-Shot Cooperation with Humans, Assuming Humans Are Biased

ICLR 2023poster

There is a recent trend of applying multi-agent reinforcement learning (MARL) to train an agent that can cooperate with humans in a zero-shot fashion without using any human data. The typical workflow is to first repeatedly run self-play (SP) to build a policy pool and then train the final adaptive…

2023

Master: Meta Style Transformer for Controllable Zero-Shot and Few-Shot Artistic Style Transfer

CVPR 2023poster

Transformer-based models achieve favorable performance in artistic style transfer recently thanks to its global receptive field and powerful multi-head/layer attention operations. Nevertheless, the over-paramerized multi-layer structure increases parameters significantly and thus presents a heavy bu…

Cited by 20SourcePDFScholar
2023

Object Reprojection Error (ORE): Camera pose benchmarks from lightweight tracking annotations

NeurIPS 2023poster

3D spatial understanding is highly valuable in the context of semantic modeling of environments, agents, and their relationships. Semantic modeling approaches employed on monocular video often ingest outputs from off-the-shelf SLAM/SfM pipelines, which are anecdotally observed to perform poorly or…

Cited by 0SourcePDFScholar
2023

PI-Trans: Parallel-Convmlp and Implicit-Transformation Based Gan for Cross-View Image Translation

ICASSP 2023accepted

For semantic-guided cross-view image translation, it is crucial to learn where to sample pixels from the source view image and where to reallocate them guided by the target view semantic map, especially when there is little overlap or drastic view difference between the source and target images. Hen…

Cited by 0SourceScholar
2023

PackQViT: Faster Sub-8-bit Vision Transformers via Full and Packed Quantization on the Mobile

NeurIPS 2023poster

While Vision Transformers (ViTs) have undoubtedly made impressive strides in computer vision (CV), their intricate network structures necessitate substantial computation and memory resources. A decision-making process for CV tasks typically entails performing computations with low latency, which is…

Cited by 21SourcePDFScholar
2023

Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training

AAAI 2023technical

Vision transformers (ViTs) have recently obtained success in many applications, but their intensive computation and heavy memory usage at both training and inference time limit their generalization. Previous compression algorithms usually start from the pre-trained dense models and only focus on eff…

2023

Pruning Parameterization With Bi-Level Optimization for Efficient Semantic Segmentation on the Edge

CVPR 2023poster

With the ever-increasing popularity of edge devices, it is necessary to implement real-time segmentation on the edge for autonomous driving and many other applications. Vision Transformers (ViTs) have shown considerably stronger results for many vision tasks. However, ViTs with the full-attention me…

Cited by 28SourcePDFScholar
2023

RZCR: Zero-shot Character Recognition via Radical-based Reasoning

IJCAI 2023poster

The long-tail effect is a common issue that limits the performance of deep learning models on real-world datasets. Character image datasets are also affected by such unbalanced data distribution due to differences in character usage frequency. Thus, current character recognition methods are limited…

Cited by 14SourcePDFScholar
2023

SMAE: Few-Shot Learning for HDR Deghosting With Saturation-Aware Masked Autoencoders

CVPR 2023poster

Generating a high-quality High Dynamic Range (HDR) image from dynamic scenes has recently been extensively studied by exploiting Deep Neural Networks (DNNs). Most DNNs-based methods require a large amount of training data with ground truth, requiring tedious and time-consuming work. Few-shot HDR ima…

Cited by 19SourcePDFScholar
2023

SpeedDETR: Speed-aware Transformers for End-to-end Object Detection

ICML 2023poster

Vision Transformers (ViTs) have continuously achieved new milestones in object detection. However, the considerable computation and memory burden compromise their efficiency and generalization of deployment on resource-constraint devices. Besides, efficient transformer-based detectors designed by ex…

Cited by 3SourcePDFScholar
2023

TINYCOD: Tiny and Effective Model for Camouflaged Object Detection

ICASSP 2023accepted

This paper introduces an effective and tiny model for real-time Camouflaged Object Detection (COD) named Tiny-COD. It achieves high performance with very low costs (Parameters < 5M, FLOPs < 1.5G), which can be applied on mobile devices. Specifically, we introduce a simple but effective Adjacent Scal…

Cited by 0SourceScholar
2023

Towards Real-Time Segmentation on the Edge

AAAI 2023technical

The research in real-time segmentation mainly focuses on desktop GPUs. However, autonomous driving and many other applications rely on real-time segmentation on the edge, and current arts are far from the goal. In addition, recent advances in vision transformers also inspire us to re-design the ne…

Cited by 14SourcePDFScholar
2023

UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View Representation

ICCV 2023poster

Jointly processing information from multiple sensors is crucial to achieving accurate and robust perception for reliable autonomous driving systems. However, current 3D perception research follows a modality-specific paradigm, leading to additional computation overheads and inefficient collaboration…

Cited by 78PDFcodeScholar
2023

Unsupervised Deep Probabilistic Approach for Partial Point Cloud Registration

CVPR 2023poster

Deep point cloud registration methods face challenges to partial overlaps and rely on labeled data. To address these issues, we propose UDPReg, an unsupervised deep probabilistic registration framework for point clouds with partial overlaps. Specifically, we first adopt a network to learn posterior…

2022

3D-Aware Semantic-Guided Generative Model for Human Synthesis

ECCV 2022poster

"Generative Neural Radiance Field (GNeRF) models, which extract implicit 3D representations from 2D images, have recently been shown to produce realistic images representing rigid/semi-rigid objects, such as human faces or cars. However, they usually struggle to generate high-quality images represen…

2022

Accurate Inference of Unseen Combinations of Multiple Rootcauses with Classifier Ensemble

ICASSP 2022accepted

Root cause analysis (RCA) of network faults is crucial to wireless network operation and management. It, however, is challenging, due to diverse feature types, diverse lengths of time slices, simultaneous occurrences of multiple root causes, and lack of training samples. In this paper, we present ou…

Cited by 0SourceScholar
2022

Compiler-Aware Neural Architecture Search for On-Mobile Real-Time Super-Resolution

ECCV 2022poster

"Deep learning-based super-resolution (SR) has gained tremendous popularity in recent years because of its high image quality performance and wide application scenarios. However, prior methods typically suffer from large amounts of computations and huge power consumption, causing difficulties for re…

2022

DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis

CVPR 2022oral

Synthesizing high-quality realistic images from text descriptions is a challenging task. Existing text-to-image Generative Adversarial Networks generally employ a stacked architecture as the backbone yet still remain three flaws. First, the stacked architecture introduces the entanglements between g…

Cited by 349PDFcodeScholar
2022

Geometry-Contrastive Transformer for Generalized 3D Pose Transfer

AAAI 2022technical

We present a customized 3D mesh Transformer model for the pose transfer task. As the 3D pose transfer essentially is a deformation procedure dependent on the given meshes, the intuition of this work is to perceive the geometric inconsistency between the given meshes with the powerful self-attention…

2022

Learning To Restore 3D Face From In-the-Wild Degraded Images

CVPR 2022poster

In-the-wild 3D face modelling is a challenging problem as the predicted facial geometry and texture suffer from a lack of reliable clues or priors, when the input images are degraded. To address such a problem, in this paper we propose a novel Learning to Restore (L2R) 3D face framework for unsuperv…

Cited by 3PDFScholar
2022

MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation

CVPR 2022poster

Estimating 3D human poses from monocular videos is a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by exploiting spatial and temporal relationships. However, those works ignore the fact that it is an inverse problem where multiple feasib…

Cited by 415PDFcodeScholar
2022

Mining Relations among Cross-Frame Affinities for Video Semantic Segmentation

ECCV 2022poster

"The essence of video semantic segmentation (VSS) is how to leverage temporal information for prediction. Previous efforts are mainly devoted to developing new techniques to calculate the cross-frame affinities such as optical flow and attention. Instead, this paper contributes from a different angl…

2022

Multi-Modal Perception Attention Network with Self-Supervised Learning for Audio-Visual Speaker Tracking

AAAI 2022technical

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of multi-modal signals remains a challenging issue. In this paper,…

2022

PPT: Token-Pruned Pose Transformer for Monocular and Multi-View Human Pose Estimation

ECCV 2022poster

"Recently, the vision transformer and its variants have played an increasingly important role in both monocular and multi-view human pose estimation. Considering image patches as tokens, transformers can model the global dependencies within the entire image or across images from other views. However…

2022

Physically-Guided Disentangled Implicit Rendering for 3D Face Modeling

CVPR 2022poster

This paper presents a novel Physically-guided Disentangled Implicit Rendering (PhyDIR) framework for high-fidelity 3D face modeling. The motivation comes from two observations: widely-used graphics renderers yield excessive approximations against photo-realistic imaging, while neural rendering metho…

Cited by 8PDFScholar
2022

Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model

CVPR 2022poster

To achieve disentangled image manipulation, previous works depend heavily on manual annotation. Meanwhile, the available manipulations are limited to a pre-defined set the models were trained for. We propose a novel framework, i.e., Predict, Prevent, and Evaluate (PPE), for disentangled text-driven…

Cited by 49PDFcodeScholar
2022

SPViT: Enabling Faster Vision Transformers via Latency-Aware Soft Token Pruning

ECCV 2022poster

"Recently, Vision Transformer (ViT) has continuously established new milestones in the computer vision field, while the high computation and memory cost makes its propagation in industrial production difficult. Considering the computation complexity, the internal data pattern of ViTs, and the edge d…

2022

Topology-Preserving Shape Reconstruction and Registration via Neural Diffeomorphic Flow

CVPR 2022poster

Deep Implicit Functions (DIFs) represent 3D geometry with continuous signed distance functions learned through deep neural nets. Recently DIFs-based methods have been proposed to handle shape reconstruction and dense point correspondences simultaneously, capturing semantic relationships across shape…

Cited by 45PDFcodeScholar
2022

Towards Interpretable Video Super-Resolution via Alternating Optimization

ECCV 2022poster

"In this paper, we study a practical space-time video super-resolution (STVSR) problem which aims at generating a high-framerate high-resolution sharp video from a low-framerate low-resolution blurry video. Such problem often occurs when recording a fast dynamic event with a low-framerate and low-re…

2021

Intrinsic-Extrinsic Preserved GANs for Unsupervised 3D Pose Transfer

ICCV 2021poster

With the strength of deep generative models, 3D pose transfer regains intensive research interests in recent years. Existing methods mainly rely on a variety of constraints to achieve the pose transfer over 3D meshes, e.g., the need for manually encoding for shape and pose disentanglement. In this p…

Cited by 33PDFcodeScholar
2021

Recurrent Mask Refinement for Few-Shot Medical Image Segmentation

ICCV 2021poster

Although having achieved great success in medical image segmentation, deep convolutional neural networks usually require a large dataset with manual annotations for training and are difficult to generalize to unseen classes. Few-shot learning has the potential to address these challenges by learning…

Cited by 151PDFcodeScholar
2021

Robust and Accurate RGB-D Reconstruction With Line Feature Constraints

RA-L 2021

Scene reconstruction with consumer-level RGB-D cameras has developed considerable momentum in both robotics and vision communities. In the literature of robotics, high-quality camera tracking, the key to accurate reconstruction, is challenging in geometric featureless scenes or under large lighting

Cited by 5SourceScholar
2021

Transformer-Based Attention Networks for Continuous Pixel-Wise Prediction

ICCV 2021poster

While convolutional neural networks have shown a tremendous impact on various computer vision tasks, they generally demonstrate limitations in explicitly modeling long-range dependencies due to the intrinsic locality of the convolution operation. Initially designed for natural language processing ta…

Cited by 239PDFcodeScholar
2020

Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHAT

ICASSP 2020accepted

This paper proposes a straightforward 2-D method to spatially calibrate the visual field of a camera with the auditory field of an array microphone by generating and overlaying an acoustic image over an optical image. Using a low-cost microphone array and an off-the-shelf camera, we show that polyno…

Cited by 0SourceScholar
2020

Belief Propagation Neural Networks

NeurIPS 2020poster

Learned neural solvers have successfully been used to solve combinatorial optimization and decision problems. More general counting variants of these problems, however, are still largely solved with hand-crafted solvers. To bridge this gap, we introduce belief propagation neural networks (BPNNs), a…

2020

Exocentric to Egocentric Image Generation Via Parallel Generative Adversarial Network

ICASSP 2020accepted

Cross-view image generation has been recently proposed to generate images of one view from another dramatically different view. In this paper, we investigate exocentric (third-person) view to egocentric (first-person) view image generation. This is a challenging task since egocentric view sometimes…

Cited by 0SourceScholar
2020

Local Class-Specific and Global Image-Level Generative Adversarial Networks for Semantic-Guided Scene Generation

CVPR 2020poster

In this paper, we address the task of semantic-guided scene generation. One open challenge widely observed in global image-level generation methods is the difficulty of generating small objects and detailed local texture. To tackle this issue, in this work we consider learning the scene generation i…

Cited by 192PDFcodeScholar
2020

Refactoring Policy for Compositional Generalizability using Self-Supervised Object Proposals

NeurIPS 2020poster

We study how to learn a policy with compositional generalizability. We propose a two-stage framework, which refactorizes a high-reward teacher policy into a generalizable student policy with strong inductive bias. Particularly, we implement an object-centric GNN-based student policy, whose input obj…

2020

Towards Scale-Invariant Graph-related Problem Solving by Iterative Homogeneous GNNs

NeurIPS 2020poster

Current graph neural networks (GNNs) lack generalizability with respect to scales (graph sizes, graph diameters, edge weights, etc..) when solving many graph analysis problems. Taking the perspective of synthesizing graph theory programs, we propose several extensions to address the issue. First, in…

Cited by 64SourcePDFScholar
2019

Multi-Channel Attention Selection GAN With Cascaded Semantic Guidance for Cross-View Image Translation

CVPR 2019oral

Cross-view image translation is challenging because it involves images with drastically different views and severe deformation. In this paper, we propose a novel approach named Multi-Channel Attention SelectionGAN (SelectionGAN) that makes it possible to generate images of natural scenes in arbitrar…

Cited by 441PDFcodeScholar
2018

Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation

CVPR 2018poster

Recent works have shown the benefit of integrating Conditional Random Fields (CRFs) models into deep architectures for improving pixel-level prediction tasks. Following this line of research, in this paper we introduce a novel approach for monocular depth estimation. Similarly to previous works, our…

2016

Adapting ASR for under-resourced languages using mismatched transcriptions

ICASSP 2016accepted

Mismatched transcriptions of speech in a target language refers to transcriptions provided by people unfamiliar with the language, using English letter sequences. In this work, we demonstrate the value of such transcriptions in building an ASR system for the target language. For different languages,…

Cited by 0SourceScholar
2016

Signer-independent fingerspelling recognition with deep neural network adaptation

ICASSP 2016accepted

We study the problem of recognition of fingerspelled letter sequences in American Sign Language in a signer-independent setting. Fingerspelled sequences are both challenging and important to recognize, as they are used for many content words such as proper nouns and technical terms. Previous work ha…

Cited by 0SourceScholar