← Search

Xuan Wang

102 accepted papers

2026

A Diagnostic Study of Multi-Agent LLMs for Real-World Debates

ICML 2026poster

Multi-agent LLM debates are increasingly deployed in domains such as policy analysis and city planning, where no objective ground truth exists. Despite this, debate quality is typically evaluated using outcome-based proxies such as LLM-as-judge scores that provide little insight into whether meaning…

Cited by 0SourceScholar
2026

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization

CVPR 2026

We introduce AvatarPointillist, a novel framework for generating dynamic 4D Gaussian avatars from a single portrait image. At the core of our method is a decoder-only Transformer that autoregressively generates a point cloud for 3D Gaussian Splatting. This sequential approach allows for precise, ada

Cited by 0SourcecodeScholar
2026

BDI-based Opponent Modeling and Strategy Generation for Multi-Issue Negotiation (Student Abstract)

AAAI 2026technical

Accurately modeling opponent behaviors and integrating strategy are key challenges for multi-issue automated negotiation. Existing approaches often isolate preference learning or trend prediction and lack a unified cognitive structure with coordinated reasoning. This paper proposes a BDI (Belief-Des

Cited by 0SourcePDFScholar
2026

BIOARC: Discovering Optimal Neural Architectures for Biological Foundation Models

ICML 2026poster

Foundation models have revolutionized AI, yet biological applications often repurpose general architectures without accounting for the intrinsic structural and functional properties of distinct modalities, such as genomic and proteomic sequences. Consequently, these architectures lack the inductive …

Cited by 0SourceScholar
2026

BeyondBench: Benchmark-Free Evaluation of Reasoning in Language Models

ICLR 2026poster

Evaluating language models fairly is becoming harder as static benchmarks available on the internet risk contamination by training data. This makes it unclear whether models are truly reasoning or just recalling answers. In this paper, we introduce $\textbf{BeyondBench}$, an evaluation framework tha…

Cited by 0SourcecodeScholar
2026

BulletTime4D: Towards High Spatio-Temporal Resolution Dynamic Scene Rendering via Spike-Guided Stereo Vision

AAAI 2026technical

High spatio‑temporal resolution novel‑view scene rendering is crucial for applications such as sports analysis and scientific experiments. However, existing Dynamic Scene Rendering (DSR) approaches typically rely on conventional RGB cameras with limited frame rates, making it difficult to achieve hi

Cited by 0SourcePDFScholar
2026

Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models

ICLR 2026poster

The rapid progress of visual autoregressive (VAR) models has brought new opportunities for text-to-image generation, but also heightened safety concerns. Existing concept erasure techniques, primarily designed for diffusion models, fail to generalize to VARs due to their next-scale token prediction…

Cited by 0SourcecodeScholar
2026

JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation

AAAI 2026technical

While small language models (SLMs) have shown promise on various reasoning tasks, their ability to judge the correctness of answers remains unclear compared to large language models (LLMs). Prior work on LLM-as-a-judge frameworks typically relies on comparing candidate answers against ground-truth l

Cited by 0SourcePDFScholar
2026

MER-Tracker: Towards High-Speed 3D Point Tracking via Multi-View Event-RGB Hybrid Cameras

CVPR 2026

This paper proposes the first task for high-speed 3D point tracking using multi-view Event-RGB hybrid cameras. We design a cuboid observation device comprising 4 RGB cameras (30fps) and 2 Event cameras to synchronously capture high-speed motions, and propose MER-Tracker, a high-frame-rate 3D point-t

Cited by 0SourceScholar
2026

MimicTalker: A Multimodal Interactive and Memory-Enhanced Framework for Real-Time Dyadic 3D Head Generation

CVPR 2026

Dyadic interactive head generation aims to synthesize realistic head motions that respond both verbally and non-verbally to an interlocutor in real-time conversation. The existing works often focus on offline scenarios, and struggle with a shallow understanding of the multimodal conversational conte

Cited by 0SourceScholar
2026

PACT: Phase-Like Transition Constraints in Adapter-Based Continual Learning of Vision-Language Models

CVPR 2026

Continual Learning (CL) enables Vision-Language Models (VLMs) to acquire new capabilities while retaining prior knowledge, for example, by employing task-specific adapters. Existing CL approaches typically optimize these adapters to convergence, often with (near-)orthogonality constraints to reduce

Cited by 0SourceScholar
2026

PhaseAlign: Complex Phase Alignment for Stable Open-Vocabulary Semantic Segmentation

ICML 2026poster

Open-Vocabulary Segmentation(OVS) aims to achieve pixel-level semantic recognition from arbitrary text queries. Existing large-scale visual-linguistic models, such as CLIP, perform well in zero-shot generalization, but their image-level training objectives and real-valued cross-modal alignment mix a…

Cited by 0SourceScholar
2026

Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs

CVPR 2026

While recent vision-language models (VLMs) demonstrate strong image understanding, their ability to "think with images," i.e., to reason through multi-step visual interactions, remains limited. We introduce VISTA-Gym, a scalable training environment for incentivizing tool-integrated visual reasoning

Cited by 0SourcecodeScholar
2026

Semantic Impact–Driven Visual Scheduling in Vision-Language Models

ICML 2026poster

Vision-Language Models (VLMs) suffer from high inference latency due to long visual sequences. To enable efficient, on-demand utilization of visual information, we argue that visual necessity should be assessed by its semantic impact on the output distribution, rather than inferred from intermediate…

Cited by 0SourceScholar
2026

UIKA: Fast Universal Head Avatar from Pose-Free Images

CVPR 2026

We present UIKA, a feed-forward animatable Gaussian head model from an arbitrary number of pose-free inputs, including a single image, multi-view captures, and smartphone-captured videos. Unlike the traditional avatar method, which requires a studio-level multi-view capture system and reconstructs a

Cited by 0SourcecodeScholar
2026

X-MoGen: Unified Motion Generation Across Humans and Animals

AAAI 2026technical

Text-driven motion generation has attracted increasing attention due to its broad applications in virtual reality, animation, and robotics. While existing methods typically model human and animal motion separately, a joint cross-species approach offers key advantages, such as a unified representatio

Cited by 0SourcePDFScholar
2026

effGen: Enabling Small Language Models as Capable Autonomous Agents

ICML 2026poster

Most existing language model agentic systems today are built and optimized for large language models (e.g., GPT, Claude, Gemini) via API calls. While powerful, this approach faces several limitations including high token costs and privacy concerns for sensitive applications. We introduce $\textbf{ef…

Cited by 0SourceScholar
2025

3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial Representations

CVPR 2025poster

Recent studies have combined 3D Gaussian and 3D Morphable Models (3DMM) to construct high-quality 3D head avatars. In this line of research, existing methods either fail to capture the dynamic textures or incur significant overhead in terms of runtime speed or storage space. To this end, we propose…

Cited by 0SourcePDFScholar
2025

A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare

EMNLP 2025

The application of large language models (LLMs) in healthcare holds significant promise for enhancing clinical decision-making, medical research, and patient care. However, their integration into real-world clinical settings raises critical concerns around trustworthiness, particularly around dimens

Cited by 0SourcePDFScholar
2025

A3: Few-shot Prompt Learning of Unlearnable Examples with Cross-Modal Adversarial Feature Alignment

CVPR 2025poster

In the age of pervasive machine learning applications, protecting digital content from unauthorized use has become a pressing concern. Unlearnable examples (UEs)--data modified with imperceptible perturbations to inhibit model training while preserving human usability--have emerged as a promising ap…

Cited by 0SourcePDFScholar
2025

Adaptive Learning of High-Value Regions for Semi-Supervised Medical Image Segmentation

ICCV 2025poster

Existing semi-supervised learning methods typically mitigate the impact of unreliable predictions by suppressing low-confidence regions. However, these methods fail to explore which regions hold higher learning value and how to design adaptive learning strategies for these regions. To address these…

2025

AniMo: Species-Aware Model for Text-Driven Animal Motion Generation

CVPR 2025poster

Text-driven motion generation has made significant strides in recent years. However, most existing works focus on human motion, largely overlooking the rich and diverse behaviors of animals. Understanding and synthesizing animal motion have important applications in wildlife conservation, animal eco…

2025

BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integration

EMNLP 2025

Mixture-of-Experts (MoE) models have become increasingly powerful in multimodal learning by enabling modular specialization across modalities. However, their effectiveness remains unclear when additional modalities introduce more noise than complementary information. Existing approaches, such as the

2025

Better to Teach than to Give: Domain Generalized Semantic Segmentation via Agent Queries with Diffusion Model Guidance

ICML 2025spotlight

Domain Generalized Semantic Segmentation (DGSS) trains a model on a labeled source domain to generalize to unseen target domains with consistent contextual distribution and varying visual appearance. Most existing methods rely on domain randomization or data generation but struggle to capture the un…

2025

CONSENSAGENT: Towards Efficient and Effective Consensus in Multi-Agent LLM Interactions Through Sycophancy Mitigation

ACL 2025finding

Multi-agent large language model (LLM) systems have shown remarkable performance in tasks such as reasoning, planning, and decision-making. However, their applicability is limited by challenges such as high computational costs and robustness issues. In this work, we identify and systematically evalu…

2025

CROSSAGENTIE: Cross-Type and Cross-Task Multi-Agent LLM Collaboration for Zero-Shot Information Extraction

ACL 2025finding

Large language models (LLMs) excel in generating unstructured text. However, they struggle with producing structured output while maintaining accuracy in zero-shot information extraction (IE), such as named entity recognition (NER) and relation extraction (RE). To address these challenges, we propos…

2025

CVIRO: A Consistent and Tightly-Coupled Visual-Inertial-Ranging Odometry on Lie Groups

IROS 2025

Ultra-Wideband (UWB) is widely used to mitigate drift in visual-inertial odometry (VIO) systems. Consistency is crucial for ensuring the estimation accuracy of a UWB-aided VIO system. An inconsistent estimator can degrade localization performance, where the inconsistency primarily arises from two ma

Cited by 0SourceScholar
2025

DEBATE, TRAIN, EVOLVE: Self‐Evolution of Language Model Reasoning

EMNLP 2025

Large language models (LLMs) have improved significantly in their reasoning through extensive training on massive datasets. However, relying solely on additional data for improvement is becoming increasingly impractical, highlighting the need for models to autonomously enhance their reasoning withou

2025

DFNeRF: Disentangled Facial Neural Radiance Fields for Text-based Editing of Free-view Talking Head

ICASSP 2025accepted

In this paper, we propose a text-based approach that can edit the speech content of a free-view talking head based on its transcript. The core of our method is to establish the relationship between phonemes and head attributes. To avoid discontinuities in head pose and facial expressions caused by e…

Cited by 0SourceScholar
2025

DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking Heads

ICCV 2025poster

In this work, we investigate the generation of high-fidelity, audio-driven 3D Gaussian talking heads from monocular videos. We present DGTalker, an innovative framework designed for real-time, high-fidelity, and 3D-aware talking head synthesis. By leveraging Gaussian generative priors and treating t…

Cited by 0SourcePDFScholar
2025

Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

ICCV 2025poster

This paper addresses the challenge of high-fidelity view synthesis of humans with sparse-view videos as input. Previous methods solve the issue of insufficient observation by leveraging 4D diffusion models to generate videos at novel viewpoints. However, the generated videos from these models often…

2025

Diffusion-based Realistic Listening Head Generation via Hybrid Motion Modeling

CVPR 2025highlight

Listening head generation aims to synthesize non-verbal responsive listening head videos that naturally react to a certain speaker, for which, both realistic head movements, expressive facial expressions, and high visual qualities are expected. Previous approaches typically follow a two-stage pipeli…

Cited by 0SourcePDFScholar
2025

DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations

CVPR 2025poster

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward…

2025

Enhanced Equilibria-Solving via Private Information Pre-Branch Structure in Adversarial Team Games

UAI 2025

In ex ante coordinated adversarial team games (ATGs), a team competes against an adversary, and team members can only coordinate their strategies before the game starts. The team-maxmin equilibrium with correlation (TMECor) is a suitable solution concept for extensive-form sequential ATGs. One class

Cited by 0SourcePDFScholar
2025

Enhancing Target-unspecific Tasks through a Features Matrix

ICML 2025poster

Recent developments in prompt learning of large Vision-Language Models (VLMs) have significantly improved performance in target-specific tasks. However, these prompting methods often struggle to tackle the target-unspecific or generalizable tasks effectively. It may be attributed to the fact that o…

Cited by 0SourcePDFScholar
2025

Fine-Grained 3D Gaussian Head Avatars Modeling from Static Captures via Joint Reconstruction and Registration

ICCV 2025poster

Recently, 3D head avatar modeling based on 3D Gaussians has demonstrated significant advantages in rendering quality and efficiency, given sufficient data. Some efforts have begun to train prior models on large datasets to develop generalizable 3D Gaussian head avatar modeling methods. Unfortunately…

Cited by 0SourcePDFScholar
2025

HERA: Hybrid Explicit Representation for Ultra-Realistic Head Avatars

CVPR 2025poster

We introduce a novel approach to creating ultra-realistic head avatars and rendering them in real time (\geq 30 fps at 2048 x1334 resolution). First, we propose a hybrid explicit representation that combines the advantages of two primitive based efficient rendering techniques. UV-mapped 3D mesh is u…

Cited by 0SourcePDFScholar
2025

Human-Robot Co-Transportation using Disturbance-Aware MPC with Pose Optimization

IROS 2025

This paper proposes a new control algorithm for human-robot co-transportation using a robot manipulator equipped with a mobile base and a robotic arm. We integrate the regular Model Predictive Control (MPC) with a novel pose optimization mechanism to more efficiently mitigate disturbances (such as h

Cited by 1SourceScholar
2025

Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Recently, open-vocabulary semantic segmentation has garnered growing attention. Most current methods leverage vision-language models like CLIP to recognize unseen categories through their zero-shot capabilities. However, CLIP struggles to establish potential spatial dependencies among scene objects…

Cited by 0SourcePDFScholar
2025

Joint Pedestrian and Vehicle Traffic Optimization in Urban Environments using Reinforcement Learning

IROS 2025

Reinforcement learning (RL) holds significant promise for adaptive traffic signal control. While existing RL-based methods demonstrate effectiveness in reducing vehicular congestion, their predominant focus on vehicle-centric optimization leaves pedestrian mobility needs and safety challenges unaddr

Cited by 6SourcecodeScholar
2025

Learning Dynamical Coupled Operator For High-dimensional Black-box Partial Differential Equations

IJCAI 2025

The deep operator networks (DON), a class of neural operators that learn mappings between function spaces, have recently emerged as surrogate models for parametric partial differential equations (PDEs). However, their full potential for accurately approximating general black-box PDEs remains underex

2025

Lie Detector: Unified Backdoor Detection via Cross-Examination Framework

NeurIPS 2025poster

Institutions with limited data and computing resources often outsource model training to third-party providers in a semi-honest setting, assuming adherence to prescribed training protocols with pre-defined learning paradigm (e.g., supervised or semi-supervised learning). However, this practice can i…

Cited by 0SourceScholar
2025

MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models

ACL 2025finding

The widespread use of Large Multimodal Models (LMMs) has raised concerns about model toxicity. However, current research mainly focuses on explicit toxicity, with less attention to some more implicit toxicity regarding prejudice and discrimination. To address this limitation, we introduce a subtler…

2025

MIAT: Maneuver-Intention-Aware Transformer for Spatio-Temporal Trajectory Prediction

IROS 2025

Accurate vehicle trajectory prediction is critical for safe and efficient autonomous driving, especially in mixed traffic environments when both human-driven and autonomous vehicles co-exist. However, uncertainties introduced by inherent driving behaviors—such as acceleration, deceleration, and left

Cited by 4SourcecodeScholar
2025

Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMs

EMNLP 2025

Recent advances in Multimodal Large Language Models (MLLMs) have enhanced their versatility as they integrate a growing number of modalities. Considering the heavy cost of training MLLMs, it is efficient to reuse the existing ones and extend them to more modalities through Modality-incremental Conti

Cited by 0SourcePDFScholar
2025

No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion Models

NeurIPS 2025poster

Enhancing the cross-domain generalization of 3D semantic segmentation is a pivotal task in computer vision that has recently gained increasing attention. Most existing methods, whether using consistency regularization or cross-modal feature fusion, focus solely on individual objects while overlookin…

Cited by 0SourcecodeScholar
2025

Robust Online Calibration for UWB-Aided Visual-Inertial Navigation with Bias Correction

IROS 2025

This paper presents a novel robust online calibration framework for Ultra-Wideband (UWB) anchors in UWB-aided Visual-Inertial Navigation Systems (VINS). Accurate anchor positioning, a process known as calibration, is crucial for integrating UWB ranging measurements into state estimation. While sever

Cited by 0SourceScholar
2025

Self-Ensembling Gaussian Splatting for Few-Shot Novel View Synthesis

ICCV 2025poster

3D Gaussian Splatting (3DGS) has demonstrated remarkable effectiveness in novel view synthesis (NVS). However, 3DGS tends to overfit when trained with sparse views, limiting its generalization to novel viewpoints. In this paper, we address this overfitting issue by introducing Self-Ensembling Gaussi…

2025

Spike4DGS: Towards High-Speed Dynamic Scene Rendering with 4D Gaussian Splatting via a Spike Camera Array

NeurIPS 2025poster

Spike camera with high temporal resolution offers a new perspective on high-speed dynamic scene rendering. Most existing rendering methods rely on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) for static scenes using a monocular spike camera. However, these methods struggle with dyna…

Cited by 0SourcecodeScholar
2025

SusGen-GPT: A Data-Centric LLM for Financial NLP and Sustainability Report Generation

NAACL 2025findings

The rapid growth of the financial sector and the increasing focus on Environmental, Social, and Governance (ESG) considerations have created a pressing need for advanced natural language processing (NLP) tools. Despite recent advancements, there is still a notable absence of open-source Large Langua…

2025

Towards Building Human-like Smart Agents in Modern 3D Video Games (Student Abstract)

AAAI 2025technical

In recent years, reinforcement learning has been widely applied in the field of games. However, most studies focus on assisting agents to achieve victory, with less attention paid to whether the agents exhibit human-like characteristics. In order to build human-like agents with high performance, we…

Cited by 0SourcePDFScholar
2024

A Pre-convolved Representation for Plug-and-Play Neural Illumination Fields

AAAI 2024technical

Recent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light mo…

Cited by 2SourcePDFScholar
2024

Bi-CL: A Reinforcement Learning Framework for Robots Coordination Through Bi-level Optimization

IROS 2024poster

In multi-robot systems, achieving coordinated missions remains a significant challenge due to the coupled nature of coordination behaviors and the lack of global information for individual robots. To mitigate these challenges, this paper introduces a novel approach, Bi-level Coordination Learning (B…

Cited by 3SourceScholar
2024

Decoupling Degradations with Recurrent Network for Video Restoration in Under-Display Camera

AAAI 2024technical

Under-display camera (UDC) systems are the foundation of full-screen display devices in which the lens mounts under the display. The pixel array of light-emitting diodes used for display diffracts and attenuates incident light, causing various degradations as the light intensity changes. Unlike gene…

2024

Learning Coordinated Maneuver in Adversarial Environments

IROS 2024poster

This paper aims to solve the coordination of a team of robots traversing a route in the presence of adversaries with random positions. Our goal is to minimize the overall cost of the team, which is determined by (i) the accumulated risk when robots stay in adversary-impacted zones and (ii) the missi…

Cited by 0SourceScholar
2024

Motion-adaptive Separable Collaborative Filters for Blind Motion Deblurring

CVPR 2024poster

Eliminating image blur produced by various kinds of motion has been a challenging problem. Dominant approaches rely heavily on model capacity to remove blurring by reconstructing residual from blurry observation in feature space. These practices not only prevent the capture of spatially variable mot…

2024

On the Approximation Risk of Few-Shot Class-Incremental Learning

ECCV 2024poster

"Few-Shot Class-Incremental Learning (FSCIL) aims to learn new concepts with few training samples while preserving previously acquired knowledge. Although promising performance has been achieved, there remains an underexplored aspect regarding the basic statistical principles underlying FSCIL. There…

2024

On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

EMNLP 2024finding

Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences. Central to RLHF is learning a reward function for scoring human preferences. Two main approaches for learning a reward model are 1) training an EXplicit Reward Model (EXRM) a…

2024

P2P: Transforming from Point Supervision to Explicit Visual Prompt for Object Detection and Segmentation

IJCAI 2024poster

Point-supervised vision tasks, including detection and segmentation, aiming to learn a network that transforms from points to pseudo labels, have attracted much attention in recent years. However, the lack of precise object size and boundary annotations in the point-supervised condition results in a…

2024

Real-time 3D-aware Portrait Editing from a Single Image

ECCV 2024poster

"This work presents , a practical method that can efficiently edit a face image following given prompts, like reference images or text descriptions, in a 3D-aware manner. To this end, a lightweight module is distilled from a 3D portrait generator and a text-to-image model, which provide prior knowle…

2024

Scaling Team Coordination on Graphs with Reinforcement Learning

ICRA 2024poster

This paper studies Reinforcement Learning (RL) techniques to enable team coordination behaviors in graph environments with support actions among teammates to reduce the costs of traversing certain risky edges in a centralized manner. While classical approaches can solve this non-standard multi-agent…

Cited by 6SourceScholar
2024

See and Think: Embodied Agent in Virtual Environment

ECCV 2024poster

"Large language models (LLMs) have achieved impressive pro-gress on several open-world tasks. Recently, using LLMs to build embodied agents has been a hotspot. This paper proposes STEVE, a comprehensive and visionary embodied agent in the Minecraft virtual environment. STEVE comprises three key comp…

Cited by 37SourcePDFScholar
2024

Team Coordination on Graphs: Problem, Analysis, and Algorithms

IROS 2024poster

Team Coordination on Graphs with Risky Edges (TCGRE) is a recently emerged problem, in which a robot team collectively reduces graph traversal cost through support from one robot to another when the latter traverses a risky edge. Resembling the traditional Multi-Agent Path Finding (MAPF) problem, bo…

Cited by 3SourceScholar
2024

TriageAgent: Towards Better Multi-Agents Collaborations for Large Language Model-Based Clinical Triage

EMNLP 2024finding

The global escalation in emergency department patient visits poses significant challenges to efficient clinical management, particularly in clinical triage. Traditionally managed by human professionals, clinical triage is susceptible to substantial variability and high workloads. Although large lang…

2023

3D GAN Inversion With Facial Symmetry Prior

CVPR 2023poster

Recently, a surge of high-quality 3D-aware GANs have been proposed, which leverage the generative power of neural rendering. It is natural to associate 3D GANs with GAN inversion methods to project a real image into the generator's latent space, allowing free-view consistent synthesis and editing, r…

Cited by 45SourcePDFScholar
2023

CiT-Net: Convolutional Neural Networks Hand in Hand with Vision Transformers for Medical Image Segmentation

IJCAI 2023poster

The hybrid architecture of convolutional neural networks (CNNs) and Transformer are very popular for medical image segmentation. However, it suffers from two challenges. First, although a CNNs branch can capture the local image features using vanilla convolution, it cannot achieve adaptive feature l…

2023

FSI: Frequency and Spatial Interactive Learning for Image Restoration in Under-Display Cameras

ICCV 2023poster

Under-display camera (UDC) systems remove the screen notch for bezel-free displays and provide a better interactive experience. The main challenge is that the pixel array of light-emitting diodes used for display diffracts and attenuates the incident light, leading to complex degradation. Existing m…

Cited by 22PDFScholar
2023

GIFD: A Generative Gradient Inversion Method with Feature Domain Optimization

ICCV 2023poster

Federated Learning (FL) has recently emerged as a promising distributed machine learning framework to preserve clients' privacy, by allowing multiple clients to upload the gradients calculated from their local data to a central server. Recent studies find that the exchanged gradients also take the r…

Cited by 40PDFcodeScholar
2023

High-Fidelity Clothed Avatar Reconstruction From a Single Image

CVPR 2023poster

This paper presents a framework for efficient 3D clothed avatar reconstruction. By combining the advantages of the high accuracy of optimization-based methods and the efficiency of learning-based methods, we propose a coarse-to-fine way to realize a high-fidelity clothed avatar reconstruction (CAR)…

2023

High-Fidelity Facial Avatar Reconstruction From Monocular Video With Generative Priors

CVPR 2023poster

High-fidelity facial avatar reconstruction from a monocular video is a significant research problem in computer graphics and computer vision. Recently, Neural Radiance Field (NeRF) has shown impressive novel view rendering results and has been considered for facial avatar reconstruction. However, th…

2023

Local Implicit Ray Function for Generalizable Radiance Field Representation

CVPR 2023poster

We propose LIRF (Local Implicit Ray Function), a generalizable neural rendering approach for novel view rendering. Current generalizable neural radiance fields (NeRF) methods sample a scene with a single ray per pixel and may therefore render blurred or aliased views when the input views and rendere…

Cited by 30SourcePDFScholar
2023

Local-to-Global Registration for Bundle-Adjusting Neural Radiance Fields

CVPR 2023poster

Neural Radiance Fields (NeRF) have achieved photorealistic novel views synthesis; however, the requirement of accurate camera poses limits its application. Despite analysis-by-synthesis extensions for jointly learning neural 3D representations and registering camera frames exist, they are susceptibl…

Cited by 77SourcePDFScholar
2023

MEGClass: Extremely Weakly Supervised Text Classification via Mutually-Enhancing Text Granularities

EMNLP 2023long findings

Text classification is essential for organizing unstructured text. Traditional methods rely on human annotations or, more recently, a set of class seed words for supervision, which can be costly, particularly for specialized or emerging domains. To address this, using class surface names alone as ex…

Cited by 0SourcecodeScholar
2023

Next3D: Generative Neural Texture Rasterization for 3D-Aware Head Avatars

CVPR 2023highlight

3D-aware generative adversarial networks (GANs) synthesize high-fidelity and multi-view-consistent facial images using only collections of single-view 2D imagery. Towards fine-grained control over facial attributes, recent efforts incorporate 3D Morphable Face Model (3DMM) to describe deformation in…

2023

ReactIE: Enhancing Chemical Reaction Extraction with Weak Supervision

ACL 2023findings

Structured chemical reaction information plays a vital role for chemists engaged in laboratory work and advanced endeavors such as computer-aided drug design. Despite the importance of extracting structured reactions from scientific literature, data annotation for this purpose is cost-prohibitive du…

Cited by 8SourcePDFScholar
2023

SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

CVPR 2023poster

Generating talking head videos through a face image and a piece of speech audio still contains many challenges. i.e., unnatural head movement, distorted expression, and identity modification. We argue that these issues are mainly caused by learning from the coupled 2D motion fields. On the other han…

2023

Team Coordination on Graphs with State-Dependent Edge Costs

IROS 2023poster

This paper studies a team coordination problem in a graph environment. Specifically, we incorporate “support” action which an agent can take to reduce the cost for its teammate to traverse some high cost edges. Due to this added feature, the graph traversal is no longer a standard multi-agent path p…

Cited by 9SourceScholar
2023

Text Augmented Open Knowledge Graph Completion via Pre-Trained Language Models

ACL 2023findings

The mission of open knowledge graph (KG) completion is to draw new findings from known facts. Existing works that augment KG completion require either (1) factual triples to enlarge the graph reasoning space or (2) manually designed prompts to extract knowledge from a pre-trained language model (PLM…

2023

UV Volumes for Real-Time Rendering of Editable Free-View Human Performance

CVPR 2023poster

Neural volume rendering enables photo-realistic renderings of a human performer in free-view, a critical task in immersive VR/AR applications. But the practice is severely limited by high computational costs in the rendering process. To solve this problem, we propose the UV Volumes, a new approach t…

2022

Deblur-NeRF: Neural Radiance Fields From Blurry Images

CVPR 2022poster

Neural Radiance Field (NeRF) has gained considerable attention recently for 3D scene reconstruction and novel view synthesis due to its remarkable synthesis quality. However, image blurriness caused by defocus or motion, which often occurs when capturing scenes in the wild, significantly degrades it…

Cited by 289PDFcodeScholar
2022

FENeRF: Face Editing in Neural Radiance Fields

CVPR 2022poster

Previous portrait image generation methods roughly fall into two categories: 2D GANs and 3D-aware GANs. 2D GANs can generate high fidelity portraits but with low view consistency. 3D-aware GAN methods can maintain view consistency but their generated images are not locally editable. To overcome thes…

Cited by 169PDFcodeScholar
2022

Hallucinated Neural Radiance Fields in the Wild

CVPR 2022poster

Neural Radiance Fields (NeRF) has recently gained popularity for its impressive novel view synthesis ability. This paper studies the problem of hallucinated NeRF: i.e., recovering a realistic NeRF at a different time of day from a group of tourism images. Existing solutions adopt NeRF with a control…

Cited by 137PDFcodeScholar
2022

Seed-Guided Topic Discovery with Out-of-Vocabulary Seeds

NAACL 2022long

Discovering latent topics from text corpora has been studied for decades. Many existing topic models adopt a fully unsupervised setting, and their discovered topics may not cater to users’ particular interests due to their inability of leveraging user guidance. Although there exist seed-guided topic…

2022

StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGAN

ECCV 2022poster

"One-shot talking face generation aims at synthesizing a high-quality talking face video from an arbitrary portrait image, driven by a video or an audio segment. In this work, we provide a solution from a novel perspective that differs from existing frameworks. We first investigate the latent featur…

2022

Underwater Small Target Detection Based on Deformable Convolutional Pyramid

ICASSP 2022accepted

Due to the problem of severe deformation, occlusion, diversified scenarios, general object detection methods cannot achieve satisfactory results in underwater object detection tasks. In this paper, we propose a two-stage Underwater Small Target Detection (USTD) network. In the proposed USTD, the Def…

Cited by 0SourceScholar
2021

COVID-19 Literature Knowledge Graph Construction and Drug Repurposing Report Generation

NAACL 2021system demonstrations

To combat COVID-19, both clinicians and scientists need to digest the vast amount of relevant biomedical knowledge in literature to understand the disease mechanism and the related biological functions. We have developed a novel and comprehensive knowledge discovery framework, COVID-KG to extract fi…

2021

ChemNER: Fine-Grained Chemistry Named Entity Recognition with Ontology-Guided Distant Supervision

EMNLP 2021main

Scientific literature analysis needs fine-grained named entity recognition (NER) to provide a wide range of information for scientific discovery. For example, chemistry research needs to study dozens to hundreds of distinct, fine-grained entity types, making consistent and accurate annotation diffic…

2021

Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-Training

EMNLP 2021main

We study the problem of training named entity recognition (NER) models using only distantly-labeled data, which can be automatically obtained by matching entity mentions in the raw text with entity types in a knowledge base. The biggest challenge of distantly-supervised NER is that the distant super…

2021

Noise Robust Named Entity Understanding for Voice Assistants

NAACL 2021industry

Named Entity Recognition (NER) and Entity Linking (EL) play an essential role in voice assistant interaction, but are challenging due to the special difficulties associated with spoken user queries. In this paper, we propose a novel architecture that jointly solves the NER and EL tasks by combining…

Cited by 5SourcePDFScholar