← Search

Hao Li

302 accepted papers

2026

AdapTok: Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space

CVPR 2026

We propose AdapTok, an adaptive temporal causal video tokenizer that can flexibly allocate tokens for different frames based on video content. AdapTok is equipped with a block-wise masking strategy that randomly drops tail tokens of each block during training, and a block causal scorer to predict th

Cited by 0SourcecodeScholar
2026

Agentmandering: A Game-Theoretic Framework for Fair Redistricting via Large Language Model Agents

AAAI 2026technical

Redistricting plays a central role in shaping how votes are translated into political power. While existing computational methods primarily aim to generate large ensembles of legally valid districting plans, they often neglect the strategic dynamics involved in the selection process. This oversight

Cited by 0SourcePDFScholar
2026

Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferences

CVPR 2026

Narrative image generation aims to create images featuring multiple distinct characters while capturing their interrelationships, posing significant challenges for current text-to-image diffusion models. As a result, general personalized methods often suffer from poor semantic alignment, identity bl

Cited by 0SourceScholar
2026

An Improved Sliding Mode Control Algorithm for Integrated Driving and Sensing Electromagnetic Actuators

RA-L 2026

The design of control algorithms in a multi-coupling physical field is critical for advancing the application of integrated driving-and-sensing electromagnetic actuators. Accordingly, we propose a discrete fast terminal sliding mode control (FTSM) algorithm enhanced with a radial basis function neur

Cited by 0SourceScholar
2026

AviaSafe: A Physics-Informed Data-Driven Model for Aviation Safety-Critical Cloud Forecasts

CVPR 2026

Current AI weather forecasting models predict conventional atmospheric variables but cannot distinguish between cloud microphysical species critical for aviation safety. We introduce AviaSafe, a hierarchical, physics-informed neural forecaster that produces global, six-hourly predictions of these fo

Cited by 0SourceScholar
2026

Beyond Extrapolation: Knowledge Utilization Paradigm with Bidirectional Inspiration for Time Series Forecasting

ICML 2026poster

Time-series forecasting is critical in various scenarios, such as energy, transportation, and public health. However, most existing forecasters rely primarily on one-way inference, \textit{i.e.}, mapping \textbf{history} to \textbf{target}, and overlook the structural information provided by a revis…

Cited by 0SourceScholar
2026

Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale

ICML 2026poster

Large Language Models (LLMs) have rapidly advanced, with Gemini-3-Pro setting a new performance milestone. In this work, we explore collective intelligence as an alternative to monolithic scaling, and demonstrate that open-source LLMs' collaboration can surpass Gemini-3-Pro. We first revisit LLM rou…

Cited by 0SourceScholar
2026

Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT

CVPR 2026

Multiple-choice question answering (MCQA) has been a popular format for evaluating and reinforcement fine-tuning (RFT) of modern multimodal language models. Its constrained output format allows for simplified, deterministic automatic verification.However, we find that the options may leak exploitabl

Cited by 0SourceScholar
2026

Boosting World Models Learning via Latent-Space Value Alignment

ICML 2026poster

Model-based reinforcement learning aims to construct world models for efficient sampling. Current mainstream algorithms can be broadly categorized into two paradigms: maximum likelihood and value-aware world models. The former employs structured Recurrent/Transformer State-Space Models to capture en…

Cited by 0SourceScholar
2026

CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT Tracking

AAAI 2026technical

RGB-Thermal (RGBT) tracking aims to exploit visible and thermal infrared modalities for robust all-weather object tracking. However, existing RGBT trackers struggle to resolve modality discrepancies, which poses great challenges for robust feature representation. This limitation hinders effective cr

Cited by 0SourcePDFScholar
2026

Chain of World: World Model Thinking in Latent Motion

CVPR 2026

Vision-Language-Action (VLA) models are promising for embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future frames, but waste capacity reconstructing redundant backgrounds. To overcom

Cited by 0SourcecodeScholar
2026

DcSplat: Dual-Constraint Human Gaussian Splatting with Latent Multi-View Consistency

AAAI 2026technical

Human Novel View Synthesis (HNVS) aims to synthesize photorealistic human images from novel viewpoints given observations from known views. Despite significant advances achieved by existing methods such as NeRF, diffusion models, and 3DGS, they still face substantial challenges in achieving stable m

Cited by 0SourcePDFScholar
2026

DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers

CVPR 2026

Recent breakthroughs in Diffusion Transformers (DiTs) have revolutionized the field of visual synthesis due to their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for represent

Cited by 0SourcecodeScholar
2026

Dual Reactive Planning for Heterogeneous Robots With Evolving Capabilities in Unknown Environments

RA-L 2026

Heterogeneous robot teams executing Linear Temporal Logic (LTL) missions are usually modeled with fixed robot capabilities. In practice, capabilities may change during execution through tool acquisition, sensor activation, or module reconfiguration, making previously infeasible tasks executable and

Cited by 0SourceScholar
2026

Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation

ICLR 2026poster

Recent advances in video generation have enabled thrilling experiences in producing realistic videos driven by scalable diffusion transformers. However, they usually fail to produce satisfactory outputs that are aligned to users' authentic demands and preferences. In this work, we introduce Dual-Ite…

Cited by 0SourcecodeScholar
2026

FDP: A Frequency-Decomposition Preprocessing Pipeline for Unsupervised Anomaly Detection in Brain MRI

AAAI 2026technical

Due to the diversity of brain anatomy and the scarcity of annotated data, supervised anomaly detection for brain MRI remains challenging, driving the development of unsupervised anomaly detection (UAD) approaches. Current UAD methods typically utilize synthetically generated noise perturbations on

Cited by 0SourcePDFScholar
2026

FaithFusion: Harmonizing Reconstruction and Generation via Pixel-wise Information Gain

CVPR 2026

In controllable driving-scene reconstruction and 3D scene generation, maintaining geometric fidelity while synthesizing visually plausible appearance under large viewpoint shifts is crucial. However, effective fusion of geometry-based 3DGS and appearance-driven diffusion models faces inherent challe

Cited by 0SourceScholar
2026

From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

CVPR 2026

Keypoint-based matching is a fundamental component of modern 3D vision systems, such as Structure-from-Motion (SfM) and SLAM. Most existing learning-based methods are trained on image pairs, a paradigm that fails to explicitly optimize for the long-term trackability of keypoints across sequences und

Cited by 0SourcecodeScholar
2026

From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty Calibration

AAAI 2026technical

Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of tasks, yet they generally lack self-awareness, often displaying overconfidence when confronted with questions beyond their knowledge boundaries. This limitation severely hinders their trustworthiness in high

Cited by 0SourcePDFScholar
2026

From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors

ICLR 2026poster

Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, o…

Cited by 0SourcecodeScholar
2026

Group Editing: Edit Multiple Images in One Go

CVPR 2026

In this paper, we tackle the problem of performing consistent and unified modifications across a set of related images. This task is particularly challenging because these images may vary significantly in pose, viewpoint, and spatial layout. Achieving coherent edits requires establishing reliable co

Cited by 9SourcecodeScholar
2026

Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence

ICML 2026oral

The pursuit of spatial intelligence fundamentally relies on access to large-scale, fine-grained 3D data. However, existing approaches predominantly construct spatial understanding benchmarks by generating question–answer (QA) pairs from a limited number of manually annotated datasets, rather than sy…

Cited by 0SourceScholar
2026

Hybrid Vector-Occupancy Field for Robust Implicit 3D Surface Reconstruction

AAAI 2026technical

We introduce the Hybrid Vector-Occupancy Field (HVOF), a new implicit 3D representation for reconstructing both open and closed surfaces from sparse point clouds. Existing approaches, such as occupancy field and signed distance fields, face severe limitations. They struggle with open surfaces, while

Cited by 0SourcePDFScholar
2026

HybridOM: Hybrid Physics-Based and Data-Driven Global Ocean Modeling with Efficient Regional Downscaling

ICML 2026poster

Global ocean modeling is vital for climate science but struggles to balance computational efficiency with accuracy. Traditional numerical solvers are accurate but computationally expensive, while pure deep learning approaches, though fast, often lack physical consistency and long-term stability. To …

Cited by 0SourceScholar
2026

HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics Prediction

CVPR 2026

Spatial Transcriptomics (ST) merges the benefits of pathology images and gene expression, linking molecular profiles with tissue structure to analyze spot-level function comprehensively. Predicting gene expression from histology images is a cost-effective alternative to expensive ST technologies. Ho

Cited by 0SourcecodeScholar
2026

ICL-Router: In-Context Learned Model Representations for LLM Routing

AAAI 2026technical

Large language models (LLMs) often exhibit complementary strengths. Model routing harnesses these strengths by dynamically directing each query to the most suitable model, given a candidate model pool. However, routing performance relies on accurate model representations, and adding new models typic

Cited by 0SourcePDFScholar
2026

IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction

ICLR 2026poster

Humans naturally perceive the geometric structure and semantic content of a 3D world as intertwined dimensions, enabling coherent and accurate understanding of complex scenes. However, most prior approaches prioritize training large geometry models for low-level 3D reconstruction and treat high-leve…

Cited by 0SourcecodeScholar
2026

IMPASTO: Integrating Model-Based Planning with Learned Dynamics Models for Robotic Oil Painting Reproduction

ICRA 2026poster

Robotic reproduction of oil paintings using soft brushes and pigments requires force-sensitive control of deformable tools, prediction of brushstroke effects, and multi-step stroke planning, often without human step-by-step demonstrations or faithful simulators. Given only a sequence of target oil p…

2026

Identity-Aware Vision-Language Model for Explainable Face Forgery Detection

AAAI 2026technical

Recent advances in generative artificial intelligence have enabled the creation of highly realistic image forgeries, raising significant concerns about digital media authenticity. While existing detection methods demonstrate promising results on benchmark datasets, they face critical limitations in

Cited by 0SourcePDFScholar
2026

Large Language Models as Topological Thinkers: A Benchmark on Graph Persistent Homology

ICML 2026poster

Large language models (LLMs) are increasingly used in scientific discovery, system modeling, and decision-making, prompting interest in their ability to reason over complex structured data. Existing benchmarks primarily focus on static or local graph reasoning, overlooking the high-order structures …

Cited by 0SourceScholar
2026

LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning

ICML 2026poster

Current chemical large language models (LLMs) predominantly rely on explicit Chain-of-Thought (CoT) to solve complex reasoning problems. However, forcing nonverbal tacit chemical logic into discrete natural language imposes a fundamental ``modality mismatch,'' creating an artificial bottleneck for r…

Cited by 0SourceScholar
2026

MHopReg: Efficient Hierarchical Multi-Hop Graph Search for Point Cloud Registration

CVPR 2026

Outlier rejection for correspondence-based point cloud registration confronts two fundamental challenges in real-world scenarios. First, low-overlap regions yield sparse and fragmented inlier distributions that are difficult to discover using conventional one-step global search strategies. Second, l

Cited by 0SourceScholar
2026

MN-Diff: Diffusion Parameterized MoE-NCDE for Continuous Time Series Generation with Irregular Observations

ICML 2026poster

Time series generation (TSG) is widely used across domains, yet most existing methods assume regular sampling and fixed output resolutions. These assumptions are often violated in practice, where observations are irregular and sparse, while downstream applications require continuous and high-resolut…

Cited by 0SourceScholar
2026

MoCa: Modeling Object Consistency for 3D Camera Control in Video Generation

ICLR 2026poster

Camera control is important in text-to-video generation for achieving realistic scene navigation and view synthesis. This control is defined by parameters that describe movement through 3D space, thereby introducing a 3D consistency into the generation process. A core challenge for existing methods…

Cited by 0SourceScholar
2026

OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer

CVPR 2026

General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., camera intrinsics, poses, and depth maps). To address this issue, we introduce OmniVGGT, a novel framework that can effecti

Cited by 0SourcecodeScholar
2026

PhenoBrain: Phenotype-Conditioned Long-Range Communication for Multi-Modal Brain Network Analysis

ICML 2026oral

Multi-modal brain network analysis aims to predict neuropsychiatric status from functional connectomes with heterogeneous phenotypes. However, most existing methods treat phenotypes as auxiliary features and perform late fusion, implicitly assuming that the connectome representation should be learne…

Cited by 0SourceScholar
2026

PhysGM: Large Physical Gaussian Model for Feed-Forward 4D Synthesis

CVPR 2026

Despite advances in physics-based 3D motion synthesis, current methods face key limitations: reliance on pre-reconstructed 3D Gaussian Splatting (3DGS) built from dense multi-view images with time-consuming per-scene optimization; physics integration via either inflexible, hand-specified attributes

Cited by 0SourcecodeScholar
2026

ProbeMDE: Uncertainty-Guided Active Proprioception for Monocular Depth Estimation in Surgical Robotics

ICRA 2026poster

Monocular depth estimation (MDE) provides a useful tool for robotic perception, but its predictions are often uncertain and inaccurate in challenging environments such as surgical scenes where textureless surfaces, specular reflections, and occlusions are common. To address this, we propose ProbeMDE…

2026

RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation

CVPR 2026

RGB-Thermal (RGBT) tracking aims to achieve robust object localization across diverse environmental conditions by fusing visible and thermal infrared modalities. However, existing RGBT trackers rely solely on initial-frame visual information for target modeling, failing to adapt to appearance variat

Cited by 0SourcecodeScholar
2026

Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction

RSS 2026poster

3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models lack the precision and metric consistency required for phys…

Cited by 0SourceScholar
2026

RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation

ICLR 2026poster

Large language and vision-language models have inspired end-to-end vision-language-action (VLA) systems in robotics, yet existing robot datasets remain costly, embodiment-specific, and insufficient, limiting robustness and generalization. Recent approaches address this by adopting a plan-then-execut…

Cited by 0SourcecodeScholar
2026

Robust Vision-Language Models via Manifold-Adversarial Adapters

ICML 2026poster

Vision-language models (VLMs) have progressed rapidly with large-scale high-quality data and adaptation strategies, yet remain brittle under real-world corruptions, where both visual recognition and language-grounded reasoning degrade. Beyond cascaded image restoration, a natural alternative is para…

Cited by 0SourceScholar
2026

SEF-MAP: Subspace-Decomposed Expert Fusion for Robust Multimodal HD Map Prediction

ICRA 2026poster

High-definition (HD) maps are essential for autonomous driving, yet multi-modal fusion often suffers from inconsistency between camera and LiDAR modalities, leading to performance degradation under low-light conditions, occlusions, or sparse point clouds. To address this, we propose SEF-MAP, a Subsp…

2026

SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models

ICML 2026poster

Preference learning has garnered extensive attention as an effective technique for aligning diffusion models with human preferences in visual generation tasks. However, existing alignment approaches such as Diffusion-DPO suffer from two fundamental challenges: training instability caused by high gra…

Cited by 0SourceScholar
2026

SRGCD: Stability-Driven Region Growth Framework for 3D Change Detection

CVPR 2026

With the growing accessibility of large-scale 3D point clouds from LiDAR and photogrammetric techniques, 3D change detection (3DCD) has become essential for understanding dynamic scenes. Existing methods typically formulate this as segmentation, treating each point independently for binary classific

Cited by 0SourceScholar
2026

SeesawNet: Towards Non-stationary Time Series Forecasting with Balanced Modeling of Common and Specific Dependencies

IJCAI 2026

Instance normalization (IN) is widely used in non-stationary multivariate time series forecasting to reduce distribution shifts and highlight common patterns across samples. However, IN can over-smooth instance-specific structural information that is essential for modeling temporal and cross-channel

Cited by 0Scholar
2026

SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models

CVPR 2026

Diffusion models have demonstrated exceptional generative capabilities but are computationally intensive, posing significant challenges for deployment in resource-constrained or latency-sensitive environments.Quantization offers an effective means to reduce model size and computational cost, with po

Cited by 0SourcecodeScholar
2026

Sheaf Neural Networks on SPD Manifolds: Second-Order Geometric Representation Learning

ICML 2026poster

Graph neural networks face two fundamental challenges rooted in the linear structure of Euclidean vector spaces: (1) Current architectures represent geometry through vectors (directions, gradients), yet many tasks require matrix-valued representations that capture relationships between directions—su…

Cited by 0SourceScholar
2026

Task-Aware Mechanism: Hybrid MoE Vision Tower Towards Holistic Video Understanding

ICML 2026poster

Does \emph{Comprehending the main idea of a 2-hour movie} and \emph{Counting the birds appearing in a 15-second clip} really warrant the same video processing pipeline? We present Task-Aware Mechanism (TAM), a hybrid-gated Mixture-of-Experts (MoE) vision tower that adapts frame count and resolution …

Cited by 0SourceScholar
2026

The Avengers: A Routing Recipe for Collective Intelligence in Language Models

AAAI 2026technical

Proprietary models are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this paper, we present the Avengers---a lightweight framework that leverages the collective intelligence of these smaller mod

Cited by 0SourcePDFScholar
2026

Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated Planning

ICML 2026poster

The integration of external tools has substantially expanded the capabilities of large language model (LLM) agents, but also introduced new attack surfaces beyond prompt injection. In particular, cross-tool description poisoning can manipulate planner-visible tool metadata to steer an agent’s trajec…

Cited by 0SourceScholar
2026

Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling

AAAI 2026technical

Recent vision-language-action (VLA) models built on pretrained vision-language models (VLMs) have demonstrated strong performance in robotic manipulation. However, these models remain constrained by the single-frame image paradigm and fail to fully leverage the temporal information offered by multi-

Cited by 0SourcePDFScholar
2026

UMI-Underwater: Learning Underwater Manipulation without Underwater Teleoperation

RSS 2026poster

Underwater robotic grasping is difficult due to degraded, highly variable imagery and the expense of collecting diverse underwater demonstrations. We introduce a system that (i) autonomously collects successful underwater grasp demonstrations via a self-supervised data collection pipeline and (ii) t…

Cited by 0SourceScholar
2026

Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vision

ICLR 2026poster

Chain-of-Thought (CoT) reasoning has proven effective in enhancing Large Language Models (LLMs) on complex tasks by decomposing problems into step-wise solutions. However, extending CoT to multi-modal settings remains challenging, as it requires modeling transitions of visual states alongside textua…

Cited by 0SourcecodeScholar
2026

VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network

CVPR 2026

The rapid advances in deep learning have significantly enhanced the accuracy of multimodal 3D human pose estimation (HPE). However, the state-of-the-art (SOTA) HPE pipelines still rely on Transformers, whose quadratic complexity makes real-time processing for long sequences impractical. Mamba addres

Cited by 0SourcecodeScholar
2026

Vision-Language-Action Instruction Tuning: From Understanding to Manipulation

ICLR 2026poster

To operate effectively in the real world, robots should integrate multimodal reasoning with precise action generation. However, existing vision-language-action (VLA) models often sacrifice one for the other, narrow their abilities to task-specific manipulation data, and suffer catastrophic forgettin…

Cited by 0SourcecodeScholar
2025

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models

ACL 2025finding

Existing multi-objective preference alignment methods for large language models (LLMs) face limitations: (1) the inability to effectively balance various preference dimensions, and (2) reliance on auxiliary reward/reference models introduces computational complexity. To address these challenges, we…

2025

AU-Blendshape for Fine-grained Stylized 3D Facial Expression Manipulation

ICCV 2025poster

While 3D facial animation has made impressive progress, challenges still exist in realizing fine-grained stylized 3D facial expression manipulation due to the lack of appropriate datasets. In this paper, we introduce the AUBlendSet, a 3D facial dataset based on AU-Blendshape representation for fine-…

2025

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding

CVPR 2025poster

Visual Document Understanding has become essential with the increase of text-rich visual content. This field poses significant challenges due to the need for effective integration of visual perception and textual comprehension, particularly across diverse document types with complex layouts. Moreove…

2025

AdvDisplay: Adversarial Display Assembled by Thermoelectric Cooler for Fooling Thermal Infrared Detectors

AAAI 2025technical

When the current physical adversarial patches cannot deceive thermal infrared detectors, the existing techniques implement adversarial attacks from scratch, such as digital patch generation, material production, and physical deployment. Besides, it is difficult to finely regulate infrared radiation.…

2025

BRIDGE: Bootstrapping Text to Control Time-Series Generation via Multi-Agent Iterative Optimization and Diffusion Modeling

ICML 2025poster

Time-series Generation (TSG) is a prominent research area with broad applications in simulations, data augmentation, and counterfactual analysis. While existing methods have shown promise in unconditional single-domain TSG, real-world applications demand for cross-domain approaches capable of contro…

Cited by 0SourcePDFScholar
2025

Breaking the Reasoning Barrier A Survey on LLM Complex Reasoning through the Lens of Self-Evolution

ACL 2025finding

The release of OpenAI’s O1 and subsequent projects like DeepSeek R1 has significantly advanced research on complex reasoning in LLMs. This paper systematically analyzes existing reasoning studies from the perspective of self-evolution, structured into three components: data evolution, model evolutio…

Cited by 0SourcePDFScholar
2025

CCIN: Compositional Conflict Identification and Neutralization for Composed Image Retrieval

CVPR 2025highlight

Composed Image Retrieval (CIR) is a multi-modal task that seeks to retrieve target images by harmonizing a reference image with a modified instruction. A key challenge in CIR lies in compositional conflicts between the reference image (e.g., blue, long sleeve) and the modified instruction (e.g., gre…

2025

CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction

ICCV 2025poster

Despite its significant achievements in large-scale scene reconstruction, 3D Gaussian Splatting still faces substantial challenges, including slow processing, high computational costs, and limited geometric accuracy. These core issues arise from its inherently unstructured design and the absence of…

Cited by 0SourcePDFScholar
2025

Convex Relaxation for Robust Vanishing Point Estimation in Manhattan World

CVPR 2025award

Determining the vanishing points (VPs) in a Manhattan world, as a fundamental task in many 3D vision applications, consists of jointly inferring the line-VP association and locating each VP. Existing methods are, however, either sub-optimal solvers or pursuing global optimality at a significant cost…

2025

Cross-Category Subjectivity Generalization for Style-Adaptive Sketch Re-ID

ICCV 2025poster

Sketch-based person re-identification (re-ID) enables pedestrian retrieval using sketches. While recent methods have improved modality alignment between sketches and RGB images, the challenge of subjective style variation, where sketches exhibit diverse and unpredictable appearances, remains largely…

Cited by 0SourcePDFScholar
2025

DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes

ICRA 2025

Novel-view synthesis approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, making them unsuitable where computational resources are limited. Additionally, few-shot methods often struggle with poor reconst

Cited by 5SourcecodeScholar
2025

DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation

ICCV 2025poster

Human-centric generative models are becoming increasingly popular, giving rise to various innovative tools and applications, such as talking face videos conditioned on text or audio prompts. The core of these capabilities lies in powerful pre-trained foundation models, trained on large-scale, high-q…

2025

DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents

NeurIPS 2025poster

Large Language Models (LLMs) are increasingly central to agentic systems due to their strong reasoning and planning capabilities. By interacting with external environments through predefined tools, these agents can carry out complex user tasks. Nonetheless, this interaction also introduces the risk…

Cited by 0SourcecodeScholar
2025

Deconfound Semantic Shift and Incompleteness in Incremental Few-shot Semantic Segmentation

AAAI 2025technical

Incremental few-shot semantic segmentation (IFSS) expands segmentation capacity of the trained model to segment new-class images with few samples. However, semantic meanings may shift from background to object class or vice versa during incremental learning. Moreover, new-class samples often lack re…

Cited by 0SourcePDFScholar
2025

DiffPortrait360: Consistent Portrait Diffusion for 360 View Synthesis

CVPR 2025poster

Generating high-quality 360-degree views of human heads from single-view images is essential for enabling accessible immersive telepresence applications and scalable personalized content creation.While cutting-edge methods for full head generation are limited to modeling realistic human heads, the l…

2025

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak

EMNLP 2025

Large Language Models (LLMs) are susceptible to generating harmful content when prompted with carefully crafted inputs, a vulnerability known as LLM jailbreaking. As LLMs become more powerful, studying jailbreak methods is critical to enhancing security and aligning models with human values. Traditi

Cited by 0SourcePDFScholar
2025

Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study

EMNLP 2025

Vision-Language Models (VLMs) are powerful yet computationally intensive for widespread practical deployments. To address such challenge without costly re-training, post-training acceleration techniques like quantization and token reduction are extensively explored. However, current acceleration eva

Cited by 0SourcePDFScholar
2025

Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing

NeurIPS 2025oral

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this…

Cited by 0SourcecodeScholar
2025

EverybodyDance: Bipartite Graph–Based Identity Correspondence for Multi-Character Animation

NeurIPS 2025poster

Consistent pose‐driven character animation has achieved remarkable progress in single‐character scenarios. However, extending these advances to multi‐character settings is non‐trivial, especially when position swap is involved. Beyond mere scaling, the core challenge lies in enforcing correct Identi…

Cited by 0SourceScholar
2025

FoundIR: Unleashing Million-scale Training Data to Advance Foundation Models for Image Restoration

ICCV 2025poster

Despite the significant progress made by all-in-one models in universal image restoration, existing methods suffer from a generalization bottleneck in real-world scenarios, as they are mostly trained on small-scale synthetic datasets with limited degradations. Therefore, large-scale high-quality rea…

Cited by 0SourcePDFScholar
2025

Frequency-Domain Guided Multiple Parallel Kernels Network for Low-Light Remote Sensing Image Enhancement

ICASSP 2025accepted

Due to dark environments, optical aberrations, etc, the remote sensing images are often submerged under low contrast degradation, which greatly hinders their practical applications for agricultural management and other related tasks. The surface features of remote sensing images are often continuous…

Cited by 0SourceScholar
2025

From Monocular Vision to Autonomous Action: Guiding Tumor Resection via 3D Reconstruction

IROS 2025

Surgical automation requires precise guidance and understanding of the scene. Current methods in the literature rely on bulky depth cameras to create maps of the anatomy; however, this does not translate well to space-limited clinical applications. Monocular cameras are small and allow minimally inv

Cited by 6SourceScholar
2025

FuXi-Ocean: A Global Ocean Forecasting System with Sub-Daily Resolution

NeurIPS 2025oral

Accurate, high-resolution ocean forecasting is crucial for maritime operations and environmental monitoring. While traditional numerical models are capable of producing sub-daily, eddy-resolving forecasts, they are computationally intensive and face challenges in maintaining accuracy at fine spatial…

Cited by 0SourceScholar
2025

FuXi-RTM: A Physics-Guided Prediction Framework with Radiative Transfer Modeling

ICCV 2025poster

Similar to conventional video generation, current deep learning-based weather prediction frameworks often lack explicit physical constraints, leading to unphysical outputs that limit their reliability for operational forecasting. Among various physical processes requiring proper representation, radi…

Cited by 0SourcePDFScholar
2025

GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation

CVPR 2025poster

Robotic manipulation in real-world settings remains challenging, especially regarding robust generalization. Existing simulation platforms lack sufficient support for exploring how policies adapt to varied instructions and scenarios. Thus, they lag behind the growing interest in instruction-followin…

Cited by 0SourcePDFScholar
2025

GIFStream: 4D Gaussian-based Immersive Video with Feature Stream

CVPR 2025poster

Immersive video offers a 6-Dof-free viewing experience, potentially playing a key role in future video technology. Recently, 4D Gaussian Splatting has gained attention as an effective approach for immersive video due to its high rendering efficiency and quality, though maintaining quality with manag…

Cited by 0SourcePDFScholar
2025

GRPose: Learning Graph Relations for Human Image Generation with Pose Priors

AAAI 2025technical

Recent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this…

2025

GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing

NeurIPS 2025poster

Current image generation and editing methods primarily process textual prompts as direct inputs without explicit reasoning about visual composition or operational steps. We present Generation Chain-of-Thought (GoT), a novel paradigm that empowers a Multimodal Large Language Model (MLLM) to first gen…

Cited by 0SourceScholar
2025

Hello Again! LLM-powered Personalized Agent for Long-term Dialogue

NAACL 2025long

Open-domain dialogue systems have seen remarkable advancements with the development of large language models (LLMs). Nonetheless, most existing dialogue systems predominantly focus on brief single-session interactions, neglecting the real-world demands for long-term companionship and personalized in…

2025

HieraFashDiff: Hierarchical Fashion Design with Multi-stage Diffusion Models

AAAI 2025technical

Fashion design is a challenging and complex process. Recent works on fashion generation and editing are all agnostic of the actual fashion design process, which limits their usage in practice. In this paper, we propose a novel hierarchical diffusion-based framework tailored for fashion design, coine…

2025

LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts

ACL 2025long

Safety concerns in large language models (LLMs) have gained significant attention due to their exposure to potentially harmful data during pre-training. In this paper, we identify a new safety vulnerability in LLMs: their susceptibility to natural distribution shifts between attack prompts and origi…

2025

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

ICCV 2025poster

Large language models have demonstrated substantial advancements in reasoning capabilities. However, current Vision-Language Models (VLMs) often struggle to perform systematic and structured reasoning, especially when handling complex visual question-answering tasks. In this work, we introduce LLaVA…

2025

LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models

NAACL 2025findings

Multi-modal Large Language Models (MLLMs) have achieved remarkable success by integrating visual and textual modalities. However, they incur significant computational overhead due to the large number of vision tokens processed, limiting their practicality in resource-constrained environments. We int…

Cited by 2SourcePDFScholar
2025

LangBridge: Interpreting Image as a Combination of Language Embeddings

ICCV 2025poster

Recent years have witnessed remarkable advances in Large Vision-Language Models (LVLMs), which have achieved human-level performance across various complex vision-language tasks. Following LLaVA's paradigm, mainstream LVLMs typically employ a shallow MLP for visual-language alignment through a two-s…

2025

LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion

ICCV 2025poster

Recovering 3D structures with open-vocabulary scene understanding from 2D images is a fundamental but daunting task. Recent developments have achieved this by performing per-scene optimization with embedded language information. However, they heavily rely on the calibrated dense-view reconstruction…

Cited by 0SourcePDFScholar
2025

Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment

EMNLP 2025

With rapid advancement and increasing accessibility of LLMs, fine-tuning aligned models has become a critical step for adapting them to real-world applications, which makes the safety of this fine-tuning process more important than ever. However, recent studies have highlighted a critical challenge:

2025

Learning Crossmodal Interaction Patterns via Attributed Bipartite Graphs for Single-Cell Omics

NeurIPS 2025poster

Crossmodal matching in single-cell omics is essential for explaining biological regulatory mechanisms and enhancing downstream analyses. However, current single-cell crossmodal models often suffer from three limitations: sparse modality signals, underutilization of biological attributes, and insuffi…

Cited by 0SourcecodeScholar
2025

MIRA: Medical Time Series Foundation Model for Real-World Health Data

NeurIPS 2025poster

A unified foundation model for medical time series—pretrained on open access and ethically reviewed medical corpora—offers the potential to reduce annotation burdens, minimize model customization, and enable robust transfer across clinical institutions, modalities, and tasks, particularly in data-sc…

Cited by 0SourceScholar
2025

MUCD: Unsupervised Point Cloud Change Detection via Masked Consistency

AAAI 2025technical

3D Change Detection (3DCD) has gradually become another research hotspot after image change detection. Recent works focus on using artificial labels for supervised or weakly-supervised training of siamese networks to segment changed points. However, labeling every points of multi-temporal point clou…

Cited by 0SourcePDFScholar
2025

NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints

NeurIPS 2025poster

Compositional training has been the de-facto paradigm in existing Multimodal Large Language Models (MLLMs), where pre-trained vision encoders are connected with pre-trained LLMs through continuous multimodal pre-training. However, the multimodal scaling property of this paradigm remains difficult…

Cited by 0SourceScholar
2025

Omni-Mol: Multitask Molecular Model for Any-to-any Modalities

NeurIPS 2025poster

In the molecular domain, numerous studies have explored the use of multimodal large language models (LLMs) to construct a general-purpose, multi-task molecular model. However, these efforts are still far from achieving a truly universal molecular model. We identify three key challenges in this endea…

Cited by 0SourceScholar
2025

PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for Free

ACL 2025long

Prompt injection attacks pose a critical threat to large language models (LLMs), enabling goal hijacking and data leakage. Prompt guard models, though effective in defense, suffer from over-defense—falsely flagging benign inputs as malicious due to trigger word bias. To address this issue, we introd…

2025

PUMA: Empowering Unified MLLM with Multi-granular Visual Generation

ICCV 2025poster

Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models for visual content generation. However, existing approaches face a trade-off between generatio…

2025

Partial Point Cloud Registration with Multi-view 2D Image Learning

AAAI 2025technical

Learning representations from numerous 2D image data has shown promising performance, yet very few works apply this representations to point cloud registration. In this paper, we explore how to leverage the 2D information to assist the point cloud registration, and propose IAPReg, an Image-Assisted…

Cited by 0SourcePDFScholar
2025

Pioneer: Physics-informed Riemannian Graph ODE for Entropy-increasing Dynamics

AAAI 2025technical

Dynamic interacting system modeling is important for understanding and simulating real world systems, e.g., meteorology and the spread of COVID. The system is typically described as a graph, where multiple objects dynamically interact with each other and evolve over time. In recent years, graph Ordi…

2025

PointTruss: K-Truss for Point Cloud Registration

NeurIPS 2025poster

Point cloud registration is a fundamental task in 3D computer vision. Recent advances have shown that graph-based methods are effective for outlier rejection in this context. However, existing clique-based methods impose overly strict constraints and are NP-hard, making it difficult to achieve both…

Cited by 0SourceScholar
2025

Political Actor Agent: Simulating Legislative System for Roll Call Votes Prediction with Large Language Models

AAAI 2025technical

Predicting roll call votes through modeling political actors has emerged as a focus in quantitative political science and computer science. Widely used embedding-based methods generate vectors for legislators from diverse data sets to predict legislative behaviors. However, these methods often conte…

Cited by 1SourcePDFScholar
2025

QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized Generation

ICCV 2025poster

Existing text-to-image models often rely on parame- ter fine-tuning techniques such as Low-Rank Adaptation (LoRA) to customize visual attributes. However, when com- bining multiple LoRA models for content-style fusion tasks, unstructured modifications of weight matrices often lead to undesired featu…

Cited by 0SourcePDFScholar
2025

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

CVPR 2025poster

Recent advancements in robotic manipulation have highlighted the potential of intermediate representations for improving policy generalization. In this work, we explore grounding masks as an effective intermediate representation, balancing two key advantages: (1) effective spatial guidance that spec…

Cited by 0SourcePDFScholar
2025

Robust Stabilization of an Autonomous Underwater Vehicle in Specified Finite-time with Disturbance Rejection

IROS 2025

This study investigates the robust finite-time stabilization of an autonomous underwater vehicle (AUV) with disturbance rejection, where the finite-time can be predetermined. The AUV is modeled as a rigid body moving within fluids, and the systems dynamics involves uncertain parameters arising from

Cited by 0SourceScholar
2025

SLIM: A Symmetric, Low-Inertia Manipulator for Constrained, Contact-Rich Spaces

RA-L 2025

Operation in constrained and cluttered spaces poses a challenge for robotic manipulators, in part due to their bulky link geometry and kinematic limitations in comparison to human hands and arms. To address these limitations, we introduce SLIM, a custom end-effector consisting of a bidirectional han

Cited by 0SourceScholar
2025

STAR: Learning Diverse Robot Skill Abstractions through Rotation-Augmented Vector Quantization

ICML 2025spotlight

Transforming complex actions into discrete skill abstractions has demonstrated strong potential for robotic manipulation.Existing approaches mainly leverage latent variable models, e.g., VQ-VAE, to learn skill abstractions through learned vectors (codebooks), while they suffer from codebook collapse…

2025

STRIDER: Navigation via Instruction-Aligned Structural Decision Space Optimization

NeurIPS 2025poster

The Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) task requires agents to navigate previously unseen 3D environments using natural language instructions, without any scene-specific training. A critical challenge in this setting lies in ensuring agents’ actions align wi…

Cited by 0SourceScholar
2025

Self-Critique Guided Iterative Reasoning for Multi-hop Question Answering

ACL 2025finding

Although large language models (LLMs) have demonstrated remarkable reasoning capabilities, they still face challenges in knowledge-intensive multi-hop reasoning. Recent work explores iterative retrieval to address complex problems. However, the absence of intermediate guidance often leads to inaccur…

2025

Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation

CVPR 2025poster

Despite the significant success of imitation learning in robotic manipulation, its application to bimanual tasks remains highly challenging. Existing approaches mainly learn a policy to predict a distant next-best end-effector pose (NBP) and then compute the corresponding joint rotation angles for m…

Cited by 2SourcePDFScholar
2025

Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency

ICASSP 2025accepted

Connectionist Temporal Classification (CTC), a non-autoregressive training criterion, is widely used in online keyword spotting (KWS). However, existing CTC-based KWS decoding strategies either rely on Automatic Speech Recognition (ASR), which performs suboptimally due to its broad search over the a…

Cited by 0SourceScholar
2025

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

CVPR 2025poster

The remarkable success of Large Language Models (LLMs) has extended to the multimodal domain, achieving outstanding performance in image understanding and generation. Recent efforts to develop unified Multimodal Large Language Models (MLLMs) that integrate these capabilities have shown promising res…

2025

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

NeurIPS 2025poster

Recent advancements in large language models have demonstrated how chain-of-thought (CoT) and reinforcement learning (RL) can improve performance. However, applying such reasoning strategies to the visual generation domain remains largely unexplored. In this paper, we present **T2I-R1**, a novel rea…

Cited by 0SourcecodeScholar
2025

TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement

NAACL 2025findings

Large Language Models (LLMs) have achieved impressive results in Machine Translation (MT). However, human evaluations reveal that LLM-generated translations still contain various errors. Notably, feeding the error information back into the LLMs can facilitate self-refinement, leading to enhanced tra…

2025

TMetaNet: Topological Meta-Learning Framework for Dynamic Link Prediction

ICML 2025poster

Dynamic graphs evolve continuously, presenting challenges for traditional graph learning due to their changing structures and temporal dependencies. Recent advancements have shown potential in addressing these challenges by developing suitable meta-learning-based dynamic graph neural network models.…

2025

TacCap: A Wearable FBG-Based Tactile Sensor for Efficient Human-to-Robot Skill Transfer

IROS 2025

Tactile sensing is essential for dexterous manipulation, yet large-scale human demonstration datasets lack tactile feedback, limiting their effectiveness in skill transfer to robots. To address this, we introduce TacCap, a wearable Fiber Bragg Grating (FBG)-based tactile sensor designed for seamless

Cited by 0SourceScholar
2025

Test-Time Code-Switching for Cross-lingual Aspect Sentiment Triplet Extraction

NAACL 2025long

Aspect Sentiment Triplet Extraction (ASTE) is a thriving research area with impressive outcomes being achieved on high-resource languages. However, the application of cross-lingual transfer to the ASTE task has been relatively unexplored, and current code-switching methods still suffer from term bou…

Cited by 0SourcePDFScholar
2025

TypeTele: Releasing Dexterity in Teleoperation by Dexterous Manipulation Types

CoRL 2025poster

Dexterous teleoperation plays a crucial role in robotic manipulation for real-world data collection and remote robot control. Previous dexterous teleoperation mostly relies on hand retargeting to closely mimic human hand postures. However, these approaches may fail to fully leverage the inherent dex…

Cited by 0SourceScholar
2025

UDSH: An Unsupervised Deep Image Stitching and De-Occlusion Method for Heavy Occlusion Scene

IROS 2025

Image stitching in heavy occlusion scenarios faces the dual challenges of accurate alignment and occlusion removal. On one hand, occlusion causes the loss of key texture and structural information in the image. On the other hand, it affects the image’s integrity. Existing stitching methods perform w

Cited by 0SourceScholar
2025

Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference

NeurIPS 2025oral

Large Language Models (LLMs) are now integral across various domains and have demonstrated impressive performance. Progress, however, rests on the premise that benchmark scores are both accurate and reproducible. We demonstrate that the reproducibility of LLM performance is fragile: changing system…

Cited by 0SourcecodeScholar
2025

VDG: Vision-Only Dynamic Gaussian for Driving Simulation

RA-L 2025

Recent advances in dynamic Gaussian splatting have significantly improved scene reconstruction and novel-view synthesis. However, existing methods often rely on pre-computed camera poses and Gaussian initialization using Structure from Motion (SfM) or other costly sensors, limiting their scalability

Cited by 23SourceScholar
2025

VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence

AAAI 2025technical

Social Intelligence Queries (Social-IQ) serve as the primary multimodal benchmark for evaluating a model’s social intelligence level. While impressive multiple-choice question (MCQ) accuracy is achieved by current solutions, increasing evidence shows that they are largely, and in some cases entire…

2025

VLSBench: Unveiling Visual Leakage in Multimodal Safety

ACL 2025long

Safety concerns of Multimodal large language models (MLLMs) have gradually become an important problem in various applications. Surprisingly, previous works indicate a counterintuitive phenomenon that using textual unlearning to align MLLMs achieves comparable safety performances with MLLMs aligned…

2025

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation

CVPR 2025poster

In 3D speech-driven facial animation generation, existing methods commonly employ pre-trained self-supervised audio models as encoders. However, due to the prevalence of phonetically similar syllables with distinct lip shapes in language, these near-homophone syllables tend to exhibit significant co…

2025

Wave-wise Discriminative Tracking by Phase-Amplitude Separation, Augmentation and Mixture

IJCAI 2025

Distinguishing key features in complex visual tasks is challenging. A novel approach treats image patches (tokens) as waves. By using both phase and amplitude, it captures richer semantics and specific invariances compared to pixel-based methods, and allows for feature fusion across regions for a ho

Cited by 0SourcePDFScholar
2025

When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning

EMNLP 2025

Large reasoning models (LRMs) achieve remarkable performance via long reasoning chains, but often incur excessive computational overhead due to redundant reasoning, especially on simple tasks. In this work, we systematically quantify the upper bounds of LRMs under both Long-Thinking and No-Thinking

Cited by 0SourcePDFScholar
2025

Whisker-Inspired Tactile Sensing: A Sim2Real Approach for Precise Underwater Contact Tracking

RA-L 2025

Aquatic mammals use whiskers to detect and discriminate objects and analyze water movements, inspiring the development of robotic whiskers for sensing contacts, surfaces, and water flows. We present the design and application of underwater whisker sensors based on Fiber Bragg Grating (FBG) technolog

Cited by 5SourceScholar
2024

3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications

ICASSP 2024accepted

Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-ti…

Cited by 0SourceScholar
2024

ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process

ICLR 2024poster

Image recognition and generation have long been developed independently of each other. With the recent trend towards general-purpose representation learning, the development of general representations for both recognition and generation tasks is also promoted. However, preliminary attempts mainly fo…

2024

ADVSV: An Over-the-Air Adversarial Attack Dataset for Speaker Verification

ICASSP 2024accepted

It is known that deep neural networks are vulnerable to adversarial attacks. Although Automatic Speaker Verification (ASV) built on top of deep neural networks exhibits robust performance in controlled scenarios, many studies confirm that ASV is vulnerable to adversarial attacks. The lack of a stand…

Cited by 0SourceScholar
2024

ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings

EMNLP 2024main

The safety defense methods of Large language models (LLMs) stays limited because the dangerous prompts are manually curated to just few known attack types, which fails to keep pace with emerging varieties. Recent studies found that attaching suffixes to harmful instructions can hack the defense of L…

Cited by 9SourcePDFScholar
2024

Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft

CVPR 2024poster

Many reinforcement learning environments (e.g. Minecraft) provide only sparse rewards that indicate task completion or failure with binary values. The challenge in exploration efficiency in such environments makes it difficult for reinforcement-learning-based agents to learn complex tasks. To addres…

Cited by 38SourcePDFScholar
2024

CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models

ACL 2024findings

The advancement of large language models (LLMs) has enhanced the ability to generalize across a wide range of unseen natural language processing (NLP) tasks through instruction-following.Yet, their effectiveness often diminishes in low-resource languages like Chinese, exacerbated by biased evaluatio…

2024

Clustering then Propagation: Select Better Anchors for Knowledge Graph Embedding

NeurIPS 2024poster

Traditional knowledge graph embedding (KGE) models map entities and relations to unique embedding vectors in a shallow lookup manner. As the scale of data becomes larger, this manner will raise unaffordable computational costs. Anchor-based strategies have been treated as effective ways to alleviate…

Cited by 0SourcePDFScholar
2024

Contrastive Learning with Audio Discrimination for Customizable Keyword Spotting in Continuous Speech

ICASSP 2024accepted

Customizable keyword spotting (KWS) in continuous speech has attracted increasing attention due to its real-world application potential. While contrastive learning (CL) has been widely used to extract keyword representations, previous CL approaches all operate on pre-segmented isolated words and emp…

Cited by 0SourceScholar
2024

DEIE: Benchmarking Document-level Event Information Extraction with a Large-scale Chinese News Dataset

COLING 2024main

A text corpus centered on events is foundational to research concerning the detection, representation, reasoning, and harnessing of online events. The majority of current event-based datasets mainly target sentence-level tasks, thus to advance event-related research spanning from sentence to documen…

2024

Diffusion-based Blind Text Image Super-Resolution

CVPR 2024poster

Recovering degraded low-resolution text images is challenging especially for Chinese text images with complex strokes and severe degradation in real-world scenarios. Ensuring both text fidelity and style realness is crucial for high-quality text image super-resolution. Recently diffusion models have…

2024

Enhancing Realism in 3D Facial Animation Using Conformer-Based Generation and Automated Post-Processing

ICASSP 2024accepted

Recent progress has propelled the development of realistic talking-face videos for avatars. Yet, animating 3D cartoon avatars remains intricate due to the imprecise nature of facial-driven data. This often manifests as inconsistent mouth configurations and rigid facial expressions, curbing the anima…

Cited by 0SourceScholar
2024

Expressiveness is Effectiveness: Self-supervised Fashion-aware CLIP for Video-to-Shop Retrieval

IJCAI 2024poster

The rise of online shopping and social media has spurred the Video-to-Shop Retrieval (VSR) task, which involves identifying fashion items (e.g., clothing) in videos and matching them with identical products provided by stores. In real-world scenarios, human movement in dynamic video scenes can cause…

Cited by 1SourcePDFScholar
2024

GGRt: Towards Generalizable 3D Gaussians without Pose Priors in Real-Time

ECCV 2024poster

"This paper presents GGRt, a novel approach to generalizable novel view synthesis that alleviates the need for real camera poses, complexity in processing high-resolution images, and lengthy optimization processes, thus facilitating stronger applicability of 3D Gaussian Splatting (3D-GS) in real-wor…

2024

GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding

CVPR 2024highlight

Applying Neural Radiance Fields (NeRF) to downstream perception tasks for scene understanding and representation is becoming increasingly popular. Most existing methods treat semantic prediction as an additional rendering task i.e. the "label rendering" task to build semantic NeRFs. However by rende…

Cited by 25SourcePDFScholar
2024

Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction

CVPR 2024poster

Whole Slide Image (WSI) classification is often formulated as a Multiple Instance Learning (MIL) problem. Recently Vision-Language Models (VLMs) have demonstrated remarkable performance in WSI classification. However existing methods leverage coarse-grained pathogenetic descriptions for visual repre…

2024

Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing

AAAI 2024technical

GAN-based image attribute editing firstly leverages GAN Inversion to project real images into the latent space of GAN and then manipulates corresponding latent codes. Recent inversion methods mainly utilize additional high-bit features to improve image details preservation, as low-bit codes cannot f…

Cited by 3SourcePDFScholar
2024

Grasp as You Say: Language-guided Dexterous Grasp Generation

NeurIPS 2024poster

This paper explores a novel task "Dexterous Grasp as You Say'' (DexGYS), enabling robots to perform dexterous grasping based on human commands expressed in natural language. However, the development of this field is hindered by the lack of datasets with natural human guidance; thus, we propose a lan…

2024

High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models

ICASSP 2024accepted

Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech representations(semantic & acoustic) and using two sequence-to-seque…

Cited by 0SourceScholar
2024

LTGC: Long-tail Recognition via Leveraging LLMs-driven Generated Content

CVPR 2024poster

Long-tail recognition is challenging because it requires the model to learn good representations from tail categories and address imbalances across all categories. In this paper we propose a novel generative and fine-tuning framework LTGC to handle long-tail recognition via leveraging generated cont…

Cited by 16SourcePDFScholar
2024

Learning Speech Representation from Contrastive Token-Acoustic Pretraining

ICASSP 2024accepted

For fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate representations extracted from speech should serve as a "bridge" between text and acoustic information, containing info…

Cited by 0SourceScholar
2024

Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation

ECCV 2024poster

"Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily focus on the direct synthesis of global motions while negle…

2024

MICRO: Model-Based Offline Reinforcement Learning with a Conservative Bellman Operator

IJCAI 2024poster

Offline reinforcement learning (RL) faces a significant challenge of distribution shift. Model-free offline RL penalizes the Q value for out-of-distribution (OOD) data or constrains the policy closed to the behavior policy to tackle this problem, but this inhibits the exploration of the OOD region.…

2024

Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding

ICASSP 2024accepted

Recently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations and using two sequence-to-sequence tasks to decouple TTS. However, existing methods suffer from three problems: the high-…

Cited by 0SourceScholar
2024

NeRFCodec: Neural Feature Compression Meets Neural Radiance Fields for Memory-Efficient Scene Representation

CVPR 2024poster

The emergence of Neural Radiance Fields (NeRF) has greatly impacted 3D scene modeling and novel-view synthesis. As a kind of visual media for 3D scene representation compression with high rate-distortion performance is an eternal target. Motivated by advances in neural compression and neural field r…

Cited by 11SourcePDFScholar
2024

On the Scalability of Diffusion-based Text-to-Image Generation

CVPR 2024poster

Scaling up model and data size has been quite successful for the evolution of LLMs. However the scaling law for the diffusion based text-to-image (T2I) models is not fully explored. It is also unclear how to efficiently scale the model for better performance at reduced cost. The different training s…

Cited by 22SourcePDFScholar
2024

Parameter-Inverted Image Pyramid Networks

NeurIPS 2024spotlight

Image pyramids are commonly used in modern computer vision tasks to obtain multi-scale features for precise understanding of images. However, image pyramids process multiple resolutions of images using the same large-scale model, which requires significant computational cost. To overcome this issue,…

2024

PointMC: Multi-instance Point Cloud Registration based on Maximal Cliques

ICML 2024poster

Multi-instance point cloud registration is the problem of estimating multiple rigid transformations between two point clouds. Existing solutions rely on global spatial consistency of ambiguity and the time-consuming clustering of highdimensional correspondence features, making it difficult to handle…

Cited by 1SourcePDFScholar
2024

RoboMP$^2$: A Robotic Multimodal Perception-Planning Framework with Multimodal Large Language Models

ICML 2024poster

Multimodal Large Language Models (MLLMs) have shown impressive reasoning abilities and general intelligence in various domains. It inspires researchers to train end-to-end MLLMs or utilize large models to generate policies with human-selected prompts for embodied agents. However, these methods exhib…

Cited by 2SourcePDFScholar
2024

Robustly Train Normalizing Flows via KL Divergence Regularization

AAAI 2024technical

In this paper, we find that the training of Normalizing Flows (NFs) are easily affected by the outliers and a small number (or high dimensionality) of training samples. To solve this problem, we propose a Kullback–Leibler (KL) divergence regularization on the Jacobian matrix of NFs. We prove that su…

Cited by 2SourcePDFScholar
2024

Sorting, Reasoning, and Extraction: An Easy-to-Hard Reasoning Framework for Document-Level Event Argument Extraction

ICASSP 2024accepted

Document-level event argument extraction is a crucial task to help understand event information. Existing methods mostly ignore the different extraction difficulties of arguments, and the lack of task planning significantly affects the extraction and reasoning abilities of the model. In this paper,…

Cited by 0SourceScholar
2024

TDT-KWS: Fast and Accurate Keyword Spotting Using Token-and-Duration Transducer

ICASSP 2024accepted

Designing an efficient keyword spotting (KWS) system that delivers exceptional performance on resource-constrained edge devices has long been a subject of significant attention. Existing KWS search algorithms typically follow a frame-synchronous approach, where search decisions are made repeatedly a…

Cited by 0SourceScholar
2024

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

ICLR 2024poster

We present the All-Seeing (AS) project: a large-scale dataset and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1.2 billion regions annota…

2024

Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition

AAAI 2024technical

Multimodal intent recognition aims to leverage diverse modalities such as expressions, body movements and tone of speech to comprehend user's intent, constituting a critical task for understanding human language and behavior in real-world multimodal scenarios. Nevertheless, the majority of existing…

2024

Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image Generation

CVPR 2024poster

Vanilla text-to-image diffusion models struggle with generating accurate human images commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limbs. Existing methods address this issue mostly by fine-tuning the model with extra images or adding additional controls --…

Cited by 9SourcePDFScholar
2024

VOODOO 3D: Volumetric Portrait Disentanglement For One-Shot 3D Head Reenactment

CVPR 2024poster

We present a 3D-aware one-shot head reenactment method based on a fully volumetric neural disentanglement framework for source appearance and driver expressions. Our method is real-time and produces high-fidelity and view-consistent output suitable for 3D teleconferencing systems based on holographi…

Cited by 13SourcePDFScholar
2024

Which Side Are You On? A Multi-task Dataset for End-to-End Argument Summarisation and Evaluation

ACL 2024findings

With the recent advances of large language models (LLMs), it is no longer infeasible to build an automated debate system that helps people to synthesise persuasive arguments. Previous work attempted this task by integrating multiple components. In our work, we introduce an argument mining dataset th…

2023

Adaptive Symmetry Reference Trajectory Generation in Shared Autonomy for Active Knee Orthosis

RA-L 2023

Gait symmetry training plays an essential role in the rehabilitation of hemiplegic patients. Robotics-based gait training has been widely accepted by patients and clinicians. Reference trajectory generation for the affected side using the motion data of the unaffected side is an important way to ach

Cited by 3SourceScholar
2023

Boosting Low-Data Instance Segmentation by Unsupervised Pre-Training With Saliency Prompt

CVPR 2023poster

Recently, inspired by DETR variants, query-based end-to-end instance segmentation (QEIS) methods have outperformed CNN-based models on large-scale datasets. Yet they would lose efficacy when only a small amount of training data is available since it's hard for the crucial queries/kernels to learn lo…

2023

Clusterformer: Cluster-based Transformer for 3D Object Detection in Point Clouds

ICCV 2023poster

Attributed to the unstructured and sparse nature of point clouds, the transformer shows greater potential in point clouds data processing. However, the recent query-based 3D detectors usually project the features acquired from a sparse backbone into the structured and compact Bird's Eye View(BEV) pl…

Cited by 15PDFScholar
2023

DiffusionRet: Generative Text-Video Retrieval with Diffusion Model

ICCV 2023poster

Existing text-video retrieval solutions are, in essence, discriminant models focused on maximizing the conditional likelihood, i.e., p(candidates|query). While straightforward, this de facto paradigm overlooks the underlying data distribution p(query), which makes it challenging to identify out-of-d…

Cited by 72PDFcodeScholar
2023

Do You Hear The People Sing? Key Point Analysis via Iterative Clustering and Abstractive Summarisation

ACL 2023long

Argument summarisation is a promising but currently under-explored field. Recent work has aimed to provide textual summaries in the form of concise and salient short texts, i.e., key points (KPs), in a task known as Key Point Analysis (KPA). One of the main challenges in KPA is finding high-quality…

2023

Guided Recommendation for Model Fine-Tuning

CVPR 2023poster

Model selection is essential for reducing the search cost of the best pre-trained model over a large-scale model zoo for a downstream task. After analyzing recent hand-designed model selection criteria with 400+ ImageNet pre-trained models and 40 downstream tasks, we find that they can fail due to i…

2023

Hybrid CNN-Transformer Feature Fusion for Single Image Deraining

AAAI 2023technical

Since rain streaks exhibit diverse geometric appearances and irregular overlapped phenomena, these complex characteristics challenge the design of an effective single image deraining model. To this end, rich local-global information representations are increasingly indispensable for better satisfyin…

2023

Intra-Event and Inter-Event Dependency-Aware Graph Network for Event Argument Extraction

EMNLP 2023long findings

Event argument extraction is critical to various natural language processing tasks for providing structured information. Existing works usually extract the event arguments one by one, and mostly neglect to build dependency information among event argument roles, especially from the perspective of ev…

Cited by 0SourceScholar
2023

Learning a Sparse Transformer Network for Effective Image Deraining

CVPR 2023highlight

Transformers-based methods have achieved significant performance in image deraining as they can model the non-local information which is vital for high-quality image reconstruction. In this paper, we find that most existing Transformers usually use all similarities of the tokens from the query-key p…

2023

MonoNeRD: NeRF-like Representations for Monocular 3D Object Detection

ICCV 2023poster

In the field of monocular 3D detection, it is common practice to utilize scene geometric clues to enhance the detector's performance. However, many existing works adopt these clues explicitly such as estimating a depth map and back-projecting it into 3D space. This explicit methodology induces spars…

Cited by 35PDFcodeScholar
2023

NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates Space

ICCV 2023poster

Monocular 3D Semantic Scene Completion (SSC) has garnered significant attention in recent years due to its potential to predict complex semantics and geometry shapes from a single image, requiring no 3D inputs. In this paper, we identify several critical issues in current state-of-the-art methods, i…

Cited by 174PDFcodeScholar
2023

Not all quantifiers are equal: Probing Transformer-based language models' understanding of generalised quantifiers

EMNLP 2023long main

How do different generalised quantifiers affect the behaviour of transformer-based language models (TLMs)? The recent popularity of TLMs and the central role generalised quantifiers have traditionally played in linguistics and logic bring this question into particular focus. The current research inv…

Cited by 0SourceScholar
2023

Point-Teaching: Weakly Semi-supervised Object Detection with Point Annotations

AAAI 2023technical

Point annotations are considerably more time-efficient than bounding box annotations. However, how to use cheap point annotations to boost the performance of semi-supervised object detection is still an open question. In this work, we present Point-Teaching, a weakly- and semi-supervised object dete…

2023

Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval

NeurIPS 2023poster

Cross-modal Retrieval methods build similarity relations between vision and language modalities by jointly learning a common representation space. However, the predictions are often unreliable due to the Aleatoric uncertainty, which is induced by low-quality data, e.g., corrupt images, fast-paced vi…

2023

Sonicverse: A Multisensory Simulation Platform for Embodied Household Agents that See and Hear

ICRA 2023poster

Developing embodied agents in simulation has been a key research topic in recent years. Exciting new tasks, algorithms, and benchmarks have been developed in various simulators. However, most of them assume deaf agents in silent environments, while we humans perceive the world with multiple senses.…

Cited by 12SourcecodeScholar
2023

SteerNeRF: Accelerating NeRF Rendering via Smooth Viewpoint Trajectory

CVPR 2023poster

Neural Radiance Fields (NeRF) have demonstrated superior novel view synthesis performance but are slow at rendering. To speed up the volume rendering process, many acceleration methods have been proposed at the cost of large memory consumption. To push the frontier of the efficiency-memory trade-off…

2023

StyleGene: Crossover and Mutation of Region-Level Facial Genes for Kinship Face Synthesis

CVPR 2023highlight

High-fidelity kinship face synthesis has many potential applications, such as kinship verification, missing child identification, and social media analysis. However, it is challenging to synthesize high-quality descendant faces with genetic relations due to the lack of large-scale, high-quality anno…

2023

TG-VQA: Ternary Game of Video Question Answering

IJCAI 2023poster

Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions, i.e., annotations or priors, current contrastive learning-based VideoQA methods remains challenging to perform fine-gr…

Cited by 17SourcePDFScholar
2023

Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment

IJCAI 2023poster

Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they fail to leverage the heterogeneous concepts in data. In this…

2023

The ObjectFolder Benchmark: Multisensory Learning With Neural and Real Objects

CVPR 2023poster

We introduce the ObjectFolder Benchmark, a benchmark suite of 10 tasks for multisensory object-centric learning, centered around object recognition, reconstruction, and manipulation with sight, sound, and touch. We also introduce the ObjectFolder Real dataset, including the multisensory measurements…

Cited by 31SourcePDFScholar
2023

Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks

CVPR 2023highlight

Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to eliminating this inconsistency is to use generalist models for general task modeling. However, existing attempts at generalis…

2023

WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image Segmentation

IJCAI 2023poster

The top-down and bottom-up methods are two mainstreams of referring segmentation, while both methods have their own intrinsic weaknesses. Top-down methods are chiefly disturbed by Polar Negative (PN) errors owing to the lack of fine-grained cross-modal alignment. Bottom-up methods are mainly perturb…

Cited by 4SourcePDFScholar
2023

XMem++: Production-level Video Segmentation From Few Annotated Frames

ICCV 2023poster

Despite advancements in user-guided video segmentation, extracting complex objects consistently for highly complex scenes is still a labor-intensive task, especially for production. It is not uncommon that a majority of frames need to be annotated. We introduce a novel semi-supervised video object s…

Cited by 33PDFcodeScholar
2022

A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal Retrieval

NeurIPS 2022accept

Cross-modal retrieval aims to build correspondence between multiple modalities by learning a common representation space. Typically, an image can match multiple texts semantically and vice versa, which significantly increases the difficulty of this task. To address this problem, probabilistic embedd…

2022

ATF-3D: Semi-Supervised 3D Object Detection With Adaptive Thresholds Filtering Based on Confidence and Distance

RA-L 2022

Performance of current point cloud-based outdoor 3D object detection relies heavily on large-scale high-quality 3D annotations. However, such annotations are usually expensive to collect and outdoor scenes easily accumulate massive unlabeled data containing rich scenes. Semi-supervised learning is a

Cited by 12SourceScholar