← Search

Yi Yang

442 accepted papers

2026

A Roadmap for Responsible Robotics

ICRA 2026poster

This document presents the outcomes of the Dagstuhl Seminar "Roadmap for Responsible Robotics," held in September 2023 at the Leibniz Centre for Informatics, Schloss Dagstuhl, Germany. The seminar brought together researchers from Robotics, Computer Science, Social and Cognitive Sciences, and Philos…

Cited by 0Scholar
2026

AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows

CVPR 2026

Training-free 3D editing aims to modify 3D shapes based on human instructions without model finetuning. It plays a crucial role in 3D content creation. However, existing approaches often struggle to produce strong or geometrically stable edits, largely due to inconsistent latent anchors introduced b

Cited by 0SourcecodeScholar
2026

Beyond Independent Genes: Learning Module-Inductive Representations for Gene Perturbation Prediction

ICML 2026poster

Predicting transcriptional responses to genetic perturbations is a central problem in functional genomics. In practice, perturbation responses are rarely gene-independent but instead manifest as coordinated, program-level transcriptional changes among functionally related genes. However, most existi…

Cited by 0SourceScholar
2026

BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment

ICLR 2026poster

Conditional image generation augments text-to-image synthesis with structural, spatial, or stylistic priors and is used in many domains. However, current methods struggle to harmonize guidance from both sources when conflicts arise: 1) input-level conflict, where the semantics of the conditioning im…

Cited by 0SourcecodeScholar
2026

Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra

AAAI 2026technical

Retrieving molecular structures from tandem mass spectra is a crucial step in rapid compound identification. Existing retrieval methods, such as traditional mass spectral library matching, suffer from limited spectral library coverage, while recent cross-modal representation learning frameworks ofte

Cited by 0SourcePDFScholar
2026

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts

CVPR 2026

Multimodal large language models (MLLMs) have shown considerable potential in chart understanding and reasoning tasks. However, they still struggle with high information density (HID) charts characterized by multiple subplots, legends, and dense annotations due to three major challenges: (1) limited

Cited by 0SourcecodeScholar
2026

CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving

ICLR 2026poster

Despite recent advances, multimodal large language models continue to struggle with visual mathematical problem solving. Some recent works recognize that visual perception is a bottleneck in visual mathematical reasoning, but their solutions are limited to improving the extraction and interpretation…

Cited by 0SourceScholar
2026

ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation

ICLR 2026poster

Multi-instance image generation (MIG) remains a significant challenge for modern diffusion models due to key limitations in achieving precise control over object layout and preserving the identity of multiple distinct subjects. To address these limitations, we introduce **ContextGen**, a novel Diffu…

Cited by 0SourcecodeScholar
2026

DIPP: A Diffusion-Based Potential Planner for Synergistic Navigation and Mapping

ICRA 2026poster

Object-Goal Navigation (ObjectNav) requires an embodied agent to search for and reach a target object category in previously unseen environments using only onboard egocentric observations, which is a fundamental capability for long-horizon autonomous robots. Current Object-Goal Navigation methods ty…

Cited by 0Scholar
2026

DLVINet: Advancing Dual-Lens Video Inpainting Beyond Parallax Constraints

AAAI 2026technical

Dual-lens video inpainting aims to simultaneously restore missing or corrupted contents in videos captured by each lens of binocular systems. Although preliminary explorations have been conducted, existing methods still face two key challenges: limited exploitation of long-range reference informatio

Cited by 0SourcePDFScholar
2026

DSSM-SG: Dynamic 3D Scene Graphs with Spatio-Semantic Memory for Long-Term Indoor Navigation Tasks

ICRA 2026poster

Dynamic indoor environments pose significant challenges for autonomous robots, as objects frequently move and scenes continuously change, requiring robust scene representation and adaptive navigation strategies. In this work, we introduce DSSM-SG, a dynamic open-vocabulary 3D scene graph framework e…

Cited by 0Scholar
2026

Echoes of Ownership: Adversarial-Guided Dual Injection for Copyright Protection in MLLMs

CVPR 2026

With the rapid deployment of multimodal large language models (MLLMs), disputes regarding model ownership have become increasingly frequent, raising significant concerns about intellectual property protection. In this paper, we propose a framework for generating copyright triggers for MLLMs, enablin

Cited by 0SourcecodeScholar
2026

Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers

ICML 2026poster

The quadratic complexity of standard attention mechanisms poses a significant scalability bottleneck for large language models (LLMs) in long-context scenarios. While hybrid attention strategies that combine sparse and full attention within a single model offer a viable solution, they typically empl…

Cited by 0SourceScholar
2026

Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World

ICLR 2026poster

In this paper, we explore how to empower general-purpose Vision-Language Models (VLMs) to control humanoid agents. General-purpose VLMs (e.g., GPT-4) exhibit strong open-world generalization, and remove the need for additional fine-tuning data. To build such an agent, two key components are required…

Cited by 0SourcecodeScholar
2026

Energy-GS: Image Energy-guided Pose Alignment Gaussian Splatting with redesigned pose gradient flow

CVPR 2026

High-quality 3D scene representation in radiance fields relies on accurate camera poses which are often difficult to acquire in real-world scenarios. An effective solution is to use RGB images for the joint optimization of radiance fields and camera poses, an approach that has been well explored in

Cited by 0SourcecodeScholar
2026

FUSAR-GPT: A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery

CVPR 2026

Research on the intelligent interpretation of all-weather, all-time Synthetic Aperture Radar (SAR) is crucial for advancing remote sensing applications. In recent years, although Visual Language Models (VLMs) have demonstrated strong open-world understanding capabilities on RGB images, their perform

Cited by 0SourceScholar
2026

FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation

ICML 2026poster

Simulation-Based Inference (SBI) is critical for scientific discovery, with generative models offering a promising path toward efficient inference. However, existing methods struggle with effective multimodal modeling. They often rely on brute-force fusion strategies that ignore the structural dispa…

Cited by 0SourceScholar
2026

FilterGS: Traversal-Free Parallel Filtering and Adaptive Shrinking for Large-Scale LoD 3D Gaussian Splatting

CVPR 2026

3D Gaussian Splatting has revolutionized neural rendering with real-time performance. However, scaling this approach to large scenes using Level-of-Detail methods faces critical challenges: inefficient serial traversal consuming over 60% of rendering time, and redundant Gaussian-tile pairs that incu

Cited by 0SourcecodeScholar
2026

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

CVPR 2026

Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they struggle with fine-grained temporal control in multi-event scenarios or when visual cues are insufficient, such as small regions, off-screen sounds, or occlud

Cited by 0SourcecodeScholar
2026

From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion

ICML 2026poster

Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a single fused image that preserves *fine local details* while maintaining *globally consistent appearance*. Most existing approaches build shared representations on 2D feature grids, which exce…

Cited by 0SourceScholar
2026

Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token Selection

CVPR 2026

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet processing long visual token sequences remains computationally expensive. Existing approaches mitigate this cost by reducing image tokens, either by discarding them after the visual encod

Cited by 0SourcecodeScholar
2026

HiMo: High-Speed Objects Motion Compensation in Point Clouds (Abstract Reprint)

AAAI 2026technical

LiDAR point cloud is essential for autonomous vehicles, but motion distortions from dynamic objects degrade the data quality. While previous work has considered distortions caused by ego motion, distortions caused by other moving objects remain largely overlooked, leading to errors in object shape a

Cited by 0SourcePDFScholar
2026

Insert Anything: Image Insertion via In-Context Editing in DiT

AAAI 2026technical

This work presents Insert Anything, a unified framework for reference-based image insertion that seamlessly integrates objects from reference images into target scenes under flexible, user-specified control guidance. Instead of training separate models for individual tasks, our approach is trained o

Cited by 0SourcePDFScholar
2026

Learning a Unified Latent Action Space from Videos with Action-centric Cycle Consistency

CVPR 2026

Video data provides a rich source beyond expensive action-labeled data for advancing robot learning. Recent approaches have demonstrated promising potential in leveraging video data by learning latent actions for policy training. The latent action tokenizer encodes latent actions between successive

Cited by 0SourceScholar
2026

LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization

ICLR 2026poster

Generating coherent and communicative visual sequences, such as image sequences and videos, remains a significant challenge for current multimodal systems. Despite advances in visual quality and the integration of world knowledge, existing models still struggle to maintain logical flow, often result…

Cited by 0SourceScholar
2026

Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective

ICLR 2026poster

Autoregressive large language models (LLMs) have unified a vast range of language tasks, inspiring preliminary efforts in autoregressive (AR) video generation. Existing AR video generators either diverge from standard LLM architectures, depend on bulky external text encoders, or incur prohibitive la…

Cited by 0SourcecodeScholar
2026

MCOO-SLAM: A Multi-Camera Omnidirectional Object SLAM System

RA-L 2026

Object-level SLAM offers structured and semantically meaningful environment representations, making it more interpretable and suitable for high-level robotic tasks. However, most existing approaches rely on RGB-D sensors or monocular views, which suffer from narrow fields of view, occlusion sensitiv

Cited by 2SourceScholar
2026

Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight

CVPR 2026

Recent advances in Vision-Language-Action (VLA) models demonstrate that visual signals can effectively complement sparse action supervisions. However, letting VLA directly predict high-dimensional visual states can distribute model capacity and incur prohibitive training cost, while compressing visu

Cited by 0SourcecodeScholar
2026

Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory Fields

CVPR 2026

Tracking-any-point (TAP) answers query-conditioned correspondence but leaves the dense, all-pairs structure of a video implicit. We formulate All-Pairs Tracking (APT): given a video, predict dense displacement and visibility for every source-target frame pair, from which per-pixel trajectories can b

Cited by 0SourceScholar
2026

Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction

ICLR 2026poster

Reconstructing visual stimuli from fMRI signals is a central challenge bridging machine learning and neuroscience. Recent diffusion-based methods typically map fMRI activity to a single neural embedding, using it as static guidance throughout the entire generation process. However, this fixed guidan…

Cited by 0SourcecodeScholar
2026

Object Fidelity Diffusion for Remote Sensing Image Generation

ICLR 2026poster

High-precision controllable remote sensing image generation is both meaningful and challenging. Existing diffusion models often produce low-fidelity objects due to their inability to adequately capture morphological details, which may affect the robustness and reliability of object detection models.…

Cited by 0SourcecodeScholar
2026

OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics

ICRA 2026poster

Robotic systems demand accurate and comprehensive 3D environment perception, requiring simultaneous capture of photo-realistic appearance (optical), precise layout shape (geometric), and open-vocabulary scene understanding (semantic). Existing methods typically achieve only partial fulfillment of th…

2026

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

CVPR 2026

Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale and semantic diversity, causing performance gaps between common and rare concepts. To overcome these limitations, we introduce OmniVTG, a new large-s

Cited by 0SourcecodeScholar
2026

Oscillation Inversion: Training-Free Image and Video Enhancement Through Oscillated Latents in Large Flow Models

AAAI 2026technical

We explore the oscillatory behavior observed in inversion methods applied to large-scale flow models, including text-to-image and text-to-video. By employing an augmented fixed-point-inspired iterative approach to invert real-world images, we observe that the solution does not achieve convergence, i

Cited by 0SourcePDFScholar
2026

PIPS: Planar Instance 3D Reconstruction Leveraging Planar Structural Priors

ICRA 2026poster

Planar structures, ubiquitous in man-made indoor environments, enable compact and accurate scene abstraction for various downstream tasks. Recent methods distill planar features into learning-based MVS geometries to obtain coherent 3D plane estimation from multi-view inputs. However, the lack of exp…

Cited by 0codeScholar
2026

PointThinker: Point-Incentivized Parallel Thinking for Multimodal Large Language Model

CVPR 2026

This paper explores parallel thinking for Multi-modal Large Language Models (MLLMs), aiming to improve Chain-of-Thought (CoT) through multiple diverse reasoning paths. We guide the model to list multiple visual key points and develop an independent reasoning path for each. Therefore, we term this me

Cited by 0SourceScholar
2026

Polyphonia: Training-Free Context-Aware Music Editing with Acoustic-Informed Attention Calibration

ICML 2026poster

The advancement of diffusion-based text-to-music generation has opened new avenues for zero-shot music editing. However, existing methods fail to achieve context-aware editing, which requires altering specific stems while strictly preserving the background accompaniment. This limitation severely hin…

Cited by 0SourceScholar
2026

Rendering Multi-Human and Multi-Object with 3D Gaussian Splatting

ICRA 2026poster

Reconstructing dynamic scenes with multiple interacting humans and objects from sparse-view inputs is a critical yet challenging task, essential for creating high-fidelity digital twins for robotics and VR/AR. This problem, which we term Multi-Human Multi-Object (MHMO) rendering, presents two signif…

2026

SSR-ZSON: Zero-Shot Object Navigation Via Spatial-Semantic Relations within a Hierarchical Exploration Framework

ICRA 2026poster

Zero-shot object navigation in unknown environments presents significant challenges, mainly due to two key limitations: insufficient semantic guidance leads to inefficient exploration, while limited spatial memory resulting from environmental structure causes entrapment in local regions. To address …

2026

Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models

ICLR 2026poster

Rigged 3D assets are fundamental to 3D deformation and animation. However, existing 3D generation methods face challenges in generating animatable geometry, while rigging techniques lack fine-grained structural control over skeleton creation. To address these limitations, we introduce Stroke3D, a no…

Cited by 0SourceScholar
2026

Structured Reasoning for LLMs: A Unified Framework for Efficiency and Explainability

ICLR 2026poster

Recent Large Language Models (LLMs) have made remarkable progress, but they still struggle with complex reasoning tasks such as logical deduction and planning. This is partly because they rely primarily on token-level probability relationships, which limits their ability to reason effectively. In t…

Cited by 0SourcecodeScholar
2026

TaRO: Temporal-Aware Reasoning Optimization for Video Temporal Grounding

ICML 2026poster

Multi-modal Large Language Models (MLLMs) have achieved remarkable progress in video temporal grounding (VTG) with the introduction of reinforcement learning (RL) for generating reasoning paths. However, existing models often produce superficial reasoning, such as providing generic video description…

Cited by 0SourceScholar
2026

TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation

ICML 2026poster

Monocular normal estimation for transparent objects is critical for laboratory automation, yet it remains challenging due to complex light refraction and reflection. These optical properties often lead to catastrophic failures in conventional depth and normal sensors, hindering the deployment of emb…

Cited by 0SourceScholar
2026

Uncertainty-Aware 3D Reconstruction for Dynamic Underwater Scenes

ICLR 2026poster

Underwater 3D reconstruction remains challenging due to the intricate interplay between light scattering and environment dynamics. While existing methods yield plausible reconstruction with rigid scene assumptions, they struggle to capture temporal dynamics and remain sensitive to observation noise.…

Cited by 0SourceScholar
2026

Uncertainty-Aware Gaussian Map for Vision-Language Navigation

ICLR 2026poster

Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter perceptual uncertainty, such as insufficient evidence for reliable grounding or ambiguity in interpreting spatial cues, yet th…

Cited by 0SourceScholar
2026

Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization

CVPR 2026

Parallel test-time scaling typically trains separate generation and verification models, incurring high training and inference costs. We propose Advantage Decoupled Preference Optimization (ADPO), a unified reinforcement learning framework that jointly learns answer generation and self-verification

Cited by 0SourcecodeScholar
2026

Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learning

CVPR 2026

Scalable robot learning is hindered by the high cost of acquiring diverse, high-quality embodied data. Existing data generation approaches partially mitigate this issue but typically depend on hard-to-access hardware and labor-intensive manual effort, with limited generalization to diverse scene con

Cited by 0SourceScholar
2025

3DIS: Depth-Driven Decoupled Image Synthesis for Universal Multi-Instance Generation

ICLR 2025spotlight

The increasing demand for controllable outputs in text-to-image generation has spurred advancements in multi-instance generation (MIG), allowing users to define both instance layouts and attributes. However, unlike image-conditional generation methods such as ControlNet, MIG techniques have not been…

2025

Achieving binary weight and activation for LLMs using Post-Training Quantization

ACL 2025finding

Quantizing large language models (LLMs) to 1-bit precision significantly reduces computational costs, but existing quantization techniques suffer from noticeable performance degradation when using weight and activation precisions below 4 bits (W4A4). In this paper, we propose a post-training quantiz…

2025

Adapting General-Purpose Embedding Models to Private Datasets Using Keyword-based Retrieval

ACL 2025finding

Text embedding models play a cornerstone role in AI applications, such as retrieval-augmented generation (RAG). While general-purpose text embedding models demonstrate strong performance on generic retrieval benchmarks, their effectiveness diminishes when applied to private datasets (e.g., company-s…

Cited by 0SourcePDFScholar
2025

Adapting Text-to-Image Generation with Feature Difference Instruction for Generic Image Restoration

CVPR 2025poster

Diffusion-based Text-to-Image (T2I) models have demonstrated significant potential in image restoration. However, existing models continue to grapple with challenges such as complex training and prompt design. We introduce a new perspective for improving image restoration by injecting knowledge from…

Cited by 0SourcePDFScholar
2025

Automated 3D-GS Registration and Fusion via Skeleton Alignment and Gaussian-Adaptive Features

IROS 2025

In recent years, 3D Gaussian Splatting (3D-GS)based scene representation demonstrates significant potential in real-time rendering and training efficiency. However, most existing methods primarily focus on single-map reconstruction, while the registration and fusion of multiple 3D-GS submaps remain

Cited by 2SourceScholar
2025

Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion

AAAI 2025technical

Human motion generative models have enabled promising applications, but the ability of text-to-motion (T2M) models to produce realistic motions raises security concerns if exploited maliciously. Despite growing interest in T2M, limited research focus on safeguarding these models against adversarial…

Cited by 2SourcePDFScholar
2025

BVINet: Unlocking Blind Video Inpainting with Zero Annotations

ICCV 2025poster

Video inpainting aims to fill in corrupted regions of the video with plausible contents. Existing methods generally assume that the locations of corrupted regions are known, focusing primarily on the "how to inpaint". This reliance necessitates manual annotation of the corrupted regions using binary…

Cited by 0SourcePDFScholar
2025

BrainGuard: Privacy-Preserving Multisubject Image Reconstructions from Brain Activities

AAAI 2025technical

Reconstructing perceived images from human brain activity forms a crucial link between human and machine learning through Brain-Computer Interfaces. Early methods primarily focused on training separate models for each individual to account for individual variability in brain activity, overlooking va…

2025

DICS: Find Domain-Invariant and Class-Specific Features for Out-of-Distribution Generalization

ICASSP 2025accepted

While deep neural networks have made remarkable progress in various tasks, their performance typically deteriorates and faces insecurity when tested in out-of-distribution (OOD) scenarios. Many OOD methods focus on extracting domain-invariant features but neglect whether these features are unique to…

Cited by 0SourceScholar
2025

DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization

EMNLP 2025

Recent advances in Emotional Support Conversation (ESC) have improved emotional support generation by fine-tuning Large Language Models (LLMs) via Supervised Fine-Tuning (SFT). However, common psychological errors still persist. While Direct Preference Optimization (DPO) shows promise in reducing su

2025

DeltaPhi: Physical States Residual Learning for Neural Operators in Data-Limited PDE Solving

NeurIPS 2025poster

The limited availability of high-quality training data poses a major obstacle in data-driven PDE solving, where expensive data collection and resolution constraints severely impact the ability of neural operator networks to learn and generalize the underlying physical system. To address this challen…

Cited by 0SourceScholar
2025

DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization

ICML 2025poster

Text-to-3D generation automates 3D content creation from textual descriptions, which offers transformative potential across various fields. However, existing methods often struggle to align generated content with human preferences, limiting their applicability and flexibility. To address these limit…

Cited by 7SourcePDFScholar
2025

DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models

ICCV 2025poster

Image-conditioned generation methods, such as depth- and canny-conditioned approaches, have demonstrated remarkable abilities for precise image synthesis. However, existing models still struggle to accurately control the content of multiple instances (or regions). Even state-of-the-art models like F…

2025

DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagery

CVPR 2025highlight

Drones have become essential tools for reconstructing wild scenes due to their outstanding maneuverability. Recent advances in radiance field methods have achieved remarkable rendering quality, providing a new avenue for 3D reconstruction from drone imagery. However, dynamic distractors in wild env…

Cited by 2SourcePDFScholar
2025

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs

EMNLP 2025

Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activated for each token, SMoE still requires loading all expert parameters, leading to high memory usage and challenges in dep

Cited by 0SourcePDFScholar
2025

Dual Reciprocal Learning of Language-based Human Motion Understanding and Generation

ICCV 2025poster

Language-based human motion understanding focuses on describing human motions using natural language descriptions. Conversely, human motion generation aims to generate human motions from textual inputs. Despite significant progress in both fields, further advancements are hindered by two primary cha…

2025

Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

NeurIPS 2025poster

Instruction-based image editing enables precise modifications via natural language prompts, but existing methods face a precision-efficiency tradeoff: fine-tuning demands massive datasets (>10M) and computational resources, while training-free approaches suffer from weak instruction comprehension.…

Cited by 0SourceScholar
2025

Evaluating Global Geo-Alignment for Precision Learned Autonomous Vehicle Localization Using Aerial Data

ICRA 2025

Recently there has been growing interest in the use of aerial and satellite map data for autonomous vehicles, primarily due to its potential for significant cost reduction and enhanced scalability. Despite the advantages, aerial data also comes with challenges such as a sensor-modality gap and a vie

Cited by 1SourceScholar
2025

Fixed-Time Variable Gain Trajectory Tracking Control for 6-DOF Manipulators With Unknown Disturbances

RA-L 2025

In this paper, fixed-time variable gain trajectory tracking control is designed for solving nonlinear unknown disturbance problems of 6-DOF manipulators. The designed control method comprises a fixed-time variable gain disturbance observer (FTVGDOB) and a fixed-time variable gain controller (FTVGC),

Cited by 0SourceScholar
2025

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

NeurIPS 2025poster

Long-form video understanding poses a significant challenge for video large language models (VideoLLMs) due to prohibitively high computational and memory demands. In this paper, We propose $\textbf{FlexSelect}$, a flexible and efficient token selection strategy for processing long videos. FlexSele…

Cited by 0SourcecodeScholar
2025

From Image to Video: An Empirical Study of Diffusion Representations

ICCV 2025poster

Diffusion models have revolutionized generative modeling, enabling unprecedented realism in image and video synthesis.This success has sparked interest in leveraging their representations for visual understanding tasks. While recent works have explored this potential for image generation, the visual…

Cited by 0SourcePDFScholar
2025

From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment

ICCV 2025poster

Multi-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process a finite number of frames in a single inference, potentially omitting crucial visual information. To address the challe…

Cited by 0SourcePDFScholar
2025

Gaussian-based World Model: Gaussian Priors for Voxel-Based Occupancy Prediction and Future Motion Prediction

ICCV 2025poster

In autonomous driving, accurately predicting occupancy and motion is crucial for safe navigation within dynamic environments. However, existing methods often suffer from difficulties in handling complex scenes and uncertainty arising from sensor data. To address these issues, we propose a new Gaussi…

2025

GaussianGraph: 3D Gaussian-Based Scene Graph Generation for Open-World Scene Understanding

IROS 2025

Recent advancements in 3D Gaussian Splatting(3DGS) have significantly improved semantic scene understanding, enabling natural language queries to localize objects within a scene. However, existing methods primarily focus on embedding compressed CLIP features to 3D Gaussians, suffering from low objec

Cited by 6SourcecodeScholar
2025

GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning

CVPR 2025poster

Learning from demonstration is a powerful method for robotic skill acquisition. However, the significant expense of collecting such action-labeled robot data presents a major bottleneck. Video data, a rich data source encompassing diverse behavioral and physical knowledge, emerges as a promising alt…

Cited by 0SourcePDFScholar
2025

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

ICCV 2025poster

In this paper, we tackle the task of online video temporal grounding (OnVTG), which requires the model to locate events related to a given text query within a video stream. Unlike regular video temporal grounding, OnVTG requires the model to make predictions without observing future frames. As onlin…

2025

High-Fidelity Stereoscopic Image Rain Removal with Texture Integrity and Disparity Consistency

ICASSP 2025accepted

This paper tackles the challenge of stereoscopic image rain removal by focusing on enhancing texture integrity and disparity consistency. Existing stereoscopic rain removal techniques often fall short due to 1) disruptions in texture coherence caused by complex rain streaks, and 2) inaccuracies in d…

Cited by 0SourceScholar
2025

High-Precision Object Pose Estimation Using Visual-Tactile Information for Dynamic Interactions in Robotic Grasping

ICRA 2025

In various robotic applications, understanding accurate object poses for robots is essential for high-precision tasks such as factory assembly or daily insertions. Tactile sensing, which compensates for visual information, offers rich texture-based or force-based data for object pose estimation. How

Cited by 0SourceScholar
2025

Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph Generation

ICLR 2025poster

DETR introduces a simplified one-stage framework for scene graph generation (SGG) but faces challenges of sparse supervision and false negative samples. The former occurs because each image typically contains fewer than 10 relation annotations, while DETR-based SGG models employ over 100 relation qu…

Cited by 4SourcePDFScholar
2025

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework

EMNLP 2025

The performance of large language models (LLMs) is closely tied to their training data, which can include copyrighted material or private information, raising legal and ethical concerns. Additionally, LLMs face criticism for dataset contamination and internalizing biases. To address these issues, th

2025

Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models

AAAI 2025technical

Diffusion models have revitalized the image generation domain, playing crucial roles in both academic research and artistic expression. With the emergence of new diffusion models, assessing the performance of text-to-image models has become increasingly important. Current metrics focus on directly m…

2025

Internal-Stably Energy-Saving Cooperative Control of Articulated Wheeled Robot with Distributed Drive Units

ICRA 2025

Articulated wheeled robots play a crucial role in the logistics industry. However, conventional tractor-driven articulated wheeled robots exhibit poor internal stability and are prone to jackknifing, while also consuming a significant amount of energy. By deploying distributed drives and coordinatin

Cited by 0SourceScholar
2025

Know the Unknown: An Uncertainty-Sensitive Method for LLM Instruction Tuning

ACL 2025finding

Large language models (LLMs) demonstrate remarkable capabilities but face challenges from hallucinations, which typically arise from insufficient knowledge or context. While instructing LLMs to acknowledge knowledge limitations by responding with “I don’t know” appears promising, we find that models…

2025

LGSDF: Continual Global Learning of Signed Distance Fields Aided by Local Updating

RA-L 2025

Implicit reconstruction of ESDF (Euclidean Signed Distance Field) involves training a neural network to regress the signed distance from any point to the nearest obstacle, which has the advantages of lightweight storage and continuous querying. However, existing algorithms usually rely on conflictin

Cited by 5SourcecodeScholar
2025

LLM Agents Can Be Choice-Supportive Biased Evaluators: An Empirical Study

AAAI 2025technical

With Large Language Model (LLM) agents taking on more evaluation responsibilities in decision-making, it is essential to recognize their possible biases to guarantee fair and trustworthy AI-supported decisions. This study is the first to thoroughly examine the choice-supportive bias in LLM agents, a…

Cited by 0SourcePDFScholar
2025

LawShift: Benchmarking Legal Judgment Prediction Under Statute Shifts

NeurIPS 2025poster

Legal Judgment Prediction (LJP) seeks to predict case outcomes given available case information, offering practical value for both legal professionals and laypersons. However, a key limitation of existing LJP models is their limited adaptability to statutory revisions. Current SOTA models are neithe…

Cited by 0SourceScholar
2025

Learning without Isolation: Pathway Protection for Continual Learning

ICML 2025poster

Deep networks are prone to catastrophic forgetting during sequential task learning, i.e., losing the knowledge about old tasks upon learning new tasks. To this end, continual learning (CL) has emerged, whose existing methods focus mostly on regulating or protecting the parameters associated with the…

2025

Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflection

ICLR 2025poster

Visual instructions for long-horizon tasks are crucial as they intuitively clarify complex concepts and enhance retention across extended steps. Directly generating a series of images using text-to-image models without considering the context of previous steps results in inconsistent images, increa…

Cited by 0SourcePDFScholar
2025

MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures - A Comprehensive Framework

EMNLP 2025

Large Language Models (LLMs)-based Multi-Agent Systems (MAS) exhibit remarkable problem-solving and task planning capabilities across diverse domains due to their specialized agentic roles and collaborative interactions. However, this also amplifies the severity of security risks under MAS attacks.

Cited by 0SourcePDFScholar
2025

MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs

ICCV 2025poster

While multimodal large language models (MLLMs) have demonstrated extraordinary vision-language understanding capabilities, their abilities to solve instance-level visual-language problems beyond a single image warrant further exploration. To assess these unproven abilities of MLLMs, this paper propo…

2025

MaGS: Reconstructing and Simulating Dynamic 3D Objects with Mesh-adsorbed Gaussian Splatting

ICCV 2025poster

3D reconstruction and simulation, although interrelated, have distinct objectives: reconstruction requires a flexible 3D representation that can adapt to diverse scenes, while simulation needs a structured representation to model motion principles effectively. This paper introduces the Mesh-adsorbed…

Cited by 0SourcePDFScholar
2025

MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh

ICCV 2025poster

We present MeshLLM, a novel framework that leverages large language models (LLMs) to understand and generate text-serialized 3D meshes. Our approach addresses key limitations in existing methods, including the limited dataset scale when catering to LLMs' token length and the loss of 3D structural in…

Cited by 0SourcePDFScholar
2025

NeRF Is a Valuable Assistant for 3D Gaussian Splatting

ICCV 2025poster

We introduce NeRF-GS, a novel framework that jointly optimizes Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). This framework leverages the inherent continuous spatial representation of NeRF to mitigate several limitations of 3DGS, including sensitivity to Gaussian initialization, li…

Cited by 0SourcePDFScholar
2025

OSDA Agent: Leveraging Large Language Models for De Novo Design of Organic Structure Directing Agents

ICLR 2025spotlight

Zeolites are crystalline porous materials that have been widely utilized in petrochemical industries as well as sustainable chemistry areas. Synthesis of zeolites often requires small molecules termed Organic Structure Directing Agents (OSDAs), which are critical in forming the porous structure. Mol…

Cited by 0SourcePDFScholar
2025

Open-RGBT: Open-Vocabulary RGB-T Zero-Shot Semantic Segmentation in Open-World Environments

ICRA 2025

Semantic segmentation is a critical technique for effective scene understanding. Traditional RGB-T semantic segmentation models often struggle to generalize across diverse scenarios due to their reliance on pretrained models and predefined categories. Recent advancements in Visual Language Models (V

Cited by 0SourcecodeScholar
2025

OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding

IROS 2025

Recent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing methods are hindered by rigid offline pipelines and the inability to provide precise

Cited by 4SourcecodeScholar
2025

OpenGS-SLAM: Open-Set Dense Semantic SLAM with 3D Gaussian Splatting for Object-Level Scene Understanding

ICRA 2025

Recent advancements in 3D Gaussian Splatting have significantly improved the efficiency and quality of dense semantic SLAM. However, previous methods are generally constrained by limited-category pre-trained classifiers and implicit semantic representation, which hinder their performance in open-set

Cited by 15SourcecodeScholar
2025

OpenMulti: Open-Vocabulary Instance-Level Multi-Agent Distributed Implicit Mapping

RA-L 2025

Multi-agent distributed collaborative mapping provides comprehensive and efficient representations for robots. However, existing approaches lack instance-level awareness and semantic understanding of environments, limiting their effectiveness for downstream applications. To address this issue, we pr

Cited by 3SourceScholar
2025

OpenObj: Open-Vocabulary Object-Level Neural Radiance Fields With Fine-Grained Understanding

RA-L 2025

In recent years, there has been a surge of interest in open-vocabulary 3D scene reconstruction facilitated by visual language models (VLMs), which showcase remarkable capabilities in open-set retrieval tasks. Although the semantic ambiguity of existing point-wise feature maps is alleviated by open-v

Cited by 12SourceScholar
2025

OpenVox: Real-time Instance-level Open-vocabulary Probabilistic Voxel Representation

IROS 2025

In recent years, vision-language models (VLMs) have advanced open-vocabulary mapping, enabling mobile robots to simultaneously achieve environmental reconstruction and high-level semantic understanding. While integrated object cognition helps mitigate semantic ambiguity in point-wise feature maps, e

Cited by 4SourcecodeScholar
2025

Origin Identification for Text-Guided Image-to-Image Diffusion Models

ICML 2025poster

Text-guided image-to-image diffusion models excel in translating images based on textual prompts, allowing for precise and creative visual modifications. However, such a powerful technique can be misused for *spreading misinformation*, *infringing on copyrights*, and *evading content tracing*. This…

2025

Parking-SG: Open-Vocabulary Hierarchical 3D Scene Graph Representation for Open Parking Environments

ICRA 2025

Automatic Valet Parking (AVP) has garnered significant attention from industry and academia due to its potential to enhance traffic efficiency, parking safety, and user experience. While AVP technologies have been successfully applied in standard parking scenarios with clear markings, real-world par

Cited by 2SourceScholar
2025

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

ICCV 2025poster

Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge. Existing visual-language models often struggle to effectively analyze and reaso…

2025

Reaction Graph: Towards Reaction-Level Modeling for Chemical Reactions with 3D Structures

ICML 2025poster

Accurately modeling chemical reactions using Artificial Intelligence (AI) can accelerate discovery and development, especially in fields like drug design and material science. Although AI has made remarkable advancements in single molecule recognition, such as predicting molecular properties, the st…

2025

Representation Learning with Mutual Influence of Modalities for Node Classification in Multi-Modal Heterogeneous Networks

IJCAI 2025

Nowadays, numerous online platforms can be described as multi-modal heterogeneous networks (MMHNs), such as Douban's movie networks and Amazon's product review networks. Accurately categorizing nodes within these networks is crucial for analyzing the corresponding entities, which requires effective

2025

RoadsideSplat: Robust 3D Gaussian Reconstruction from Monocular Roadside Surveillance

IROS 2025

Reconstructing dynamic roads from roadside traffic surveillance cameras is crucial for smart cities and digital twin applications. While the latest monocular depth estimation methods demonstrate strong performance, they exhibit instability in roadside scenarios. Existing reconstruction approaches fo

Cited by 0SourceScholar
2025

SPARK: Simulating the Co-evolution of Stance and Topic Dynamics in Online Discourse with LLM-based Agents

EMNLP 2025

Topic evolution and stance dynamics are deeply intertwined in online social media, shaping the fragmentation and polarization of public discourse. Yet existing dynamic topic models and stance analysis approaches usually consider these processes in isolation, relying on abstractions that lack interpr

2025

STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task planner

IROS 2025

The ability to perform reliable long-horizon task planning is crucial for deploying robots in real-world environments. However, directly employing Large Language Models (LLMs) as action sequence generators often results in low success rates due to their limited reasoning ability for long-horizon emb

Cited by 4SourceScholar
2025

Scene Map-based Prompt Tuning for Navigation Instruction Generation

CVPR 2025poster

Navigation instruction generation (NIG), which provides interactive feedback and guidance to humans along a trajectory, is vital for developing embodied agents capable of human-machine communication and collaboration through natural language. Early data-driven methods directly map sequences of past…

2025

Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation

CVPR 2025poster

Advances in talking-head animation based on Latent Diffusion Models (LDM) enable the creation of highly realistic, synchronized videos. These fabricated videos are indistinguishable from real ones, increasing the risk of potential misuse for scams, political manipulation, and misinformation. Hence,…

2025

Sparse Rewards Can Self-Train Dialogue Agents

ACL 2025finding

Recent advancements in state-of-the-art (SOTA) Large Language Model (LLM) agents, especially in multi-turn dialogue tasks, have been primarily driven by supervised fine-tuning and high-quality human feedback. However, as base LLM models continue to improve, acquiring meaningful human feedback has be…

2025

SparseDiT: Token Sparsification for Efficient Diffusion Transformer

NeurIPS 2025poster

Diffusion Transformers (DiT) are renowned for their impressive generative performance; however, they are significantly constrained by considerable computational costs due to the quadratic complexity in self-attention and the extensive sampling steps required. While advancements have been made in exp…

Cited by 0SourcecodeScholar
2025

TAPNext: Tracking Any Point (TAP) as Next Token Prediction

ICCV 2025poster

Tracking Any Point (TAP) in a video is a challenging computer vision problem with many demonstrated applications in robotics, video editing, and 3D reconstruction. Existing methods for TAP rely heavily on complex tracking-specific inductive biases and heuristics, limiting their generality and potent…

2025

TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models

IROS 2025

Reliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rely on low-level visual features and lack high-level textual information, which struggle with accurate segmentation when c

Cited by 1SourceScholar
2025

TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation

ICCV 2025poster

Video generation models are revolutionizing content creation, with image-to-video models drawing increasing attention due to their enhanced controllability, visual consistency, and practical applications. However, despite their popularity, these models rely on user-provided text and image prompts, a…

2025

Towards Human-like Virtual Beings: Simulating Human Behavior in 3D Scenes

ICCV 2025poster

Building autonomous agents that can replicate human behavior in the realistic 3D world is a key step toward artificial general intelligence. This requires agents to be holistic goal achievers and to naturally adapt to environmental dynamics. In this work, we introduce ACTOR, an agent capable of perf…

2025

Transformer-based Speech Model Learns Well as Infants and Encodes Abstractions through Exemplars in the Poverty of the Stimulus Environment

COLING 2025main

Infants are capable of learning language, predominantly through speech and associations, in impoverished environments—a phenomenon known as the Poverty of the Stimulus (POS). Is this ability uniquely human, as an innate linguistic predisposition, or can it be empirically learned through potential li…

2025

UDSH: An Unsupervised Deep Image Stitching and De-Occlusion Method for Heavy Occlusion Scene

IROS 2025

Image stitching in heavy occlusion scenarios faces the dual challenges of accurate alignment and occlusion removal. On one hand, occlusion causes the loss of key texture and structural information in the image. On the other hand, it affects the image’s integrity. Existing stitching methods perform w

Cited by 0SourceScholar
2025

UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis

ICCV 2025poster

Text-to-image generation has transformed content creation, yet precise visual text rendering remains challenging for generative models due to blurred glyphs, semantic inconsistencies, and limited style controllability. Current methods typically employ pre-rendered glyph images as conditional inputs,…

Cited by 0SourcePDFScholar
2025

Vehicle Drifting Planning and Control Framework for Flexible U-turns in Space-limited Environments

IROS 2025

Space-limited U-shape bend is a safety-critical scenario that requires the high maneuverability of vehicles. However, due to the non-holonomic nature of the vehicle, it is difficult to perform flexible U-turns without intricate adjustments, which is detrimental to the efficient execution of tasks. T

Cited by 0SourceScholar
2025

Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents

EMNLP 2025

Role-playing agents (RPAs) have attracted growing interest for their ability to simulate immersive and interactive characters. However, existing approaches primarily focus on static role profiles, overlooking the dynamic perceptual abilities inherent to humans. To bridge this gap, we introduce the c

Cited by 0SourcePDFScholar
2025

VideoGrain: Modulating Space-Time Attention for Multi-Grained Video Editing

ICLR 2025poster

Recent advancements in diffusion models have significantly improved video generation and editing capabilities. However, multi-grained video editing, which encompasses class-level, instance-level, and part-level modifications, remains a formidable challenge. The major difficulties in multi-grained ed…

2025

ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning

AAAI 2025technical

Zero-shot learning (ZSL) aims to recognize unseen classes by transferring semantic knowledge from seen classes to unseen ones, guided by semantic information. To this end, existing works have demonstrated remarkable performance by utilizing global visual features from Convolutional Neural Networks (…

2024

A Rigid-Flexible Coupling Oscillator for Pneumatic Autonomous Robots

RA-L 2024

The traditional oscillators are limited by their material characteristics and structural configuration, exhibiting a low oscillation frequency and low output power under pressure input. This results in a slow locomotion speed for the robot. To address these limitations, this work presents a compact,

Cited by 1SourceScholar
2024

Automated Tone Transcription and Clustering with Tone2Vec

EMNLP 2024finding

Lexical tones play a crucial role in Sino-Tibetan languages. However, current phonetic fieldwork relies on manual effort, resulting in substantial time and financial costs. This is especially challenging for the numerous endangered languages that are rapidly disappearing, often compounded by limited…

2024

Beyond Surface Similarity: Detecting Subtle Semantic Shifts in Financial Narratives

NAACL 2024findings

In this paper, we introduce the Financial-STS task, a financial domain-specific NLP task designed to measure the nuanced semantic similarity between pairs of financial narratives. These narratives originate from the financial statements of the same company but correspond to different periods, such a…

2024

CQIL: Inference Latency Optimization with Concurrent Computation of Quasi-Independent Layers

ACL 2024long

The fast-growing large scale language models are delivering unprecedented performance on almost all natural language processing tasks. However, the effectiveness of large language models are reliant on an exponentially increasing number of parameters. The overwhelming computation complexity incurs a…

2024

CapHuman: Capture Your Moments in Parallel Universes

CVPR 2024poster

We concentrate on a novel human-centric image synthesis task that is given only one reference facial photograph it is expected to generate specific individual images with diverse head positions poses facial expressions and illuminations in different contexts. To accomplish this goal we argue that ou…

2024

Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts

ECCV 2024poster

"Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for 3D scene editing still exhibit certain shortcomings, hindering their further interactive design. Such schemes typically…

2024

Clustering Propagation for Universal Medical Image Segmentation

CVPR 2024poster

Prominent solutions for medical image segmentation are typically tailored for automatic or interactive setups posing challenges in facilitating progress achieved in one task to another. This also necessitates separate models for each task duplicating both training time and parameters. To address abo…

2024

Connecting the Dots: Inferring Patent Phrase Similarity with Retrieved Phrase Graphs

NAACL 2024findings

We study the patent phrase similarity inference task, which measures the semantic similarity between two patent phrases. As patent documents employ legal and highly technical language, existing semantic textual similarity methods that use localized contextual information do not perform satisfactoril…

2024

Controllable Navigation Instruction Generation with Chain of Thought Prompting

ECCV 2024poster

"Instruction generation is a vital and multidisciplinary research area with broad applications. Existing instruction generation models are limited to generating instructions in a single style from a particular dataset, and the style and content of generated instructions cannot be controlled. Moreove…

2024

DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval

AAAI 2024technical

Text-video retrieval is a critical multi-modal task to find the most relevant video for a text query. Although pretrained models like CLIP have demonstrated impressive potential in this area, the rising cost of fully finetuning these models due to increasing model size continues to pose a problem. T…

2024

DRIP: Unleashing Diffusion Priors for Joint Foreground and Alpha Prediction in Image Matting

NeurIPS 2024poster

Recovering the foreground color and opacity/alpha matte from a single image (i.e., image matting) is a challenging and ill-posed problem where data priors play a critical role in achieving precise results. Traditional methods generally predict the alpha matte and then extract the foreground through…

Cited by 2SourcePDFScholar
2024

DSVT: Dynamic 3D Surround View for Tractor-Trailer Vehicles Based on Real-Time Pose Estimation with Drop Model

IROS 2024poster

In recent years, 3D surround view systems have attracted a lot of attention in the field of advanced driver assistance systems (ADAS). However, the foundational assumption of unchanging camera poses in traditional 3D surround view systems, which is designed for single-unit vehicles, results in a fai…

Cited by 4SourceScholar
2024

DataStealing: Steal Data from Diffusion Models in Federated Learning with Multiple Trojans

NeurIPS 2024poster

Federated Learning (FL) is commonly used to collaboratively train models with privacy preservation. In this paper, we found out that the popular diffusion models have introduced a new vulnerability to FL, which brings serious privacy threats. Despite stringent data management measures, attackers can…

2024

DeFlow: Decoder of Scene Flow Network in Autonomous Driving

ICRA 2024poster

Scene flow estimation determines a scene’s 3D motion field, by predicting the motion of points in the scene, especially for aiding tasks in autonomous driving. Many networks with large-scale point clouds as input use voxelization to create a pseudo-image for real-time running. However, the voxelizat…

Cited by 20SourcecodeScholar
2024

Deep SE(3)-Equivariant Geometric Reasoning for Precise Placement Tasks

ICLR 2024poster

Many robot manipulation tasks can be framed as geometric reasoning tasks, where an agent must be able to precisely manipulate an object into a position that satisfies the task from a set of initial conditions. Often, task success is defined based on the relationship between two objects - for instanc…

Cited by 14SourcePDFScholar
2024

Depth-Aware Blind Image Decomposition for Real-World Adverse Weather Recovery

ECCV 2024poster

"In this paper, we delve into Blind Image Decomposition (BID) tailored for real-world scenarios, aiming to uniformly recover images from diverse, unknown weather combinations and intensities. Our investigation uncovers one inherent gap between the controlled lab settings and the complex real-world e…

2024

DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

ICML 2024poster

Recent LLM-driven visual agents mainly focus on solving image-based tasks, which limits their ability to understand dynamic scenes, making it far from real-life applications like guiding students in laboratory experiments and identifying their mistakes. Hence, this paper explores DoraemonGPT, a comp…

2024

Enhancing Hallucination Detection through Perturbation-Based Synthetic Data Generation in System Responses

ACL 2024findings

Detecting hallucinations in large language model (LLM) outputs is pivotal, yet traditional fine-tuning for this classification task is impeded by the expensive and quickly outdated annotation process, especially across numerous vertical domains and in the face of rapid LLM advancements. In this stud…

2024

Entangled View-Epipolar Information Aggregation for Generalizable Neural Radiance Fields

CVPR 2024poster

Generalizable NeRF can directly synthesize novel views across new scenes eliminating the need for scene-specific retraining in vanilla NeRF. A critical enabling factor in these approaches is the extraction of a generalizable 3D representation by aggregating source-view features. In this paper we pro…

2024

Epipolar-Free 3D Gaussian Splatting for Generalizable Novel View Synthesis

NeurIPS 2024poster

Generalizable 3D Gaussian splitting (3DGS) can reconstruct new scenes from sparse-view observations in a feed-forward inference manner, eliminating the need for scene-specific retraining required in conventional 3DGS. However, existing methods rely heavily on epipolar priors, which can be unreliable…

Cited by 0SourcePDFScholar
2024

Exploring the Relationship between In-Context Learning and Instruction Tuning

EMNLP 2024finding

In-Context Learning (ICL) and Instruction Tuning (IT) are two primary paradigms of adopting Large Language Models (LLMs) to downstream applications. However, they are significantly different. In ICL, a set of demonstrations is provided at the inference time, but the LLM’s parameters are not updated.…

2024

Fast and Robust Point Cloud Registration with Tree-based Transformer

ICRA 2024poster

Point cloud registration is essential in computer vision and robotics. Recently, transformer-based methods have achieved advanced point cloud registration performance. However, the standard attention mechanism utilized in these methods considers many low-relevance points, and it has difficulty focus…

Cited by 2SourcecodeScholar
2024

Fine-tuning the Diffusion Model and Distilling Informative Priors for Sparse-view 3D Reconstruction

IROS 2024poster

3D reconstruction methods such as Neural Radiance Fields (NeRFs) are capable of optimizing high-quality 3D representation from images. However, NeRF is limited by the requirement for a large number of multi-view images, making its application to real-world scenarios challenging. In this work, we pro…

Cited by 0SourcecodeScholar
2024

FragRel: Exploiting Fragment-level Relations in the External Memory of Large Language Models

ACL 2024findings

To process contexts with unlimited length using Large Language Models (LLMs), recent studies explore hierarchically managing the long text. Only several text fragments are taken from the external memory and passed into the temporary working memory, i.e., LLM’s context window. However, existing appro…

Cited by 1SourcePDFScholar
2024

FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention

NeurIPS 2024poster

Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computational and data resources, posing a challenge to developing long video diffusion models. This paper investigates a strai…

Cited by 22SourcePDFScholar
2024

HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting

ECCV 2024poster

"Creating digital avatars from textual prompts has long been a desirable yet challenging task. Despite the promising results achieved with 2D diffusion priors, current methods struggle to create high-quality and consistent animated avatars efficiently. Previous animatable head models like FLAME have…

2024

Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models

NeurIPS 2024poster

Prevalent human-object interaction (HOI) detection approaches typically leverage large-scale visual-linguistic models to help recognize events involving humans and objects. Though promising, models trained via contrastive learning on text-image pairs often neglect mid/low-level visual cues and strug…

Cited by 7SourcePDFScholar
2024

Improving Bird's Eye View Semantic Segmentation by Task Decomposition

CVPR 2024poster

Semantic segmentation in bird's eye view (BEV) plays a crucial role in autonomous driving. Previous methods usually follow an end-to-end pipeline directly predicting the BEV segmentation map from monocular RGB inputs. However the challenge arises when the RGB inputs and BEV targets from distinct per…

2024

Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition Learning

ICML 2024poster

Previous efforts using frozen Large Language Models (LLMs) for visual understanding, via image captioning or image-text retrieval tasks, face challenges when dealing with complex multimodal scenarios. In order to enhance the capabilities of Multimodal Large Language Models (MLLM) in comprehending th…

2024

Interpretable3D: An Ad-Hoc Interpretable Classifier for 3D Point Clouds

AAAI 2024technical

3D decision-critical tasks urgently require research on explanations to ensure system reliability and transparency. Extensive explanatory research has been conducted on 2D images, but there is a lack in the 3D field. Furthermore, the existing explanations for 3D models are post-hoc and can be mislea…

2024

Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval

CVPR 2024poster

We study the zero-shot Composed Image Retrieval (ZS-CIR) task which is to retrieve the target image given a reference image and a description without training on the triplet datasets. Previous works generate pseudo-word tokens by projecting the reference image features to the text embedding space. H…

2024

LCP-Fusion: A Neural Implicit SLAM with Enhanced Local Constraints and Computable Prior

IROS 2024poster

Recently the dense Simultaneous Localization and Mapping (SLAM) based on neural implicit representation has shown impressive progress in hole filling and high-fidelity mapping. Nevertheless, existing methods either heavily rely on known scene bounds or suffer inconsistent reconstruction due to drift…

Cited by 0SourcecodeScholar
2024

LSK3DNet: Towards Effective and Efficient 3D Perception with Large Sparse Kernels

CVPR 2024poster

Autonomous systems need to process large-scale sparse and irregular point clouds with limited compute resources. Consequently it is essential to develop LiDAR perception methods that are both efficient and effective. Although naively enlarging 3D kernel size can enhance performance it will also lead…

2024

Learning from One Continuous Video Stream

CVPR 2024poster

We introduce a framework for online learning from a single continuous video stream - the way people and animals learn without mini-batches data augmentation or shuffling. This poses great challenges given the high correlation between consecutive video frames and there is very little prior work on it…

Cited by 3SourcePDFScholar
2024

MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis

CVPR 2024highlight

We present a Multi-Instance Generation (MIG) task simultaneously generating multiple instances with diverse controls in one image. Given a set of predefined coordinates and their corresponding descriptions the task is to ensure that generated instances are accurately at the designated locations and…

2024

MS-DETR: Efficient DETR Training with Mixed Supervision

CVPR 2024poster

DETR accomplishes end-to-end object detection through iteratively generating multiple object candidates based on image features and promoting one candidate for each ground-truth object. The traditional training procedure using one-to-one supervision in the original DETR lacks direct supervision for…

2024

MS2SL: Multimodal Spoken Data-Driven Continuous Sign Language Production

ACL 2024findings

Sign language understanding has made significant strides; however, there is still no viable solution for generating sign sequences directlyfrom entire spoken content, e.g., text or speech. In this paper, we propose a unified framework for continuous sign language production, easing communication bet…

Cited by 3SourcePDFScholar
2024

Moving Off-the-Grid: Scene-Grounded Video Representations

NeurIPS 2024spotlight

Current vision models typically maintain a fixed correspondence between their representation structure and image space. Each layer comprises a set of tokens arranged “on-the-grid,” which biases patches or tokens to encode information at a specific spatio(-temporal) location. In this work we present…

Cited by 2SourcePDFScholar
2024

Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion

ECCV 2024poster

"Visual acoustic matching (VAM) is pivotal for enhancing the immersive experience, and the task of dereverberation is effective in improving audio intelligibility. Existing methods treat each task independently, overlooking the inherent reciprocity between them. Moreover, these methods depend on pai…

Cited by 1SourcePDFScholar
2024

Navigation Instruction Generation with BEV Perception and Large Language Models

ECCV 2024poster

"Navigation instruction generation, which requires embodied agents to describe the navigation routes, has been of great interest in robotics and human-computer interaction. Existing studies directly map the sequence of 2D perspective observations to route descriptions. Though straightforward, they o…

2024

OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments

RA-L 2024

Environment representations endowed with sophisticated semantics are pivotal for facilitating seamless interaction between robots and humans, enabling them to effectively carry out various tasks. Open-vocabulary representation, powered by Visual-Language models (VLMs), possesses inherent advantages,

Cited by 38SourcecodeScholar