← Search

Yifan liu

64 accepted papers

2026

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models

CVPR 2026

Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and embodied agents. We attribute this to a modality gap between the 3D tasks and the 2D training of VLM, which led to ineffi

Cited by 0SourceScholar
2026

Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented Reasoning

AAAI 2026technical

Understanding and adhering to traffic regulations is essential for autonomous vehicles to ensure safety and trustworthiness. However, traffic regulations are complex, context-dependent, and differ between regions, posing a major challenge to conventional rule-based decision-making approaches. We pre

Cited by 0SourcePDFScholar
2026

LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic Pyramid

AAAI 2026technical

Vision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-grained visual perception. We attribute this to the inner limitations of ViTs in capturing diverse visual semantic level

Cited by 0SourcePDFScholar
2026

Learning Goal-Directed Rolling: Spherical Robot Point-to-Point Control Through Reinforcement Learning

RA-L 2026

Point-to-point navigation is an important ability for spherical robots. Traditional methods usually use a planner and a tracker for short-range target control. However, this hierarchical method suffers from a mismatch issue. In this work, we propose an end-to-end controller based on reinforcement le

Cited by 0SourceScholar
2026

RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment Graph

CVPR 2026

Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on top of 2D visual backbones and depend heavily on labeled data for training, which is often scarce in real-world scenarios, causing a sim-to-real gap. Moreover

Cited by 0SourceScholar
2026

TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control

ICML 2026spotlight

Large Language Models (LLMs) training is prohibitively expensive, driving interest in low-precision fully-quantized training (FQT). While novel 4-bit formats like NVFP4 offer substantial efficiency gains, achieving near-lossless training at such low precision remains challenging. We introduce **Tetr…

Cited by 0SourceScholar
2026

WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting

ICML 2026poster

We present WorldMirror, a unified feed-forward model for comprehensive 3D geometric prediction tasks. Unlike existing methods constrained to image-only inputs or customized for a specific task, our framework flexibly integrates diverse geometric priors, including camera poses, intrinsics, and depth …

Cited by 0SourceScholar
2025

Anti-Tamper Protection for Unauthorized Individual Image Generation

ICCV 2025poster

With the advancement of personalized image generation technologies, concerns about forgery attacks that infringe on portrait rights and privacy are growing. To address these concerns, protection perturbation algorithms have been developed to disrupt forgery generation. However, the protection algori…

2025

AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity

EMNLP 2025

Recent advancements in multimodal large language models (MLLMs) have garnered significant attention, offering a promising pathway toward artificial general intelligence (AGI). Among the essential capabilities required for AGI, creativity has emerged as a critical trait for MLLMs, with association se

Cited by 0SourcePDFScholar
2025

ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian Splatting

ICASSP 2025accepted

As 3D Gaussian Splatting (3D-GS) emerges as a promising technique for 3D reconstruction and novel view synthesis, offering superior rendering quality and efficiency, it becomes crucial to ensure secure transmission and copyright protection of 3D assets in anticipation of widespread distribution. Whi…

Cited by 0SourceScholar
2025

DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility

NAACL 2025findings

Video-to-speech (V2S) synthesis, the task of generating speech directly from silent video input, is inherently more challenging than other speech synthesis tasks due to the need to accurately reconstruct both speech content and speaker characteristics from visual cues alone. Recently, audio-visual p…

2025

Enhancing Speech Emotion Recognition with Speech Dynamic Modeling and Multi-Modal Knowledge Distillation

ICASSP 2025accepted

Complementary semantic information from the text modality, obtained through runtime transcription, plays a crucial role in Speech Emotion Recognition (SER). However, it introduces additional computational overhead and potential errors. To address these issues, we propose the SDMMKD framework, which…

Cited by 0SourceScholar
2025

GaussianReg: Rapid 2D/3D Registration for Emergency Surgery via Explicit 3D Modeling with Gaussian Primitives

ICCV 2025poster

Intraoperative 2D/3D registration, which aligns preoperative CT scans with intraoperative X-ray images, is critical for surgical navigation. However, existing methods require extensive preoperative training (several hours), making them unsuitable for emergency surgeries where minutes significantly i…

2025

Hide-in-Motion: Embedding Steganographic Copyright Information into 4D Gaussian Splatting Assets

ICRA 2025

As 4D extensions of 3D Gaussian Splatting (4D-GS) emerge as groundbreaking techniques for dynamic scene reconstruction and novel view synthesis in robotics and computer vision, ensuring the security and trustworthiness of these assets becomes crucial. While steganography has advanced significantly i

Cited by 9SourcecodeScholar
2025

InfoBridge: Balanced Multimodal Integration through Conditional Dependency Modeling

ICCV 2025poster

Developing systems that interpret diverse real-world signals remains a fundamental challenge in multimodal learning. Current approaches face significant obstacles from inherent modal heterogeneity. While existing methods attempt to enhance fusion through cross-modal alignment or interaction mechanis…

2025

InstantSplamp: Fast and Generalizable Stenography Framework for Generative Gaussian Splatting

ICLR 2025poster

With the rapid development of large generative models for 3D, especially the evolution from NeRF representations to more efficient Gaussian Splatting, the synthesis of 3D assets has become increasingly fast and efficient, enabling the large-scale publication and sharing of generated 3D objects. Howe…

2025

Learning Efficient and Generalizable Human Representation with Human Gaussian Model

ICCV 2025poster

Modeling animatable human avatars from videos is a long-standing and challenging problem. While conventional methods require per-instance optimization, recent feed-forward methods have been proposed to generate 3D Gaussians with a learnable network.However, these methods predict independent Gaussian…

2025

MA-DPR: Manifold-aware Distance Metrics for Dense Passage Retrieval

EMNLP 2025

Dense Passage Retrieval (DPR) typically relies on Euclidean or cosine distance to measure query–passage relevance in embedding space, which is effective when embeddings lie on a linear manifold. However, our experiments across DPR benchmarks suggest that embeddings often lie on lower-dimensional, no

Cited by 0SourcePDFScholar
2025

MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models

CVPR 2025poster

Recent advances in generalizable 3D Gaussian Splatting have demonstrated promising results in real-time high-fidelity rendering without per-scene optimization, yet existing approaches still struggle to handle unfamiliar visual content during inference on novel scenes due to limited generalizability.…

2025

PDCE: Patch-wise Dynamic Curve Estimation for Low-Light Image Enhancement

ICASSP 2025accepted

Low-light image enhancement (LLIE) can be reformulated as an image-specific curve estimation (CE) problem. Traditional CE-based methods struggle with issues such as uniform processing across different regions, static parameter estimation, and lack of effective global semantic enhancement. To address…

Cited by 0SourceScholar
2025

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

ICCV 2025poster

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and ignore that AI assistants perceive and act based on first-per…

2025

Socially-Aware Robot Navigation Enhanced by Bidirectional Natural Language Conversations Using Large Language Models

IROS 2025

Robotic navigation plays a pivotal role in a wide range of real-world applications. While traditional navigation systems focus on efficiency and obstacle avoidance, their inability to model complex human behaviors in shared spaces has underscored the growing need for socially aware navigation. In th

Cited by 6SourcecodeScholar
2025

U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation

AAAI 2025technical

U-Net has become a cornerstone in various visual applications such as image segmentation and diffusion probability models. While numerous innovative designs and improvements have been introduced by incorporating transformers or MLPs, the networks are still limited to linearly modeling patterns as we…

2024

Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask Completion

AAAI 2024technical

Amodal scene analysis entails interpreting the occlusion relationship among scene elements and inferring the possible shapes of the invisible parts. Existing methods typically frame this task as an extended instance segmentation or a pair-wise object de-occlusion problem. In this work, we propose a…

2024

Entwined Inversion: Tune-Free Inversion For Real Image Faithful Reconstruction and Editing

ICASSP 2024accepted

Text-conditional image editing is a very practical AIGC task that has recently emerged with great commercial and academic research value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing, but DDIM Inversion often results in reconstructio…

Cited by 0SourceScholar
2024

GTP-4o: Modality-prompted Heterogeneous Graph Learning for Omni-modal Biomedical Representation

ECCV 2024poster

"Recent advances in learning multi-modal representation have witnessed the success in biomedical domains. While established techniques enable handling multi-modal information, the challenges are posed when extended to various clinical modalities and practical modality-missing setting due to the inhe…

2024

HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects

ECCV 2024poster

"Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of multiple objects. Thus, we propose HIMO, a large-scale MoCap da…

Cited by 3SourcePDFScholar
2024

High-Fidelity 3D Head Avatars Reconstruction through Spatially-Varying Expression Conditioned Neural Radiance Field

AAAI 2024technical

One crucial aspect of 3D head avatar reconstruction lies in the details of facial expressions. Although recent NeRF-based photo-realistic 3D head avatar methods achieve high-quality avatar rendering, they still encounter challenges retaining intricate facial expression details because they overlook…

2024

ICGNet: A Unified Approach for Instance-Centric Grasping

ICRA 2024poster

Accurate grasping is the key to several robotic tasks including assembly and household robotics. Executing a successful grasp in a cluttered environment requires multiple levels of scene understanding: First, the robot needs to analyze the geometric properties of individual objects to find feasible…

Cited by 13SourcecodeScholar
2024

Inter-X: Towards Versatile Human-Human Interaction Analysis

CVPR 2024poster

The analysis of the ubiquitous human-human interactions is pivotal for understanding humans as social beings. Existing human-human interaction datasets typically suffer from inaccurate body motions lack of hand gestures and fine-grained textual descriptions. To better perceive and generate human-hum…

2024

MirageRoom: 3D Scene Segmentation with 2D Pre-trained Models by Mirage Projection

CVPR 2024highlight

Nowadays leveraging 2D images and pre-trained models to guide 3D point cloud feature representation has shown a remarkable potential to boost the performance of 3D fundamental models. While some works rely on additional data such as 2D real-world images and their corresponding camera poses recent st…

Cited by 2SourcePDFScholar
2024

Retrieval Augmented Fact Verification by Synthesizing Contrastive Arguments

ACL 2024long

The rapid propagation of misinformation poses substantial risks to public interest. To combat misinformation, large language models (LLMs) are adapted to automatically verify claim credibility. Nevertheless, existing methods heavily rely on the embedded knowledge within LLMs and / or black-box APIs…

2024

Robust Lightweight Depth Estimation Model via Data-Free Distillation

ICASSP 2024accepted

Existing Monocular Depth Estimation (MDE) methods often use large and complex neural networks. Despite the advanced performance of these methods, we consider the efficiency and generalization for practical applications with limited resources. In our paper, we present an efficient transformer-based m…

Cited by 0SourceScholar
2024

SOMTP: A Self-Supervised Learning-Based Optimizer for MPC-Based Safe Trajectory Planning Problems in Robotics

RA-L 2024

Model Predictive Control (MPC)-based trajectory planning has been widely used in robotics, and incorporating Control Barrier Function (CBF) constraints into MPC can greatly improve its obstacle avoidance efficiency. Unfortunately, traditional optimizers are resource-consuming and slow to solve such

Cited by 4SourceScholar
2023

3DPPE: 3D Point Positional Encoding for Transformer-based Multi-Camera 3D Object Detection

ICCV 2023poster

Transformer-based methods have swept the benchmarks on 2D and 3D detection on images. Because tokenization before the attention mechanism drops the spatial information, positional encoding becomes critical for those methods. Recent works found that encodings based on samples of the 3D viewing rays c…

Cited by 25PDFcodeScholar
2023

CTVIS: Consistent Training for Online Video Instance Segmentation

ICCV 2023poster

The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/nega…

Cited by 46PDFcodeScholar
2023

Dynamic Token Pruning in Plain Vision Transformers for Semantic Segmentation

ICCV 2023poster

Vision transformers have achieved leading performance on various visual tasks yet still suffer from high computational complexity. The situation deteriorates in dense prediction tasks like semantic segmentation, as high-resolution inputs and outputs usually imply more tokens involved in computations…

Cited by 28PDFcodeScholar
2023

Prior-Enhanced Temporal Action Localization Using Subject-Aware Spatial Attention

ICASSP 2023accepted

Temporal action localization (TAL) aims to detect the boundary and identify the class of each action instance in a long untrimmed video. Current approaches treat video frames homogeneously, and tend to give background and key objects excessive attention. This limits their sensitivity to localize act…

Cited by 0SourceScholar
2023

QuantSR: Accurate Low-bit Quantization for Efficient Image Super-Resolution

NeurIPS 2023spotlight

Low-bit quantization in image super-resolution (SR) has attracted copious attention in recent research due to its ability to reduce parameters and operations significantly. However, many quantized SR models suffer from accuracy degradation compared to their full-precision counterparts, especially at…

2023

SegPrompt: Boosting Open-World Segmentation via Category-Level Prompt Learning

ICCV 2023poster

Current closed-set instance segmentation models rely on predefined class labels for each mask during training and evaluation, limiting their ability to detect novel objects. Open-world instance segmentation (OWIS) models address this challenge by detecting unknown objects in a class-agnostic manner.…

Cited by 21PDFcodeScholar
2023

Segment Anything in High Quality

NeurIPS 2023poster

The recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction quality falls short in many cases, particularly when dealing with…

2023

The Effects of Robot Motion on Comfort Dynamics of Novice Users in Close-Proximity Human-Robot Interaction

IROS 2023poster

Effective and fluent close-proximity human-robot interaction requires understanding how humans get habituated to robots and how robot motion affects human comfort. While prior work has identified humans' preferences over robot motion characteristics and studied their influence on comfort, we are yet…

Cited by 3SourceScholar
2023

Towards Hierarchical Policy Learning for Conversational Recommendation with Hypergraph-based Reinforcement Learning

IJCAI 2023poster

Conversational recommendation systems (CRS) aim to timely and proactively acquire user dynamic preferred attributes through conversations for item recommendation. In each turn of CRS, there naturally have two decision-making processes with different roles that influence each other: 1) director, whic…

2023

Video Task Decathlon: Unifying Image and Video Tasks in Autonomous Driving

ICCV 2023poster

Performing multiple heterogeneous visual tasks in dynamic scenes is a hallmark of human perception capability. Despite remarkable progress in image and video recognition via representation learning, current research still focuses on designing specialized networks for singular, homogeneous, or simple…

Cited by 7PDFScholar
2023

ZegCLIP: Towards Adapting CLIP for Zero-Shot Semantic Segmentation

CVPR 2023poster

Recently, CLIP has been applied to pixel-level zero-shot learning tasks via a wo-stage scheme. The general idea is to first generate class-agnostic region proposals and then feed the cropped proposal regions to CLIP to utilize its image-level zero-shot classification capability. While effective, suc…

2022

Controllable Shadow Generation Using Pixel Height Maps

ECCV 2022poster

"Shadows are essential for realistic image compositing. Physics based shadow rendering methods require 3D geometries, which are not always available. Deep learning-based shadow synthesis methods learn a mapping from the light information to an object’s shadow without explicitly modeling the shadow g…

Cited by 30SourcePDFScholar
2022

Direction and Trajectory Tracking Control for Nonholonomic Spherical Robot by Combining Sliding Mode Controller and Model Prediction Controller

RA-L 2022

A spherical robot is a nonlinear, nonholonomic, and unstable system which increases the difficulty of the direction and trajectory tracking problem. In this study, we propose a new direction controller Hierarchical Terminal Sliding Mode Controller (HTSMC), an instruction planning controller called M

Cited by 33SourceScholar
2022

Multi-Terrain Velocity Control of the Spherical Robot by Online Obtaining the Uncertainties in the Dynamics

RA-L 2022

One controller cannot work on multiple and unknown terrains in the velocity control of the spherical robot, because the dynamic models of the robot vary on different terrains, and unmodeled dynamics and uncertainties exist in estimated dynamic models. Based on the above problem, a new velocity contr

Cited by 24SourceScholar
2022

SegViT: Semantic Segmentation with Plain Vision Transformers

NeurIPS 2022accept

We explore the capability of plain Vision Transformers (ViTs) for semantic segmentation and propose the SegViT. Previous ViT-based segmentation networks usually learn a pixel-level representation from the output of the ViT. Differently, we make use of the fundamental component—attention mechanism, t…

2021

A Simple Baseline for Semi-Supervised Semantic Segmentation With Strong Data Augmentation

ICCV 2021poster

Recently, significant progress has been made on semantic segmentation. However, the success of supervised semantic segmentation typically relies on a large amount of labeled data, which is time-consuming and costly to obtain. Inspired by the success of semi-supervised learning methods in image class…

Cited by 150PDFcodeScholar
2021

Channel-Wise Knowledge Distillation for Dense Prediction

ICCV 2021poster

Knowledge distillation (KD) has been proven a simple and effective tool for training compact dense prediction models. Lightweight student networks are trained by extra supervision transferred from large teacher networks. Most previous KD variants for dense prediction tasks align the activation maps…

Cited by 366PDFcodeScholar
2021

Dynamic Neural Representational Decoders for High-Resolution Semantic Segmentation

NeurIPS 2021poster

Semantic segmentation requires per-pixel prediction for a given image. Typically, the output resolution of a segmentation network is severely reduced due to the downsampling operations in the CNN backbone. Most previous methods employ upsampling decoders to recover the spatial resolution. Various de…

Cited by 16SourcePDFScholar
2021

Fuzzy PID Controller Based on Yaw Angle Prediction of a Spherical Robot

IROS 2021poster

In this paper, a fuzzy PID controller based on yaw angle prediction is applied to design an attitude controller for a spherical rolling robot. The robot consists of a 2-DOF pendulum located inside a spherical shell with freedom to rotate about the transversal and longitudinal axis. The proposed cont…

Cited by 24SourceScholar
2020

Efficient Semantic Video Segmentation with Per-frame Inference

ECCV 2020poster

For semantic segmentation, most existing real-time deep mod-els trained with each frame independently may produce inconsistent results when tested on a video sequence. A few methods take the correlations in the video sequence into account, e.g., by propagating the results to the neighboring frames u…

Cited by 172SourcePDFScholar
2020

Instance-Aware Embedding for Point Cloud Instance Segmentation

ECCV 2020poster

Although recent works have made significant progress in encoding meaningful context information for instance segmentation in 2D images, the works for 3D point cloud counterpart lag far behind. Conventional methods use radius search or other similar methods for aggregating local information. However,…

Cited by 24SourcePDFScholar
2019

Pixel Level Data Augmentation for Semantic Image Segmentation Using Generative Adversarial Networks

ICASSP 2019accepted

Semantic segmentation is one of the basic topics in computer vision, it aims to assign semantic labels to every pixel of an image. Unbalanced semantic label distribution could have a negative influence on segmentation accuracy. In this paper, we investigate using data augmentation approach to balanc…

Cited by 0SourceScholar
2019

Structured Knowledge Distillation for Semantic Segmentation

CVPR 2019oral

In this paper, we investigate the issue of knowledge distillation for training compact semantic segmentation networks by making use of cumbersome networks. We start from the straightforward scheme, pixel-wise distillation, which applies the distillation scheme originally introduced for image classif…

Cited by 929PDFScholar
2017

Multi-task deep neural network with shared hidden layers: Breaking down the wall between emotion representations

ICASSP 2017accepted

Emotion representations are psychological constructs for modelling, analysing, and recognising emotion, being one essential element of affect. Due to its complexity, the boundaries between different emotion concepts are often fuzzy, which is also reflected in the diversification of emotion databases…

Cited by 0SourceScholar