← Search

Tao Hu

33 accepted papers

2026

BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning

AAAI 2026technical

Class-Incremental Learning (CIL) aims to continually learn new classes without forgetting previously acquired knowledge. Vision-language models such as CLIP offer strong transferable representations via multi-modal supervision, making them a promising choice for CIL. However, applying CLIP to CIL po

Cited by 0SourcePDFScholar
2026

MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second

CVPR 2026

We present MoVieS, a Motion-aware View Synthesis model that reconstructs 4D dynamic scenes from monocular videos in one second. It represents dynamic 3D scenes with pixel-aligned Gaussian primitives and explicitly supervises their time-varying motions. This allows, for the first time, the unified mo

Cited by 0SourcecodeScholar
2026

NSF-HRPT: Neural Semantic Field Meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment

ICRA 2026poster

The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research has made progress in collision prediction, accurately quantifying risk levels from monocular vision inputs remains challenging due to the complex dyna…

Cited by 0Scholar
2026

SenseSearch: Empowering Vision-Language Models with High-Resolution Agentic Search-Reasoning via Reinforcement Learning

CVPR 2026

Vision-Language Models (VLMs) are limited by static knowledge and insufficient fine-grained visual analysis, hindering their performance on knowledge-intensive and visually complex tasks. While recent research has explored VLMs that employ external tools like search or cropping to enhance model perf

Cited by 0SourcecodeScholar
2025

Kinematic Model and Trajectory Tracking Algorithm for High-Speed Spherical Robots

IROS 2025

This paper proposes a new turning theory for spherical robots, which better describes the turning mechanism of spherical robots under turning constraints, using a pendulum-driven spherical robot as an example. Compared to the previous turning theory, the new theory shows greater alignment with real-

Cited by 0SourceScholar
2025

Low-Light Video Enhancement via Spatial-Temporal Consistent Decomposition

IJCAI 2025

Low-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition strategy that incorporates view-independent and view-dependent components to enhance the performance of LLVE. We leverage

Cited by 0SourcePDFScholar
2025

PRM: Photometric Stereo based Large Reconstruction Model

ICCV 2025poster

We propose PRM, a novel photometric stereo based large reconstruction model to reconstruct high-quality meshes with fine-grained details. Previous large reconstruction models typically prepare training images under fixed and simple lighting, offering minimal photometric cues for precise reconstructi…

Cited by 0SourcePDFScholar
2025

Seg-diffusion: Text-to-Image Diffusion Model for Open-Vocabulary Semantic Segmentation

ICASSP 2025accepted

Open-vocabulary semantic segmentation (OVSS) is a challenging computer vision task that labels each pixel within an image based on text descriptions. Recent advancements in OVSS are largely attributed to the increased model capacity. However, these models often struggle with unfamiliar images or uns…

Cited by 0SourceScholar
2024

EiffHDR: An Efficient Network for Multi-Exposure High Dynamic Range Imaging

ICASSP 2024accepted

While recent progress in Multi-exposure HDR imaging is promising, the growing complexity of state-of-the-art (SOTA) methods poses challenges for their analysis and comparison. In this paper, we analyze the motivations and approaches behind previous SOTA works and introduce EiffHDR, an efficient Mult…

Cited by 0SourceScholar
2024

Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier Domain

AAAI 2024technical

RAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area of research. However, current methods often ignore the difference between cell phone RAW images and DSLR camera RGB image…

2023

NeRFLix: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-Viewpoint MiXer

CVPR 2023poster

Neural radiance fields(NeRF) show great success in novel-view synthesis. However, in real-world scenes, recovering high-quality details from the source images is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation…

2023

Point2Pix: Photo-Realistic Point Cloud Rendering via Neural Radiance Fields

CVPR 2023poster

Synthesizing photo-realistic images from a point cloud is challenging because of the sparsity of point cloud representation. Recent Neural Radiance Fields and extensions are proposed to synthesize realistic images from 2D input. In this paper, we present Point2Pix as a novel point renderer to link t…

Cited by 20SourcePDFScholar
2023

Ref-NeuS: Ambiguity-Reduced Neural Implicit Surface Learning for Multi-View Reconstruction with Reflection

ICCV 2023oral

Neural implicit surface learning has shown significant progress in multi-view 3D reconstruction, where an object is represented by multilayer perceptrons that provide continuous implicit surface representation and view-dependent radiance. However, current methods often fail to accurately reconstruct…

Cited by 62PDFcodeScholar
2022

Direction and Trajectory Tracking Control for Nonholonomic Spherical Robot by Combining Sliding Mode Controller and Model Prediction Controller

RA-L 2022

A spherical robot is a nonlinear, nonholonomic, and unstable system which increases the difficulty of the direction and trajectory tracking problem. In this study, we propose a new direction controller Hierarchical Terminal Sliding Mode Controller (HTSMC), an instruction planning controller called M

Cited by 33SourceScholar
2022

Enhancing Unsupervised Domain Adaptation via Semantic Similarity Constraint for Medical Image Segmentation

IJCAI 2022poster

This work proposes a novel unsupervised cross-modality adaptive segmentation method for medical images to tackle the performance degradation caused by the severe domain shift when neural networks are being deployed to unseen modalities. The proposed method is an end-2-end framework, which conducts a…

Cited by 6SourcePDFScholar
2022

Human-Object Interaction Detection via Disentangled Transformer

CVPR 2022poster

Human-Object Interaction Detection tackles the problem of joint localization and classification of human object interactions. Existing HOI transformers either adopt a single decoder for triplet prediction, or utilize two parallel decoders to detect individual objects and interactions separately, and…

Cited by 77PDFScholar
2022

MixFormer: Mixing Features Across Windows and Dimensions

CVPR 2022oral

While local-window self-attention performs notably in vision tasks, it suffers from limited receptive field and weak modeling capability issues. This is mainly because it performs self-attention within non-overlapped windows and shares weights on the channel dimension. We propose MixFormer to find a…

Cited by 161PDFcodeScholar
2022

Multi-Terrain Velocity Control of the Spherical Robot by Online Obtaining the Uncertainties in the Dynamics

RA-L 2022

One controller cannot work on multiple and unknown terrains in the velocity control of the spherical robot, because the dynamic models of the robot vary on different terrains, and unmodeled dynamics and uncertainties exist in estimated dynamic models. Based on the above problem, a new velocity contr

Cited by 24SourceScholar
2021

EgoRenderer: Rendering Human Avatars From Egocentric Camera Images

ICCV 2021poster

We present EgoRenderer, a system for rendering full-body neural avatars of a person captured by a wearable, egocentric fisheye camera that is mounted on a cap or a VR headset. Our system renders photorealistic novel views of the actor and her motion from arbitrary virtual camera locations. Rendering…

Cited by 16PDFScholar
2021

Fuzzy PID Controller Based on Yaw Angle Prediction of a Spherical Robot

IROS 2021poster

In this paper, a fuzzy PID controller based on yaw angle prediction is applied to design an attitude controller for a spherical rolling robot. The robot consists of a 2-DOF pendulum located inside a spherical shell with freedom to rotate about the transversal and longitudinal axis. The proposed cont…

Cited by 24SourceScholar
2015

Online computation of sparse representations of time varying stimuli using a biologically motivated neural network

ICASSP 2015accepted

Natural stimuli are highly redundant, possessing significant spatial and temporal correlations. While sparse coding has been proposed as an efficient strategy employed by neural systems to encode sensory stimuli, the underlying mechanisms are still not well understood. Most previous approaches model…

Cited by 0SourceScholar