← Search

Rui Hu

37 accepted papers

2026

Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training

ICML 2026poster

Generative Flow Networks (GFlowNets) excel at sampling diverse, high-reward objects. In many practical applications where active reward queries are infeasible, these models must be trained using static offline datasets. Prevailing training methods typically rely on a proxy model to provide reward fe…

Cited by 0SourceScholar
2026

Finite-Time Convergence Analysis of ODE-based Generative Models for Stochastic Interpolants

ICLR 2026poster

Stochastic interpolants offer a robust framework for continuously transforming samples between arbitrary data distributions via ordinary or stochastic differential equations (ODEs/SDEs), holding significant promise for generative modeling. While previous studies have analyzed the finite-time converg…

Cited by 0SourceScholar
2026

From Contrast to Consistency: Rethinking Event-based Continuous-Time Optical Flow Estimation

CVPR 2026

Estimating continuous optical flow is a fundamental yet challenging problem in dynamic visual perception. Event-based cameras, with microsecond latency and high dynamic range, capture brightness changes asynchronously, offering a unique opportunity to model motion with fine temporal precision. Howev

Cited by 0SourceScholar
2026

LENS: Learning to Segment Anything with Unified Reinforced Reasoning

AAAI 2026technical

Text-prompted image segmentation enables fine-grained visual understanding and is critical for applications such as human-computer interaction and robotics. However, existing supervised fine-tuning methods typically ignore explicit chain-of-thought (CoT) reasoning at test time, which limits their ab

Cited by 0SourcePDFScholar
2026

Q Cache: Visual Attention Is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model

AAAI 2026technical

Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual tokens engenders a substantial computational load and key-value (KV) cache footprint bottleneck. Existing approaches focus

Cited by 0SourcePDFScholar
2026

WavefrontDiffusion: Dynamic Decoding Schedule for Improved Reasoning

ICLR 2026poster

Diffusion Language Models (DLMs) have shown strong potential for text generation and are becoming a competitive alternative to autoregressive models. The denoising strategy plays an important role in determining the quality of their outputs. Mainstream denoising strategies include Standard Diffusio…

Cited by 0SourceScholar
2025

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding

CVPR 2025poster

Visual Document Understanding has become essential with the increase of text-rich visual content. This field poses significant challenges due to the need for effective integration of visual perception and textual comprehension, particularly across diverse document types with complex layouts. Moreove…

2025

Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks

ICLR 2025spotlight

Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have found applications in various important tasks, attracting great research interest in their training algorithms. In general, GFlowNets are trained by fitting the for…

Cited by 1SourcePDFScholar
2025

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

CVPR 2025poster

The emergence and growing popularity of multimodal large language models (MLLMs) have significant potential to enhance various aspects of daily life, from improving communication to facilitating learning and problem-solving. Mobile phones, as essential daily companions, represent the most effective…

2025

Detecting Backdoor Attacks in Federated Learning via Direction Alignment Inspection

CVPR 2025highlight

The distributed nature of training makes Federated Learning (FL) vulnerable to backdoor attacks, where malicious model updates aim to compromise the global model's performance on specific tasks. Existing defense methods show limited efficacy as they overlook the inconsistency between benign and mali…

2025

FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models

NeurIPS 2025poster

Large Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raising concerns about data scarcity and the lack of access to domain-specific, sensitive information. Federated Learning (FL)…

Cited by 0SourceScholar
2025

GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding

ICCV 2025poster

Pixel grounding, encompassing tasks such as Referring Expression Segmentation (RES), has garnered considerable attention due to its potential for bridging the gap between vision and language modalities. However, advancements in this domain are currently constrained by limitations inherent in existin…

2025

Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models

ACL 2025finding

Omnimodal Large Language Models (OLLMs) have shown significant progress in integrating vision and text, but still struggle with integrating vision and audio, often exhibiting suboptimal performance when processing audio queries compared to text queries. This disparity is primarily due to insufficien…

2025

MEraser: An Effective Fingerprint Erasure Approach for Large Language Models

ACL 2025long

Large Language Models (LLMs) have become increasingly prevalent across various sectors, raising critical concerns about model ownership and intellectual property protection. Although backdoor-based fingerprinting has emerged as a promising solution for model authentication, effective attacks for rem…

2025

ST3: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming

AAAI 2025technical

Multimodal large language models (MLLMs) enhance their perceptual capabilities by integrating visual and textual information. However, processing the massive number of visual tokens incurs a significant computational cost. Existing analysis of the MLLM attention mechanisms remains shallow, leading t…

Cited by 2SourcePDFScholar
2025

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents

NeurIPS 2025poster

In this paper, we introduce UI-Genie, a self-improving framework addressing two key challenges in GUI agents: verification of trajectory outcome is challenging and high-quality training data are not scalable. These challenges are addressed by a reward model and a self-improving pipeline, respectivel…

Cited by 0SourcecodeScholar
2024

A Dynamic Calibration Framework for the Event-Frame Stereo Camera System

RA-L 2024

The fusion of event cameras and conventional frame cameras is a novel research field, and a stereo structure consisting of an event camera and a frame camera can incorporate the advantages of both. This letter develops a dynamic calibration framework for the event-frame stereo camera system. In this

Cited by 5SourceScholar
2024

Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion

ICLR 2024poster

Learning world models can teach an agent how the world works in an unsupervised manner. Even though it can be viewed as a special case of sequence modeling, progress for scaling world models on robotic applications such as autonomous driving has been somewhat less rapid than scaling language models…

Cited by 56SourcePDFScholar
2024

FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance

ECCV 2024poster

"CLIP has achieved impressive zero-shot performance after pretraining on a large-scale dataset consisting of paired image-text data. Previous works have utilized CLIP by incorporating manually designed visual prompts like colored circles and blur masks into the images to guide the model’s attention,…

2024

FashionR2R: Texture-preserving Rendered-to-Real Image Translation with Diffusion Models

NeurIPS 2024poster

Modeling and producing lifelike clothed human images has attracted researchers' attention from different areas for decades, with the complexity from highly articulated and structured content. Rendering algorithms decompose and simulate the imaging process of a camera, while are limited by the accura…

Cited by 1SourcePDFScholar
2023

Reconstructing Objects in-the-wild for Realistic Sensor Simulation

ICRA 2023poster

Reconstructing objects from real world data and rendering them at novel views is critical to bringing realism, diversity and scale to simulation for robotics training and testing. In this work, we present NeuSim, a novel approach that estimates accurate geometry and realistic appearance from sparse…

Cited by 17SourceScholar
2022

Hybrid Local SGD for Federated Learning with Heterogeneous Communications

ICLR 2022spotlight

Communication is a key bottleneck in federated learning where a large number of edge devices collaboratively learn a model under the orchestration of a central server without sharing their own training data. While local SGD has been proposed to reduce the number of FL rounds and become the algorithm…

Cited by 60SourcePDFScholar
2022

Rethinking Closed-Loop Training for Autonomous Driving

ECCV 2022poster

"Recent advances in high-fidelity simulators [22,82,44] have enabled closed-loop training of autonomous driving agents, potentially solving the distribution shift in training v.s. deployment and allowing training to be scaled both safely and cheaply. However, there is a lack of understanding of how…

2021

Federated Learning with Sparsification-Amplified Privacy and Adaptive Optimization

IJCAI 2021poster

Federated learning (FL) enables distributed agents to collaboratively learn a centralized model without sharing their raw data with each other. However, data locality does not provide sufficient privacy protection, and it is desirable to facilitate FL with rigorous differential privacy (DP) guarante…

Cited by 59SourcePDFScholar
2020

Conditional Entropy Coding for Efficient Video Compression

ECCV 2020poster

We propose a very simple and efficient video compression framework that only focuses on modeling the conditional entropy between frames. Unlike prior learning-based approaches, we reduce complexity by not performing any form of explicit transformations between frames and assume each frame is encoded…

Cited by 73SourcePDFScholar
2020

Learning Lane Graph Representations for Motion Forecasting

ECCV 2020poster

We propose a motion forecasting model that exploits a novel structured map representation as well as actor-map interactions. Instead of encoding vectorized maps as raster images, we construct a lane graph from raw map data to explicitly preserve the map structure. To capture the complex topology and…

2020

PnPNet: End-to-End Perception and Prediction With Tracking in the Loop

CVPR 2020poster

We tackle the problem of joint perception and motion forecasting in the context of self-driving vehicles. Towards this goal we propose PnPNet, an end-to-end model that takes as input sequential sensor data, and outputs at each time step object tracks and their future trajectories. The key component…

Cited by 221PDFScholar
2020

PolyTransform: Deep Polygon Transformer for Instance Segmentation

CVPR 2020poster

In this paper, we propose PolyTransform, a novel instance segmentation algorithm that produces precise, geometry-preserving masks by combining the strengths of prevailing segmentation approaches and modern polygon-based methods. In particular, we first exploit a segmentation network to generate inst…

Cited by 218PDFScholar
2019

DeepPruner: Learning Efficient Stereo Matching via Differentiable PatchMatch

ICCV 2019poster

Our goal is to significantly speed up the runtime of current state-of-the-art stereo algorithms to enable real-time inference. Towards this goal, we developed a differentiable PatchMatch module that allows us to discard most disparities without requiring full cost volume evaluation. We then exploit…

Cited by 342PDFScholar
2019

UPSNet: A Unified Panoptic Segmentation Network

CVPR 2019oral

In this paper, we propose a unified panoptic segmentation network (UPSNet) for tackling the newly proposed panoptic segmentation task. On top of a single backbone residual network, we first design a deformable convolution based semantic segmentation head and a Mask R-CNN style instance segmentation…

Cited by 548PDFcodeScholar