← Search

Kai Kang

16 accepted papers

2026

LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer

AAAI 2026technical

In controllable image synthesis, generating coherent and consistent images from multiple references with spatial layout awareness remains an open challenge. We propose LAMIC, a Layout-Aware Multi-Image Composition framework that, for the first time, extends single-reference diffusion models to multi

Cited by 0SourcePDFScholar
2025

Analysis and Calibration of Nonlinear Power Amplifiers in Wideband OFDM-Based LEO Satellite Communication System

ICASSP 2025accepted

Low earth orbit (LEO) satellite communication system is vital due to its global coverage and low latency. To meet higher data rates, orthogonal frequency division multiplexing (OFDM) technology is recommended for adoption. In this paper, we analyze the nonlinear behavior of high-power amplifier (HPA…

Cited by 0SourceScholar
2025

MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

ICCV 2025poster

Multimodal large language models (MLLMs) excel at 2D visual understanding but remain limited in their ability to reason about 3D space. In this work, we leverage large-scale high-quality 3D scene data with open-set annotations to introduce 1) a novel supervised fine-tuning dataset and 2) a new evalu…

2025

Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping

NeurIPS 2025poster

We revisit scene-level 3D object detection as the output of an object-centric framework capable of both localization and mapping using 3D oriented boxes as the underlying geometric primitive. While existing 3D object detection approaches operate globally and implicitly rely on the a priori existence…

Cited by 0SourceScholar
2024

Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable Avatars

NeurIPS 2024poster

To build photorealistic avatars that users can embody, human modelling must be complete (cover the full body), driveable (able to reproduce the current motion and appearance from the user), and generalizable (_i.e._, easily adaptable to novel identities). Towards these goals, _paired_ captures, that…

2024

LMDX: Language Model-based Document Information Extraction and Localization

ACL 2024findings

Large Language Models (LLM) have revolutionized Natural Language Processing (NLP), improving state-of-the-art and exhibiting emergent capabilities across various tasks. However, their application in extracting information from visually rich documents, which is at the core of many document processing…

2020

Greedy Hybrid Rate Adaptation in Dynamic Wireless Communication Environment

ICASSP 2020accepted

High data throughput is desired in the wireless communication system design. Rate adaptation is an efficient way to update the data rate in the dynamic wireless environment. Conventional rate adaptation algorithms rely on the feedback of acknowledgment/negative acknowledgment (ACK/NACK) messages or…

Cited by 0SourceScholar
2019

Online Learning for Computation Peer Offloading with Semi-bandit Feedback

ICASSP 2019accepted

Fog computing is emerging as a promising paradigm to perform distributed, low-latency computation. Efficient computation peer offloading is critical to fully utilize the computational resources in fog networks. In this paper, we consider computation peer offloading problem in a fog network with time…

Cited by 6SourceScholar
2018

Distributed Censoring with Energy Constraint in Wireless Sensor Networks

ICASSP 2018accepted

In wireless sensor networks (WSN s), energy is always precious for sensor nodes. To save energy, censoring is introduced to cut the total number of transmission by only transmitting informative data. This algorithm, however, ignores the energy consumption during the delivery of parameters, which can…

Cited by 0SourceScholar
2017

Object Detection in Videos With Tubelet Proposal Networks

CVPR 2017poster

Object detection in videos has drawn increasing attention recently with the introduction of the large-scale ImageNet VID dataset. Different from object detection in static images, temporal information in videos is vital for object detection. To fully utilize temporal information, state-of-the-art me…

Cited by 255PDFScholar
2016

Object Detection From Video Tubelets With Convolutional Neural Networks

CVPR 2016spotlight

Deep Convolution Neural Networks (CNNs) have shown impressive performance in various vision tasks such as image classification, object detection and semantic segmentation. For object detection, particularly in still images, the performance has been significantly increased last year thanks to powerfu…

Cited by 513PDFcodeScholar
2016

Slicing Convolutional Neural Network for Crowd Video Understanding

CVPR 2016spotlight

Learning and capturing both appearance and dynamic representations are pivotal for crowd video understanding. Convolutional Neural Networks (CNNs) have shown its remarkable potential in learning appearance representations from images. However, the learning of dynamic representation, and how it can b…

Cited by 104PDFScholar