← Search

Chia-Wen Kuo

9 accepted papers

2025

D-Attn: Decomposed Attention for Large Vision-and-Language Model

ICCV 2025poster

Large vision-and-language models (LVLMs) have traditionally integrated visual and textual tokens by concatenating them into a single homogeneous input for large language models (LLMs), thereby maximally preserving the pre-trained language capabilities. However, this constrained architecture for visu…

2024

CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

NeurIPS 2024poster

Recent advancements in Multimodal Large Language Models (LLMs) have focused primarily on scaling by increasing text-image pair data and enhancing LLMs to improve performance on multimodal tasks. However, these scaling approaches are computationally expensive and overlook the significance of efficien…

2022

Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning

CVPR 2022poster

Significant progress has been made on visual captioning, largely relying on pre-trained features and later fixed object detectors that serve as rich inputs to auto-regressive models. A key limitation of such methods, however, is that the output of the model is conditioned only on the object detector…

Cited by 84PDFcodeScholar
2021

Unbiased Teacher for Semi-Supervised Object Detection

ICLR 2021poster

Semi-supervised learning, i.e., training networks with both labeled and unlabeled data, has made significant progress recently. However, existing works have primarily focused on image classification tasks and neglected object detection which requires more annotation effort. In this work, we revisit…

2020

FeatMatch: Feature-Based Augmentation for Semi-Supervised Learning

ECCV 2020poster

Recent state-of-the-art semi-supervised learning (SSL) methods use a combination of image-based transformations and consistency regularization as core components. Such methods, however, are limited to simple transformations such as traditional data augmentation or convex combinations of two images.…

Cited by 162SourcePDFScholar
2020

Who2com: Collaborative Perception via Learnable Handshake Communication

ICRA 2020poster

In this paper, we propose the problem of collaborative perception, where robots can combine their local observations with those of neighboring agents in a learnable way to improve accuracy on a perception task. Unlike existing work in robotics and multi-agent reinforcement learning, we formulate the…

Cited by 180SourceScholar
2019

Learning Pose-aware 3D Reconstruction via 2D-3D Self-consistency

ICASSP 2019accepted

3D reconstruction, inferring 3D shape information from a single 2D image, has drawn attention from learning and vision communities. In this paper, we propose a framework for learning pose-aware 3D shape reconstruction. Our proposed model learns deep representation for recovering the 3D object, with…

Cited by 0SourceScholar
2015

Model-based 3D object recognition and fetching by a 7-DoF robot with online obstacle avoidance for factory automation

ICRA 2015poster

The objective of this paper is to present the model-based 3D object recognition and fetching by a 7-DoF robot with online obstacle avoidance for factory automation. The robot can fetch the random type of 3D objects with arbitrary pose using a 3D visual camera. Object recognition pipeline based on di…

Cited by 18SourceScholar