FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding
Consistent offline folding for native low-bit VLA inference, from calibration to integer execution on desktop and edge GPUs.
Embodied AI · Robot learning
Research in efficient embodied intelligence.
Exploring how robots can learn, reason, and act reliably under limited computational resources.

I conduct my research under the guidance of Dr. An Thai Le. I graduated from the University of Science, VNU-HCM (HCMUS), where I was supervised by Assoc. Prof. Ly Quoc Ngoc.
I’m Hung Thinh Ho, a researcher focused on efficient embodied intelligence. My research lies at the intersection of robot learning, vision-language-action (VLA) models, and efficient machine learning. I’m interested in how robots can translate multimodal understanding into reliable actions under limited computational resources.
My current work investigates low-bit quantization and efficient inference for VLA models, with an emphasis on preserving policy behavior while reducing computational cost. I also explore world models and retrieval-based approaches to robotic execution, connecting structured representations of the environment with action generation.
My broader goal is to develop embodied AI that is both computationally efficient and robust in the physical world, combining algorithmic research with evaluation on real robotic systems.
Consistent offline folding for native low-bit VLA inference, from calibration to integer execution on desktop and edge GPUs.
A shared C++ runtime for VLA model execution across heterogeneous hardware, with no PyTorch dependency for inference.