Vision-Language-Action
Building compact VLA models that connect visual perception, language understanding, and physical action.
- Small VLA architecture
- Embodied intelligence
- Multimodal policies
Senior Edge AI Engineer
I’m Ho Thinh Hung, an Edge AI engineer building compact Vision-Language-Action models and the inference systems that make them useful in the physical world.
HO THINH HUNG
01 / MY POINT OF VIEW
I like the moment when a model leaves the lab and meets real constraints: limited memory, strict latency, noisy sensors, and a machine that has to make the right move.
That’s where I work—between model architecture and hardware, turning ambitious multimodal ideas into dependable edge systems.
02 / WHAT I DO
Three areas where I spend most of my time.
Building compact VLA models that connect visual perception, language understanding, and physical action.
Making capable models fit real memory, power, and latency budgets without losing what makes them useful.
Designing reliable, observable systems for real-time computer vision and high-throughput inference.
03 / SELECTED WORK
Reproducible FP8, NVFP4, AWQ, and TensorRT workflows for LLM and speech models—from preparation to deployment-oriented validation.
Hands-on notes and benchmarks for X-VLA, SmolVLA, quantization, pruning, and edge deployment.
FP16 → INT8 → INT4Controlled experiments separated Myelin and CASK fusion failures from lossy AWQ repacking and a silent InternVL3 exporter failure.
04 / PRECISION LAB
Real measurements from my ModelOpt recipes—including BF16, FP8, NVFP4, INT8 SmoothQuant, and INT4 AWQ.
Balanced throughput and model footprint with a small WER change.
05 / LEARNING IN PUBLIC
Paper notes and first-hand engineering experiments, rewritten as practical English articles about the decisions behind fast, reliable edge AI.
Controlled experiments that isolated four garbled-generation failures across FP16, NVFP4, AWQ, and InternVL3—with NVIDIA maintainer confirmation.
Inference performance across H100, RTX, and Jetson—covering model bottlenecks, deployment splits, and action generation strategies.
Activation-aware weight quantization, salient channels, and why TinyChat matters for real runtime gains.
Read article ↗Moving activation difficulty into weights for efficient, accurate post-training quantization of large models.
Read article ↗06 / TOOLKIT
Python · C++ · CUDA
TensorRT · Triton · vLLM · ONNX
DeepStream · YOLO · OpenCV · VLMs
Docker · Kafka · Redis · WebRTC