This blog post explores the development of a custom Vision LLM tailored for document processing at Grab, specifically addressing the challenges of accurate information extraction from various Southeast Asian document formats. It details the shortcomings of traditional Optical Character Recognition (OCR), the selection of a suitable base model (Qwen2-VL), and the extensive training and fine-tuning processes undertaken to create a lightweight model achieving high accuracy with low latency. Key improvements include better handling of non-Latin scripts, and the importance of quality data in model training is emphasized. Future plans include expanding capabilities and market reach.