O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
Published in 40th AAAI Conference on Artifical Intelligence, 2026
Tackled the gap in LVLMs’ ability to understand hand-drawn sketches by introducing the first large-scale dataset linking sketches, real images, and natural-language instructions, and an LVLM trained on it. Across tasks such as localization, counting, retrieval, and VQA, our model achieves state-of-the-art performance, significantly advancing sketch-based visual reasoning.