Projects

Conference Papers


O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model

Published in 40th AAAI Conference on Artifical Intelligence, 2026

Tackled the gap in LVLMs’ ability to understand hand-drawn sketches by introducing the first large-scale dataset linking sketches, real images, and natural-language instructions, and an LVLM trained on it. Across tasks such as localization, counting, retrieval, and VQA, our model achieves state-of-the-art performance, significantly advancing sketch-based visual reasoning.

Download Paper | Download Bibtex

Other Projects


Conditional LiDAR Point Cloud Generation using Diffusion Models

This project explores diffusion models for generating LiDAR point clouds with conditional control over object positions, enabling scene manipulation and data augmentation. It is the first work to achieve controllable object manipulation in diffusion-based LiDAR generation

Few-Shot Learning for Intent Detection

Developed a few-shot episodic learning pipeline for intent detection on the banking77 and clinc150 datasets on PyTorch. Implemented a cross-encoder and parametric network for similarity scoring.