M.Tech Research Thesis {Colloquium}: 202: CDS: 23, July 2026 “Advancing Visual Reasoning Capabilities of Multimodal LLMs Beyond Natural Images”

When

23 Jul 26    
11:30 AM - 12:30 PM

Event Type

DEPARTMENT OF COMPUTATIONAL AND DATA SCIENCES
M.Tech Research Thesis {Colloquium}


Speaker: Mr. Rishi Gupta
S.R. Number: 06-18-01-10-22-24-1-24504
Title: “Advancing Visual Reasoning Capabilities of Multimodal LLMs Beyond Natural Images”
Research Supervisor: Prof. Anirban Chakraborty
Date & Time : July 23, 2026 (Thursday), 11:30 AM
Venue : CDS #202


ABSTRACT
Artificial Intelligence is undergoing a paradigm shift with the emergence of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Although many of these models demonstrate remarkable capabilities on natural image understanding tasks (e.g., Liu et al., NeurIPS 2023; Deitke et al., CVPR 2025), they remain fundamentally grounded in dense, photorealistic pixel representations. Human visual reasoning, however, extends beyond natural images and routinely relies on structured visual modalities such as sketches and explicit 3D geometry of objects. These image-adjacent representations preserve semantic and geometric structure while discarding much of the textural information present in natural images, making them intuitive for human communication and creative workflows. Therefore, enabling MLLMs to reason with such structured modalities remains an important step toward more flexible and holistic visual intelligence. This thesis explores how MLLMs can comprehend and generate visual modalities beyond dense natural images while preserving intuitive interaction through text, hand-drawn sketches, and natural images.

The thesis explores this problem from two complementary axes: understanding sparse visual inputs and generating semantically interpretable geometric representations. Firstly, we focus on sketch understanding, as sketches provide an intuitive means of expressing concepts that are difficult to describe textually. However, current MLLMs struggle to understand hand-drawn sketches due to the substantial distribution gap between natural images and hand-drawn sketches, together with the scarcity of large-scale open-source multimodal training data suitable for learning sketch-language alignment. To bridge this gap, we introduce SketchVCL, a large-scale synthetic dataset and develop O3SLM, the first open-weight MLLM to achieve unified three-way alignment across hand-drawn sketches, natural images and textual instructions. Trained via a two-stage protocol consisting of pretraining and instruction tuning, O3SLM serves as a unified model for multiple sketch-based tasks: (a) object localization, (b) counting, (c) image retrieval (i.e., SBIR and fine-grained SBIR), and (d) visual question answering (VQA). Comprehensive evaluations across three existing sketch datasets, namely QuickDraw!, Sketchy, and Tu Berlin, along with our synthetic SketchVCL dataset, show that O3SLM achieves state-of-the-art performance, substantially outperforming existing MLLMs in sketch comprehension and reasoning.

The second contribution addresses MLLM-based 3D asset generation. We argue that the primary limitation of existing approaches is representational rather than architectural. Existing 3D representations either expose low-level geometric parameters that are not naturally aligned with the capabilities of pretrained MLLMs or impose restrictive structural abstractions that limit geometric expressiveness. To address this, we introduce the Hierarchical Patch Scaffold (HPS), a novel representation designed around semantically meaningful, interpretable part-level abstractions that are naturally understood by both humans and LLMs. By design, it enables expressive, editable, and open-domain 3D generation without imposing restrictive geometric priors or requiring expensive finetuning. Building upon HPS, we propose SurfScaff3D, an interactive framework for text and image to 3D generation. HPS inherits the flexibility of B-spline patches while exposing an interface that leverages the open-world capabilities of foundation models to reliably generate 3D from diverse input prompts. Its hierarchical organization enables targeted natural-language edits at any level of granularity without regenerating the entire object. Simultaneously, HPS offers explicit per-patch control over surface detail, enabling faithful representation of arbitrary object shapes. Extensive evaluations and comparisons with state-of-the-art MLLM-based 3D generation methods demonstrate strong geometric fidelity, semantic alignment and flexible editing capability across diverse object categories.

This thesis advances the visual capabilities of MLLMs by expanding their scope beyond traditional, texture-heavy pixel grids toward structured, intuitive image-adjacent visual modalities. By introducing frameworks for unified sketch-image-text alignment and interpretable, part-level 3D geometric generation, this work successfully bridges the gap between low-level pixel processing and high-level structural reasoning. These contributions demonstrate that MLLMs can move past dense photorealism to effectively reason with both sparse and explicit geometric representations. Ultimately, these frameworks provide practical tools for multimodal models to operate beyond standard images, opening up more flexible possibilities for visual intelligence and creative workflows.


ALL ARE WELCOME