BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//wp-events-plugin.com//7.4.0.1//EN
TZID:Asia/Kolkata
X-WR-TIMEZONE:Asia/Kolkata
BEGIN:VEVENT
UID:215@cds.iisc.ac.in
DTSTART;TZID=Asia/Kolkata:20260723T113000
DTEND;TZID=Asia/Kolkata:20260723T123000
DTSTAMP:20260721T123712Z
URL:https://cds.iisc.ac.in/events/m-tech-research-thesis-colloquium-202-cd
 s-23-july-2026-advancing-visual-reasoning-capabilities-of-multimodal-llms-
 beyond-natural-images/
SUMMARY:M.Tech Research Thesis {Colloquium}: 202: CDS: 23\, July 2026 “Ad
 vancing Visual Reasoning Capabilities of Multimodal LLMs Beyond Natural Im
 ages”
DESCRIPTION:DEPARTMENT OF COMPUTATIONAL AND DATA SCIENCES\nM.Tech Research 
 Thesis {Colloquium}\n\n\n\nSpeaker: Mr. Rishi Gupta\nS.R. Number: 06-18-01
 -10-22-24-1-24504\nTitle: “Advancing Visual Reasoning Capabilities of Mu
 ltimodal LLMs Beyond Natural Images”\nResearch Supervisor: Prof. Anirban
  Chakraborty\nDate &amp\; Time : July 23\, 2026 (Thursday)\, 11:30 AM\nVen
 ue : CDS #202\n\n\n\nABSTRACT\nArtificial Intelligence is undergoing a par
 adigm shift with the emergence of Large Language Models (LLMs) and Multimo
 dal Large Language Models (MLLMs). Although many of these models demonstra
 te remarkable capabilities on natural image understanding tasks (e.g.\, Li
 u et al.\, NeurIPS 2023\; Deitke et al.\, CVPR 2025)\, they remain fundame
 ntally grounded in dense\, photorealistic pixel representations. Human vis
 ual reasoning\, however\, extends beyond natural images and routinely reli
 es on structured visual modalities such as sketches and explicit 3D geomet
 ry of objects. These image-adjacent representations preserve semantic and 
 geometric structure while discarding much of the textural information pres
 ent in natural images\, making them intuitive for human communication and 
 creative workflows. Therefore\, enabling MLLMs to reason with such structu
 red modalities remains an important step toward more flexible and holistic
  visual intelligence. This thesis explores how MLLMs can comprehend and ge
 nerate visual modalities beyond dense natural images while preserving intu
 itive interaction through text\, hand-drawn sketches\, and natural images.
 \n\nThe thesis explores this problem from two complementary axes: understa
 nding sparse visual inputs and generating semantically interpretable geome
 tric representations. Firstly\, we focus on sketch understanding\, as sket
 ches provide an intuitive means of expressing concepts that are difficult 
 to describe textually. However\, current MLLMs struggle to understand hand
 -drawn sketches due to the substantial distribution gap between natural im
 ages and hand-drawn sketches\, together with the scarcity of large-scale o
 pen-source multimodal training data suitable for learning sketch-language 
 alignment. To bridge this gap\, we introduce SketchVCL\, a large-scale syn
 thetic dataset and develop O3SLM\, the first open-weight MLLM to achieve u
 nified three-way alignment across hand-drawn sketches\, natural images and
  textual instructions. Trained via a two-stage protocol consisting of pret
 raining and instruction tuning\, O3SLM serves as a unified model for multi
 ple sketch-based tasks: (a) object localization\, (b) counting\, (c) image
  retrieval (i.e.\, SBIR and fine-grained SBIR)\, and (d) visual question a
 nswering (VQA). Comprehensive evaluations across three existing sketch dat
 asets\, namely QuickDraw!\, Sketchy\, and Tu Berlin\, along with our synth
 etic SketchVCL dataset\, show that O3SLM achieves state-of-the-art perform
 ance\, substantially outperforming existing MLLMs in sketch comprehension 
 and reasoning.\n\nThe second contribution addresses MLLM-based 3D asset ge
 neration. We argue that the primary limitation of existing approaches is r
 epresentational rather than architectural. Existing 3D representations eit
 her expose low-level geometric parameters that are not naturally aligned w
 ith the capabilities of pretrained MLLMs or impose restrictive structural 
 abstractions that limit geometric expressiveness. To address this\, we int
 roduce the Hierarchical Patch Scaffold (HPS)\, a novel representation desi
 gned around semantically meaningful\, interpretable part-level abstraction
 s that are naturally understood by both humans and LLMs. By design\, it en
 ables expressive\, editable\, and open-domain 3D generation without imposi
 ng restrictive geometric priors or requiring expensive finetuning. Buildin
 g upon HPS\, we propose SurfScaff3D\, an interactive framework for text an
 d image to 3D generation. HPS inherits the flexibility of B-spline patches
  while exposing an interface that leverages the open-world capabilities of
  foundation models to reliably generate 3D from diverse input prompts. Its
  hierarchical organization enables targeted natural-language edits at any 
 level of granularity without regenerating the entire object. Simultaneousl
 y\, HPS offers explicit per-patch control over surface detail\, enabling f
 aithful representation of arbitrary object shapes. Extensive evaluations a
 nd comparisons with state-of-the-art MLLM-based 3D generation methods demo
 nstrate strong geometric fidelity\, semantic alignment and flexible editin
 g capability across diverse object categories.\n\nThis thesis advances the
  visual capabilities of MLLMs by expanding their scope beyond traditional\
 , texture-heavy pixel grids toward structured\, intuitive image-adjacent v
 isual modalities. By introducing frameworks for unified sketch-image-text 
 alignment and interpretable\, part-level 3D geometric generation\, this wo
 rk successfully bridges the gap between low-level pixel processing and hig
 h-level structural reasoning. These contributions demonstrate that MLLMs c
 an move past dense photorealism to effectively reason with both sparse and
  explicit geometric representations. Ultimately\, these frameworks provide
  practical tools for multimodal models to operate beyond standard images\,
  opening up more flexible possibilities for visual intelligence and creati
 ve workflows.\n\n\n\nALL ARE WELCOME
CATEGORIES:Events,MTech Research Thesis Colloquium
END:VEVENT
BEGIN:VTIMEZONE
TZID:Asia/Kolkata
X-LIC-LOCATION:Asia/Kolkata
BEGIN:STANDARD
DTSTART:20250723T113000
TZOFFSETFROM:+0530
TZOFFSETTO:+0530
TZNAME:IST
END:STANDARD
END:VTIMEZONE
END:VCALENDAR