DEPARTMENT OF COMPUTATIONAL AND DATA SCIENCES
Ph.D. Thesis Colloquium
Speaker: Mr. Ankit Dhiman
S.R. Number: 06-18-00-11-12-21-1-20342
Title: “From Reconstructing Scenes to Geometry-Guided Synthesis”
Research Supervisor: Prof. Venkatesh Babu
Date & Time : September 25, 2026 (Friday), 11:00 AM
Venue : #102, CDS Seminar Hall
ABSTRACT
Recent 3D representations such as Neural Radiance Fields and 3D Gaussian Splatting have advanced the field of photorealistic novel-view synthesis from posed images. Beyond the task of novel-view synthesis, these 3D representations aggregate 2D foundation model predictions across views and provide geometric conditional signals to generative models for complex tasks. This thesis explores this broader role, progressing from (1) improving the underlying 3D representations, to (2) resolving inconsistencies in cross-view predictions, and finally (3) using recovered geometry to guide generative models. Specifically, the thesis is organized into three parts:
- Neural Scene Representations: The first part focuses on three limitations of current neural scene representations: (i) modeling stratified scenes, (ii) the high computational cost for reconstruction form high-resolution multi-view images, and (iii) aliasing in dynamic scenes. Strata-NeRF models stratified scenes using a vector-quantized latent-conditioned radiance field that learns transitions between different levels without supervision. Turbo-GS accelerates the fitting of 3D Gaussian Splatting for high-resolution multi-view inputs. For dynamic scenes, we show that motion changes the effective sampling rate and derive a motion-aware filter to reduce aliasing across scale and time. Together, these works improve the efficiency and robustness of neural scene representations.
- Lifting 2D Cues into 3D: The second part uses a 3D representation to aggregate 2D predictions that are consistent for a single view but inconsistent across views. ChromaDistill colorizes scenes captured without color (legacy grayscale and Infrared image sequences) by distilling a pretrained colorization model into the 3D representation during training. This makes the predicted color consistent across views. UniC-Lift lifts per-image instance masks into a 3D Gaussian representation and directly predicts consistent instance labels, avoiding the post-processing clustering stage used by earlier methods.
- Geometry-Conditioned Generation: The third part explores whether explicit scene geometry can constrain generative models. Mirror reflection synthesis is a challenging test for generative models as the reflected content on the mirror should be determined by the scene geometry and an incorrect reflection is immediately apparent. Reflecting Reality introduces a depth-conditioned approach to the aforementioned task together with the large-scale synthetic SynMirror dataset. MirrorVerse improves its generalization by introducing greater diversity in synthetic scenes through SynMirrorV2. GeoMirror takes a further step by computing the reflected scene geometry explicitly and using it to guide generation. Finally, GeoNoise extends this idea beyond reflections, using scene geometry to guide the generative process itself for training-free, geometry-consistent novel-view synthesis.
This thesis advances neural scene representations by making them more efficient and robust to aliasing artifacts, and shows the importance of making geometry explicit for reliable visual prediction and generation. This is important when correctness is defined in physical terms, and plausibility is insufficient. As geometry-consistent novel-view synthesis and high-fidelity dynamic reconstruction continue to advance, these ideas open new avenues toward monocular dynamic novel-view synthesis and geometrically consistent video and world models.
ALL ARE WELCOME



