Ph.D: Thesis Colloquium: 102: CDS: 07, August 2026“Systems for Federated Learning across the Edge–Fog–Cloud Continuum”

When

7 Aug 26    
12:00 PM - 1:00 PM

Event Type

DEPARTMENT OF COMPUTATIONAL AND DATA SCIENCES
Ph.D. Thesis Colloquium


Speaker : Ms. Roopkatha Banerjee
S.R. Number : 06-18-01-10-12-21-1-19488
Title : “Systems for Federated Learning across the Edge–Fog–Cloud Continuum”
Research Supervisor : Prof. Yogesh Simmhan
Date & Time : August 07, 2026 (Friday), 12:00 Noon
Venue : #102, CDS Seminar Hall


ABSTRACT

“Federated Learning (FL) has emerged as a privacy-preserving paradigm for training deep neural networks(DNNs) across distributed clients that own private data, sharing only model updates rather than raw data with a coordinating server. Its promise is amplified by accelerated edge devices, such as NVIDIA Jetson, which pack hundreds to thousands of CUDA cores into compact, low-power form factors co-located with data sources in smart cities, homes, and industrial settings. Yet most FL research is validated in pseudo-distributed simulation, under idealized assumptions: homogeneous and reliable clients, unconstrained energy, trusted infrastructure, and labeled, stationary data from a fixed set of classes. These assumptions rarely hold when FL is deployed on real, heterogeneous edge and fog hardware in the field.

This dissertation takes a systems approach to making FL practical under the constraints that real-world edge deployments impose. Motivated throughout by smart-city deployments, we build a modular framework as a common substrate and layer upon it optimizations and scheduling strategies that each confront a constraint simulation hides: device heterogeneity and failure, energy budgets, continuous drift and data sovereignty, and unlabeled, open-world data. All are validated on real edge accelerators and single-board computers spanning the edge–fog–cloud continuum, at the scale of hundreds of clients.

We first develop Flotilla, a scalable, modular, and resilient FL framework that serves as the deployment substrate for this dissertation. Its “user-first” design lets researchers rapidly compose synchronous and asynchronous FL strategies while remaining agnostic to the DNN architecture, and its stateless clients and checkpointed server state enable rapid recovery from failures. Across five FL strategies and five DNN models, Flotilla shows sub-second failover on 200+ clients and a resource footprint comparable to or better than Flower, OpenFL, and FedML.

Building on this substrate, FedJoule trains within a global energy budget over heterogeneous edge accelerators. We formulate a client-selection problem that maximizes accuracy within an overall energy limit while reducing training time, solved through a bi-level Integer Linear Program using approximate Shapley values and energy–time prediction models. Across diverse budgets and non-IID distributions, FedJoule outperforms state-of-the-art and simple baselines by 15% on accuracy and 48% on time.

We then design SURGE, a continuous FL framework that remains dependable and data-sovereign under real-world drift. SURGE decouples data storage from computation using decentralized W3C Solid pods, and delegates local training to trusted, GPU-equipped fog nodes through drift-aware pod selection, trust-constrained pod-to-fog assignment, and pipelined scheduling. Across deployments of 20–160 devices, SURGE improves global-model accuracy by 3–17% over FedDrift and reduces drift-detection overhead by up to 23x, adding only sub-second Solid overhead and roughly 7% orchestration overhead for Flotilla.

Finally, in ODIN, we support incremental FL and open-world class discovery over unlabeled, continuously evolving data streams. ODIN maps data into a hyperspherical feature space using a lightweight Vision Transformer, filters known from out-of-distribution samples, and coordinates a server-side merge-and-discover protocol to detect novel categories, while diversity-based coreset replay mitigates catastrophic forgetting. On two real-world Indian vehicle datasets, ODIN outperforms the FC²DL and Fed-GCD baselines by 7–32% and 19–41%, respectively, approaching SAM-based zero-shot labeling to within 12% at orders-of-magnitude lower latency.

Together, these contributions offer a systems foundation for federated learning that is deployable, efficient, dependable, and adaptive under real-world edge constraints, giving practitioners concrete frameworks, optimization methods, and scheduling strategies to leverage accelerated edge and fog hardware for privacy-preserving distributed intelligence.”


ALL ARE WELCOME