| Abstract | Machine learning (ML) models generally require large datasets for better accuracy. However, data privacy and security concerns impose constraints on centralized data collection and model training. Federated Learning (FL) addresses this by training models without data leaving its origin, but FL still encounters challenges regarding training efficiency, model convergence and model accuracy due to data heterogeneity among clients. This work proposes a Hierarchical Federated Learning (HFL) framework that introduces an intermediate aggregation layer that dynamically groups similar clients into clusters, allowing analysis of privacy and accuracy trade-off through two alternative approaches: HFL+NoDataMove (within the cluster model weight aggregation for better privacy) and HFL+DataMove (within the cluster data pooling for improved accuracy). Both alternatives employ Locality Sensitive Hashing (LSH) for lightweight cluster building and optionally quantization for communication efficiency. Our experiments to evaluate the proposed framework were conducted on a real-life Fitbit activity dataset of 33 users using linear regression as the model for prediction of calories burnt during daily activity. Experiments across 10 independent experiment runs demonstrate that HFL+DataMove reaches centralized accuracy (mean MSEs of 0.070 vs 0.057, respectively), significantly outperforming three FL benchmarks FedAvg, FedProx, SCAFFOLD, and one proposed method HFL+NoDataMove (p<0.001, Cohen’s d > 2.5). Experiments were conducted to evaluate the proposed framework in terms of model performance, communication overhead, and computational cost, with Mean Squared Error used as the primary performance metric. The results suggest that the proposed methodology is feasible for IoT (Internet of Things) deployment, with total communication overhead remaining below 400 KB across 10 global rounds. |
|---|