Please use this identifier to cite or link to this item:
http://repository.iiitd.edu.in/xmlui/handle/123456789/2001Full metadata record
| DC Field | Value | Language |
|---|---|---|
| dc.contributor.author | Singh, Ramanjeet | - |
| dc.contributor.author | Shah, Rinku (Advisor) | - |
| dc.date.accessioned | 2026-08-20T14:02:10Z | - |
| dc.date.available | 2026-08-20T14:02:10Z | - |
| dc.date.issued | 2024-12-01 | - |
| dc.identifier.uri | http://repository.iiitd.edu.in/xmlui/handle/123456789/2001 | - |
| dc.description.abstract | Distributed machine learning (ML) workloads, particularly the training of deep neural networks (DNNs) and large language models (LLMs), rely heavily on efficient network communication to achieve scalability. However, network limitations, such as bandwidth bottlenecks, latency, and traffic imbalances, often limit the linear scalability of distributed training systems. This report investigates the network characteristics and communication behaviors of distributed ML workloads through simulation, with a focus on identifying bottlenecks and deriving insights for improving network infrastructure. Using Astra-Sim, a distributed ML workload simulator integrated with NS-3 for network modeling, we simulate various scenarios, including the introduction of network jitter and variations in NIC parameters and congestion control protocols. The experiments analyze the impact of topology, de lay, and protocol selection on job completion times (JCT), revealing key trade-offs and performance bottlenecks. The results demonstrate that network jitter significantly affects performance, particu larly in hierarchical topologies, while lower-latency NICs and efficient congestion control protocols such as DCQCN and DCTCP provide substantial improvements. Looking ahead, this work proposes the simulation of large-scale models across high-dimensional topologies involving thousands of GPUs to analyze network characteristics at scale. These insights aim to guide the design of next-generation network protocols and infrastructure, addressing the challenges of scaling distributed ML workloads. | en_US |
| dc.language.iso | en_US | en_US |
| dc.publisher | IIIT-Delhi | en_US |
| dc.subject | Machine Learning | en_US |
| dc.subject | Large Language Models | en_US |
| dc.subject | Network Technologies for AI | en_US |
| dc.title | Data center network technologies for AI/ML workloads | en_US |
| dc.type | Other | en_US |
| Appears in Collections: | Year-2024 | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| BTP_Report_2021085 - Ramanjeet Singh.pdf Restricted Access | 1.13 MB | Adobe PDF | View/Open Request a copy |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.