Beyond Data Points: Regionalizing Crowdsourced Latency Measurements

Sharma, Taveesh; Schmitt, Paul; Bronzino, Francesco; Feamster, Nick; Marwell, Nicole

Computer Science > Networking and Internet Architecture

arXiv:2405.11138 (cs)

[Submitted on 18 May 2024 (v1), last revised 26 Oct 2024 (this version, v4)]

Title:Beyond Data Points: Regionalizing Crowdsourced Latency Measurements

Authors:Taveesh Sharma, Paul Schmitt, Francesco Bronzino, Nick Feamster, Nicole Marwell

View PDF HTML (experimental)

Abstract:Despite significant investments in access network infrastructure, universal access to high-quality Internet connectivity remains a challenge. Policymakers often rely on large-scale, crowdsourced measurement datasets to assess the distribution of access network performance across geographic areas. These decisions typically rest on the assumption that Internet performance is uniformly distributed within predefined social boundaries. However, this assumption may not be valid for two reasons: crowdsourced measurements often exhibit non-uniform sampling densities within geographic areas; and predefined social boundaries may not align with the actual boundaries of Internet infrastructure. In this paper, we present a spatial analysis on crowdsourced datasets for constructing stable boundaries for sampling Internet performance. We hypothesize that greater stability in sampling boundaries will reflect the true nature of Internet performance disparities than misleading patterns observed as a result of data sampling variations. We apply and evaluate a series of statistical techniques to: aggregate Internet performance over geographic regions; overlay interpolated maps with various sampling unit choices; and spatially cluster boundary units to identify contiguous areas with similar performance characteristics. We assess the effectiveness of the techniques we apply by comparing the similarity of the resulting boundaries for monthly samples drawn from the dataset. Our evaluation shows that the combination of techniques we apply achieves higher similarity compared to directly calculating central measures of network metrics over census tracts or neighborhood boundaries. These findings underscore the important role of spatial modeling in accurately assessing and optimizing the distribution of Internet performance, to inform policy, network operations, and long-term planning decisions.

Comments:	24 pages
Subjects:	Networking and Internet Architecture (cs.NI); Computers and Society (cs.CY)
Cite as:	arXiv:2405.11138 [cs.NI]
	(or arXiv:2405.11138v4 [cs.NI] for this version)
	https://doi.org/10.48550/arXiv.2405.11138

Submission history

From: Taveesh Sharma [view email]
[v1] Sat, 18 May 2024 01:39:22 UTC (3,220 KB)
[v2] Tue, 21 May 2024 13:10:51 UTC (3,234 KB)
[v3] Mon, 12 Aug 2024 11:54:24 UTC (33,650 KB)
[v4] Sat, 26 Oct 2024 13:36:57 UTC (19,961 KB)

Computer Science > Networking and Internet Architecture

Title:Beyond Data Points: Regionalizing Crowdsourced Latency Measurements

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Networking and Internet Architecture

Title:Beyond Data Points: Regionalizing Crowdsourced Latency Measurements

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators