Source-linked AI summary
Petascale Computational Systems
Gordon Bell, Jim Gray, Alex Szalay
TL;DR
Computational science is becoming data intensive, requiring systems that balance processing with storage, I/O, networking, and semantic data management. The paper argues for balanced cyberinfrastructure across Tier-1 through Tier-3 facilities, illustrated by a Tier-2 implementation close to Amdahl’s disk-capacity law and within a factor of 3 in bandwidth.
Problem
Scientific instruments, simulations, and accessible archives are producing petascale datasets whose storage, I/O, analysis, and management requirements challenge CPU-focused infrastructure.
Method
The paper applies Amdahl’s balanced-system laws to system design and combines scaling estimates, data-locality principles, infrastructure recommendations, and a Tier-2 case study.
Results
The proposed design requires petascale I/O and networking alongside computation, while the JHU Tier-2 facility approached Amdahl’s disk-capacity balance and achieved bandwidth within a factor of 3.
Takeaways & Limitations
Funding agencies should support balanced systems and allocate resources across Tier-1 through Tier-3 cyberinfrastructure rather than concentrating resources on CPU farms.
Abstract
from arXiv · showhide
Computational science is changing to be data intensive. Super-Computers must be balanced systems; not just CPU farms but also petascale IO and networking arrays. Anyone building CyberInfrastructure should allocate resources to support a balanced Tier-1 through Tier-3 design.