Key Takeaways

  • NASA is expanding an open science data estate that already serves more than 53 million unique users.
  • AWS and Microsoft Azure are becoming important infrastructure layers for storage, discovery, and computing.
  • Rising public resistance to AI data centers could reshape how government agencies and cloud providers plan future capacity.

NASA’s space data operation is increasingly starting to resemble a hyperscale technology platform. The agency’s Science Mission Directorate managed about 150 PB across 54,532 datasets and 10 repositories in FY 2024, serving more than 53 million unique users.

Those figures, published through NASA Science Data Repository Metrics, underscore a shift in what constitutes critical space infrastructure. Rockets, satellites, and scientific instruments still attract most of the attention. But storage capacity, search systems, network connections, and computing environments now determine how quickly observations can become useful science.

Earth science represents much of that weight. NASA archived 128.6 PB of Earth science information and distributed over 450 TB per day to 8.35 million users globally in FY 2024, according to research published by CODATA Data Science Journal. This is not a static digital library. Researchers continually access, combine, and process information for work spanning climate monitoring, disaster response, agriculture, and atmospheric science.

Commercial cloud infrastructure is taking on a larger role. NASA’s Earth Science Data and Information System operates most of its components in AWS, while more than 90% of its roughly 170 PB archive resides in Amazon S3. Full migration is targeted for the end of 2026.

Moving information into cloud storage does not, by itself, make that information easier to use. Scientists still need consistent metadata, interoperable formats, manageable computing costs, and ways to locate relevant records among tens of thousands of datasets. Data gravity also matters. At this scale, moving information repeatedly between systems can become slow and expensive, making it more practical to bring computation to the archive.

NASA’s response includes the Science Cloud, which brings several internal cloud environments under standardized technical and governance guardrails while connecting workloads to commercial services including AWS and Microsoft Azure. The approach gives research teams a more consistent route to computing resources without forcing every mission to design its own operating model.

Discovery is another target. NASA recently rebuilt its Science Discovery Engine using AWS OpenSearch, combining keyword and vector search, and the organization reported reducing operating costs by about sixfold. The agency’s Science Discovery Engine infrastructure update shows how AI-adjacent technologies can improve access before more ambitious scientific models are introduced.

That said, the physical footprint behind cloud computing is becoming politically harder to ignore. Data centers supporting AI and cloud services have lost favor in some communities, where residents associate new facilities with pressure on electricity prices, water supplies, transmission infrastructure, and available land. Whether every proposed facility produces those effects is a more complicated question. Public perception can still influence permits, utility planning, and construction schedules.

What happens when open science depends on infrastructure that communities increasingly view with suspicion? For NASA and its suppliers, the answer may involve greater transparency around energy use, regional capacity, water consumption, and the public value created by computing workloads. A facility processing wildfire observations or hurricane data may have a different public-interest case than one devoted largely to commercial AI inference, even if both draw from the same grid.

SpaceX and Starlink add another layer through launch services and satellite connectivity, while Google Cloud and Alphabet bring AI and analytics capabilities to the broader market surrounding public-sector science. The chief executives of these technology companies therefore sit near the commercial edges of an ecosystem that NASA is trying to keep open, interoperable, and scientifically accountable. Their companies do not control NASA’s data policy, but their infrastructure decisions can shape the cost and availability of the technologies on which modern research increasingly relies.

Governance will be as important as capacity. NASA’s Science Data and Computing Strategy 2025–2030 prioritizes an open, interoperable data ecosystem, more efficient computing infrastructure, and faster scientific work using AI and quantum computing. It also advances a standard open data policy and supports FAIR principles, under which data should be findable, accessible, interoperable, and reusable.

For enterprise technology leaders, NASA offers a useful preview. Large data estates are becoming ecosystems rather than archives, and search, governance, computing location, and infrastructure legitimacy are converging into one architectural problem. Cloud migration remains part of the answer. The harder work is ensuring that scale produces broader access and measurable value without losing public trust along the way.