AI-Ready Data and Research Cyberinfrastructure
NSF NCAR provides national-scale high-performance computing (HPC) systems, scalable high-bandwidth storage, modern software tools designed to support AI and ML research, and open access to large public Earth system datasets. From data discovery and preprocessing to distributed training, evaluation, inference, and deployment, this integrated environment enables end-to-end model development within a unified, performance-optimized ecosystem. Here, deployment refers to running a trained and validated machine learning model within an operational or user-facing system. Multi-GPU clusters, high-speed interconnects, and parallel file systems support large-scale training and fine-tuning of deep learning models. At the same time, optimized ML frameworks and workflow tools allow researchers to operationalize inference at scale efficiently. By co-locating curated Earth system science datasets with advanced computing infrastructure, NSF NCAR reduces data movement and enables reproducible, high-throughput experimentation.
We work throughout the research and data lifecycle to transform large and diverse datasets into formats optimized for scalable analysis and AI workflows. Centralized access, grounded in FAIR (Findable, Accessible, Interoperable, Reusable) principles, ensures that researchers can discover, access, and reuse data with clarity and confidence. Researchers can analyze data where it resides using streaming and subsetting services, reducing the need to download massive files. Standardized metadata, provenance tracking, and documentation support transparent interpretation, reproducibility, and citation. Together, these capabilities provide a robust foundation for trustworthy AI development, accelerate scientific discovery, and maintain rigor and operational reliability.
Data
Computing for AI Research
