NSF Just Put $183 Million Behind the Data Layer of AI Science — Inside the $83M IDSS Awards and the New $100M Unlocking Dataset Value Program
July 22, 2026 · 6 min read
Granted Research Team · Editorial policy
Everyone talks about compute. When the White House and the national labs describe the race to build AI-driven science, the imagery is always the same: exascale supercomputers, GPU clusters, the physical machinery of intelligence. But anyone who has actually trained a model or run an AI-enabled analysis knows the machine is the easy part. The hard part — the part that quietly determines whether a discipline can use AI at all — is the data: whether it is findable, whether it is clean, whether it carries the metadata a model needs, whether it can move from an instrument to a compute node without a graduate student spending six months writing glue code.
On July 22, 2026, the U.S. National Science Foundation put real money behind that unglamorous truth. It announced $83 million in awards through its Integrated Data Systems and Services (IDSS) program, and simultaneously launched a companion initiative — Unlocking Dataset Value for AI-Enabled Scientific Discovery — offering up to $100 million in awards ranging from $2 million to $5 million each. Taken together, these two moves represent roughly $183 million aimed squarely at the data layer of American science, and they are among the clearest signals yet of how NSF intends to operationalize the AI-for-science agenda in the Genesis Mission era.
What IDSS actually funds
The Integrated Data Systems and Services program is not about collecting new data. It is about the infrastructure that lets existing scientific data be discovered, shared, and analyzed at scale — the connective tissue between NSF facilities, national laboratories, instruments, software, and AI resources. As NSF Acting Director Brian Stone framed it: "America's leadership in artificial intelligence depends not only on advanced computing resources but also on the data infrastructure that enables researchers to discover, access, share and analyze scientific data at scale."
The $83 million was distributed across two categories, and the distinction matters if you plan to compete in a future round.
Category I — National-Scale Integrated Data Systems funds the most ambitious, ground-up platforms:
- Fabric for AI-Driven Science (FabAID) — Morgridge Institute for Research, Madison, Wisconsin
- National Data Platform (NDP) — UC San Diego
Category II — Transition to National-Scale Operations funds efforts that have already proven a concept and are ready to scale into durable national services:
- Interactive Discovery Laboratory (iDLab) — UCLA
- BRIDGE National Center — University of California, Irvine
- National Science Data Fabric (NSDF) — University of Tennessee, Knoxville
- Multidisciplinary Environment for Scientific Advancement (MESA) — University of Arizona
NSF also awarded a set of smaller planning grants to seed future full proposals — a detail worth underlining, because planning grants are the on-ramp for institutions that were not ready to compete for a national center this cycle.
The common thread across all six flagship awards is that each one builds a system that collects, stores, shares, and analyzes scientific data across NSF facilities and national labs; creates shared platforms for discovering and reusing data across projects and disciplines; provides access to AI tools and advanced computing; and enables secure, reproducible workflows. These are not domain science grants. They are grants to build the roads that domain science will drive on.
The companion play: Unlocking Dataset Value
If IDSS builds the highways, the Unlocking Dataset Value for AI-Enabled Scientific Discovery program — call it AI Datasets — resurfaces the cargo. Its premise is that the United States is already sitting on enormous scientific datasets that are effectively unusable by modern AI because they lack the structure machines need: consistent metadata, extractable features, interoperable formats, and automated pipelines.
Rather than fund new data collection, the program pays teams to make existing datasets AI-ready. Specifically, it funds:
- Novel methods for feature extraction, metadata generation, and dataset integration
- Robust data pipelines for automated AI analysis
- Increased accessibility and interoperability of existing datasets
- Work that enables AI-driven insights and interdisciplinary research
The numbers are generous by NSF standards: up to $100 million total, with individual awards of $2 million to $5 million, plus planning grants of up to $200,000. NSF's framing is blunt: "high-quality scientific data are foundational to advancing an AI-enabled research and innovation ecosystem." The program is explicitly aligned with America's AI Action Plan and is designed to support AI-enabled discovery across the same national science and technology challenges that anchor the Genesis Mission.
Why these two programs launched together
The simultaneity is not an accident. IDSS and AI Datasets are two halves of the same theory of the case. IDSS builds the platforms — the data fabrics, the national data platform, the interoperable operations centers. AI Datasets fills those platforms with datasets that are worth pointing a model at. A national data fabric with no clean, well-labeled datasets flowing through it is a highway to nowhere; a beautifully curated dataset with no infrastructure to serve it at scale is a museum piece.
Both connect upward to the Genesis Mission, the DOE-led national effort to fuse AI, exascale compute, and the 17 national labs into a single scientific platform. NSF's IDSS awards are explicitly described as supporting the Genesis Mission — which tells you something important about the funding landscape ahead: the agencies are coordinating, and proposals that speak the shared vocabulary of AI-readiness, interoperability, and national-scale infrastructure are going to travel further than ones that treat data as an afterthought.
The strategic read for research teams
Most readers of this analysis will not be the Morgridge Institute or UC San Diego. So the useful question is: what do these announcements mean if you run a lab, a data center, or a mid-sized research institution and you were not in the first cohort?
1. Planning grants are the real entry point. Both programs awarded planning grants, and IDSS explicitly seeded them for future full proposals. A planning grant is a low-risk, high-signal way to enter the pipeline: it funds you to build the partnerships, governance model, and technical architecture that a competitive national-center proposal requires. If you are anywhere near this space, a planning grant is the move — not a moonshot full proposal you are not yet resourced to win.
2. AI-readiness is now a fundable activity in its own right. For years, "cleaning the data" was overhead you buried inside a science grant. The Unlocking Dataset Value program makes it a first-class, $2M-$5M funded objective. If your institution holds a valuable, underused dataset — instrument archives, longitudinal study data, sensor networks, specimen collections — the work of making it AI-ready is now the deliverable, not a chore. Frame the dataset's national value, its interdisciplinary reach, and the specific pipelines you would build.
3. Speak the interoperability language. Every winning IDSS award shares a vocabulary: shared platforms, reuse across disciplines, secure and reproducible workflows, connection to national labs and AI resources. Reviewers in this program are not rewarding clever single-lab tools. They are rewarding infrastructure that many communities can use. Proposals that show real cross-institutional demand and a credible path to national-scale operations will win; single-PI, single-domain tools will not.
4. Watch the transition category. Category II — "Transition to National-Scale Operations" — is a tell. NSF is signaling that it wants to graduate proven pilots into durable services rather than fund an endless parade of prototypes. If you have a working data tool with real users, the fundable narrative is the transition story: how you go from a research project to sustained national infrastructure.
5. Position for the Genesis Mission halo. Because IDSS is tied to the Genesis Mission, and because a dozen agencies are now attaching funding to the same national science and technology challenges, data-infrastructure work has unusual cross-agency optionality right now. A single strong capability — say, an AI-ready pipeline for materials or biology data — can be pitched to NSF here, to DOE through its Genesis Mission SBIR/STTR topics, and to mission agencies that need the same capability. Build the capability once; tell the story in the vocabulary each funder uses.
The bigger signal
For a decade, the prestige and the dollars in computational science flowed to compute. IDSS and Unlocking Dataset Value are NSF's acknowledgment that the binding constraint has moved. The models are good enough; the compute is being built; what is missing is data that machines can actually use, and infrastructure that can serve it at national scale. That is now a fundable frontier — $183 million worth, in a single day, with more clearly coming.
The teams that internalize this shift early — that treat data infrastructure and AI-readiness as fundable science rather than back-office overhead — are the ones who will be positioned when the next, larger round opens. The data layer just became the main event.
Tracking NSF's AI infrastructure programs and want a shortlist of the specific solicitations, deadlines, and funder contacts that match your institution's data assets? Granted maps your research profile against active federal and foundation opportunities so you spend your time writing, not hunting.