NSF's AI Datasets Program (NSF 26-512): $60–100M, Three Award Tiers, and a November 4 Deadline for Making Scientific Data 'AI-Ready'
July 27, 2026 · 6 min read
Granted Research Team · Editorial policy
There is a quiet assumption inside most AI-for-science funding: that the hard part is the model. Build a better neural network, the thinking goes, and discovery follows. NSF's new AI Datasets program — solicitation NSF 26-512, formally titled Unlocking Dataset Value for AI-Enabled Scientific Discovery — bets the opposite. Its premise is that the models are increasingly commoditized and the real bottleneck is data that AI can actually use: curated, harmonized, well-documented, governed, machine-readable scientific datasets. The program puts $60 million to $100 million behind that premise, with a first full-proposal deadline of November 4, 2026.
For research teams sitting on valuable but messy data — a decades-long field survey, an instrument archive, a collection of clinical or ecological or materials records — this is one of the more directly actionable solicitations of the year. Here is how it is built and how to approach it.
The three tiers, and what each is really for
NSF 26-512 funds 25 to 50 awards across three distinct tiers, and the tiers are not just size buckets — they encode different theories of impact.
Flagship Awards — up to $5,000,000 each, 5 to 10 awards. These are the anchor projects. A Flagship is for a dataset (or federated collection of datasets) with the scale and scientific centrality to become a durable community resource — the kind of asset that many research groups will build on for years. Winning at this tier means demonstrating not just that your data is valuable but that transforming it into an AI-ready resource will move an entire field.
Impact Awards — up to $2,000,000 each, 10 to 20 awards. These fund focused, high-value efforts: taking a specific important dataset and doing the real work of making it usable by AI — feature extraction, metadata generation, building the pipelines that let automated analysis run over it. An Impact award is narrower than a Flagship but still expects a demonstrated payoff for a research community.
Planning Grants — up to $200,000 each, 10 to 20 awards. These are the on-ramp. A Planning grant funds the assessment, partnership-building, and design work needed to scope a future Flagship or Impact proposal. If your data is promising but your governance model, community partnerships, or technical pipeline are not yet mature, this is the honest place to start — and, strategically, a Planning award establishes a track record that strengthens a later full proposal.
The tiering rewards self-awareness. The single most common way to lose here is to over-reach — to pitch a Flagship when the underlying dataset and partnerships would credibly support only an Impact award, or an Impact when the work is really still at the Planning stage. Reviewers can tell. Match the tier to reality.
What NSF actually wants the money to buy
The solicitation names three technical goals, and every competitive proposal should speak to them explicitly:
- Apply AI-based capabilities to feature extraction and metadata generation — using AI itself to enrich the data, automatically labeling, describing, and structuring it so downstream models can consume it.
- Develop robust data pipelines for automated AI analysis — the plumbing that lets AI run over the data reliably and at scale, not as a one-off but as an ongoing service.
- Augment and harmonize datasets for AI use — cleaning, aligning, standardizing, and where appropriate combining datasets so they interoperate.
Crucially, NSF requires every proposal to address dataset governance, security, integrity, and community-contribution processes. This is not boilerplate. A proposal that treats governance as an afterthought — who controls the data, how quality is maintained, how the community contributes and benefits, how security and privacy are handled — will read as naive. The program is trying to build lasting public research infrastructure, and infrastructure without a governance model does not last.
Eligibility is unusually broad — use that
NSF 26-512 accepts proposals from a wide field: institutions of higher education (including two-year and community colleges), non-profit non-academic organizations, U.S.-based for-profit companies, state and local governments, Tribal nations, and federal agencies and FFRDCs. There is no limit on the number of proposals per organization or per PI/co-PI.
That breadth is a strategic opening. The most compelling AI-ready-data proposals often pair an organization that holds important data (a government agency, a field station, a company, a Tribal nation) with a group that has the technical muscle to make it AI-ready (a university lab, an FFRDC). Because eligibility spans all of those, cross-sector teaming is not just allowed — it is arguably the archetype the program was designed to fund. If you hold data but lack AI engineering, or you have AI engineering but no distinctive data, the solicitation is inviting you to find the other half.
The cost-sharing rule — read it correctly
The solicitation states plainly: "Inclusion of voluntary committed cost sharing is prohibited." This trips people up, so be precise about what it means. It does not mean you should propose to do the work for less. It means you must not offer institutional matching funds as a way to look more competitive — NSF does not want cost sharing to become a back-door bidding war that favors wealthy institutions. Build your budget to the full, honest cost of the work within the tier ceiling, and do not gesture at matching contributions. A community college and a flagship research university are meant to compete on the merit of the data and the plan, not on who can put up match.
There is no letter of intent and no preliminary proposal required — you go straight to the full proposal. That lowers the barrier to entry but also removes an early filter, which means the full-proposal field may be crowded. Quality of framing will decide it.
Where this sits in the bigger picture
NSF 26-512 does not stand alone. It explicitly aligns with America's AI Action Plan and the Genesis Mission — the $5 billion, 15-agency AI-for-science push the White House announced on July 22, 2026 (see our Genesis Mission analysis). It also complements NSF's separate $83 million IDSS data-infrastructure awards and the NAIRR compute resource. The through-line across all of these is a single conviction: that the next decade of American science will be gated less by algorithms than by whether the nation's scientific data can be fed to those algorithms at all.
For a proposal, that context is leverage. A dataset project that connects to NAIRR-style compute, or that positions its data as fuel for one of the Genesis challenge areas — living-systems prediction, materials design, drug discovery — is telling reviewers that it understands the mission, not just the mechanics.
A build plan for the November 4 deadline
Pick the honest tier. Assess whether your data, partnerships, and technical readiness support a Flagship, an Impact, or a Planning grant today. When in doubt, drop a tier — a strong Impact beats a weak Flagship, and a Planning award is a legitimate, strategic first move.
Lead with the data's scientific value, not the AI. The program funds datasets that matter because AI can unlock them. Establish why the data is important and who needs it before you describe the pipelines.
Write the governance section as if it were a specific aim. Ownership, quality control, security, privacy, and community contribution deserve real design, not a paragraph. This is where serious proposals separate from hopeful ones.
Team across sectors deliberately. Broad eligibility exists so data-holders and AI-builders can pair up. If you are only one of those, spend your first weeks finding the other.
Budget the reviews in. Awardees must join bi-annual virtual meetings and budget for annual in-person meetings — a small but real cost and time commitment to include from the start.
The premise of NSF 26-512 is that data, not models, is the scarce input in AI-driven science. If you are sitting on the kind of data the rest of a field wishes it could use, this program was written for you — and the clock runs to November 4, 2026.