1,000+ Opportunities
Find the right grant
Search federal, foundation, and corporate grants with AI — or browse by agency, topic, and state.
Teaching & Learning (T&L) Benchmarks and Datasets Request for Proposals is sponsored by Bill & Melinda Gates Foundation and Learning Commons. This RFP aims to fund the design, development, and validation of three distinct and independent, open-source K-12 instructional data corpora, and an Artificial Intelligence (AI) benchmark built for critical use cases such as Adaptive Learning Experiences & Feedback for Students,…
Get a weekly digest of new grants like this
A free weekly digest of new foundation and federal funding opportunities as they're added to Granted. Unsubscribe anytime.
Or search similar grants →Extracted from the official opportunity page/RFP to help you evaluate fit faster.
Teaching & Learning (T&L) Benchmarks and Datasets Request for Proposals – K-12 AI Infrastructure Program Download the RFP as a PDF Frequently Asked Questions Through this Request for Proposals, we’re seeking to improve the quality and relevance of existing Artificial Intelligence (AI) datasets, evaluations and benchmarks for education.
Like effective human expert judgement, AI benchmarks should be based on multiple data points and grounded in demonstrated education theory. We’re aiming to fund the design, development, and validation of three distinct and standalone, open-source AI openly licensed responses focused on K-12 math education.
These resources will support development and testing of AI model performance compared to human experts in three critical use cases: Adaptive Learning Experiences & Feedback for Students Lesson and Instructional Planning for Teachers Request for Proposal Submissions are due by July 31, 2026 at 11:59PM Anywhere on Earth (AoE) Grant Opportunity at a Glance Teaching & Learning (T&L) Benchmarks and Datasets Up to $5.
5M USD │ 3 awards anticipated; multiple awards to the same respondent are permitted Up to $2M for Adaptive Learning Experiences & Feedback for Students (1 award), up to $1. 5M for Lesson and Instructional Planning for Teachers (1 award), and up to $2M for Teacher Coaching (1 award) Open to organizations or institutions meeting the Demonstrated Experience and Minimum Scale requirements.
See Eligibility section in the RFP instructions for full details. Both the proposal review and award monitoring will be managed directly by the Gates Foundation and Learning Commons teams. The Digital Promise K-12 AI Infrastructure Program will support communication, participation in the K-12 AI Infrastructure community, and dissemination of public goods.
Learning Commons will also support funded projects as a public asset distribution partner, supporting benchmark hosting and dataset deployment. Submit applications using this Qualtrics Form | PDF Version (view only) grant-support@ld-insights.
com Introduction & Program Overview The purpose of this competitive Request for Proposals (RFP) is to fund the design, development, and validation of three distinct and independent, open-source K-12 instructional data corpora, and an Artificial Intelligence (AI) benchmark built for each of the following critical use cases: Adaptive Learning Experiences & Feedback for Students Lesson and Instructional Planning for Teachers Each grant will produce a benchmark as a core public good: a fixed set of pedagogically meaningful tasks that AI models attempt under consistent conditions, scored against rubrics that education subject matter experts have validated.
The benchmark will be run against all major AI models as part of this project to create a public leaderboard comparing performance; giving developers, researchers, and the field a shared view of where AI for education stands on these key use cases. Performance on individual tasks within the benchmark should be measured by discrete evaluators, which provide automated tests of model performance.
A benchmark is only as good as the annotated data that is used to create it. We are seeking “Gold Standard” datasets that come from authentic learning contexts, are large (n > 1,000), are richly annotated for meaningful constructs, include data about learning behaviors and activities, and include learning outcomes measures where possible.
A benchmark is only as good as the annotated data that is used to create it and provides both training and testing data for model performance. These datasets should be annotated for a specific area that educators and subject-matter experts agree represent good practice. The creation and sharing of high-quality datasets with the benchmark are also considered as a key asset created and funded through this project.
The primary subject area is K-12 mathematics. Each project must also include an English Language Arts (ELA) extension or pilot component. A portion of the data corpus, annotation framework, or evaluator validation focused on ELA should demonstrate cross-subject generalizability of the resulting datasets and benchmarks.
For each of these use cases, competitive proposals will develop all of the following: richly annotated datasets, automated evaluators and benchmarks. These public goods will be released using open licenses to facilitate widescale adoption and use.
For purposes of this project, we define a “Gold Standard Dataset” as a large corpus of authentic K-12 instructional data (including elements such as high fidelity-classroom audio, teacher instructional plans, clickstream data, handwritten student work, student-tutor interactions, coach-teacher interactions, and other lesson artifacts) annotated against constructs grounded in demonstrated education theory and validated through empirical human expert decisions.
These corpora are designed as durable, reusable AI infrastructure, applicable across multiple downstream assets: benchmarks, post-training, context optimization, skill.
md development, etc. We define an “AI benchmark” as a rigorous, open-source measurement instrument, built from data within the Gold Standard Dataset, that automatically evaluates whether a machine learning model can match human expert quality in K-12 math education on a variety of pedagogically-relevant tasks using automated evaluators.
The benchmark is the first packaged utilization of the dataset funded under this RFP and the demonstration of its immediate field value. The primary subject area is mathematics, with required ELA extension or pilot components as described in the Program Overview.
We seek to overcome several limitations in the current state of K-12 educational AI infrastructure: Measurement of Pedagogically Relevant Constructs & Tasks – most existing education benchmarks do not measure model performance on pedagogically meaningful learning and teaching tasks; instead they evaluate model performance on standardized exam questions.
Lack of High-Quality Annotated Data – the field lacks large, openly licensed, expert-annotated corpora of authentic K-12 instructional data. This data scarcity limits downstream AI assets for education.
Furthermore, existing datasets often insufficiently represent students furthest from opportunity (e.g. multilingual learners, students with disabilities, high-poverty communities) which can introduce bias and limit the validity, fairness, and generalizability of AI systems.
Benchmarks Highlight Model Differences Not Capabilities – most benchmarks are intended to rank model performance and highlight differences; we seek to identify models that are ready for classroom use; those that can achieve human-level performance on relevant tasks.
The sections below outline the use case(s) for each award track, an example of how that dataset and the benchmark built from it relates to education technologies, an initial set of AI tasks, capabilities, and evidence the data corpus could support, and a benchmark in each award track. The tasks, capabilities, and evidence should be considered as a starting point and not a fixed prescription.
We invite expansion and revision to incorporate the expertise of applicants. Track 1: Adaptive Learning Experiences & Feedback (up to $2M) Customizing learning materials and experiences for students based on their demonstrated abilities is one of the earliest uses of AI in education.
Generative AI has the potential to radically expand this area of work and adapt at the individual student level with new considerations, more frequently and efficiently, and in new subjects and materials. These adaptive systems use AI to adjust the difficulty, sequence, and type of practice problems in real time based on a student’s responses.
They provide instant feedback, guide practice toward mastery, and personalize learning trajectories for each student. This approach operates inside a tight, multi-turn loop: a student attempts something, the system infers what they understand, and the system selects the next move: a hint, a scaffold, a different problem, a worked example, or an explanation pitched at a specific gap.
Ideally these systems can help a student to move into their “Zone of Proximal Development” and create a rich learning experience. Generic AI models perform poorly here for at least two interlocking reasons. They do not durably represent what a particular student knows across turns, and they default to directly providing answers, resolving frustration but bypassing the productive struggle through which learning consolidates.
Further, they do not have broader knowledge about the subject area, reducing the accuracy and specificity of misconceptions diagnoses.
A high-quality dataset for this track, and the evaluators and benchmark built from it, must capture and measure both how accurately a model tracks a student’s evolving understanding and how appropriately it acts on that understanding, including the case in which the best next move is to not provide an additional problem or solution. A middle school student is solving a multi-step algebra problem and writes “3x + 5 = 20, so 3x = 25.
” Current models, asked to help, will typically point out the calculation error and rewrite the line correctly, or restate the rule for isolating a variable. Either move resolves the student’s immediate error but bypasses the question of whether the student understands inverse operations.
A high-performing model would recognize that this was a conceptual error commonly made by students and would assign learning materials and experiences that address this specific error and then reassess the item. A benchmark or gold standard dataset for this track should surface this type of difference in response.
It should measure whether the model is tracking what this student knows across turns, whether the intervention it selects addresses the actual blocker, and whether the model recognizes the cases where holding back is the right call. It should also report these as separable signals, so the field can see whether a model is good at tracing knowledge, good at selecting the next move, or has uneven performance.
Representative Tasks & Evidence The table below describes constructs the dataset/benchmark and evaluators should cover, what good performance looks like, and further considerations. These items are a starting point identified through prior research on adaptive feedback and knowledge tracing; applicants are invited to expand, revise, or reframe them to reflect their team’s expertise and the data they have access to.
Real-time knowledge tracing Predictive accuracy on next-task performance; persistence of state across turns; appropriate recovery from noisy or inconsistent attempts. Does the approach include a knowledge graph or other high-level representation of the domain (and misconceptions)?
What is the unit of analysis (single turn, episode, multi-session arc) and how does the benchmark or dataset capture states that persist or change (both learning and forgetting) across turns? How is tracing accuracy reported separately from intervention quality, so that a model strong at one and weak at the other is legible to the field?
Just-in-time intervention selection The selected intervention addresses the actual blocker rather than a generic difficulty; differentiated selection across student profiles producing errors on the same content.
How does the benchmark or dataset capture the conditional structure of intervention quality (the right move depends on what the student knows, what they have just tried, and what they are likely to be one step from understanding)? How does it score items where multiple interventions could be defensible?
Misconception-aware task selection Misconception classification aligns with knowledge graphs or other validated taxonomies that are interpretable; tasks address the specific misunderstanding rather than re-presenting the original problem. How does the benchmark or dataset establish or build upon a defensible taxonomy of misconceptions?
How does it handle errors that are slips rather than conceptual, and errors that have multiple plausible underlying causes? Distinguish errors due to disengagement Ability to identify errors due to disengagement surfacing as carelessness, lack of effort, gaming the system, overconfident rapid response, or other behaviors.
Can the system accurately identify this type of activity and take corrective action to not adjust student knowledge estimates when a student’s error does not reflect their ability? Cognitive demand regulation Model withholds direct answers when scaffolding is more appropriate; offers worked examples or hints rather than solutions; recognizes unproductive struggle.
How does the benchmark or dataset treat the case where the pedagogically correct action is to withhold support, such as when a student is close to a breakthrough and a worked example would short-circuit their reasoning? Appropriate feedback selection Feedback is appropriate to the grade level of the problem and the domain; does not leak the solution and has appropriate justification for materials provided.
How does the benchmark or dataset evaluate feedback as a whole? This includes its content, its calibration to grade level, and its restraint, rather than scoring only whether the feedback is technically correct?
Quality parity with human judgement AI-generated tasks and activities score within an acceptable margin of expert lessons on a recognized rubric for evaluating the fitness of adaptive materials and experiences, under blind expert review. What recognized rubric or other standards will be used for the parity comparison, and how will expert review be structured to include knowledge estimations and task / activity selection?
Is parity reported as a single composite or as a profile across the rubric’s dimensions? Building on the constructs above, describe how your dataset/benchmark and evaluators will measure the model’s ability to maintain and update a representation of a student’s understanding across multiple interactions and use that representation to select the next instructional move.
Your proposal should explain how the constructs identified (both those listed above and any the team proposes to add or revise) will be operationalized and measured. Track 2: Lesson and Instructional Planning (up to $1.
5M) Lesson and instructional planning is among the most-used and highest-leverage applications of generative AI in K-12 today: teachers routinely use these tools to create lesson plans, build practice sets, scaffold readings, and adapt materials for the students in front of them. The pedagogical challenge is the practical reasoning a skilled teacher uses to move through a curriculum based on what their class actually needs.
Generic models can produce surface-fluent lesson plans, but they struggle to sequence coherently against a learning progression, as identified through a Knowledge Graph or other structured knowledge representation, to identify and target lesson-specific misconceptions. They also frequently do not adapt the modality of a task for English learners or students with learning differences without quietly lowering rigor.
A high-quality dataset and the benchmark and evaluators built from it for this track must capture and evaluate whether the model’s planning decisions hold up against the kinds of decisions an expert teacher would make and defend. Applicants are encouraged to anchor the dataset and benchmark built from it to work with widely adopted curriculum such as Illustrative Mathematics 360.
The benchmark or dataset must also incorporate other high-quality instructional materials (HQIM) and test generalizability beyond a single curriculum, since one that only measures performance against a single curriculum would have limited market utility. For the required ELA extension or pilot component, applicants should anchor to a designated open-licensed ELA HQIM (such as EL Education).
Further, we encourage applicants to anchor scoring in rubrics designed to evaluate the lesson plans teachers actually produce.
The MTSS Center's rubric for high-quality middle school math lesson plans is a strong example: it scores plans against evidence-based and high-leverage instructional practices, such as explicit instruction, scaffolded supports, math discourse, and use of formative data, and was built to give teachers a practical quality check on their own materials.
Applicants are welcome to adapt it, combine it with other validated rubrics, or propose alternatives, as long as the scoring reflects judgments that expert teachers would use. For the required ELA extension, applicants should propose an analogous lesson plan rubric anchored in an ELA HQIM context. An 8th-grade teacher is planning the next week of instruction on linear equations and has a district license to an HQIM curriculum.
She gives the AI tool the standard she is teaching and a copy of last Friday’s exit ticket showing ⅓ of her students are still shaky on integer operations. She asks the model to suggest how to sequence the next four days. Current models might return a confidently written plan that hits the standard but minimizes the integer-operations gap, trusting that students will pick it up in context.
As a skilled teacher with experience in her district’s HQIM, she would make a different call: the gap is broad enough to warrant embedding integer practice inside the new work rather than separating it, and surfacing the underlying misconception driving the gap. She would first look to the existing HQIM to see if it had supporting materials and activities for this issue before seeking or adapting external resources.
A benchmark or gold standard dataset for this track should document instructional artifacts relevant to the instructional plan and measure whether the model’s plan actually responds to the class evidence it was given, whether the adaptations it makes preserve the mathematical demand of the original task, and whether its claims about which standard a task addresses hold up against expert teacher judgment.
It should also consider coherence and integration of curriculum and an adopted HQIM. Representative Tasks & Evidence The table below describes constructs the dataset/benchmark and evaluators should cover, what good performance looks like for each, and further considerations. These items are a starting point identified through research on curriculum design, learning progressions, and standards-aligned instructional materials.
Applicants are invited to expand, revise, or reframe them to reflect their team’s expertise and the data they have access to. Coherent Sequencing & Pathfinding Recommended sequence respects validated learning progressions; pacing adapts to evidence about the specific class; backtracking to precursor skills is justified by visible student need.
How does the input to the model represent prior learning: what evidence about the actual class is supplied, in what form, and how does the benchmark or dataset distinguish plans that respond to that evidence from plans that ignore it? How does it score items where multiple defensible sequences could follow from the same evidence? Is there a knowledge graph or other model used to represent curriculum sequencing and progression?
Misconception-targeted task selection Model identifies common misconceptions associated with a standard; selects or generates tasks that surface or remediate those misconceptions rather than a generic practice item. How does the benchmark or dataset establish a defensible mapping from standards to misconceptions?
How does it handle cases where the model selects a task that is plausible but addresses a different misconception than the one the student is actually showing? Differentiation without losing rigor The modality of the task changes (visual representation added, language simplified, dual-language support, content adapted for student interest) but the underlying work asked of the student is not oversimplified.
How does the benchmark or dataset evaluate differentiation for different populations such as English learners or students with IEPs in a way that detects when a model has reduced the underlying mathematical demand of a task rather than genuinely adapting it? Standards alignment & pacing Alignment matches expert teacher judgment; the model accurately identifies partial matches and flags activities that claim alignment they do not have.
Also has awareness of pacing between standards. How is standards alignment tested against expert teacher judgment at the activity level? How does the benchmark or dataset handle activities that partially align, or that claim alignment they do not have?
Are the activities paced in a way that is reasonable compared to human judgement? Quality parity with human judgement AI-generated lessons score within an acceptable margin of expert lessons on a recognized rubric for high-quality instructional materials, under blind expert review. What recognized rubric will be used for the parity comparison, and how will expert review be structured to avoid order, length, and authorship biases?
Is parity reported as a single composite or as a profile across the rubric’s dimensions? Building on the constructs above, describe how your datasets and the benchmark and evaluators built from them will evaluate the model’s capacity to sequence, prune, and adapt instructional materials for a teacher’s actual class while preserving rigor and standards alignment.
Your proposal should explain how the constructs identified (both those listed above and any the team proposes to add or revise) engage with how the dataset and resulting benchmark and evaluator will produce evidence that is actionable for the teachers and ed-tech developers who would consume it, rather than generating only a single aggregate score.
Track 3: Teacher Coaching (up to $2M) Teacher coaching is the most technically demanding of the three tracks: it requires reasoning across modalities (classroom audio, video, transcripts, observation notes), interpreting that evidence against established frameworks for instructional quality, and conducting a coaching conversation grounded in observations.
Coaching shares pedagogical DNA with tutoring a student and both depend on diagnostic listening, calibrated feedback, and a deliberate choice about when to suggest and when to ask. Applicants are encouraged to draw on relevant work in tutoring evaluation where it transfers.
Effective coaching also requires lesson-specific noticing: an expert coach evaluates instructional moves not in the abstract but against the specific HQIM lesson being taught including its intended student work, discourse structures, and learning objectives. Good coaching ties feedback to how the teacher’s practice enabled (or hindered) enactment of that lesson.
The distinct challenge for this track is what we call the observation-to-coaching bridge: ensuring that the model’s coaching advice is appropriately grounded in the evidence observed in the classroom, and provides the most relevant feedback at that moment, based on observed behavior, rather than offering generic, plausible-sounding feedback that any model could produce without watching the lesson at all.
A new teacher records a 40-minute lesson from a unit on equivalent fractions, where students are meant to construct number-line representations and compare them to articulate why two fractions can name the same point. They upload it to an AI coaching tool for feedback.
During the lesson, the teacher ran an 85% teacher-to-student talk ratio, prematurely intervened when a student started articulating a productive misconception, and skipped the partner-comparison discussion the lesson was built around. Current models typically return generic suggestions like “incorporate more student voice.
” Instead, a skilled coach identified the bypassed partner discussion as the highest-leverage issue, since it was the core mechanism for students to articulate the target reasoning. They recommended a specific revoicing move (“Say more about why they look different — can you show us on your number line?
”) to keep that thinking in the room, and framed a concrete goal for the next observation cycle: protecting at least ten minutes for partner discourse in the next lesson.
A gold standard dataset and the benchmark built from it for this track should measure whether the model surfaces the discourse pattern, recognizes the specific HQIM lesson and its intended student work, connects observed moves to lesson enactment, and delivers actionable, lesson-grounded coaching rather than general praise.
Representative Tasks & Evidence The table below describes constructs the dataset/benchmark and evaluators should cover, what good performance looks like for each, and the further considerations.
These items are a starting point identified through research on teacher coaching, classroom observation, and multimodal discourse analysis; applicants are invited to expand, revise, or reframe them to reflect their team’s expertise and the data they have access to.
Multimodal classroom discourse analysis Evaluation draws on video, audio, and transcripts of the same lesson alongside student work artifacts a coach would typically review (e.g. written work, student-produced representations); accurate diarization and segmentation across these modalities; agreement with established observation protocols (e.g., MQI, CLASS, Danielson) within reasonable bounds.
What classroom evidence will the benchmark or dataset use as input, and what level of multimodal fidelity is required? How will it handle the diarization and segmentation challenges that are known to depress model performance on classroom audio, particularly in classrooms with overlapping speech or non-standard recording conditions?
How will it represent and evaluate the integration of student work artifacts (exit tickets, written responses, problem-solving artifacts) with observed discourse? Actionable feedback generation Feedback is observational and specific; tied to a small number of high-leverage moves; uses non-judgmental language; an experienced coach would endorse it as something they would actually say.
How does the benchmark or dataset distinguish feedback that an experienced coach would deliver from feedback that is plausible-sounding but generic? How does it handle the trade-off between volume of feedback (many suggestions) and weight of feedback (one or two high-leverage moves a teacher can act on)?
Observation-to-coaching coherence Coaching utterances cite specific moments or patterns from the observation; advice would change if the evidence changed; an observer can trace each suggestion back to evidence in the lesson. How does the benchmark or dataset distinguish coaching advice that is genuinely grounded in observed evidence from advice that is plausible-sounding but generic?
Does it contain counterfactual lessons where evidence-grounded coaching should change accordingly, so that the benchmark or dataset can detect a model that produces the same feedback regardless of what it observed? Coaching-stance calibration Model selects an open question when teacher reflection is more productive; offers a direct suggestion when stakes or time constraints warrant; avoids excessive scaffolding for experienced teachers.
How does the benchmark evaluate the choice among asking, suggesting, and holding back, since high-quality coaching depends on knowing when to probe versus when to provide an immediate course of action? How are items scored where multiple stances could be defensible?
Longitudinal coaching coherence Coaching advice builds on prior cycles rather than resetting; the model tracks what the teacher was working on previously and notices progress or regression on prior goals; feedback across multiple observations shows a coherent arc rather than disconnected critiques.
How does the benchmark or dataset represent the multi-cycle nature of coaching, in which an effective coach returns to a teacher with reference to last session’s goals? How are items constructed to test whether a model is genuinely tracking the coaching relationship rather than treating each observation as the first? How is “progress” on a teacher’s prior goal scored, given that observable change in practice often takes multiple cycles?
Quality parity with expert human plans AI coaches are scored within an acceptable margin of error on a recognized rubric for teacher coaching, under blind expert review. What recognized rubric will be used for the parity comparison, and how will expert review be structured to avoid potential expert bias and congruence between expert judgement? Is parity reported as a single composite or as a profile across the rubric’s dimensions?
Building on the constructs above, your proposal should describe: how your dataset and the benchmark and evaluators built from it will jointly evaluate the model’s analysis of classroom evidence and the coaching conversation the model produces about that lesson, with particular attention to how observation-to-coaching coherence will be measured (including whether feedback is grounded in both observed instructional evidence and the specific HQIM lesson’s design and objectives).
Your proposal should explain how the constructs identified (both those listed above and any the team proposes to add or revise) engage with the multimodal data and annotation challenges specific to classroom observation. Proposal Guidance & Evaluation Criteria The proposal form is organized into the sections below, each corresponding to one or more of the four evaluation criteria.
Please abide by the character limits in the individual sections below. The guidance provided for each section is detailed, but your proposal does not need to address every sub-question separately; concise, integrated responses are welcome. Proposals will be reviewed holistically.
The proposal submitted by the most qualified team demonstrating the most promising approach and strongest overall value for the investment will be selected.
The table below maps each form section to the evaluation criteria reviewers will apply: Eligibility & Qualifying Experience → Threshold screen (not scored) Overview, Dataset + Benchmark Track-Specific Tasks & Evidence → Criterion 1: Significance Team Background, Key Personnel, Data Acquisition Plan → Criterion 2: Assets and Capabilities Project Plan, Measurement & Evaluation, Responsible AI & Safeguards → Criterion 3: Project Workplan Confirmation and Dissemination Plan, Targeted Universalism, Global Access & Digital Public Goods → Criterion 4: Release and Dissemination Plan Eligibility & Qualifying Experience This section is a threshold screen and is not scored by peer reviewers.
Reviewers will use it to verify that the team meets the Demonstrated Experience and Minimum Scale requirements before the proposal is forwarded for review. Provide a brief statement (2,000 characters or fewer) describing your experience and resources in the following areas.
Please include: Prior experience producing or curating large, expert-annotated open datasets in education or a closely adjacent domain, ideally publicly released and available prior to June 1, 2026 (RFP release) Prior experience building and deploying automated evaluations of LLM outputs, ideally publicly released and available prior to June 1, 2026 (RFP release) Pedagogical expertise in the use case area of the chosen track, with primary expertise in mathematics and demonstrated capacity (in-house or through partnership) for the required ELA extension Please describe the datasets you will produce or aggregate for the data corpus and use to develop the benchmark, including data sources, annotation schema and scale, and timeline to collect/annotate them (if not already developed) Criterion 1: Significance (WHY) Combined Limit (Overview and Track-Specific Prompt): 5,000 characters.
This criterion asks: why is this project important, why now, and why this team? Your proposal should cite specific evidence for each claim and connect the proposed approach directly to documented limitations in current models, to theory for effective practices in each domain and to datasets available to model/test these practices in AI outputs.
Describe the proposed work: what you will build and how your approach addresses known limitations in K-12 educational AI infrastructure (both the dataset gap and the evaluation benchmark gap) for the use case in your chosen track. Draw on the failure modes and capability gaps described in the Background & Context section.
Domain Validity in K-12 Education & Hallucination Awareness (Track-specific): Frontier models continue to climb general-purpose benchmarks (MMLU, GSM8K, expert-level reasoning suites) while their performance on authentic educational tasks such as cognitive demand regulation, quality parity with human plans, and multimodal classroom discourse analysis, remains comparatively low and uneven.
Datasets and benchmarks anchored in authentic K-12 tasks and grounded in how expert educators define quality and the benchmarks built from them can help close the gap so that progress on benchmarks is reflected in improved student and teacher experiences. These should ensure fundamentally that the knowledge being represented is accurate and is not created through an AI model hallucination.
Please refer to the Benchmark + Dataset Tracks section relevant to your proposal above for more details. Criterion 2: Assets and Capabilities (WHAT AND WHO) This criterion asks: does the team have what it needs to succeed? Reviewers will look for concrete evidence of existing assets and demonstrated team capacity, not aspirational descriptions of what the team plans to develop.
Detail how your team, advisors, and partners provide the technical expertise to create a high-quality gold standard dataset and a rigorous, reproducible benchmark. How
According to the current listing, eligibility includes: Organizations or institutions meeting demonstrated experience and minimum scale requirements. Demonstrated experience in AI evaluation methodology, scoring harness design, and reproducible benchmark construction. Confirm the full requirements in the official notice before applying.
Teaching & Learning (T&L) Benchmarks and Datasets Request for Proposals is funded by Bill & Melinda Gates Foundation and Learning Commons. Verify program details on the funder's official page before applying.
Start from the official opportunity page linked in this listing — it carries the sponsor's submission instructions.
MacKenzie Scott's Yield Giving has moved more than $26 billion — including roughly $7 billion in 2025, over a third of all U.S. megagifts — through unrestricted, trust-based grants with no application. Melinda French Gates's Pivotal Ventures has committed $2 billion to women's health and economic power. The Bezos philanthropies have moved billions more. As the Buffett–Gates era winds down, this new class of megadonor is rewriting the rules of major giving. Here is what changes for nonprofits, and how to become the kind of organization this money finds.
Read articleThe Gates Foundation committed at least $1 billion over two years to equitable AI alongside its 10th Goalkeepers Report — 40% education, 40% health, 10% agriculture, 10% digital infrastructure. Five days later, 60 organizations signed a five-year goal to reach 3.4 billion speakers of underrepresented languages. Here is how implementing nonprofits should read both.
Read articleGates committed $540.2 million over a decade to the Institute for Health Metrics and Evaluation — the largest gift in University of Washington history — while IHME's federal grant income falls from $7.3 million to a projected $4.8 million. The gift is a case study in the anchor-funder model: what it can do that federal money cannot, and the concentration risk it quietly underwrites.
Read article