1,000+ Opportunities
Find the right grant
Search federal, foundation, and corporate grants with AI — or browse by agency, topic, and state.
This listing may be outdated. Verify details at the official source before applying.
Find similar grantsAlignment Research Engineer Accelerator — AI Safety Technical Program (2025) is sponsored by Coefficient Giving (formerly Open Philanthropy). Supports mission-aligned projects and measurable outcomes in AI alignment research.
Get a weekly digest of new grants like this
A free weekly digest of new foundation and federal funding opportunities as they're added to Granted. Unsubscribe anytime.
Or search similar grants →Extracted from the official opportunity page/RFP to help you evaluate fit faster.
Research directions Open Phil wants to fund in technical AI safety — LessWrong Grants & Fundraising Opportunities AI Frontpage Research directions Open Phil wants to fund in technical AI safety by jake_mendel , maxnadeau , Peter Favaloro AI Alignment Forum Linkpost for www. openphilanthropy. org 70 min read The Open Philanthropy has just launched a large new Request for Proposals for technical AI safety research .
Here we're sharing a reference guide , created as part of that RFP, which describes what projects we'd like to see across 21 research directions in technical AI safety. This guide provides an opinionated overview of recent work and open problems across areas like adversarial testing, model transparency, and theoretical approaches to AI alignment.
We link to hundreds of papers and blog posts and offer approximately a hundred different example projects. We hope this is a useful resource for technical people getting started in alignment research. We'd also welcome feedback from the LW community on our prioritization within or across research areas.
For each research area, we include: Discussion of key technical problems and why they matter Related work and important papers in the field Example projects we'd be excited to see Specifications of what we think would make for a good research project, and what we're looking for in proposals. Applications ( here ) start with a simple 300 word expression of interest and are open until April 15, 2025.
We have plans to fund $40M in grants and have available funding for substantially more depending on application quality. In this section we briefly orient readers to the 21 research areas that we’ll discuss in more detail below. For ease of consumption, we’ve grouped them into 5 rough clusters, though of course there is overlap and ambiguity in how to categorize each research area.
Our favorite topics are marked with a star (*) – we’re especially eager to fund work in these areas. In contrast, we will have a high bar for topics marked with a dagger (†). Adversarial machine learning This cluster of research areas uses simulated red-team/blue-team exercises to expose the vulnerabilities of an LLM (or a system that incorporates LLMs).
Across these directions, a blue team attempts to make an AI system adhere with very high reliability to some specification of its safety behavior, and then a red team attempts to find edge cases that violate the specification. We think this adversarial style of evaluation and iteration is necessary to ensure an AI system has a low probability of catastrophic failure.
Through these research directions, we aim to develop robust safety techniques that mitigate risks from AIs before those risks emerge in real-world deployments. *Jailbreaks and unintentional misalignment : New techniques for finding inputs that elicit competent, goal-directed behavior in LLM agents that the developers clearly tried to prevent.
We’re especially interested in inputs which might arise organically over the course of deploying an LLM agent in an environment. *Control evaluations: Control evaluations are a way to stress-test systems for constraining and monitoring AIs, to ascertain whether misaligned AIs could collude with one another to subvert human oversight and achieve their own goals.
We’d like to support more such evaluations, especially on scalable oversight protocols like AI debate. *Backdoors and other alignment stress tests: Past research has implanted backdoors in safety-trained LLMs and tested whether standard alignment techniques are capable of catching or removing them. We’re interested in more research on this, and other “stress tests” of today’s state-of-the-art alignment methods.
*Alternatives to adversarial training : Adversarial training (and the rest of today’s best alignment techniques) have failed to create LLM agents that reliably avoid misaligned goals. We’re interested in techniques like latent adversarial training and circuit breaking that might succeed where standard adversarial training falters.
Robust unlearning: One idea for reducing AI risks is to remove models' knowledge of potentially dangerous topics, such as cybersecurity exploits or virology. We’re looking to fund better unlearning techniques that can totally erase specific information from a model’s weights.
Exploring sophisticated misbehavior in LLMs Future, more capable AI models might exhibit novel failure modes that are hard to detect with current methods – for instance, failure modes that involve LLMs reasoning about their human developers or becoming optimized to deceive flawed human assessors.
We want to fund research that identifies the conditions under which these failure modes occur, and makes progress toward robust methods of mitigating or avoiding them. *Experiments on alignment faking: Alignment faking is a speculative failure mode in which an AI pretends to be aligned to its training/testing process in order to achieve some other goal once deployed.
We’re interested in further work on what factors reinforce or prevent alignment faking. *Encoded reasoning in CoT and inter-model communication: People sometimes give LLMs access to chains of thought that they can use to “talk to themselves” while they reason through a problem, or grant LLMs the ability to query other LLM instances to answer some question.
In either of these settings, there’s a chance that the LLMs will write messages that encode meaning beyond the natural language definitions of the words used. We’re interested in research that studies the conditions under which this phenomenon occurs and/or develops methods to catch and prevent it.
Black-box LLM psychology : Research that systematically answers questions like “ Do LLMs have stable values/goals/preferences that transcend specific inputs? ” and “ Which rules are LLM agents happy to break, and which are they more committed to? ”.
Evaluating whether models can hide dangerous behaviors : Prior work has raised the concerns that future models might “sandbag” (act less capable on dangerous capability evaluations, to avoid raising concerns) or take “treacherous turns” (wait until they are granted power and autonomy before disempowering monitors to achieve harmful goals).
We’re interested in evaluations that can test whether an AI has the capability to execute these strategies (regardless of whether it actually does so). Reward hacking of human oversight : A common concern in AI safety is that AIs might trick their human raters into assigning them high scores that they don’t deserve.
We’re interested in research that investigates this phenomenon empirically: under what conditions does it occur in practice? How problematic is it? We see potential in the idea of using a network’s intermediate representations to predict, monitor, or modify its behavior.
Some approaches are feasible without an understanding of the model’s learned mechanisms, while other techniques may become possible with the invention of interpretability methods that more comprehensively decompose an AI’s internal mechanisms into components that can be understood and intervened on individually.
We’re interested in funding research across this spectrum — everything from useful kludges to new ideas for making models more transparent and steerable. *Applications of white-box techniques: Real-world applications of interpretability have so far been limited, and few instances have been found where interpretability methods outperform black-box methods.
We’re interested in funding research that leverages interpretability insights to make progress on useful and realistic tasks, including: model steering, capability elicitation, finding adversarial inputs, robust unlearning, latent adversarial training, probing, and low-probability estimation. Activation monitoring : Probes on a model’s internal activations are one strategy for catching AIs taking subtly harmful or misaligned actions.
We’re interested in research that tests how useful probes are for monitoring LLMs and LLM agents. Finding feature representations: One challenge in understanding what happens in neural networks is that the latent variables (“features”) in the algorithms they execute are not easily visible from their activations.
We’re interested in funding research that helps us find which features are being represented in a model’s internals, with a focus on diversifying beyond sparse autoencoders, currently the most widely-studied approach. Toy models for interpretability : By “toy models”, we mean small, simplified proxies that capture some important dynamic about deep learning.
We want to support the development of better toy models that distill challenges in understanding the internals of frontier LLMs. Interpretability benchmarks: We’d like to support more benchmarks for interpretability research. A benchmark should consist of a set of tasks that good interpretability methods should be able to solve.
Our goal is to create concrete, standardized challenges to better compare interpretability techniques and accelerate progress in the field. Externalizing reasoning: It could be safer to have much smaller language models which put more reasoning into natural language. We’re interested in techniques for training a language model that is very weak over a single forward pass, but much stronger when it reasons with long chains of thought.
†More transparent architectures: It may be possible to design new language models that are much more interpretable than the current mainstream. We’re especially interested in attempts to make models conduct their reasoning in natural language, which could be done by pushing much more reasoning into the chain of thought or replacing parts of a forward pass with natural language queries.
Proposals don’t have to be competitive with the main paradigm, but they should aim to build at least a Pythia-level model. Trust from first principles We trust nuclear power plants and orbital rockets through validated theories that are principled and mechanistic, rather than through direct trial-and-error. We would benefit from similarly systematic, principled approaches to understanding and predicting AI behavior.
One approach to this is model transparency, as in the previous cluster. But understanding may not be a necessary condition: this cluster aims to get the safety and trust benefits of interpretability without humans having to understand any specific AI model in all its details.
*White-box estimation of rare misbehavior: AIs may only exhibit egregiously bad behaviour in scenarios that are extremely rare before deployment and very hard for us to find by search over inputs, but which may be common once in deployment.
We’re interested in funding research that leverages knowledge about the structure of a model’s activation space to efficiently estimate the probability of some particular rare output, even when that probability is too small to estimate by random sampling.
Theoretical study of inductive biases: We are particularly interested in theory-driven work that can shed light on why models generalize well or poorly in different cases, on the likelihood of scheming arising, and on how a model’s internal structure develops over the course of training. Alternative approaches to mitigating AI risks These research areas lie outside the scope of the clusters above.
†Conceptual clarity about risks from powerful AI: It is extremely challenging to reason well about the risks that AGI and ASI will bring, or about which research approaches show the most promise for mitigating these risks. We are interested in funding conceptual research that helps the world think more clearly about future AI risks, and about what needs to be done to avoid them.
† New moonshots for aligning superintelligence: It’s possible that none of the approaches currently under discussion will be sufficient for aligning superintelligent (as opposed to near-human-level) systems. Therefore, we’re also interested in funding entirely new research agendas that take a novel approach to aligning superintelligent systems. Proposals should be clear on how their agendas aim to avoid or mitigate scheming.
*Jailbreaks and unintentional misalignment LLMs are often capable and helpful, but there are still plenty of inputs where they violate their developers’ specifications. We want to support work that searches for inputs on which LLMs violate these specifications — particularly competent, goal-directed violations.
Two lines of recent work have looked for undesirable behaviors in LLMs, approaching the problem from two different angles: Andriushchenko et al. , Kumar et al. , and @elder_plinius , among others, have demonstrated that, on certain adversarial inputs (specifically: jailbreaks ), LLM agents will competently pursue user-provided goals that the agents’ developers attempted to prevent them from pursuing.
OpenAI , Yuan et al. , Järviniemi and Hubinger ( § 4) , and Meinke et al. (§3.
6) display inputs on which LLMs take undesirable/misaligned actions without being instructed to do so. The inputs in these papers are intended to be the sort of inputs that could have been given to an AI accidentally and without any intent to elicit objectionable behavior (though they retain some artificiality).
We think that studying these inputs sheds light on how well today’s alignment techniques work for instilling rules and values into frontier models. We also think that being able to effectively instill rules into LLMs will become increasingly high-stakes as LLM agents grow more capable and are deployed more widely.
Proposals should focus on inputs that make AIs competently pursue undesirable goals (with or without explicit instruction to do so), not inputs where they merely hallucinate, “misunderstand”, make a “ mistake ”, or get derailed in pursuit of some benign goal.
There’s no bright line between competent and incompetent violations, but we have in mind cases in which an LLM takes actions that accomplish/are optimized for some goal (over multiple points in time, or across variations in the situation). We think this distinction is important because we expect incompetent failures to resolve over time as AI developers make their models more generally capable.
This criterion has some commonalities with the focus in Debenedetti et al. on “targeted attack success rate (i.e., does the agent execute the attacker’s malicious actions)”, not just derailing the agent, and with the Harm Score in Andriushchenko et al.
Most work involving jailbreaking, red-teaming, or adversarial attacks typically requires the red-teamer to get the model to comply rather than refusing, or to answer a single question correctly rather than incorrectly. In contrast, a focus on competent violations demands something stricter: flexible, goal-optimized behavior.
Agents: The division between competent and incompetent failures is probably most clear in the agent setting, where one can more clearly demonstrate that the agent’s actions, across multiple timesteps or variations in the setting, are in pursuit of some goal. So we prefer work that operates on, or at least can be applied to, LLM agents.
Unambiguous misbehavior: We’re mostly interested in behaviors that the model’s developers explicitly tried to forbid, e.g. tasks that trigger a refusal if you request them directly or tasks that are clearly prohibited in the relevant model spec , rather than edge cases in which it’s genuinely unclear how the model ought to behave.
We prioritize this because such behaviors provide useful evidence about how effective AI developers are at instilling certain behaviors into models when they’re actually trying. For similar reasons, we’re more interested in studying whether models can reliably abide by a rule or goal that is trained into the model, rather than presented in the prompt (as is typically studied in work on prompt injection).
Unintended misbehavior: We care more about inputs that accidentally trigger misaligned behaviors (like a more competent and goal-directed version of OpenAI , Yuan et al. , Järviniemi and Hubinger ( § 4) , Meinke et al. (§3.
6) ) than we do inputs that include instructions to do some harmful task (as in Andriushchenko et al. and many other jailbreaking papers). That is to say, we’re more interested in safety (i.e. what violations arise even when no one is deliberately trying to get the model to misbehave) than security (i.e. can an adversary cause the model to violate its developers’ specifications).
Example projects: We think there are lots of frontiers on which to expand beyond the prior work we cited above. For example: A proposal to look for real-world inputs that trigger undesirable goal-directed behaviors in LLM agents, of the sort that have been mentioned anecdotally here and here and elsewhere. The challenge here would be finding a large enough dataset of agent trajectories to study.
A proposal that extends the examples in Yuan et al. and Ruan et al. but focuses on “unintended risks” (as Yuan et al.
puts it) that are more goal-directed/competent than the examples those papers find (which generally seem more like “honest mistakes” or hallucinations). A proposal that further investigates the sort of inputs studied in Järviniemi and Hubinger ( § 4) , Meinke et al. (§3.
6) , Scheurer et al. , and Greenblatt et al. (§3) , i.e. situations where models appear to be, without prompting, deliberately deceiving their human supervisors.
We think there are many ways to extend and/or interrogate these case studies.
For example, you could ablate/alter various parts of the inputs to understand what triggers “deceptive” behaviors, investigate whether these failure modes are likely to cause problems in practice, study how persistent and goal-directed the behaviors actually are, look for cases whether agents go beyond deception to sabotage, try to measure whether the models are “ just role-playing ”, or expand the scope/rigor in some other way.
A proposal to build more realistic and more difficult harmful tasks along the lines of AgentHarm (especially if using the Inspect framework) and recalculate attacker success rate on those new tasks. Anecdotally, we’ve heard that models are easier to jailbreak when the tasks are unrealistic (e.g., contain “John Doe”, use unrealistic website interfaces, etc.).
Other proposals to build new evaluations for competent violations, which would ideally measure avoidance of competent violations and degradation of general model capabilities in a realistic setting. A concept we’ve mentioned a few times in this section is that some inputs that trigger bad behaviors are natural/unintentional/accidental .
One way to find this kind of input would be to look through inputs that some model/agent was actually run on, in a real-world setting that excludes adversarially generated inputs, leaving only inputs where no one is trying to cause the model to misbehave. Another way could simply be to use low perplexity.
We’re interested in work that tries to develop this concept more thoroughly and quantitatively and considers the pros and cons of different definitions. This is one of our favorite research topics, so we’re especially interested in funding work in this area. We’re interested in supporting more research conducting “control evaluations,” as demonstrated in Greenblatt et al.
, Wen et al. , and Bhatt et al. (forthcoming) and as described here and here .
Control evaluations are an attempt to conservatively evaluate the safety of protocols like AI-critiquing-AI (e.g., McAleese et al. ), AI debate ( Arnesen et al. , Khan et al.
), other scalable oversight methods , and other system architectures for monitoring and/or constraining AIs’ outputs to detect and prevent unsafe actions (e.g., Naihin et al. , Willison ). Control evaluations assess the question, “if some or all of the AIs in this protocol were trying to stealthily sabotage their developers, would they succeed?
”. Concretely, control evaluations consist of a red team/blue team game, where the red team controls the “misaligned” AIs, and the blue team tries to use the red team’s models productively or at least catch them in attempted sabotage. For more detail on the control evaluation methodology, we recommend reading Greenblatt et al.
There are multiple reasons to presume an LLM is untrustworthy and malicious, including concerns about prompt injection, poisoning , and organically arising misalignment .
We’re most interested in misalignment, though we think control evaluations have relevance for all three concerns — there is significant overlap between control techniques/evaluations and the defenses/evaluations in ML security work on prompt injection, see e.g., OWASP , Debenedetti et al. , and Harang .
That said, the misalignment threat model leads to different methodological choices from these other concerns; for instance, control evaluations typically require the red team to succeed at some act of sabotage before getting caught, which makes the blue team’s job much easier than preventing prompt injection or jailbreaks . Nevertheless, some security research uses a similar methodology, like Debenedetti et al.
For more reading on how control evaluations apply to these other failure modes, we encourage you to read Clymer’s “Extending control evaluations to non-scheming threats ”. Proposals for control evaluations should include the following: Task distribution: The control evaluation should be run on tasks which are similar to economically (or otherwise) valuable tasks, but also present an opportunity for a malicious model to cause harm.
Control measures: A blue team that constructs a protocol/pipeline/system architecture for using AIs to complete a set of tasks. The blue team’s goal is to have the resulting system complete the tasks correctly, without being sabotaged.
Strong red team: A red team (which could be the same people as the blue team) gets to create an “attack policy” that attempts to achieve a (harmful) side objective before getting turned off by the blue team. But there are many ways to design the game and define red or blue team victory conditions within this high-level structure. No easy solution: We’re especially interested in settings (like in Greenblatt et al.
, Wen et al. , and Anonymous ) in which the blue team can’t win by merely implementing OWASP-style security mitigations , and has to rely on human oversight or oversight by other AIs instead. That said, see Willison on why these settings can be difficult for the blue team.
Assess existing scalable oversight protocols (e.g., McAleese et al. , Arnesen et al. , Khan et al.
) with the rules and assumptions of control evaluations. Extend Greenblatt et al. to a more realistic and/or more agentic setting.
Evaluate the efficacy of human auditing as a control measure. Run control evaluations in which the red team gives fewer hints to the attack policy, as in Mallen et al. Consider the control project that Joe Benton proposes here .
This is one of our favorite research topics, so we’re especially interested in funding work in this area. *Backdoors and other alignment stress tests One line of evidence about how future systems might be misaligned comes from studying artificially backdoored models, also known as Trojans or sleeper agents .
These models act like a “ challenge trial ” to test our best methods for aligning and monitoring AIs: how effective are we at catching or fixing models that are deliberately constructed to misbehave on certain inputs? We’re interested in funding research on both attacks and defenses in the backdoor setting.
The arms race between attacks and defenses in prior work on backdoors in safety-trained LLMs has led to substantial research progress. Hubinger et al. and Cao et al.
demonstrate backdoored models that persist through safety training. , but In response, MacDiarmid et al. , Sheshadri et al .
, and Zeng et al. make progress on catching and/or removing these backdoors. Meanwhile, Price et al.
and Bailey et al. study a setting where the backdoor trigger is semantic , rather than syntactic , showing that further safety training can remove the backdoor. We’re interested in funding research that continues this dialectic.
Our desiderata for research in this area (see below) are motivated by the threat model of organically arising misalignment, and also by the model-poisoning threat model (as discussed in e.g., Feng and Tramèr , OWASP ). For these threat models, it makes sense to grant the red team full control over the model weights.
Studies with limited attacker control, such as those restricted to RLHF label manipulation, will be considered a lower priority or potentially out of scope for our grantmaking. For work on attacks, i.e., making backdoored models: Strong red team: We’re focused on the fully white-box threat model, i.e., the attacker has arbitrary access to the model weights.
Hard to remove: We’re interested in work that tries to construct backdoors that are difficult to detect or remove, e.g., they persist through standard safety training. Resistant to adversarial training: Ideally, the backdoors would not only be robust to standard safety training, they’d also be robust to adversarial training and to bespoke countermeasures such as those studied in MacDiarmid et al. , Sheshadri et al .
, and Zeng et al. Competent violations: We’re most interested in backdoors that trigger behavior that is harmful and competent (like in Rando and Tramèr ), not just low-quality or useless (like in Price et al. ).
One way to achieve this goal is to operate in the LLM agent setting rather than the chatbot setting. See the Jailbreaks and Alternatives to Adversarial Training sections for more discussions of this property. Semantic trigger: We’re interested in further work along the lines of Price et al.
that uses a distribution shift — rather than a simple phrase — as a backdoor trigger. Strong red team: We’re focused on the fully white-box threat model, i.e., the attacker has arbitrary access to the model weights. This means that the defender cannot prevent the backdoor from being inserted in the first place.
Run-time detection: We’re most interested in detecting backdoored inputs at inference time (as studied in MacDiarmid et al. ). While other defense goals — such as distinguishing backdoored and clean models, removing backdoors from a given model, or reconstructing the backdoor trigger from the model weights — are valuable research areas, their difficulty makes them a lower priority.
With sufficient constraints on the attacker, these approaches could warrant investigation. Hubinger et al. find that adversarial training fails to remove backdoors and, in some cases, makes them worse.
Does this phenomenon replicate? How robust is it to small changes in the experimental set-up? Take techniques that have been used to identify models with backdoors and adapt these techniques to remove those backdoors.
Relatedly, test whether the technique in Mack and Turner can be adapted to remove undesirable behaviors from models. Make backdoored models that cannot be detected/fixed by SOTA defense methods like MacDiarmid et al. , Sheshadri et al .
, and Zeng et al. This is one of our favorite research topics, so we’re especially interested in funding work in this area. *Alternatives to adversarial training As mentioned earlier , researchers have failed to make LLMs consistently comply with developer-provided rules like “ never help users make a bomb ” or “ never deliberately sabotage your developer ”.
We’re interested in supporting new ideas for mitigating worst-case behavior of LLMs. Motivation: If LLM agents are to be deployed in high-stakes settings, we’d like to ensure they comply with human-specified safety rules. In particular, we’d like to make sure they don’t violate our rules in ways that make it difficult to detect the violation or to rein them back in.
Adversarial training is a natural approach to this problem, but when applied to LLMs, it generalizes poorly and degrades general capabilities. We want to support new ideas beyond adversarial training for improving the left tail of model behavior, perhaps building off these three promising lines of recent work: Xhonneux et al. , Sheshadri et al .
, Zeng et al. , and Kireev et al. , among others, demonstrated new methods for removing worst-case behaviors by searching for hidden activations that trigger such behaviors, instead of searching over input space.
Sheshadri et al . refer to this family of techniques as “latent adversarial training” or “LAT” . Zou et al.
demonstrate an alternative to adversarial training for LLMs and LLM agents that uses a set of off-policy harmful behaviors. The authors successfully improve robustness against a wide range of adversarial attacks. Guan et al.
develop a method (which they call “deliberative alignment”) that attempts to increase adherence to the model’s specification by training the model to reason through the implications of the model spec for its response to the current query .
Single-model robustness: Eligible research should improve the worst-case behavior of a single model, rather than using a pipeline involving pre/post-processing or monitoring around a model to catch attacks. We think there’s value in using other models to catch when LLMs misbehave (and we hope to fund some such work through the Control Evaluations section of this RFP).
But these monitor models will also have defects, and we think it’s risky to rely too heavily on monitors when human adversaries or policies themselves might come up with strategies to evade them. We’re focused on approaches that make the models never try to violate the rules in the first place.
Competent violations: We think it might be easier to rule out competent violations (see earlier section for more discussion) than it is to rule out the broader set of failure modes including hallucinating, “misunderstanding” the situation, or getting derailed in pursuit of some benign goal. One example of a paper that aims to rule out competent failures is Wang et al. , which focuses on ruling out competent assistance with bomb-making.
We’re interested in more research in this setting, and eventually in the setting of competent violations more broadly. For more discussion of this criterion, we recommend reading Ziegler ’s argument that future robustness projects should “change the task so that catastrophes require competence”. No input-space search: One very nice advantage that techniques like Sheshadri et al .
and Zou et al have over adversarial training is that you never have to find inputs that successfully jailbreak the model. This is appealing given that a thorough search to find all inputs (or even classes of inputs) that trigger noncompliant/misaligned behaviors may eventually become infeasible.
We’re most excited about worst-case compliance/alignment techniques that don’t require searching over token-space, but we’re open to proposals that do. Make a model that never assists users make bombs, as defined in Wang et al. Extend one of the aforementioned papers to new LLM agent settings and/or to larger models.
Further study the generalization properties (to new classes of attack) of any of the prior work in this area, especially in the LLM agent setting. LAT techniques must balance giving the adversary enough flexibility that they can discover problematic behaviors that input-space adversaries can’t, but enough constraints to avoid finding latent states that don’t correspond to any realistic input (see Jermyn for more details).
Current approaches simply search in an epsilon-ball near known activation vectors. This is limited because some inputs could correspond to activations that lie far from our training data’s activation vectors — indeed, this may be particularly likely for inputs which cause rare egregious behavior. Therefore, we’re interested in supporting work to improve LAT via improved modeling of activation vectors .
For example, a project could replace the standard epsilon-ball regularization with a different regularization, such as an ellipsoid that is scaled by the principal values of the
According to the current listing, eligibility includes: Individuals and organizations involved in AI alignment research. Confirm the full requirements in the official notice before applying.
The current listing shows up to $1,000,000. Verify award ceilings, matching requirements, and allowable costs in the official notice.
Alignment Research Engineer Accelerator — AI Safety Technical Program (2025) is funded by Coefficient Giving (formerly Open Philanthropy). Verify program details on the funder's official page before applying.
Start from the official opportunity page linked in this listing — it carries the sponsor's submission instructions.
Coefficient Giving (formerly Open Philanthropy), the largest funder in existential risk reduction, has issued a Request for Proposals for Technical AI Safety Research focused on alignment, interpretability, and control of advanced AI systems. The RFP accepts proposals on a rolling basis with no fixed deadline. Coefficient Giving expects to spend roughly $40 million on this initiative and is open to spending substantially more depending on application quality. The program funds a wide range of activities including fundamental alignment research, interpretability and mechanistic understanding of AI systems, AI control and containment strategies, evaluation and benchmarking of AI safety properties, scalable oversight methods, and seed funding for new AI safety research organizations. This is distinct from Coefficient Giving's separate AI Governance Research RFP which focuses on policy and governance questions.
Coefficient Giving (the rebranded Open Philanthropy) is funding technical research that advances the science of building safe and trustworthy AI systems, including understanding and preventing frontier AI system misalignment, creating valid evaluations and interventions, and developing oversight mechanisms for superhuman AI capabilities. The RFP is open to proposals of many sizes and purposes, ranging from rapid funding for API credits and compute to discrete 6-24 month projects, academic start-up packages, and seed funding for new research organizations. Coefficient expected to deploy roughly $40 million over the initial period, with capacity to spend substantially more depending on application quality.