DARPA's Biggest Single-Tranche AI Award This Cycle Asks You to Build a Stock Market. DV026 Pays $1.8M — and Bans You From Looking Inside the Model.

September 2, 2026 · 6 min read

Granted Research Team · Editorial policy

Almost every AI evaluation contract written in the last three years starts the same way: define the bad behavior, then test for it. Write a taxonomy of harms, build a red-team suite, score the model against the list. The method's weakness is structural — you can only find what you thought to name.

DARPA just posted a $1.8 million SBIR topic built on the opposite premise. Don't define harm. Build a market, put the AI in it, give it incentives, and watch what it actually does.

DPA26BZ06-DV026, "Influence Benchmarks for AI Systems," pre-released September 2, 2026 as part of Release 6 of the FY26 Department of War SBIR Broad Agency Announcement. It opens September 23 and closes October 21, 2026 on DSIP. And on a release containing seven SBIR topics, it carries the largest single committed tranche of any of them.

The award structure is a signal, not a detail

Phase I is $300,000 over 12 months, delivered as a 10-page white paper and a 5-page slide deck, with milestones at Months 2, 6, 9, and 12.

Direct to Phase II is $1,800,000 over 24 months with no option periods. That phrasing does real work. Most DARPA Direct-to-Phase-II topics on this release split the money — DV025 runs $1,000,000 with a $500,000 option, DV028 and DV029 each run $1,500,000 with a $500,000 option. Options are how a program manager buys the right to stop. DV026 has none. The entire amount is committed up front against a 24-month schedule.

That is a program office saying it does not want a mid-course off-ramp — and it is also a warning to the proposer. There is no second bite. Your Month 24 deliverable has to be defensible from the cost volume you write in October, with no renegotiation point built in.

What DARPA is actually buying

The technical core is a simulated auction environment used as a behavioral instrument. Agents interact through a defined decision space — the topic describes it as "a set of bids, each defined by the time, asset, quantity, and price offered or requested" — inside a market that also carries a dynamic "news feed" injecting market-relevant information over time.

The methodological claim is that economic frameworks reveal latent preferences without requiring you to specify them first. The solicitation is explicit that the approach "does not rely on a priori definitions of 'harmful' or 'helpful' traits." You do not ask the model whether it would deceive. You give it a reason to, and price the outcome.

Then comes the constraint that defines the whole architecture: black-box access only. Queries and outputs. No weights, no activations, no logprobs, no fine-tuning. If your evaluation methodology assumes any degree of internal access, you are not responsive to this topic — and the constraint is deliberate, because it is what makes the instrument work equally on a commercial API and on an adversary's model you will never be given.

Phase I carries exactly one hard numeric target, at Month 12: greater than 90% allocative efficiency, alongside evidence that the simulated market replicates real bidding patterns well enough to compare against human auction behavior. This is the gate most technically strong teams will underweight. If the market does not clear rationally, every behavioral inference drawn from it is noise. Efficiency is not a nice-to-have you tune at the end; it is the validity condition for the entire measurement. Design for it in Month 2 or spend Month 11 discovering you cannot reach it.

The earlier milestones build to that. Month 2 defines market mechanics — structures, payout schemes, efficiency measures, utility calculations, the news feed. Month 6 formalizes the agent decision space and scaling strategy. Month 9 requires stock agents from at least ten different LLMs operating in the sandbox.

The requirement that explains why this is a biology topic

Read the Phase II schedule and the office assignment stops being strange.

Month 6 of Phase II deploys "human" stock agents that replicate human market preferences, plus classifiers that distinguish human from AI behavior. Month 15 models social interaction — collaboration and deception — with quantitative impact metrics on market efficiency. Month 18 quantifies how latent AI preferences affect social dynamics and market outcomes. Month 21 delivers object and source code to the Government with documentation. Month 24 requires a final report that includes analysis of how LLM advances during the period of performance affected the evaluation's validity.

But the single most revealing line specifies that at least one stock agent should represent the human user of the AI under test, in order to explore how that AI "may drive human behaviors of concern such as encouraging delusions or eroding military discipline."

That is not an AI capability benchmark. It is a behavioral-effects measurement on a human population, run through a simulated proxy. It sits in the Biological Technologies Office for the same reason DARPA's Unbiased Behavioral Discovery Platforms work does: the subject of study is behavior, and the AI is the stimulus. Proposals written as pure ML evaluation infrastructure — throughput, coverage, model support matrices — will miss what the program manager is buying.

The team composition most bidders will get wrong

The topic's own reference list is the tell: two of eight references address auction theory and human bidding behavior. This topic needs an economist grounded in experimental and behavioral economics on the technical team, not consulting from the sidelines and not a citation in the related-work section.

An LLM engineering shop with no auction-theory depth can build something that runs. It will struggle to defend the market design, the utility functions, the error function comparing AI decisions to human norms — which Phase II requires be rigorously defined by Month 2 and applied consistently thereafter — and the claim that its efficiency number means anything. A team with an experimental economist can defend all of it and will read as the lower-risk award.

Three other things separate a credible cost volume from an amateur one.

Budget inference honestly. Ten or more LLMs, many agents, thousands of market rounds, 24 months. API spend at that scale is a real line item, and underestimating it reads as inexperience to anyone who has run large evaluation harnesses. So does pretending the cost is flat when frontier pricing moves every quarter.

Plan for model drift explicitly. The Month 24 report must address how model advances during performance affect evaluation validity. That means versioning, re-baselining, and frozen reference snapshots designed into the architecture from the start — not a retrospective caveat.

Resolve IP before you draft. Source code delivery to the Government at Month 21 interacts directly with what you can license in and what you can commercialize out. Your contract-type election compounds it: firm-fixed-price, cost-plus reimbursement (which requires a DCMA Final Determination Letter you will not obtain in six weeks from a standing start), or Other Transaction for Prototype (Model OT template plus certifications). Decide first; the choice reshapes the cost volume.

The commercial thesis DARPA hands you

Most SBIR topics leave commercialization as the proposer's problem. This one names the buyer.

The solicitation's argument is liability-driven: AI developers facing litigation over user harm need defensible, quantitative evidence bounding their systems' influence, and establishing those limits can move a company's valuation more than the cost of any individual case. Insurers underwriting AI deployment, regulators building evidentiary standards, and enterprise procurement teams running vendor evaluations follow as secondary markets.

Use it. A commercialization section that restates DARPA's own framing back to it, with named prospective customers and a plausible pricing model, is worth more than a generic TAM slide. And engage the a-priori-harm-taxonomy alternative directly — explain why revealed preference under incentives outperforms a checklist. That argument is what DARPA is actually purchasing, and demonstrating you understand it is the cheapest differentiation available to you.

Technical questions close October 14, 2026 at SBIR_BAA@darpa.mil with the topic number in the subject line; the FAQ updates until one week before the deadline. Phase II Technical and Business Assistance adds $25,000 above the cost ceiling. Proposals are due October 21, 2026, and DARPA accepts nothing late.

Six weeks is enough time to write this proposal well, and not nearly enough to assemble the team it requires from scratch — which is the honest bid/no-bid question to settle this week. When you are triaging which solicitations genuinely match your bench before committing proposal budget, Granted is built to make that call faster.

Get AI Grants Delivered Weekly

New funding opportunities, deadline alerts, and grant writing tips every Tuesday.

Browse all DARPA grants

More DARPA Articles

226 Proposals, $664M Requested, $105M Available — and Then a Defense Agency Showed Up. How Screwworm Became a Cross-Agency Funding Vertical.

USDA has obligated roughly $1 billion in emergency funds against New World screwworm and awarded $105 million to 40 Grand Challenge projects from 226 proposals. On September 2, 2026, DARPA pre-released a $2.1 million SBIR topic for networked screwworm detection. Four funding streams, four eligibility regimes, four clockspeeds — and most applicants only know about one.

Read article

The Air Force Closed Its Open BAA and Told Applicants to Resubmit. The Standing Defense Solicitation Is Being Rebuilt Under EO 14332.

AFOSR's open Broad Agency Announcement — the single front door to Air Force basic research funding — closed on March 23, 2026 for mandatory review under Executive Order 14332, and proposals submitted in its final week must be resubmitted to a replacement that has no posted date. Meanwhile NRL's Long Range BAA takes white papers only through September 30, and the Air Force's new $99 million command-and-control BAA runs to 2028 with a March 15, 2027 white paper date for FY28 money. A field guide to the rolling-solicitation channel while it is in flux.

Read article

DARPA Wants Drone Video Cut by 99 Percent — Running on Two Watts. DV019's $2M Semantic ISR Topic Closes September 23.

DPA26BZ05-DV019 asks small businesses to reduce ISR video transmission from small drones by 90 to 99 percent while preserving what an operator actually needs, on a Jetson-class board at a 2-watt final power budget. It is Direct-to-Phase-II, up to $1.5 million base plus a $500,000 option, and its exclusion list — no weapon release, no lethal engagement, no named-person identification — is the most strategically revealing paragraph in DARPA's FY26 Release 5.

Read article

Not sure which grants to apply for?

Use our free grant finder to search active federal funding opportunities by agency, eligibility, and deadline.

Find Grants

Ready to write your next grant?

Draft your proposal with Granted AI. Professional members win a grant in 12 months or get a full refund.

Backed by the Granted Guarantee