←
AI for Government
Visionary · M7 · lesson 7 of 46 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI in Defense and National Security
📖
now learning

AI in Defense and National Security

15 min

An assistant secretary at a defense agency, Beatriz Okonjo-Lindqvist, sat through a vendor briefing for an AI system that promised to fuse satellite imagery, signals data, and open-source reporting into automated threat assessments, billed as "decision-ready intelligence in seconds." The demo was dazzling. Then a colonel in the back asked one question that emptied the room of its enthusiasm: "When this thing tells a commander a building is a legitimate target and it is wrong, who pulled the trigger, the algorithm or the human who trusted it?" Nobody had a clean answer. That question sits at the center of AI in defense and national security, where the technology's promise is enormous and the cost of error is final. This lesson is an unclassified framework for leaders who must govern AI in the highest-stakes domain there is.

The scope of this unclassified treatment

Everything below stays at the unclassified, conceptual level. It teaches judgment about governance, and it deliberately states no handling rule, no access condition, no marking, and no control identifier of any kind. Your organization's security authorities own those, and they own them for reasons that do not survive being summarized in a training lesson. If a question here touches how a specific dataset, system, model, or output must actually be handled, that question belongs to your security officer and your counsel before it reaches your program plan.

Read that as a hard constraint on how you use this material, not a disclaimer. Leaders get into trouble in this domain by carrying a plausible recollection of a rule into a decision meeting instead of the rule itself. The governance questions in this lesson are designed to be asked out loud in an unclassified room. The answers, in many cases, will not be. That asymmetry is normal, and it is one more reason the internal governance described here has to be unusually disciplined.

Why national security AI is its own category

Every lesson about government AI talks about stakes. Here the stakes include lethal force, the security of the nation, and decisions made under adversarial pressure where an opponent is actively working to deceive your systems. That combination changes what good governance looks like, and it changes it before you write a single requirement. Three features set the domain apart, and each one breaks an assumption that civilian AI governance quietly relies on.

First, adversaries adapt. A fraud model faces fraudsters who probe for weakness opportunistically. A defense model may face a nation-state intelligence service that studies how the model behaves and then deliberately feeds it false data to manipulate its conclusions. Second, errors can be irreversible and lethal. There is no version of "we corrected it in the next release" that returns a destroyed building or an unrecoverable escalation. Third, secrecy constrains oversight. The normal checks of public transparency, journalistic scrutiny, and open audit are limited here by necessity, so internal governance has to be far stronger to compensate for what the outside world cannot see.

Hold that third point carefully, because it is the one most often inverted in practice. Restricted visibility is treated as a reason that review is impractical. It is the opposite. When the public cannot check your work, the burden shifts entirely onto the people who can, and the standards those people apply have to be higher than anything a civilian agency would need. Compensating governance is the price of operating with reduced external scrutiny, not an optional extra you fund when the schedule allows.

Three application areas, plainly

Strip away the terminology and the current wave of defense AI concentrates in three places. Each has a real capability case, and each has a characteristic way of going wrong that a leader can recognize from a briefing slide without any technical depth at all. Learn the three failure signatures rather than the three technologies, because the technologies will be replaced within a few procurement cycles and the failure signatures will not. What follows is deliberately brief on how each capability works and deliberately specific about where each one has historically disappointed the people who bought it.

Autonomous systems

These are systems that sense, decide, and act with reduced human input, ranging from logistics drones to defensive systems that must react faster than a person can. The promise is speed, endurance, and keeping people out of physical danger. The danger is a machine making a consequential decision that a human should own. The organizing principle in defense policy is that a person exercises appropriate judgment over the use of force, especially over lethal decisions. The hard engineering and policy question is what that judgment can actually mean when the system operates faster than human reaction time.

That question does not have a clean general answer, and you should be suspicious of anyone who gives you one. What you can insist on is that the answer be specific to the system in front of you: at what point in this system's operation does a named human decide, how much time does that person have, what information do they see, and what happens if they say no. A vendor who cannot answer those four questions concretely has not designed for human judgment. They have designed a fast machine and reserved a seat for a person who will be too late to use it.

Intelligence analysis

AI can triage oceans of data, including imagery, intercepts, and documents, and surface what merits a human analyst's attention. This is genuinely transformative, because no human team can read everything, and the material that goes unread is where the missed indicator usually sits. The risk is the same hallucination and bias problem as everywhere else, amplified by consequence: a model that confidently misidentifies a pattern, or one that an adversary has deliberately fed misleading inputs to. The correct framing is that AI narrows the haystack for human judgment. It does not deliver verdicts.

The practical governance move is to make that framing visible in the product itself. If the tool's output looks like a conclusion, analysts will treat it as a conclusion, whatever the training slide said. If it looks like a ranked list of things worth examining, with the underlying material one click away and the model's uncertainty on the face of it, the analytic habit survives contact with the tool. Interface design is governance in this domain, and it is usually decided by people who were never told that.

Cybersecurity

This is the most symmetric battlefield of the three. AI defends networks at machine speed by detecting intrusions, prioritizing patching, and responding faster than human defenders could. But adversaries use the same class of capability to find vulnerabilities and craft attacks, which makes this an arms race where standing still means falling behind. It has one feature the other two do not: your own AI tools are themselves targets, valuable precisely because compromising the thing that watches the network is better than evading it. Defensive AI expands the attack surface it is deployed to protect.

The governance implication is that you cannot evaluate a defensive AI tool purely on how well it defends. You also have to ask what happens when it is the thing being attacked: what an adversary gains by degrading it, by feeding it inputs that train it toward blindness, or by simply learning what it does and does not alert on. Tools that watch the network are high-value targets for exactly the reason they are useful, and a defensive capability procured without that assumption has been evaluated against only part of its risk.

What connects them

In all three, the failure mode is the same: a human treating an AI output as ground truth when it is a probabilistic estimate that an adversary may have shaped. The applications differ. The governance question does not. When the system says "legitimate target" and it is wrong, who pulled the trigger, the algorithm or the human who trusted it? In this domain that question never has a comfortable answer, so the governance has to.

The ethics of military AI as design specification

Ethics here is not a soft addition to a technical program. It is operational risk management carrying moral weight, and it is most useful when translated out of principle language and into requirements a program manager can be held to. The Department of Defense has adopted ethical principles for AI, and they are worth knowing precisely because they can be read as design obligations rather than as slogans on a poster.

  • Responsible. Humans exercise appropriate judgment and remain accountable. The algorithm is never the one who pulled the trigger.
  • Equitable. Deliberate steps to minimize unintended bias, because biased targeting or screening is simultaneously a moral failure and an operational one.
  • Traceable. Personnel understand how the system works, what data it uses, and where its limits sit, well enough to use it responsibly and to investigate it when it fails.
  • Reliable. The system has an explicit, well defined use, and its safety and effectiveness are tested across that use throughout its life.
  • Governable. The system is built to detect and avoid unintended consequences, with the ability to disengage or deactivate it when it behaves wrongly.

Governability deserves particular emphasis, because it is the principle most often satisfied on paper and least often satisfied in fact. The capacity to turn a system off when it misbehaves is a first-class design requirement, not a fallback. A system you cannot stop is one you should not field. That standard sounds obvious until you apply it honestly: an off-switch that exists in the architecture diagram but cannot be reached inside the timeline the system actually operates on is not an off-switch. It is a comfort object.

Traceability has a similar trap. It is usually procured as documentation, and documentation is necessary but not sufficient. The test is not whether a manual exists. It is whether the specific people who will use this system under pressure can say, in their own words, what it is good at, where it degrades, and what a wrong answer from it would look like. If they cannot, the system is not traceable to them, whatever the program office holds on file.

Governing for an adversary who is trying to fool you

The feature that most distinguishes this domain from civilian AI is the active adversary. In tax fraud detection, the data is messy but it is not maliciously engineered against your model. In national security, an opponent may conduct data poisoning by corrupting training data, or craft inputs designed to make your model reach a wrong and dangerous conclusion, or simply observe which behaviors your system rewards and then perform those behaviors. National security AI governance must assume the system will be attacked and design for that from the beginning.

The practical implication is that a model evaluated only against benign data is untested for the conditions that actually matter. Adversarial testing, input validation, and human cross-checks aimed specifically at manipulation belong in the acceptance criteria, not in a research backlog. But be precise about what that testing buys you. Adversarial testing demonstrates resistance to the attacks you thought to try. It is evidence, not a guarantee, and an opponent who is well resourced and patient is under no obligation to attack in a way your red team anticipated. Treat a clean adversarial evaluation as a floor you have cleared, never as a ceiling an adversary cannot climb.

The same discipline applies to data provenance. A model is only as trustworthy as the pipeline that feeds it, and the pipeline is usually longer, more automated, and less examined than the model. Ask where each training and inference input comes from, who could alter it, and what would happen if someone did. That question is uncomfortable because the honest answer is often that nobody has traced it end to end. Discovering that in a governance review is inconvenient. Discovering it after an operation is not.

A usable artifact: the national security AI governance screen

Beatriz adopted the following screen for any AI capability proposed in her portfolio. It forces the domain's hard questions before money or trust is committed. Five of its rows are the ethical principles above, applied as gates rather than aspirations. The remaining two rows exist because this domain adds two demands that a general ethics framework does not make explicit: the assumption of an active adversary, and the need to compensate for reduced external scrutiny.

GateThe leadership questionRed flag
ResponsibleFor any consequential or lethal decision, does a named human exercise judgment and bear accountability?"The system decides; the operator confirms."
EquitableHave we tested for bias in targeting, screening, or assessment outputs?No disparate-impact testing at all.
TraceableDo the people using it understand its data, logic, and limits well enough to know when not to trust it?Operators treat outputs as ground truth.
ReliableIs the use narrowly defined and tested across it, including degraded and contested conditions?Tested only on clean, benign data.
GovernableCan we detect misbehavior and disengage or deactivate the system inside the time available?No reachable off-switch at operational speed.
AdversarialHave we assumed an opponent will try to poison or deceive it, and tested against that?No adversarial or data-poisoning testing.
OversightWith external scrutiny necessarily limited, is internal governance correspondingly stronger?Restricted visibility used as a reason to skip review.

The screen is deliberately answerable by a non-technical executive. Every question can be put to a program manager or a vendor in a meeting, and every red flag is something you can hear in a sentence rather than something you have to find in a test report. That is the point. A governance instrument that only a data scientist can operate will be operated by data scientists, which means it will never be operated by the person who signs.

Reading a vendor briefing in this domain

The briefing Beatriz sat through is the standard shape: an impressive fusion of sources, a compressed timeline, and a phrase like decision-ready intelligence in seconds. Nothing about that pitch is dishonest, and that is what makes it hard. The claims are usually true about what the system produces and silent about what the system is right about. Your job in the room is to convert every temporal claim into an accuracy claim and every capability claim into an accountability claim, out loud, while the people who built it are still present to answer.

Three questions do most of the work. What exactly does this system output, stated as a sentence a person would act on rather than a category label. How often is that output wrong, measured how, on data resembling what we will actually feed it. And when it is wrong, what does the wrongness look like, because a system that fails obviously is a different risk from one that fails plausibly. The second question is the one vendors are least prepared for, and the answer you should refuse to accept is a benchmark figure with no stated evaluation conditions.

Then ask the colonel's question directly, in the vendor's own terms. If this output is wrong and a commander acts on it, who decided. Watch whether the answer describes a person or a process. A good answer names a role, a moment, and a piece of information that role holds which the system does not. A bad answer describes a workflow in which the human step is a confirmation screen. That distinction, audible in a single sentence, separates a capability you can govern from one you will be explaining to an investigator.

The record you will need afterward

The requirement that a human must always be able to say they decided has a practical consequence that programs routinely discover too late: you have to be able to reconstruct the decision after the fact, and reconstruction is only possible if somebody designed for it in advance. Traceability and responsibility are the two ethical principles that meet here. Traceability says the people using the system understand it well enough to investigate a failure. Responsibility says a named human bears accountability. Neither survives an inquiry if the system kept no usable record of what it presented and what the person did with it.

What that record has to contain is a design question for your organization and your authorities, not something a lesson can specify. What a lesson can tell you is the test to apply: imagine the investigation. Somebody sits down a year from now, with no memory of the day and no access to the people involved, and asks what the system said, what the human saw, what the human chose, and why. If your architecture cannot answer those four questions from durable evidence, you do not have accountability. You have an org chart, and an org chart has never once resolved an inquiry.

The leader's burden

Beatriz killed the dazzling vendor system. Not because the technology was bad, but because it could not answer the colonel's question. It positioned the AI as the decider and the human as a confirmer, and it had no story at all for an adversary trying to fool it. She approved a redesigned version a year later with the same underlying capability, but with human judgment genuinely in the loop, adversarial testing built into acceptance, and a governable stop that a person could actually reach. The principle she leads by is unglamorous and absolute: in this domain, a human must always be the one who decided, and must always be able to say so afterward.

Notice what the year cost and what it bought. It cost a delay that her program office resented and that a competitor agency did not incur. It bought a system whose failure, when it eventually comes, will be investigable, attributable, and survivable as an institution. That trade is the whole job. Nobody will thank you for the incident that did not happen, and the discipline of buying insurance against unattributable failure has to come from your own conviction, because the incentives around you will not supply it.

Anti-patterns to watch for

Treating the human in the loop as a control rather than a role. Putting a person between the model and the action does nothing on its own. If that person sees only the model's conclusion, has seconds to respond, faces no consequence for approving and real friction for refusing, and has never been shown what a wrong output looks like, then the loop is decorative. The control is not the human's presence. It is the human's realistic capacity to decide otherwise, and that capacity has to be designed, measured, and defended against schedule pressure.

Treating a boundary as a guarantee. Separation of systems and restriction of access are load-bearing protections and you should insist on them. They are not proof that misuse cannot occur. Boundaries are crossed by people, by process exceptions, and by data aggregated somewhere upstream. Describe your protections as what they are, barriers that raise cost and create detectable events, then instrument for the case where one is crossed anyway.

Letting adversarial testing become a certificate. A red team result is a snapshot of resistance to attacks somebody imagined, on a date, against a model version. Programs quote it forever, across retrains and integrations and shifts in the threat picture. Ask when the test ran, against which version, by whom, and what they were not permitted to try.

Buying the demo instead of the system. A demonstration runs on curated data, in benign conditions, with the vendor's engineer in the room. None of those three conditions will hold in operation, and all three are exactly what determines whether the capability survives contact. Insist on evaluation under degraded inputs and contested conditions, with your own people at the controls, before commitment rather than after.

Using necessary secrecy to shrink the review. The reasoning starts reasonably, that fewer people can see the work, and ends badly, that therefore fewer people will review it. The correct response to a smaller reviewing population is a more rigorous review by that population, more documented rationale, and more deliberate rotation of who examines what. Reduced external visibility increases the internal obligation rather than reducing it.

Confusing speed of output with quality of judgment. The pitch in this domain is almost always temporal: decisions in seconds, warning in minutes, response at machine speed. Speed is genuinely valuable, and it is also the easiest property to demonstrate and the least correlated with correctness. Ask what the system does faster, then ask separately how often it is right and what the error costs. Programs that only answer the first question have not been evaluated.

Practice prompts

  1. Run one AI capability proposed or fielded in your portfolio through the seven gates in the governance screen, one paragraph per gate. Where you cannot answer, write "unknown" rather than a guess, and treat the unknowns as your governance backlog.
  2. For a system that involves any consequential decision, write down the name of the human who decides, the number of seconds or minutes they have, the information in front of them at that moment, and what mechanically happens if they decline. If you cannot fill in all four, draft the requirement that would let you.
  3. Draft the five questions you would ask a vendor about how their system behaves when its inputs have been deliberately manipulated. Then draft the answers you would accept and the answers that would end the conversation.
  4. Write a one-page memo specifying which internal mechanisms compensate for your program's necessarily limited external oversight. Be concrete about who reviews, how often, and with what authority to stop work.
  5. Trace the data provenance of one model your organization uses, from original collection to inference input. Identify every point at which the data could be altered and who would have to be compromised for that to happen. Note where the trace goes cold.

Reflection

These questions are worth sitting with rather than answering quickly. They determine whether the frameworks above become practice or paperwork, and most have an honest answer less comfortable than the one you would give in a meeting. Write your answers privately first, so what you believe is on record before the room reshapes it.

  • Where in your organization would a decision made by a model be indistinguishable, after the fact, from a decision made by a person?
  • If a system you approved produced a wrong and consequential output tomorrow, could you reconstruct why, and could you name who decided?
  • What would it cost you personally, professionally and politically, to delay a capability by a year for governance reasons the way Beatriz did?
  • Which of your controls have you been describing as things that prevent harm, when they are actually things that raise the cost of harm and make it detectable?
  • Who in your organization is empowered to say that a system cannot be fielded, and when did they last actually say it?

Glossary

  • Autonomous system. A system that senses, decides, and acts with reduced human input. Autonomy is a spectrum, and the governance question is which decisions the human retains.
  • Meaningful human judgment. The requirement that a person exercise real judgment over consequential decisions, particularly the use of force. Its content is set by whether the person has the time, information, and authority to decide otherwise.
  • Data poisoning. Deliberate corruption of the data a model learns from, so that the model behaves as the attacker intends while appearing to function normally.
  • Adversarial input. An input crafted to make a model reach a wrong conclusion while looking unremarkable to a reviewer.
  • Governability. The designed ability to detect that a system is misbehaving and to disengage or deactivate it within the time available in operation.
  • Traceability. The condition in which people using a system understand its data, logic, and limits well enough to use it responsibly and to investigate a failure.
  • Provenance. The documented history of where data came from and who could have changed it along the way.
  • Compensating governance. Internal review strengthened deliberately to offset external scrutiny that is limited for legitimate reasons.

Closing

The uncomfortable truth about this domain is that its governance cannot be delegated downward. In civilian programs, a technically strong team can often compensate for a disengaged executive. Here the questions that matter are questions about accountability, authority, and acceptable risk, and those are precisely the questions a program office is not positioned to answer. If you lead in this space, the screen above is not a document you commission. It is a conversation you personally hold, repeatedly, with people who would prefer you did not.

Nothing in this lesson argues against using AI in defense and national security. The capability case is real, the analytic volumes are genuinely beyond human reach, and an adversary that adopts these tools faster will have an advantage. The argument is narrower and harder: that the value of these systems is realized only when a human being remains the one who decided, and can still say so a year later in front of people who are not inclined to be generous. Build for that day, and the rest of the governance follows.

Key takeaways

  • National security AI is its own category. Adversaries actively try to deceive your systems, errors can be lethal and final, and necessary secrecy limits the usual external oversight.
  • Keep humans meaningfully in control of force. A named person must exercise judgment and bear accountability for consequential and lethal decisions, never the algorithm.
  • AI should narrow the haystack, not deliver verdicts. In intelligence work, treat outputs as probabilistic leads for human judgment, and design the interface so that framing survives contact with the tool.
  • Governability is a design requirement. If you cannot reliably detect misbehavior and stop the system inside the time available, you should not field it.
  • Assume an active adversary. Build adversarial and data-poisoning testing into acceptance, and read a clean result as evidence of resistance to attacks you imagined, never as a guarantee.
  • The ethical principles are design specifications. Responsible, equitable, traceable, reliable, and governable each translate into a gate with a red flag you can hear in a meeting.
  • Stronger internal governance offsets limited external scrutiny. Restricted visibility is never a reason to skip review; it is a reason to review harder and document more.
  • Handling rules belong to your security authorities. This lesson gives you governance judgment, not handling detail, and no training material should ever substitute for the people who own those rules.

Frequently Asked Questions

Does this lesson tell me how to handle sensitive data or systems in a defense context?

No, and that is deliberate. It states no handling rule, marking, access condition, or control identifier, because those are set by your organization's security authorities and are not safely summarized in training material. Use this lesson for governance judgment, which questions to ask and which answers should worry you, and take every handling question to your security officer and counsel. Where anything here appears to conflict with their direction, their direction governs.

Is a human in the loop enough to make an autonomous system acceptable?

Not on its own. A human in the loop is a role, and it becomes a control only when that person has enough time to think, enough information to judge, real authority to refuse, and no professional penalty for using it. Many systems satisfy the letter of human involvement while making refusal practically impossible. Test the loop by asking what happens mechanically when the person says no, and how often that has actually occurred.

How do I evaluate a vendor claim that a system has been adversarially tested?

Ask four questions: when the testing was performed, against which model version, by whom, and what the testers were not permitted to attempt. The last is the most revealing and is almost never in the summary. Then ask whether the result was revalidated after the model was retrained or integrated. A stale certificate against a changed system produces confidence without producing protection.

Why does the lesson insist that AI should not deliver conclusions in intelligence work?

Because a conclusion invites acceptance and a ranked list invites examination, and the difference shows up in what analysts do under time pressure. AI is genuinely valuable at reducing unreadable volumes of material to a prioritized set, and unreliable as a source of verdicts, particularly when an adversary may have shaped its inputs. Designing the output as leads rather than answers preserves the analytic habit that catches the model's mistakes.

What does it actually mean for a system to be governable?

It means you can detect that the system is behaving wrongly and stop it within the time the operation allows. Both halves matter. Detection without a reachable stop is an alarm on a door you cannot close, and a stop that exists in the architecture but not inside the engagement timeline is not a stop. Test governability against the clock the system actually runs on, not against the clock in the requirements document.

How should limited external oversight change my internal governance?

It should make it heavier in every dimension: more documented rationale, more frequent review, clearer authority to halt work, and deliberate rotation of who examines what. The instinct to review less because fewer people can see the work is exactly backwards. The reviewing population is smaller, so the burden on each reviewer is larger, and the standard that population applies has to rise to cover what no outside party is in a position to check.