←
AI for Government
Visionary · M12 · lesson 12 of 46 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI in Infrastructure
📖
now learning

AI in Infrastructure

15 min

Linda Castellano is the chief engineer for a state department of transportation responsible for 28,000 lane-miles of road and 4,200 bridges. Her inspection budget lets her physically check each bridge once every two years, which means a bridge can develop a problem the day after an inspection and stay unflagged for nearly 24 months. In 2023 a culvert washed out on a rural highway eleven months after a clean inspection, and the after-action review found that traffic-load and rainfall data, both already collected by her agency, would have ranked that culvert as high-risk months earlier. The data existed. The capacity to act on it did not. That gap is exactly where AI belongs in public infrastructure.

Infrastructure is where AI's value to government is most tangible, because the assets are physical, the failures are visible, and the data is already pouring in from sensors, cameras and meters. The controlling idea for a leader is this: AI does not build the bridge or fix the pipe. It tells you, out of thousands of assets, which one to send your limited crews to first, before it fails on the evening news. Everything else in this lesson is a consequence of taking that sentence literally.

An attention-allocation engine, not a labor-saving device

Most infrastructure agencies are not short on data. They are short on the attention required to act on it. Linda's after-action review is the whole argument in miniature: the traffic-load and rainfall records that would have flagged the culvert were already sitting inside her agency, collected for other reasons, read by nobody in a way that connected them to that structure. Nothing was hidden. The information simply never arrived in front of a person who could do something about it, in time, in a form that ranked it against everything else competing for the same crew.

That reframing matters for how you scope a program and how you defend it. If you sell infrastructure AI as automation, you invite the question of which jobs it replaces, and the honest answer is none of the ones that matter, because a crew's hands on the asset are still what make the repair real. If you sell it as attention allocation, the promise is testable: does the ranked list surface the failing asset earlier than the calendar did? That is a question your own records can answer within a year, and it is the question your legislature will eventually ask you anyway.

The framing also settles an argument that otherwise consumes an entire program's first year, which is how accurate the model has to be before it is useful. Against a perfect standard, no ranking of thousands of heterogeneous assets will look good. Against the standard it actually replaces, a fixed calendar that pays no attention to traffic load, materials, age or weather, the bar is considerably lower and considerably clearer. Linda's culvert was not missed because her inspection regime was careless. It was missed because a calendar cannot notice that one structure is taking more water and more load than its neighbours. Any ranking that notices that is an improvement on the incumbent, and the incumbent, not perfection, is the comparison you are entitled to make.

Predictive maintenance: the central use

This is Linda's direct fix and the highest-value infrastructure use across government. The old model is calendar-based: inspect everything on a fixed cycle whether it needs it or not. That approach wastes effort on healthy assets and misses the one that degrades fast, and it does both at the same time, which is why agencies feel simultaneously overworked and unprotected. Predictive maintenance flips the question from "when is this scheduled?" to "which asset is most likely to fail soonest, given everything we know about it?"

A model trained on an asset's age, materials, traffic load, weather exposure and sensor readings can rank all 4,200 bridges every week by failure risk. Linda then sends her crews to the top of that ranked list instead of working a fixed calendar. The payoff is documented across utilities and transit: catching a failure before it happens is far cheaper than the emergency repair, and a water utility that predicts which of its aging mains will burst can cut both emergency excavations and service disruptions sharply. The saving is not only money. It is the disruption that never happens to the people who depend on the asset.

Notice what has and has not changed in that arrangement. The inspection capacity is the same. The crews are the same. The budget is the same. What changed is the order of the queue, which is the one variable an agency can move without an appropriation. That is why predictive maintenance is usually the first infrastructure AI investment that pays for itself, and why it is a poor argument for a budget increase and an excellent argument for keeping the budget you have.

The honest limit of a risk score

A risk score is a prioritization, not a diagnosis. The model says "inspect this one first." A human inspector confirms the actual condition, and that division of labor is not a formality you can streamline away later. Acting on a score without verification means you will occasionally tear up a healthy road, and the cost of that is not just the wasted crew day. It is crew trust, which you spend once and rebuild slowly. An inspector sent to a run of sound structures stops reading the list.

The failure runs in the other direction too, and it is the one that should worry a chief engineer most. A model that rates a failing asset as safe produces no alarm, no crew visit and no record that anything went wrong until the asset fails. There is no natural feedback loop for that error, because nobody goes to look. You have to build the loop deliberately: whenever an asset fails, go back and ask where the model had ranked it the week before, and keep that tally as a standing number rather than an incident report.

Both errors point to the same organisational requirement, which is that somebody owns the model after launch. Prioritization systems are unusually prone to being adopted informally: a list appears, schedulers start using it, and within a year the ranking is effectively setting the work plan without anyone having decided that it should. Name the owner, give that person the false-negative tally and the override records, and make the decision to rely on the list an explicit one with a date attached. An informal dependency is the hardest kind to unwind, because there is no document to revise and no decision to revisit.

Where it works across infrastructure sectors

Transportation

Beyond bridge maintenance, AI optimizes traffic-signal timing to cut congestion, analyzes crash patterns to target the intersections that need redesign, and processes drone imagery of pavement to grade road condition faster than a survey crew can. The safety-critical line is easy to state and worth writing into the operating procedure: AI can recommend a signal-timing change, and a traffic engineer signs off before it goes live. The pavement-grading case is the most comfortable of the three, because it substitutes for a measurement rather than for a judgment, and a wrong grade is corrected by the next survey rather than by a collision.

Energy grid

AI forecasts electricity demand, balances renewable supply that swings with the weather, and predicts which transformers or lines are most likely to fail. For a public power authority that means fewer outages and better integration of solar and wind, which are the two things a utility board asks about. Grid operations are safety-critical and are a target for attack, so human operators retain control and the system has to be hardened against tampering. A forecasting model that quietly becomes a control system, because operators started following it without a documented decision to do so, is a governance failure that will not appear in any procurement record.

Water systems

AI detects leaks from pressure and flow data, predicts main breaks, and helps monitor water quality. For an agency managing aging pipes, the leak-detection win alone can save millions of gallons and dollars a year, and it has the useful property of being measurable in the same units the public already understands. Water quality is the more sensitive application, because a monitoring model that misses a problem has a direct public health consequence, and the appropriate posture there is that the model raises a flag and the established testing regime confirms it.

Broadband and smart cities

AI helps map where broadband gaps actually are, as opposed to where providers claim coverage, which directs public investment to the genuinely unserved. In smart-city programs it coordinates parking, lighting and waste collection, all of which are low-stakes uses with visible benefits. The caution here is surveillance creep. A smart-city sensor network can quietly become a mass-monitoring system, one camera and one retention-policy extension at a time, so privacy limits must be designed in from the start rather than bolted on after a public outcry. Decide what is collected, how long it is kept and who may query it before the first sensor goes up, because those questions are far harder to answer once the data exists.

Sorting these uses by what happens if they are wrong

Reading those four sectors together, a useful ordering falls out, and it is not by sector but by consequence. At one end sit the uses where a wrong output costs a wasted trip or a stale map: pavement grading from drone imagery, broadband coverage mapping, parking and waste-collection coordination. A wrong answer there is corrected by the next survey or the next week's data, nobody is hurt in the interval, and the appropriate governance is ordinary project discipline. These are the uses to ship first, precisely because they build the institutional habit of running a model in production before anything important depends on one.

In the middle sit the prioritization uses, including Linda's bridge ranking, leak detection and main-break prediction. A wrong answer here misdirects a crew, and the protection is the human inspection that follows the score rather than any property of the model. At the far end sit the uses whose output moves something physical or reaches the public directly: signal timing, grid balancing and water-quality monitoring. There the requirement is not just that a human confirms, but that a named human has authority to refuse, that the refusal is workable in practice, and that the system keeps running when the model does not. Write that ordering down before you procure anything, because a vendor's product roadmap will not respect it and your own program will drift toward whatever demonstrates best.

Bad data poisons the model

Infrastructure data is notoriously messy. Inspection records are inconsistent between districts and between decades, sensors drift out of calibration and nobody notices, and gaps span years where a system was replaced or a program lapsed. A predictive model is only as good as its inputs, and the specific danger is not that a model trained on dirty data will refuse to produce an answer. It is that it will produce a confident ranking anyway, and put the wrong asset first.

The practical rule is to budget for data cleanup before you budget for the model, and to treat the cleanup as the deliverable rather than as a prerequisite to be compressed when the schedule slips. Agencies consistently invert this, because a model is procurable and a data inventory is not, and because the vendor conversation is more exciting than the records conversation. If the only thing your program produces in its first year is a trustworthy asset register with known provenance and known gaps, you have built the asset that every subsequent model depends on, and you have made the next procurement cheaper rather than more expensive.

Equity in where you invest

A model trained on historical maintenance data inherits historical neglect. If certain neighborhoods' infrastructure was never well maintained, the agency has sparser records there, fewer sensors, fewer inspection histories and fewer logged repairs. A model reading that thinner record can under-rank the real risk, and it will do so with the same confident interface it uses everywhere else. The pattern perpetuates the gap while presenting itself as neutral prioritization, which is precisely what makes it hard to argue with in a budget meeting.

So test whether your prioritization tool directs investment fairly across communities, and not just toward the assets that were always watched. The test is concrete: take the ranked list, break it down by the geographies your agency already reports on, and compare the distribution to what your inspection history would have produced. If the model's attention lands where the sensors are, you have measured your sensor deployment rather than your risk. That finding is useful. It just is not the finding the dashboard claims to be showing you.

Security when the output moves something physical

AI connected to grids, water systems and traffic controls is a cyberattack surface where a breach has physical consequences. These are critical-infrastructure systems and must be protected accordingly, with human operators able to override and with a hardened operating path that does not depend on the AI staying online. The design question to ask early is what the agency does on the day the model is unavailable, wrong or compromised, and whether the staff who would have to run the manual process have done it recently enough to still be able to.

There is a quieter dependency risk alongside the security one. As a ranking system becomes routine, the institutional knowledge that used to sit with senior inspectors migrates into the model, and the manual fallback degrades from a procedure into a memory. Keeping the fallback exercised is not nostalgia. It is the thing that makes the override authority real rather than nominal, because an operator who cannot run the system without the model will not, in practice, overrule it.

The frameworks that apply

The NIST AI Risk Management Framework, structured around Govern, Map, Measure and Manage, gives the program its spine. It is voluntary guidance rather than binding regulation, which means an agency adopts it by policy and then has to enforce it on itself. Its Measure function is where you document the model's accuracy and, specifically, its false-negative rate: the assets it ranks safe that actually fail. That is the number that should keep a chief engineer awake, and it is the number a model vendor is least likely to volunteer, because it is the hardest one to look good on.

The Office of Management and Budget memo M-24-10, issued in 2024, treats AI affecting public safety as safety-impacting, which triggers requirements for testing, monitoring and human oversight. A bridge prioritization model sits inside that category whether or not anyone has classified it, so classify it deliberately and write the reasoning down. And for any sensor network touching the public, privacy and surveillance limits belong in the design documents, set before deployment rather than negotiated after the first records request. None of these frameworks tells you which bridge to inspect. They tell you what you owe the public about how you decided.

An infrastructure AI program checklist

  1. Data first. Audit and clean the asset and sensor data before training, and budget for that work explicitly as a deliverable rather than as a prerequisite.
  2. Score directs, human confirms. A risk ranking sends a crew to inspect. The inspector verifies condition before any repair, closure or public statement.
  3. Know your false-negative rate. Track how often the model rates a failing asset as safe, build the feedback loop that lets you see it, and set an acceptable threshold for a safety-critical system.
  4. Test for investment equity. Check whether prioritization fairly covers historically neglected communities, and compare the model's attention to your sensor and inspection coverage.
  5. Harden physical-control systems. Treat grid, water and traffic AI as critical infrastructure, ensure human override, and keep the offline fallback exercised rather than merely documented.
  6. Set privacy limits up front. For any smart-city sensor network, define what is collected, how long it is retained, and who may query it before it goes live.
  7. Pilot on one corridor or district. Prove the ranking catches real failures earlier than the calendar before scaling statewide, and agree in advance what result would end the pilot.

What Linda builds first

Linda does not need to instrument all 4,200 bridges overnight. She needs to take the traffic-load, age, materials and rainfall data she already has and produce a weekly ranked risk list for one high-traffic corridor. She funds a data cleanup, runs the ranking alongside her existing two-year calendar for a year, and measures whether it surfaces high-risk structures earlier than the calendar did. Running the two in parallel is the point: it costs her nothing in safety, because the calendar still governs, and it produces the comparison she will need.

If the ranking catches the next would-be culvert washout months ahead of failure, she has a defensible case to scale and a clean number to bring to the legislature instead of an after-action report. If it does not, she has learned that cheaply, on one corridor, with her inspection regime untouched. Both outcomes are worth the year. The version of this program that fails is the one that goes statewide first, replaces the calendar immediately, and has no baseline left to compare against when someone asks whether it worked.

Anti-patterns

  • Buying the model before cleaning the data. A model trained on inconsistent inspection records and drifting sensors will not fail loudly. It will rank confidently and wrongly, and the confidence is what makes it dangerous.
  • Treating the score as a diagnosis. Acting on a ranking without physical verification tears up healthy assets and burns the crew trust that makes the whole system work.
  • Measuring only what the model found. Counting the failures caught while never counting the failures the model ranked safe. The second number is the safety number, and it is invisible unless you build the loop that surfaces it.
  • Letting a forecasting tool drift into control. Operators begin following a recommendation as though it were an instruction, without any documented decision to change the tool's role. Nothing in the procurement record will show this happened.
  • Mistaking sensor coverage for risk. A prioritization that concentrates where the data is dense is measuring your instrumentation, not your infrastructure, and it will systematically under-serve communities whose assets were never well recorded.
  • An override authority nobody can exercise. Human override written into policy while the manual fallback goes unexercised for years. On the day it is needed, the authority exists and the capability does not.
  • Retrofitting privacy onto a sensor network. Deploying smart-city sensing first and deciding collection, retention and query rules after the data exists, which is when those decisions become politically expensive and technically hard.
  • Going statewide before you have a baseline. Replacing the calendar everywhere at once leaves nothing to compare against, so the program can never prove it worked and can never be defended when it is questioned.

Practice prompts

  • Pick one asset class your organization maintains and write down where its condition data actually lives, who owns each source, how current it is, and where the gaps are. Do this before any conversation with a vendor, and expect the inventory itself to change what you would buy.
  • Take your most recent unplanned asset failure and reconstruct what your agency already knew about that asset beforehand. Was the information present but unread, present but unconnected, or genuinely absent? The answer tells you whether your problem is a model problem or a data problem.
  • Write the false-negative feedback loop for a prioritization system: when an asset fails, who checks where the model had ranked it, where that gets recorded, and who reviews the running tally. If no one owns this, the model has no safety metric.
  • Take a ranked risk list and break it down by the geographies your agency already reports on. Compare that distribution to your sensor and inspection coverage, and decide whether you are looking at risk or at instrumentation.
  • For one system whose output moves something physical, run a tabletop exercise on the day the model is unavailable or wrong. Note who has actually performed the manual procedure recently, and treat any gap as a finding rather than as a training note.

Reflection

Think about the last infrastructure failure your organization had to explain publicly, and ask the Linda question about it: did the agency already hold information that would have changed the priority order beforehand? In most cases the answer is yes, and the useful follow-up is not why the model was missing. It is why the information never reached a person who could act on it. That path, from a sensor reading to a crew's work order, is the thing an AI program actually rebuilds, and it can be described without mentioning AI at all.

Then ask what your organization would do with a ranking it disagreed with. If a senior inspector's judgment and the model's list conflict, who wins, and is that written down anywhere? Agencies that have not answered this question in advance answer it by default, usually in favour of whichever is easier to defend afterwards. Deciding it deliberately, and recording the disagreements when they happen, is how the program gets better instead of merely getting used.

Glossary

  • Predictive maintenance. Prioritizing inspection and repair by an asset's estimated likelihood of failure rather than by a fixed calendar cycle.
  • Calendar-based inspection. The traditional approach of inspecting every asset on a fixed interval regardless of its condition, which spends effort evenly on assets that are not equally at risk.
  • Risk score. A model's ranking of an asset's relative failure likelihood. It directs attention. It is not a condition assessment and does not substitute for one.
  • False negative. An asset the model ranks as safe that subsequently fails. In a safety-critical system this is the error that matters most, and the one with no natural feedback loop.
  • Sensor drift. The gradual loss of calibration in a deployed sensor, which corrupts a model's inputs silently and is easy to miss in an automated pipeline.
  • Safety-impacting AI. Under OMB M-24-10, AI whose output could meaningfully affect human safety, which triggers requirements for testing, monitoring and human oversight.
  • Surveillance creep. The gradual expansion of a sensor network's collection, retention or querying beyond its original purpose, typically through a series of individually small extensions.
  • Offline fallback. The operating procedure an agency runs when the model is unavailable, wrong or compromised. It is real only if the staff who would run it have practiced it.

Closing

Infrastructure AI is unusually easy to explain and unusually easy to oversell. The honest version fits in a sentence: your agency already knows more about its assets than it can act on, and a ranking model closes part of that gap by putting the most likely failure in front of a person while there is still time to prevent it. Nothing in that sentence promises fewer inspectors, a smaller budget or an asset that fixes itself.

What it does promise is a change in the order of the queue, and that is enough. The programs that work start with the records rather than the model, keep the inspector at the point of decision, count the failures the ranking missed as carefully as the ones it caught, check whether the attention is landing fairly across communities, and prove all of it on one corridor before touching the rest of the network. Linda's culvert is the argument. The data was there. The next one is the test of whether anything changed.

Key takeaways

  • AI allocates attention, not labor. Its core infrastructure value is ranking which of thousands of assets to inspect or repair first. Crews still do the physical work.
  • Predictive maintenance beats the calendar. Prioritizing by failure risk catches problems before they become expensive emergencies, and it changes the queue rather than the budget.
  • A score is a prioritization, not a diagnosis. The model directs a crew to look. The inspector confirms condition before any repair, closure or public statement.
  • Clean the data before buying the model. Messy inspection records and drifting sensors do not produce an error message. They produce a confident ranking that puts the wrong asset first.
  • Track the false-negative rate. The assets a model rates safe that actually fail are the number that matters most in a safety-critical system, and it stays invisible unless you build the loop that surfaces it.
  • Test for investment equity. Models inherit historical neglect and can under-rank communities whose assets were never well recorded. Check whether the ranking is measuring risk or measuring your sensor coverage.
  • Harden anything touching physical controls. Grid, water and traffic AI are cyberattack surfaces with physical consequences. Keep human override real by keeping the offline fallback exercised.
  • Set privacy limits before the sensors go up. Collection, retention and query rules are far cheaper to decide in the design documents than after the data exists.
  • Pilot one corridor and keep the baseline. Run the ranking alongside the existing regime so you can prove it surfaces failures earlier, and bring that number rather than an after-action review.

Frequently Asked Questions

Do we need sensors on every asset before this is worth doing?

No, and assuming so is the most common reason these programs never start. Linda's corridor pilot runs on data her agency already collects for other purposes: age, materials, traffic load and rainfall. New instrumentation improves a model and is rarely the constraint at the beginning. The constraint is usually that existing records are scattered across systems that were never designed to be read together.

How do we know whether the model is actually working?

By keeping the old regime running alongside it for long enough to compare. The question a ranking system has to answer is whether it surfaces genuinely high-risk assets earlier than the existing cycle did, and you can only answer that if the existing cycle is still producing results. Replacing the calendar immediately feels decisive and destroys the only baseline you had.

What happens when an inspector disagrees with the ranking?

The inspector's assessment of physical condition governs, because a score is a prioritization and an inspection is a diagnosis. The important part is what happens next: record the disagreement, and review those records periodically. A pattern of experienced inspectors overriding the model in the same direction is a finding about the model, and it is the highest-quality feedback the program will ever receive.

Is a bridge prioritization model a rights-impacting system?

The classification that clearly applies is safety-impacting, because the output affects decisions about physical assets the public relies on. That triggers testing, monitoring and human oversight obligations under federal guidance. The equity question is separate and still applies: if the ranking systematically under-serves particular communities, that is a governance problem regardless of which label the system carries. Classify deliberately, write the reasoning down, and expect to defend the reasoning rather than the label.

Where does surveillance risk enter an infrastructure program?

Through the sensing layer, and gradually. A camera installed to count vehicles, a network extended to a new district, a retention period lengthened because storage got cheaper, and a query interface opened to a new department are each individually reasonable and collectively a mass-monitoring capability. Set collection, retention and access rules in the design documents and require a documented decision to change any of them.

Our data has multi-year gaps. Is the program still viable?

Usually yes, provided you know where the gaps are and say so. A documented gap is a manageable limitation, because you can exclude affected assets from the ranking, weight them conservatively, or prioritize filling the gap. An undocumented gap is the dangerous case, because the model will still produce a ranking for those assets and nothing in the output will indicate that it was reasoning from almost nothing.