←
AI for Government
Proficient · M38 · lesson 38 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Prioritization Frameworks for Government AI
📖
now learning

Prioritization Frameworks for Government AI

15 min

Solange Osei-Bonsu, IT Director for a county Department of Social Services, had seven AI proposals on her desk and a capital budget line of $1.4 million, which would not stretch across all of them. Her deputy wanted the chatbot for benefits intake. The County Administrator wanted predictive analytics for child welfare caseloads. Two supervisors had each championed a separate translation tool for different language communities. And the federal compliance team had just flagged that any automated decision-support touching eligibility would need an equity impact assessment before it could be approved. The meeting was in four days. She needed a defensible ranking, not a gut call.

Solange's problem is the ordinary condition of government AI rather than an unlucky quarter. Every project an agency wants to run competes for the same appropriated funds, the same small population of technical talent, the same backlog of authorizations, and the same limited attention from whoever holds the senior AI role. Meanwhile the opportunity surface keeps expanding, because use-case inventories surface work nobody had catalogued and shared services arrive faster than governance does. In that environment a prioritization framework is not an artifact for the strategy team. It is the operating mechanism by which scarce resources get pointed at the highest-value, highest-equity, most compliant work.

Why Scoring Matters More in Government

In the private sector, a product team can pivot fast. In government, every decision is public record. If Solange funds the wrong AI tool and it fails, the failure appears in a board audit, an investigative news story, or a Freedom of Information Act request. The cost is not just money. It is institutional trust, and trust does not appear in the capital budget line that got spent.

A scoring framework does three things for a public-sector leader. First, it forces explicit assumptions, which means the guesses get written down where someone can challenge them. Second, it creates a paper trail that survives staff turnover, so the reasoning outlasts the reasoner. Third, it gives you something to show a skeptical elected official who suspects you picked a vendor for reasons you would rather not put in writing. Think of it as a routing table for limited bandwidth: every request that arrives gets evaluated against the same rules before it gets forwarded anywhere, with no special paths and no unexplained exceptions, or at least none that go undocumented.

The Three Failure Patterns a Framework Prevents

Without a framework, three dysfunctional patterns take over, and each one is recognisable enough that most public servants have watched at least one of them run.

Loudest-voice prioritization funds the project of the most senior person or the most aggressive vendor, regardless of impact. Compliance-panic prioritization lurches from one audit finding to the next, always reacting and never building anything that compounds. Shiny-object prioritization lets whichever technology is currently exciting dominate the roadmap, even when the agency's real value sits in back-office automation, benefits processing or records management. All three feel like decisiveness from the inside. All three produce a portfolio nobody would have chosen deliberately.

The case record is not gentle about the consequences. A federal identity-verification rollout in 2022 advanced on prominence and urgency without the disparate-impact testing that an equity-weighted scoring exercise would have flagged before award. Michigan's MIDAS unemployment fraud detection system ran for years on a prioritization logic that weighted recovering improper payments far above avoiding false accusations against legitimate claimants, producing roughly forty thousand false positives and more than twenty million dollars in settlements before the state legislature restricted fully automated adjudication. The Dutch childcare benefits scandal and the SyRI risk-scoring system show what happens when fraud-recovery politics drive prioritization with no equity weight at all. Houston Federation of Teachers v. Houston ISD, decided in the Southern District of Texas in 2017, shows the cost when institutional priorities favour deployment over explainability, and State v. Loomis, decided in Wisconsin in 2016, demonstrates the same pattern in criminal justice at state level.

Read that list as a description of prioritization failures rather than technology failures. In each case the system did roughly what it was built to do. What went wrong upstream was the weighting: somebody decided, implicitly and without writing it down, which harms counted and which did not.

The Five Things Government Frameworks Must Add

RICE, ICE and MoSCoW all come from product management, and none of them handles government conditions out of the box. A sound public-sector framework has to add five things.

  • An explicit compliance lane. Statutory mandates, OMB memoranda, governance obligations, continuous monitoring duties, Section 508 accessibility and paperwork clearances cannot be traded away against feature work. They are not competing candidates; they are conditions of operating.
  • An equity weight. Rights-impacting and safety-impacting systems carry obligations to citizens who cannot take their business elsewhere, and a framework that scores them like any other project has quietly decided those obligations are optional.
  • Credit for capability building. Early projects should pay forward by building evaluation harnesses, synthetic data pipelines, model documentation practices and incident-response playbooks that compound across the whole portfolio.
  • Political viability as a constraint, not a score. A project without support from oversight bodies or civil society faces a real obstacle. That obstacle should be recorded as a constraint on feasibility rather than dressed up as merit.
  • Auditability. The output has to be legible to auditors, inspectors general, appropriations staff and oversight boards, which means the scoring sheet has to survive being read by someone who was not in the room.

Those five are the test any framework has to pass before you let it allocate money. The three that follow are the raw material you adapt to meet it.

RICE: Reach, Impact, Confidence, Effort

RICE is a scoring method developed in product management. The formula is (Reach × Impact × Confidence) ÷ Effort. Each factor gets a numeric score, and you divide the weighted numerator by the effort required. Higher scores rise to the top. The scales are conventions your agency sets and records in advance, and the one thing that matters is that everyone scores every project the same way.

One common convention runs like this. Reach on a scale of 1 to 10, where 1 is a handful of people, 5 is thousands, and 10 is millions or a fundamental capability. Impact on a scale of 1 to 4, where 1 is minimal improvement, 2 is some improvement, 3 is significant and 4 is transformational. Confidence as a percentage, where 50% is very uncertain, 75% is moderate and 100% is high. Effort in weeks, where 4 weeks is minimal, 12 weeks is moderate and 26 weeks is substantial.

Here is how each component translates to government work in practice:

  • Reach. How many residents or staff does this tool affect in one budget cycle? A benefits chatbot handling 40,000 annual intake calls sits far higher on the scale than an analytics dashboard used by 12 analysts. Count the people the tool touches, then place that count on your agreed band.
  • Impact. How much does it move a measurable outcome? Use your agency's existing performance metrics rather than inventing new ones. A meaningful reduction in processing time sits in the middle of the scale; eliminating an entire category of appeals sits at the top of it.
  • Confidence. What is your evidence base? A pilot with your own data supports a high percentage. Extrapolating from a vendor's case study in a different jurisdiction supports a low one. In government procurement, vendor claims are marketing until validated, and Confidence is the factor where that distinction gets priced.
  • Effort. Measure in person-months or procurement cycles, not dollars alone. A cheaper tool that requires a long procurement, an equity assessment, a data-sharing agreement and staff retraining costs more than it looks, and a pricier tool that extends an existing contract and runs on current infrastructure may score better despite the sticker.

Two worked examples show the arithmetic. A benefits processing automation project with Reach 8, Impact 3, Confidence 85% and Effort 20 weeks scores (8 × 3 × 0.85) ÷ 20, which is 20.4 ÷ 20, or 1.02. A tax fraud detection project with Reach 3, Impact 4, Confidence 70% and Effort 24 weeks scores (3 × 4 × 0.70) ÷ 24, which is 8.4 ÷ 24, or 0.35. On RICE alone, benefits processing wins by a wide margin, and it wins because it reaches far more people with moderately high confidence, not because anyone in the room preferred it.

Solange ran RICE on her seven proposals. The translation tool for Spanish speakers, serving 35,000 residents, with low integration complexity and a vendor already under contract, scored highest. The predictive child welfare tool, despite its political appeal, scored low on Confidence because the vendor's evidence came from a county with very different demographics.

ICE: Fast Triage for Early-Stage Ideas

Impact, Confidence, Ease is a lighter method that drops the Reach factor. It is simpler and faster than RICE, and it explicitly does not account for scale or effort, which is the trade you are making when you choose it. Use ICE when you need to cut a long list down quickly, say from twenty ideas to seven, before investing time in deeper RICE scoring. It works well at the beginning of an annual planning process when proposals are still rough.

Two conventions for combining the three factors are in circulation, and you should pick one deliberately and write it down. Scoring each factor from 1 to 10 and multiplying all three produces totals running from 1 to 1,000, which spreads projects out and makes a weak factor extremely punishing. Averaging the three instead, adding them and dividing by three, produces a score on the same 1 to 10 scale as the inputs, which is easier for a room to interpret and much more forgiving of one low factor. Neither is more correct. They rank projects differently, so switching between them mid-cycle destroys comparability, and that is the actual risk.

The trap in government is that Ease gets conflated with technically simple. A technically easy tool can be procurement-hard. Before scoring Ease, ask whether this requires a new vendor contract; whether it touches personally identifiable information and triggers privacy review; and whether it requires approval from a labor union or civil service board. Those hurdles belong inside the Ease score. A tool that rates 9 on technical simplicity and 3 on procurement complexity cannot honestly be recorded as a 9. Write down how you combined the two before you score anything, because a combining rule chosen after you see the numbers is not a rule, it is a preference.

MoSCoW: Getting Stakeholders to the Same Table

MoSCoW stands for Must have, Should have, Could have, Won't have for this cycle. Must have means a regulatory requirement or a critical mission failure if it is not done. Should have means important but not critical, needed in the medium term. Could have means it would improve things but is not essential. Won't have means explicitly not prioritized this period. It was developed for requirements management and adapts well to government AI planning because it trades false precision for genuine alignment.

Unlike RICE or ICE, MoSCoW produces categories rather than numbers. That matters when your audience includes elected officials, department heads and community advocates who distrust numeric scores they did not help create. It also matters because the Must have category maps cleanly onto the compliance lane: mandated work goes there by definition, and does not have to win an argument against discretionary projects it was never comparable to.

Run a MoSCoW session as a structured workshop. Give each stakeholder group sticky notes and a constraint: no more than 30% of items can be Must have. That cap forces the real conversation about tradeoffs that numeric scoring can hide. Solange ran a 90-minute session with department heads and got the child welfare predictive tool moved from Must have to Should have without a single argument about numbers, because the heads of other departments simply pointed out that the equity assessment timeline would push it past the fiscal year. MoSCoW also makes the "won't have this cycle" category explicit and visible, which is far easier to defend publicly than a ranked list where the bottom items quietly disappear.

Balancing Impact, Feasibility, and Compliance

Most private-sector prioritization guides talk about impact and feasibility. Government adds a third leg: compliance readiness. An AI tool with high impact and high feasibility still fails if it cannot clear the required review processes before the budget window closes, and the review calendar does not care how good the idea is.

Build compliance cost into your scoring from the start. In Solange's county, any AI tool touching eligibility determinations must complete an Algorithmic Impact Assessment, a process that takes 60 to 90 days and requires sign-off from the County Counsel's office. That goes into Effort. Any tool collecting biometric data triggers a separate ordinance review. Any tool replacing a staff function requires a labor-management consultation under the memorandum of understanding with the relevant union. None of these are optional, and none of them shorten because a project is popular.

A practical checklist before finalizing scores:

  1. Does this tool make or influence decisions about individuals? If yes, what equity review is required?
  2. Does this require new data-sharing agreements with state or federal agencies?
  3. Can it be procured under an existing contract vehicle, or does it need a new solicitation?
  4. Does it affect positions covered by civil service or a union agreement?
  5. Can the vendor provide a data processing addendum that meets your jurisdiction's privacy requirements?

Score compliance readiness on a simple three-point scale: ready now, some months to clear, or unknown. Solange's rule is that any tool rated unknown drops below every tool in the first two categories regardless of its RICE score. That is a rule her agency sets for itself and records in advance. It is not a legal test, it does not separate lawful from unlawful, and its whole value comes from being written down before anyone knows which project it will disadvantage.

The unknown category deserves more respect than it usually gets. A project whose compliance path nobody has traced is not a low-risk project, it is an unmeasured one, and the most expensive surprises in government technology arrive from that box. Treating unknown as worse than a long but understood clearance path creates the right incentive: somebody goes and finds out. Some unknowns resolve into ready now as soon as anyone asks the question. The rest turn out to be the reason the project would have failed late in the year, which is precisely when you least want to discover it.

Four Dimensions to Add to Any Scoring Method

Whichever method you pick, four government-specific dimensions have to be applied on top of it, and each one changes a ranking in a different direction.

DimensionQuestion it asksHow it acts on the ranking
Compliance imperativeDoes this address a regulatory requirement, federal mandate or policy directive?Raises priority; mandated work should be separated out rather than competing
Equity considerationDoes this disproportionately serve underserved populations?Increases priority weight
Capability buildingDoes this build capability that enables future projects?Increases priority weight, even where immediate impact is lower
Political viabilityIs there stakeholder support, especially from elected officials?A constraint on whether the project can succeed, not a scoring factor

A worked portfolio shows how those dimensions change an answer that RICE alone would have given. Project A, benefits processing automation, scores RICE 1.02, with medium compliance relevance, high equity value because it serves vulnerable populations, medium capability building and high political viability. It is the clear first priority. Project B, benefits fraud detection, scores RICE 0.35, but carries high compliance relevance because it addresses improper payment concerns, low equity value, high capability building because it creates detection infrastructure the agency will reuse, and high political viability. It ranks second on a strong case despite the lower RICE score. Project C, HR recruitment optimization, scores RICE 0.24 with low compliance relevance, neutral equity, low capability building and medium viability, and it gets deferred.

Notice what made Project B's case: compliance and capability, not politics. Leadership enthusiasm was present, and it was not the reason. That distinction is the difference between a framework that survives an audit and one that gets described as cover.

Defending Your Ranking to Skeptical Stakeholders

A scoring framework only works if you can explain it out loud to someone who is angry that their project ranked third. The score does not decide. The score documents what was decided and why, so the next team can audit the decision.

When presenting results to leadership, lead with the problem rather than the methodology. "We cannot fund all seven proposals this cycle. Here is how we evaluated what to fund first, and here is the evidence behind each factor." Attach the scoring sheet as an appendix. Do not spend twenty minutes explaining RICE. Spend that time explaining the outcomes the top-ranked projects are expected to produce and the evidence behind those expectations, because that is the only part your audience can actually evaluate.

When a supervisor pushes back, return to the routing table. "We applied the same rules to every proposal. The translation tool scored higher on Confidence because we have a working pilot with verified data. The analytics tool scored lower because the vendor's evidence does not match our population. We can revisit it next cycle with a local pilot." That answer works because it is falsifiable: it names the factor, the evidence and the condition under which the ranking would change.

Document every adjustment made after stakeholder feedback. If someone in authority overrides your score, note it in writing: what changed, who directed it, and on what basis. This is not defensive paperwork, it is institutional memory. Two budget cycles from now, someone will ask why the overridden project underperformed, and the record is the difference between an answer and a guess. Plan to revisit the whole prioritization quarterly, because circumstances change and a ranking treated as permanent stops being a decision and becomes an inheritance.

When Politics and Procurement Override Good Scoring

Sometimes a well-scored proposal still loses, because a board member committed to a vendor at a conference, or a state grant requires a specific solution, or procurement rules effectively eliminate all but one supplier. This happens. It is not a failure of your framework; it is a feature of democratic governance. Elected officials have authority you do not have and grant conditions are real constraints. None of that means you abandon scoring.

When a politically directed choice is unavoidable, do three things. Score it anyway and put the score in the record alongside the direction you received. Negotiate scope, so that a low-scoring tool arrives as a bounded pilot before anyone commits the full amount. And set explicit success criteria in the contract and the board memo. If the tool hits its metrics within the agreed window you have validation; if it misses them you have the documentation to course-correct without embarrassment.

Procurement rules are a separate constraint, and a tool you cannot legally buy is not a viable option regardless of its score. Know your own thresholds: most jurisdictions set a dollar level above which competitive bidding is required and a higher one above which the governing body must approve, and sole-source justifications are usually available but must be documented and publicly posted. Those figures vary by jurisdiction and the only ones that govern you are the ones in your own purchasing ordinance, which is a document worth reading rather than a number worth remembering from a lesson. What travels is the pattern: if a vendor offers a pilot priced just below your competitive-bid threshold, that is a flag rather than a convenience.

The routing table holds even when external routing rules apply. You document the external override, you bound its scope, and you keep scoring everything else by the same rules.

Anti-Patterns

  • Overweighting political viability. A senior official wants a project, so it goes to the top despite lower objective scores. The agency invests in lower-value work while better projects wait, and over a few cycles the prioritization process loses the credibility that made it useful. Keep political viability as a constraint, asking whether the project has enough support to succeed, and where it lacks support, work on building support rather than adjusting the framework until it produces the desired answer.
  • Underweighting compliance. Compliance gets scored as one factor among several, so a must-do obligation loses to a more interesting discretionary project and the problem surfaces later at a worse moment. Separate mandated compliance work out of the competition entirely. Do it first, then prioritize the discretionary portfolio with whatever capacity remains.
  • Ignoring capability building. Scoring purely on immediate impact means the agency never builds evaluation harnesses, data pipelines or playbooks, stays dependent on external expertise, and cannot sustain adoption. Score capability building explicitly and let early projects earn priority for the infrastructure they leave behind, even at lower immediate impact.
  • Using the framework as cover. The real decision gets made through politics and the scoring exercise is run afterwards to justify it. People work this out quickly, and once they do, every future ranking is read as theatre. If the framework ranks Project A first and you choose Project B for other reasons, say so plainly and explain the reasons. Credibility comes from honest application, not from the framework always agreeing with you.
  • Treating scores as objective. A number carries an authority its inputs have not earned. Reach bands, Impact scales, Confidence percentages and compliance cutoffs are all conventions your agency chose, and two projects separated by a fraction of a point have not been shown to differ at all. Use scores to structure an argument and to make assumptions visible, never to end a discussion.
  • Setting the cutoff after seeing the results. Deciding where the funding line falls, or how two Ease sub-scores combine, once you already know which projects it advantages, converts a framework into a rationalization. Fix the rules in writing before scoring, and if you later change them, record the change and rescore everything.
  • Switching methods mid-cycle. RICE and the two ICE conventions rank projects differently, and moving between them part-way through destroys comparability across the portfolio while looking like a methodological improvement. Choose one, document it, and finish the cycle on it.
  • Scoring Ease on technical difficulty alone. The procurement calendar, the privacy review, the union consultation and the data-sharing agreement are all part of how hard a thing is to actually do. A tool that is simple to build and impossible to buy this fiscal year is not an easy tool.

Practice Prompts

  • Choose a framework and justify the choice. RICE is comprehensive, ICE is faster, MoSCoW gives category clarity where numbers would be contested. Which matches your organization's decision-making style and your audience's tolerance for scoring?
  • Take five competing AI projects and score each with RICE. Calculate the scores, rank them, and then ask the honest question: do the rankings make sense, and where they do not, is the problem the projects or your scales?
  • For each of those projects, assess the four government dimensions: compliance imperative, equity consideration, capability building and political viability. Note every place where a dimension moves a project past one that outscored it on RICE.
  • Write a one-page defence of your top three priorities. Then hand it to a colleague whose project is not on the list and ask them to argue against it.
  • Draft the communication for people whose projects were not funded. How will you show the process was fair, and what specifically would have to change for their project to rank higher next cycle?
  • Run the compliance checklist against one proposal you expect to fund. Time each review it triggers, add the total to Effort, and rescore. Watch whether it still ranks where it did.
  • Write down your combining rules before your next scoring session: the Reach bands, the Impact scale, the ICE convention, and how sub-scores merge. Date the document.

Reflection

Think about the last significant AI or technology investment your agency made. Could you reconstruct today why it was chosen over the alternatives, and could someone who was not in the room reconstruct it from what was written down? Then ask the harder version: if that project underperformed, would the record show whether the failure was in the scoring, in an override, or in an assumption that turned out to be wrong? Those three have different remedies, and only one of them is fixed by a better framework. The value of scoring is largely that it makes the difference between them visible two years later, when everybody who made the decision has moved on.

Glossary

  • RICE: A prioritization method scoring Reach, Impact, Confidence and Effort, combined as (Reach × Impact × Confidence) ÷ Effort to produce a ranked list.
  • ICE: A lighter method scoring Impact, Confidence and Ease. Two combining conventions are in use, multiplication and averaging, and they rank projects differently.
  • MoSCoW: A categorical method sorting work into Must have, Should have, Could have and Won't have for the current cycle, producing alignment rather than a numeric ranking.
  • Reach: How many residents or staff a tool affects within one budget cycle, placed on an agreed band rather than compared as raw counts across projects.
  • Confidence: The strength of the evidence behind a Reach and Impact estimate, which is where the difference between your own pilot data and a vendor's case study gets priced.
  • Effort: The work required, measured in person-months or procurement cycles rather than dollars, and inclusive of review and clearance time.
  • Compliance readiness: A separate score for whether a project can clear its required reviews inside the available window, rated ready now, some months to clear, or unknown.
  • Capability building: The degree to which a project leaves behind infrastructure, skills or reusable artifacts that make subsequent projects cheaper.
  • Political viability: Whether a project has sufficient support from those with authority to affect its success. A constraint on feasibility, deliberately not a scoring factor.
  • Algorithmic Impact Assessment: A documented review of a system's likely effects on the people it touches, frequently a precondition for deploying decision-support that affects eligibility.
  • Override record: The written note of a decision that departed from the scored ranking, capturing what changed, who directed it and on what basis.

Closing

Prioritization frameworks are tools for making fair, transparent and defensible decisions about where to invest limited resources. They do not make the decisions for you. They make the decision-making rigorous enough to explain, and explicit enough that the assumptions inside it can be challenged while there is still time to change the answer.

Solange went into her meeting with a ranked list, a scoring sheet and a one-page explanation of the evidence behind each factor. She did not get everything she recommended, which is normal. What she got was a record: of what was scored, what was overridden, who directed the override and on what basis. Two budget cycles later, that record is the reason her department can say why it built what it built. Staff turnover, leadership changes and public records requests all arrive eventually. A framework that lives in a shared document survives them. A gut call does not.

Key Takeaways

  • RICE quantifies tradeoffs. Reach × Impact × Confidence ÷ Effort gives a defensible ranking across proposals that are otherwise hard to compare, provided Effort includes procurement and compliance time rather than implementation hours alone.
  • ICE is for speed, not precision. Use it to trim a long list early and RICE for final budget decisions. Pick one combining convention, multiplication or averaging, write it down, and do not switch mid-cycle, because the two rank projects differently.
  • MoSCoW builds shared ownership. Category-based workshops get stakeholders to accept tradeoffs voluntarily, which matters where no single person has unilateral authority, and the Must have category is where mandated compliance work belongs.
  • Government frameworks must add five things. A compliance lane, an equity weight, credit for capability building, political viability treated as a constraint, and auditability by people who were not in the room.
  • Compliance readiness is a first-class factor. An equity impact assessment, a labor consultation or a new data-sharing agreement adds months to a timeline and must appear in Effort before rankings are finalized, not after.
  • Scores are conventions, not measurements. Every band, scale and cutoff is something your agency chose. Scores structure an argument and surface assumptions; they do not settle questions, and a narrow gap between two projects settles nothing at all.
  • Fix the rules before you see the results. A cutoff or a combining rule chosen after the numbers are known is a preference wearing a framework's clothes, and readers of your audit trail will notice.
  • Three failure patterns are what you are preventing. Loudest voice, compliance panic and shiny object. The case record shows each of them producing real harm, and in every instance the technology did what it was built to do while the weighting was wrong.
  • Document every override. When political direction or procurement rules displace a score, record the instruction, the source and the rationale. This protects the institution and creates the audit trail future decisions depend on.
  • Scope political mandates down. If a low-scoring project must proceed, negotiate a bounded pilot with fixed money, a fixed timeline and explicit success metrics before full commitment.
  • The framework's value is institutional, not individual. Staff turnover, leadership changes and records requests all arrive eventually. A scoring framework in a shared document survives them; a gut call does not.

Frequently Asked Questions

Doesn't a scoring framework just launder political decisions?

It can, and that is the fastest way to destroy one. The failure mode is running the numbers after the decision has already been made and presenting them as the reason. People work this out, usually sooner than the person doing it expects, and afterwards every ranking you produce is read as theatre. The defence is honesty about departures: when the framework ranks one project first and you fund another, say so explicitly, record who directed it and why, and keep scoring everything by the same rules.

Which framework should we actually use?

Match it to the decision and the audience. RICE for final budget allocation, where the stakes justify the effort and you have evidence to put behind Confidence. ICE for early triage, when you are cutting twenty rough ideas down to a handful worth analysing properly. MoSCoW when your audience includes elected officials or community advocates who will distrust numbers they did not help produce, and where alignment matters more than ranking. Many agencies run ICE first, MoSCoW for stakeholder alignment, and RICE for the final cut.

How do we score a project whose benefits are hard to quantify?

Score it honestly and let the Confidence factor carry the uncertainty, which is exactly what that factor exists for. The temptation is to inflate Impact because the project feels important, and the discipline is to leave Impact where the evidence puts it and record low Confidence instead. That combination produces the right next action, which is usually a small pilot to generate the evidence rather than a full commitment made on hope.

Our RICE scores put two projects within a few percent of each other. Which wins?

Neither, on the scores. A gap that small is inside the noise of your own estimating, because every input was a judgment placed on a band your agency invented. Treat close scores as a signal that RICE has done all it can and the decision now needs the government dimensions: compliance imperative, equity, capability building, and whether either project can actually clear its reviews this cycle. Record which of those broke the tie.

What do we do about a project leadership wants that scores poorly?

Score it anyway, put the score in the record next to the direction you received, and then negotiate the shape rather than the fact. A bounded pilot with fixed money, a fixed timeline and explicit success criteria in the contract and the board memo gives you a decision point that arrives on a date rather than never. If it hits its metrics you have validation and you were wrong, which is fine. If it misses them you have documentation that lets you stop without a fight.

How often should we rerun prioritization?

Quarterly is a reasonable default, because circumstances change faster than annual budget cycles acknowledge: evidence arrives, a review clears, a grant condition appears, a vendor fails. The point of rerunning is not to churn the portfolio but to catch the cases where a Confidence score was low and is now high, or where a compliance path rated unknown has resolved. A ranking that is never revisited stops being a decision and becomes something the agency inherited.