Executive summary
Federal agencies increasingly obtain artificial-intelligence capabilities through contracts, cloud platforms and software subscriptions. That makes procurement a governance system. The decisive rules are often written before a model is deployed: what evidence a vendor must provide, which performance thresholds apply, whether an agency can test the system independently, what records are retained, who may override an output and how the government can leave a failing supplier. Current Office of Management and Budget guidance addresses competition, data rights, vendor lock-in, high-impact uses and ongoing testing. WSOI recommends converting those principles into a standard lifecycle control for high-impact AI acquisitions. No such system should influence a consequential decision until the agency has defined the decision, measured a non-AI baseline, established disaggregated acceptance tests, preserved human authority, secured reproducible logs and funded an exit path. Contracts should govern model changes and incidents after award rather than treating acceptance as a one-time event.
Key findings
- An AI acquisition is not only a software purchase. It can delegate evidence collection, classification, recommendation and workflow control to a vendor-managed system.
- The most effective safeguards are requirements, data rights, testing access and remedies written before award—not ethics language added after deployment.
- Accuracy alone is not an adequate acceptance criterion. Agencies need measures tied to the real decision, relevant subgroups, failure severity and comparison with the existing process.
- High-impact systems require traceable versions, inputs, outputs, human interventions and incidents so that a decision can be reconstructed later.
- Human review is meaningful only when reviewers have time, authority, information and a practical ability to disagree with the system.
- Vendor portability and government access to records are continuity controls. Without them, an agency can become unable to audit, improve or replace a system it remains responsible for.
Procurement is where public AI becomes governable
Public discussion about artificial intelligence frequently begins with model behavior: hallucinations, bias, privacy, security or automation. Agencies encounter those risks through a more ordinary mechanism. A program office defines a need, an acquisition team writes a solicitation, vendors make claims, evaluators select an offer and a contract determines what happens next.
That sequence sets the boundaries of later oversight. If the solicitation requires only a generic accuracy percentage, the agency may not receive the subgroup results or error analysis needed to understand harm. If a commercial license restricts testing, the agency may lack the right to reproduce a vendor’s evaluation. If a subscription provides no exportable logs, a later appeal may depend on screenshots or memory. If a model changes without notice, a test performed at award may describe a system that no longer exists.
The federal government remains accountable for a public function even when a contractor supplies the model. Federal Acquisition Regulation 37.114 already recognizes the underlying principle for service contracts: work that can influence government authority or decision-making requires sufficient government oversight and must not displace inherently governmental functions. AI makes this boundary easier to blur because a system can shape a decision without formally signing it.
Current federal guidance provides a foundation, not a complete operating standard
OMB Memorandum M-25-22, issued April 3, 2025, directs agencies to improve competition and manage risks when acquiring AI. It addresses intellectual-property rights, government data, privacy, vendor lock-in, compliance with separate high-impact AI requirements, ongoing testing, performance remedies and notice of new features. The memorandum specifically contemplates independent evaluations using agency-defined data that vendors cannot access and calls for sufficient detail to verify or reproduce vendor testing when practicable.
OMB Memorandum M-25-21 applies governance and minimum risk-management practices to federal AI, including systems acquired on an agency’s behalf. It treats some uses as high impact when outputs serve as a principal basis for decisions or actions with significant effects on rights or safety. The exact classification depends on use, not simply on whether a vendor markets the product as “AI.”
The National Institute of Standards and Technology’s AI Risk Management Framework organizes responsible practice around governance, mapping context, measuring risk and managing it. The Government Accountability Office’s accountability framework similarly emphasizes governance, data, performance and monitoring. These frameworks are useful because they shift attention from a single model score to a managed system.
The remaining problem is operational consistency. An acquisition workforce needs clauses, evidence packages, review gates and named decision rights. A principles document cannot by itself determine whether a particular agency has the data, expertise or contractual leverage to act.
Define the decision before describing the tool
A solicitation should begin with the public decision or workflow, not a request to “add AI.” The requirement should state who is affected, what information is considered, what output the system produces, how that output is used and what authority remains with a government official.
This definition allows an agency to measure the existing process. A baseline might include processing time, cost, disagreement between reviewers, false positives, false negatives, appeal rates, accessibility failures or unequal error patterns. Without a baseline, a vendor can demonstrate that its model performs a task while the agency cannot show that the acquisition improves the public service.
The baseline also exposes cases in which automation is unnecessary. A simpler rules engine, improved form, additional staff or better data exchange may solve the problem with lower operational risk. Market research under FAR Part 10 is intended to identify tradeoffs and suitable acquisition approaches; it should compare non-AI alternatives rather than assume the desired technology in advance.
Classify risk by use and consequence
The same technical model can present different risks in different settings. A summarization tool used to organize publicly available comments is not equivalent to a system that ranks benefit applications, flags alleged fraud or recommends an enforcement action. Procurement controls should scale with the consequence of error and the degree of reliance.
WSOI recommends three acquisition tiers:
- Assistive: low-consequence drafting, search or administrative support where a trained user reviews the complete work product.
- Operational: systems that prioritize cases, allocate resources, control access or materially shape staff attention, even when a person makes the nominal final decision.
- High impact: systems used as a principal basis for decisions affecting rights, benefits, safety, eligibility, employment, enforcement or access to essential public services.
Classification should occur before solicitation and be reviewed when the use changes. A tool purchased for assistive drafting can become operational if staff begin relying on its rankings. Contract scope and oversight must follow actual use, not the original label.
Require an evidence package, not a marketing claim
Vendors should provide a standardized evidence package proportionate to risk. For high-impact systems, the package should include:
- The model and product versions evaluated, including relevant dependencies and retrieval components.
- The intended use, prohibited uses and operating conditions.
- Data provenance, coverage, known exclusions and lawful-use constraints.
- Evaluation methods, test-set design, performance by relevant subgroup and uncertainty intervals where appropriate.
- Known failure modes, adversarial or security testing and mitigation limits.
- Accessibility, privacy and records-management implications.
- Human-review workflow and escalation design.
- System architecture sufficient to identify which component produced a disputed result.
- Change-management, incident-response and end-of-service plans.
Trade-secret protection may limit public disclosure of some details. It should not prevent authorized government evaluators from obtaining enough information to assess performance and reproduce material tests. The agency should distinguish transparency to the public, access for auditors and access for operational staff rather than using confidentiality as an all-or-nothing concept.
Test the system against the public task
Vendor benchmarks are rarely sufficient for acceptance. The agency should test with data that resembles deployment conditions and that the vendor did not use to tune the system. This principle appears directly in M-25-22. The contract must supply time, access and technical support for independent evaluation.
Acceptance criteria should connect to the actual harm. In a screening system, false negatives and false positives may have different costs. In a language system, a fluent answer that invents a legal authority can be more dangerous than an obvious failure. For an operational model, latency and availability may be as important as predictive performance. Aggregate scores should be accompanied by distributions, edge cases and results for legally or operationally relevant groups.
The agency should compare the system with the measured baseline and a simpler alternative. A model that performs well in isolation may still be inferior when staff time, appeals, data preparation, security controls and vendor fees are counted.
Human review must be designed as a control
Stating that a person remains “in the loop” does not establish meaningful review. Reviewers may see only the model’s recommendation. They may lack time to inspect source material, believe disagreement will be penalized or learn that managers expect the automated queue to be followed.
For consequential uses, the contract and operating procedure should specify what evidence the reviewer sees, when review is mandatory, who can override the output, how disagreement is recorded and how an affected person can obtain reconsideration. The agency should measure override and reversal patterns. Extremely low disagreement may indicate exceptional performance; it may also indicate automation bias or a review process with no practical independence.
Some decisions should remain nondelegable. A contractor or model may assemble information or identify inconsistencies, but officials exercising sovereign authority must retain responsibility for final enforcement, adjudication and policy judgments where law or public accountability requires it.
Preserve the evidence needed to reconstruct a decision
An appeal or investigation may occur months after an output. By then, a hosted model may have changed. Prompts, retrieval indexes, system instructions, thresholds or safety layers may also differ. A defensible record therefore requires more than the final answer.
For each consequential transaction, the system should retain an appropriate record of the model or service version, relevant input, material retrieved evidence, output, confidence or score where meaningful, rules applied, human reviewer, override, final action and later correction. Retention must comply with privacy, security and records law; more logging is not automatically better. The objective is a minimum sufficient record that can explain the decision without creating an unnecessary repository of sensitive information.
Treat model changes and incidents as contract events
Cloud AI changes continuously. A vendor may update a model, alter a retrieval system or introduce a feature without the traditional release cycle used for installed software. M-25-22 recognizes this problem by addressing new-feature notification, monitoring and rollback.
Contracts should define which changes require notice, regression testing or renewed approval. A material change includes any update reasonably capable of altering performance, legal compliance, security, data use or the explanation available to affected people. Agencies should reserve the ability to freeze a version, disable a feature or roll back after a failed test.
Incident clauses should define severity levels and notification periods. A serious incident may involve unauthorized disclosure, discriminatory performance, systemic fabrication, security compromise, a harmful automated action or evidence that the system operated outside approved conditions. The vendor should preserve relevant records, cooperate with investigation and disclose whether similar incidents affect other government customers, subject to lawful limits.
Build an exit before dependency becomes irreversible
Vendor lock-in is not only a pricing concern. It can become an accountability failure if the government cannot export its data, reproduce decisions or transition a public service without losing operational knowledge. M-25-22 identifies knowledge transfer, data and model portability, rights in code or models developed under contract, and transparent licensing and pricing as possible protections.
Every high-impact acquisition should include an exit plan with export formats, documentation, transition assistance, retention and deletion obligations, continued access to historical logs and a tested continuity procedure. The government does not need ownership of every commercial model. It needs sufficient rights and information to continue its mission, meet records obligations and verify decisions made while the contract was active.
The WSOI federal AI acquisition gate
No high-impact deployment until seven conditions are met
- Purpose: The affected decision, user, population and legal authority are documented.
- Baseline: The existing process and at least one simpler alternative have been measured.
- Evidence: An authorized evaluator has received the required system evidence package.
- Performance: Independent, deployment-relevant tests meet disaggregated acceptance thresholds.
- Authority: Human review, override, notice and appeal pathways are operational.
- Traceability: Versioned records can reconstruct a consequential output and final action.
- Continuity: Change control, incident remedies and a funded exit path are contractually enforceable.
The senior program official, contracting officer, chief information officer or designee, privacy official and responsible AI official should sign the deployment gate according to agency structure. Approval should expire or require review after a material change, serious incident or defined operating period.
Implementation without freezing innovation
A standard gate can slow low-risk experiments if applied without judgment. Agencies should therefore provide lightweight pathways for assistive tools using nonsensitive information and no consequential reliance. Sandboxes should use representative but protected data, clear time limits and an explicit rule that pilot outputs do not silently become operational decisions.
Reusable contract language and shared evaluation services can reduce duplication. Smaller agencies should not need to build independent model-testing laboratories for every purchase. Government-wide acquisition vehicles can provide baseline clauses, but each agency must still define its use, data, thresholds and public responsibilities.
Procurement officials also need technical support. Contracting expertise, program expertise, legal review, cybersecurity, privacy, accessibility and model evaluation are distinct capabilities. Cross-functional review should occur while requirements can still change, not after proposals arrive.
Notice and redress are part of system performance
When an AI-supported process affects an identifiable person, the agency should decide before award what that person will be told. Notice should identify the role of automation in terms a reader can use, the principal information considered, the channel for correcting inaccurate data and the route to human reconsideration. It need not reveal source code or security-sensitive controls.
Appeals supply operational evidence. Reversal rates, recurring data errors, inaccessible notices and differences among offices can reveal problems that a predeployment benchmark misses. Contracts should require vendors to support investigations and corrections without charging the government as though every disputed output were a new customization.
Agencies should publish aggregate performance and incident information for high-impact uses when law permits. Public reporting can include volume, purpose, evaluation dates, major limitations, appeals and corrective actions without exposing personal records. Transparency should be useful enough to test accountability, not a catalog of product names.
Learn across the federal portfolio
A government-wide repository should collect approved clauses, evaluation designs, serious incident patterns and exit lessons. The repository should distinguish reusable controls from agency-specific results and protect procurement-sensitive information. Its purpose would be institutional memory: a failure already identified by one agency should not need to be rediscovered through another public deployment.
OMB should publish an annual acquisition scorecard showing how many high-impact systems passed, failed or conditionally passed their deployment gates; how often material changes triggered reevaluation; and whether agencies successfully exported data or changed vendors. The scorecard should measure governance outcomes without creating a public ranking that encourages agencies to conceal difficult uses.
Limitations
This framework addresses civilian federal acquisition and does not resolve classified, intelligence or weapons-system requirements. It does not prescribe one universal metric because error costs and lawful decision criteria differ across programs. It also cannot make poor data fit for use or supply technical staff an agency does not have.
AI definitions and products will continue to change. The standard is intentionally based on function, evidence and consequence rather than a fixed list of model types. Agencies should update implementation details as NIST revises the AI Risk Management Framework and as OMB or the Federal Acquisition Regulation changes.
Conclusion
The federal government does not need to choose between using AI quickly and governing it seriously. It needs procurement that makes speed conditional on inspectable evidence and reversible commitments.
A public agency can outsource software, hosting and technical support. It cannot outsource responsibility for the decision. The contract should make that fact operational from the first requirement through the final retained record.
The resulting standard is intentionally compatible with innovation: vendors can compete on performance and approach, while the government defines the evidence, rights and public safeguards that every acceptable solution must provide.
Sources and methodology
This paper compares current federal AI memoranda, government accountability frameworks and acquisition rules. Recommendations are WSOI Institute’s synthesis and do not represent the position of any cited agency.
- Office of Management and Budget, M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government.
- Office of Management and Budget, M-25-21, Accelerating Federal Use of AI through Innovation, Governance, and Public Trust.
- National Institute of Standards and Technology, AI Risk Management Framework.
- U.S. Government Accountability Office, Artificial Intelligence Accountability Framework.
- Federal Acquisition Regulation, Part 10, Market Research.
- Federal Acquisition Regulation 37.114, Special Acquisition Requirements.
- Federal Acquisition Regulation, Part 39, Acquisition of Information Technology.
Publication information
Published: July 31, 2026
Author: WSOI Institute
Funding: Independently self-funded; no external sponsor supported this paper.
Suggested citation: WSOI Institute, “Buying AI for the Public: A Federal Procurement Standard for High-Impact Systems,” July 31, 2026.
