2 min read

The Federal Government Is Buying AI at Scale. What could go wrong?

The Federal Government Is Buying AI at Scale. What could go wrong?
The Great Wave off Kanagawa by Katsushika Hokusai - Metropolitan Museum of Art

AI agents operating entirely outside anyone's awareness, doing things nobody authorized, on systems nobody was watching. That’s what.

That scenario is closer than agencies realize, because they have no evaluation framework to detect it. In April 2026, GAO released Artificial Intelligence Acquisitions: Agencies Should Collect and Apply Lessons Learned to Improve Future Procurements (GAO-26-107859), examining AI acquisitions at several agencies (Citation #1).

The GAO findings are striking.

On the testing and evaluation side, GAO found that agencies have struggled with early testing and continuous evaluation because AI systems vary widely and are difficult to assess, and that agencies consistently lack the data scientists and technical experts needed to evaluate contractor proposals before award.

The result is agencies buying systems they cannot fully evaluate, deploy them without baseline performance standards, and retire them without recording what they learned.

The absence of evaluation infrastructure is compounded by a cost problem.

Officials from Project Linchpin (a DoD AI infrastructure initiative) separately warned GAO that federal agencies are chronically underestimating the ongoing infrastructure costs required to sustain AI systems after initial deployment.

These anecdotes reflect a structural gap between a traditional system acquisition: discrete, well-specified products with knowable lifecycles and what AI procurement actually requires: continuous evaluation, adaptive contracting, and sustained technical capacity.

The government has a pre-release evaluation framework through NIST's Center for AI Standards and Innovation (Citation #2), which conducts 30-to-90-day security reviews of frontier models from the major AI labs before public release. But that framework operates at the model level, before deployment.

Inside agencies, once a system goes live against casework, financial data, or operational decisions, there is no equivalent evaluation framework.

For contractors and advisors supporting federal AI programs, the GAO findings are both a warning and an opportunity.

Agencies that accept GAO's recommendations to update policies will mandate lessons-learned documentation. The lessons could become compliance requirements that will flow downstream into contracts, statements of work, and program management requirements.

Agencies that have not built evaluation frameworks will need help. Contractors can bring that rigor to federal AI.

Citations:

1. GAO-26-107859, Artificial Intelligence Acquisitions: Agencies Should Collect and Apply Lessons Learned to Improve Future Procurements, April 13, 2026. Coverage via Biometric Update and Nextgov/FCW.

2. US Government AI Testing: CAISI and Frontier Models 2026, DDR Innova, May 2026.