top of page

The Federal AI Incident Registry: Standardizing Documentation and Oversight for Federal AI Systems

  • Nathan Russell
  • Jun 3
  • 17 min read

Updated: 6 days ago

Section 1: The Problem

The federal government is deploying AI at a rapid pace. The GAO reported that AI use cases across 11 federal agencies nearly doubled from 571 to 1,110 in a single year. (GAO, 2025) CDT’s separate review found a 200% increase across the full federal inventory, from 710 to 2,133 use cases. (CDT, 2024) These systems are making decisions in benefit eligibility, law enforcement, medical diagnosis, and immigration processing, all domains where a wrong output has a direct effect on someone’s life.


Fracture One: Classification

The Center for Democracy and Technology found that the Department of Housing and Urban Development classified an AI translation tool as “rights-impacting,” which triggers testing, documentation, and oversight obligations under OMB Memorandum M-25-21. The Department of Homeland Security did not classify a functionally similar translation tool in the same category, meaning the obligations required for HUD did not apply to DHS. (CDT, 2024) Same technology, two agencies, opposite risk determinations. This happens because agencies are not required to document or publish the reasoning behind their risk determinations. CDT found that inventory reporting is “incredibly variable within and between federal agencies, including different levels of detail and different approaches to reporting and categorizing the risk level of use cases.” (CDT, 2024) The inconsistency exists even within agencies, where some subcomponents completed required inventory fields while others in the same department left them blank. The Department of Justice’s inventory entries contain no information about risk mitigation or governance procedures for any of its use cases, and many DOJ cases were designated “too new to fully assess.” (CDT, 2024) These designations included AI tools used for criminal investigations, license plate reading, gunshot detection, and recidivism prediction. Without consistent classification, the accountability requirements that attach to “high-impact” AI apply differently at different agencies with no mechanism to reconcile the differences.


Fracture Two: Documentation

Fifty federal agencies were required to submit AI compliance plans. Only 31 did. (EPIC, 2024) Of those that were submitted, EPIC found that “the most pervasive and problematic themes across all compliance plans is their abstract and high-level nature. Some compliance plans are essentially regurgitations of the M-24-10 requirements, with no detailed specifics of how the requirements will be applied and addressed within the agency’s processes.” (EPIC, 2024) The Department of Homeland Security’s compliance plan requires 41 additional pages across four separate policy documents to understand their approach. (EPIC, 2024) The SEC, when asked to describe barriers to responsible AI use, said it planned to form a working group to identify them. EPIC noted that “the SEC doesn’t yet have a working group that will assess such barriers, much less a plan to mitigate or remove any identified barriers.” (EPIC, 2024) There is no common form, no required fields, and no common standard for what an accountability record should contain. When something goes wrong, there is no consistent trail to follow.


Fracture Three: Incident Reporting

Covington and Burling’s October 2024 analysis of OMB’s AI procurement memo found that while OMB requires vendors to report “serious AI incidents and malfunctions” within 72 hours, (Covington & Burling, 2024) the memo “authorizes individual agencies to determine the criteria for what constitutes a serious AI incident or malfunction.” Every agency defines that threshold differently. Reports go to different destinations in different formats on different timelines. Covington noted that “harmonization has not been achieved for pre-existing cybersecurity incident reporting requirements that have been in place for a significant period of time now.” (Covington & Burling, 2024) The federal government already tried this approach in cybersecurity and it failed. When a system fails at one agency, the knowledge stays inside that agency. There is no shared repository, no cross-agency notification, and no mechanism for the government to learn from its own AI failures at scale.


None of this is the result of negligence or bad faith. It is what happens when you build accountability frameworks independently, agency by agency, without asking whether they add up to a coherent system. The OMB directed agencies to classify, document, and report, but never specified a shared standard for any of the three. The result is a federal government generating governance artifacts at every stage of AI deployment and learning nothing from any of them.


Section 2: Existing Frameworks and Gaps

OMB Memorandum M-25-21, issued April 3, 2025, is the federal government’s primary framework for managing AI. (OMB, 2025) It has real requirements. Agencies have to put risk management practices in place for high-impact AI. Every CFO Act agency has to name a Chief AI Officer to oversee compliance and keep track of high-impact systems. Agencies have to build AI strategies within 180 days, share custom code and models across government, and have a plan to shut down any high-impact system that is not working. This paper does not argue that nothing exists.

The problem is that M-25-21 tells agencies what to do but not how to do it. Each agency is left to figure out the process on its own, and that is where things fall apart.


Where Classification Breaks Down

M-25-21 creates two categories of high-impact AI: “rights-impacting” and “safety-impacting.” Those are the right categories. But the memo tells agencies to decide for themselves whether their systems qualify, with no shared process for making that call. (CDT, 2024) The labels exist. The method for applying them does not.


Where Documentation Breaks Down

M-25-21 tells agencies to document their AI governance but says nothing about what that documentation should look like. No standard format, no required fields, no minimum level of detail. (EPIC, 2024) So agencies produce compliance plans that check the box without actually describing what their systems do. The VA’s plan says nothing about a mental health tracking app running passively on veterans’ phones. (CDT, 2023) The SSA’s plan says nothing about an AI anti-fraud system that was deployed based on a misrepresented statistic and later found to have flagged only two potentially fraudulent claims out of over 110,000 processed. (Alms, 2025)


Where Incident Reporting Breaks Down

M-25-21 and the October 2024 OMB procurement memo both require agencies to report AI incidents, but each agency gets to decide what counts as serious enough to report. Covington found that the federal government has not even managed to standardize cybersecurity incident reporting after decades of trying, so there is no reason to expect AI reporting will be any different without intervention. (Covington & Burling, 2024) The GAO confirmed as much in March 2026 with report GAO-26-107681, finding that OMB’s AI guidance does not fully address privacy-related risks. (GAO, 2026) Of 10 expert-identified challenges, the guidance fully covers 2 and only partially covers the other 8. Both of GAO’s recommendations remain open.


M-25-21 gave agencies the right goals but none of the tools to reach them. The requirements to classify, document, and report are all there. What is missing is anything that would make classification consistent, documentation useful, or reporting comparable from one agency to the next. That is the gap this paper is trying to close.


Section 3: Strategic and Political Stakes

The structural gap documented in Sections 1 and 2 is not an abstract policy problem. It has concrete costs. The government repeats its own mistakes because no mechanism exists to share what went wrong. Accountability becomes nearly impossible when something fails.


The Government Repeats its Own Mistakes

In September 2023, the GAO examined how seven federal law enforcement agencies used commercial facial recognition services, including Clearview AI, Marinus Analytics, and Thorn. The findings showed a pattern of deployment without basic safeguards. (GAO, 2023) None of the seven agencies required their staff to complete facial recognition training before using the tools. Across six of them, roughly 60,000 searches had been conducted before training requirements even existed. Only three agencies had developed policies addressing the civil rights implications of facial recognition. The FBI had 10 out of 196 staff trained. ATF headquarters did not know its own staff had been sending photos to Clearview AI. (GAO, 2023)


None of that knowledge carried over. In 2025, ICE and CBP deployed Mobile Fortify, a facial recognition application built by NEC, and classified it as high-impact. Neither agency published the Privacy Impact Assessment that OMB guidance required before deployment. (Cushing, 2026) WIRED found that DHS had accelerated approval by dismantling centralized privacy reviews and removing department-wide limits on facial recognition use. (Cameron & Varner, 2026) Internal assessments acknowledged the application does not actually verify identities. The agency deployed it anyway. (Cameron & Varner, 2026)

The consequences showed up quickly. In April 2025, Juan Carlos Lopez-Gomez, a U.S. citizen carrying a Social Security card, was arrested on suspicion of being an unauthorized immigrant. ICE detained him for 30 hours based on what the agency called “biometric confirmation of his identity.” (Taper, 2025) Jensy Machado, another U.S. citizen, was held at gunpoint by ICE agents in a separate case of mistaken identity. (Taper, 2025) These are U.S. citizens who lost their liberty because a facial recognition system got it wrong and nothing in the federal government’s accountability structure caught the error before it reached them. The Secret Service, CBP, ICE, FBI, ATF, and DEA all ran into similar problems with facial recognition independently over seven years. If a shared incident registry had existed after the GAO’s 2023 findings, every agency deploying facial recognition would have known about the documented failures before their own systems went live. Mobile Fortify, Lopez-Gomez, and Machado were all preventable. None of them were prevented.


The IRS Algorithm

Facial recognition is not the only domain where this pattern shows up. In January 2023, a Stanford research team led by Daniel Ho found that Black taxpayers were audited at 2.9 to 4.7 times the rate of non-Black taxpayers. (Ho et al., 2023) The most extreme disparity was among single Black men with dependents claiming the Earned Income Tax Credit, who were nearly 20 times as likely to be audited as non-Black married taxpayers claiming the same credit. (Ho et al., 2023) This was not intentional discrimination. The IRS used internal algorithms that were never publicly disclosed, and those algorithms prioritized the likelihood of underreporting over the dollar amount. That design choice systematically targeted lower-income filers who are disproportionately Black. IRS commissioner Danny Werfel confirmed the findings in a letter to Congress in May 2023 and pledged to be “laser-focused” on addressing the disparities. (Ho et al., 2023) A subsequent GAO review in May 2024 found no evidence the IRS had ever conducted a comprehensive review of the rules and filters in its Dependent Database, and discovered that some risk scores varied by sex and had not been updated since 2001. (GAO, 2024) For over two decades the IRS ran an algorithm with measurable bias and never reviewed it. No monitoring metric was defined when the system was deployed, no reporting mechanism existed, and the structural bias went unchecked until outside researchers found it.


Michigan’s MiDAS: The Cautionary Tale

Michigan’s MiDAS system is the most extreme example of what happens when none of these accountability components exist. From October 2013 to September 2015, the Michigan Unemployment Insurance Agency used the Michigan Integrated Data Automated System to detect unemployment fraud. MiDAS processed 40,195 fraud cases by algorithm alone. No human review, no appeal process before penalties were imposed, no documentation of how the system made its determinations. The error rate was 85%. Over 34,000 people were falsely accused of unemployment fraud. (Charette, 2018; AI Incident Database, n.d.) The state imposed 400% penalties on the claimed fraud amounts, the highest in the nation, and immediately garnished wages and seized tax refunds. People lost their homes, filed for bankruptcy, and were referred for criminal prosecution based on outputs that were wrong the vast majority of the time. The UIA defended the system for two years before a federal lawsuit forced it to stop in September 2015. The agency never explained why MiDAS failed or why it ignored internal warnings. The class action was not settled until January 2024, over a decade later, for $20 million. (University of Michigan Ford School, 2024) MiDAS was a state system, not a federal one. But it is exactly what this paper’s proposed registry is designed to prevent: an automated system making consequential decisions about people’s lives with no classification of its risk level, no accountability record, and no way to catch its failures before they became a crisis.


The Political Costs of an Ungoverned Failure

The current administration has made AI leadership a central part of its technology agenda. The Sacks National AI Legislative Framework, published in March 2026, calls for removing barriers to innovation and establishing American dominance in artificial intelligence. (The White House, 2026) The America’s AI Action Plan, published in July 2025, declares that “the United States is in a race to achieve global dominance in artificial intelligence.” (The White House, 2025) That ambition comes with risk. If a federal AI system fails publicly and the administration cannot show that it had a real accountability structure in place, Congress has every reason to intervene. When that happens, the first questions will be who authorized the deployment and what evidence supported the decision. Right now there is no standardized record of who signed off, what testing was done, what the known error rates were, or when the system would be rolled back. The registry this paper proposes does not slow down AI adoption. It gives the administration something to point to when things go wrong.


Section 4: The Federal AI Incident Registry

The three fractures documented in this paper are not three separate problems. They are one structural failure showing up at three stages of the AI deployment cycle. The proposed Federal AI Incident Registry closes all three with a single mechanism built on three components: a standardized classification rubric, a minimum accountability record, and a shared incident registry at NIST. Each component maps to one fracture. Each depends on the others to function. Together they create something that does not exist anywhere in the federal government right now: a continuous loop where classification informs documentation, documentation supports incident response, and incident data feeds back into how the government classifies and manages risk going forward.


Component One: The Standardized Classification Rubric

M-25-21’s definitions of “rights-impacting” and “safety-impacting” are the right ones. The problem, as documented in Sections 1 and 2, is that agencies are told to apply those definitions but given no shared process for doing so. (CDT, 2024) The proposed rubric gives agencies a structured process for determining whether a system qualifies as rights-impacting or safety-impacting under M-25-21. It asks two questions:


1. What happens if the system gets it wrong? Does the system’s output feed into a decision that has a legal, material, or significant effect on someone’s rights, liberties, or safety? If a wrong output can lead to a denied benefit, a false arrest, a missed diagnosis, or a wrongful investigation, the system meets the first threshold. The nature of the harm determines the category. If the effect falls on an individual’s civil rights, privacy, equal opportunities, or access to government services, the system is rights-impacting. If it affects human life, the environment, or critical infrastructure, it is safety-impacting. A system can be both.

2. How much human review stands between the output and the action? Is a person reviewing every consequential output before it is acted on, or is the system operating with partial or no human oversight? If the system can trigger action without full human review, it meets the second threshold.

If a system meets both, it is classified as high-impact under M-25-21’s existing definitions. This does not replace M-25-21’s definitions. It gives agencies a consistent process for applying them. Two agencies can still reach different conclusions about a similar system, but the rubric ensures they are working through the same analysis and documenting how they got there.


The critical addition is transparency. Right now an agency can classify a system as “not high-impact” and never publish the reasoning behind that determination. The rubric requires agencies to publish the completed analysis, not just the label but the logic behind it. If two agencies evaluate similar systems and reach different conclusions, the published reasoning makes it possible to see where and why they diverged. Without that, classification is just an internal checkbox. With it, Congress, inspectors general, and the public can actually evaluate whether the determination makes sense.


Component Two: Minimum Accountability Record

The documentation problem is not that agencies fail to produce records. It is that the records they produce contain fundamentally different information depending on the agency. (EPIC, 2024; GAO, 2025) The minimum accountability record fixes this by establishing 10 required fields that every agency must complete for every system classified as high-impact:

1. System name and unique identifier — so the same system cannot appear under different names at different agencies.

2. Deploying agency and sub-component — identifying the specific office operating the system, not just the department.

3. Vendor and version information — for tracking when multiple agencies deploy the same commercial tool.

4. Operational domain — such as law enforcement, benefits adjudication, or fraud detection.

5. Date of deployment — establishing how long the system has been operational.

6. Classification rubric results with published reasoning — from the two-factor analysis in Component One.

7. Data sources and training methodology — for identifying potential bias in how the system was built.

8. Human oversight mechanism and review rate — documenting what human review actually occurs, not just whether it exists on paper.

9. Known limitations and failure modes — requiring agencies to disclose what the system cannot do and where it is known to fail.

10. Point of contact for public inquiries — Congress, inspectors general, researchers, and the public have a named person to reach.


Agencies already collect most of this information under M-25-21. What changes is the format. When the IRS deploys a fraud detection algorithm and CBP deploys a risk assessment tool, the documentation would follow the same structure and cover the same categories. Stanford’s 2023 research on the IRS found that algorithmic tools flagged Black taxpayers for audits at three to five times the rate of other filers, but that finding required outside researchers to reconstruct information the agency had never systematically recorded. (Ho et al., 2023) A standardized accountability record makes that information available before a disparity turns into a scandal.


Component Three: A Shared Federal Registry Housed at NIST

The registry would be housed at the National Institute of Standards and Technology. NIST already manages the AI Risk Management Framework, (NIST, 2023) it operates as a standards body rather than a regulatory authority, and it has existing relationships with every federal agency through its technical standards work. Those three qualities make it the right home. An agency is more likely to participate in a system managed by a standards organization than one run by the office that oversees its budget.


The federal precedent for this model is the FAA’s Aviation Safety Reporting System. ASRS has operated for 49 years, collected over 2.1 million reports, and maintained zero confidentiality breaches. (NASA, 2024) Three design features make it work, and this registry replicates all three:


1. Third-party management. ASRS is managed by NASA, not the FAA, separating the entity that collects reports from the entity that enforces regulations. Housing the AI registry at NIST rather than OMB replicates that independence.

2. Standardized intake. Every ASRS report follows the same format regardless of who submits it. The minimum accountability record serves the same function for AI systems.

3. Safe harbor. Pilots and controllers who submit reports within 10 days receive protection from punitive enforcement action. This is what drives participation. Without it, reporting is a liability.


The safe harbor provision for the AI registry works the same way. Agencies that submit classification rubrics, accountability records, and incident reports within established timelines receive protection from punitive oversight action based on the disclosed information. If reporting a problem leads to punishment, agencies will not report problems. Forty-nine years and 2.1 million reports prove that safe harbor works.


Section 5: Political Alignment and Feasibility

A proposal for a federal AI accountability registry invites an obvious objection: that it amounts to more regulation at a time when the administration has made clear it wants less. But the registry does not restrict what AI systems agencies can build, procure, or deploy. It does not create a new agency or a new approval process. It standardizes how the federal government documents and reports what it is already doing.


The political foundation for this proposal already exists. Executive Order 13960, signed in December 2020 during President Trump’s first term, established nine principles for federal AI use including transparency, accountability, and responsible governance. (Executive Order 13960, 2020) The AI in Government Act, signed into law in the same period, required agencies to inventory their AI use cases and report to Congress through the GAO. (P.L. 116-260, 2021) This paper does not propose a new law. It proposes amending an existing one to add the three components described in Section 4, which is a significantly easier legislative path than a standalone bill. America’s AI Action Plan, released in July 2025, reinforced both removing barriers to AI adoption and maintaining public trust. (The White House, 2025) The Sacks Framework, released in March 2026, called for promoting American competitiveness and removing barriers to innovation while ensuring the public’s trust in how AI is developed and used. (The White House, 2026) The registry satisfies all three. It creates a consistent federal standard that vendors and agencies can build to instead of navigating 24 different agency interpretations of the same definitions. It replaces fragmented documentation practices with a single process. And it makes classification logic, accountability records, and incident data publicly available through NIST, which is the transparency mechanism the current framework lacks.


Agencies are more likely to deploy AI confidently when they know exactly what is required of them rather than navigating ambiguous guidance that changes from department to department. The registry gives them that predictability. It also gives the public visibility into what their government is doing with AI systems that affect their rights and safety. Every piece of this proposal follows through on a position the administration or Congress has already taken.


Section 6: Recommendations

This paper proposes a two-track implementation strategy. The two tracks are not alternatives. They are sequential. The first delivers immediate operational improvement, the second makes it permanent.


Track One: OMB Supplemental Guidance (90-Day Timeline)

OMB already has the authority to issue supplemental guidance to M-25-21 without new legislation. (OMB, 2025) Within 90 days, OMB should require all federal agencies to adopt the standardized classification rubric and the minimum accountability record described in Section 4. M-25-21 already requires agencies to classify their systems and maintain documentation. The supplemental guidance specifies how: the two-factor test as the classification method and the 10-field record as the documentation standard.

Track Two: Statutory Amendment to the AI in Government Act (180-Day Timeline)

The shared registry at NIST requires a legislative foundation. Congress should amend the AI in Government Act of 2020 (P.L. 116-260, 2021) to formally establish the registry, designate NIST as the managing entity, codify the safe harbor provision, and mandate annual reporting to Congress on registry data and cross-agency findings. The AI in Government Act already requires agencies to inventory their use cases and report to the GAO. A centralized registry with standardized intake is a natural extension of what the law already demands. The 180-day timeline reflects that legislation moves slower than executive guidance, but this is an amendment to an existing law, not a standalone bill.


Why Two Tracks?

If Congress moves slowly, the OMB guidance track still delivers two of the three components within 90 days. Agencies start classifying and documenting consistently right away. By the time the legislative track delivers the NIST registry, the first two components are already generating comparable data ready to populate it. Partial implementation still improves the status quo. Full implementation closes the loop.


Section 7: Conclusion

The federal government’s use of AI is operational. GAO documented AI use cases across federal agencies nearly doubling from 571 to 1,110 in a single year, with CDT’s broader review finding over 2,100 use cases across the full federal inventory. (GAO, 2025; CDT, 2024) These systems are making decisions that affect people’s rights, liberties, and safety every day. The question is no longer whether the government should use AI. The question is whether the government has the infrastructure to know what its systems are doing, how they are classified, and what happens when they fail.


Right now it does not. The evidence in this paper shows classification that is inconsistent across agencies, documentation that varies so widely it cannot be compared, and incident reporting with no centralized mechanism for the government to learn from its own failures. These are not three separate problems. They are one structural gap at three stages of the deployment cycle.


The Federal AI Incident Registry closes that gap. The classification rubric, the minimum accountability record, and the shared registry at NIST give the federal government something it does not currently have: a way to track what its AI systems are doing, hold agencies to a consistent standard, and learn from failures before they repeat. The design is modeled on 49 years of proven federal precedent through the Aviation Safety Reporting System, and the implementation path runs through existing executive authority and existing legislation. The foundation is already there. This proposal builds on it.


References

Alms, N. (2025, May 15). DOGE went looking for phone fraud at SSA — and found almost none. Nextgov/FCW. https://www.nextgov.com/digital-government/2025/05/doge-went-looking-phone-fraud-ssa-and-found-almost-none/405346/


AI Incident Database. (n.d.). Incident 373: Michigan’s MiDAS automated unemployment fraud detection system. Retrieved from https://incidentdatabase.ai/cite/373


Cameron, D., & Varner, M. (2026, February 5). ICE and CBP’s face-recognition app can’t actually verify who people are. WIRED. https://www.wired.com/story/cbp-ice-dhs-mobile-fortify-face-recognition-verify-identity/


Center for Democracy and Technology. (2023, August 15). More than meets the AI: A look at some trends in federal AI inventories. https://cdt.org/insights/more-than-meets-the-ai-a-look-at-some-trends-in-federal-ai-inventories/


Center for Democracy and Technology. (2024). Needle in the AI-stack: An analysis of the 2024 federal AI use case inventory update. https://cdt.org/insights/needle-in-the-ai-stack/

Charette, R. N. (2018). Michigan’s MiDAS unemployment system: Algorithm alchemy created lead. IEEE Spectrum.


Covington & Burling LLP. (2024, October). OMB releases requirements for responsible AI procurement by federal agencies. https://www.cov.com/en/news-and-insights/insights/2024/10/omb-releases-requirements-for-responsible-ai-procurement-by-federal-agencies


Cushing, T. (2026, February 12). ICE, CBP knew facial recognition app couldn’t do what DHS says it could, deployed it anyway. Techdirt.


Electronic Privacy Information Center. (2024). An analysis of federal agency AI compliance plans. EPIC.


Exec. Order No. 13960, 85 Fed. Reg. 78939 (2020). Promoting the use of trustworthy artificial intelligence in the federal government.


Government Accountability Office. (2023). Facial recognition services: Federal law enforcement agencies should have better awareness of systems used by employees (GAO-23-105607). https://www.gao.gov/products/gao-23-105607


Government Accountability Office. (2024). Artificial intelligence: IRS should address potential bias in selecting tax returns for audit. https://www.gao.gov/products/gao-24-106649

Government Accountability Office. (2025). Artificial intelligence: Agencies’ inventory of use cases has grown, but better practices needed (GAO-25-107653). https://www.gao.gov/products/gao-25-107653


Government Accountability Office. (2026). Artificial intelligence: OMB guidance does not fully address privacy-related risks and challenges (GAO-26-107681). https://www.gao.gov/products/gao-26-107681


Ho, D. E., Huq, A., & Kovacs-Goodman, D. (2023). Algorithmic fairness and the IRS: Racial disparities in income tax auditing. Stanford Institute for Economic Policy Research.

National Aeronautics and Space Administration. (2024). Aviation Safety Reporting System: Program overview. https://asrs.arc.nasa.gov/


National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework (AI RMF 1.0). U.S. Department of Commerce. https://www.nist.gov/artificial-intelligence/ai-risk-management-framework


Office of Management and Budget. (2025). Memorandum M-25-21: Advancing the responsible acquisition and use of artificial intelligence in the federal government. Executive Office of the President.


P.L. 116-260, AI in Government Act of 2020, 134 Stat. 1182 (2021).

Taper, J. (2025, June 6). Facial recognition is getting a lot more invasive under Trump. Slate.

The White House. (2025). America’s AI Action Plan.


The White House. (2026). The Sacks framework: A national AI legislative framework.


University of Michigan Ford School of Public Policy. (2024). MiDAS class-action settlement reached in Michigan unemployment fraud case.

Recent Posts

See All
bottom of page