UNAUTOMATEDTHE HUMAN INTELLIGENCE PROJECT

Automate the work. Never automate the human.

HUMAN JUDGMENT. REAL CONSEQUENCE. A BETTER FUTURE.

  • DISCERNMENT. Examine what is true.
  • PRUDENCE. Decide what is worth doing.
  • WISDOM. Act with good judgment.

UNAUTOMATED.ORG

All Briefs

The Unautomated Brief

An AI Audit Is Not a Stamp of Approval

Sep 12, 20268 min read
Editorial image of a magnifying glass examining a small paper document, illustrating careful independent review rather than automatic approval.
Image: Magnifying glass examines small paper with text, Alejandro García Cordero / Unsplash, licensed under the Unsplash License.

This article is independent commentary by UNAUTOMATED. The Office of Governor Gavin Newsom and the State of California did not sponsor, review, or endorse it.

An audit can be a valuable moment of truth. It can require an organization to slow down, gather evidence, invite an outside view, and face questions that may otherwise be ignored. Those are meaningful benefits. But an AI audit is not a stamp of approval, and it should never become an excuse to stop asking responsible questions.

That distinction matters as more organizations make claims about AI safety, assurance, and independent review. On September 9, California Governor Gavin Newsom announced that he had signed Senate Bill 813 and Assembly Bill 1405. According to the Governor’s announcement, the bills establish standards connected to independent verification organizations, third-party audits, and assessments of AI systems. AB 1405 also creates a state registry for AI auditors and sets standards around their independence, transparency, and integrity.

That is the reported development. It is a framework for stronger outside review and clearer expectations around who performs that work. It does not mean every AI system must now be audited. It also does not mean that an audit, by itself, proves a particular system is safe, fair, accurate, or appropriate for every setting. Those are different claims, and responsible leaders should keep them separate.

Independent review can make accountability harder to avoid

Organizations are often tempted to grade their own work generously. A team that built or bought an AI tool may understand its promise better than its weak points. Tight deadlines, budget pressure, and a desire to show progress can make uncomfortable evidence easy to discount. An independent reviewer can bring distance to those pressures.

At its best, an outside review asks simple but demanding questions. What is this system actually being used for? What information does it rely on? Who may be affected if it makes a mistake? What was tested, and what was not? Are the organization’s claims supported by records that someone else can examine?

That kind of review can improve accountability because it turns broad promises into specific questions. It can expose missing documentation. It can reveal that a system was tested in one setting but deployed in another. It can show that a vendor’s reassuring language is not the same as evidence. It can also give employees and customers a clearer basis for asking leaders to explain their choices.

California’s new framework is notable for recognizing that the quality and independence of the reviewer matter. A review that is vague, conflicted, or hidden from scrutiny does little to build trust. Standards for the people and organizations doing the assessment can help make the work more credible.

Still, credibility is not certainty. An audit is a snapshot of a defined piece of work at a particular time. The quality of the result depends on the scope, the information made available, the method used, and the willingness of the organization to address what the review finds.

Scope is where confidence can become misleading

Every assurance claim has boundaries. A review may examine a model’s documented controls, but not how employees use it in practice. It may test a limited set of examples, but not the unusual cases that appear after launch. It may assess whether a policy exists, but not whether people have the training and authority to follow it. None of those limits makes the review worthless. They do mean the limits should be clear.

Leaders should be cautious when they hear a simple statement such as, “The AI passed an audit.” Passed what, exactly? For which use? Against which standard? With what evidence? A responsible answer names the system, its intended purpose, the date of review, the reviewer, the criteria, and the known exclusions. A vague answer turns a useful process into marketing.

This is especially important when an AI tool influences consequential decisions. A system used to draft internal notes carries a different level of risk from one that affects employment, housing, credit, healthcare, education, public benefits, safety, or access to essential services. A review suited to the first situation may not be enough for the second. The right level of care rises with the possible impact on people.

Evidence must be documented, not implied

Good assurance work leaves a trail. That does not require releasing private information or publishing every technical detail. It does require enough documentation for responsible people to understand what was examined and how conclusions were reached.

For an organization, that record may include the system’s purpose, the data it is allowed to use, the decisions it may influence, testing records, reported problems, review findings, and the actions taken afterward. It should also identify who accepted the remaining risk. If no one is prepared to put their name next to that decision, the organization may not be ready to make it.

Documentation matters because memory fades and teams change. A new manager should not have to reconstruct why an AI tool was approved. An employee responding to a concern should not have to guess whether the system was tested for the situation in front of them. Clear records make it possible to learn, correct, and explain.

Findings need correction paths, not filing cabinets

An audit that identifies a problem but produces no change is not accountability. It is paperwork. The real test comes after the report: who receives the findings, who is responsible for responding, what must be fixed, and when will progress be checked again?

Some findings may call for a small change in instructions or staff training. Others may require tighter access, a new approval step, more frequent monitoring, a pause in use, or a decision not to deploy the system at all. Leaders should create those correction paths before a problem arrives. Otherwise, the pressure to keep moving can overwhelm the evidence that says a change is needed.

People affected by an AI-supported decision also need a meaningful route to question it. When a result matters, they should not be left with a machine’s answer and no responsible person to contact. Human review is not a ceremonial button at the end of a process. It means someone has the information, authority, and time to look at the case and make a real judgment.

The responsible person is still in the room

An auditor can examine a system. A consultant can recommend controls. A vendor can provide documentation. None of them can take the place of the leader who chooses to use the tool in a real organization with real people.

That leader must decide whether the use is appropriate, whether the scope of review matches the stakes, whether the evidence is sufficient, and whether the organization will act when problems appear. Responsibility cannot be outsourced with a contract, a certificate, or a polished assurance statement.

This should not make organizations dismiss independent review. It should make them use it well. A thoughtful audit can be a disciplined checkpoint. It can surface blind spots and give a team a stronger basis for improvement. But the purpose is not to collect a reassuring label. The purpose is to make better decisions before harm occurs and to respond honestly when the evidence changes.

Five questions leaders should ask about any AI audit claim

Before accepting an assurance claim at face value, ask these five practical questions:

  1. What exactly did the review cover? Name the system, use case, date, criteria, and important exclusions. Do not treat a review of one feature or setting as approval for every possible use.
  2. Was the reviewer genuinely independent and qualified? Ask about conflicts of interest, relevant experience, and whether the reviewer had enough access to examine the evidence rather than simply repeat claims.
  3. What evidence supports the conclusion? Look for documented testing, records, limitations, and findings—not only a summary statement or a badge.
  4. What happens when the review finds a concern? Confirm there is a named owner, a correction plan, a timeline, and a way to pause or change use when the risk requires it.
  5. Where is meaningful human oversight? Identify the person who can question a result, intervene in a real case, and explain the organization’s decision to those affected.

The strongest assurance claim is not, “Nothing can go wrong.” It is, “We know what this system is for, we have examined the evidence, we will be honest about its limits, and responsible people will remain accountable for what happens next.”

That is a higher standard than a stamp of approval. It is also the standard trust requires.

Source

This Brief draws on the following primary source. Read the original source for its full account and context.

Primary source: Office of Governor Gavin Newsom, “Governor Newsom signs first-in-the-nation AI safeguards to protect Californians, calls on the federal government to do its part” — September 9, 2026

See something that needs review?

We welcome corrections and additional context. Please do not include confidential or sensitive information.

Report a correction
    An AI Audit Is Not a Stamp of Approval | The Unautomated Brief | UNAUTOMATED