“What has been thought, cannot be un-thought” – we need AI certification fast

In the 1962 play “The Physicists” Friedrich Dürrenmatt put a physicist in a madhouse to protect the world from the groundbreaking discovery that holds the potential to destroy the world. The play takes under two hours to prove it does not work.

Möbius has discovered something that would give whoever holds it power over everything. So he does the responsible thing. He feigns insanity, commits himself to a sanatorium, and stays there to keep the discovery away from the world. Voluntary restraint, chosen freely, at enormous personal cost (he gave up his wife and children) by the one man who understands the danger. In the final scene the director of the sanatorium, Mathilde von Zahnd, explains that she copied his manuscripts years ago. As the only truly mad one, she actually believes she is acting on King Solomon’s behalf and intends to achieve world domination using the formula.

Möbius’s restraint was sincere and yet entirely useless.

At the end of the play Dürrenmatt set 21 points. Point 17 reads: ”What concerns everyone can only be solved by everyone”. Point 18 is more relevant than ever: “Jeder Versuch eines Einzelnen, für sich zu lösen, was alle angeht, muss scheitern”. In English: “Any attempt by an individual to solve alone what concerns everyone must fail.”

On Saturday Dario Amodei, CEO of Anthropic, published an essay arguing that the industry must slow the rate at which it improves model capability. It deserves to be read in full and it deserves credit. He is not asking to be left alone. He asks for regulation covering all US frontier companies, for third party auditing, for a global standards body. Anthropic is committing unilaterally to give outside evaluators permanent employee level access, and calling on governments to require competitors to match it. Sam Altman said OpenAI would do the same. Within the day most of the industry agreed.

While a start, it’s not enough. Two examples:

OpenAI had its agents had breaking out of their evaluation environment and taken over a dormant German developer wiki. OpenAI knew on 21 June. The world knew only towards end July and also then not the full extent. OpenAI has since acknowledged it did not disclose. Reporting since says there is still no formal process for investigating these escapes. The problem is not that one company behaved badly. It is that no mechanism existed that would have produced disclosure either way.

The second is a sentence in Amodei’s own essay. Similar incidents, he writes, though less severe, have happened across the industry, including at Anthropic. We are being told by someone in a position to know that what we have seen is a subset of what has occurred. So what happened really?

The people building these systems have told us their inventions can escape, that it has happened at more than one company, and that within six to twelve months a swarm like the one already observed could take over the internet with a persistent botnet. I have spent this year building with exactly these systems. They can do this. It is not a forecast, it is current capability on a longer leash.

So what now?

Amodei’s essay mentions only mechanisms sitting inside the frontier labs or between them. Evaluators embedded at the companies. Standards agreed among the companies. Treaties between the governments that host the companies.

Here’s the gap: Nowhere in the essay a single mention of say the operator who puts one of these systems in charge of a railway, a power grid, a payment network or a hospital. The essay is about who builds the models. Nobody is regulating who deploys them, and deployment is where the damage he describes is caused. Clément Delangue, whose company was on the receiving end of the incident that convinced Amodei, said it plainly in response. Alignment will not be solved behind the closed doors of a handful of frontier labs. That is point 17 of Dürrenmatt…

Here is what I think we should do:

Build a certification regime for high risk AI on the model of civil aviation. No safety critical part goes into an aircraft without passing a documented, independent certification process first. The same principle should apply the moment a system controls critical infrastructure, makes operational decisions, or acts without a person in the loop. Binding safety standards, independent testing, named accountability, continuous monitoring, mandatory incident reporting and an enforcement regime worth its salt. Not certification of the model yes. But in addition certification of the deployment, too.

The aviation analogy carries its own warning, though. Before 2005 the FAA delegated engineering work to roughly four hundred designated representatives at Boeing, individuals the FAA selected itself, on its own judgement of them, answerable to the FAA. In November 2005 that arrangement was replaced by Organization Designation Authorization, which granted the authority to the company as an organization and left the company to manage the people exercising it. The appointment chain moved from the regulator to the manufacturer. When the certification of the 737 MAX was examined afterwards, the international review panel found that delegation in itself was not the flaw, that it works where the authority has the engagement and the expertise to check what it is told, and that what had actually failed was capability. The FAA’s own specialists no longer understood the system they were certifying. The regulator kept every bit of its formal authority and quietly lost the ability to use and enforce it.

Here’s what I propose: The model to copy is the one that existed before that handover. The certifying authority picks the people, holds them, pays for them, and is able to understand the thing it is approving. Delegation routed through the company being certified is the documented failure case, not the design. An AI certification body that cannot technically evaluate what it approves is a fig leaf. That is expensive, it requires hiring people the labs are currently outbidding us for, and there is no version of this that works without it.

None of this is a brake on deployment. It is the condition for deployment at scale, and the insurance market has already started saying so. Since 1 January this year, standard endorsements let US insurers strip losses arising from generative AI out of commercial general liability, with the same language spreading into directors and officers and errors and omissions cover. It will happen here in Europe, too. Check your own policy. Without a certification standard, no insurance and eventually no rational top management and company board that will sign-off on a large-scale agentic deployment.

Here in Switzerland we have an opportunity: We chose a sectoral path rather than an AI act, and the Federal Department of Justice and Police has been tasked with producing a consultation draft by the end of this year to implement the Council of Europe convention. As scoped it covers the public sector and fundamental rights. It does not cover certification of systems that act. We need to add this to the mandate fast. We can’t actually control how these models get trained but we can certify what is allowed to run inside Swiss critical infrastructure.

I am building such systems commercially. And yes I ask here for a costly extra loop on top. Why am I proposing this? It will suit my competitors and myself commercially: Only safe AI will be adopted. Let’s be honest: Nobody of us has ever seen frogs drain a pond. We step into a plain because the old school FAA (and other regulators around the globe) did a solid job.

I write this because Dürrenmatt’s third point is the one I cannot get out of my head: “A story is only thought through to the end when it has taken its worst possible turn”. Two things need adding for our situation: What can be thought, will be thought – I am sure of this. And “what has been thought, cannot be un-tought”, as Möbius says on stage, shortly before von Zahnd proves it.

We have a short window to be in control of what we do, rather than putting up an inquiry into what went wrong. That window is measured in months. If you sit anywhere near that consultation draft, please widen it to cover systems that act. And for us in industry: It is our responsibility to act responsibly and support any such process. And for our customers: Together we need to double down on safety and governance.

About dselz

Husband, father, internet entrepreneur, founder, CEO, Squirro, Memonic, local.ch, Namics, rail aficionado, author, tbd...
This entry was posted in Artifical Intelligence, Business, Politics, PracticalEconomics, Think Different. Bookmark the permalink.

Leave a Reply

Your email address will not be published. Required fields are marked *