tl;dr: SFC is offering a prize of $5MM - $20MM for initiating “Independently Supervised Peer-Assisted” (ISPA) inspections between frontier AI labs, specifically for pre-deployment security risks like lab leaks or sandbox exfiltration. Jaan Tallinn will contribute a minimum of $5MM to the prize, plus another $5MM at 1:2 matching rate with other funders.
What if an independent, broadly trusted, non-governmental auditor could hire frontier AI labs as contractors to help inspect each other for pre-deployment security risks, with the supervision and participation of the independent auditor?
Lab leaks, sandbox exfiltration, and other pre-deployment AI risks are an increasing concern in AI development, whereby every frontier AI lab poses some risk to public safety if AI escapes or is stolen from the lab. Every frontier lab is thus also a risk to every other lab, as members of the public.
Fortunately, this yields a natural incentive for each lab to oversee other labs and protect the public from pre-deployment risks. However, institutional boundaries between labs prevent mutual oversight — mostly for good reasons — which makes it very tricky for public safety incentives to fully take effect.
Governments are the “hammer” typically expected to break through those institutional boundaries, to audit security measures against pre-deployment risks, especially catastrophic exfiltration. However, the use of government authority has its own drawbacks, being an implicit threat of force backed by police and military might.
We therefore envision a complementary approach that is initiated, coordinated, and supervised by non-governmental expert auditors, who collectively “ask nicely” on behalf of public interest for AI labs to assist as contractors in inspecting one another’s pre-deployment security. So, for each “direction” of the mutual inspections, there should be roles for three institutions:
1) the Independent Supervisor, an independent, non-governmental institution with AI auditing expertise, who initiates and is primarily in charge of the security inspection, and acts on behalf of public interest;
2) the Peer Assistant frontier lab, an industry peer to the inspected lab, which serves as a contractor enlisted by the Independent Supervisor to increase the breadth and depth of the inspection, providing additional expert personnel, computing resources, and AI resources;
3) the Inspected frontier lab, which agrees to be inspected to improve their pre-deployment security, and to build trust in those improvements.
In one sentence: the Independent Supervisor enlists the Peer Assistant frontier lab to help inspect the Inspected frontier lab’s pre-deployment security on behalf of the public. Hence, ISPA stands for Independently Supervised and Peer-Assisted.
The direction of the inspection could also be reversed, where the Inspected lab becomes the Peer Assistant lab and conversely, yielding a mutual inspection. It may even be easier to achieve both directions of the inspection at once. On these details the prize is agnostic.
Crucially, an ISPA inspection should be conducted with limited mutual information access on a need-to-know basis, to avoid flows of intellectual property between the labs that are unnecessary for the inspection. This yields a creative challenge for the Independent Supervisor to design and facilitate information flows that are task-appropriate to mitigating pre-deployment risks. In particular, the inspection itself should be careful to avoid leaking vulnerable security information that would undermine the intended security benefits of the inspection.
The ISPA Inspection Prize will be paid out to a combination of natural persons and institutions who initiate, coordinate, and supervise the inspection(s), if anyone is identified as succeeding at all. The sharing proportions will be decided by an after-the-fact credit assessment, where prospective winners are asked to fairly assign mutual credit scores to each other for initiating, coordinating, and supervising the inspection(s). Credit assignments that seem more fair will be taken more seriously in our final allocation, so there will be little or no incentive for any prospective winner to assign full credit to only itself.
The prize money can be donated or received as compensation. The total prize money may be increased somewhat based on our assessment of the audit quality, with the following considerations:
The Independent Supervisor(s) should possess expertise in AI auditing, should participate directly in the inspection process, and should come away with justified confidence that the inspection was productive and informative.
Ideally, all of Anthropic, Google, OpenAI, and xAI should be tasked as Peer Assistants to inspect each other for pre-deployment risks. If they are not all involved, more involvement is better.
Each frontier lab should bring unique resources to the table for inspecting each other, which would not otherwise have been available to the Independent Supervisors, beyond API access.
The public should have access to as much information as possible about the inspection process and results, without jeopardizing the security of the labs.
Whose job should it be to protect the public from AI risks? We argue it should be a mixture of various entities, with a mixture of advantages and disadvantages that limit the effectiveness and desirability of each:
| Entity | Comparative Advantage | Disadvantages |
|---|---|---|
| Frontier AI labs | Engineering talent. Frontier AI labs are staffed with many talented engineers with expertise in developing, testing, deploying, and monitoring AI systems. | Imperfect incentives. Frontier AI labs have some incentives to protect their corporate interests and missions from pre-deployment risks, but are not perfectly aligned with public safety. |
| Business customers | Commercial demand. Business customers of AI labs pay for AI, and are thus well-positioned to provide detailed feedback and demand for safety measures. | Post-deployment bias. Customer businesses are not responsible for what frontier AI companies do before the AI is deployed in a product or service. |
| Governments | Authority. Governments can unilaterally place binding demands on AI labs, backed by the use of force. | Authoritarianism. Government interference with private enterprise can be heavy handed, given their backing with police and military authorities, which can be stifling. |
| Nonprofits & individual researchers | Independence. These entities don’t answer to shareholders or voters, thus bringing unique perspectives on risk evaluations. | Resource constraints. Nonprofits and individuals have less money than big businesses or governments. |
The ISPA prize aims to mix the advantages of the first and last rows, by mobilizing resources and expertise already recruited and developed by frontier AI labs, to audit each other for pre-deployment security risks. This differs from a non-profit or individual being solely responsible for the technical details of the audit, precisely because of the additional resources and expertise the frontier developers can bring to the table.
To inquire about eligibility, potential winners may optionally request to be vetted in advance by SFC for relevant supervisory expertise, independence, and institutional integrity, using the following form:
Eligibility Inquiry Form
For grants to support a prospective or ongoing ISPA inspection, interested parties may optionally request funding from https://survivalandflourishing.fund/.
Q: Can you provide more details on how to conduct the inspection?
A: We’d rather not be too prescriptive about how it’s done, since it’s a politically difficult task, and the people running the audit will have opinions and face constraints of their own. The key is strong coverage of pre-deployment risks, including audits of both the lab’s AI models as well as surrounding infrastructure and processes for preventing catastrophic pre-deployment incidents.
Q: Is this tractable?
A: Maybe! We’ve already seen third-party audits of frontier AI labs for pre-deployment security, and we’ve already seen frontier AI labs collaborate on issues of public safety. An ISPA inspection would be a sort of “combo move”, involving more than one AI lab and a third-party supervisor. Thus, an ISPA inspection would be a combination of already-realistic auditing activities, that hasn’t been accomplished before, hence the prize for initiating and supervising it.
Q: Do antitrust laws make this impossible?
A: No. Businesses can collaborate to establish industry standards that benefit the public. Also, the Peer Assistant frontier lab should be acting under the supervision of the Independent Supervisor for the purposes of carrying out the inspection, not on their own initiative. Finally, the Independent Supervisor institution should retain real legal authority to prevent needless exchange of information that would collude against public interests or other companies.
Q: Do AI labs even want this?
A: Yes. To varying degrees, frontier AI labs and their employees are afraid of each other losing control of their AI systems and harming the public, and would therefore like a chance to prevent that.