Before an AI system reaches the public, developers and oversight bodies increasingly rely on structured testing rather than new statute. This piece sets out what that testing involves, where the approach came from, and what it cannot do.
Pre-deployment testing in plain terms
Testing an AI system means running it against a prepared set of inputs and recording how it behaves, then deciding whether that behaviour is acceptable for the use intended. In practice this covers three different activities that are often bundled together. The first is measurement: does the system produce accurate, consistent output on tasks resembling its intended job. The second is adversarial probing, usually called red-teaming: people or automated tools deliberately try to make the system fail, leak information it should protect, or produce content its operator has ruled out. The third is the construction of safeguards — filters, refusal behaviour, rate limits, monitoring, human sign-off on consequential decisions — and then testing whether those safeguards hold when someone attacks them.
None of this is unique to AI. Pharmaceuticals, aircraft components and financial models are all released on the basis of documented testing regimes. What is distinctive about modern AI systems is that their behaviour is learned from data rather than specified line by line, so the only reliable way to find out what a system does is to try it.
Origins of the testing-first argument
Machine learning research has long relied on benchmarks: fixed datasets against which competing systems are scored. That culture produced a habit of measurement, but benchmarks were designed to compare research prototypes, not to certify products for public release.
Two pressures changed that. The first came from within computer security, where the practice of hiring people to attack a system before adversaries do was already established. As general-purpose models were connected to real tools, data stores and customers, security teams began treating them as attack surfaces and applying familiar methods: threat modelling, penetration testing, incident response planning.
The second pressure came from governments. Once general-purpose systems became widely available, legislators faced a problem of pace: a statute takes years to draft and longer to amend, while model capabilities shift between releases. Several jurisdictions responded by building technical evaluation capacity — public bodies whose job is to test systems and publish methods — alongside, and often in advance of, binding rules. The United Kingdom established a state institute for this purpose, and comparable initiatives exist elsewhere. Standards bodies, including the US National Institute of Standards and Technology, produced voluntary risk-management frameworks that describe processes rather than prohibitions.
The argument now heard from parts of the financial and regulatory world runs along these lines. The BBC reports that a senior figure in UK central banking has argued regulation is not the right starting point for AI, and that stringent testing and safeguards are what is needed to contain the risks. The precise scope of that position — which systems, tested by whom, to what standard — is not set out in the material available here.
Current practice
A serious testing programme today has several layers. Internal evaluation happens continuously during development, with results recorded against a fixed set of measures so that regressions are visible between versions. Before a significant release, dedicated teams run structured adversarial exercises: attempts at prompt injection, where hostile instructions are smuggled into data the model reads; attempts to extract training data or system instructions; attempts to elicit assistance with harmful tasks; and tests of whether the system behaves differently for different groups of users.
Beyond the model itself, attention has shifted to the surrounding system. A model with no tools can only produce text. A model that can send email, execute code, query a customer database or move money can cause direct harm, so the testing boundary has widened to include permissions, logging, and what happens when a component fails. Deployers increasingly document these arrangements formally, and third-party auditors have emerged to assess them — an activity sometimes called AI assurance.
External testing is the newest layer. Some developers grant researchers or government institutes access to systems before public release; some run bug-bounty style programmes inviting outsiders to report failures. The depth of this access, and whether findings are published, varies considerably and is largely determined by contract rather than law.
Common misconceptions
The most persistent error is treating testing and regulation as alternatives. They interact: a rule that requires nothing measurable is hard to enforce, and testing with no legal consequence attached is a matter of goodwill. When someone argues for testing “rather than” regulation, the substantive question is usually not whether rules should exist but who sets the standard, who verifies compliance, and what follows a failure.
A second error is assuming a passed test transfers. Evaluation results are specific to the version tested, the prompts used and the configuration in place. Change the system prompt, connect a new tool, fine-tune on fresh data, or deploy in another language, and prior results may not hold.
Third, benchmarks can be gamed, sometimes inadvertently. If material resembling a test set appears in training data, scores rise without capability improving. This is why methodology and contamination checks matter more than headline numbers.
Fourth, red-teaming establishes that a failure is possible, not that the absence of a failure is proof of safety. It is evidence of what attackers found in the time available, nothing more.
Finally, safeguards are frequently confused with capability limits. A filter that blocks a request has not removed the underlying ability; it has placed an obstacle in front of it, and obstacles can be circumvented.
Where to look next
Readers wanting the technical grounding should start with the published risk-management frameworks from national standards bodies, which are written for practitioners and are free to read. The publications of government AI testing institutes set out evaluation methods in more detail and are the clearest available account of what state-run testing actually measures.
For the legal dimension, the European Union’s AI Act is the most developed example of binding, risk-tiered rules and repays direct reading rather than summary. Sector regulators in finance, health and transport have issued their own guidance, which tends to be more concrete about accountability than general AI policy documents.
Anyone following the regulation debate should watch for one detail in particular: whether a proposal specifies who verifies the testing. That single point usually distinguishes competing positions more sharply than any disagreement about risk itself.
Frequently asked questions
What does red-teaming an AI model mean?
Red-teaming means assigning people, or automated systems, to attack an AI system deliberately in order to find failures before real adversaries or ordinary users do. Attempts typically include smuggling hostile instructions into input data, trying to extract confidential system instructions, and seeking assistance with tasks the operator has prohibited. Findings feed back into safeguards. A red-team exercise shows what testers found in the time they had, not that a system is safe.
Is AI testing legally required?
It depends on jurisdiction and use. Some frameworks, such as the European Union’s AI Act, impose obligations that vary with the risk of the application, and sector regulators in areas like finance and healthcare apply existing duties to AI systems used within their remit. Elsewhere, testing rests on voluntary standards and commercial agreements. Whether a general legal testing requirement should exist is precisely what is currently being debated.
What is the difference between an AI safeguard and a capability limit?
A safeguard is an obstacle placed in front of a behaviour: a content filter, a refusal rule, a permission check, a human approval step. A capability limit would mean the system genuinely cannot do the thing. Most deployed protections are safeguards, which is why circumvention is possible and why monitoring matters. Treating a safeguard as though it removed the underlying capability is a common and consequential mistake.
Why do benchmark scores not prove an AI system is safe?
Benchmarks measure performance on a fixed set of tasks under specific conditions. Results are tied to the exact version, configuration and prompts tested, so they may not survive an update, a new tool connection or a different language. Scores can also rise if similar material appears in training data, without real improvement. Methodology and contamination checks tell you more than the number itself.
Who carries out AI testing at present?
Mostly the developers themselves, through internal evaluation and adversarial exercises. Beyond that, some government institutes test systems and publish methods, independent auditors assess deployment arrangements for organisations that commission them, and outside researchers report failures where developers grant access. How much external access is given, and whether results are published, is currently determined largely by private arrangement.
Sources and further reading
- The BBC’s technology reporting, which carried the comments on AI regulation and testing referenced above.
- Published AI risk-management frameworks from national standards institutes, which describe testing and governance processes for practitioners.
- Official material from government AI testing and security institutes, setting out evaluation methods and scope.
- The text of the European Union’s AI Act, as the most developed example of risk-tiered statutory rules for AI systems.
Surfaced from the rss:bbc_tech signal “debate over AI regulation”. AI-assisted draft, editorially reviewed.

