Why regulators need structured data, explainability, and governance before ML can deliver operational value.
Regulators everywhere are under pressure. They are dealing with more cases, higher public expectations, and flat budgets, while still facing regular scrutiny from ministers and Parliament. Machine learning could help across many areas: improving risk models, supporting investigations, spotting possible fraud earlier, finding unusual patterns in registration data, and helping teams focus on the most important needs.
But there is a deeper reason this is difficult, and it goes beyond capacity or budget. Regulatory processes are built on fairness and determinism. When a regulator makes a decision, whether to investigate, to escalate, or to close a case, that decision follows a defined process, can be traced step by step, and can be defended on its merits. The entire operating model is designed around consistency, transparency, and accountability. Machine learning, by its nature, is probabilistic rather than deterministic. It identifies patterns and assigns likelihoods rather than following fixed rules. That represents a fundamental cultural shift for organisations whose legitimacy depends on being able to explain precisely why a decision was made. The fear is not irrational: adopting ML means accepting a degree of uncertainty in processes that have historically been built to eliminate it.
The question for most regulators is not whether ML has potential, but whether it can go through rigorous scrutiny. That is, whether it can be used in a way that survives an audit, a tribunal, a Freedom of Information request, or a bad headline.
The instinct to start with experimentation and worry about data, governance, and assurance later is understandable. In a regulatory setting, however, that approach is expensive. Moving first and governing later does not create speed. It creates rework, usually under pressure, usually visible, and usually at a moment the organisation can least afford it.
ML projects in regulation rarely fail because the model is wrong. They fail because the data underneath wasn’t understood, consistent, or didn’t answer the question being asked.
For example, if a model that is used to identify subjects for investigation suggests deploying resources incorrectly, this isn’t a AI or modelling problem, it most likely is a data problem. The same applies to fraud detection sitting on top of inconsistent case records, or prioritisation models drawing on reference data nobody quite trusts.
The UK Government’s AI Playbook treats data quality, lineage, and accessibility as prerequisites, and the bar is higher again for regulators. A model that ranks a casework queue is, in effect, making a regulatory judgement at speed. The data feeding it has to be auditable end to end, with clear structures, one agreed source of truth for the entities being regulated, reference data that’s actively maintained rather than historically preserved, and metadata treated as part of how the organisation operates rather than a documentation exercise.
Skip this and the data layer gets rebuilt mid-deployment, under pressure, at greater cost.
The pattern we see most often isn’t too little ambition. It’s too much, spread too thin.
The highest-value uses of ML in regulation are narrower than the wider AI conversation suggests. For example: anomaly detection to flag returns or filings for human review, case triage to surface the most serious complaints first, and fraud signals in structured data that human investigators can act on. Each is a well-bounded problem where small, defensible improvements compound quickly, and each can be explained in a sentence to someone who has never read a model card.
What regulatory decision is this model supporting?
What’s the baseline of the existing process and how will an ML solution provide value?
What does good look like?
All three need to be defined before any real work begins. Regulators will need to walk away from use cases that can’t answer them.
Explainability in the private sector is a “nice to have”. For regulators, it’s a necessity. The Algorithmic Transparency Recording Standard (ATRS) requires central government departments to publish information about the algorithmic tools they use. The ICO expects anyone affected by a significant automated decision to receive a meaningful explanation of it. In a regulatory setting, a risk score that cannot be explained and defended under scrutiny is not deployable, however impressive its aggregate performance may be.
Sometimes accuracy is the price of explainability. That means choosing a more interpretable model over a marginally more accurate one. For regulators, where decisions affect organisations, funding, individuals, and public trust, that trade-off is almost always worth it.
Explainability is not only a governance requirement. It is what enables operational teams to use a model with confidence and step away from it when judgement demands.
Governance done well isn’t a committee at the end of a project. It’s what lets ML run safely at scale: model risk management, drift and bias monitoring in production, human-in-the-loop checks proportionate to the consequence of the decision, and clear ownership across data, technology, risk, policy, and the business.
In a regulator, every model will eventually be challenged. It may be by leadership, policy colleagues, internal audit, an FOI request, or the need to explain why one case was escalated ahead of another. Ungoverned ML is fragile because it cannot answer those questions with confidence. Governed ML is resilient: it has a documented purpose, an accountable owner, known limitations, and clear points for human intervention.
By the time a model reaches formal sign-off, the important governance decisions have already been taken (or failed to be taken). The teams that succeed here co-own the outcome from the outset, rather than throwing artefacts from one business function to another.
Large transformation programmes have their place. Machine learning capability usually isn’t built well through them. It’s built by small teams that ship to production, learn from real users – caseworkers, investigators, analysts – and iterate.
That means collaborating for depth rather than scale: ensuring the data scientist and the regulatory expert work side by side rather than through layers of governance, and keeping delivery cycles short enough for something usable to land within weeks.
It also means a willingness to stop work that isn’t working. Most ML use cases look different at the end compared to where they started, and the teams that reach production are the ones willing to redesign or walk away rather than push a struggling use case over the line.
The argument running through this paper is straightforward: regulators need structured data, explainability, and governance in place before ML can deliver operational value. Without trusted data foundations, models cannot be audited. Without explainability, decisions cannot be defended. Without governance built into delivery, ML remains fragile under the scrutiny that regulatory life guarantees.
These five priorities are less a technology agenda than an operating model for using ML in a regulatory environment. Where trusted data, carefully chosen problems, explainable models, embedded governance, and close-knit delivery teams are in place, ML can improve speed, consistency, and targeting without weakening accountability. Where they are absent, progress is slower, costlier, and harder to sustain.
Regulators that invest in these foundations now will be best placed to adopt ML safely and at pace. Those that do not will face slower delivery, higher risk, and greater scrutiny.
Our approach to AI and ML reflects that context. We focus on what needs to be true before a model reaches production: trusted data, clear lineage, active governance, and explainability that holds up when it matters. It is less about building clever models and more about making sure the data and the decision-making framework around them are fit for purpose.