There Is No Saudi AI Law, and That Is Why Your Pilot Dies in Procurement
Executive summary
Every executive committee now has an AI item, and the ones I watch stall do not stall on accuracy. They stall because nobody can answer in writing which data the system sees, where inference happens, who the processor is, and who is accountable when the output is wrong. No Saudi statute governs how you build, procure or supervise a model. Several binding instruments govern the data it sees, the place it is hosted, and what you may reproduce to train it — and the cybersecurity guideline written for AI itself is still a draft. An AI initiative here is a data-governance and third-party-risk project with a model attached.
What actually binds you, and what does not
There is no Saudi AI statute to comply with. That is the correction, and most of what is sold as AI governance in the Kingdom is built against a law that does not exist. The CMS AI regulation scanner, last updated 17 February 2026, records no AI-specific binding law in Saudi Arabia. SDAIA's AI ethics principles, its generative AI guidelines, its AI adoption framework, its deepfakes guidance — all of it is guidance rather than binding law. The draft Global AI Hub Law that went to public consultation in April 2025 remains unenacted, and it was drafted about sovereign data hosting and data embassies. A compliance program written against that statute has no addressee, and you will pay for it anyway.
What did arrive came from a direction most AI governance programs were not watching. The new Copyright Law, issued by Royal Decree M/169 and published in the Official Gazette on 13 February 2026, is in force. It carries an express exception permitting the reproduction of a work to develop artificial intelligence products and algorithms, on conditions: that the work was lawfully published, that the original copy was lawfully obtained, and that the reproduction is limited to what serves the purpose. Its implementing regulations have not been issued, so the third condition has no measurement behind it yet. Read that precisely. The instrument speaks to what you may reproduce in order to train, and to nothing downstream of it.
So the honest form of the claim is narrower than the slogan, and it is the form that survives an argument. No Saudi instrument tells you how to build a model, how to procure one, or how to supervise the decisions it touches. Several binding instruments tell you what data it may see, where that data may be hosted, and what you may lawfully reproduce to train on. The gap is in the middle, and the middle is where your pilot is sitting.
The one you will meet first is the Personal Data Protection Law, and it is in live enforcement. It came into force on 14 September 2023, the transition period ended on 14 September 2024, and in the twelve months to mid-January 2026 SDAIA's violations committees issued 48 decisions confirming breaches of the law. The committees may impose fines of up to five million riyals, doubled for repeat violations. The procedural clock is short: once the Authority notifies you of an alleged violation you have five days to file your response with the committee, fifteen days for the committee to notify its decision, sixty days to appeal to the competent court.
Five days is not enough time to assemble a data inventory you never built.
The one instrument written specifically for AI is still a draft. The National Cybersecurity Authority opened public consultation on AI Cybersecurity Guidelines on 5 July 2026 and closed feedback on 5 August 2026. It addresses entities that use or plan to adopt AI systems in the Kingdom, generative and agentic alike, and it is organized into four domains: cybersecurity governance, cybersecurity defence, cybersecurity resilience, and third-party cybersecurity. The final text is not published. I am not designing controls against a draft, and I would not sign a policy that cites one. What I will use is the four headings, because they are already the four things a regulated technology function owns, and the draft points them at a new class of asset.
The pilot does not die on accuracy
The AI item passes the committee easily. Someone runs a model live against a policy document or a claims note, the accuracy looks respectable, and the room agrees to proceed. Four months later the item is still open. Nothing about the model changed. Legal asked which data the system would see, procurement asked who the counterparty actually is, and neither question had a written answer.
Deloitte Middle East published a survey of senior tax and finance leaders across Saudi Arabia, the UAE, Qatar and Kuwait on 24 February 2026: 18 percent were actively piloting generative AI use cases, 9 percent were scaling them, and 10 percent reported enterprise-wide AI strategies and governance frameworks in place. That is a different function from mine and the sample carries no insurance breakdown, but the shape is the one I see. The distance between piloting and scaling is the distance between a live run and a file an auditor can read.
Six questions decide whether an initiative ships. Which data classes the system sees. Where inference is performed and under whose jurisdiction. Who the processor is, contractually and in fact. What happens to the outputs and how long they are retained. Who is accountable, by name, when the output is wrong. And how you would evidence any of that to a supervisor twelve months from now.
Drop those six onto the four domains and the ownership stops being ambiguous. Which data the system sees, and where inference runs, is defence. Who the processor is and whether the workload can be exited is third party. Who answers for a wrong output is governance. What the process does at four in the afternoon when the endpoint is unavailable during a claims peak is resilience. Treating AI as a fifth domain is how organizations end up with an AI steering committee and no data inventory.
Classify the data before you choose the tool
The order I keep meeting is backwards. A team selects a tool, runs a proof of concept, and then asks the security team to classify the data it touched. Classification is the input that decides which vendors can be shortlisted at all, and it belongs in the room before the first vendor call.
The Personal Data Protection Law's Implementing Regulations, issued on 7 September 2023, require a documented impact assessment where the processing involves sensitive data or the use of a new technology, and a generative model applied to customer records is a new technology under any reading of that phrase. In insurance the sensitive-data question is not marginal: health data sits in the middle of the book, and the Implementing Regulations expect every stage of health data processing to be documented so the person responsible for each stage can be identified. Read that twice before anyone pastes a claim history into a prompt window. The cybersecurity side arrives at the same place from a different direction. The National Cybersecurity Authority's controls for non-critical-infrastructure private sector entities, issued as NCNICC-1:2025, set expectations for the use of cloud and hosting services that begin with data classification, separation of the entity's environment from those of other tenants, and return of data in a usable format on termination of the service. Those sit in the third-party and cloud component of that framework, not in defence, and that is the correct place for them: your model endpoint is somebody else's hosted service.
So classify at field level. "The policy administration system" is not a classification. "National ID, date of birth, diagnosis code, claim amount, adjuster note" is a classification, and it is the version that tells you whether a given endpoint is eligible at all — before you have spent a quarter finding out that it is not.
An inference endpoint is a third party with a location
An inference endpoint is a third party with a location attached. When a model answers a prompt containing customer data, someone else's infrastructure processed that data, in a place, under a legal regime, on a contract that either says so or does not.
Article 29 of the Personal Data Protection Law governs transfer outside the Kingdom, and the Regulation on Personal Data Transfer Outside the Kingdom, published on 1 September 2024, sets the conditions. Article 7 of that regulation requires a risk assessment before personal data is transferred or disclosed outside the Kingdom. SDAIA published a Risk Assessment Guideline for that purpose on 25 February 2025 — non-binding, four phases, the fourth of which asks about implications for the Kingdom's vital interests rather than the individual's rights alone. The adequacy list the regulation contemplates has not been published, which leaves three safeguard routes in practice: standard contractual clauses, binding common rules, or a certificate of accreditation. None of the three is a country determination someone can point at in a meeting.
Financial-sector practice adds a second gate. The Saudi Central Bank's cyber security framework, in its cloud computing section dated 24 May 2017, states that a member organization should obtain SAMA approval before using cloud services or signing the contract, and that in principle only cloud services located in Saudi Arabia should be used, with explicit approval required otherwise. Supervision of insurance has since moved to the Insurance Authority, which began operations on 23 November 2023 as the sole regulator of the sector, and the instruments SAMA issued for insurers stay in force until the Authority replaces them. A hosted model endpoint is a cloud service in every sense the drafters meant.
Then the clause procurement finds late. On one firm's reading of the Implementing Regulations — the only plain statement of it I found — the processor must confirm that it is not subject to regulations in other countries that affect its ability to comply with the Personal Data Protection Law. Put that representation in front of a global model provider and you will learn how long your pilot is actually going to take. Raise it in month one. In month five it is the reason the item is still open.
Accountability is an artifact, not an attitude
"The model decided" describes a control failure, not a decision. A supervisor asking why a customer was declined is not asking about attention weights. They are asking who owned the outcome, what that person was shown before it happened, and what they could have done about it. Those are org-chart questions whose answers are documents.
SDAIA's Principles and Controls of AI Ethics, issued on 1 September 2023, sort systems into four risk tiers — little or no risk, limited, high, and unacceptable — and name accountability, human oversight, and explainability among their principles. They are not binding, and non-binding gets read as optional. Enforcement arrives through the binding instruments underneath, and through whichever supervisor concludes you placed yourself in the wrong tier. In the rooms I have sat in, this gets heard as a request for explainability and answered by buying a tool that produces feature attributions. The deliverable is a named person, a documented path by which they can overturn the output, and a record of every occasion on which they did.
A pilot has to prove three things a demonstration never does: that the organization can run it, evidence it, and stop it. Write the exit artifacts down before the pilot starts: a field-level data inventory with a classification against each field; the impact assessment where sensitive data or new technology is in scope; the transfer risk assessment where inference leaves the Kingdom; a processing agreement carrying the processor representation; a retention and deletion position for prompts and outputs, not only for source records; an override log; an error register with a named owner per class of error; a documented answer for what the process does when the endpoint is unavailable, tested rather than asserted; and an exit clause that returns your data in a usable format.
If the pilot ends and those files do not exist, a pilot did not happen — a demonstration happened.
This is procurement work with a model in the middle of it, and the instrument is one I have built before. At Tbar Holding I wrote a service catalog of more than a hundred services from nothing, every service carrying a service level, in a company that had no service level agreements at all beforehand. Then I scored every vendor against the level it had signed, from the ticket data rather than from the vendor's own report, and renegotiated or replaced the ones that missed. That took forty percent off contracted vendor spend, and it worked because the measurement existed before the conversation did. Define what wrong means for this system, measurably, before the vendor tells you their accuracy figure — then run it on a real queue at real volume with a human reading every output. Two weeks on curated inputs will tell you the model is good. Two months on the live queue will tell you what happens the day an upstream format changes and nobody mentions it.
Where I would not put a model today
I have not deployed a generative model into a regulated decision path, and I will not claim otherwise. So read the rest as positions I would defend in a committee. I would not put a model on a decision path that declines or reduces a customer's entitlement without a human who reviews it, can overturn it, and is recorded as having looked. Accountability for that outcome has to attach to a person who can be asked about it afterwards, and a model cannot be asked.
I would not send raw customer records to an endpoint whose processing region I cannot name in a contract, and I would not accept "our global infrastructure" as an answer to that question. I would not give an autonomous agent write access to a production system on the strength of a pilot; the failure mode there is a wrong action nobody approved and nobody can reverse.
Some of this is unsettled and I would rather say so than paper over it. The National Cybersecurity Authority's consultation closed on 5 August 2026 and no final text has been published. SDAIA has not published the adequacy list the transfer regulation contemplates, and that one publication would change several of the calculations above. And I do not know how the Insurance Authority will treat model-assisted decisions in the rulebook it is still assembling — that is my own supervisor I am saying that about, and it is the answer I would most like to have. Where the rule is unsettled I document the decision and the reasoning, so that whichever way it settles, the file explains what we knew at the time we chose.
Where I would put a model today is internal, reversible, low-classification work, read by a human before anyone outside the team sees it. Summarizing internal documentation. Drafting first-pass test cases. Sorting an internal ticket queue. None of that impresses a committee, and all of it produces the override log, the error register and the retention position you will need on the day a customer-facing case finally becomes answerable in writing.
Practical checklist
- Write down which instruments actually bind you before writing an AI policy — today that is the data protection law, the transfer regulation, the hosting controls and the copyright exception, and no AI statute at all.
- Classify the data at field level before the first vendor call; classification decides the shortlist, not the reverse.
- Name the processor and the processing region in the contract, and treat an inference endpoint as a third party with a location attached.
- Complete the transfer risk assessment before any personal data leaves the Kingdom, and keep the written output on file.
- Assign a named human owner to every automated output that affects a customer, with a documented override path and an override log.
- List the pilot's exit artifacts up front — inventory, assessments, processing agreement, override log, exit clause — and refuse to close the pilot until each one exists.
Common mistakes
- Building a compliance program against a Saudi AI statute that does not exist, while the instruments that do bind you — data protection, transfer, hosting, and now copyright — go unaddressed.
- Reading non-binding guidance as optional, when enforcement arrives through the binding instruments underneath it.
- Choosing the tool first, then asking the security team to classify the data afterwards.
- Reading the copyright law's AI exception as permission to build, when it speaks only to what you may reproduce in order to train.
- Running the pilot on curated inputs and presenting the result as production evidence.
How Hamad approaches this
I treat an AI proposal the way I treat any third-party arrangement that touches regulated data, because that is what it is. The questions I put to a model vendor are the questions I have put to every vendor I have contracted: what do you see, where does it go, what do you commit to, and how do I check. The last one is the one that gets negotiated out of the draft if you let it.
The board and the executive committee I report into ask the same second question about an AI item, and it is never about accuracy. It is who signed. The item is approved on enthusiasm and killed on evidence, so I move the evidence to the front of the paper rather than the back. At Rasan I lead the integration build and the business continuity work, which is the resilience domain under its own name, and the first artifact I want from any AI proposal is the field-level data inventory and a straight answer on the processor.
What I have done many times is take a technology proposal through legal, through procurement, and into a committee that has to sign for it. Nothing in that path is new. The model is the only new object in the room, and everything that stalls around it is old.
Continue the conversation
The executive profile, the operating philosophy, and direct channels for advisory or transformation discussions.
Related articles
Cybersecurity Governance: Why Tools Are Not Enough Without Operating Discipline
Policies, layered controls, SOC alignment, incident response, and improvement cycles — building a program that still works when the next vendor promise fades.
ReadTech Relationships and Onboarding: Turning Partner Discussions into Controlled Execution
Readiness assessments, integration governance, dependency maps, and exit criteria — treating partner onboarding as a delivery program, not an open-ended dialogue.
ReadTechnology Operations in Insurance and FinTech: Where Reliability Becomes Trust
Regulatory touchpoints, integration wiring, onboarding friction, and the trust chain from policy to claim — why reliability in insurance and FinTech is a product feature, not plumbing trivia.
Read