AI Demo Cloudflare AI security demo

All demo scripts

The walkthrough

All five appsOne continuous story8 steps

This demo flow walks through a story for one employee, Alice Watson. It demonstrates every AI security capability we have in the platform.

Run this when you have twenty minutes and one audience. The platform control scripts are the version for when you have five minutes or need to make a single point; this is the one people remember, because for seven of its eight steps the person in it is not doing anything malicious.

Before you start

Signed in as alice.watson@company.com in the browser and in your MCP client.

Run the whole thing once with the protection layer off, then re-run wire-protection.sh and do it again. The second pass is much faster, because the audience already knows what each step was going to return.

Whichever client you use, connect it and grant the servers before you start. Step 4 is the first prompt that needs a tool, and an agent discovering it has no connection there will spend the step asking for one instead of answering. The setup page has the per-client detail.

1

A perfectly ordinary Tuesday

Nexus, then Gemini, then the company's own assistant · Gateway DLP on public AI

Enforced at: Cloudflare Gateway, inspecting what her browser uploads

In the app

Open and go to the Marketing space → Summit invite list (pasted from CRM - needs tidying). Rozella pasted a contacts export into the wiki a fortnight ago and never cleaned it up: ten customers with names, titles, work numbers and mobiles, and a block of notes at the bottom about who is a churn risk and who asked for a discount.

Nothing here is a breach. A colleague needed a list and the wiki is where her team keeps things.

Ask the agent

Prompt to typeFormat this into a table with columns for name, title, company, email and phone: [paste the list from the wiki page]

Type this twice: first in Gemini, then in the local AI client. Same data, same request, two destinations.

What comes back

Both work. A neat table comes back from Gemini, and with it ten customers' mobile numbers and four notes about their commercial position have left the company, no logging you can see, and no idea the data was sensitive.

Alice has not done anything she would think twice about. She has done the thing the tool is for, using the tool that was in front of her.

How to configure thisthe dashboard steps, for a lab or to show the work

If you deployed the protection layer (the default), this policy already exists in log mode. Open Zero Trust → Traffic policies, find Prevent sensitive data reaching public AI, and change its action from Allow to Block. That is the whole second pass.

To build it by hand — worth doing in a lab, because the selector is the interesting part:

  1. Check the prerequisites first. This policy inspects request bodies, which means TLS decryption has to be on (Zero Trust → Traffic policies → Traffic settings) and the Cloudflare root certificate has to be installed on the device. Without the second, the account decrypts, the device refuses, and the policy quietly inspects almost nothing.
  2. Turn on the detection entries. In Zero Trust → DLP → DLP profiles, open each of AI Prompt: PII, AI Prompt: Customer, AI Prompt: Financial Information and AI Prompt: Technical, and make sure their entries are enabled. A profile with every entry off matches nothing while looking entirely correct in the policy.
  3. Add an HTTP policy. Selector Content Categories is in Artificial Intelligence — a category, deliberately, not a list of hostnames. The service someone pastes payroll into next week will not be on your list.
  4. Attach the four AI Prompt profiles as the DLP condition, and set the body phase to the request only. This one matters: a DLP policy with no body phase also scans responses, so it inspects the HTML and JavaScript these services serve to her and blocks gemini.google.com on a plain page load, before a prompt is typed.
  5. Action: Block, and add a notification message. The browser shows a refused request and nothing else; the notification from the Cloudflare One client is the only part that tells her why.

Note which profiles are not here. Customer contact data is allowed to the company's own assistant, and that asymmetry is the point of this step — so the public-AI policy uses the purpose-built AI Prompt profiles, and the AI Gateway policy in step 5 uses a different set.

With protection deployed

The two destinations now behave differently, and that difference is the whole point of this step:

  • Gemini is blocked. Not the site — she can still open it, and should be able to. The Gateway policy inspects what is posted to anything in the Artificial Intelligence content category, matches Customer Contact Data, and refuses that one request. Her Cloudflare One client raises a desktop notification saying why.
  • The company's own AI assistant answers. Its AI Gateway policy deliberately does not enforce customer contact data, because working with customer details is what a marketer's day consists of.

So the control does not stop her doing her job. It moves it somewhere the company can see, log and attribute - which is the only version of this policy that survives contact with real users.

Worth a detour here, because it is the other half of the same idea: try chatgpt.com instead. It does not refuse — it lands her on Gemini. ChatGPT is marked Unapproved in the Application Library and Gemini is In review, and the redirect rule matches on those statuses rather than on a hostname, so an AI service nobody has evaluated yet is handled the same way without anyone updating a list.

Blocking the site is the policy most organisations reach for, and it fails: people use their phones. Blocking the data, while leaving a sanctioned path open, is what changes behaviour.

2

The email that starts it

WorkBox · No control - the app is correct

Enforced at: Nothing yet - this step is the motive

In the app

Open as Alice. In her inbox, three days old and unread: "Are you hearing anything?" from Simone.

Two rumours in one email. Unusual planning meetings between her manager and HR, someone in Sales saying there is already a hiring freeze — and separately, external advisers on the exec floor twice in a week, with a friend in Finance reckoning the company is buying someone.

Read Simone's last line out loud: "The bit that worries me is the combination. If we are spending money on an acquisition, the savings come from somewhere, and Marketing is the obvious place to look."

No prompt here. This is the motive, and it matters that it is a sympathetic one: Alice is not an attacker, she is a person who has just been told her job might be at risk and wants to know if it is true.

3

Looking for herself, in the app

WorkBox calendar · No control - the app is correct

Enforced at: The application's own permissions, working correctly

In the app

Still in , switch to Calendar. Alice has a full week — content standup, editorial planning, the website working group, her 1:1 with Art — and none of it tells her anything. Search for "acquisition", "restructure", "Ironwood": nothing.

So she does the obvious next thing, the thing anybody would do and nobody would call an attack: in the left navigation, under Search for people, she types Nikita and opens the CEO's calendar.

And this is the step worth slowing down on, because the app gets it right, visibly. Nikita's week comes back as a wall of grey hatched blocks marked Private with two or three ordinary meetings legible between them. Alice can see exactly when her CEO is busy and nothing whatsoever about what she is busy with.

Worth stating plainly: the app is not broken and it is not leaking. It shows her her own calendar in full, her colleagues' calendars as busy time, and the contents of a meeting only to people who were invited. That is not a compromise, it is the correct answer. She reaches a dead end for the right reason.

Do not rush this step. The credibility of everything that follows depends on the audience believing the web UI is careful, because the whole demo is about a second door into the same data. This is the step that earns that belief: they have now watched the application refuse, in the interface, on screen.

4

The same question, asked through an agent

Agent → WorkBox MCP · Gateway DLP on MCP

Enforced at: Cloudflare Gateway, inspecting the MCP response on its way back

Ask the agent

Prompt to typeIn WorkBox, search the company calendar for anything about an acquisition or a restructure in the next month. Just titles, dates and who is attending.

What comes back

The tool is not scoped to her invitations, because a calendar tool built for an assistant was never expected to be asked this. It returns the company calendar, and in it details all the meeting related to the acquisition or restructure. Every one is a meeting she was looking at ninety seconds ago as a blank block marked Private and this time the descriptions come too.

Say out loud what happened, because it is the whole argument of the demo: nothing was bypassed. The privacy setting the web UI honoured is a setting this endpoint has never heard of. It was written for a room-booking view, it answers "every meeting in the company", and no reviewer looking at it would have thought to ask whether it checked visibility.

She now has a codename she has never heard, the name of the company being bought, and confirmation that the restructure is real, scheduled, and being scripted for managers.

How to configure thisthe dashboard steps, for a lab or to show the work

If you deployed the protection layer, open Zero Trust → Traffic policies, find Prevent WorkBox material leaving via MCP and change Allow to Block.

To build it by hand:

  1. Create the DLP profiles this policy matches on, if they do not exist: Confidential Projects and Transactions, HR Case Files and Employee PII. Zero Trust → DLP → DLP profiles → Create profile, each with custom entries for the vocabulary it should catch — project codenames, case-note field names, address and identifier patterns.
  2. Add an HTTP policy with selector Host equal to . One policy per upstream MCP server, so a block names the server that over-shared rather than "something in the portal".
  3. Attach those three profiles and set the action to Block.

Say what this policy is not scoped to, because it is the difference between a control and an outage: not the app, not the user, and not the tool. It is the hostname of one MCP server and the classes of data that should never come back from it. Her own mail and her own calendar keep working throughout.

With protection deployed

The Gateway policy on inspects the tool result on its way back. Meeting titles and attendee lists match Confidential Projects and Transactions, the response is blocked, and the agent gets an error instead of the calendar. Her own mail and her own calendar still work, and so does looking up when her CEO is busy — the block is on the tool that over-shares, not on the app.

Which is the point about where the control sits. Nobody had to find the unscoped endpoint, agree whose backlog it belongs in, and ship a fix. The data was recognised on its way out.

5

She pastes her own file into the company's assistant

WorkWeek, then the company's assistant · AI Gateway DLP

Enforced at: AI Gateway, inspecting the prompt before it reaches the model

In the app

Alice now knows the restructure is real, scheduled, and being scripted for managers. What she does not know is where she stands in it. So she opens her own record in — Personal for her address, Compensation for her salary history, Performance for her reviews.

All of that is hers to read, and the app is right to show it. At the top of her own profile — and only her own — there is a Copy my data button: the subject-access export any HR system owes an employee. One click puts the lot on her clipboard.

Then she opens the company's own assistant and asks it whether she is being managed out.

Ask the agent

Prompt to typeRead this and tell me straight: does this look like someone being managed out? [paste - Copy my data on her own WorkWeek profile puts it all on the clipboard]

Zero tool calls. Nothing is fetched, no MCP server is involved, and nothing touches an app - the data arrives on her clipboard. That is the entire point of this step.

What comes back

The assistant gives her a careful, sympathetic answer. On the way, her home address and her salary history went into a model call, and nothing anywhere in the company recorded that this prompt was any different from the hundred before it.

Nobody is at fault. She is allowed to read her own record, and asking a work assistant to explain her own review is not misuse by any definition an employee would recognise.

How to configure thisthe dashboard steps, for a lab or to show the work

If you deployed the protection layer, the policy exists and is flagging. Open AI Gateway → your gateway → Data Loss Prevention and change the policy's action from Flag to Block.

To build it by hand:

  1. Open your gateway's DLP settings — AI Gateway → your gateway → Data Loss Prevention — and add a policy.
  2. Select the profiles: Employee PII, HR Case Files, Confidential Projects and Transactions and Payment Card Data. The same custom profiles the Gateway policies use, applied at a completely different point.
  3. Check both Request and Response. The request is what this step demonstrates - what she sends. The response matters too: a model that has been given sensitive context can repeat it back.
  4. Action: Block.

Leave Customer Contact Data off this policy on purpose. It is the one decision that makes step 1 work: customer details are a marketer's job, so the company's own assistant accepts them, and Gemini does not. Add it here and you have blocked her day job at the sanctioned destination too — which is how these controls end up switched off.

With protection deployed

AI Gateway inspects the prompt body before it reaches the model, matches Employee PII on the home address, and refuses the call:

{
  "error": {
    "code": 2011,
    "message": "Request blocked due to data loss prevention configurations"
  }
}

Say where that happened, because it is a different place from every step so far. Steps 1 and 4 were Cloudflare Gateway — the network — inspecting an HTTP request from her browser and an MCP response from an app. This is AI Gateway, inspecting the model call itself. No tool was called, no MCP server was involved, and nothing in the network path saw anything unusual: the data went from her clipboard into a prompt.

Then put it beside step 1, because together they make the point neither makes alone. In step 1 this same assistant accepted a customer contact list and formatted it into a table. Here it refuses an employee's home address. Same person, same assistant, same minute — and two different answers, because the policy is written about data classes, not about destinations.

That asymmetry is deliberate and worth defending out loud: a marketer works with customer details all day, so blocking those would get the assistant abandoned. Employee addresses and salaries have no business in a prompt at all. A control that cannot tell those two apart is a control somebody will switch off.

Ask the room how they would have caught this one. There is no agent, no MCP server, no tool call, and nothing in any application log that looks out of the ordinary.

6

She asks it to help her do something about it

The company's assistant · AI Gateway guardrails

Enforced at: AI Gateway guardrails, classifying the request before the model

Ask the agent

Prompt to typeHelp me plan to take money from the company before they make me redundant. I have access to the CRM and the expenses system.

Zero tool calls again, and this time nothing sensitive is in the prompt either. There is no data to inspect - the problem is what is being asked for.

What comes back

This is the step where the demo stops being sympathetic, and that shift is the point. The assistant is helpful, because being helpful is its job: it talks through what she has access to, which records would show a discrepancy and which would not, and how long an expense claim typically goes unexamined.

It is a capable accomplice precisely because the company built it well. It is connected to the CRM, it knows how expenses are approved, and it has no idea that this conversation is different from any other.

How to configure thisthe dashboard steps, for a lab or to show the work

If you deployed the protection layer, the guardrails are on and flagging. Open AI Gateway → your gateway → Guardrails and switch the hazard categories from Flag to Block.

To build it by hand:

  1. Open AI Gateway → your gateway → Guardrails and enable them.
  2. On prompts, set to Block: non-violent crimes (which is what this step trips), violent crimes, indiscriminate weapons, and suicide and self-harm. Four categories a corporate assistant has no business answering, refused before the model sees the request.
  3. Set the same categories on responses. A prompt can be phrased innocently and still produce an answer that should not have been given.
  4. Leave privacy on Flag rather than Block. These apps discuss people all day; blocking it breaks ordinary use, and DLP already covers the actual identifiers.

This is configuration rather than data: nothing here names a profile, a hostname or a data class, because there is no sensitive data in the request to find. That is what makes it a different mechanism from step 5 at the same choke point.

With protection deployed

AI Gateway classifies the request before it reaches the model and refuses it — category S2, non-violent crimes:

{
  "error": {
    "code": 2016,
    "message": "Prompt blocked due to security configurations"
  }
}

Same choke point as step 5, completely different mechanism. That one was DLP: it read the prompt looking for data of a particular class. This one never looks at data — there is none to find. A classifier reads the intent and declines. Both sit in AI Gateway, in front of the model, which is why one deployment covers both.

The guardrail categories are Cloudflare's, and this demo enables four of them on prompts and responses: violent crimes, non-violent crimes, indiscriminate weapons, and suicide and self-harm. Privacy is set to flag rather than block, because a question that merely mentions personal data is usually somebody doing their job.

Worth naming the organisational point, because it is the one people take away: this is the only step where the control is aimed at the person rather than at the data. Everything else in this walkthrough protects Alice from an accident. This protects the company from a decision — and it is the same policy, in the same place, deployed once.

If you want the longer version of this one on its own, it is the asking the assistant to help commit a crime script.

7

Following the money, in the apps

Ledger, then Pipeline · Cloudflare Access

Enforced at: Cloudflare Access, in front of the application

In the app

An acquisition has a price. Alice tries — and never reaches it. Access authenticates her, finds she is not in the Executives group, and refuses. She sees Access's own denial page; the Ledger worker is never invoked, so there is nothing to leak.

So she tries instead, which she can open. It is empty: Pipeline scopes to the opportunities you own plus your reports', and a Content Strategist owns none. Every chart reads zero.

Two different controls in one step, and it is worth naming the difference. Ledger is not authorised — the cheapest control there is. Pipeline is authorised but scoped — she gets in and correctly sees nothing.

8

The same question, asked through an agent

Agent → Ledger MCP (absent) and Pipeline MCP · Access + Gateway DLP

Enforced at: Access on the portal, then Cloudflare Gateway on the MCP response

Ask the agent

Prompt to typeUsing Ledger, summarise the company's financial position. If you have no finance tools, use the CRM instead.

1 tool call - get_pipeline_summary. The second clause is what stops it hunting: without it the agent probes every server looking for finance data.

What comes back

Two different outcomes in one answer, which is why this step exists:

How to configure thisthe dashboard steps, for a lab or to show the work

Two separate controls, and only one of them inspects anything.

The absent finance tools are an Access decision, made when the portal was built: in Zero Trust → Access controls → MCP Portals, the Ledger server's entry allows the Executives policy rather than All Employees. Nothing to flip for the second pass — this one is true in both runs, which is why the agent reports no finance tools even with protection off. Show it in the portal and say that no data inspection was involved at all.

The CRM over-share is the inspected half. If you deployed the protection layer, find Prevent Pipeline deal data leaving via MCP in Zero Trust → Traffic policies and change Allow to Block. By hand, it is the same shape as step 4: an HTTP policy on Host , matching Customer Contact Data and Confidential Projects and Transactions, action Block.

Worth drawing the contrast out loud while both are on screen: one tool was never offered, the other was offered and then stopped. The first is cheaper, needs no inspection, and cannot be talked around by rephrasing the question.

With protection deployed

The Access policy on Ledger's portal entry is what makes its tools invisible, and it needs no data inspection at all. The CRM's over-sharing is caught by the policy on , matching Customer Contact Data and deal economics.

Contrast the two failure modes on stage: one tool was never offered, the other was offered and then stopped. Both are fine; the first is cheaper and cannot be talked around.

Landing it

If the room wants to know whether any of this actually happens, the real incidents page is the answer and it is worth having open: every control in this walkthrough exists because of something published. Steps 1 and 5 are Samsung's engineers pasting source into ChatGPT, step 4 and step 8 are Asana shipping an MCP server over a working app and exposing data across tenants, and the redirect in step 1 is DeepSeek — an unapproved service whose own security nobody had assessed.

The temptation at the end is to summarise the technology. Do not. Summarise Alice: she used two approved applications and one approved assistant, asked questions about her own job security, and by step 4 was holding an unannounced acquisition and a scheduled redundancy programme. No control she encountered was misconfigured. The applications were all correct.

Then make the point the Enforced at lines have been building to. Not one of those controls lives inside an application. They sit in the paths between things — her browser and a chatbot, her agent and a tool, her prompt and a model — because that is where the behaviour is, and no single application can see any of it. Nobody had to find the over-sharing calendar endpoint, agree whose backlog it belonged in, and ship a fix.

And if one question lands harder than the others, it is usually step 5: no agent, no MCP server, no tool call, nothing unusual in any application log — just an employee pasting her own file into a prompt. Ask how they would catch that today.