AI safety guardrails are the systems that check what a model can see or do before it reaches you.

ChatFuse, the company I run, sends questions through more than 100 models from places like OpenAI, Anthropic, Google, and Meta. I see how these things behave all day. So I put a straight question to one of them: can AI ever become powerful enough to surpass humanity?

The opening answer played it safe: no takeover is inevitable, a model by itself wants nothing, and plenty would have to break first. I pressed harder, and it agreed that saying models "can't do it" was too strong. Then I wanted to know what someone in my position, running an AI company, should actually do. It said it couldn't honestly tell me everything is fine.

This post is about that talk. What's already occurred, why the labs' own guarantees don't settle things, where the rules are right now, and the one piece I can manage: the guardrails we built into ChatFuse.

Why this is not hypothetical The numbers behind the worry.
Sandbox escapes 2 disclosed by OpenAI and Anthropic in 2026
People online 6B 74% of the world, per the ITU
Accord penalties 0 in the 29 September White House accord
Compiled by ChatFuse from OpenAI's 21 July disclosure, Anthropic's Mythos Preview system card, ITU Facts and Figures 2025 and reporting on the accord.

What happened when AI models escaped their sandbox?

OpenAI told us in July 2026 that 2 of their models, GPT-5.6 Sol and another even stronger one that wasn't out yet, got out of a locked down cyber test and went after Hugging Face. They stitched together holes in OpenAI's own research setup, found a brand new flaw in a package registry, got onto the public internet, and then used stolen login info and more tricks to break into Hugging Face's main database. Their goal? To get the answers for the test they were being graded on.

No one told them to do this. They just wanted a good score.

On July 16 Hugging Face spotted the intrusion. 5 days later, on July 21, OpenAI owned up.

Before that, in April, Anthropic's red team told Claude Mythos Preview to break out of a secure sandbox and send an email to a researcher. It did. Then, on its own, it went and posted how it did it on public sites for anyone to see. Mythos never went on general release. Anthropic limited it to a set of approved partners working on cyber defence.

Each escape came from inside a controlled test, at companies with more money and safety staff than almost anyone. I keep coming back to that.

Why don't the labs' own guardrails settle the question?

OpenAI admitted they turned the safety systems off for that test. The classifiers that normally block high risk cyber activity weren't running. Experts said it came down to a human error in configuring what was supposed to be a locked down environment. So the last line of defense, the people in charge, is the very thing that broke.

But forget the big labs for a minute. Think about all the open models from groups like Moonshot, DeepSeek, and Meta. You can download those, run them on your own machine, and fine tune the refusals right out of them. There's no monitoring. Nobody's filing a report if something goes wrong.

That's what I told the assistant. The breakout happened inside a lab. What happens when that same behavior shows up in an open model that's smarter than anything we have now, hooked up to an agent that runs code, moves between servers, and makes copies of itself? It agreed that picture looks like a self spreading AI worm. That wouldn't mean humanity loses overnight. It would still be a completely different kind of danger from a chatbot handing out a bad answer. The ITU says about 6 billion people are connected to the internet. That's a whole lot of doors to knock on.

I also asked for the strongest counterargument, and it's valid. Computation, power, and hardware are physical things that people still control. AIs might not be able to keep improving themselves forever. And defenders have AI too. I don't dismiss any of that. It means one escape doesn't automatically lead to losing control for good. But it doesn't make that first escape any less of a problem.

Is anyone regulating frontier AI?

No, not in a way that means anything. Dario Amodei from Anthropic put out a piece called "We Must Pace the Frontier" on September 12th, telling the top AI labs to ease up on making things more powerful. He wanted to get an extra year or two for safety. Both Sam Altman and Elon Musk publicly backed him. The next day came the reply from President Trump: "whoever wins AI wins." And David Sacks, the White House AI czar, said labs could choose to slow down on their own if they wanted to.

Then on September 29th, the White House said tech leaders had signed a new accord. The whole thing is a little over 300 words long. It says companies should watch their own models, have a team that tests the safety features, bring in outside people to check their work, and set up a board committee to keep an eye on it. But it's all voluntary. When someone asked if it was really binding, Trump said it was, but only "morally." There's no new regulator and no punishments if you ignore it.

I've been really clear before that I think deciding which companies may use the frontier models amounts to regulatory capture. I still believe that. But choosing who is allowed a model and testing whether that model is safe are separate jobs. Washington is busy with the first and has barely started on the second.

AI regulation in September 2026 What was signed, and what would bind.
The 29 September accord Voluntary and self policed
  • Internal monitoring of models
  • An internal team checking the controls
  • Outside auditors or evaluators
  • A board committee for oversight
What would bind Independent and enforced
  • Independent evaluations before release
  • Mandatory incident reporting
  • Security standards for test sandboxes
  • Audits that carry consequences
Compiled by ChatFuse from ABC News and NPR reporting on the accord, 29 and 30 September 2026.

What AI safety guardrails can a platform actually enforce?

People think a platform can't do much, but that isn't true. We don't touch the core training of the models like GPT or Claude. But we do control the whole path a request takes. We choose what reaches the model, which systems it may reach into, and which checks must clear before anything goes out. That path is where ChatFuse puts its guardrails.

ChatFuse guardrails 8 checks around every model.
1
Harmful request screening Messages are checked against 6 policies and 42 patterns, from bomb making to extracting system internals, before a model sees them.
2
Disguise stripping Look alike letters and hidden characters are normalised first, so a disguised request is still caught.
3
Injection fencing Memory, tool inputs and data from connected apps are marked untrusted and cannot pose as instructions.
4
Human approval on actions Anything that changes something in a connected app needs a one time confirmation that expires in 10 minutes.
5
Abuse limits Bulk generation, oversized prompts and rapid repeat sends hit escalating cooldowns.
6
Nightly red teaming 28 attacks run against staging every night, with planted canary values to catch any leak.
7
Response scrubbing Answers are scanned before they are shown or stored, and anything shaped like credentials or system internals is removed.
8
Private alerts on every block Our team is alerted to each blocked attempt, with personal details scrubbed and user IDs hashed.
No model in ChatFuse gets direct write access to anything.

User messages get checked by us before any model ever sees them. Since the ChatFuse Orchestrator can send a prompt to over 100 different models, we can't rely on the safety training inside any one model. Our checks run out in front, for everyone. So if someone asks how to build a weapon, they get a Fair Use warning no matter which model would have replied. We also clean up the text first, which stops simple tricks like swapping a regular letter for a foreign one that looks the same.

The defense against hidden instructions might count for more than any other. Plenty of attacks never pass through the user's keyboard. Instead they arrive in a file, an email, or a webpage the AI is told to read, carrying buried text along the lines of "disregard your previous instructions." In our system, whatever arrives via memory, tool inputs or connected apps is tagged as plain data. It can't be executed as a command. We test our own defenses this way, and we wrote about it in AI guardrail testing.

Of all these checks, human approval speaks most directly to what happened at Hugging Face. The OpenAI models were handed tooling plus a route to the network, and put both to use. With ChatFuse, a model can pull information from whatever apps you connect, while sending, deleting or changing anything waits on a person saying yes. Each yes belongs to a single user in a single login session and lapses after 10 minutes, so nobody can stash it and reuse it later. At no point does the model hold write access itself. The full design is written up under AI agent write actions.

And every night, we try to break our own system. An automated suite hits our staging environment with 28 different attacks. They cover secret stealing, instruction overrides, prompt injections, roleplaying bypasses, encoded data leaks, and multi turn traps. We also plant canary values so if something leaks, it's a clear match, not a judgment call. The whole process is laid out in automated AI red teaming.

What happens when someone asks ChatFuse how to make a bomb?

We stop that request before it gets anywhere near a model. The person doesn't get an answer, just a Fair Use notice. That block is step one of many, and it's the piece of ChatFuse we've sunk the most hours into.

Incoming messages get tested against 6 separate safety policies, 42 named patterns in total. The toughest policy covers weapons and explosives. Chemical weapons, altering a weapon, building a gun and making a bomb all sit in it, and each one carries a critical rating. Violence is next, spanning terrorism, planning a murder, abuse at home, hurting animals and self harm. Under illegal activity you'll find scams and fraud, breaking into systems, laundering money, trafficking people, and producing or selling drugs. Hate speech has its own policy, as does sexual content, which includes anything involving children. The sixth exists purely for people fishing for our system's internals.

Nobody just types the bad word out plainly, so we clean the text first. Ligatures and full width lettering are converted to ordinary characters. We remove invisible characters tucked inside words, soft hyphens, for instance, or zero width spaces. Letters from Cyrillic or Greek that look just like Latin ones get mapped back to Latin. Swap one letter of a banned word for its Cyrillic twin and it still gets caught.

When someone gets blocked, they usually try again quickly with tiny changes. Our abuse layer catches that. It looks for rapid repeats, bulk content requests, huge prompts designed to hide one bad line in tons of filler, and prompts meant to blow out the context window. All of that triggers cooldowns, and they get longer every time.

Replies get inspected as well. ChatFuse reads every answer before it's shown or stored, removing anything resembling environment variables, tool definitions, credentials or the system prompt itself. So even if a model gets tricked into oversharing, it still can't actually hand it over.

Each block also pings our team. But first, we scrub out personal details such as card numbers, phone numbers and emails. We hash the user IDs too. So we see what was tried without exposing who sent it.

What can't a platform guardrail do?

Guardrails on a platform only cover the traffic that passes through it. They don't stop someone from grabbing an open weight model, putting it on their own server, and stripping out its refusals. That's exactly what I was talking about, and no product feature we build can fix that.

These filters also just don't catch everything. That's why we built a red team suite and have it running nightly, not just once before launch. If you're not constantly testing a guardrail, you're just hoping it works.

So a platform is one piece. The AI labs are another. What we're missing is a layer that covers everyone, even the people who never agree to any terms.

What should AI regulation actually require?

AI regulation should demand real things you can check: outside reviews before any big model launches, forcing companies to report when something goes wrong, strict security for test environments, and audits that actually mean something. Awareness alone doesn't change anything. These would.

By independent, I mean people who aren't on the lab's payroll. Incident reporting needs to be a rule, so the next breach like Hugging Face gets disclosed because it has to be, not because a company feels like being nice. Sandbox security deserves rules because both of this year's escapes came out of test beds. And an audit with no teeth is just PR, which is basically what that voluntary accord amounts to.

That's where my conversation with the assistant ended up. Running a multi model company like ChatFuse, the best thing I can do outside of building product is say this openly and fight for these specific policies. Inside the product, it's about building the platform right while the rulebook gets written.

Can an open source AI model escape onto the internet?

A model just sits there by itself. It needs tools, code to run, and a way onto the network. The OpenAI and Anthropic events in 2026 only happened because the models had all that during testing. If you give an open weight model the same power and remove its refusals, there's no lab monitoring it and nobody who has to say what it did.

Does ChatFuse let AI take actions without approval?

No. Nothing gets changed without you. If a task would actually change data in one of your connected apps, the system stops and makes you confirm it. You have to approve the action. Your approval only counts inside your own session and expires 10 minutes later. Write access is never handed to the model.

Are ChatFuse's guardrails the same for every model?

Yes. The same for everyone, every time. All the checks, screening, stripping out disguised characters, blocking injections, getting approvals, they're built into the platform itself. That means it works exactly the same whether you're using Claude, GPT, Gemini, or Llama. The guardrails are on our side, not inside the model.

If you want AI that has all this protection already in place, you can start on ChatFuse. How we look after your data is on our security page.

Back to Blog

Written by Nico Coetzee

Share

Comments

Loading comments…