Unfiltered
Sign in Open chat

Three layers, and only one of them is the model

2 min read

Three layers, and only one of them is the model

A model can be trained to refuse things at the weight level, a system prompt can add or remove caution on top of that, and a separate runtime filter can screen input or output independently of both. "Which AI is unfiltered" usually means the first layer, and it is the one a product cannot talk its way around โ€” the other 2 can be swapped without touching it.

"Unfiltered AI models" and "what AI is unfiltered" both assume unfiltered is one property a model either has or does not. It sits at 3 separate layers instead, and a product can be unfiltered at 2 of them and still refuse at the third.

On this page
  1. The model's own training
  2. The instructions layer, on top
  3. The runtime filter, separate from both
  4. Why the same claim can be true and false at once
  5. Reasonable questions

The model's own training

Before a model ever answers a user, it is trained on examples of refusing certain requests, and that training becomes part of its weights the same way grammar and facts do. A heavily trained refusal is not a setting anyone can flip at runtime โ€” it has to be trained out, which is a different and much larger undertaking than editing a prompt.

The instructions layer, on top

A system prompt sets tone and caution level for whatever model sits underneath it, and swapping the prompt can make the same model feel like a different product. It changes how carefully the model treats a subject; it does not undo training the model received before the prompt ever existed.

The runtime filter, separate from both

A filter that screens the request before it reaches the model, or screens the reply before it reaches you, is a third layer again. It can be strict on a lightly trained model and produce something that reads as heavily filtered, or absent on a heavily trained model and produce something that still refuses, because the training is still there underneath.

Why the same claim can be true and false at once

A product built on a lightly trained model with a permissive prompt and no runtime filter is genuinely unfiltered at all 3 layers, and this is what the phrase is usually reaching for. A product built on a heavily trained model, with the prompt and the filter both stripped away, is unfiltered at 2 layers and still hits the training underneath โ€” which is why some services marketed as unfiltered still refuse things, and the marketing is not exactly lying, just describing 2 of the 3.

Reasonable questions

Which layer matters most for what I actually notice?

The model's training, because it is the one no prompt or missing filter reaches. The other 2 layers change how often you notice it and how the refusal is worded, not whether it exists.

Can a permissive prompt make a heavily trained model behave unfiltered?

It can remove the hedging and caveats that sit on top, which is a real change. It generally cannot remove a flat refusal the model was trained to give, because that refusal does not come from the prompt in the first place.

How do I tell which layer is responsible for a given refusal?

A refusal that softens with a different wording of the same request usually comes from the prompt or filter layer. A refusal that holds however the question is rephrased usually comes from training, and no amount of rewording changes that.

Ask the same question here 2 different ways and see which layer answers.

Open unfiltered chat

How an answer gets shaped