Moderation
AI moderation
Rules catch patterns. AI judges the message you cannot list: sarcasm, coded harassment, a scam dressed as a friendly DM — against a policy you write per server.
- Per-server policy prompt
- Fails open if unsure/down
- Daily call budget
- Free + Premium ceilings
Why teams turn it on
Your policy, plain English
The judge prompt is yours. Set the tone for each community instead of shipping one global blacklist that never fits.
Conservative by design
Borderline and benign traffic is left alone. If the model is slow, down, or over budget, ModShield fails open — it never punishes on a guess.
Bounded cost
Pre-filters skip short noise. A hard daily cap per server keeps spend predictable; Premium raises the ceiling when volume needs it.
Sits on top of rules
AI is the layer for context. Link filter, automod, and flood still handle the cheap, obvious cases first.
How it works
1. Pre-filter
Trivial or too-short messages never reach the model. Obvious rule hits should already be handled by cheaper filters.
2. Policy judge
Eligible messages are classified against your per-server policy prompt with conservative thresholds.
3. Act or pass
Clear violations can delete/warn/strike like any other filter. Uncertainty or outage → pass (fail open).
Real situations
Friendly-looking scams
“Hey can you help me verify my wallet” passes keyword lists but fails a scam-aware policy.
Coded harassment
Dogwhistles and pile-ons that never use banned words still violate a clear community policy.
Multilingual communities
Policy can describe intent in English while messages arrive in several languages the model can still judge.
FAQ
Does AI ban people automatically by default?
You choose actions. The default philosophy is conservative; many servers start with delete/warn only.
What if the model is down?
Fail open: messages go through rather than random punishments.
Is Premium required?
Core AI may run on Free with a lower daily cap; Premium raises headroom. See billing docs.