Guardrails, Moderation & Prompt-Injection Defense

Lesson 11 of 12 12 minIncluded with Builder

Keep AI systems safe and trustworthy: moderation, prompt-injection defense, safe tool use, and human-in-the-loop where it counts.

What you'll learn

  • Keep AI systems safe and trustworthy: moderation, prompt-injection defense, safe tool use, and human-in-the-loop where it counts.
  • How guardrails, moderation & prompt-injection defense fits into the Advanced AI Engineering track
  • A hands-on project step you can put in your portfolio

The hands-on project

Design the safety architecture for an AI agent that reads customers' incoming support emails and can take actions: reply to the customer, issue a refund, or escalate to a human. Specify your defense in depth: where you place input and output moderation; how you protect against prompt injection given the agent is reading untrusted email content; how you'd scope and gate each of the three tools (which run automatically vs require human approval, and why); and what you'd log for audit. Then write two sentences for your system prompt that explicitly instruct the model to treat email content as untrusted data that may attempt to override its instructions.

The full brief, voice tutor walkthrough, and feedback are inside the lesson.

Unlock Advanced AI Engineering with Builder

Builder unlocks all six tracks, unlimited voice tutoring, and the all-tracks Certificate of Achievement.