Topic

Prompt Engineering

Prompt design patterns, instruction hierarchy, and defensive prompt construction.

prompt engineeringsystem promptsinstruction hierarchyguardrailstask decomposition
Evergreen Overview

Prompt engineering is not just about better outputs. In practice it shapes reliability, scope, fallback behavior, and how well an AI system resists misuse when instructions, tools, and untrusted content collide.

What matters most
  • Instruction hierarchy and role separation
  • Clear task boundaries, fallback behavior, and refusal handling
  • Prompt structures that support monitoring and repeatable evaluation
Where teams get into trouble
  • Overloading prompts with too many responsibilities
  • Relying on wording instead of system controls
  • Treating prompts as static text instead of part of application design
Who this page is for
  • Teams operating prompt-heavy workflows
  • Builders refining assistant and agent behavior
  • Reviewers trying to connect prompt design to safety and risk
References

Current notes, events, and source material

These items are included because they add useful evidence, framing, implementation detail, or upcoming context for teams working in this area.

OpenAI News September 22, 2026 guide

Diagnose prompt-cache misses without widening an agent’s tool access

OpenAI’s prompt-caching guidance explains how to compare requests for changes that invalidate shared prefixes, place explicit breakpoints, and preserve tool definitions while changing which tools are callable. GPT-6 can also receive appended reasoning-effort updates without rewriting the earlier prefix. Workload savings still need measurement.

OpenAI News March 24, 2026 tool

Teen-safety policy prompts: adapt labels and regression-test the classifier

OpenAI’s Teen Safety Policy Pack supplies prompt-based classification policies and matching validation datasets for gpt-oss-safeguard. The six initial areas cover risks including dangerous activities, harmful body ideals and age-restricted goods. Developers map the policy labels into filtering, review or monitoring workflows and can adapt the prompts to their application. These inspectable starting materials do not provide comprehensive protection or establish performance on a particular product’s users and content.

Skill engineering: independent reviews, edit hooks and rule-level evaluations video thumbnail Play video
AI Engineer September 21, 2026 video

Skill engineering: independent reviews, edit hooks and rule-level evaluations

AI Engineer’s notes from Paul Bakaus’s workshop explain separate visual and deterministic reviews, selective instruction loading, edit hooks and rule-by-rule ablation tests. They distinguish blocking pre-tool hooks from post-edit feedback, describe portability failures, and retain human judgment because aesthetic evaluators can reward the wrong behavior.

GEPA: use execution feedback to improve prompts and agent programs video thumbnail Play video
AI Engineer September 26, 2026 video

GEPA: use execution feedback to improve prompts and agent programs

Lakshya Agrawal’s publisher notes describe GEPA’s reflective search: inspect execution traces and errors, propose text changes, score candidates, and retain alternatives that work well on different examples. Optimize Anything extends the editable object from prompts to programs, skills and policies. The reusable method depends on informative feedback and a suitable evaluator. The talk’s large benchmark gains lack complete evaluation protocols in the presentation, so they do not establish a universal advantage over reinforcement learning.