Prompting Playbook: Best Practices for LLM Engineers
Summary
This video presents a 'Prompting Playbook' with best practices for working with Large Language Models (LLMs), focusing on two scenarios: migrating existing prompts and building new agentic use cases. It emphasizes the importance of rigorous evaluation suites to identify and address failure modes systematically. Key techniques include applying general prompt hygiene, structuring prompts with clear sections (role, guidelines, policy, tone), using XML tags, defining output contracts, and introducing tools for complex calculations. For new agents, it compares approaches like using larger models, enabling adaptive thinking, and adopting an agentic generate-evaluate-repair loop, highlighting the efficiency gains of the latter for complex tasks and soft constraints.
Key Insights
Use structured formatting (e.g., XML tags) to delineate prompt sections for clarity.
The example prompt is improved by adding XML tags to clearly define sections for the bot's role, general guidelines, policy, and tone of voice. This structure helps the model differentiate between different types of instructions, preventing confusion and improving performance.
A clear separation of guidelines, policy, and data is crucial for model comprehension.
If a human cannot easily distinguish between guidelines, policy, and data within a prompt, the model likely cannot either. Implementing clear structural separation improves the model's ability to process and adhere to the intended instructions.
Provide tools to LLMs for reliable execution of complex tasks like calculations.
Instructions alone do not add capability. To address failures in complex calculations, like proration for billing, provide the model with a specific tool (e.g., a ‘calculate proration tool’). This involves defining the tool schema and implementing its logic for reliable execution.
State both sides of trade-offs when instructing models on escalation decisions.
When an agent needs to decide whether to escalate a case, instruct it on both sides of the trade-off. For example, stating the cost of escalation ($8) without mentioning the cost of getting it wrong (refunds, lost customer trust) leads to over-optimization against escalation. Providing both benefits and costs ensures a more balanced decision.
An agentic generate-evaluate-repair loop can efficiently solve complex tasks with fewer tokens and lower latency.
Breaking down a complex task into an agentic loop (generator, evaluator, repairer) using simple, independent prompts proved more efficient than a single, highly complex prompt. This approach solved all test cases with significantly lower token counts and latency compared to optimized single prompts or larger models with adaptive thinking.
Sections
Introduction to Prompting and Scenarios
Prompting is a critical skill for building effective AI systems with Large Language Models (LLMs).
The presenter, Margot van Laar, an Applied AI Engineer at Anthropic, introduces the session on the Prompting Playbook. She highlights that prompting is one of the first and most critical skills for engineers working with LLMs and essential for building effective AI systems.
Two common LLM prompting scenarios are prompt migration and building new agentic use cases.
The presentation will cover two practical scenarios: 1. Migrating an existing production prompt to a new model or architecture, where it may no longer perform as well. 2. Building an entirely new agentic use case from scratch, requiring prompt creation from zero to one.
A practical example of a customer support bot prompt will illustrate best practices.
To demonstrate best practices, a miniaturized example, inspired by real customer prompts for a telco company's customer support bot (Meridian Mobile), will be used. This example is representative of common problems encountered when maintaining complex prompts developed collaboratively without a clear owner.
The Crucial Role of Evaluations
Evaluations are necessary to rigorously understand prompt performance changes.
Evaluations are essential for providing the rigor needed to understand whether changes to a prompt correlate with improvements in its performance. Different models have different capabilities and behaviors, and a new model might perform worse due to behavioral differences or lower capability.
Eval suites should cover control cases, edge cases, and capability boundaries.
A representative eval suite should include: 1. Control cases: Unambiguous tests the model should always pass. 2. Edge cases: Scenarios where the model has previously failed, ensuring new instructions prevent recurrence. 3. Capability boundaries: Testing the model's understanding of when to hand off to a human or refuse a request.
Systematically targeting failure modes with an eval suite is key to prompt improvement.
The process involves running an initial eval (v0) to identify failure modes, then systematically targeting these modes one at a time by iterating on the prompt to resolve them. This approach is often more practical than writing prompts from scratch.
General prompting hygiene should be applied before targeting specific failure modes.
Before diving into specific failure modes, it's best practice to apply general prompting 101 hygiene to clean up the prompt. This includes removing redundant information and improving its structure, which can provide an initial performance uplift.
Prompt Cleanup and Structure Refinement
Use structured formatting (e.g., XML tags) to delineate prompt sections for clarity.
The example prompt is improved by adding XML tags to clearly define sections for the bot's role, general guidelines, policy, and tone of voice. This structure helps the model differentiate between different types of instructions, preventing confusion and improving performance.
Remove redundant or outdated information, especially model-specific patches.
Redundant information, such as references to hero images or cookies, was removed. Additionally, outdated instructions or 'patches' intended for previous models, like telling the bot it's a human or specific deflection instructions, should be removed as they can lead to over-optimization or incorrect behavior in newer models.
A clear separation of guidelines, policy, and data is crucial for model comprehension.
If a human cannot easily distinguish between guidelines, policy, and data within a prompt, the model likely cannot either. Implementing clear structural separation improves the model's ability to process and adhere to the intended instructions.
Define an output contract to ensure consistent and predictable response formats.
An output contract, specifying the desired output format (e.g., using XML tags), is a key best practice for ensuring consistency, especially for complex output structures like nested JSONs. For conversational agents, this might be less critical but remains a good practice.
Utilize stop sequences in the API call to enforce output format consistency.
In addition to prompt engineering, harness-level controls like stop sequences can help ensure output consistency. A stop sequence, such as a closing XML tag, can signal the model to stop generating text, enforcing the defined output structure programmatically.
Structured outputs can programmatically enforce consistency for complex schemas.
For more complex output schemas, utilizing structured output features within the API can provide a more programmatic and robust way to ensure consistency and accuracy in the model's responses.
Targeting Specific Failure Modes
Models can withhold information they possess due to outdated defensive instructions.
A failure mode was observed where the model withheld specific hotspot data for a legacy plan, instead deflecting the user to check a URL. This was due to an old instruction, 'Never give a customer the wrong plan details,' which newer, more capable models overfitted to, causing them to withhold accessible information.
Update prompts to reflect model evolution and remove redundant defensive patches.
Newer, more intelligent models are better at instruction following. Instructions designed as patches for previous model behaviors may become redundant or cause over-optimization. Prompts should be updated to guide the model's balanced reasoning rather than relying on old, overly restrictive directives.
Use version control to track defensive prompt changes and their eventual effects.
To manage the impact of defensive changes (patches) introduced for previous models, using version control is recommended. This practice helps track the reason for these changes and allows for backtracking if they produce unwanted effects in newer model generations.
Provide tools to LLMs for reliable execution of complex tasks like calculations.
Instructions alone do not add capability. To address failures in complex calculations, like proration for billing, provide the model with a specific tool (e.g., a ‘calculate proration tool’). This involves defining the tool schema and implementing its logic for reliable execution.
Giving models tools enhances their ability to reason over and execute complex problems reliably.
Instead of just telling a model to perform a calculation correctly, giving it access to a tool that can perform the calculation is the correct approach. This provides the model with the actual capability to reason over harder problems and execute them reliably.
State both sides of trade-offs when instructing models on escalation decisions.
When an agent needs to decide whether to escalate a case, instruct it on both sides of the trade-off. For example, stating the cost of escalation ($8) without mentioning the cost of getting it wrong (refunds, lost customer trust) leads to over-optimization against escalation. Providing both benefits and costs ensures a more balanced decision.
As models become more intelligent, clearly articulate trade-offs for nuanced decision-making.
Sophisticated models are better at making trade-offs themselves. It is crucial to explicitly state both sides of any trade-off in the instructions to guide the model's decision-making process effectively, preventing it from optimizing for only one aspect.
Building New Agentic Use Cases from Scratch
Consider model choice, prompt structure, and harness configurations for new agents.
When building a new agent from scratch, such as a retail staff scheduler, consider not only the prompt but also the chosen model and the surrounding harness (API call parameters, tools) to explore their impact on performance.
Larger, more capable models can reduce violations even with basic prompts.
In an example of building a staff scheduler, using a more capable model like Opus 4.7 significantly reduced the number of violations compared to a smaller model like Sonnet 4.6, even when initial prompts were basic and failing.
Adaptive thinking can improve compliance but may increase cost and latency.
Enabling adaptive thinking in a capable model (Opus 4.7) reliably generated compliant schedules without changing the prompt. However, this approach significantly increased token usage and latency, highlighting a trade-off between capability and efficiency.
Prompt optimization on smaller models may not overcome inherent capability limitations.
Improving a prompt for a smaller model (Sonnet 4.6) by adding instructions to check work showed some improvement but still failed. Attempts to fix this by increasing max tokens led to even higher token usage and latency, suggesting prompt alone might not be enough for complex tasks with less capable models.
An agentic generate-evaluate-repair loop can efficiently solve complex tasks with fewer tokens and lower latency.
Breaking down a complex task into an agentic loop (generator, evaluator, repairer) using simple, independent prompts proved more efficient than a single, highly complex prompt. This approach solved all test cases with significantly lower token counts and latency compared to optimized single prompts or larger models with adaptive thinking.
Agentic loops allow for incorporating soft constraints at runtime, offering flexibility.
A key benefit of the generate-evaluate-repair loop is the ability to incorporate soft requirements (e.g., employee preferences, specific shift needs) at runtime via the evaluation prompt. This avoids needing to modify backend logic for case-by-case constraints.
Conclusion and Key Takeaways
Rigorous evaluation and systematic failure mode targeting are essential for prompt improvement.
The video covered maintaining prompts and building new agents, emphasizing the need for evaluation suites to rigorously assess prompt changes. Targeting failure modes one-by-one with techniques like adding structure and using tools improved model behavior.
Structuring prompts, removing redundancy, and providing tools are fundamental best practices.
General hygiene principles, such as clear prompt structure using XML tags, removing outdated instructions, and integrating tools for reliable task execution, immediately uplift prompt performance.
Agentic approaches, like generate-evaluate-repair loops, offer efficiency for complex tasks.
For new agentic bots, splitting tasks into separate prompts within an agentic loop (generate-evaluate-repair) is more efficient than a single, monolithic prompt, offering benefits in token usage, latency, and flexibility for soft constraints.
Ask a Question
*Uses 1 Wisdom coin from your coin balance









