Prompt engineering from scratch fits into a 4-week plan with 30–40 minutes of practice a day. Week one covers prompt structure, week two examples and chaining, week three documents and model differences, and week four agents, workflows and defense against prompt injection. Below is the plan, with exercises from the AI Academy prompt library and techniques from the official Anthropic, OpenAI and Google guides as they stand for the fall 2026 models, including Claude Opus 5.5.
The Prompt Engineering from Scratch course follows the same logic: 4 lessons of about 15 minutes each, from prompt anatomy to system roles. The first 2 lessons are free, and the rest is part of the subscription with a 7-day free trial.
Prompt engineering is the skill of writing instructions for a model so that it consistently produces output that meets your requirements. That is how the OpenAI guide defines it, and the key word is "consistently": one good answer proves nothing.
In 2026 the skill comes down to designing prompts and testing them. Models change every few months: on September 22, 2026, Anthropic released Claude Opus 5.5, where thinking is always on and its depth is set by the effort parameter (Anthropic docs). OpenAI states that even different snapshots of the same model can behave differently, so it recommends pinning model versions and keeping a test suite. Anthropic opens its guide with three prerequisites: clear success criteria, a way to test against them, and a first draft prompt (Prompt engineering overview).
The three official guides agree on the basics and differ on details:
| Technique | Anthropic (Claude) | OpenAI (ChatGPT, API) | Google (Gemini 3) |
|---|---|---|---|
| Structure | XML tags for instructions, context and examples | Identity, Instructions, Examples, Context sections; Markdown and XML | XML tags or Markdown headings, one format per prompt |
| Examples | 3–5 relevant, diverse examples in tags | A dedicated section with sample inputs and desired outputs | Recommends always including few-shot examples |
| Long documents | Documents first, question last: up to 30% better responses in tests | Context usually works best near the end of the prompt | All context first, instructions and question at the very end |
| Testing | Success criteria and empirical tests before tuning | Pinned model snapshot and an eval suite | Iteration: prompt design can take a few attempts |
Sources: Anthropic Prompting best practices, OpenAI Prompt engineering, Google Prompt design strategies.
A working prompt has four blocks: who is answering (role), what the model needs to know (context), what to do (task) and how to present the result (format). The test is Anthropic's "golden rule": show the prompt to a colleague who doesn't know the task. If they would be confused, the model will be too.
Role: you are a financial analyst at a B2B SaaS company.
Context: quarterly revenue, team goals, budget limits (below).
Task: propose 3 OKRs for next quarter.
Format: a table "objective | key result | metric",
no more than 120 words of explanation.
This week's exercises:
Theory: Structure: Role, Task, Context, Format and System Prompts and Stable Roles.
Examples (few-shot prompting) are one of the most reliable ways to set a model's output format, tone and structure. Anthropic recommends 3–5 examples that mirror the real task and differ from each other, each wrapped in an <example> tag so the model can tell them apart from instructions.
The second technique is prompt chaining: a large task is split into steps, and each step's output becomes the next step's input. The third is self-correction: in a separate step the model checks its draft against a list of criteria. With built-in reasoning models such as Claude Opus 5.5, the effort parameter controls thinking depth in the API, while in chat, clearly written quality criteria matter most.
This week's exercises:
Theory: Examples (Few-Shot) and Answer Tone and Chain-of-Thought and Step-by-Step Reasoning.
After two weeks you will have a set of personal templates and your first chains. Prompt Engineering from Scratch reinforces them with quizzes and progress tracking, and you get a certificate on completion. Unlock full access free for 7 days.
With long documents, the order of prompt parts matters most: documents first, then instructions, question at the very end. In Anthropic's tests this order improves response quality by up to 30%, especially with complex multi-document inputs. Google gives the same advice for Gemini 3 and adds a bridge phrase after the data block, such as "Based on the information above…".
Wrap each document in its own tag with its source and date so the model can refer to a specific file. For contracts, reports and data exports, ask for quotes from the source text before conclusions: that makes every claim easy to verify.
Models differ noticeably in their default behavior, which matters when you pick a tool:
This week's exercises:
Model strengths and weaknesses by task are collected in Model comparison.
Adversarial prompting means attacking an AI system through text: trying to make the model break its instructions, leak data or follow someone else's command. OWASP ranks prompt injection first among 2025 risks for LLM applications (LLM01) and splits it into direct attacks from the user's prompt and indirect attacks from websites, files and other external sources the model reads.
Once a prompt becomes part of a workflow, the agent starts reading external data, and every such source can carry someone else's instructions. For Opus 5.5, Anthropic suggests a concrete defense: wrap text the user pasted in a <pasted_content> tag with a random ID, and state in the system prompt that instructions inside it are followed only when the user's own message asks for it. The docs add that tags can be imitated, so this is one layer among several (Prompting Claude Opus 5.5).
This week's exercises:
Theory: Adversarial: injections, jailbreak, red-teaming and Adversarial Prompting and Defending AI Systems. What changed in the model itself is covered in our news post on the Claude Opus 5.5 release.
Progress is measured on a fixed task set: the same inputs are run before and after training and scored against the same criteria. This is how Anthropic and OpenAI recommend working with prompts in production, and the same method works for personal learning.
| Week and skill | Library prompt or workflow | Outcome |
|---|---|---|
| Week 1: role, context, task, format | "Quarterly OKR Draft", "Job Description Without the Bullshit" | 5 personal templates |
| Week 2: few-shot, chaining, self-check | "Objection Handling", "A Cold Email Without a Template" | A 4-step chain with a checklist |
| Week 3: documents, choosing a model | "A Brief Contract Summary", "Clustering Customer Feedback" | 3 models compared on your own task |
| Week 4: agents, injection defense | "Meeting Notes → Tasks", "Lead Qualification" | A workflow that passes an injection test |
For automated scoring across dozens of examples, see the lesson Prompt evaluation: LLM-as-judge and test harnesses and the Prompt Evaluations course. You can check your current level with the level quiz. If a whole department is taking the plan, connect everyone through a team subscription: one subscription for the whole team and centralized access management.
A solid working level in 4 weeks is realistic with daily practice on your own tasks: structure, examples, chaining, document work and security basics. How fast you get there depends on how many real tasks you run through a model each week.
For the first three weeks a chat interface is enough. Coding helps in week four if you build agents through an API, though ready-made workflows also run in no-code tools such as n8n.
The core is shared: a clear task, context, examples, format. Details differ: Anthropic recommends XML tags and documents at the top of the prompt, OpenAI splits instructions into Identity, Instructions, Examples and Context, and Google recommends always adding examples for Gemini 3 and asking explicitly for detailed answers.
It is an attempt to trick a model with text: make it break its rules, leak data or follow an instruction hidden in an email, file or web page. The main defenses are separating data from instructions, limiting the agent's permissions and confirming risky actions manually.
The first 2 lessons of every course and most of the prompt library are free. Full access is $9 a month or $79 a year, with the first 7 days free and cancellation anytime.
Do week one of the plan right now: lessons 1 and 2 of Prompt Engineering from Scratch are free. For weeks 2–4, unlock all courses, prompts and workflows: 7 days free, then $9 a month.