Escalate on policy, low confidence or irreversible actions
Preparation
A two-week study plan
Days 1–6: one domain per session, weighted by exam %. Read the notes, then do 10 practice questions
Day 7: first mock exam to get a baseline
Days 8–11: work on your weakest domains
Days 12 & 14: full mocks. Aim for a comfortable margin above 720
On every wrong answer, read the explanation and re-read the source
Must-read sources
Claude API: Messages & Tool Use
Message Batches (cost/latency trade-off)
Claude Code docs: CLAUDE.md, hooks, subagents, headless
Claude Agent SDK overview
MCP: tools vs resources
The official exam guide + Architect's Playbook
Demo: a practice-question app built with Claude Code that tracks mock scores and weak domains over time.
Test yourself · 5 minutes
Six practice questions
Real exam-style questions from the study app question bank
Shout out A, B, C or D, or raise your hand for each letter
Click an option to mark the room's pick, then press → to reveal the answer
17 questions with answers to take home: examples/test-questions.md
Quiz timer
Question 1 / 6 · Domain 1 · Agent architecture
Your support agent runs a tool-use loop. What should decide that it has finished and the answer can be shown to the customer?
AThe reply contains a phrase like “All done”
Bstop_reason equals end_turn
CA fixed cap of 5 iterations has been reached
DThe last turn contained text but no tool call
✓ B.end_turn is the only reliable completion signal. Text can appear in the middle of a task, and iteration caps are a safety net, not a way to detect completion.
Question 2 / 6 · Domain 1 · Multi-agent research
A coordinator starts a subagent with the prompt “Analyze the document.” The output is generic and off-topic. What's the most likely cause?
AThe subagent has its own isolated context, so the document and goal must be passed in the prompt
BThe subagent inherited conflicting history from the coordinator
CThe subagent needs tool_choice: any
DThe subagent's max_tokens is too high
✓ A. Subagents don't inherit the coordinator's history. Everything they need (the document, the goal, the output format) has to be in their prompt.
Question 3 / 6 · Domain 2 · Tool design
When Claude “uses a tool”, what actually happens?
AClaude runs the function internally and returns the result
BClaude emits a structured tool_use request, your code runs it and returns a tool_result
CAnthropic's servers run your function
DClaude rewrites the tool as Python and runs it in a sandbox
✓ B. The model proposes and your code executes. The result goes back in a tool_result block matched by tool_use_id.
Question 4 / 6 · Domain 3 · Claude Code
A new engineer's Claude Code ignores the team's coding standards. The lead set them up months ago in their own ~/.claude/CLAUDE.md. What's the fix?
ACopy the lead's file to the new engineer's machine
BPaste the standards into every session
CMove them to the project's CLAUDE.md and commit it
DIncrease the context window
✓ C. User-level ~/.claude/CLAUDE.md applies only to that person. Team conventions belong in the project's CLAUDE.md, committed to version control.
Question 5 / 6 · Domain 4 · Structured output
An extraction pipeline sometimes returns broken JSON, and separately, line items that don't add up to the stated total. Which fixes match the two problems?
ABoth are syntax errors, so a JSON schema fixes both
BBoth are semantic errors, so validate and retry for both
CBroken JSON: use a JSON schema. Wrong totals: validate and retry with feedback
DBroken JSON is semantic, wrong totals are syntax
✓ C. A schema guarantees valid JSON with the required fields, not correct values. Semantic errors need validation plus retries or self-correction.
Question 6 / 6 · Domain 5 · Context management
lookup_order returns 40+ fields, but the conversation only needs 4. Long sessions run out of context. What's the best fix?
ATell Claude in the system prompt to ignore irrelevant fields
BCall the tool less often
CUse a model with a bigger context window
DFilter the result in code (e.g. a PostToolUse hook) so only the needed fields enter the context
✓ D. Telling the model to ignore fields still costs the tokens. Trim tool results in code before they reach the context.
02
Your first agent
A beginner's guide, from a markdown file to your own loop
Definition
An agent is a model using tools in a loop
Goal"fix the failing test"
→
Claudedecides the next step
→
Tool callread, search, run…
→
Resultgoes back to Claude
↺ repeat until stop_reason == "end_turn"
The model proposes
Claude asks to call a tool with some inputs. It never runs anything itself.
Your code executes
You run the tool, or the harness does, and send the result back.
The model is stateless
The full history is sent on every turn, so context is a budget you have to manage.
Before you build
Do you actually need an agent?
Start with the simplest option that works:
One API call: classify, summarise, extract
Workflow: fixed steps, and your code decides the order
Agent: open-ended, and the model decides the steps
Build an agent only if all four are true
Complex: the steps are hard to write down in advance
Valuable: the result is worth the extra cost and time
Viable: Claude is good at this kind of task
Recoverable: tests, review or rollback can catch mistakes
Four ways in
Pick the lowest level that works
L1
Claude Code subagent.claude/agents/*.md
No code. A markdown file with a prompt and a tool list. Good for your own dev workflow.
L2
Claude Agent SDKpip install claude-agent-sdk
Claude Code as a library: built-in Read/Edit/Bash/Grep, MCP, hooks. You host it.
L3
Claude API + Tool Runnerclient.beta.messages.tool_runner
Your own tools, and the SDK runs the loop. Good for putting an agent inside your product.
L4
Claude Managed Agentsagents → sessions
Anthropic runs the loop and a sandbox. Good for long-running, scheduled or hosted agents.
Level 1 · no code
A subagent is a markdown file
<!-- .claude/agents/test-fixer.md -->
---
name: test-fixer
description: Runs the test suite and fixes
failing tests. Use after code changes.
tools: Read, Edit, Grep, Glob, Bash
model: sonnet
---
You are a careful test engineer.
1. Run the tests and collect the failures.
2. For each failure, find the root cause
before changing anything.
3. Fix the code, not the test, unless the
test itself is wrong.
4. Re-run the tests and report what you
changed and why.
description tells Claude when to delegate, so write it like a trigger
tools limits what it can touch (least privilege)
It runs in its own context window, so your main chat stays clean
Create one interactively with /agents
Try it: "use the test-fixer agent on this branch"
Level 1 · examples to steal
Seven subagents, ready to copy
repo-explainer
Explains an unfamiliar repo: layout, modules, open TODOs
read-only
code-reviewer
Reviews changes: max 5 findings, each with a breaking input and a fix
never edits
security-auditor
Scans for secrets, injection, unsafe defaults, prompt-only rules
read-only
commit-writer
Drafts a commit message from the diff
Haiku · never commits
test-fixer
Runs the tests, finds the root cause, fixes the code
edits + Bash
test-writer
Adds missing tests without touching the code under test
writes tests/ only
docs-writer
Writes a README from what the code actually does
writes docs
Try one
"Is this secure? Use the security-auditor"
.claude/agents/
Level 2 · Claude Agent SDK
Claude Code in your own script
# pip install claude-agent-sdk
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions
async def main():
async for message in query(
prompt="Find every TODO in this repo and group them by theme",
options=ClaudeAgentOptions(
tools=["Read", "Grep", "Glob"], # the only tools it has
allowed_tools=["Read", "Grep", "Glob"], # no permission prompts
),
):
print(message)
asyncio.run(main())
You get Claude Code's built-in tools, context management and permissions for free
Add MCP servers, hooks and subagents as your needs grow
Level 3 · Claude API
Your own tools, the SDK runs the loop
import anthropic
from anthropic import beta_tool
client = anthropic.Anthropic()
@beta_tool
def lookup_order(order_id: str) -> str:
"""Look up the status of a customer order by its order number."""
return f"Order {order_id}: shipped, arrives Friday" # call your real system here
runner = client.beta.messages.tool_runner(
model="claude-opus-5",
max_tokens=16000,
tools=[lookup_order],
messages=[{"role": "user", "content": "Where is my order A-1042?"}],
)
for message in runner: # one message per turn, stops when Claude is done
print(message)
Under the hood
The loop is driven by stop_reason
stop_reason
What your loop does
tool_use
Run the tool, append a tool_result, call again
end_turn
Done, so show the answer
max_tokens
Truncated: raise the limit or restructure
refusal
Handle it and don't read the content blindly
Append the wholeresponse.content to the history, not just the text
Parallel tool calls: send all results back in one user message
If a tool fails, return is_error: true. Don't drop the result
Agent architecture is the biggest exam domain, at 27%.
Beginner tips
What I wish I'd known
01
Tool descriptions are prompts
Say what the tool does, when to use it and what it returns. Most "bad agent" problems start here.
02
Return less
Tool output fills the context. Return the 5 fields you need, not 40.
03
Enforce in code
Rules that must never break belong in hooks or app logic, not in the prompt.
04
Least privilege
Start read-only. Add write access and Bash only when you need them.
05
Log every turn
Print each tool call and result. You can't debug a loop you can't see.
06
Escalate on purpose
Hand off to a human for irreversible actions, policy cases or low confidence.
Live demo · 5 minutes
From a markdown file to code
0:00
cd examples/demo-repo && claude: a tiny shop with a bug and TODOs
0:30
L1: “use the repo-explainer agent”. Read-only, runs in its own context
1:30
L1: “the tests are red, use test-fixer”. Finds the quantity bug and fixes it
2:30
L2: python 02_agent_sdk.py. The same kind of agent from a script
3:30
L3: python 03_tool_runner.py. Custom tools, with the refund rule enforced in code
4:30
Show 04_manual_loop.py: the stop_reason loop the runner hides