Fast revision notes for Claude architecture design questions.
Workflow vs Agentic Pattern
| Scenario | Choose |
|---|---|
| Known fixed steps | Workflow |
| Predictable sequence | Workflow |
| Predictable token cost | Workflow |
| Open-ended planning | Agentic |
| Dynamic tool selection | Agentic |
Example:
Extract → Classify → Summarise → Persist = Workflow
Exam rule: Do not choose an agent just because it is more flexible.
Prompting
Zero-shot
Best starting point when categories and descriptions are already well defined.
Few-shot
Useful when you want to demonstrate exact output structure, field names, or examples.
Chain-of-thought
Not the default for every task. Simple routing or closed-set classification does not automatically need multi-step reasoning.
Leading Questions
Reduce bias by asking for:
Even-handed comparison + explicit evaluation criteria.
RAG
Use Retrieval-Augmented Generation when answers must be grounded in an authoritative knowledge source.
Typical architecture:
User → Retrieval → Relevant Documents → Claude → Answer + Citations
When RAG Is Strong
Choose RAG when you need:
Authoritative answers
Citations
Auditability
Predictable cost
Predictable latency
For compliance Q&A, prefer retrieval from the approved internal corpus instead of an unrestricted web-search agent.
Grounding Problems
If Claude starts contradicting retrieved sources:
Check model-version changes.
Check grounding instructions.
Require source-supported responses.
Require citations.
Add verification.
Model Selection
Correct sequence:
Requirements → Candidate Model → Representative Evaluation → Final Selection
Define Before Testing
Quality bar
Latency tolerance
Expected volume
Cost constraints
Model Selection Traps
Do not automatically choose:
The newest model.
The largest model.
The most capable model.
Choose the lightest model that consistently meets requirements.
Tiered Routing
Routine Request → Smaller/Faster Model
Complex Request → More Capable Model
This can reduce both cost and latency.
High-Volume Classification
For well-defined labels and tight latency requirements, prefer a lighter model rather than deep reasoning on every request.
Fast Exam Recall
Predictable steps? Workflow.
Dynamic planning? Agent.
Defined categories? Zero-shot first.
Need format examples? Few-shot.
Authoritative answers? RAG.
Simple traffic? Smaller model.
Complex traffic? Higher-capability model.