Claude, Nutshell Series

CCAR-P Revision 2: Architecture Patterns, Prompting, RAG & Model Selection

Fast revision notes for Claude architecture design questions.

Workflow vs Agentic Pattern

Scenario Choose
Known fixed steps Workflow
Predictable sequence Workflow
Predictable token cost Workflow
Open-ended planning Agentic
Dynamic tool selection Agentic

Example:

Extract → Classify → Summarise → Persist = Workflow

Exam rule: Do not choose an agent just because it is more flexible.

Prompting

Zero-shot

Best starting point when categories and descriptions are already well defined.

Few-shot

Useful when you want to demonstrate exact output structure, field names, or examples.

Chain-of-thought

Not the default for every task. Simple routing or closed-set classification does not automatically need multi-step reasoning.

Leading Questions

Reduce bias by asking for:

Even-handed comparison + explicit evaluation criteria.

RAG

Use Retrieval-Augmented Generation when answers must be grounded in an authoritative knowledge source.

Typical architecture:

User → Retrieval → Relevant Documents → Claude → Answer + Citations

When RAG Is Strong

Choose RAG when you need:

Authoritative answers
Citations
Auditability
Predictable cost
Predictable latency

For compliance Q&A, prefer retrieval from the approved internal corpus instead of an unrestricted web-search agent.

Grounding Problems

If Claude starts contradicting retrieved sources:

Check model-version changes.
Check grounding instructions.
Require source-supported responses.
Require citations.
Add verification.

Model Selection

Correct sequence:

Requirements → Candidate Model → Representative Evaluation → Final Selection

Define Before Testing

Quality bar
Latency tolerance
Expected volume
Cost constraints

Model Selection Traps

Do not automatically choose:

The newest model.
The largest model.
The most capable model.

Choose the lightest model that consistently meets requirements.

Tiered Routing

Routine Request → Smaller/Faster Model

Complex Request → More Capable Model

This can reduce both cost and latency.

High-Volume Classification

For well-defined labels and tight latency requirements, prefer a lighter model rather than deep reasoning on every request.

Fast Exam Recall

Predictable steps? Workflow.
Dynamic planning? Agent.
Defined categories? Zero-shot first.
Need format examples? Few-shot.
Authoritative answers? RAG.
Simple traffic? Smaller model.
Complex traffic? Higher-capability model.

Leave a comment