Claude, Nutshell Series

CCAR-P Revision 4: Evaluation, Observability, Cost & Performance

Quick revision for evaluation, testing, monitoring and optimisation questions.

Evaluation Sequence

Define Metrics → Build Dataset → Run Evaluation → Analyse → Decide

Evaluation Dataset

Include:

Representative cases
Edge cases
Adversarial cases

The evaluation set should represent the actual user population and workload.

Stable Reference Set

Keep a stable evaluation/reference set across:

Prompt changes
Model-version upgrades
Architecture iterations

This helps detect quality drift.

Testing Types

Test Purpose
Regression Detect previously working behaviour that has broken.
Adversarial Prompt injection and malformed/malicious inputs.
Integration Validate end-to-end behaviour across components.
Red-team Probe security and access-control weaknesses.

Model Upgrade

Do not assume a new model behaves the same.

Re-run the stable evaluation set before promotion.

Observability

Capture:

Latency
Token usage
Distributed traces
Tool-call outcomes
Redacted payload information
Model identity/version
Request correlation ID

Important Observability Gaps

Missing:

Model identity/version

or

Request correlation ID across agent/tool calls

creates a real diagnostic blind spot.

SLA

A useful SLA contains:

Metric + Threshold + Evaluation Window + Breach Consequence

Example:

p95 latency < 800 ms over a 28-day window

A statement such as “the system should be fast” is not a measurable SLA.

Cost Optimisation

Common optimisations:

Prompt caching
Trim irrelevant retrieval context
Cache repeatedly used documents/chunks
Tiered model routing
Avoid repeated processing

Prompt Caching

Best for:

Long static content repeated across many requests.

Place stable cacheable content before request-specific content.

Token Lifecycle

Input Preparation

Trim irrelevant retrieved passages.
Summarise old conversation history.

Prompt Construction

Place stable content first.
Use a cacheable repeated prefix.

Output Handling

Validate structured schema.
Persist validated results.

Accuracy Must Stay Unchanged

If the requirement says accuracy must remain unchanged:

Remove redundant cost first.

Prefer caching over immediately switching to a weaker model or deleting information required for reasoning.

Fast Exam Recall

Prompt changed and old tests fail? Regression test.
Prompt injection? Adversarial test.
Whole pipeline? Integration test.
Model upgraded? Stable evaluation set.
Repeated prompt? Prompt caching.
Good SLA? Metric + threshold + window + consequence.

Leave a comment