Quick revision for evaluation, testing, monitoring and optimisation questions.
Evaluation Sequence
Define Metrics → Build Dataset → Run Evaluation → Analyse → Decide
Evaluation Dataset
Include:
Representative cases
Edge cases
Adversarial cases
The evaluation set should represent the actual user population and workload.
Stable Reference Set
Keep a stable evaluation/reference set across:
Prompt changes
Model-version upgrades
Architecture iterations
This helps detect quality drift.
Testing Types
| Test | Purpose |
|---|---|
| Regression | Detect previously working behaviour that has broken. |
| Adversarial | Prompt injection and malformed/malicious inputs. |
| Integration | Validate end-to-end behaviour across components. |
| Red-team | Probe security and access-control weaknesses. |
Model Upgrade
Do not assume a new model behaves the same.
Re-run the stable evaluation set before promotion.
Observability
Capture:
Latency
Token usage
Distributed traces
Tool-call outcomes
Redacted payload information
Model identity/version
Request correlation ID
Important Observability Gaps
Missing:
Model identity/version
or
Request correlation ID across agent/tool calls
creates a real diagnostic blind spot.
SLA
A useful SLA contains:
Metric + Threshold + Evaluation Window + Breach Consequence
Example:
p95 latency < 800 ms over a 28-day window
A statement such as “the system should be fast” is not a measurable SLA.
Cost Optimisation
Common optimisations:
Prompt caching
Trim irrelevant retrieval context
Cache repeatedly used documents/chunks
Tiered model routing
Avoid repeated processing
Prompt Caching
Best for:
Long static content repeated across many requests.
Place stable cacheable content before request-specific content.
Token Lifecycle
Input Preparation
Trim irrelevant retrieved passages.
Summarise old conversation history.
Prompt Construction
Place stable content first.
Use a cacheable repeated prefix.
Output Handling
Validate structured schema.
Persist validated results.
Accuracy Must Stay Unchanged
If the requirement says accuracy must remain unchanged:
Remove redundant cost first.
Prefer caching over immediately switching to a weaker model or deleting information required for reasoning.
Fast Exam Recall
Prompt changed and old tests fail? Regression test.
Prompt injection? Adversarial test.
Whole pipeline? Integration test.
Model upgraded? Stable evaluation set.
Repeated prompt? Prompt caching.
Good SLA? Metric + threshold + window + consequence.