Claude, Nutshell Series

Claude Exam Quick Revision 12: Human Review & Provenance

Aggregate Accuracy Trap

Overall accuracy = 97%

This may still hide poor accuracy for a particular document type or field.

Measure Accuracy By

  • Document type.
  • Individual field.
  • Relevant segment.

Stratified Random Sampling

Do not review only low-confidence results.

Sample some high-confidence outputs too.

Why? Detect confidently wrong predictions and new failure patterns.

Confidence Calibration

Field confidence
+
Labeled validation set
↓
Calibrated threshold
↓
Human review

Memory: Confidence must be validated against actual correctness.

Claim-Source Mapping

Every research finding should preserve:

Claim
Evidence
Source URL / document
Publication date

Conflicting Sources

Do NOT arbitrarily choose one value.

Source A → 42%
Source B → 51%

Report both with attribution.

Temporal Context

Preserve:

  • Publication date.
  • Data-collection date.

Statistics from different periods may not actually conflict.

Coverage Gaps

If sources are unavailable:

  • Explicitly report the gap.
  • Identify affected topic.
  • Do not pretend research is complete.

Report Structure

Well-supported findings

Contested findings

Coverage gaps

Content-Specific Formatting

  • Financial data → tables.
  • News → prose.
  • Technical findings → structured lists.

Fast Revision


Overall accuracy can hide weak segments.
Sample high-confidence outputs.
Confidence must be calibrated.
Claim must keep source.
Conflicting sources → preserve both.
Missing coverage → report it.

Claude, Nutshell Series

Claude Exam Quick Revision 11: Advanced Batch & Extraction

Schema Error vs Semantic Error

  • Schema error: Invalid structure/type.
  • Semantic error: Valid structure but wrong meaning/value.

Remember: JSON Schema does not guarantee semantic correctness.

Self-Validation Fields

{
  "stated_total": 120,
  "calculated_total": 110,
  "conflict_detected": true
}

Expose inconsistencies explicitly.

detected_pattern

Add fields such as:

detected_pattern

Use them to analyse which code patterns repeatedly cause false positives.

Batch custom_id

  • Assign each request a custom_id.
  • Correlate request with response.
  • Identify individual failures.

Resubmit Only Failed Items

100 documents
↓
7 fail
↓
Identify failed custom_id values
↓
Fix only those inputs
↓
Resubmit 7

Context-Limit Failure

If an oversized document fails:

Chunk document
↓
Resubmit failed document

Before Large Batch

Small sample
↓
Refine prompt
↓
Validate
↓
Run large batch

Benefit: Better first-pass success and fewer expensive retries.

SLA Planning

Batch processing can take up to 24 hours.

Work backwards from the required SLA when deciding submission frequency.

Fast Revision


Schema valid ≠ semantically correct.
custom_id = correlation.
Failed subset = resubmit subset.
Oversized document = chunk.
Large batch = test sample first.

Claude, Nutshell Series

Claude Exam Quick Revision 10: Iterative Refinement Patterns

Concrete Examples

If prose instructions produce inconsistent results:

Give 2–3 concrete input/output examples.

Interview Pattern

Ask Claude to question you before implementation when important requirements may be missing.

Useful for discovering:

  • Failure modes.
  • Edge cases.
  • Cache invalidation.
  • Security requirements.
  • Performance constraints.

Memory: Unknown design space → interview first.

Interacting Problems

If fixes affect each other:

Explain issues together
↓
Claude reasons across dependencies

Independent Problems

Fix issue 1
↓
Validate
↓
Fix issue 2

Iterative Test Refinement

Tests
↓
Run
↓
Share failure
↓
Improve implementation
↓
Repeat

CI Review Re-Runs

When reviewing after new commits:

  • Include prior findings.
  • Report only new issues.
  • Report still-unresolved issues.
  • Avoid duplicate PR comments.

Test Generation

Provide existing test files before generating new tests.

This helps avoid duplicate scenarios.

Fast Revision


Inconsistent transformation → examples.
Unknown design → interview.
Interacting issues → together.
Independent issues → sequential.

Claude, Nutshell Series

Claude Exam Quick Revision 9: Advanced Claude Configuration

@import

Use @import to keep CLAUDE.md modular.

@import ./standards/testing.md
@import ./standards/api.md

Use: Reference focused standards instead of creating one huge CLAUDE.md.

/memory

  • Shows which memory/instruction files are loaded.
  • Useful for diagnosing inconsistent behaviour.

Skill: context: fork

context: fork
  • Runs skill in isolated subagent context.
  • Useful for verbose analysis.
  • Useful for brainstorming alternatives.
  • Prevents main-context pollution.

Skill: allowed-tools

allowed-tools:
  - Read
  - Grep

Restricts which tools the skill can use.

Skill: argument-hint

Shows which arguments the user should provide when invoking the skill.

Memory Trick

context: fork → isolate
allowed-tools → restrict
argument-hint → guide input

Explore Subagent

  • Use for verbose codebase exploration.
  • Return concise findings to main session.
  • Useful during complex planning.

Plan + Execute Pattern

Plan
↓
Explore
↓
Choose approach
↓
Direct execution

Plan Mode and direct execution can be used together.

Claude, Nutshell Series

Claude Exam Quick Revision 8: Advanced MCP Design

Split Overly Generic Tools

Instead of:

analyze_document

Prefer:

extract_data_points
summarize_content
verify_claim_against_source

Reason: Clear tool boundaries improve selection reliability.

Avoid Overlapping Tool Names

Bad:

analyze_content
analyze_document

Better:

extract_web_results
analyze_pdf_document

System Prompt Can Bias Tool Selection

  • Tool descriptions may be correct.
  • System-prompt keywords can still create unwanted tool associations.
  • If tool routing remains wrong, inspect both descriptions and system instructions.

MCP isError

Use structured failure information such as:

{
  "isError": true,
  "errorCategory": "transient",
  "isRetryable": true,
  "message": "Service temporarily unavailable"
}

Empty Result vs Failure

Successful query + no matches
≠
Query failed

Exam trap: Never hide an access failure by returning an empty successful result.

Local Recovery

  • Subagent handles transient failures locally where possible.
  • Only unresolved failures go to coordinator.
  • Include partial results and attempted recovery.

MCP Resources

Resources expose content/catalogs such as:

  • Database schemas.
  • Documentation hierarchies.
  • Issue catalogs.
  • Available datasets.

Benefit: Reduces unnecessary exploratory tool calls.

Existing vs Custom MCP Server

  • Standard integration → prefer existing/community MCP server.
  • Team-specific workflow → custom MCP server.

Fast Revision


Generic tool → split.
Ambiguous name → rename.
Resource → expose information.
Failure ≠ empty result.
Transient error → local recovery.

Claude, Nutshell Series

Claude Exam Quick Revision 7: Session State & Resumption

Resume vs Fork

  • –resume <session-name> → Continue an existing session.
  • fork_session → Create an independent branch from shared prior context.

Memory: Resume = continue | Fork = diverge.

When to Resume

  • Previous analysis is still valid.
  • Files have not changed significantly.
  • You are continuing the same investigation.

When to Start Fresh

  • Old tool results are stale.
  • Code changed substantially.
  • Previous context may mislead Claude.

Best pattern: New session + structured summary.

Changed Files

When resuming after code changes, tell Claude exactly which files changed.

This enables targeted re-analysis instead of repeating the whole investigation.

Crash Recovery

Agent structured exports
↓
Manifest / checkpoint
↓
Coordinator reloads state
↓
Relevant state injected into agents

Fast Revision

  • Resume → continue valid context.
  • Fork → explore alternative.
  • Stale context → fresh session + summary.
  • Crash recovery → structured state + manifest.
Claude, Nutshell Series

Claude Exam Quick Revision 6: Tools, Configuration & Rules

Built-In Tools

  • Read: Read a known file.
  • Edit: Targeted replacement inside an existing file.
  • Write: Create or overwrite an entire file.
  • Glob: Find files by filename/path pattern.
  • Grep: Search text inside files.
  • Bash: Run shell commands.
  • Agent: Delegate work to a separate-context subagent.

Memory: Glob finds files. Grep finds text.

Configuration Scopes

  • User: ~/.claude/ → me across projects.
  • Project: repository configuration → team/repository.
  • Local: personal settings for one repository.
  • Managed: organization-enforced configuration.

Important Files

  • CLAUDE.md: Instructions and project guidance.
  • settings.json: Claude Code configuration.
  • settings.local.json: Personal project overrides.
  • .claude/rules/*.md: Modular/path-specific instructions.
  • .mcp.json: Shared MCP configuration.
  • SKILL.md: Reusable skill definition.
  • .claude/agents/: Custom subagents.

Path-Specific Rules

---
paths:
  - "**/*.tf"
---

Use path-scoped rules when standards should apply only to particular files.

Global standards → CLAUDE.md
File-specific standards → .claude/rules/

Important Principle

CLAUDE.md provides instructions but does not provide deterministic security enforcement.

Claude, Nutshell Series

Claude Exam Quick Revision 5: Multi-Agent Workflow Patterns

Prompt Chaining

Use when the sequence is predictable and known in advance.

Example: Extract → Validate → Transform → Publish.

Dynamic Adaptive Decomposition

Use when intermediate discoveries determine the next step.

Example: Debug issue → inspect logs → discover database problem → inspect queries.

Routing

Classify the input and send it to the appropriate specialist or workflow.

Parallel Sectioning

Run different independent subtasks simultaneously.

Example: Security review + performance review + maintainability review.

Parallel Voting

Run the same task multiple times and aggregate results for greater confidence.

Orchestrator-Workers

A coordinator dynamically creates and delegates subtasks to specialized workers.

Evaluator-Optimizer

Generate → critique → improve → repeat.

Context Efficiency

Let subagents perform verbose exploration and return concise synthesized findings to the coordinator.

Memory:
Chain = fixed path.
Route = choose path.
Parallel = many paths.
Orchestrator = create paths dynamically.
Agent = decide next step dynamically.

Claude, Nutshell Series

Claude Exam Quick Revision 4: Hooks, Enforcement & Escalation

PreToolUse

Runs before a tool executes.

Best for:

  • Blocking unauthorized transactions.
  • Enforcing refund limits.
  • Checking prerequisites.
  • Preventing policy violations.

Example: Refund > $500 → block and escalate before process_refund executes.

PostToolUse

Runs after a tool returns.

Best for:

  • Trimming verbose results.
  • Normalizing MCP outputs.
  • Keeping only fields needed by the agent.

Example: Tool returns 40 fields → retain status, shipping date and tracking number.

Escalation: Good Triggers

  • Explicit request for a human.
  • Policy gap.
  • Permission boundary.
  • High-risk operation.
  • Deterministic hook blocks an action.

Escalation: Bad Triggers

  • Negative sentiment alone.
  • Profanity alone.
  • Self-reported model confidence alone.

Memory: Prompts guide. Hooks enforce.

Claude, Nutshell Series

Claude Exam Quick Revision 3: Agentic Loops & Tool Calling

Agentic Loop

Flow:

Claude → tool_use → execute tool → tool_result → Claude → next action → end_turn

When Should the Loop Stop?

  • stop_reason = tool_use: Execute tool and continue.
  • stop_reason = end_turn: Claude has finished the turn.

Do not: Parse phrases such as “task complete” or “I am done.”

A maximum-turn limit is useful as a safety guard, but it is not the primary completion signal.

Returning Tool Results

  • Return the result in a user-role message.
  • Use a tool_result block.
  • Match the original tool_use_id.
  • The result does not have to be JSON-only.

Multi-Turn Tool Calling

Each tool result can influence Claude’s next decision.

Example: Search → Read → Analyse → Call API → Final answer.

Memory: tool_use = continue. end_turn = finish.