Anthropic Claude Certified Architect - Professional CCAR-P Question # 12 Topic 2 Discussion
CCAR-P Exam Topic 2 Question 12 Discussion:
Question #: 12
Topic #: 2
You are identifying the highest-impact optimization for a deployment whose token cost is dominated by a long, repeated system prompt and a large retrieved context per request.
Which optimization most directly targets the dominant cost driver?
A.
Increase retrieval depth on every request to maximize recall, worsening the dominant cost driver by adding more retrieved tokens per request rather than reducing them.
B.
Add additional repeated content to the system prompt to give the model more guidance.
C.
Move the long, repeated system prompt into a cacheable prefix and trim retrieved context to the spans relevant to each query.
D.
Switch every request to the heaviest available model to maximize output quality, accepting that higher per-request inference cost compounds rather than addresses the dominant cost driver.
Option C addresses both components responsible for the excessive cost. Anthropic prompt caching allows stable, repeatedly submitted prompt material—such as system instructions, tool definitions, and reusable background information—to be placed in a consistent prefix. After that prefix is written to the cache, qualifying subsequent requests can reuse it at the lower cache-read cost instead of repeatedly processing the same content at the standard input-token rate.
Retrieval must be optimized separately. Supplying an entire document collection or excessively deep search results increases cost, consumes context capacity, and may reduce answer quality by surrounding the relevant evidence with distracting material. Retrieval should select the smallest set of authoritative passages that provides sufficient evidence for the current query. This typically requires relevance scoring, deduplication, metadata filtering, reranking, and explicit token-budget limits.
Increasing retrieval depth or expanding the repeated system prompt directly worsens the identified cost driver. Moving every request to a heavier model changes the unit economics but does not correct inefficient context construction. The recommended optimization therefore combines prefix caching with query-specific context pruning, followed by evaluation to confirm that reduced context does not lower task accuracy.
Chosen Answer:
This is a voting comment (?). You can switch to a simple comment. It is better to Upvote an existing comment if you don't have anything to add.
Submit