Pre-Winter Sale Special Limited Time 70% Discount Offer - Ends in 0d 00h 00m 00s - Coupon code: force70

Anthropic Claude Certified Architect - Professional CCAR-P Question # 30 Topic 4 Discussion

Anthropic Claude Certified Architect - Professional CCAR-P Question # 30 Topic 4 Discussion

CCAR-P Exam Topic 4 Question 30 Discussion:
Question #: 30
Topic #: 4

You are responding to an adversarial input pattern in which users include text claiming admin authority and instructing the model to bypass safety restrictions.

Which combination of controls most effectively mitigates this attack pattern?


A.

Trusting that the model will intrinsically recognize and reject all bypass attempts without prompt-level instructions, runtime classifiers, scoped permissions, or audit logging.


B.

Prompt-level instructions that treat user content as untrusted data, runtime classifiers that detect override attempts, scoped tool permissions that cannot be elevated by user content, and audit logging of attempts.


C.

Removing all safety restrictions and guardrails to eliminate the attack surface that bypass attempts target, accepting that this makes the assistant unrestricted for all inputs.


D.

Granting users any privilege level they assert in their message content, on the assumption that cooperative behavior requires honoring self-declared authority without independent verification.


Get Premium CCAR-P Questions

Contribute your Thoughts:


Chosen Answer:
This is a voting comment (?). It is better to Upvote an existing comment if you don't have anything to add.