The existing evaluation set does not represent the population or operating conditions the hiring-support system will encounter. Option B corrects this coverage defect by expanding geographic and tenure slices and requiring the system to be rescored before broader deployment.
Aggregate performance on one narrow group can conceal substantial differences among regions, experience levels, job families, languages, and other relevant cohorts. The expanded dataset should therefore support disaggregated metrics, not merely one combined score. It should also include representative, edge, and adversarial cases, with labels and grading procedures reviewed for consistency and potential bias.
Anthropic’s evaluation guidance states that evaluations should mirror the real-world task distribution and include edge cases. Success criteria should also be relevant to the application’s actual purpose and users. Define Success Criteria and Build Evaluations
Option A replaces one unrepresentative method with another. Option C deliberately narrows coverage further. Option D mistakes high performance on a restricted subset for evidence of generalization.
Because hiring is a consequential domain, the organization should combine representative quantitative evaluation, subgroup analysis, expert review, monitoring, access controls, transparency, and meaningful human responsibility for final decisions.
Study Guide references/topics: Dataset representativeness; evaluation slices; subgroup performance; bias detection; consequential-use governance; pre-release rescoring.
===============
Submit