The verified answer is A. Compare pre-deployment and post-deployment metrics such as time saved in documentation, number of actionable tasks created, and employee adoption rates. The question is not asking whether the FM produces technically accurate summaries. It is asking whether the FM meets company business objectives, and the stated objective is productivity improvement. AWS AI Practitioner guidance identifies business objective alignment metrics for AI applications, including task completion rate, user satisfaction, and cost per interaction. AWS also identifies business value metrics for generative AI applications such as ROI, efficiency, conversion rate, accuracy, and customer lifetime value. These are business-impact measurements, not just model-quality measurements.
Option A is best because it compares business outcomes before and after deployment. If the tool is intended to improve productivity, the company should measure actual productivity signals: reduced documentation time, more actionable tasks created, higher employee adoption, and other operational improvements. AWS Prescriptive Guidance separates generative AI monitoring into application health, business and user-interaction health, and model quality health. It states that business and user-interaction health evaluates whether the application meets business objectives by tracking adoption, customer satisfaction, productivity improvements, cost savings, and task automation efficiency.
Option B is incomplete because precision, recall, and BLEU are technical evaluation metrics. They can help assess output quality but do not prove productivity improved. Option C may improve the system, but adding RAG does not determine whether the current FM meets business objectives. Option D is also incomplete because employee sentiment alone does not measure actual productivity. Therefore, the correct solution is to compare pre-deployment and post-deployment business metrics.
Submit