Microsoft Azure AI Fundamentals (Updated Version) AI-901 Question # 31 Topic 4 Discussion
AI-901 Exam Topic 4 Question 31 Discussion:
Question #: 31
Topic #: 4
You are deploying a generative Al model to a Microsoft Foundry project. You assign a higher tokens per minute (TPM) allocation to the model. What is the result of this change?
A.
The model generates shorter responses and additional source citations.
B.
The model generates longer responses.
C.
The speed and scale at which the model deployment can process inputs changes
D.
The Azure region availability of the model changes.
Tokens per minute, or TPM , is a throughput allocation for a model deployment. Assigning a higher TPM allocation increases the amount of token processing capacity available to the deployment, which affects the speed and scale at which the deployment can process requests.
It does not directly make responses longer, change region availability, or automatically add citations.
Contribute your Thoughts:
Chosen Answer:
This is a voting comment (?). You can switch to a simple comment. It is better to Upvote an existing comment if you don't have anything to add.
Submit