You have a multimodal model deployed in Microsoft Foundry
You need to configure the model to analyze an image and provide both a written summary and an audio summary of the image.
What is the minimum number of user prompts that should be sent to the model?
Submit