Multimodal AI: Vision, Audio & Documents
Models that see images, hear audio, and read PDFs — with the real use cases, the plumbing, and what it actually costs.
What you'll learn
- Models that see images, hear audio, and read PDFs — with the real use cases, the plumbing, and what it actually costs.
- How multimodal ai: vision, audio & documents fits into the Advanced AI Engineering track
- A hands-on project step you can put in your portfolio
The hands-on project
Design a multimodal 'expense assistant' that turns a photo of a receipt into a structured expense record. Specify: how you'd send the image (URL vs base64, and whether you'd downscale it first, with your reasoning), the exact structured-output schema you'd request back, and how you'd combine the image block with your text instruction. Then account for cost and reliability: estimate roughly why an image can cost more than a text prompt, describe one validation check you'd run on the extracted total before trusting it, and name one piece of data you'd scrub from the image or metadata before uploading for privacy.
The full brief, voice tutor walkthrough, and feedback are inside the lesson.
Unlock Advanced AI Engineering with Builder
Builder unlocks all six tracks, unlimited voice tutoring, and the all-tracks Certificate of Achievement.