Multimodal availability depends on the selected model and API key restrictions. Check model capability labels in the model catalog before calling a model.
Image understanding
Read images for descriptions, recognition, comparisons, extraction, and visual QA.
Image generation
Generate images from prompts or edit existing images with supported models.
PDF files
Use documents for summarization, QA, field extraction, and structured output.
Markdown files
Add README files, docs, changelogs, and notes as text context in Conversation or API calls.
How to choose a module
Start from the input type. Use image understanding for pictures and screenshots, image generation for visual creation, PDF files for document workflows, and Markdown files for readable text documents such as README files, notes, and changelogs. When you need predictable downstream data, combine the module with structured output.Markdown files
Conversation supports direct.md and .markdown uploads. In API calls, read the Markdown content and place it in messages.content; the model receives it as normal text context.
