Multimodal Prompt Optimization
Independent PiSkill directory guide. The original prompt remains hosted by Google Cloud Generative AI.
What does this prompt do?
Optimizes Gemini prompts that combine text with visual or other multimodal inputs and evaluates the resulting structured task behavior.
Primary use case
Improve a prompt that interprets text and multimodal evidence together.
Expected output
An optimized multimodal prompt with evaluation results.
Inputs or variables
- prompt
- multimodal examples
- target output
Related prompts
AI-Assisted Data Science
Uses Gemini to explore a data question, inspect a dataset, develop analysis steps, produce code, and explain findings with assumptions and validation checks.
Gemini and Document AI Entity Extraction
Combines Document AI and Gemini to extract defined entities, reconcile OCR context, and assess whether structured results match the underlying document.
Gemini Document Processing
Processes documents with Gemini by defining the target schema, extracting grounded information, handling multiple formats, and validating outputs against source evidence.
Gemini Controlled JSON Output
Uses controlled generation to produce responses that conform to a declared JSON schema.
Gemini Text Classification
Shows reusable prompt patterns for assigning text to a controlled set of labels with Gemini.
Gemini Text Extraction
Demonstrates prompts that extract named fields and facts from unstructured text into a predictable result.