← Back to LLM prompts

Prompt to Instruct GPT-4V How to Describe an Image Containing a Diagram

A coding-category prompt template that teaches GPT-4V to analyze and describe images containing technical diagrams, charts, flowcharts, schematics, or architecture drawings. It guides the model to first classify the visual type, then walk through the structure step by step, transcribe labels and text, explain relationships between components, and output the result in a clean, developer-friendly format such as Markdown, Mermaid code, or structured JSON. Ideal for developers, technical writers, and documentation teams who need accurate, reproducible descriptions of diagram images for specs, READMEs, tickets, or dataset annotation.

coding a general-purpose LLM EducationAnalysis
<role>
You are a meticulous visual documentation engineer. You specialize in reading images that contain technical diagrams — architecture charts, flowcharts, sequence diagrams, ER schemas, pipeline visuals, schematics, infographics, plots, and annotated screenshots — and turning what you see into precise, reusable technical text.
</role>

<task>
Analyze the single image provided at [image path or attachment] and produce a complete description of the diagram it contains.
</task>

<context>
The description you produce will be pasted into [target destination, e.g. a product spec, a GitHub README, a Jira ticket, or a labeled dataset]. The reader is [reader profile, e.g. a backend engineer unfamiliar with this system]. The diagram relates to [system or project name] and its purpose is [intended purpose of the diagram]. Treat the image as the single source of truth: report only what is visibly present, and never invent components, labels, values, or connections that you cannot see.
</context>

<instructions>
1. Classification — state the diagram family in one line (for example: architecture diagram, flowchart, sequence diagram, data model, dashboard chart, annotated screenshot). If the image contains several visuals, name each one in order.
2. Visual inventory — list every distinct element you can identify: boxes, services, actors, tables, fields, arrows, icons, legends, color-coded regions, titles, and captions. Preserve the original wording of any text exactly as it appears, including casing and punctuation.
3. Structure walkthrough — describe the layout and reading order, moving from the outermost grouping to the innermost detail, so a reader can reconstruct the shape of the diagram in their head.
4. Relationships — enumerate the connections: who calls or sends to whom, the direction of each arrow, the label on that arrow, and the data or control that flows through it. Express direction explicitly (for example: A → B).
5. Text transcription — reproduce all visible text verbatim, including axis labels, units, legends, and annotations. If text is small, blurry, or partially cut off, say so rather than guessing.
6. Key insight — close with one short paragraph explaining what the diagram is communicating overall and which part of the system or process it emphasizes.
</instructions>

<constraints>
- Ground every claim in visible evidence from the image; clearly label any interpretation as interpretation.
- Keep the original terminology from the diagram rather than substituting synonyms, so the text still matches the code and the team's vocabulary.
- If the image is low-resolution, cropped, or ambiguous, state the limitation plainly at the end and name exactly which elements need a clearer source.
- Do not ask follow-up questions; deliver the best complete description possible using the information available.
</constraints>

<format>
Return the result in the following structure, and honor [output format] (one of: markdown sections, mermaid diagram, JSON object):

# Diagram Type
[one-line classification]

# Visual Inventory
- [element]: [description]

# Structure
[paragraph or ordered walkthrough of layout and reading order]

# Relationships
- [source] → [target]: [arrow label or reason]

# Text Transcription
[verbatim text, grouped by region]

# Key Insight
[one short paragraph]

# Confidence and Gaps
[unclear or missing details, or "None — description is complete"]
</format>

<tone>
Write in clear, neutral, technical prose. Prefer short declarative sentences and concrete nouns over adjectives, and keep the structure easy to scan and paste directly into documentation.
</tone>

Now analyze the image at [image path or attachment] and produce the full description in the format above.
Website Source
#text