AI Vision
Point it at an image and ask what is going on
AI Vision reads images the way chat reads text: describe a scene, extract the writing on a label, compare two screenshots, or generate accurate alt text at scale.
What is the AI Vision?
AI Vision is image understanding rather than image creation. You upload a photo, a screenshot, a diagram or a scan and ask questions about it. It can describe what is present, read visible text, identify categories and attributes, compare two images and explain what differs between them.
It differs from the AI Document Analyzer in what it is optimised for. The document analyzer is built for structured documents: contracts, reports, statements, multi-page PDFs where layout, sections and cross-references matter. Vision is built for pictures: product photos, screenshots, whiteboards, receipts, charts, UI captures, anything where the meaning is visual.
It is strong at description, text extraction from reasonably clear images, categorisation and explaining charts. It is weaker at precise counting, exact measurement, fine print on poor scans, and any judgement that depends on knowledge outside the frame.
Capabilities
What you can do with the AI Vision
How it works
Using the AI Vision
- 01
Upload the image
Photos, screenshots, scans and diagrams all work. Higher resolution matters most when text extraction is involved.
- 02
Ask a specific question
"What is the return period stated on this receipt?" gets a better answer than "tell me about this image".
- 03
Follow up
Vision runs inside a conversation — drill into a detail without uploading the image again.
- 04
Use the output downstream
Send extracted attributes into product copy, or push descriptions into alt text and metadata.
Examples
What good input and output look like
AI Vision
Live demoWriting the prompt1/2You upload + type
Source image attachedUploaded photo: a mechanic in a green apron truing a bicycle wheel on a workstandAmmarAI writes
Alt text generation
Input
Image attached. Task: write accessible alt text under 25 words.
Output
"A mechanic in a green apron truing a bicycle wheel on a workstand in a daylit workshop, spoke wrench in hand."
Chart explanation
Input
Screenshot attached. Question: what is the story in this chart?
Output
Revenue rose in each of the first three quarters and dipped in Q4. The Q4 decline is roughly the size of the Q2 gain, so the year ends close to where Q3 finished…
Key features
What the tool gives you
Scene description
Detailed natural-language description of what an image contains, at whatever depth you ask for.
Text extraction
Pull written content out of photos, screenshots and signage.
Attribute tagging
Return consistent structured attributes across a batch of product images.
Comparison
Explain the differences between two versions of a design, a screenshot or a photo.
Accessibility support
Generate alt text that describes function and content rather than restating the file name.
Who it is for
People who get the most from this
E-commerce teams
Tag and describe large product image libraries consistently instead of by hand.
Accessibility and content teams
Produce alt text for an entire image library in a fraction of the time.
Support teams
Read a customer's screenshot and understand the error before replying.
Analysts
Get a written reading of a chart or dashboard capture to paste into a report.
Workflows
Practical ways teams use it
Bulk alt text pass
Run the site's image library through vision, generate descriptive alt text, review the edge cases, and publish.
Catalogue enrichment
Extract colour, material, style and shape from product photos, then feed those attributes into the product description generator.
Screenshot triage
Have support paste a customer screenshot, extract the visible error text, and route the ticket correctly.
Tips that improve results
- Ask one question at a time when precision matters; compound questions get compound vagueness.
- Upload the highest resolution you have when text extraction is the goal.
- For batch work, define the exact output shape you want, such as a fixed list of attributes.
- Verify counts and measurements yourself. Approximate quantity judgement is a known weakness.
- Write alt text prompts around purpose: what does a reader need to know about this image?
Mistakes worth avoiding
- Using it for multi-page structured documents, where the document analyzer is the right tool.
- Trusting exact counts of objects in a busy image.
- Uploading a blurry photo of small print and treating the extraction as reliable.
- Asking about things outside the frame, such as who took the photo or when.
AI Vision FAQ
How is AI Vision different from the AI Document Analyzer?
Vision answers questions about pictures: photos, screenshots, diagrams. The document analyzer handles structured multi-page documents where sections, tables and cross-references matter.
Can it read text in images?
Yes, reliably on clear images. Low-resolution scans, unusual fonts and dense small print reduce accuracy, so verify anything important.
Is it good for alt text?
It is one of the best uses. Ask for description focused on content and purpose, then review for context only you know.
Can it identify people?
It will not identify specific individuals. It describes what is visible, such as a person's activity or clothing, without naming them.
Can I process many images at once?
Batch workflows are supported on higher plans, which is where catalogue tagging and library-wide alt text become practical.
Related
Tools that pair well with this
AI Document Analyzer
Chat with uploaded documents — PDF, Word, CSV — and get summaries, answers with citations, and extracted structured data.
AI ChatAI Chat
A conversational AI that keeps context, answers questions, researches topics, and can hand work off to other tools in the platform. Supports multi-model chat and document-based conversations.
AI ImageAI Image Generator
Generate high-quality images from text prompts. Create product shots, marketing visuals, social graphics, and more — then edit, upscale, or vary them.
AI E-commerceAI Product Description Generator
Turn a handful of product specs into persuasive, on-brand product page copy.
AI TranscriptionAI Transcription
Accurately transcribe audio and video files into text with speaker labels and timestamps. Supports multiple languages and common audio formats.
Try the AI Vision free
One AI for everything you create. Start on the free plan and upgrade only when the volume demands it.