AI Vision

Point it at an image and ask what is going on

AI Vision reads images the way chat reads text: describe a scene, extract the writing on a label, compare two screenshots, or generate accurate alt text at scale.

What is the AI Vision?

AI Vision is image understanding rather than image creation. You upload a photo, a screenshot, a diagram or a scan and ask questions about it. It can describe what is present, read visible text, identify categories and attributes, compare two images and explain what differs between them.

It differs from the AI Document Analyzer in what it is optimised for. The document analyzer is built for structured documents: contracts, reports, statements, multi-page PDFs where layout, sections and cross-references matter. Vision is built for pictures: product photos, screenshots, whiteboards, receipts, charts, UI captures, anything where the meaning is visual.

It is strong at description, text extraction from reasonably clear images, categorisation and explaining charts. It is weaker at precise counting, exact measurement, fine print on poor scans, and any judgement that depends on knowledge outside the frame.

Capabilities

What you can do with the AI Vision

Describe an image in detail, including composition and context
Extract visible text from photos, screenshots and signs
Generate accurate alt text for accessibility at scale
Categorise and tag product images with consistent attributes
Explain a chart, diagram or whiteboard photo in words
Compare two images and articulate what changed
Check images against simple rules, such as whether a logo is present

How it works

Using the AI Vision

  1. 01

    Upload the image

    Photos, screenshots, scans and diagrams all work. Higher resolution matters most when text extraction is involved.

  2. 02

    Ask a specific question

    "What is the return period stated on this receipt?" gets a better answer than "tell me about this image".

  3. 03

    Follow up

    Vision runs inside a conversation — drill into a detail without uploading the image again.

  4. 04

    Use the output downstream

    Send extracted attributes into product copy, or push descriptions into alt text and metadata.

Examples

What good input and output look like

AI Vision

Live demo1/2
You

You upload + type

Uploaded photo: a mechanic in a green apron truing a bicycle wheel on a workstandSource image attachedUploaded photo: a mechanic in a green apron truing a bicycle wheel on a workstand

AI

AmmarAI writes

Alt text generation

Input

Image attached. Task: write accessible alt text under 25 words.

Output

"A mechanic in a green apron truing a bicycle wheel on a workstand in a daylit workshop, spoke wrench in hand."

Chart explanation

Input

Screenshot attached. Question: what is the story in this chart?

Output

Revenue rose in each of the first three quarters and dipped in Q4. The Q4 decline is roughly the size of the Q2 gain, so the year ends close to where Q3 finished…

Key features

What the tool gives you

Scene description

Detailed natural-language description of what an image contains, at whatever depth you ask for.

Text extraction

Pull written content out of photos, screenshots and signage.

Attribute tagging

Return consistent structured attributes across a batch of product images.

Comparison

Explain the differences between two versions of a design, a screenshot or a photo.

Accessibility support

Generate alt text that describes function and content rather than restating the file name.

Who it is for

People who get the most from this

E-commerce teams

Tag and describe large product image libraries consistently instead of by hand.

Accessibility and content teams

Produce alt text for an entire image library in a fraction of the time.

Support teams

Read a customer's screenshot and understand the error before replying.

Analysts

Get a written reading of a chart or dashboard capture to paste into a report.

Workflows

Practical ways teams use it

Bulk alt text pass

Run the site's image library through vision, generate descriptive alt text, review the edge cases, and publish.

Catalogue enrichment

Extract colour, material, style and shape from product photos, then feed those attributes into the product description generator.

Screenshot triage

Have support paste a customer screenshot, extract the visible error text, and route the ticket correctly.

Tips that improve results

  • Ask one question at a time when precision matters; compound questions get compound vagueness.
  • Upload the highest resolution you have when text extraction is the goal.
  • For batch work, define the exact output shape you want, such as a fixed list of attributes.
  • Verify counts and measurements yourself. Approximate quantity judgement is a known weakness.
  • Write alt text prompts around purpose: what does a reader need to know about this image?

Mistakes worth avoiding

  • Using it for multi-page structured documents, where the document analyzer is the right tool.
  • Trusting exact counts of objects in a busy image.
  • Uploading a blurry photo of small print and treating the extraction as reliable.
  • Asking about things outside the frame, such as who took the photo or when.

AI Vision FAQ

How is AI Vision different from the AI Document Analyzer?

Vision answers questions about pictures: photos, screenshots, diagrams. The document analyzer handles structured multi-page documents where sections, tables and cross-references matter.

Can it read text in images?

Yes, reliably on clear images. Low-resolution scans, unusual fonts and dense small print reduce accuracy, so verify anything important.

Is it good for alt text?

It is one of the best uses. Ask for description focused on content and purpose, then review for context only you know.

Can it identify people?

It will not identify specific individuals. It describes what is visible, such as a person's activity or clothing, without naming them.

Can I process many images at once?

Batch workflows are supported on higher plans, which is where catalogue tagging and library-wide alt text become practical.

Try the AI Vision free

One AI for everything you create. Start on the free plan and upgrade only when the volume demands it.