Module 2 · Building Agents · scripted

Exercise: The Agent's Eyes

25 min

The goal: let the agent see what you see

Your first agents drove the conversation and you reported back in words. Now upgrade the channel: in this exercise you answer the agent with photographs. The Agent Lab's input bar has a 📷 button — attach an image from your computer or phone (on a phone, it opens the camera), and it is sent into the conversation as part of your message. The agent sees it.

The exercise at a glance

  1. 1
    Use a picture to get a task done.
  2. 2
    Use a picture to get feedback on your work.
  3. 3
    Turn a picture into a simulation.

The exercise playground

The playground is the same Agent Lab you used for your first agents — everything you built earlier is still there (your saved agents and conversations included). The one new thing you need: the input bar's 📷 button attaches an image to your message. The Prompts view works too: an attached image shows up in the log as [image attached], one more thing on the growing prompt.

You can also do the exercise in ChatGPT, Claude, or Gemini instead — all of them accept photos. If you're going to use one of those tools, click here for the instructions →.

How the playground works

  1. 1
    Open it: /exercises/agent-lab — or scan the QR code on the next card.
  2. 2
    Attach photos with the 📷 button in the input bar.
    On a phone it opens the camera; on a computer it picks a file.
  3. 3
    The photo goes into the conversation as part of your message — the agent sees it.

Scan to open the Agent Lab

Scan to open the Agent Lab
ai-agents-seminar.vercel.app/exercises/agent-lab

Step 1 · Use a picture to get a task done

Step 1 · What you'll do

  1. 1
    Take a picture of something around you — your desk, a bookshelf, the contents of a bag, a schedule on your screen.
  2. 2
    Send it with a task: "plan the cleanup" · "pick my next three reads and an order" · "what am I forgetting for this trip?" · "turn this schedule into a plan for tomorrow".
  3. 3
    Watch for the details it uses that you never typed — that's the picture doing work words would have dropped.

The picture is the context; the task tells the agent what the context is for. The test of a good pairing: the agent's answer uses things from the image you would never have thought to mention.

Capture — end of Step 1

💾 Save this before you move on

Copy into your course document and save:

  1. 1The photo and the task you gave with it.
  2. 2One thing the agent used from the image that you had not mentioned — and would not have typed.

Step 2 · Use a picture to get feedback on your work

Step 2 · What you'll do

  1. 1
    As a group, draw a plan or an idea — on paper or a whiteboard — or screenshot something you're working on.
  2. 2
    Photograph it and send it three times, changing the role:
    · "Act as a skeptic. Poke holes in this — how does it fail in ways we haven't thought of?" · "What are the gaps and ambiguities in this?" · "What are the critical questions we should be asking?"
  3. 3
    Answer its follow-up questions and see how the critique sharpens.

This is the whiteboard move from the lesson, aimed at your own work. The same drawing, prompted three ways, gives you a skeptic, a gap-finder, and a facilitator — and a critique session your group can actually argue with.

Capture — end of Step 2

💾 Save this before you move on

Copy into your course document and save:

  1. 1The drawing you photographed.
  2. 2The strongest single criticism or question the agent produced — the one your group had not thought of.

Step 3 · Turn a picture into a simulation

Step 3 · What you'll do

  1. 1
    Draw a process diagram or a user interface sketch and photograph it. (A process from your research works well.)
  2. 2
    Send the photo with the simulation prompt (below).
  3. 3
    Interact with it: step through the process, or "click" around the interface, and see whether the simulation stays true to your drawing.

The simulation prompt

user prompt
Act as the system in this image. Tell me how I'm allowed to interact with you, then let me interact with you. Simulate what the system does, with outputs.

A drawing of a system is enough for the agent to become the system: a persona-based simulation, generated from a photograph. This is prototyping with a pen — you can test a process or an interface before anything exists.

Before we regroup

💾 Save this before you move on

Keep one document for this course. Copy each item into it and save — we will compare the best of each, and the best paper-drawn system of the session.

Copy into your course document and save:

  1. 1The diagram from Step 3 and your opening prompt.
  2. 2One exchange from the simulation — and one place where it followed your drawing exactly, or departed from it.

Be ready to discuss:

  • The photos went into the conversation like any other message. What does that mean for the cost of a conversation full of photos — and for what should happen to old ones?
  • Where in your research would an agent's eyes replace a measurement you currently type in by hand?