Camera Input & Identification

The Camera Input agent lets you use your phone camera as a direct input during a guided task. Instead of describing what you see, point the camera at the work and Atanzo AI reads the frame — confirming a step, identifying a part or label, assessing a condition, or surfacing relevant knowledge from your content library.

What It's For

  • Confirm that a step has been completed correctly based on what the camera sees
  • Identify a part, component, or label against your uploaded reference images
  • Surface relevant documentation or instructions tied to the identified item
  • Assess a visible condition (e.g., wear, damage, orientation) and flag it if needed

You stay in the flow of the task. There is no separate lookup step — the camera feeds directly into the guidance you're given.

When to Use It

Reach for camera input whenever what you're looking at matters more than what you can say about it: identifying a part from its markings, confirming a label matches the right unit, or checking a condition that's easier to show than describe. It also saves time on any step where searching for the right reference document would otherwise slow you down.

Starting a Camera Session

Camera input runs as its own session — you start it as its own agent rather than switching into it partway through a voice or text session.

  1. From your list of agents, choose the camera-based option (shown on screen as Video Task).
  2. Turn on your camera when prompted. You'll see a live preview once it's active.
  3. Review your plan, then start the task.
  4. Point the camera at the work as you go — the agent reads each frame in context with your current step.

You can turn the camera on or off at any point within that session, but the session itself starts as a camera session.

Visual Identification

When your Media Manager contains reference images — part photographs, diagrams, label photos, equipment placards — the system uses those images to recognise what the camera sees. There is no manual lookup step; the right reference is surfaced automatically during the task.

Example: you scan a serial plate on a piece of equipment. The system identifies the unit against your uploaded part images and immediately surfaces the relevant section of the service manual. You don't need to search; the match is surfaced in context.

To make visual identification work well:

  • Upload clear, well-lit reference images for each part or item you're likely to encounter
  • Include multiple angles or variants where applicable (e.g., worn vs. new, different mounting orientations)
  • Use descriptive filenames and tags when uploading so the system has context alongside the image

See Media Manager for instructions on uploading and organizing reference images.

Picking the Right Item Out of a Cluttered View

When multiple components are visible in frame, Atanzo AI focuses on the relevant item rather than treating the whole image as a single input. This is useful in environments where parts are tightly packed or the background is visually busy — you don't need to isolate the item by hand before the camera can identify it.

How to Tell It's Working

While the camera is active you'll see a live preview and a brief indicator each time a frame is captured and analyzed. If the agent confirms a step, identifies a part, or surfaces a document without you asking, that's the camera input working as intended.

If identification doesn't match what you're pointing at:

  • Check that reference images for that part exist in your Media Manager and are well-lit and in focus
  • Try a different angle — square-on shots of labels and plates work best
  • If the camera shows as off or the preview is blank, turn it back on — the agent can't see your task with the camera off

Accessing Camera Input

Camera input is available when your plan includes it. If the camera-based agent doesn't appear in your list of agents, contact your account administrator to confirm your access.

Camera Permissions

The browser must be allowed to access the camera. On first use, the browser will display a permission prompt. If you deny it or need to re-enable it:

  • Chrome (Android/desktop): tap the lock icon in the address bar and set Camera to Allow
  • Safari (iOS): go to Settings > Safari > Camera and set to Allow or Ask
  • Other browsers: consult your browser's site permissions settings

Camera permission is per-site. Granting permission once is typically remembered for future sessions on the same device.

  • Choosing an Input Mode — comparison of voice, text, and camera modes (plus smart glasses, coming soon)
  • Media Manager — uploading and managing reference images for visual identification