THE MINDS
SECTION 01
ISSUE 001
Visual Reasoning Leaves the Caption Box
Observed now: Built-in computer use lets Gemini 3.5 Flash interpret interface state and take action across browser, mobile, and desktop environments; Gemini Robotics-ER 1.6 adds spatial planning, success detection, tool calls, and instrument reading. Visual evidence is becoming operational without implying that every diagram, interface, or physical workflow is solved.
Why this idea is here
What the evidence establishes.
Google documents built-in computer use in Gemini 3.5 Flash and embodied spatial reasoning, planning, success detection, tool calls, and instrument reading in Gemini Robotics-ER 1.6. The sources establish current visual-action capability, not universal visual-workflow competence.
Source ledger
Read the sources.
- S01Computer use in Gemini 3.5 Flash
official model release / published 2026-06-24 / retrieved 2026-07-10
- S02Gemini Robotics-ER 1.6
official research release / published 2026-04-14 / retrieved 2026-07-09