Google's Gemini Deep Research Agent is now accessible via the Interactions API, enabling autonomous long-horizon research tasks that plan, search, and synthesize cited reports. Two model variants are available: a speed-optimized version for streaming to client UIs and a max-comprehensiveness version for automated synthesis. Key new features include collaborative planning (review and refine the research plan before execution), native chart and infographic generation, remote MCP server integration for external tools, extended tooling options (Google Search, URL Context, Code Execution, File Search), and multimodal research grounding with images, PDFs, and audio. Tasks run asynchronously in the background and can be polled for results, with optional real-time streaming of progress and intermediate reasoning.

3m read timeFrom philschmid.de
Post cover image
Table of contents
What's newSetupRun your first Deep Research taskCollaborative planningNative charts and infographicsRemote MCP serversTool configurationMultimodal research groundingReal-time streaming with visuals and thought summariesWhere to go next