Google’s Managed Agents for the Gemini API have a new Antigravity preview, and developers using the original May preview have a near-term migration deadline. Google’s official lifecycle documentation lists antigravity-preview-09-2026 as released on September 17, 2026. The earlier antigravity-preview-05-2026 is scheduled to shut down on October 5, 2026, with the September preview as the recommended replacement.
This matters because Managed Agents are not just another Gemini model endpoint. A single interaction can launch a Google-hosted Linux environment where an agent can reason, run code, work with files, search and fetch web content, and continue work across interactions. The current Antigravity agent is built with Gemini 3.8 Flash by default and is available through the Gemini Interactions API and Google AI Studio.
Source check: September 24, 2026. This guide is based on Google’s current Gemini API documentation, model lifecycle pages, and the official Antigravity SDK announcement. It does not claim independent hands-on benchmarks where Google has not published comparable data.
September 23 update: Antigravity SDK now runs local AI agents offline
Google has added local-model support to the Antigravity SDK, creating a second execution path alongside its hosted Managed Agents stack. The initial optimized path uses Gemma 4 26B A4B through LiteRT, and Google says local agentic workflows can run completely offline. Google recommends a machine with more than 24 GB of VRAM or unified memory for the documented setup.
The practical difference is architectural: teams no longer have to choose only between a fully hosted agent runtime and building a separate local orchestration layer. Antigravity can now keep code and prompts on-device for privacy-sensitive or offline work, while still supporting hybrid patterns where a cloud model such as Gemini 3.8 Flash plans a task and local Gemma workers execute the heavier steps. Google also documents plug-and-play support for OpenAI-compatible local servers such as Ollama, LM Studio, and vLLM through LocalOpenAIAgentConfig.
This does not replace the hosted antigravity-preview-09-2026 Managed Agent or change its current preview pricing. It expands the Antigravity SDK into local and hybrid execution, which is especially relevant for proprietary code, compliance-sensitive environments, rate-limit avoidance, and workloads where cloud token cost is a concern.
What changed in September 2026?
| Area | Previous state | Current state |
|---|---|---|
| Base agent ID | antigravity-preview-05-2026 |
antigravity-preview-09-2026 |
| Lifecycle | May preview | September preview released September 17 |
| Old-preview shutdown | No immediate deadline at launch | October 5, 2026 for the May preview |
| Default underlying model | The May launch began on Gemini 3.5 Flash and later moved to 3.6 Flash | Gemini 3.8 Flash |
| Access | Gemini API / AI Studio preview | Gemini API and AI Studio, free and paid projects |
| Compute billing | Preview environment compute not billed | Still not billed during preview; model tokens and tool usage drive paid usage |
The most important action item is simple: if production or test code still specifies antigravity-preview-05-2026, plan the move now rather than waiting for the October 5 shutdown.
What is a Gemini Managed Agent?
Managed Agents let developers delegate the execution layer to Google instead of assembling a model, tool loop, sandbox, file layer, and web retrieval stack independently. The default Antigravity agent can provision a remote Linux environment and iterate through planning, tool execution, observation, and follow-up reasoning until a task finishes or hits a configured limit.
Google currently documents these built-in capabilities:
- Code execution: Bash, Python, and Node.js commands, including installing packages and running tests.
- File work: read, write, edit, search, and list files inside the environment.
- Web retrieval: Google Search and URL Context are available to the agent.
- Long-running tasks: background execution is supported for workflows that take minutes.
- Context compaction: the agent can compact context during long sessions rather than simply failing at a normal chat-style context boundary.
- Custom functions and remote MCP: developers can add their own application tools or connect supported remote MCP servers.
That makes Managed Agents closer to a hosted execution runtime than a conventional text-generation request. If you are evaluating the broader protocol layer around agent tools, see our Model Context Protocol (MCP) guide. For a different Gemini real-time developer surface, see Gemini 3.8 Live API pricing and model IDs.
Current agent ID and supported models
For the base agent, Google’s current documentation uses:
antigravity-preview-09-2026
The default underlying model is gemini-3.8-flash. Google also currently allows these values through agent_config.model:
| Model | API value | Typical reason to choose it |
|---|---|---|
| Gemini 3.8 Flash | gemini-3.8-flash |
Current default for balanced reasoning, coding, and tool use |
| Gemini 3.7 Flash | gemini-3.7-flash |
Previous-generation compatibility or controlled comparisons |
| Gemini 3.6 Flash | gemini-3.6-flash |
Older Flash baseline for agentic workflows |
| Gemini 3.5 Flash | gemini-3.5-flash |
Lighter general workflows |
| Gemini 3.5 Flash-Lite | gemini-3.5-flash-lite |
Latency- and cost-sensitive workloads |
For a named managed agent created with agents.create, Google says the model is fixed in the agent definition rather than overridden on each interaction. That is useful for predictable debugging and security boundaries, but it also means model choice should be treated as deployment configuration rather than a casual per-request switch.
Pricing: there is no single flat “agent request” price
Google describes Managed Agents as pay-as-you-go based on the underlying Gemini model tokens and the tools the agent uses. Free-tier projects also receive a free rate limit and usage quota. During the preview period, Google says environment compute — the CPU, memory, and sandbox execution layer — is not billed.
The important budgeting difference versus ordinary chat is that an agent interaction can create many reasoning and tool-use steps. Google says a single managed-agent interaction typically consumes roughly 100,000 to 3 million tokens, with complex workflows sometimes reaching 3–5 million tokens.
Google’s current documentation gives these task-level estimates based on its own runs:
| Task type | Google’s typical cost estimate |
|---|---|
| Research and information synthesis | $0.30–$1.00 |
| Document and content generation | $0.30–$1.30 |
| Process and system design | $0.25–$0.80 |
| Data processing and analysis | $0.70–$3.25 |
Google also notes that especially complex workflows can approach about $5. These are vendor estimates, not guaranteed prices for your application. Real cost depends on model selection, input size, agent loop length, tool activity, cache behavior, and how aggressively the task is bounded.
Use token budgets before you trust an autonomous loop
The current Antigravity agent supports max_total_tokens inside agent_config. The cap covers input, output, and thinking tokens; cached tokens do not count toward that limit. When the limit is reached, the interaction can return as incomplete, and developers can continue from the preserved state with another interaction.
For production design, this is more important than treating the agent like a normal completion endpoint. A useful control stack is:
- choose the cheapest model that reliably completes the task;
- set a maximum token budget per interaction;
- stream or monitor progress for long jobs;
- cancel runs that are no longer economically useful;
- measure cost by completed business outcome rather than by API call count.
Sandbox networking is a security boundary, not a convenience setting
Google-hosted isolation reduces the risk of running arbitrary agent code on your own application server, but it does not remove agent security concerns. Google’s agents documentation says remote environments can have outbound network access and recommends explicit network controls for sensitive workloads.
In practice, an agent that can browse, run code, install packages, read files, and call external services should be treated as a privileged workload. For production use:
- restrict outbound access to the domains the workflow genuinely needs;
- use least-privilege API keys or service accounts;
- prefer short-lived credentials where possible;
- separate testing credentials from production credentials;
- require human review before high-impact side effects;
- log the agent’s external calls and changes to important systems.
Google’s managed-credentials design can inject secrets through an egress proxy so the secret does not need to appear directly inside the sandbox or model context. That is a better pattern than dropping a long-lived production key into a prompt or workspace file.
For a broader threat model covering prompt injection, excessive permissions, runtime isolation, and connector risk, see our AI Agent Security guide.
Persistent environments: useful, but understand the lifecycle
Managed Agents can preserve files and state between interactions by reusing an environment. Google currently says environments are deleted after seven days of inactivity. This makes multi-step work practical, but it also means teams should decide deliberately what data is allowed to persist in an agent workspace and what must be removed sooner.
Managed Agents are still in preview. Google’s current custom-agent documentation also notes several platform limits, including a maximum of 1,000 managed agents, no agent versioning/rollback yet, and no nested subagent delegation for named managed agents. Preview schemas and capabilities can change, so production integrations should isolate API-specific assumptions behind a small adapter layer.
Migration checklist: May preview to September preview
- Find every hard-coded base agent ID. Search application code, infrastructure configuration, tests, examples, and stored templates for
antigravity-preview-05-2026. - Change to
antigravity-preview-09-2026. Google’s lifecycle page names this as the recommended replacement before the October 5 shutdown. - Retest tool behavior. Preview agents can change. Regression-test code execution, filesystem operations, Search/URL Context, custom functions, remote MCP, streaming, and background jobs that your product depends on.
- Retest model-dependent behavior. The current default is Gemini 3.8 Flash. If consistent behavior matters more than following the default, explicitly configure a supported model and maintain your own evaluation set.
- Check budget controls. Add or review
max_total_tokensfor long-running tasks so a regression cannot quietly turn into runaway usage. - Review network policy and credentials. Migration is a good moment to replace broad outbound access and long-lived secrets with allowlists and least-privilege managed credentials.
- Run representative end-to-end tasks before cutover. Measure completion quality, latency, token usage, tool failures, and cost on your actual workflows.
Managed Agents vs building your own agent runtime
| If you need… | Managed Agents are attractive when… | A custom runtime may fit better when… |
|---|---|---|
| Sandbox execution | You want Google-hosted Linux environments without operating the execution layer | You require custom kernels, networking, hardware, or compliance controls |
| Tool loop | You want a prebuilt Antigravity harness | You need complete orchestration control |
| State | Persistent files and resumed interactions are enough | You need your own durable state model and lifecycle |
| Model choice | The supported Gemini Flash family fits the workload | You need multi-provider routing or unsupported models |
| Operations | You prefer managed infrastructure | You need deep observability or custom runtime policy enforcement |
The practical question is not whether “managed” is universally better. It is whether outsourcing the execution environment and harness removes enough operational work to justify accepting a preview API and Google’s current runtime constraints.
Who should care about this update?
Existing Managed Agents users have the clearest action: migrate off the May agent ID before October 5. Teams evaluating agent infrastructure now have a current Gemini 3.8 Flash-based managed option with free-tier access and explicit token-budget controls. Security-sensitive teams should focus less on the convenience of “one API call” and more on network policy, credentials, human approval, and what the agent can modify once it is running.
If you are comparing costs across Google’s consumer subscriptions rather than developer APIs, use our separate Gemini pricing guide. Managed Agents are an API/runtime product and should not be confused with Gemini app subscription limits.
FAQ
What is the current Gemini Managed Agents Antigravity ID?
Google’s current documentation uses antigravity-preview-09-2026.
When does the old Antigravity preview shut down?
Google’s lifecycle page lists October 5, 2026 as the shutdown date for antigravity-preview-05-2026, with the September preview as its replacement.
Which model does the current Antigravity agent use?
The current default is Gemini 3.8 Flash. Google also documents configurable 3.7 Flash, 3.6 Flash, 3.5 Flash, and 3.5 Flash-Lite options.
Are Gemini Managed Agents free?
They are available to free-tier projects with a free rate limit and quota. Paid usage is pay-as-you-go based on Gemini model tokens and tool usage. Google says environment compute is not billed during the preview period.
Can Managed Agents use MCP?
Yes. Google documents remote MCP server support, subject to transport and naming limitations in the current preview. For the protocol itself, see our MCP guide.
Are Managed Agents generally available?
No. Google currently labels Managed Agents and the Antigravity agent as preview products, so features and schemas can change.
Primary sources
- Google Developers Blog — Local AI models in the Antigravity SDK (September 23, 2026)
- Google AI for Developers — Antigravity agent
- Google AI for Developers — Agents overview
- Google AI for Developers — Building Managed Agents
- Google AI for Developers — Gemini deprecations and lifecycle
- Google — Introducing Managed Agents in the Gemini API
- Google — Managed Agents: 3.6 Flash, hooks, budget controls, and free tier
AI-XBlog Weekly Brief
Keep up with AI that actually works
Join the AI-XBlog Weekly Brief for major AI updates, practical workflows, useful tools, and editor’s picks. No daily noise.
Reader discussion
Join the discussion
Have you tried this tool or workflow? Share your experience, corrections, or questions. Useful reader feedback may help us improve this article.
All comments are reviewed before publication. Your email address will not be published. Promotional links and low-value spam are removed.
