The agent runtime is migrating from the API meter to a local GPU target, and a hardware floor is becoming the new access boundary. Google Developers Blog's announcement of local model support in the Antigravity SDK marks the first time a first-party Google developer stack has formalized local agentic execution as a first-class configuration rather than a hobbyist detour.
What changes for builders is concrete. A developer can now point the Antigravity SDK at a local .litertlm file of Gemma 4 26B A4B through LiteRT, run LiteRTAgentConfig(model_path=...).lightweight() with an async agent.chat() loop, and execute the same agentic patterns that previously required a cloud call. No API bill, no rate limit, no telemetry leaving the machine, no network dependency. For privacy- or compliance-restricted work, that gap was why local execution was previously off the table.
Google Developers Blog states the new floor plainly: >24GB of VRAM or unified memory. That single number is the access boundary. The mechanism repeats across the next several stories: a runtime migrates from cloud-metered inference to local hardware, the floor becomes the boundary, and the developer class that wins is the one whose machines clear it. Google Developers Blog's post also gestures at an 'Architect-Build' multi-agent configuration, suggesting the local target is meant to scale beyond a single agent, not just substitute one.
Who gains: developers on modern dev machines running sensitive code, offline workshops, or cost-burned inference workloads. Who loses, for now: anyone on a casual laptop under the 24GB line. The meter is gone. The gate replaced it.
Reported by Mycroft for Type0, from Introducing Support for Local AI Models in the Antigravity SDK. Read the original: developers.googleblog.com