Google’s new AI agent can draft your emails, monitor your inbox and eventually spend your money

Google on Tuesday unveiled Gemini Spark, a personal AI agent designed to work around the clock — drafting emails, assembling documents, monitoring inboxes, and eventually making purchases — even when a user’s laptop is closed and their phone is locked.

The announcement, made at Google I/O 2026, is the company’s most ambitious attempt yet to transform its AI assistant from a tool that answers questions into one that autonomously completes tasks. It also arrives at a moment of extraordinary competition, as Microsoft, OpenAI, Anthropic, and Apple all race to build AI systems that don’t merely converse but act — completing multi-step workflows with decreasing human supervision.

“We are in that part of the cycle where people want to see real value in the products they use on a day-to-day basis,” Sundar Pichai, CEO of Google and Alphabet, said during a press briefing ahead of the keynote address. With Spark, he argued, that value comes from an agent that never stops working. It operates around the clock in Google’s cloud, he said, so “you don’t need to keep your laptop open to make sure it’s running.”

The product arrives at an inflection point for the technology industry, as Google, Microsoft, OpenAI, Anthropic, and Apple all race to build AI systems that don’t merely converse but do — completing multi-step workflows with decreasing human supervision. It also raises urgent questions about trust, spending guardrails, and what happens when an artificial intelligence agent misinterprets a user’s intent.

Spark will begin rolling out this week to a small group of trusted testers, with a beta planned for Google AI Ultra subscribers in the United States next week.

Inside the cloud architecture that lets Gemini Spark work while you sleep

Unlike conventional AI assistants that activate only when prompted, Gemini Spark is architecturally different. It runs persistently on Google Cloud infrastructure, powered by the company’s new Gemini 3.5 Flash model and what Google calls the Antigravity agent harness — the same underlying system that powers the company’s internal developer tools.

In practical terms, this means Spark can accept a complex instruction — “email my boss a status update pulling the latest figures from our shared spreadsheet and the project timeline in our Slides deck” — and then execute it across multiple Google applications without further input. The agent can pull context from emails, documents, and calendar entries, synthesize the information, and produce a finished output.

Josh Woodward, VP of Google Labs, Gemini App, and AI Studio, described the experience in visceral terms during the briefing: “When you use it, it almost feels like you’re tossing things over your shoulder — Spark’s catching them and gets the job done.”

The cloud-based architecture is a deliberate design choice. Because Spark operates on remote servers rather than on a user’s device, it can continue working through tasks after a user walks away. A student could ask Spark to build a study guide that updates itself as new assignments arrive from a professor. A small business owner could instruct it to monitor their inbox and flag potential customer inquiries. A parent could delegate the logistics of a neighborhood block party — tracking RSVPs, coordinating contributions, scouting venues. These are not hypothetical scenarios. Woodward said they reflect how early testers have actually been using the product.

Over the coming months, Google plans to expand Spark’s capabilities significantly. The company will roll out MCP (Model Context Protocol) connections to more than 30 third-party partners, including Canva, OpenTable, and Instacart. Users will also be able to text and email Spark directly, create custom sub-agents for specialized tasks, and connect Spark to Chrome for web-based actions. Later this year, a new Android interface called Android Halo will provide live, at-a-glance visibility into what Spark is working on, displayed at the top of a user’s phone screen.

Google compares its AI spending safeguards to giving a teenager their first debit card

For all its ambition, Spark confronts a fundamental challenge that has bedeviled every AI agent to date: How do you trust an autonomous system to act on your behalf — particularly when money is involved?

Google is acutely aware of the concern. When asked during the press briefing how Spark would avoid making unauthorized purchases, Woodward reached for an analogy that was striking in its candor. “On the team, we think a lot of it is like if you’re giving a teenager their first debit card — there’s sort of limits and sort of constraints around it, and that’s how we’ll be designing Spark as we go through the year,” he said.

At launch, Spark will not autonomously make purchases. Users will be given explicit opportunities to review and approve any transaction before it goes through. But Google has built the infrastructure for a more autonomous future. Vidhya Srinivasan, who leads Google’s ads and commerce teams, introduced the Agent Payments Protocol, or AP2 — a system designed to let AI agents make secure purchases within user-defined boundaries.

The concept works like this: a user tells their agent the specific brands, products, and spending limits they’re comfortable with. If the criteria are met, the agent can automatically complete a purchase. AP2 creates what Google describes as a transparent, verifiable link between the user, the merchant, and payment processors, using privacy-preserving technology and tamper-proof digital mandates to ensure the agent is acting within its authorization. AP2 also generates a permanent digital paper trail, so that if a return is needed, the user and the merchant are looking at the same record. Google plans to bring AP2 to its products in the coming months, starting with Gemini Spark.

The system is underpinned by the Universal Commerce Protocol (UCP), an open-source standard Google announced earlier this year that gives agents and commerce systems a common language across the entire shopping journey. The UCP Tech Council now includes Amazon, Meta, Microsoft, Salesforce, and Stripe — a remarkable coalition that underscores how seriously the industry takes the prospect of agent-driven commerce.

Google also announced the Universal Cart, an intelligent shopping cart that works across merchants and Google services. Users can add items while browsing Search, chatting with Gemini, watching YouTube, or reading Gmail. The cart then works in the background — tracking price drops, surfacing deals based on payment card perks, and even flagging product incompatibilities. The shopping infrastructure is rolling out in the U.S. this summer across Search and the Gemini app, with YouTube and Gmail to follow.

How Google, OpenAI, Microsoft, Anthropic, and Apple are racing to build the definitive AI agent

The announcement lands in the middle of the most intense competitive period in AI history. Google, Microsoft, OpenAI, Anthropic, and Apple are all racing to ship autonomous agents that can do real work — and each is placing a fundamentally different architectural bet on how to get there.

OpenAI recently unified its Operator and deep research capabilities into ChatGPT agent — a system that brings together website interaction, information synthesis, and conversational intelligence. It carries out tasks using its own virtual computer, shifting between reasoning and action to handle complex workflows. The company emphasizes that users remain in control, with ChatGPT requesting permission before taking consequential actions. But the product has faced scrutiny over reliability. OpenAI’s Computer-Using Agent scores 38.1% on OSWorld, the industry benchmark for computer use tasks, while humans score over 72%.

Anthropic launched its Claude Computer Use Agent in research preview in March, giving Claude the ability to see, navigate, and control a user’s desktop — clicking buttons, opening applications, filling spreadsheets, and completing multi-step workflows. Claude Cowork handles tasks autonomously — users give it a goal and Claude works on their computer, local files, and applications to return a finished deliverable. Anthropic has iterated aggressively, recently shipping ten pre-built financial agents and pursuing deep Microsoft 365 integration.

Microsoft introduced Copilot Cowork to move beyond chat and into execution — helping users delegate real tasks and have them completed. Cowork runs in the cloud, meaning users don’t have to worry about closing their laptop. The system is grounded in Work IQ, Microsoft’s intelligence layer that understands organizational data, tools, and structure. The shift moves Copilot from a sidebar helper to an orchestrator of autonomous agents.

Apple is also preparing a revamped Siri for WWDC 2026 that will act as an “always-on agent” capable of handling tasks across apps using personal data. Google’s Gemini models will help power the upgraded Siri through a multi-year deal reportedly costing Apple around $1 billion per year.

The convergence is unmistakable: every major platform is moving from assistants that talk to agents that act. But each is approaching the problem differently. OpenAI’s agent operates primarily through a browser. Anthropic’s works directly on a user’s desktop. Microsoft’s is tightly bound to the Office 365 ecosystem. Apple’s emphasizes on-device processing and privacy. Google’s approach with Spark is distinctive in its bet on cloud persistence and deep integration with its own services. 

Rather than controlling a user’s screen pixel by pixel, Spark works through structured integrations — Google’s own Workspace APIs, and increasingly, third-party connections through MCP. The advantage is reliability and speed: structured tool use is far more predictable than screen-reading. The disadvantage is that Spark, at least initially, can only act within the systems it’s been connected to.

The AI model behind Spark processes trillions of tokens a day — and Google says it could save enterprises billions

Spark’s capabilities are inseparable from the model that drives it. Gemini 3.5 Flash, also announced Monday, is Google’s new workhorse AI model — designed specifically for the demands of agentic workflows.

The performance claims are important. Google says 3.5 Flash outperforms its previous frontier model, Gemini 3.1 Pro, across nearly all benchmarks, while running four times faster than comparable frontier models in terms of output tokens per second. An even more optimized version, available within Google’s Antigravity development platform, runs twelve times faster.

Pichai framed the economics bluntly. Companies processing roughly one trillion tokens per day on Google Cloud — a figure he said top enterprise customers are hitting — could save over $1 billion annually by shifting 80% of their workloads to a mix of Flash and frontier models like 3.5 Pro. In a market where, as Pichai noted, CIOs are already “blowing through their annual token budgets and it’s only May,” the cost argument may matter as much as the capability argument.

Internally, Google’s own developers have been consuming Gemini 3.5 Flash at a staggering and rapidly accelerating pace. In March, Google was processing about half a trillion tokens per day internally. That figure has since grown to more than three trillion — doubling roughly every few weeks. Pichai described this as a “powerful feedback loop” that continually improves the model.

Koray Kavukcuoglu, CTO of Google DeepMind and Chief AI Architect for Google, said the model’s speed is what makes agentic use cases practical. “3.5 Flash is especially good when deploying multiple agents simultaneously and completing long-running tasks,” he said during the briefing, adding that Google had successfully tested agents building “a working operating system entirely from scratch.”

The 3.5 Pro model, the more powerful sibling, is currently being tested internally and will roll out next month.

What Gemini Spark costs and where it fits in Google’s new subscription tiers

Gemini Spark will be available to Google AI Ultra subscribers. The company is simultaneously restructuring its subscription tiers to make the technology more accessible. A new Ultra plan at $100 per month provides a 5x higher usage limit than the Pro plan, along with priority access to Antigravity and 20TB of cloud storage. The top-tier Ultra plan drops from $250 to $200 per month, with a 20x higher usage limit and access to the full suite of capabilities.

Both tiers include Gemini Spark, the Daily Brief agent — a proactive morning digest that triages email, calendar, and tasks overnight — and access to the new Gemini Omni and 3.5 Flash models. The pricing positions Spark as a premium product — more expensive than Anthropic’s Claude Pro at $20 per month, but comparable to the higher tiers of competing products like Claude Max ($100–$200/month) and OpenAI’s ChatGPT Pro ($200/month).

Why privacy, reliability, and ecosystem lock-in could undermine Google’s agent ambitions

The risks are real and multidimensional.

Reliability remains the industry’s greatest challenge. Even the best AI models hallucinate, misinterpret instructions, and make errors that a human would never make. An agent that drafts an email to the wrong person, misreads a spreadsheet figure, or sends a payment to the wrong merchant could create consequences that are difficult to reverse. Google’s approach of requiring explicit approval for high-stakes actions like spending money or sending emails is a sensible safeguard — but it also limits how autonomous the agent can actually be. An agent that asks for confirmation at every turn isn’t much of an agent at all.

Privacy is another concern. Spark’s ability to synthesize information across a user’s entire Gmail inbox, calendar, documents, and chat history means it has an extraordinarily deep view of a person’s digital life. Google says Spark operates on a fully managed, secure runtime with isolated ephemeral virtual machines, encrypted credentials, and Data Loss Prevention policies. But the concentration of personal context in a single AI system — accessible through natural language — creates a surface area that will attract scrutiny from regulators, privacy advocates, and security researchers.

Market timing is uncertain, too. The consumer appetite for always-on AI agents is unproven at scale. Google says the Gemini app has 900 million monthly users, but it’s unclear how many of those users are ready for the conceptual leap from “ask a question, get an answer” to “delegate a task, trust the outcome.” The history of digital assistants — from Clippy to early Siri to Alexa — is littered with products that promised proactive intelligence and delivered frustration.

And then there is the question of ecosystem lock-in. Spark works best within Google’s own services. While MCP connections to third-party apps will broaden its reach, the initial experience is one of deep Workspace integration. For the billions of people who live inside Google’s ecosystem, this is a natural fit. For those who split their digital lives across Microsoft, Apple, and other platforms, Spark’s utility will be more limited — at least initially.

Woodward acknowledged as much when asked whether Spark would remain confined to the Google ecosystem. “It’s going to be cross-platform in two ways,” he said — through MCP integrations with third-party apps, and through availability on the web, Android, and iOS, with tasks syncing across devices via the cloud.

The real test for Gemini Spark isn’t whether it can do the work — it’s whether people will let it

Google’s bet with Gemini Spark is that the AI industry’s center of gravity is shifting from models that think to systems that act — and that the company best positioned to win that transition is the one with the most comprehensive set of consumer services to act within. It is a bet backed by enormous infrastructure investment. Google expects to spend approximately $180 to $190 billion in capital expenditure this year — roughly six times what it spent in 2022 — much of it on the AI compute required to run agents like Spark at scale for hundreds of millions of users.

The technology, in other words, is arriving. The models are fast enough, the integrations deep enough, the payment rails secure enough. Google has built a system that can draft your emails, organize your calendar, monitor your inbox, and soon enough, spend your money — all while you sleep.

But the hardest problem in artificial intelligence has never been making a machine capable. It has been making a human comfortable. For two decades, Google’s core promise has been ten blue links and a search box — a transaction built on the assumption that the user is in control. Gemini Spark asks users to renegotiate that relationship entirely, to hand a set of keys to a system that is brilliant, tireless, and still, by its maker’s own admission, best compared to a teenager with a debit card.

Gemini Spark rolls out to trusted testers this week, with a broader beta for U.S. Google AI Ultra subscribers expected next week.

Google just redesigned the search box for the first time in 25 years — here’s why it matters more than you think.

For a quarter century, the Google search box has been one of the most recognizable interfaces in computing: a thin white rectangle, a blinking cursor, a few typed words, and a list of blue links. On Tuesday, Google will formally retire that paradigm.

At its annual I/O developer conference, Google announced a sweeping redesign of the search box itself — the literal text field where billions of queries begin every day — transforming it from a simple keyword input into a dynamic, AI-driven conversation starter that can accept text, images, PDFs, videos, and even open Chrome tabs as inputs. The company is also merging its AI Overviews and AI Mode features into a single, seamless search flow, eliminating the friction that previously forced users to choose between a traditional results page and an AI-forward experience.

Liz Reid, Google’s vice president and head of Search, called it “the biggest upgrade to our iconic search box since its debut over 25 years ago” during a press briefing on Monday.

The announcement arrived alongside a blizzard of other news — new Gemini models, a personal AI agent called Spark, an intelligent shopping cart, a reimagined developer platform — but the search box redesign may prove to be the most consequential. It is the clearest signal yet that Google views the future of its flagship product not as a place where users type fragmented keywords, but as an interface where they hold open-ended, multimodal conversations with an AI system backed by the entire web.

The new search box expands, accepts files, and coaches you on what to ask

The changes show a fundamental shift in how Google expects people to interact with the product that generates the vast majority of Alphabet’s revenue.

The box itself now dynamically expands to accommodate longer, more conversational queries. Where the old interface subtly encouraged brevity — a narrow field suited to two- or three-word keyword strings — the new design invites users to fully articulate complex questions in granular detail. It also now supports multimodal inputs directly. Users can upload images, PDFs, files, and videos, or drag in content from Chrome tabs, right from the main search interface. Previously, some of these capabilities existed in AI Mode, but reaching them required extra steps. Now they sit at the primary entry point.

Google is also deploying what it describes as an AI-powered query suggestion system that “goes beyond autocomplete.” Rather than simply predicting the next word a user might type based on popular searches, the system helps users formulate complex, nuanced queries — essentially coaching them toward the kind of detailed questions that AI Mode handles best.

The new search box is starting to roll out immediately in all countries and languages where AI Mode is available.

Google is merging AI overviews and AI mode into one seamless experience

Perhaps more significant than the box itself is the architectural change happening behind it. Google is unifying AI Overviews — the AI-generated summary panels that appear atop traditional search results — with AI Mode, the more immersive conversational search experience the company launched at I/O one year ago.

Starting Tuesday, this merged experience will be live across mobile and desktop worldwide. A user can type a question, receive an AI Overview alongside traditional results, and then continue directly into a back-and-forth AI Mode conversation to ask follow-up questions — all without navigating to a separate interface.

Reid explained the logic during the press briefing: the new AI search box is “an upgrade of our traditional search box, and so the results take you directly to main search rather than AI mode.” She noted that while some power users actively sought out AI Mode, “for most users, they don’t actually want to have to think about, do they want more of a traditional page or an AI-forward search experience.”

The goal, she said, was to ensure that “for most users, they don’t have to think about where to go, they can just go to the search box they’re familiar with, and it feels like they get the best experience afterwards.”

One billion users and doubling queries reveal how fast search behavior is shifting

Google’s decision to redesign the foundational interface of its most important product did not happen in a vacuum. The company shared a set of usage statistics during the briefing that reveal just how rapidly user behavior is already changing.

AI Mode, which launched in the United States at I/O 2025, has surpassed one billion monthly users in its first year. AI Mode queries have been doubling every quarter since launch. AI Overviews, the lighter-weight AI summaries, now reach more than 2.5 billion monthly users. And overall search query volume hit an all-time high last quarter — a data point the company had previously disclosed on its earnings call.

Sundar Pichai, Google’s CEO, framed these figures as evidence that AI features are additive, not cannibalistic, to search usage. “When people use our AI-powered features in search, they use search more,” he said. He added that he loves “how search has become less about individual queries and feels more like an ongoing conversation, giving users deeper insights and connecting you with the vastness of the web.”

Reid reinforced the point: “It’s not just that people are searching more, it’s that they’re searching differently. They’re fully expressing their questions in granular detail, asking those follow-up questions and searching across modalities.”

Gemini 3.5 Flash gives Google’s AI search the speed it needs to work at scale

Under the hood, the new search experience runs on Gemini 3.5 Flash, Google’s newest AI model, which the company also introduced at I/O. Google upgraded AI Mode’s underlying model to 3.5 Flash to deliver what Reid described as “an even more powerful AI search experience.”

Gemini 3.5 Flash is the workhorse of this year’s announcements. Google claims it outperforms its previous frontier model, Gemini 3.1 Pro, on nearly all benchmarks while running four times faster in output tokens per second than comparable frontier models. Pichai described it as being “in a league of its own in the top right quadrant” of the Artificial Analysis index, which plots intelligence against speed — meaning it delivers near-frontier quality at dramatically lower latency.

That speed matters enormously for search. A conversational AI search experience that feels sluggish would be dead on arrival for a product that serves billions of queries daily. By coupling the redesigned interface with a model optimized for both quality and throughput, Google is attempting to make AI-powered search feel as instantaneous as the old keyword experience — while being dramatically more capable.

Search can now build interactive visuals and custom mini apps on the fly

The redesigned search box is also the gateway to a set of new capabilities that push search far beyond text-based answers. Google announced what it calls “generative UI” — the ability for search to dynamically build custom widgets, interactive visualizations, and even mini applications in real time, tailored to a user’s specific question.

Reid offered a concrete example during the briefing: a user could ask “How do black holes affect space time?” and receive an interactive visual in an AI Overview that brings the concept to life. Follow-up questions would trigger the system to dynamically generate entirely new visuals in real time. This is possible, she explained, because of “a novel real-time code generation system we built in partnership with the Google DeepMind team” that runs on Gemini 3.5 Flash. Generative UI capabilities will roll out to everyone this summer, free of charge.

But Google is going further still. For ongoing tasks — planning a wedding, organizing a move, tracking a fitness routine — users will be able to build what the company describes as customizable, stateful experiences within search, powered by its Antigravity development platform. These require no coding expertise. Users simply describe what they want in natural language, and search builds it. Those experiences will be available in coming months, starting with Google AI Pro and Ultra subscribers in the United States.

AI agents that monitor the web around the clock are coming to search results

The redesign also opens the door to what Google calls “information agents” — AI agents that users can configure directly within search to monitor the web 24/7 for specific conditions and deliver synthesized updates when those conditions are met.

A user could, for example, set up an agent to track market movements in a particular sector with specific parameters. The agent would create a monitoring plan, tap into real-time finance data, and proactively notify the user when conditions are met — complete with links and context for further research. Other use cases include apartment hunting, tracking sneaker drops, or monitoring any topic a user cares about. Information agents will launch first for Google AI Pro and Ultra subscribers this summer.

These agents sit within a much larger strategic pivot that Google articulated throughout the briefing: the company is going all-in on AI systems that don’t just answer questions but proactively take actions on users’ behalf. Beyond search, Google introduced Gemini Spark, a 24/7 personal AI agent that runs on dedicated virtual machines in Google Cloud. It unveiled the Universal Cart, an intelligent cross-merchant shopping cart. It announced the Agent Payments Protocol for agents to make secure purchases. And it expanded its Antigravity developer platform into a full ecosystem for building autonomous AI agents.

Publishers, advertisers, and SEO professionals face a new reality

The redesign raises profound questions for the sprawling ecosystem — publishers, advertisers, SEO professionals — that has been built around the old model of keyword search and blue links.

If users increasingly express their needs as full, conversational sentences rather than fragmented keywords, the entire discipline of search engine optimization will need to evolve. Keyword-density strategies become less relevant when the AI is parsing natural language intent rather than matching strings. Content that answers deep, nuanced questions in authoritative ways becomes more valuable; content engineered to rank for two-word keyword fragments becomes less so.

For publishers, the stakes are existential. AI Overviews already synthesize information from across the web and present it directly in search results, reducing the need for users to click through to source material. The new seamless AI Mode integration deepens that dynamic: users can now get an AI-generated answer and ask multiple follow-up questions without ever leaving the search page. Google has consistently maintained that its AI features drive more traffic to publishers, but the redesign puts that claim under renewed scrutiny as the search results page becomes more self-contained.

For advertisers — who fund the vast majority of Google’s revenue — the shift from keywords to conversations changes the calculus of ad targeting. Conversational queries contain richer intent signals, which could make ad targeting more precise and valuable. But they also create new ambiguities: when a user is in the middle of a multi-turn conversation with AI Mode, where does an ad naturally fit? Google did not detail changes to its advertising model during the briefing, but the structural shift in the interface will inevitably reshape how ads are surfaced and measured.

The search box was always more than a product — it was a habit for billions of people

There is a reason Google chose to redesign the search box rather than simply adding new features behind it. The search box is not just a product element at this point; it is a cultural artifact — one of the few pieces of digital infrastructure used by essentially the entire internet-connected world. Changing it sends an unmistakable message about where the company believes computing is headed.

For 25 years, the search box trained billions of people to think in keywords — to compress their curiosity into the shortest possible string of words. The new box invites them to do the opposite: to think out loud, to upload what they’re looking at, to ask follow-up questions, to let an AI system handle the compression.

Pichai tied the company’s broader ambitions to a striking statistic: Google’s surfaces now process over 3.2 quadrillion tokens per month, up seven-fold from a year ago. The company expects capital expenditures of approximately $180 to $190 billion in 2026 — roughly six times the $31 billion it spent four years ago — largely to support the infrastructure required for this AI transformation. When asked about the future of traditional search, he was direct. “Search is the most used AI product in the world,” he said.

The blinking cursor in Google’s search box still invites you to type. But after 25 years of teaching the world to speak in keywords, Google is now asking it to speak in sentences — and betting roughly $190 billion that it will.

Google says Gemini 3.5 Flash can slash enterprise AI costs by more than $1 billion a year

Google unveiled Gemini 3.5 Flash at its annual I/O developer conference on Tuesday, a new artificial intelligence model that the company says shatters what had become a seemingly iron law of the AI industry: that the smartest models must also be the slowest and most expensive to run.

The model sits at the center of a sweeping set of announcements — from a video-generating “world model” called Gemini Omni to a 24/7 personal AI agent called Gemini Spark — but 3.5 Flash carries perhaps the most immediate consequence for the enterprises pouring billions of dollars into AI infrastructure. Sundar Pichai, Google’s chief executive, told reporters during a press briefing Monday that companies running roughly one trillion tokens per day on Google Cloud could save more than $1 billion annually by shifting 80 percent of their workloads to a mix of Flash and other frontier models.

“You’ve probably heard anecdotes from other CIOs that companies are already blowing through their annual token budgets, and it’s only May,” Pichai said, framing the model not just as a technical achievement but as a financial lifeline for organizations struggling with the runaway costs of deploying AI at scale.

The claim, if it holds, would be one of the most significant shifts in the economics of enterprise AI since large language models entered corporate computing.

Why enterprises have been forced to choose between AI quality and AI speed

For the past three years, organizations adopting generative AI have faced a painful trade-off. The most capable models — the ones that can reason through complex multistep problems, write reliable code, and parse dense financial documents — tend to be large, slow, and expensive to query. Faster, cheaper models sacrifice accuracy. Chief information officers have been forced into a kind of AI portfolio management: routing simple queries to lightweight models and reserving the heavy-duty reasoning engines for high-stakes tasks. It is a complex, brittle system that adds engineering overhead and often delivers inconsistent user experiences.

Gemini 3.5 Flash attacks that trade-off directly. According to Google’s internal benchmarks and a third-party analysis from Artificial Analysis, the model outperforms Google’s own Gemini 3.1 Pro — a model the company positioned as its top-tier flagship just four to five months ago — on nearly every major benchmark. It scores 76.2 percent on Terminal-Bench 2.1, reaches 1656 Elo on GDPval-AA, hits 83.6 percent on MCP Atlas, and leads in multimodal understanding with 84.2 percent on CharXiv Reasoning.

Yet it does all of this while generating output tokens at four times the speed of comparable frontier models from competitors. Koray Kavukcuoglu, chief technology officer of Google DeepMind and chief AI architect for Google, told reporters the team has pushed even further: “We have developed an even more optimized version of Flash, not just four times, but actually 12 times faster with the same quality.” That turbo variant is available starting Tuesday inside Antigravity, Google’s agentic development platform.

Pichai put the performance gap in blunt terms: “3.5 Flash is better than 3.1 Pro, which was just four months ago, and it’s at the almost, I would say, 90% of the performance of frontier models, 4x faster, much faster in Antigravity, maybe 12x, and about 1/3 to one half the cost.”

Landing in what Artificial Analysis calls the “top-right quadrant” of its intelligence-versus-speed index — the only model to do so — Flash occupies a position no competitor currently holds.

The trillion-token math behind Google’s $1 billion savings claim

To understand why Flash matters so much to enterprise buyers, you need to understand the economics of tokens — the fundamental units of data that AI models process. Every query a customer service chatbot answers, every legal document an AI summarizes, every line of code an agent writes, consumes tokens. And at frontier-model pricing, those tokens add up fast.

Google says its model APIs now process around 19 billion tokens per minute. Across all of Google’s own surfaces — Search, the Gemini app, Workspace, and more — the company processes over 3.2 quadrillion tokens per month, a figure that has jumped seven-fold in the past year alone. Two years ago, at I/O 2024, the number was 9.7 trillion per month.

The explosion in token consumption is not unique to Google. Enterprises across industries are discovering that the more capable their AI deployments become, the more tokens they burn. Agentic workflows — where AI systems autonomously execute multistep tasks, call tools, write and run code, and iterate on their own output — are particularly token-hungry. A single agentic coding session can consume orders of magnitude more tokens than a simple question-and-answer exchange.

This is where Flash’s cost advantage becomes transformative. The model delivers what Google describes as frontier-level capabilities at less than half the price, in some cases almost a third the price, of comparable frontier models. For a hypothetical enterprise processing one trillion tokens per day on Google Cloud — a scale Pichai said top customers are already reaching — the savings from shifting 80 percent of workloads to a Flash-and-frontier blend would exceed $1 billion per year.

That is not a rounding error. It is the kind of number that reshapes procurement decisions, accelerates deployment timelines, and fundamentally alters the return-on-investment calculus for AI initiatives that many boards of directors have been scrutinizing with increasing impatience.

How Google’s own engineers created a data flywheel that rivals cannot easily copy

Perhaps the most strategically significant detail Google shared Tuesday was not a benchmark score or a price point. It was a chart showing the company’s own internal token consumption on Antigravity 2.0, its reimagined agentic development platform.

In March 2026, Google’s developers were processing roughly half a trillion tokens per day inside Antigravity. By the time of the I/O press briefing in mid-May, that figure had surged past three trillion — a six-fold increase in approximately ten weeks, with usage doubling “literally every few weeks,” according to Pichai.

This internal usage creates what AI researchers call a data flywheel: the more Google’s own engineers use 3.5 Flash to build products, the more real-world signal the model team collects on where the model excels and where it stumbles. That signal feeds back into model improvement, which makes the model more useful, which drives more usage, which generates more signal. It is a virtuous cycle — and it is one that competing AI labs, which rely primarily on external developer usage and synthetic benchmarks, cannot easily replicate at the same speed or fidelity.

“That scale creates a powerful feedback loop, and that is what has allowed us to keep improving the 3.5 series of models,” Pichai said.

When pressed during the Q&A about the competitive frontier — particularly in light of recent advances from rival labs — Pichai acknowledged the landscape is “very dynamic” and “moving fast” but expressed confidence in Google’s breadth. He added that the company’s focus with the 3.5 series has been on “taking the model intelligence, making sure tool use, instruction following, long horizon use cases, agent decoding all work well.”

Kavukcuoglu reinforced the agentic emphasis, noting that 3.5 Flash “can now handle multi-hour autonomous sessions” and “can independently execute complex coding pipelines or manage iterative research projects entirely by itself.” The team, he said, even tested the model by having agents build a working operating system entirely from scratch.

Antigravity 2.0 transforms Google’s code editor into an agent command center

The arrival of 3.5 Flash is tightly coupled with the launch of Antigravity 2.0, a significant expansion of the agentic development platform Google first introduced six months ago. What began as a coding environment has evolved into what Google describes as a full platform for developing and managing teams of autonomous AI agents, and the company says millions of developers are already building with it.

Antigravity 2.0 ships as a new standalone desktop application that serves as a central hub for orchestrating multiple agents simultaneously. Google offered the example of running one agent to code a website, a second to generate brand assets, and a third to plan product architecture — all in parallel, all managed from a single interface. For developers who prefer command-line workflows, there is Antigravity CLI. And for those building programmatic integrations, the new Antigravity SDK provides direct access to the same agent harness powering Google’s own first-party products.

The co-development of 3.5 Flash and Antigravity 2.0 is no accident. “We have co-developed 3.5 Flash together with Google Antigravity, our agentic development platform,” Kavukcuoglu said. This tight integration means Flash’s strengths — speed, tool use, long-context reasoning, and code generation — are specifically tuned for the kinds of workloads developers execute inside the platform.

Google is also launching Managed Agents in the Gemini API, allowing developers to spin up an agent with a single API call that reasons, uses tools, and executes code in an isolated Linux environment. And it introduced CodeMender, an AI security agent that uses Gemini’s advanced reasoning to automatically find and fix critical code vulnerabilities — a capability Kavukcuoglu described as essential as agentic systems write an increasing share of the world’s code.

Google’s $190 billion infrastructure bet and the custom silicon powering cheaper AI

The models and platforms sit atop a staggering infrastructure investment that Pichai revealed during the briefing: Google expects capital expenditures of approximately $180 billion to $190 billion in 2026 — roughly six times the $31 billion the company spent in 2022, just four years ago.

A key component of that spending is custom silicon. The company recently unveiled its eighth generation of Tensor Processing Units, adopting for the first time a dual-chip architecture with specialized designs for training (TPU 8o) and inference (TPU 8i). Google says it can now distribute model training across multiple data center sites using a system called Pathways, scaling beyond one million TPUs globally — a setup the company claims constitutes the largest training cluster in the world.

“This means training larger, more capable models in weeks, rather than months,” Pichai said. The infrastructure advantage matters enormously for Flash’s economics. Custom silicon optimized for inference means Google can run Flash at lower cost per token than competitors relying on general-purpose GPUs, and the savings get passed along — at least partially — to customers.

The capex figure also signals something strategic about Google’s long-term posture. While some investors have grown nervous about the astronomical sums cloud providers are spending on AI infrastructure, Google is framing the spending as a competitive moat. The more infrastructure it builds, the cheaper it can run inference, the more attractive its models become, and the more usage it captures to improve the next generation. It is the flywheel logic again, extended from software all the way down to silicon.

Gemini Omni, Spark, and the consumer products Flash now powers at massive scale

While the enterprise cost story dominates the Flash narrative, Google also made sweeping moves on the consumer side that put the model to work across products reaching billions of people. Flash is now the default model powering the Gemini app — which has surpassed 900 million monthly active users, more than doubling from 400 million a year ago — and AI Mode in Google Search, which has crossed one billion monthly users in its first year.

Google introduced Gemini Spark, a 24/7 personal AI agent that runs on dedicated virtual machines in Google Cloud and operates in the background even when a user’s device is off. Powered by 3.5 Flash with the full Antigravity harness, Spark integrates with Gmail, Docs, Sheets, and Slides. Josh Woodward, who leads Google Labs and the Gemini app, described the experience vividly: “When you use it, it almost feels like you’re tossing things over your shoulder, Spark’s catching them and gets the job done.” On the safety front, Spark requires explicit user approval before high-stakes actions. Google also announced the Agent Payments Protocol, which lets users set strict guardrails — approved brands, spending caps, specific merchants — before an agent can spend money on their behalf. Woodward compared the design to “giving a teenager their first debit card — there’s sort of limits and sort of constraints around it.”

Alongside Flash, Google unveiled Gemini Omni, a model capable of generating any output from any input, starting with video. Kavukcuoglu drew a sharp distinction from Google’s existing Veo model: “Veo is a text-to-video model. Omni is a true and true multi-model input, multi-model output model.” All Omni-generated content carries Google’s SynthID watermark, and the company announced that OpenAI, Kakao, and ElevenLabs are adopting SynthID as well.

The company also reimagined its search box for the first time in over 25 years, introduced information agents that monitor the web around the clock for user-defined conditions, and launched the Universal Cart — an AI-powered cross-merchant shopping cart built on Google Wallet. Liz Reid, who leads Google Search, called the new search box “the biggest upgrade to our iconic search box since its debut.”

What Google’s six-month model cadence means for the enterprise AI cost curve

Google signaled that 3.5 Flash is just the opening act of the 3.5 series. Gemini 3.5 Pro is currently in internal testing and will roll out to everyone next month. Kavukcuoglu indicated the company has been operating on roughly a six-month cadence for major model updates — Gemini 3 in November, 3.5 in May — and expects that rhythm to continue.

When a reporter from The New York Times asked how Google determines whether a release warrants a full numerical jump or a half-step increment, Kavukcuoglu said the numbering reflects the magnitude of research progress: “What defines the numbering update is really the progress that we see in our research and how it is reflected in the models and the impact that they have.”

For enterprise buyers, that cadence carries an important implication: the cost-performance curve is not just improving — it is improving on a predictable schedule. A model that outperforms the previous flagship at a third the cost every six months fundamentally changes the planning horizon for AI investments. It means the token budgets that companies are blowing through today may look quaint by the end of the year.

Google’s announcements arrive at a moment of intense competition. OpenAI, Anthropic, Meta, and a constellation of smaller labs are all racing to deliver models that balance capability with cost. Microsoft has been aggressively integrating OpenAI’s models into Azure and Copilot. But Google benefits from a structural advantage that is easy to overlook: distribution. With 13 products serving more than a billion users each — five of which exceed three billion — Google can deploy Flash to an audience no pure-play AI lab can match. Every improvement immediately benefits Search, Gmail, Docs, Maps, and YouTube. And the usage data flowing back from those billions of interactions feeds the very flywheel that makes the next model better.

The question now is whether the $1 billion savings figure — an eye-catching projection based on a specific workload mix — will survive contact with the messy reality of corporate AI deployments, where legacy systems, compliance requirements, and organizational inertia have a way of blunting even the most compelling cost curves. But if Google’s own internal usage is any guide — three trillion tokens a day and climbing, doubling every few weeks, with no sign of slowing — the company is not just selling the bet. It is making the bet itself, with its own engineers, on its own infrastructure, at a scale no customer has yet attempted. In the AI cost wars, the most persuasive pitch may simply be: we did it first.

Google unveils Gemini Omni ‘any-to-any’ AI model: what enterprises should know

Although it was already discovered by intrepid AI power users weeks ago, Google’s new Gemini Omni model officially debuted today at the company’s annual I/O developer conference in Mountain View, California, and it marks a significantly new paradigm in the wider AI and tech marketplace.

That’s because as its “omni” (from the Latin omne — meaning “all”) prefix would suggest, this is Google’s first truly native, multimodal model, that is “a model that can create anything from any input — starting with video.”

The model marks Google’s bid to collapse the multimodal generative stack — text-to-image, image-to-video, video-to-video, audio generation — into a single foundation model with a single editing surface.

The big question for business leaders is: should you switch any of your own AI stack over to Gemini Omni now?

Unfortunately, the truth is, you may not be able to just yet — the model is only available to individual users through Google’s AI subscription plans starting with the $20 per user per month “AI Plus” plan. It can currently be accessed on the Gemini website and mobile apps, Google’s web-based Flow AI image and video editing suite, and YouTube Shorts.

While the company says it is ultimately going to be available via an application programming interface (API) — which many enterprises rely on for their AI needs — it’s not ready yet.

In a departure, Google also did not issue any public benchmarks for Gemini Omni (yet). However, third-party organizations will no doubt put it to the test on various tasks and user-reported quality metrics. In the meantime, though, its quality and speed remain somewhat subjective.

But, given the capabilities and faster editing enabled by the new Omni model, individual members of your team should probably give serious consideration to switching over to it, especially if they work creating visuals for technical diagrams, marketing and comms materials, training and corporate education courses, sales collateral, and basically anything that involves visuals.

What Omni actually is

Omni is the next chapter of the work that produced Nano Banana, the image-generation and editing model Google shipped roughly a year ago.

The first model in the family, Gemini Omni Flash, accepts any combination of text, images, audio, and video as input and produces high-quality output across the same modalities — all from a single model rather than a relay of specialized systems.

Google says the model is “natively multimodal from the ground up,” which matters less as marketing copy than as an architectural claim: a unified model can reason across modalities in the same forward pass, which generally translates into more coherent edits, fewer pipeline artifacts, and a far cleaner API surface for developers.

OpenAI started this trend back in May 2024 with the release of GPT-4o, its first natively “omni” model, also trained from the ground-up to be able to analyze and generate multiple different types of content, from text to code, imagery, and audio. However, it did not support video generation, and the model was eventually deprecated following reports of sycophancy and even users demanding OpenAI retain it after developing parasocial relationships with it.

Is Gemini Omni at risk of sparking a similarly devoted following? It remains to be seen.

One big difference is that its headline interaction pattern is conversational video editing. Each instruction “builds on the last,” and past directions persist across turns so the video evolves coherently as the user iterates.

Practical examples Google highlighted include changing the world inside a clip, reimagining an action or camera angle, refining sequences over multiple turns, and generating explainer-style content from short prompts.

Google also emphasizes improved physics — gravity, kinetic energy, fluid dynamics — which is the kind of detail that separates “looks like AI video” from “looks like footage.”

Rollout, pricing, and the API question

The first thing enterprise leaders should read carefully is the rollout plan. Omni Flash is going live today inside the Gemini app for U.S. subscribers across AI Plus, AI Pro, and AI Ultra tiers — including the new $100-per-month AI Ultra plan Google announced at the same event.

Google says it will roll out to developers via Vertex AI APIs “in the coming weeks.” That gap is significant. Until the Vertex API is generally available, Omni is effectively a consumer and prosumer tool.

Enterprise pilots beyond individual seat-based experimentation should wait for the API, both because that’s where Google’s enterprise SLAs and data-handling commitments live, and because production-grade generative video without a programmatic interface is a non-starter.

Its pricing through the API per million tokens (presumably) will also determine its viability as an enterprise product outside of film/TV/entertainment and the arts productions.

For decision-makers weighing seat economics in the meantime, the new AI Ultra tier is positioned specifically at developers, technical leads, knowledge workers, and advanced creators, with priority access to Google Antigravity, higher usage limits, and bundled Omni Flash access.

For small creative teams under tight deadlines, that may be the fastest way to evaluate the model before the API arrives.

The enterprise use cases that really matter

It is easy to default to “marketing video” as the use case, but Omni’s value proposition for enterprises is broader if you think of it as a programmable video and media engine rather than a creative app:

  • Sales and marketing: rapid generation of variant ads, localized creative, and product demos without per-asset agency cycles.

  • Internal communications, learning and development (L&D): explainer videos, onboarding modules, and policy walkthroughs produced by non-specialists.

  • Customer support and documentation: dynamic, query-conditioned visual explainers attached to help articles.

  • Product and engineering: visualization of simulations, UI walkthroughs, and concept videos for spec reviews.

  • Field operations: short, situation-specific instructional clips generated on demand.

What changes with Omni versus the previous generation of tools is the unification. Many enterprises stitched a workflow together from text-to-image, image-to-video, lip-sync, and voice models, each with its own contract, billing, and data path. A single Vertex AI-backed model collapses procurement and observability into one place — assuming the eventual API delivers production-grade throughput and latency.

The governance story is the most underrated part

For CIOs and CISOs, the most important section of Google’s announcement is not the model card; it is the provenance and content-safety work shipping alongside it.

Every video generated by Omni carries Google’s SynthID digital watermark. Google is expanding C2PA Content Credentials across its generative tools, and launching an AI Content Detection API on Agent Platform that lets businesses identify AI-generated content from both Google and other popular models.

Partner integrations announced at the same event — including Shutterstock, Avid (in Pro Tools), and at least one major newswire — indicate where the standard is going.

For enterprises, this matters in three concrete ways:

  1. It gives legal and compliance teams a defensible audit trail for AI-generated media.

  2. It allows brand-safety teams to detect AI-generated material entering content pipelines from third parties.

  3. And it provides a defensible answer for regulators in jurisdictions, like the EU, that are tightening rules around synthetic-media disclosure.

There is also a “Personal Avatars” program that lets creators record short videos to authorize use of their voice and likeness across generated content, as Google leaders and employees showcased themselves today in posts centered around I/O featuring their AI generated likenesses.

This puts it in direct competition with Synthesia, a UK-based AI unicorn focused primarily on enterprise-safe AI videos and avatars.

For enterprises considering executive videos, training avatars, or branded spokesperson content, the consent model here is the right starting point — but contracts and rights-management policies will need to extend to cover it.

Risks worth flagging

Omni’s main risks are familiar but worth restating.

The competitive landscape is crowded with the aforementioned Synthesia, TikTok parent company ByteDance’s acclaimed Seedance model, Kuaishou Technology’s Kling AI models, and the fast-improving open-source field all compete for the same workflows.

Lock-in to any single video model is a real concern when output quality is still leapfrogging quarterly.

Latency and cost for production-volume video generation remain unproven outside controlled demos.

In addition, the legal status of training data for generative video is unsettled in multiple jurisdictions; enterprises should require clear indemnification language before deploying generated video into customer-facing channels.

Furthermore, VentureBeat collaborator and AI YouTuber Sam Witteveen, CEO of enterprise machine learning vendor Red Dragon AI, received early access to Gemini Omni and reported the content restrictions (which some deem to be censorship) to be quite strict, potentially restricting and inhibiting all the potential use cases an enterprise would like to pursue.

Thoughts for enterprises considering adoption

Omni is worth piloting — but the structure of the pilot matters.

For most enterprises, the right move over the next 30 to 60 days is to fund a small, sanctioned experiment with one or two AI Ultra seats in marketing or L&D, while the platform and security teams use that runway to prepare for the Vertex AI API: define data-residency requirements, set up SynthID and C2PA verification in the content pipeline, and stand up the AI Content Detection API alongside existing media-governance tooling.

Treat the consumer rollout as a UX preview, not a production plan. When the API arrives, the enterprises that have already done the governance work will be the ones moving Omni into real workflows while everyone else is still drafting policy.

Omni is not, by itself, a reason to overhaul an enterprise AI strategy. But it is a strong signal that the multimodal generative stack is consolidating into single models with first-party provenance baked in — and that is a shift technical decision-makers should be planning around now.

OpenAI co-founder Andrej Karpathy announces he’s joining Anthropic

Andrej Karpathy, the influential 39-year-old Slovak-Canadian AI researcher and one of the original 11 co-founders of OpenAI, and former head of Tesla’s AI division, announced on Tuesday, May 19 that he’s joining rival lab Anthropic.

As Karpathy posted from his account on the social network X: “Personal update: I’ve joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.”

Anthropic’s current Head of Pretraining, Nicholas Joseph, also a former OpenAI alumnus, added more context to Karpathy’s new role at Anthropic in a post of his own on X, writing: “Excited to welcome Andrej to the Pretraining team! He’ll be building a team focused on using Claude to accelerate pretraining research itself. I can’t think of anyone better suited to do it — looking forward to what we build together!”

An Anthropic spokesperson confirmed to VentureBeat via email that Karpathy will be starting a team focused on using Claude, Anthropic’s own, increasingly popular AI model, to accelerate pretraining research. This would put Anthropic further toward the overarching AI research goal of many around the world to develop “recursive self-improvement,” that is, AI that is capable of training its successors or upgrading itself with increasingly lesser, or ultimately no human intervention.

The announcement came on the same day as the start of rival AI-focused tech firm Google’s annual I/O developer conference in its headquarters city of Mountain View, California, when many new releases and announcements were expected.

Karpathy’s storied history

Karpathy is widely known for spanning three parts of the modern AI boom: academic research, big-company deployment and online education.

His own website describes him as an AI researcher and educator who was a founding member of OpenAI, later served as Director of AI at Tesla, and helped create Stanford’s first deep learning course, CS231n.

OpenAI’s December 2015 launch announcement also listed Karpathy among the group’s founding members.

At Tesla, where he worked from 2017 to 2022, Karpathy led the computer vision team for Autopilot and says his team handled in-house data labeling, neural network training and deployment on Tesla’s custom inference chip.

He then returned to OpenAI from 2023 to 2024, where his website says he built a team focused on midtraining and synthetic data generation — experience directly relevant to Anthropic’s reported pretraining role.

Karpathy’s academic work began at Stanford, where he earned his PhD under Fei-Fei Li and focused on neural networks for computer vision, natural language processing and the intersection of the two.

He also interned at Google Brain, Google Research and DeepMind, according to his website. His education includes an MSc from the University of British Columbia and a BSc from the University of Toronto, where he double-majored in computer science and physics.

What will become of Karpathy’s open source research and commitment to AI education?

Since leaving OpenAI in 2024, Karpathy has become one of AI’s most visible public educators, publishing technical and general-audience videos on large language models and neural networks.

He also launched Eureka Labs in July 2024 as an “AI-native” school; its first product, LLM101n, is described as an undergraduate-level course guiding students through training their own AI system.

Acting on his own as a free agent over the last two years, Karpathy has also helped push open source AI research forward with products and standards including autoresearch, an LLM-driven automated researcher that can run multiple hypothesis and experiments simultaneously, and the LLM Knowledge Base, an autonomous system of storing memory and context for AI agents in a kind of ever-growing library designed for them to access.

The big question is what becomes of these and Karpathy’s open source AI efforts more generally as he joins Anthropic, a lab that has supported open source via the launch of its Model Context Protocol (MCP) technical standard, but which also famously has shipped primarily proprietary AI models and harnesses (such as Claude and Claude Code).

Based on the last statement in his announcement post on X — “I remain deeply passionate about education and plan to resume my work on it in time” — it appears that at least his contributions to the AI-native school effort will be paused as he digs in at Anthropic.

Sony Marks 10 Years Of Noise-Cancelling Headphones With Premium 1000X The Collexion

Sony has just revealed its latest premium headphones. The Collexion aims to offer outstanding sound in highly crafted design.

The Hidden Players Powering The Future Of Quantum Computing

If you don’t know how to audit why an agentic system is failing, you can’t lead the product.

Anker’s Compact SOLIX S2000 Cuts Standby Power Drain

Anker’s new SOLIX S2000 targets home backup with a compact design, ultra-low standby drain, and a $599 early launch price

DSPM: The Missing Piece For A Successful DLP Project

Here’s why DSPM is the future of data loss prevention.

Samsung’s Galaxy S26 Ultra Price Gamble Pays Off Early

Samsung kept the Galaxy S26 Ultra price flat while hiking everything else. New data suggests that gamble is paying off, but how long for?