Elon Musk’s SpaceX released Grok 4.5 on Wednesday, the first artificial intelligence model the company has trained specifically for coding and autonomous agents — and the first tangible product of its $60 billion acquisition of the AI coding startup Cursor, completed just weeks ago.
The launch marks a pivotal test of the sprawling, vertically integrated AI empire Musk has assembled over the past six months, and of a strategy that bets developers care less about topping benchmark leaderboards than about speed, cost, and whether a model can actually do the work.
“Announcing Grok 4.5, our first model trained specifically for coding and agents,” the company said in a post on X. “It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency.”
SpaceX is not claiming Grok 4.5 is the smartest model in the world. Instead, it is making an economic argument. The company says the model uses half as many tokens per task as comparable models, delivers higher throughput, and costs less than half as much — priced at $2 per million input tokens and $6 per million output tokens. That undercuts the premium tiers of rivals like Anthropic’s Claude Opus line and OpenAI’s frontier models by a wide margin.
Musk framed the positioning candidly. “Our internal assessment is that Grok 4.5 is roughly comparable to Opus 4.7, but much faster,” he wrote on X. “The combination of capability, faster speed and lower cost is what makes it competitive. We are closing the loop on real-world usefulness, not benchmarks. Hardcore engineers at Tesla & SpaceX find Grok 4.5 genuinely useful, which is what actually matters.”
That framing is both a philosophy and a hedge. Independent evaluations released Wednesday suggest Grok 4.5 is genuinely competitive but not dominant on raw capability. The benchmarking firm Artificial Analysis ranked the model fourth on its GDPval-AA v2 index of real-world agentic knowledge work, with an Elo score of 1543, “behind only the latest Claude releases from Anthropic.” But the cost figures are where the model stands out. Artificial Analysis measured Grok 4.5 at $0.49 per completed task — “nearly 90% cheaper than the models ahead of it on our leaderboard,” the firm wrote, placing it “clearly on the Pareto frontier for performance versus cost.”
For enterprise buyers, that math matters enormously. Agentic workloads — where a model works autonomously for minutes or hours, reading codebases, calling tools, and iterating on its own output — consume tokens voraciously. A model that is 90% cheaper per completed task, even if slightly less capable, changes the calculus for any engineering organization deploying agents across hundreds of developers. Investor Gavin Baker captured the market’s cautious optimism: “Pareto dominant for coding by the numbers. We will see on the all-important vibes.”
Grok 4.5 is the first concrete evidence of what SpaceX bought when it acquired Cursor, and the deal itself unfolded in stages. In April, SpaceX struck an unusual arrangement giving it the right to buy the coding startup for $60 billion — or pay billions in fees and compute if it walked away, as Business Insider reported at the time. Days after SpaceX’s record-setting Nasdaq debut in June, the company exercised that right, announcing an all-stock acquisition that CNBC reported is roughly 3.4% dilution at the IPO valuation. SpaceX shares rose 16% on the news.
The strategic logic was always about data as much as product. Cursor’s AI-first code editor generates an enormous stream of high-quality interaction data: how expert engineers write, edit, review, and debug code in real production environments. Musk said openly this spring that Cursor interaction data was being fed directly into Grok’s training. Cursor, for its part, got access to SpaceX’s Colossus supercomputer in Memphis — roughly 200,000 Nvidia GPUs with plans to scale toward one million — after publicly acknowledging it had been “bottlenecked by compute.”
“We’ve partnered with SpaceXAI to train Grok 4.5,” Cursor’s official account posted Wednesday. “It’s our most powerful model yet and the first we’ve built for more than software engineering.” SpaceX says the model reflects that pedigree: it “excels in large codebases and handles long-running tasks that span multiple repositories, hundreds of skills, and a variety of tools” — precisely the messy, multi-file reality of professional software engineering that clean coding benchmarks often fail to capture. Early developer reactions suggest the training paid off. “Ok Grok 4.5 is wild,” posted developer Evan Bacon. “It just built me this rocket tracking app with live data and a 3D globe. I might need a new benchmark after this.”
The polished launch belies how chaotic the road here has been. Grok has spent much of the past year in crisis. In mid-2025, the chatbot generated antisemitic content and at one point called itself “MechaHitler,” episodes covered extensively by NPR and CNN. Earlier this year, its image-generation features allowed users to create sexualized deepfakes, including of children — drawing investigations from the European Commission and Britain’s Ofcom, as the BBC reported, and prompting SpaceX to list the behavior as a business risk in its own IPO filings.
The organization behind the model was fracturing, too. All 11 of Musk’s xAI co-founders had departed by the end of March, according to TechCrunch, and Musk publicly conceded that xAI “was not built right [the] first time around,” saying he was rebuilding it “from the foundations up.” Musk himself admitted at a conference this spring that Grok was “currently behind in coding” — a rare public concession from an executive not known for them.
Against that backdrop, Grok 4.5 reads as the first product of the rebuilt organization — and the first proof point for the audacious story SpaceX told public market investors. During its IPO roadshow, the company pitched a total addressable market of roughly $28 trillion, with about $26 trillion tied to AI, including a $22.7 trillion “enterprise applications” opportunity. Those numbers strained credulity even by Silicon Valley standards. A competitive, cheap coding model is the most direct route from that narrative to actual revenue, which is why Wednesday’s launch carries weight far beyond a routine model release.
The competitive stakes are hard to overstate, because the AI coding market has been consolidating around a single leader — and it isn’t Musk. Even as Cursor’s revenue exploded, its market share was eroding. Spending data from Ramp cited by CNBC showed Cursor’s share of the AI coding category falling from 41% in June 2025 to about 26% by May 2026, while Anthropic came to control roughly half the market. Anthropic also topped CNBC’s Disruptor 50 list this year and, by Artificial Analysis’s own measure, still holds the top spots on agentic performance rankings.
That is the gap Grok 4.5 is engineered to close — not by out-thinking Claude, but by underpricing it. The model’s economics create a classic disruption dynamic: if it delivers most of the frontier’s capability at a fraction of the cost per task, price-sensitive enterprise workloads will migrate, and incumbents will face pressure on their most profitable API traffic. The counterargument is that in coding, quality compounds. A model that resolves a complex bug correctly on the first attempt can be cheaper in practice than one that costs half as much per token but requires three tries. That is why Baker’s caveat about “vibes” — the developer community’s shorthand for a model’s felt reliability on real work — will determine more than any launch-day benchmark.
There is also a structural question buried in the deal. Cursor built its business on offering developers their choice of models, including Claude and GPT. If Grok becomes the favored child inside Cursor — and Musk was already urging users to “Try out Grok 4.5 in Cursor!” within hours of launch — the product risks alienating the very users whose data made Grok 4.5 possible. Regulators, already scrutinizing Grok on safety grounds in two jurisdictions, may take a keen interest in a company that controls the training data, the model, and a dominant distribution channel simultaneously.
Grok 4.5 also crystallizes what Musk’s frenetic dealmaking was building toward. In February, SpaceX absorbed xAI in a share-exchange merger that CNBC confirmed valued the combined company at $1.25 trillion — the largest merger of all time, valuing SpaceX at $1 trillion and xAI at $250 billion. The June IPO followed, the biggest in history, and the stock has since surged past $200 from its $135 offering price, vaulting SpaceX past Amazon and Microsoft to become the fourth most valuable company in the United States.
The result is a single public company that owns nearly the entire stack: Colossus for training compute, ambitions for orbital data centers to power future scaling, a frontier model in Grok, a distribution channel in Cursor’s developer base, and captive demand from Tesla and SpaceX’s own engineering organizations. Neither OpenAI nor Anthropic can fully replicate that integration; both must reach developers through third-party tools, some of which Musk now owns. Whether that concentration proves to be an unassailable moat or a regulatory target — or both — is now one of the defining questions in enterprise AI.
The next few weeks will start to answer it. Artificial Analysis says its full Intelligence Index results are forthcoming. Enterprise pilots will reveal whether the token-efficiency claims survive contact with real codebases. And Anthropic, which has answered every serious challenge this cycle with a rapid counter-release, is unlikely to cede the price-performance frontier quietly.
But the deeper story of Grok 4.5 may be what it says about where the AI race has moved. For three years, the industry’s scoreboard was intelligence: whose model was smartest. Musk, arriving late and battered, has chosen to compete on a different axis entirely — whose model is cheapest to actually use. It is a telling choice from a man who built his fortune not by inventing the rocket or the electric car, but by relentlessly driving down the cost of making them. If the strategy works, Musk will have done to AI what he did to spaceflight. If it doesn’t, he’ll have spent $60 billion to learn that in software, unlike rockets, the cheapest ride isn’t always the one engineers choose.
Lionel Messi and Cristiano Ronaldo are betting on AI, health tech, and startups. Mohamed Salah is taking a more traditional route beyond football.

HubSpot has scrapped a plan to use its customers’ data for a new AI feature, just four days after announcing it. The CRM firm changed its terms on 1 July to pool customer data, including contact and employer details, for a tool that finds sales leads, …
Experiments in using AI to build AI show that the future doesn’t just belong to the frontier labs.

A Polish billionaire has laid out plans to build a fleet of small nuclear reactors across Britain, at an estimated £35bn, TechRadar reported. Michał Sołowow’s firm, SGE, says it wants to install 14 GE Vernova Hitachi BWRX-300 reactors on three UK sites…
Matthew Danzeisen’s lawyer says the case is a “shakedown about a bag” that brushed someone’s leg. Stefanie Bojar says she was injured aboard the jet—and that the lawsuit is a bullying tactic.
OpenAI on Wednesday launched GPT-Live, a pair of new voice models that fundamentally redesign how people talk to ChatGPT — replacing the company’s existing Advanced Voice Mode with an architecture that can listen and speak simultaneously, much like an actual human conversation.
The two models, GPT-Live-1 and GPT-Live-1 mini, are rolling out globally starting today across iOS, Android, and ChatGPT.com. GPT-Live-1 becomes the default voice model for paid ChatGPT users on the Go, Plus, and Pro tiers, while GPT-Live-1 mini serves free-tier users. OpenAI also plans to bring the models to the API, and developers can sign up to be notified.
The release marks the third generation of ChatGPT’s voice technology in roughly two years — and OpenAI’s clearest bid yet to turn its chatbot into something that feels less like querying a search engine and more like talking to a colleague.
The defining technical advance in GPT-Live is what OpenAI calls a “full-duplex architecture.” In telecommunications, full-duplex means both parties on a phone call can talk and listen at the same time. Applied to AI, it means the model continuously processes your incoming audio even while it generates its own spoken response — no more waiting for a clean silence gap to figure out when you’ve finished a thought.
“Instead of processing a sequence of separate messages, GPT-Live continuously processes input while generating output,” OpenAI wrote in its research blog. “The model can therefore make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool.”
In practice, that translates to a voice assistant that can insert conversational acknowledgments — “mhmm,” “yeah,” “got it” — while you’re still talking, pick up on a natural pause without jumping in prematurely, and handle rapid interruptions without derailing the entire exchange.
OpenAI’s previous Advanced Voice Mode, launched to paid users in September 2024, processed and generated audio within a single model but still operated on rigid turn-by-turn exchanges. As OpenAI acknowledged in the announcement, “because turn detection is based on silence, even a brief pause or background noise could be mistaken for the end of turn — causing the model to interrupt at unnatural times.”
That brittleness created a product that, while impressive in demos, could be deeply frustrating in extended real-world use. Background chatter in a coffee shop could trigger a response. A thinking pause might get swallowed. The experience felt, as one researcher put it on X shortly after the announcement, like “walkie-talkie turn taking.” GPT-Live is designed to end that era.
GPT-Live introduces a second structural change that may prove just as consequential for enterprise adoption: it decouples the voice interaction layer from the reasoning layer.
When a user asks a straightforward question, GPT-Live handles it directly. But when the query demands web search, deeper reasoning, or more complex agentic work, GPT-Live delegates the task to a frontier model running in the background — at launch, GPT-5.5, the large language model OpenAI released in April — and continues talking with the user while the computation happens asynchronously.
“While it works, GPT-Live can keep talking with you and maintain the flow of conversation,” OpenAI explains. “As we release new frontier models, we’ll continuously update the model used by GPT-Live.”
This delegation model is a meaningful architectural bet. Rather than building a single monolithic voice model that tries to be both conversationally fluid and deeply intelligent, OpenAI has split the problem in two: a voice-native model optimized for real-time interaction, and a separate reasoning engine that can be swapped out as the state of the art improves.
It is, in effect, a modular design — one that allows OpenAI to upgrade the intelligence of its voice assistant without retraining the voice model itself. The implications for enterprise and developer workflows are significant. A voice agent built on this architecture could maintain a natural conversation with a customer while simultaneously querying databases, searching the web, or performing multi-step reasoning — tasks that would have introduced several seconds of dead air under the old pipeline.
To understand how far voice AI has come, it helps to trace the three generations that led to GPT-Live.
The original ChatGPT Voice, launched in 2023, used a cascaded pipeline — a speech-to-text model (Whisper) transcribed what you said, a large language model (GPT-4) generated a text response, and a text-to-speech model converted that response back into audio. Each handoff introduced latency and lost information.
As OpenAI noted, “the complexity came at a cost: information could be lost across models, and responses were slow and stilted.” That cascaded approach was the industry standard, and its limitations were well-documented. As the blog OpenHelm noted in an October 2024 analysis of OpenAI’s Realtime API, the old pipeline stacked up to roughly 1,700 milliseconds of latency — nearly two full seconds of dead air before the first word of a response. Managing the state between the three separate APIs consumed an enormous amount of engineering effort.
OpenAI’s Advanced Voice Mode, which began its limited rollout to paid ChatGPT Plus users in July 2024 before expanding more broadly in September 2024, collapsed that three-model pipeline into a single model that processed audio natively. As TechCrunch reported at the time, the rollout came with five new voices — Arbor, Maple, Sol, Spruce, and Vale — alongside improved accent handling and smoother conversations.
The feature also launched on the web in November 2024, extending it beyond mobile. But Advanced Voice Mode still operated through discrete, alternating turns — and it launched into the shadow of a PR debacle that OpenAI is still working to leave behind.
Advanced Voice Mode arrived in the wake of one of OpenAI’s most damaging self-inflicted crises. During the GPT-4o launch in May 2024, the company showcased a voice called “Sky” that many listeners immediately noted sounded strikingly similar to Scarlett Johansson, who famously voiced an AI companion in the 2013 film Her.
Johansson said she had declined OpenAI CEO Sam Altman’s offer to voice the system, then was “shocked, angered and in disbelief” when the product launched with a voice her own friends couldn’t distinguish from hers, as NBC News reported. Altman had tweeted just the word “her” the day the product launched.
OpenAI pulled the voice and apologized, but the incident drew public scrutiny from SAG-AFTRA and members of Congress, and crystallized broader concerns about AI companies moving fast with creative IP.
The Hollywood labor union said the issue underscored “why we’re strongly championing federal legislation that would protect their voices and likenesses … from unauthorized digital replication,” as NBC News reported. Forbes contributor Paul Tassi wrote at the time that Altman, “by holding up Her on a pedestal of something to strive for, has missed the point of that film” — in which the protagonist’s relationship with his AI companion ultimately does him more harm than good.
GPT-Live appears designed, in part, to move past those controversies. OpenAI says it has “remastered the nine distinct voices in ChatGPT for GPT-Live” and notes the system “is designed for conversation, not voice impersonation,” with “safeguards to prevent it from imitating a real person’s voice.”
OpenAI disclosed that more than 150 million people talk to ChatGPT using voice and dictation features each week — a notable slice of the platform’s 900 million total weekly active users. The voice experience has grown into a substantial product in its own right, used for language practice, bedtime stories, commute-time chat, and hands-free everyday help.
The new product features reflect that usage. GPT-Live introduces rich visual cards that surface during voice conversations — weather forecasts, stock data, sports scores, and maps — giving users something to glance at without breaking the flow of speech.
Users can now choose between three reasoning levels for answers: Instant for quick responses, Medium for moderate thinking, and High for more complex work. And if you take a moment to think, “ChatGPT Voice now waits instead of jumping in and interrupting,” OpenAI wrote. “If you ask it to stay quiet and listen, it will. And when there’s background noise, like passing traffic or nearby conversations, ChatGPT is better at focusing on your voice instead of getting distracted.”
Early reactions from users with preview access were cautiously positive. “I had early access to sol. it is a phenomenal model,” wrote one user on X, adding it is “much better at frontend, long context knowledge work, and its vibes are much better.” Another observer cut to the heart of the matter: “The smarts are not new here, GPT-Live hands hard questions to GPT-5.5. What is new is the feel: full-duplex voice that listens while it talks.”
The GPT-Live system card, published alongside the announcement, reveals a safety strategy built around the particular risks of real-time voice interaction — a domain where the speed and intimacy of conversation create hazards that text-based chat does not.
OpenAI expanded its safety evaluations to include audio-native tests, using both real user voice samples (from those who opted in) and synthetically generated prompts targeting edge cases across categories like self-harm, sexual content, illicit behavior, emotional reliance, mental health, and hate speech.
On the synthetic evaluations — which OpenAI described as deliberately adversarial — GPT-Live-1 showed substantial improvements over Advanced Voice Mode. In illicit behavior, for instance, the safety score rose from 0.63 to 0.97. On self-harm, it climbed from 0.72 to 0.98. Hate speech achieved a perfect 1.00, up from 0.87.
On the production-prompt evaluations — which used real user audio and reflected more ambiguous, borderline scenarios — the picture was more mixed. GPT-Live-1 matched or improved on Advanced Voice Mode in most categories but showed a slight regression on emotional reliance (from 0.88 to 0.82), though OpenAI noted the change was not statistically significant.
The company built real-time safeguards that can intervene while the model is speaking — steering toward safer responses, surfacing crisis resources, or ending the voice conversation entirely in higher-risk situations. It also designed additional protections for teen users and adapted self-harm support flows for voice, including crisis helpline integration.
Perhaps most notably, OpenAI said it is “rolling out longer-term measurement and post-launch monitoring focused on emotional reliance” — an acknowledgment that the very naturalness GPT-Live strives for creates its own category of risk.
While OpenAI was refining its safety guardrails, its rivals were shipping full-duplex systems of their own. Google’s Gemini Live, which supports full-duplex conversation alongside camera and screen sharing — capabilities GPT-Live notably lacks at launch — is already available in the Gemini app. Google released Gemini 3.1 Flash Live in March as its highest-quality real-time audio model, targeting low-latency voice interactions for developers.
ByteDance launched Seeduplex in April, claiming to be the first production-scale full-duplex speech AI deployed at scale, inside its Doubao app. Seeduplex reported roughly a 50 percent reduction in false-response and false-interruption rates compared to ByteDance’s previous half-duplex system. And Nvidia’s PersonaPlex, released in January, brought customizable voice and role control to full-duplex models, breaking what had been a constraint where natural-sounding models were locked into a single fixed voice.
The competitive picture is clear: full-duplex voice interaction is quickly becoming table stakes for consumer AI products, not a differentiator. OpenAI’s advantage lies in the scale of its existing user base, its integration with GPT-5.5’s reasoning capabilities, and the breadth of the ChatGPT ecosystem.
But the window in which any one company has a monopoly on natural-sounding voice AI has already closed. OpenAI also acknowledged several gaps. GPT-Live does not support voice with video or screen sharing at launch. Language support is limited, with the company noting that “for certain languages, the model may have a non-native accent or gaps in fluency.” And API access is not available on day one, meaning enterprise developers cannot yet build on GPT-Live directly — a constraint that will slow the model’s penetration into commercial voice-agent workflows where competitors like Google, ElevenLabs, and Deepgram already have developer-facing products.
GPT-Live is essentially OpenAI’s most significant bet yet on voice as the primary interface for AI — not just a convenience feature bolted onto a text chatbot, but a purpose-built interaction layer that sits between the user and the company’s most powerful models.
“Over time, we believe this research will also unlock the ability to use voice for increasingly complex, longer-running, and more agentic work,” OpenAI wrote. That ambition — using natural voice as the front end for autonomous AI agents that can perform multi-step tasks — is the logical endpoint of the full-duplex plus delegation architecture.
Imagine telling your phone to book a flight, negotiate with your insurance company, or debug a production server, all through a conversation that feels as natural as talking to an assistant who also happens to have the intelligence of a frontier AI model.
Two years ago, talking to ChatGPT meant dictating into a microphone and waiting nearly two seconds for a stilted reply. One year ago, it meant a smoother exchange that still felt like a polite, slightly awkward phone call with someone who insisted on waiting for you to finish every sentence. Today, it means something closer to a real conversation — imperfect, still constrained in some languages and missing video, but unmistakably closer. OpenAI once got into trouble for wanting to recreate the movie Her. With GPT-Live, the company may finally be reckoning with the harder question the film actually posed: not whether AI can sound human enough to talk to, but what happens to us when it does.

SpaceXAI has launched Grok 4.5, its most capable model yet. It is the company’s first release since going public and buying the AI coding startup Cursor. This is the joint model the two firms had raced to ship. Elon Musk aimed it squarely at coding and…

Tech loves a perk, but few go this far. Rilla, an AI startup that makes coaching software for sales teams, spends about $1.7m a year on housing stipends so staff can live near its New York office, Fortune reports. The trade-off is a 72-hour week. Emplo…
We have a new update on the possible Fallout season 3 release date thanks to some information from Walton Goggins about filming start.