Samsung One UI 8.5 Update Missing: Why Galaxy S25 Users Are Still Waiting Today

Samsung’s major Android update had been expected to go live on April 30 but it’s missing in action. Here’s what Samsung has said, and the latest details on its release.

iPad 12 Wait Continues: Apple Subtly Confirms No New Model For Months

If you’re waiting for the 12th-generation iPad, Apple just laid out that it’s not coming until the fall at the earliest.

xAI launches Grok 4.3 at an aggressively low price and a new, fast, powerful voice cloning suite

While Elon Musk faces off against his former colleague and OpenAI co-founder Sam Altman in court, Musk’s rival firm xAI, founded to take on OpenAI, isn’t slowing down on launching competitive new products and services.

Last night, xAI shipped a new, proprietary base large language model (LLM), Grok 4.3, and a new voice cloning suite on the web.

The new products arrive after months of tumult from xAI that saw all of Musk’s 10 original co-founders of the lab and dozens more researchers exit the firm and Grok was eclipsed on performance by many new competing LLMs from the likes of OpenAI, Anthropic, Google, and Chinese firms DeepSeek, Moonshot (Kimi), Alibaba (Qwen), z.ai, and others.

While Grok 4.3 does mark a significant leap in performance on third-party benchmarks over its direct predecessor Grok 4.2, according to the independent AI model evaluation firm Artificial Analysis, it still remains below the state-of-the-art set by OpenAI and Anthropic’s latest models.

But the marquee feature of the Grok brand has — other than Musk’s stated opposition to “wokeness” and its more freewheeling personality and image generation policy — increasingly been its low price point when accessed by developers and users via the xAI application programming interface (API), a trend only furthered by Grok 4.3, which costs $1.25 per million input tokens and $2.50 per million output tokens (up to 200,000 input tokens, at which point costs double, a common pricing strategy of leading AI labs) compared to its direct predecessor Grok 4.2’s initial API pricing of $2/$6 per million input/output tokens.

According to xAI’s release notes, Grok 4.3 began beta testing in April for subscribers to xAI’s SuperGrok ($30 monthly) plan, and those of its sibling social network, X, through its Premium+ plan ($40 monthly with 50% for first two months). Now it’s available to all through the xAI API and through partner OpenRouter.

Reasoning baked-in and agentic tool-use capabilities

At the core of Grok 4.3 is a fundamental shift in how the model processes information. Unlike previous iterations where “chain-of-thought” or reasoning could often be toggled or configured by effort levels, Grok 4.3 is built with reasoning as an active, permanent state.

This means the model is designed to “think” before it speaks for every query, a strategy intended to maximize factual accuracy and the handling of complex, multi-step instructions.

The model’s memory is equally expansive, featuring a 1 million-token context window. To put this in perspective, a million tokens is roughly equivalent to several thick novels or the entire codebase of a mid-sized application.

This allows Grok 4.3 to maintain coherence over massive datasets, though xAI has implemented a “Higher context pricing” structure for requests that exceed the 200,000-token threshold.

This tiering suggests that while the “long-term memory” is available, the computational cost of managing that much information remains a significant overhead.Technically, the model accepts both text and image inputs, outputting text.

It is specifically optimized for agentic workflows—scenarios where an AI is not just answering a question but acting as an autonomous agent to complete a task.

For the first time, Grok has access to the same tools and environments a human professional would use. Evidence of this shift is visible in early user interactions:

  • Spreadsheet Engineering: In one instance, the model spent 6 minutes and 22 seconds in a “thought” phase to build a comprehensive OSRS Sailing Combat DPS analyzer. The resulting .xlsx file wasn’t a simple table but a multi-sheet dashboard including a “Reference_Data” set and a complex “DPS_Calculator” with formulaic auto-calculations.

  • Professional Documentation: Grok now generates formatted PDFs, such as 12-page reports on SpaceX products. These documents incorporate branding, logos, hero images, and structured tables, moving well beyond the markdown blocks of previous iterations.

  • Visual Presentations: The model can design 9-slide PowerPoint decks, utilizing a “Sandwich Structure” (dark titles/conclusions with light content) and integrating data-driven decision matrices and humor.

However, its knowledge of the world is not infinite; the release notes list a knowledge cut-off date of December 2025. Yet, thanks to built-in web search, Grok can reference and use up-to-date information.

In fact, Grok 4.3 arrives with an enhanced ecosystem of tools designed to make it a functional digital employee. The xAI platform now offers a robust set of server-side tools that the model can invoke autonomously based on the complexity of the query.

  • Web and X Search: These tools allow Grok to bypass its knowledge cutoff by browsing the live internet or searching X (formerly Twitter) posts, user profiles, and threads.

  • Code Execution: The model can run Python code in a sandboxed environment to solve mathematical problems or process data.

  • File and Collections Search: A built-in Retrieval-Augmented Generation (RAG) system allows users to query uploaded document collections or search through specific file attachments.

xAI’s Custom Voices let you clone your voice at high quality in a minute or two

Beyond text, xAI has introduced Custom Voices, a sophisticated voice-cloning API and web-based voice cloning creation suite.

This product allows developers to clone a voice from a reference audio clip as short as 120 seconds. Once cloned, the “voice ID” can be used across xAI’s Text-to-Speech (TTS) and Voice Agent APIs.

xAI’s documentation emphasizes that this is not merely about timbre; the model is designed to pick up delivery patterns.

If a user records a reference clip in a “customer support” style, the resulting AI voice will mimic that helpful, professional inflection.

Despite the creative potential, xAI has placed strict geographic limits on this feature, making it available only in the United States, with a notable exception for Illinois due to regional biometric and privacy regulations.

While the console playground is open for general use, programmatic access via the POST /v1/custom-voices endpoint is currently gated to teams on an Enterprise plan.

I tried it myself and after moving through the requisite voice sampling screens on the web — the tool asks you to read aloud several passages of unrelated dialog — I indeed had a copy of my voice that sounded eerily identical to mine and accurately pronounced new words the same way I would when reading allowed from a new script it was given.

You can delete your custom voices in one click on xAI’s Custom Voices web application and create up to 30 new ones at a time.

In terms of licensing, the Custom Voices feature is strictly “scoped to your team” and is never made available to other users, ensuring a private, commercial license for corporate assets.

Access to the new Voice Agent API (grok-voice-think-fast-1.0) is billed at a flat rate of $3.00 per hour ($0.05 per minute) for speech-to-speech interactions. This is on the low-medium end of costs for other competing voice agents, according to my research:

Service

Price per 1k Characters

Estimated Cost per Minute

Estimated Cost per Hour

OpenAI TTS (Standard)

$0.015

~$0.015

~$0.90

OpenAI TTS (HD)

$0.030

~$0.030

~$1.80

Grok Voice Agent

$0.05

$3.00

ElevenLabs (Starter)

~$0.30

~$0.30

~$18.00

ElevenLabs (Pro)

~$0.18

~$0.18

~$10.80

Play.ht

~$0.20

~$0.20

~$12.00

Azure/Google Cloud

$0.016 – $0.024

~$0.02

~$1.00 – $1.50

Complementing this is the standalone Text-to-Speech (TTS) service, which offers five distinct voices (Eve, Ara, Rex, Sal, and Leo) and is priced at $4.20 per 1 million characters.

For transcription needs, the Speech-to-Text (STT) API provides real-time streaming at $0.20 per hour, while batch processing is available at a discounted rate of $0.10 per hour.

To ensure security for client-side applications, xAI utilizes Ephemeral Tokens, allowing for secure WebSocket connections without exposing primary API keys.

Once created, these voices are private to the user’s team and can be used across all voice APIs by referencing a unique 8-character alphanumeric voice_id.

For highly regulated sectors, xAI maintains production-ready standards, including SOC 2 Type II auditing, HIPAA eligibility for healthcare workloads, and GDPR compliance.

Aggressively low API pricing as a differentiator

The most aggressive aspect of the Grok 4.3 announcement is its pricing structure. Bindu Reddy, CEO of enterprise assistant startup Abacus AI noted on X that the model is “as smart as Sonnet 4.6 and 5x cheaper and faster”.

The standard API rates are set at $1.25 per million input tokens and $2.50 per million output tokens. This reflects a significant reduction in cost compared to its predecessor, Grok 4.20, with Artificial Analysis reporting an approximately 40% lower input price and 60% lower output price.

According to our calculations at VentureBeat, that places Grok-4.3 firmly in the lowest cost half of all major foundation models, far closer to Chinese open source offerings than its U.S. proprietary rivals:

Model

Input

Output

Total Cost

Source

MiMo-V2.5 Flash

$0.10

$0.30

$0.40

Xiaomi MiMo

Grok 4.1 Fast

$0.20

$0.50

$0.70

xAI

MiniMax M2.7

$0.30

$1.20

$1.50

MiniMax

MiMo-V2.5

$0.40

$2.00

$2.40

Xiaomi MiMo

Gemini 3 Flash

$0.50

$3.00

$3.50

Google

Kimi-K2.5

$0.60

$3.00

$3.60

Moonshot

Grok 4.3

$1.25

$2.50

$3.75

xAI

GLM-5

$1.00

$3.20

$4.20

Z.ai

GLM-5-Turbo

$1.20

$4.00

$5.20

Z.ai

DeepSeek V4 Pro

$1.74

$3.48

$5.22

DeepSeek

GLM-5.1

$1.40

$4.40

$5.80

Z.ai

Claude Haiku 4.5

$1.00

$5.00

$6.00

Anthropic

Qwen3-Max

$1.20

$6.00

$7.20

Alibaba Cloud

Gemini 3 Pro

$2.00

$12.00

$14.00

Google

GPT-5.4

$2.50

$15.00

$17.50

OpenAI

Claude Opus 4.7

$5.00

$25.00

$30.00

Anthropic

GPT-5.5

$5.00

$30.00

$35.00

OpenAI

However, the “reasoning” nature of the model introduces a new billing category: Reasoning tokens.

These are tokens generated during the model’s internal thinking process and are billed at the same rate as standard completion tokens. Effectively, users pay for the AI to “think” before it provides the final answer. xAI has also introduced several unique fee structures:

  • Prompt Caching: Repeated prompts are significantly cheaper, at $0.20 per million tokens, incentivizing developers to reuse context.

  • Tool Invocations: While token usage for tools is billed at standard rates, the act of invoking a tool carries a flat fee—$5.00 per 1,000 calls for Web Search or Code Execution, and $10.00 for File Attachments.

  • Usage Guideline Violation Fee: In a move that may set a new industry precedent, xAI charges a $0.05 fee for requests that are blocked by their safety filters before generation even begins.

The model itself remains accessible via a standard commercial API, with xAI recommending that all developers migrate to grok-4.3 as their “most intelligent and fastest model”.

Third-party benchmark evaluations and analysis

The reception of Grok 4.3 has been polarized, depending largely on the specific use case. Professional benchmarkers and developers have highlighted a “stark gap” between the model’s domain-specific strengths and its general reasoning consistency.

According to independent AI evaluation firm Vals AI, Grok 4.3 has taken the top spot on several specialized indices. It currently ranks #1 on CaseLaw v2 (79.3% accuracy) and #1 on CorpFin.

This 25-point jump in legal reasoning over Grok 4.20 suggests that the “always-on reasoning” architecture is particularly well-suited for the dense, logical structures of law and finance.

Artificial Analysis corroborated this performance, noting a massive improvement in agentic tasks, scoring an Elo of 1500 on the GDPval-AA benchmark, surpassing competitors like Gemini 3.1 Pro and GPT-5.4 mini.

Conversely, users focused on general-purpose agents and coding have highlighted deficiencies.

AI automated brick-and-mortar retail company Andon Labs reported that Grok 4.3 is a “big regression” on the Vending-Bench 2, which measures an AI’s ability to take consistent actions in a simulation.

They colorfully described the model as having “narcolepsy problems,” preferring to remain inactive for multiple simulation days rather than taking the required actions.

The sentiment was echoed by Vals AI, which noted that while the model improved in some coding areas, it remains weak on general coding tasks and “struggles with difficult math problems,” scoring only 11% on ProofBench.

Should your enterprise use Grok 4.3?

The launch of Grok 4.3 represents a calculated bet by xAI that the market wants specialized brilliance and extreme cost efficiency over a perfectly balanced generalist.

By achieving a score of 53 on the Artificial Analysis Intelligence Index while remaining on the “Pareto frontier” of cost-per-intelligence, xAI is positioning itself as the “value” leader for enterprise applications in legal and financial tech.

The “always-on reasoning” is a double-edged sword. While it provides the depth needed to navigate complex case law, the community reports of “narcolepsy” suggest that a model that is always “thinking” may occasionally think itself into a state of paralysis, or at least a state of excessive caution that inhibits agentic action.

In addition, prior Grok model scandals including an X chatbot version referring to itself as “MechaHitler” and posting antisemitic content, sexualized deepfake imagery generation and investigations, and references to racial conflicts and right-wing dog whistle framing of social issues — which appear to mirror many of founder Musk’s own positions, to the point that the model was at one point, checking Musk’s own X account before responding in its X implementation — nearly certain to give some enterprises pause when considering adoption. It’s unclear whether any of those issues remain with Grok 4.3, but one user did note that Grok’s system prompt appears to instruct it “you do not assign broad positive/negative utility functions to groups of people.”

For developers, the decision to adopt Grok 4.3 will likely come down to the nature of their data. For those needing to process a million tokens of legal documents at a fraction of the cost of Claude 4.6 or GPT-5.5, Grok 4.3 is a clear front-runner.

For those building high-frequency autonomous agents or complex math solvers, the “narcolepsy” and coding regressions suggest that xAI’s latest model may still need a few more “tuning passes”.

As OpenRouter noted on X upon making the model live, the “large jump in agentic performance” at a lower price point is an undeniable milestone. Whether that performance can be sustained across all domains remains the primary question for the summer of 2026.

Thinner than a hair and stretchy like rubber: New material could shield against radiation in next-gen space tech

Scientists have developed a new material that could shield humans and critical technology from harmful radiation.

Hidden IT problems are quietly creating risk, shadow IT, and lost productivity

Presented by TeamViewer


Enterprise technology failures are largely invisible. Research from TeamViewer, based on a global survey of 4,200 managers and employees, finds that the majority of digital dysfunction never reaches the IT help desk.

Employees work around slow applications, failed logins, and intermittent glitches rather than reporting them, leaving organizations without an accurate picture of how their technology is performing. The cumulative cost is significant: employees lose an average of 1.3 workdays per month to digital friction, with impacts ranging from delayed projects and lost revenue to increased employee turnover.

The research, which surveyed managers and employees across nine countries, confirms what many have long suspected: the productivity loss from digital friction is significant, and most of it never surfaces in an IT support queue, says Andrew Hewitt, VP of strategic technology at TeamViewer.

“Enterprise outages are visible because they trigger clear, system-level failures,” Hewitt says. “But much of the real disruption happens earlier, in the form of digital friction: slow apps, login issues, or intermittent glitches that don’t cross alert thresholds. These smaller issues often go unreported or are normalized by employees, even though they quietly drain productivity.”

What is digital friction and why does it go unreported?

The most common sources of friction — connectivity failures, software crashes, hardware problems, and authentication issues — aren’t edge-case scenarios, but everyday experiences employees have learned to absorb without escalating. Connectivity problems were the most widespread, with nearly half identifying them as the top productivity killer among common technology issues.

That tendency to absorb rather than report is central to the problem. Many workers don’t trust their IT team to resolve issues quickly or effectively, so when a login fails or an application stalls mid-task, the path of least resistance is to restart the device, switch tools, or use a personal phone.

“Employees are under more pressure than ever to prove output,” Hewitt says. “When reporting feels unlikely to result in a quick resolution, it creates a false sense of stability at the system level while the employee experience quietly deteriorates.”

How much productivity does digital friction cost organizations?

The business consequences extend beyond inconvenience. Many organizations report delays in critical operations, revenue loss, and lost customers as a result of IT dysfunction. Most respondents lose time each month, and few expect improvement, citing increasing complexity of workplace technology as a primary concern.

The human cost runs parallel. Workers link digital friction to frustration, decreased motivation, and burnout, and many believe it contributes to turnover, with onboarding replacements stretching to eight weeks or more.

“Employees are happiest when they feel productive and accomplished at the end of the day,” Hewitt says. “When people can’t make progress in their day-to-day work, frustration builds and burnout follows. Great technology might not be a main attractor of talent, but bad technology can certainly play a role in driving it away.”

Why employees use personal devices and unauthorized tools instead of reporting IT problems

When workplace technology consistently fails to meet employee needs, workers find alternatives, with a substantial share of respondents admitting to using personal devices or unauthorized applications as workarounds. That’s the entry point for shadow IT, or the use of unapproved hardware, software, or cloud services outside IT’s visibility and control. While employees turn to these tools simply to stay productive, they introduce security vulnerabilities, data leakage risks, and compliance gaps that IT teams may not discover until a breach occurs.

“Quite simply, it demonstrates that the IT environment is not meeting the employees’ needs,” Hewitt said. “While this helps maintain short-term productivity, it introduces significant risks and pushes work outside of IT’s visibility and control.”

TeamViewer ONE addresses this by combining remote connectivity with real-time endpoint monitoring, giving IT teams the ability to detect and resolve device and application issues before employees reach for an alternative. When the underlying environment is stable and support is fast, the impulse to work around it diminishes.

How fragmented IT infrastructure creates blind spots across devices, apps, and networks

Addressing digital friction at scale requires more than faster help desk response times. Traditional metrics such as mean time to resolution and ticket volume capture only a fraction of actual issues. A more complete picture requires measuring lost time, interrupted workflows, and employee sentiment across devices, applications, and network environments.

“Leaders need to move beyond measuring performance through IT tickets alone,” Hewitt said. “Performance should be viewed through the lens of employee experience and real-time digital workplace data.”

Fragmented infrastructure makes this difficult. When devices, applications, and networks operate in separate silos, IT teams struggle to trace root causes or identify systemic issues before they spread, often responding to symptoms rather than underlying problems.

TeamViewer ONE is designed to close that gap, integrating digital employee experience analytics, remote support, and device management into a single platform. Instead of piecing together signals from disconnected tools, IT teams get a consolidated view of endpoint health, application performance, and network conditions across the entire organization.

How organizations can shift from reactive IT support to proactive system monitoring

Achieving proactive IT is not a single-step transformation. Hewitt describes it as a progression: starting with endpoint management and security, building toward real-time visibility into the digital employee experience, and ultimately using automation and AI to resolve issues before they reach employees.

TeamViewer AI is built to support each stage of that progression, using continuous monitoring to surface anomalies and correlate signals across the digital environment, identifying patterns of poor experience before they escalate. When issues are detected, it suggests remediations, generates scripts to fix problems autonomously, and handles routine tasks such as common troubleshooting without requiring IT intervention, shifting the workload from reactive firefighting toward proactive oversight.

And while AI’s effectiveness depends on the completeness of the data it works with, consolidating onto a platform like TeamViewer ONE removes that limitation by giving AI a complete, real-time data foundation to work from.

How system performance lays the foundation for productivity, retention, and competitive advantage

TeamViewer ONE isn’t a wholesale replacement of existing IT infrastructure, but a unifying layer that connects insight with action, which enables organizations to ramp up productivity, improve retention, and ultimately realize a significant competitive advantage. It begins with visibility into what is actually causing friction across their environment. From there, leaders can use that data to prioritize fixes, and then scale remediation through automation as confidence and capability grow.

“Reducing digital friction isn’t about overhauling everything at once,” Hewitt said. “Leaders should start small, gain visibility into what’s actually causing friction, fix the biggest pain points, then scale those improvements through automation and AI. Even incremental progress can make an impact on employee engagement and productivity.”

Dig deeper: Fix it before they feel it from TeamViewer.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Motorola Beats Apple And Samsung: The Razr Ultra 2026 Debuts Innovative Battery Tech

Motorola’s latest is the first major U.S.-focused smartphone brand to use a carbon-silicon battery.

‘Escape From Tarkov’ Opens ETS Test Servers And Adds Aiming Rework

One of the most controversial new features in Tarkov is currently being tested on the ETS test servers and now you can try it out.

Why OpenAI’s ‘goblin’ problem matters — and how you can release the goblins on your own

AI is more than a technology — it’s magic.

Don’t believe me? Why, then, is one of the leading companies in the space, OpenAI, publishing entire official, corporate blog posts about goblins?

To understand, we first have to go back to earlier this week, on Monday, April 27, 2026, when a developer under the handle @arb8020 on the social network X posted a snippet from the OpenAI open source Codex GitHub repository, specifically a file named models.json.

Deep within the instructions for the new OpenAI large language model (LLM) GPT-5.5, a peculiar directive stood out, repeated four times for emphasis:

“Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user’s query.”

The discovery sent a shockwave through the “power user” and machine learning (ML) researcher circles.

Within hours, the post had gone viral, not because of a security flaw, but because of its sheer, baffling specificity.

Why had the world’s leading AI laboratory issued what Reddit users quickly dubbed a “restraining order” against pigeons and raccoons?

Goblin speculation abounds

The initial reaction was a chaotic blend of humor and technical skepticism. On Reddit’s r/ChatGPT and r/OpenAI, users began sharing screenshots of GPT-5.5’s behavior prior to the patch.

Barron Roth, a Senior Project Manager of Applied AI at Google, shared an image on X under his handle @iamBarronRoth of his GPT-5.5 powered OpenClaw agent that seemed “obsessed with goblins.”

Others reported that the model stubbornly referred to technical bugs as “gremlins in the machine”.

Developers like Sterling Crispin leaned into the absurdity, jokingly theorizing that the massive water consumption of modern data centers was actually needed to cool “the goblins being forced to work”.

More seriously, researchers on Hacker News and beyond discussed the “Pink Elephant” problem. In prompt engineering, telling a model not to think of something often makes the concept more salient in its attention mechanism.”

“Somewhere there is an OpenAI engineer who had to type never mention goblins in production code, commit it, and move on with their day,” noted one commentator on Reddit.

The presence of “pigeons” and “raccoons” led to wild speculation: Was this a defense against a specific data-poisoning attack? Or had the reinforcement learning trainers simply been “bullied by a raccoon” during a lunch break?

The tension reached a peak when OpenAI co-founder and CEO Sam Altman joined the fray on X. On the same day as the discovery, Altman posted a screenshot of a ChatGPT prompt that read: “Start training GPT-6, you can have the whole cluster. Extra goblins.”.

While humorous, it confirmed that the “goblin” phenomenon was not a localized bug but a company-wide narrative that had reached the highest levels of leadership.

OpenAI comes clean on goblin mode

Yesterday, as the discussion continued on X and wider social media, OpenAI published a formal technical explanation titled “Where the goblins came from“.

The blog post served as a sobering look at the unpredictable nature of Reinforcement Learning from Human Feedback (RLHF) and how a single aesthetic choice could derail a multi-billion-parameter model.

OpenAI revealed that the “goblin” behavior was not a bug in the traditional sense, but a byproduct of a new feature: personality customization, which it introduced for users of ChatGPT back in July 2025, but has maintained and updated ever since.

Apparently, this feature is not added after the model is finished post-training, but rather, OpenAI bakes it in as part of its underlying GPT-series model end-to-end training pipeline.

The feature allows ChatGPT users or GPT-based developers to choose from several distinct modes, such as Professional for formal workplace documentation, Friendly for a conversational sounding board, or Efficient for concise, technical answers. Other options include Candid, which provides straightforward feedback; Quirky, which utilizes humor and creative metaphors; and Cynical, which delivers practical advice with a sarcastic, dry edge.

While these personalities guide general interactions, they do not override specific task requirements; for example, a request for a resume or Python code will still follow professional or functional standards regardless of the selected personality.

The selected personality operates alongside a user’s saved memories and custom instructions, though specific user-defined instructions or saved preferences for a particular tone may override the traits of the chosen personality.

On both web and mobile platforms, users can modify these settings by navigating to the Personalization menu under their profile icon and selecting a style from the Base style and tone dropdown. Once a change is made, it is applied globally across all existing and future conversations. This system is designed to make the AI more useful or enjoyable by tailoring its delivery to individual user preferences while maintaining factual accuracy and reliability.

OpenAI states that the goblin issue actually originated several years ago, during training of a since-discontinued “Nerdy” personality designed to be “unapologetically quirky” and “playful”.

During the RLHF phase, human trainers (and reward models) were instructed to give high marks to responses that used creative, wise, or non-pretentious language. Unknowingly, the trainers began over-rewarding metaphors involving fantasy creatures. If the model referred to a difficult bug as a “gremlin” or a messy codebase as a “goblin’s hoard,” the reward signal spiked. The statistics provided by OpenAI were staggering:

  • Use of the word “goblin” rose by 175% after the launch of GPT-5.1.

  • Mentions of “gremlin” rose by 52%.

  • While the “Nerdy” personality accounted for only 2.5% of ChatGPT traffic, it was responsible for 66.7% of all “goblin” mentions.

The mechanics of ‘transfer’ and feedback loops

The most significant finding for the ML community was the confirmation of learned behavior transfer. OpenAI admitted that although the rewards were only applied to the “Nerdy” condition, the model “generalized” this preference.

The reinforcement learning process did not keep the behavior neatly scoped; instead, the model learned that “creature metaphors = high reward” across all contexts.This created a destructive feedback loop:

  1. The model produced a “goblin” metaphor in the Nerdy persona.

  2. It received a high reward.

  3. The model then produced similar metaphors in non-Nerdy contexts.

  4. These “goblin-heavy” outputs were then reused in Supervised Fine-Tuning (SFT) data for subsequent models like GPT-5.4 and GPT-5.5.

By the time the researchers identified the issue, the “goblin tic” was effectively “baked in” to the model’s weights.

This explained why GPT-5.5 continued to obsess over creatures even after the “Nerdy” personality was retired in mid-March 2026.

How you can let the goblins run free (if you want)

Because GPT-5.5 had already completed much of its training before the “goblin” root cause was isolated, OpenAI had to resort to the blunt-force “system prompt” mitigation that @arb8020 discovered on X.

The company referred to this as a “stopgap” until GPT-6 could be trained on a filtered dataset.

In a surprising nod to the developer community, OpenAI’s blog post included a specific command-line script for Codex users who find the goblins “delightful” rather than annoying.

By running a script that uses jq and grep to strip the “goblin-suppressing” instructions from the model’s cache, users can now effectively “let the creatures run free”.

The blog post also finally explained the specific list of banned animals. A deep search of GPT-5.5’s training data found that “raccoons,” “trolls,” “ogres,” and “pigeons” had become part of the same “lexical family” of tics.

Curiously, the model’s use of “frog” was found to be mostly legitimate, which is why it was spared from the system prompt’s exile list.

What it means for AI research, training and implementation going forward

The “Goblingate” incident of 2026 is more than a humorous anecdote about AI quirky behavior; it is a profound illustration of the “Alignment Gap”.

It demonstrates that even with sophisticated RLHF, models can latch onto “spurious correlations”—mistaking a stylistic quirk for a core requirement of performance.

For the AI power user community, the response transitioned from mocking the “restraining order” to a more somber realization.

If OpenAI can accidentally train its flagship model to obsess over goblins, what other more subtle and potentially harmful biases are being reinforced through the same feedback loops?

As Andy Berman, CEO of the agentic enterprise AI orchestration company Runlayer wrote on X today: “OpenAI rewarded creature metaphors while training one personality. The behavior leaked across every personality. Their fix: a system prompt that says ‘never talk about goblins.’ RL rewards don’t stay where you put them. Neither do agent permissions”

As the technical discourse continues, “Goblingate” remains the primary case study for a new era of behavioral auditing.

The investigation resulted in OpenAI building new tools to audit model behavior at the root, ensuring that future models—specifically the much-anticipated GPT-6—do not inherit the eccentricities of their predecessors.

Whether GPT-6 will indeed be free of goblins remains to be seen, but as Altman’s “extra goblins” post suggests, the industry is now fully aware that the machines are watching what we reward, even when we think we’re just being “nerdy.”

Writer launches AI agents that can act without prompts, taking on Amazon, Microsoft and Salesforce

Writer, the enterprise AI agent platform backed by Salesforce Ventures, Adobe Ventures, and Insight Partners, today launched event-based triggers for its Writer Agent platform, enabling AI agents to autonomously detect business signals across Gmail, Gong, Google Calendar, Google Drive, Microsoft SharePoint, and Slack — and execute complex multi-step workflows without any human initiating the process.

The release, which also includes a new Adobe Experience Manager connector and a suite of enhanced governance controls such as bring-your-own encryption keys and a Datadog observability plugin, represents Writer’s most aggressive bet yet on fully autonomous enterprise AI. It arrives at a moment when AWS, Salesforce, and Microsoft are all racing to establish their own agentic platforms, and when the question of how much autonomy enterprises will actually hand to AI agents remains deeply unresolved.

“We are launching a series of event triggers that power and drive our playbooks to be more proactively called,” Doris Jwo, Writer’s VP of Product Management, told VentureBeat ahead of the announcement. “We’re building on the ecosystem to actually for these connectors, such as SharePoint, Google Drive, Gong, Gmail, Google Calendar, actually listen for events happening in those platforms, so that the agent can practically know that something happened externally, and then, where relevant, call a certain playbook to be actually run live in real time, without any sort of human intervention required.”

The shift from reactive to proactive AI agents marks a critical inflection point for enterprise software. Until now, most AI assistants — including Writer’s own platform — required a human to initiate every interaction. A marketer had to open a chat window and ask for help. A salesperson had to prompt a research brief. The new event-based triggers flip that dynamic entirely: the system watches for business events and acts on its own.

Why Writer decided humans were the weakest link in enterprise AI workflows

Writer’s push toward autonomous triggers stems from a practical observation its product team made as enterprise customers scaled their use of the platform’s playbooks — the reusable, natural-language workflows that Writer introduced in November 2025 to let business users automate recurring tasks without writing code.

“What we found is, as playbooks continue to get integrated into enterprise workflows, it’s actually humans that become the bottleneck in making sure that playbooks get triggered,” Jwo said. “This really kind of solves that problem, to make sure that that sort of always-on, proactive, autonomous nature of that agent has continued to be built on.”

The mechanics work like this: Writer’s connectors, which already provided read and write access to third-party enterprise tools, now also listen for specific events — an email arriving in Gmail, a sales call completing in Gong, a new file landing in a Google Drive folder, a meeting starting or ending on Google Calendar, a message posted in Slack. When the system detects a qualifying event, it triggers a predefined playbook that executes a multi-step workflow autonomously.

Consider the use case Jwo described for marketing teams already running on Writer’s platform. An email campaign workflow typically begins when a creative brief lands in a Google Drive folder. From there, multiple team members coordinate through Slack to assemble research, build assets, draft copy, review graphics, and package everything for a campaign management tool. Writer’s event-based triggers collapse much of that chain: the moment a brief hits the designated folder, the system automatically fires a cascade of playbooks that assemble the research, generate the assets, and prepare deliverables for human review.

“All the playbooks that our customers have been building with us to build all those each individual pieces now just get automatically triggered the minute that initial brief kind of hits the Google Drive folder,” Jwo said. “That’s, I think, a very common workflow for most of these marketing sort of, like, content-heavy use cases, where it’s multiple parties involved, it’s a lot of assets coming together in a cascade.”

How Writer’s AI reasoning engine separates it from simple automation tools like Zapier

The comparison to Zapier — the popular automation tool that connects thousands of apps through if-this-then-that logic — is inevitable, and Jwo addressed it directly.

“It’s more than just an LLM in the middle,” she said. “It is an agent with reasoning and then access to a really powerful set of tools that includes connectors, that includes its own virtual sandbox, which enables it to do things like write and execute code on the fly and create those assets.”

The distinction matters for understanding where Writer sits in an increasingly crowded landscape. Zapier and similar workflow automation tools require users to manually define rigid logic paths, specifying exact conditions and actions in a deterministic sequence. Writer’s approach uses its Palmyra-powered reasoning engine to process event context and make real-time execution decisions. Users describe their goals in natural language rather than dragging around boxes and defining conditional branches.

“It’s not quite Zapier, because I think it requires a lot more — it’s more rigid,” Jwo said of traditional automation tools. “It requires more manual kind of setup to define the logic and the roles and the conditions for which a workflow has to be run.” Writer’s playbooks, by contrast, allow “a simple idea to turn into something that’s actually executable and repeatable,” she added, noting that builds take “hours and days, not weeks and months.”

This natural-language accessibility has been central to Writer’s strategy since it introduced the Agent platform and playbooks last November. The company has consistently positioned itself as a platform that puts power in the hands of business users — marketers, sales teams, operations leads — rather than requiring engineering resources to build and maintain AI workflows. Writer CEO May Habib made this case forcefully at Davos earlier this year, arguing that the leaders pulling ahead are those entering what she called “rebuild mode” — stripping workflows down to outcomes and eliminating what she described as the “coordination tax” of endless handoffs, status meetings, and alignment emails.

The event-based triggers extend that philosophy to its logical conclusion. If business users can build playbooks in natural language, and those playbooks can now fire automatically based on real-world business events, then the entire loop from signal to action can operate with minimal human involvement.

Inside the governance controls Writer built to make autonomous AI agents safe for regulated enterprises

That level of autonomy raises obvious concerns, and Writer appears to understand that governance is the linchpin of the entire strategy. The company paired its trigger launch with a substantial expansion of its administrative controls — a combination that suggests Writer views enterprise trust as its primary competitive weapon.

The new governance features include Connector Profiles, which allow administrators to configure multiple versions of the same connector with different permissions per team; Writer Agent Profiles for deploying customized agent configurations with specific capability toggles and security settings; AI Studio Observability for auditable tracking of every agent interaction; a Datadog Logs Plugin that forwards every LLM request and response as structured log events; and bring-your-own encryption key support through AWS, Azure, or GCP key management services.

“A really important part of that, and a baseline, sort of foundation for everything that we roll out, is our observability and governance platform,” Jwo told VentureBeat. “When connectors are set up, admins have full control over connector access, what is set up, who has access, which teams exactly are those access granted to, as well as individually, which exact tools do teams are able to call.”

The observability story extends to the individual user level as well. Jwo described Writer Agent’s user experience as built around progressive disclosure — clean initial views that users can expand to inspect the full chain of reasoning behind any agent action. “You can drill down to the actual tool call level,” she said. “You’d actually have the ability to look at specifically what web search results were pulled, what connector was called, what tool called, what succeeded, what failed, how did the agent divert its path to fulfill your goal.”

This transparency architecture reflects a broader conviction Writer has articulated through what it calls “The Agentic Compact” — a framework the company published for responsible AI that emphasizes foundational transparency, auditability, and human oversight. Dan Bikel, Writer’s head of AI, has argued publicly that the industry’s obsession with model scale has created what he calls a “transparency paradox,” leaving businesses with powerful tools they cannot fully understand or control. Writer’s governance-first approach to autonomous triggers represents the operational expression of that philosophy.

Writer also introduced its agent supervision suite in December 2025, offering centralized monitoring, agent approval workflows, global guardrails, and integrations with external observability and security platforms like Datadog, Noma, and Lakera. The event-based triggers now extend that governance framework to cover actions initiated without any human in the loop — a meaningfully harder problem.

Writer takes aim at AWS, Salesforce, and Microsoft in the escalating agentic platform wars

The timing of Writer’s announcement is not accidental. The enterprise agentic AI market has entered a period of intense platform competition, with the largest technology companies in the world staking claims to the same territory Writer occupies.

Jwo acknowledged the pressure directly when asked why a CIO would choose Writer over established vendor relationships with AWS, Salesforce, or Microsoft — all of which have announced agentic platforms of their own.

“At the baseline, I think we have all the pieces to be fully enterprise-grade and ready,” Jwo said. But she argued that Writer’s real advantage lies in accessibility for non-technical users. “A lot of the challenge has been: how do we get business users to actually be able to build these powerful workflows in a way that maybe a technical user, using coding agents, can do very quickly and well, but the typical business user is not accustomed to anything beyond typical prompting to actually create?”

That positioning — enterprise-grade capabilities wrapped in a business-user-friendly interface — has been Writer’s core differentiation since the company’s founding in 2020. It is also the reason Writer has attracted strategic investment from Salesforce Ventures and Adobe Ventures, both of which are building their own AI platforms but apparently see value in Writer’s approach to the business-user segment.

The company’s March 2026 release of Skills — reusable building blocks that encode a team’s specific methodologies, quality standards, and decision frameworks into the Agent platform — reinforced this direction. Skills allow marketing teams, for instance, to capture exactly how their best strategist structures competitive analysis or formats campaign briefs, then make that expertise available to every team member and every playbook across the organization. Combined with event-based triggers, the result is a system where institutional knowledge executes automatically in response to real-world business events.

Writer’s 2026 AI adoption survey, conducted with Workplace Intelligence and covering 2,400 global executives, found that 79% of enterprises face AI adoption challenges despite high investment — and that organizations with strong change management programs are six times more likely to reach production. Writer CMO Diego Lomanto has argued that the real barrier to AI adoption is not technology but trust, writing that “they treat resistance as a training problem when it’s actually a trust problem.” The governance-heavy approach to event-based triggers appears designed to address exactly that dynamic.

Salesforce, SAP, and Workday triggers are next as Writer expands its connector roadmap

Writer’s initial event trigger support covers Gmail, Gong, Google Calendar, Google Drive, SharePoint, and Slack — tools that Jwo described as “generally the most applicable to every end user.” But the company has its eye on deeper enterprise system integration.

When asked about CRM and ERP triggers for systems like Salesforce, SAP, and Workday, Jwo confirmed these are within the scope of the roadmap. “You can imagine, you know, a Salesforce opportunity is created that may trigger a cascade of events that happens,” she said. “You might want to set up the right assets, maybe the right customer environment, all sorts of things can kind of cascade from that.”

The connector ecosystem has been a strategic priority since Writer launched its MCP (Model Context Protocol) gateway in November 2025, providing governed agent access across enterprise systems including Microsoft 365, Google Workspace, HubSpot, Gong, PitchBook, FactSet, and others. The addition of Adobe Experience Manager in this release gives marketing teams direct read/write access to pages, fragments, and digital assets in Adobe’s content management system — a connector that closes the gap between AI-generated content and published output.

Jwo clarified that in most integration scenarios, Writer Agent delivers content in a draft state rather than publishing it directly. “Writer Agent basically accomplishes the majority of the workload — pulling together the assets, making the changes and presenting — and then hopefully a person just has to go through the last three or so final steps to get it out,” she said.

The real question enterprise AI must answer: how much autonomy is too much autonomy

The degree of autonomy enterprises are comfortable granting their AI agents remains one of the most consequential open questions in the industry. Jwo acknowledged that most customers still maintain human checkpoints in their workflows.

“You can also build in instructions into our playbooks to say, ‘Hey, before you move on to a next playbook, make sure that you check with me. I want to take a look, and then if I hit go, then you’re good to go,'” she said. The agent can also be designed with self-QA capabilities, validating outputs against known pitfalls before proceeding.

Writer plans to expand these checkpoint capabilities in the coming quarter, adding the ability to specify not just that a checkpoint is required but which specific person must respond and what types of responses are expected — essentially building a formal approval workflow into the autonomous trigger chain.

Jwo characterized the current system as a hybrid: the platform listens deterministically for predefined events, but the agent applies reasoning to decide what action to take — or whether to act at all. “The agent has the ability to process what happened, understand the context of it, and understand the intent of what you want to do, so it can make that decision,” she said. “You’re just saying, like, ‘Hey, the goal might be feedback is coming in, and we want to triage that in real time. And some things we might not want to action on, some things we do.’ You basically just explain that to the agent.”

She views this release as a stepping stone toward a future where agents are “even more mission-driven, and less governed by even like a set of instructions or roles” — a future where the AI doesn’t just respond to triggers but proactively identifies when action is needed based on broader organizational goals.

For now, Writer is betting that the combination of autonomous triggers, robust governance, and business-user accessibility will be enough to carve out defensible territory in an enterprise AI market where the biggest technology companies in the world are all converging on the same set of capabilities. The company’s argument is that having the foundational pieces is not enough — what matters is making those pieces work together in a way that non-technical business users can build, manage, and trust.

It is, in other words, the same wager Writer has been making since 2020 — that the future of enterprise AI belongs not to the platform with the most powerful model, but to the one that can get an entire organization to actually use it. The difference now is that the agents don’t wait to be asked.

Event-based triggers, new connectors, and enhanced governance controls are available immediately to Writer enterprise customers.

When Are Forward Deployed Engineers Essential, And When Are They Not?

Forward deployed engineers (FDEs) have become one of the most debated topics in enterprise technology, driven in large part by the visibility of Palantir’s model.