Presented by F5
Enterprise AI teams have spent years solving for compute, securing GPU allocations, negotiating cloud capacity, and benchmarking training throughput. The assumption embedded in that work is that the path between storage and compute will keep up. In production, that assumption increasingly does not hold. Real traffic introduces latency spikes, network jitter, and node degradation that controlled benchmarks fail to capture, resulting in pipelines that perform well in the lab but stall in deployment. A growing response is AI data delivery, deploying an application delivery controller (ADC) or application delivery and security platform (ADSP) in front of storage as a resilient and secure control point.
“Provisioning solves for capacity but not for delivery, and that is where the constraint now hides,” says Hunter Smit, senior manager of product marketing at F5. “Enterprises buy enough GPUs and enough storage, then assume the path between them will keep up, but AI traffic is bursty, highly concurrent, and random in its reads in ways ordinary storage networking was never built to absorb.”
Standard benchmark methodology compounds the problem, says Paul Pindell, principal solutions architect for technology alliances at F5.
“Benchmark testing is usually built to produce the best possible performance or security result, not the most realistic one,” he says. “With S3, latency is a known factor in degrading performance, so meaningful testing has to introduce consistent latency into the path.”
Most benchmark environments never do that, which means the performance numbers enterprises rely on for infrastructure decisions are drawn from conditions that production systems will never replicate. To test this assumption, F5 and MinIO conducted throughput testing under degraded network conditions.
“What stood out was how quickly S3 throughput falls off once you introduce latency,” Pindell says. “Even modest latency takes a real bite out of it, and as latency climbs toward long-haul distances, the degradation gets severe.”
The testing also showed latency mattered far more than jitter as a driver of throughput loss, which inverted what the team had expected going in. The upshot for enterprise architects is that S3 object storage deployments cannot be designed around clean-room assumptions; they have to be engineered for the degraded network conditions they will actually face.
“In AI infrastructure, people naturally focus on GPUs because they’re the most visible and expensive resource,” says Tanu Mutreja, senior director of product management at F5. “But in production environments, GPUs generate only as much value as the data path that feeds them.”
That path runs through storage, networking, databases, security, and orchestration layers, often stitched together from multiple vendors. Customers experience none of those seams; they experience the output of the whole system.
When the data path degrades, the effects compound. GPU underutilization is the most immediate and visible symptom, but Mutreja pointed to a wider set of consequences: degraded inference performance, poor-quality AI outputs, higher egress costs from unnecessary data replication, and growing operational complexity.
“At scale, data-path efficiency becomes a strategic business lever rather than technical optimization,” she says. “When the data path is engineered well, GPUs remain productive, AI applications stay responsive and trustworthy, operations scale efficiently, and organizations maximize the return on their AI investments.”
AI workloads are structurally more exposed to these failures than traditional enterprise applications. Databases, ERP systems, and web services absorb transient storage delays through caching and buffering. AI workloads running across massively parallel GPU clusters have no equivalent protection. As Mutreja noted, even minor latency spikes or bandwidth bottlenecks can cascade across large GPU clusters, simultaneously hitting utilization, training efficiency, and the customer experience.
For decades, storage and intelligence operated as sequential concerns in enterprise architecture: data was stored first, then analyzed downstream. Mutreja argued that this model no longer fits the demands of AI.
“Competitive advantage is determined not only by the volume of data, but also by relevance, lineage, security, and performant delivery of data,” she says. “Across the industry, from NVIDIA and AWS to enterprise storage providers, the movement is toward embedding intelligence directly into data infrastructure rather than stacking it on top.”
F5’s integration with MinIO instantiates this approach at the layer where storage and compute actually interact. As part of the F5 ADSP, BIG-IP sits in the data path, continuously monitoring the health of MinIO’s distributed storage nodes and directing requests only to those that remain available.
The operational impact of that capability becomes clear when nodes degrade, which is expected in distributed storage clusters. Without intelligent routing, clients that land on an unhealthy node must retry and may land on another degraded node, dragging down overall performance.
“F5 makes sure traffic only goes to healthy nodes, or even the least busy ones, so S3 client traffic is always processed in the most efficient way,” Pindell says.
The challenge grows at scale, when AI pipelines stretch across multiple locations, clouds, or edge environments.
“Once an AI pipeline crosses regions and clouds, the question stops being about performance and becomes about control,” Smit says. “You are operating under different rules in every jurisdiction, and digital sovereignty is now a design constraint. Where your data is allowed to live, who is permitted to touch it, and which borders it cannot cross now shapes the architecture before anyone talks about speed.”
That pressure is driving a visible trend of enterprises repatriating AI workloads from public cloud onto infrastructure they own and govern directly. The architecture Smit described resolves this by decoupling applications from any single storage location and placing a unified control point between them that enforces consistent policy across all of them.
“Sovereignty, resilience, and cost stop being trade-offs you manage one region at a time,” he explains. “They become a capability you run as a system.”
To solve for these issues, enterprise teams need to stop treating the storage-to-compute path as a direct connection and start treating it as a managed control point, Smit says. SecureIQLab’s independent validation of F5 BIG-IP in storage deployments has confirmed the approach delivers resilience without surrendering throughput.
“Insert a full-proxy ADC between the two, and the path becomes observable, programmable, and failure-aware, with health-based routing, quality of service, and security enforced inline,” he explains. “That single move converts data delivery from an assumption into an engineered discipline, which is what keeps GPUs fed when conditions degrade.”
Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.
Presented by Capital One Enterprises aren’t struggling to experiment with AI; they’re struggling to make it work in the real world. Moving from promising prototypes to reliable, production-scale systems is where most efforts stall.In my role within Cap…
Enterprise AI teams face a dilemma: The best models today might not be the best models a year from now. MassMutual’s answer is to stop making long-term bets — and build infrastructure that can swap models as the market shifts.
“The world of AI today is extremely dynamic,” Sears Merritt, MassMutual CIO, explained in a new VB Beyond the Pilot podcast. “We wanted to make sure we were positioned to ride that wave of dynamism.”
The strategy appears to be paying off in a big way. MassMutual has measured a roughly 30% increase in developer productivity, while AI-powered contact center workflows have reduced resolution times from 10 minutes to one and cut costs from dollars to cents.
But the broader lesson for IT leaders may be less about the results and more about how the company is thoughtfully building its AI infrastructure and keeping users at the center.
MassMutual works with vendors at the leading edge, but keeps those relationships on a clock. “Those relationships are capped so that we maintain optionality for best-of-breed tools as things mature in this space, and at some point, settle down and stabilize,” Merritt said.
That philosophy extends to open-source models. Merritt says his team is “100%” looking at open-source tools, and sees the technology playing a big role in how MassMutual (and similar companies) use AI.
“We’re certainly going to need frontier models and leading edge capabilities to do what today is impossible, and tomorrow will be possible,” he said.
MassMutual’s AI efforts fall into two broad categories.
The first focuses on enablement: Putting productivity-enhancing tools such as Copilot and virtual assistants into the hands of all employees. The second involves what Merritt describes as “deepen and focus” initiatives, where teams target a specific workflow or business process that will have a strong impact on advisors, policyholders, or employees.
Rather than focusing on adoption metrics, these projects begin with predefined success criteria. “Everything we do is measured,” Merritt said. “There’s always a success metric that we define upfront to determine whether or not we’re going to scale up some of these things.”
The company is also deliberately encouraging experimentation, giving employees access to a range of best-in-class models, “token-consumptive workflows” and other possible capabilities so they can weigh the benefits relative to “simpler, lower cost” large language models (LLMs).
At the same time, MassMutual is collecting increasingly detailed analytics around usage patterns, developer workflows, model performance, and costs. The goal is to reduce spending while also building operational intelligence to eventually route workloads to the right model based on cost, response quality, and user experience.
Those insights will eventually drive optimization decisions around model routing, prompt selection, response times, and infrastructure design.
“We’re gaining access to analytics that let us, in a very granular way, look at usage patterns, developer workflows, and begin to make sense of who’s using what, when, and for what types of tasks,” Merritt said.
Another interesting aspect of MassMutual’s approach is how it evaluates AI quality. Rather than focusing exclusively on benchmarks or token costs, the company uses what Merritt calls a “trust score” framework.
The process combines user feedback with operational metrics to understand how employees perceive AI-generated responses and whether those responses actually improve outcomes.
The contact center rebuild put that framework to the test. During development, employees were given access to two different LLMs. One generated responses in near-real-time but the quality was noisier. The other more expensive option took several additional seconds to respond but consistently delivered higher-quality answers.
Conventional wisdom and the speed of business might suggest users would prefer the former; but they overwhelmingly chose quality. Merritt’s team asked users about the quality of response, their preferred model, and their overall thoughts on the experience.
Most of the time, users said: “We want the more expensive one. We’re willing to wait, but the quality difference is so high that the two extra seconds actually is worth it to us.”
That feedback ultimately determined which model MassMutual deployed.
“We factored that experience piece into the decision-making, and that led us to say, on a relative basis, the costs were immaterial, so we’re going to use the more complex model,” Merritt said.
Listen to the full podcast to hear more about:
Why Mythos “completely changed” the cybersecurity landscape — not the type of threats, but the rate at which those threats appear;
How a team of AI engineers modernized MassMutual’s mainframe in 7 days (a process that previously would have taken 3 months);
Why MassMutual specifically avoided tokenmaxxing to rein in AI use and spending and has been going “unlimited,” to shield from cost blowups.
How a “multi-harness type of environment” will support agentic AI.
You can also listen and subscribe to Beyond the Pilot on Spotify, Apple or wherever you get your podcasts.
Presented by Snowflake
As AI agents become capable of reasoning across systems and taking action, software is evolving from something employees operate into something that understands intent. Instead of navigating disparate applications and dashboards, a single system will increasingly ask: What are you trying to accomplish?
That sounds like a user experience breakthrough. It is. But the more important implication is organizational. When software no longer relies on humans to provide context, companies can no longer assume that knowledge lives in employees’ heads or is buried inside disconnected applications. The company itself has to become machine-readable.
The winners in the AI era won’t simply deploy more intelligent models. They’ll build the data foundations, semantic context, and governance frameworks that allow machines to understand how the business works and act on that understanding with confidence.
For years, companies treated context as a human layer on top of data. The data platform held the records, then the BI tool visualized them, and the analyst interpreted them. And finally, the business leader made the judgment call. Agents collapse those layers.
When an executive asks, “Why is customer churn rising in our enterprise segment?” an effective agent needs to know far more than where the customer data lives. It needs to understand how the company defines churn, which accounts count as enterprise, whether product usage data is more reliable than survey data, which renewal events matter, what the sales team has logged, what support tickets suggest, and whether the answer differs by geography or product line.
This is why semantics — the definitions, relationships, rules, and assumptions that give data meaning — are moving from a technical concern to a boardroom issue. A semantic layer used to sound like plumbing for data teams. In an agentic enterprise, it becomes the shared language between humans and machines.
If every department teaches its own agent a different version of the business, companies will get inaccuracy at scale. The organizations that pull ahead will be the ones that create a common business knowledge base: consistent definitions, governed access, documented workflows, clear lineage, and enough flexibility to evolve as the business changes. In that world, context is treated as infrastructure, rather than just a nice-to-have.
The first wave of enterprise AI largely gave us assistants and copilots that answer questions. Useful, but still limited. You ask a question, get a response, and then return to the work of stitching systems together yourself.
The next era of AI will be different. Agents will move beyond coordinating answers, and start getting actual work done. A sales leader starting the day will not need to open a CRM, a forecasting tool, a support dashboard, and a Slack thread to understand what changed overnight. They will simply ask an agent what needs attention. The agent will identify which accounts are at risk, explain why, summarize recent customer interactions, draft follow-up actions, and perhaps initiate the next workflow.
The dashboard does not disappear because charts become useless. It disappears because static reporting becomes too slow for how businesses need to operate. The center of gravity shifts from “show me what happened” to “help me decide what to do next.”
As long as AI is mostly answering questions, governance is about controlling what it can access. That is already difficult. Employees have different permissions, sensitive data needs protection, and answers must be traceable to trusted sources. As agents begin taking action, governance becomes even more consequential.
It’s one thing for an agent to summarize a customer complaint. It’s another for it to issue a refund, reorder inventory, or send an email to a customer. This is where many companies will be tempted to choose between two imperfect paths.
One path is to tightly constrain agents from the start: define the data sources, tools, workflows, and actions they can access. This is easier to manage and measure. It also risks limiting the creativity of employees who understand their workflows best.
The other path is to let teams experiment freely: connect agents to the tools and data they use every day, and allow new use cases to emerge organically. This can produce faster adoption and unexpected innovation. It can also create real risk: stale data, inappropriate access, duplicated workflows, runaway costs, or automated actions no one fully understands.
The right answer is not maximum control or maximum freedom. It’s to prioritize governed flexibility. Companies need architectures where governance is embedded from the beginning. An agent should know not only what it can read, but what it can do, when it needs approval, how its reasoning is inspected, and how its performance is evaluated over time. In other words, governance cannot be a review meeting after the pilot. It has to be part of the system design.
One of the least appreciated consequences of agentic AI is that it will blur the line between people who use software and people who create it. When employees can describe a workflow in natural language and have an agent help build it, software development becomes less confined to engineering teams. A marketer can create a campaign analysis workflow. A finance manager can automate variance explanations. An HR leader can build a policy assistant. A support manager can design a triage process.
These employees are not becoming software engineers in the traditional sense, but they are becoming builders. That changes the talent model. Technical fluency will matter more because employees need to understand what’s possible, what’s risky, and how to evaluate an AI-generated result. Judgment becomes the most important skill.
The winners will be the people who know how to ask better questions, inspect evidence, refine workflows, and combine domain expertise with enough technical understanding to move from idea to execution.
For business leaders, this means AI adoption extends beyond an IT rollout, and is actually an organizational redesign. The distance between insight and action will shrink, and companies will need to rethink who is empowered to build, approve, and operate the workflows that run the business.
The shift from interfaces to agents will also challenge how companies buy and measure software, and change how software is priced. Per-seat licensing is giving way to consumption models, where costs reflect actual usage. For most organizations this is a better deal. You pay for value delivered, not licenses that may sit idle.
But it also changes the accountability calculus. When costs are fixed per seat, budget conversations happen once a year. When costs scale with usage, they require continuous oversight. Without visibility into how agents are used and what they produce, costs can rise quickly.
The answer is to build measurement in from the start, connecting AI usage to business outcomes, whether that is deals closed, tickets resolved, or cycle times reduced.The companies that succeed will treat AI cost management as part of operational excellence, not procurement cleanup. The question should not be, “How many tokens did we use?” It should be, “What business outcome did that intelligence produce?”
While the internal implications of agents are significant, the external ones may be even larger. Today, companies obsess over the customer experience inside their applications: the homepage, the navigation, the checkout flow, the dashboard, the mobile screen. Those things will still matter. But increasingly, customers may interact with businesses through their own agents rather than directly through a company’s app or website.
If a procurement agent compares suppliers, a travel agent books a trip, or a financial agent evaluates products, the customer may never see the interface a company spent years perfecting. The agent will care less about visual design and more about whether the company’s data, policies, pricing, inventory, documentation, and transaction systems are accessible, structured, trustworthy, and machine-readable.
That means the competitive surface area changes. A company’s brand may still be emotional, but its operational interface will increasingly be data. Businesses that expose confusing, inconsistent, or poorly governed information will be harder for agents to work with. Businesses with clean semantics, reliable APIs, governed data, and clear policies will become easier to choose, easier to transact with, and easier to trust.
The interface does not vanish only inside the enterprise. It may vanish between enterprises, too.
Most executives know they need an AI strategy, but fewer have internalized what that really requires. AI readiness is not the number of pilots launched, the number of models tested, or the number of employees with access to a chatbot. It is whether the organization’s knowledge, data, permissions, workflows, and decision logic are ready for machines to reason over them safely.
For decades, enterprise software forced humans to become translators between business intent and machine logic. AI is reversing that relationship. Machines are beginning to adapt to human intent. But they can only do that if the enterprise has done the work to make its own context legible.
The future of software is not another screen. It is a system that understands the business well enough to help run it. And that means the next great interface will not look like an interface at all.
Baris Gultekin is VP of AI at Snowflake.
Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.
A joint research collaboration between researchers at the University of Illinois at Urbana-Champaign (UIUC), UC Berkeley, and the open source AI-native vector database platform Chroma unveiled Harness-1, a 20-billion parameter open-source search agent built atop OpenAI’s gpt-oss-20B open source model that fundamentally redesigns how AI executes complex retrieval tasks.
Harness-1 achieves a massive leap in performance, scoring 73% average on its ability to recall relevant information correctly from a curated dataset, outperforming even GPT-5.4 (70.9%) and the next, most accurate open source search agent, Tongyi DeepResearch 30B, by 11.4 percentage points. (While GPT-5.5 has also been out for more than a month, the researchers didn’t test against this model as it wasn’t available when they were building theirs.)
Crucially for developers, the model and its environment are available immediately under the highly permissive Apache 2.0 license and model code/weights on Hugging Face.
Harness-1 also serves as proof-of-efficacy of another effort, Tinker, the distributed, web-based AI model training and fine-tuning API developed by Thinking Machines. Tinker was used specifically to train and run inference for Harness-1, highlighting how interactive infrastructure is actively enabling the next generation of autonomous models.
So how did the researchers do it?
To actually put these models to the test, the researchers evaluated Harness-1 and its competitors across eight highly complex search benchmarks. Rather than asking simple trivia questions, these tests required the AI to act like a real researcher sifting through diverse, dense data sources.
The benchmarks spanned several different domains, including open web searches, complex financial filings from the SEC, technical patent databases from the USPTO, and “multi-hop” question-answering tasks where the AI had to logically piece together scattered clues from multiple different documents to arrive at the correct answer.
When the results came in, Harness-1 dominated the open-source competition in its ability to successfully find and curate the right facts. Even more impressively, this relatively small 20-billion parameter model went toe-to-toe with massive, expensive proprietary AI systems. It actually outperformed heavyweights like GPT-5.4, Sonnet-4.6, and Kimi-K2.5 — thought to be the hundreds of billions or trillions of parameters. Only one giant frontier model—Opus-4.6 — managed to narrowly edge it out in overall average performance.
Harness-1 achieves its performance gains by offloading the exhaustive “bookkeeping” of a search session out of the model’s working memory and into a structured software environment.
As enterprise use cases grow more sophisticated, demanding that models autonomously sift through thousands of corporate documents or financial filings, these systems frequently succumb to “search amnesia”—forgetting their original queries, looping over rejected documents, or losing track of the specific claims they are trying to verify.
Until now, the prevailing solution to this amnesia has been brute force. Engineers typically force models to constantly reread an ever-expanding, append-only transcript of their own actions, piling every search, read, and thought back into a massive context window.
Harness-1 introduces a paradigm shift away from this method, proving that the bottleneck for true artificial autonomy isn’t necessarily the size of the model, but how efficiently its working environment manages state. It highlights once more, as Anthropic’s Claude Code has also done, that the raw model is arguably less important than the harness — or set of conditions — through which it runs.
To understand the technical leap of Harness-1, consider a real-world analogy.
Imagine hiring a brilliant research assistant and placing them in an empty room without a desk, notepads, or filing cabinets. You ask them to write a comprehensive report on a highly complex topic, which requires them to read dozens of books while keeping every single quote, citation, and dead-end search perfectly memorized in their own head. Eventually, no matter how intelligent the assistant is, their cognitive load will max out, and they will start dropping facts or losing the thread of the assignment.
This is exactly how traditional search agents operate today. They are trained as policies over growing transcripts, meaning the model searches, reads, searches again, and appends everything into its own context window.
As lead researcher Patrick (Pengcheng) Jiang of the University of Illinois noted on X: “At some point the model is not just ‘searching’ anymore. It is also being asked to be a memory system, a note taker, a verifier, and a librarian.”
Harness-1 solves this by giving the AI a desk and a filing cabinet—what the research team calls a “state-externalizing harness.”
This harness is an active, surrounding environment that takes over the routine bookkeeping, maintaining a recoverable working memory that includes a candidate pool of documents, an importance-tagged curated evidence set, compact evidence links, and verification records.
By separating semantic choices from structural state management, the AI is freed up to do what it does best.
The policy still decides what to search, determines which documents to keep, and knows when to stop, while the environment simply holds the state.
Here is a subsection breaking down the training methodology and how it differs from prior agentic search models:
The training pipeline for Harness-1 represents a fundamental shift in how the AI industry approaches agentic learning.
Historically, developers have treated search agents as policies operating over massive, ever-growing transcripts, forcing reinforcement learning (RL) algorithms to simultaneously optimize both semantic reasoning and the raw memorization of a search state.
Harness-1’s creators took a radically different approach: because their custom “harness” handles all the routine bookkeeping—like maintaining evidence links, candidate pools, and verification records—the training process only needed to teach the model how to operate this structured interface.
This division of labor drastically simplified what the underlying 20-billion parameter model actually needed to learn.
The process began with a remarkably narrow Supervised Fine-Tuning (SFT) stage. Rather than scraping petabytes of new behavioral data, the team generated just 899 filtered trajectories using a GPT-5.4 teacher agent that was plugged into the exact same harness environment the student model would eventually use.
The goal of this SFT phase was not to inject vast amounts of domain knowledge into the model, but simply to teach it the mechanical rhythms of a good researcher: how to format tool calls, how to tag documents by importance, and the discipline of verifying a claim before promoting it to the final curated set.
Following SFT, the model underwent Reinforcement Learning (RL) using an algorithm called CISPO, applied over full search episodes capping at 40 turns.
The team designed a highly specific terminal reward function that explicitly separated discovery from selection. The model was rewarded not just for finding a relevant document, but for successfully promoting it into the final answer set, while being penalized if it found the answer but failed to curate it.
The researchers also instituted a “tool diversity” bonus; without this specific incentive, they found the policy would quickly collapse into a lazy, search-heavy strategy where it spammed queries but bypassed the harder work of reading and verifying the text.
What makes Harness-1 truly innovative compared to prior work is its unprecedented data efficiency. The entire model was trained on roughly 4,400 unique items—899 SFT trajectories and 3,453 RL queries.
In stark contrast, competing open-source models required vastly larger datasets to achieve worse results: Context-1 utilized over 17,200 training items, while Search-R1 relied on a staggering 221,300 items to learn search behaviors.
By proving that a smarter external cognitive architecture can replace brute-force data scaling, Harness-1 suggests that the future of agentic AI lies in building better environments for models to work within, rather than just training larger models on more data.
From a product perspective, Harness-1 is delivered as a highly capable 20B agent merged into the openai/gpt-oss-20b base architecture.
For enterprise tech stacks, the applicability is massive because businesses need AI to execute multi-step research across proprietary databases without hallucinating or running up exorbitant compute bills.
Harness-1 manages its frontier-level performance at what the creators describe as “Context-1-level cost and latency.” Because the context window is strictly managed by the budget-aware harness rather than continuously expanding, enterprises can deploy this agent autonomously without incurring the exponential token costs typically associated with long-horizon AI tasks.
Even more impressively, Harness-1 proves it can generalize well beyond its training data. According to the research team, it was incredibly cheap to train, utilizing just 899 filtered supervised fine-tuning (SFT) trajectories and a mere 3,453 reinforcement learning (RL) queries.
“Instead of training the model to survive a giant append-only transcript, we train it to use a structured search interface: search, curate, revisit, verify, and submit,” Jiang explained.
This leanness proves a critical point for the AI industry: developers do not necessarily need petabytes of new behavioral data if they build a better cognitive framework for the model to operate within.
One of the most significant aspects of the Harness-1 release is its licensing. In plain language, Apache 2.0 is a highly permissive, enterprise-friendly software license that fundamentally enables commercialization.
Unlike “copyleft” licenses (such as the GPL) that can force companies to open-source their own proprietary software if they integrate the code, or “research-only” licenses that ban commercial use entirely, Apache 2.0 gives businesses the green light to freely build, modify, and monetize the technology.
For developers and startups, this means Harness-1 can be seamlessly integrated into commercial enterprise search products, internal data retrieval tools, or customer-facing AI applications without fear of legal reprisal.
The only major requirement is that users must include the original copyright notice and explicitly state any significant modifications they make to the source code, positioning Harness-1 as a highly viable foundational building block for the enterprise.
The announcement has clearly struck a nerve within the developer community, validating the very real pain points engineers face when building agentic systems. Jiang’s multi-part announcement thread on X quickly garnered massive traction, pulling in over 256.1K views, 3.7K likes, 2.9K bookmarks, and nearly 300 reposts within a matter of days.
This high engagement underscores a growing consensus in the AI space that brute-forcing context windows is a losing battle.
When Jiang posted on X, “I’ve been wondering: maybe search agents are bad at search partly because we make them do all the paperwork in their head,” the resonance was immediate.
For developers who have spent the last year wrestling with AI agents that confidently forget their primary instructions halfway through a database search, the Harness-1 approach feels like a desperately needed course correction.
Ultimately, the community sentiment highlights a shift in industry priorities. Developers are moving away from asking how large an AI model’s context window can get, and instead asking how efficiently an AI model’s environment can manage that context for it. By offloading the paperwork, Harness-1 is proving that smaller, smarter systems can outmaneuver the giants—provided they have the right desk to work at.
Our system did one thing, and it did it well: It turned natural-language questions into API calls.
The users were analysts, account managers, and operations leads. They knew what data they needed, but assembling it manually meant pulling from four dashboards, two BI tools, and a Salesforce report builder. With our system, they typed the request in plain English. A request like “Compile a report on sales volume for January through March 2026 for the Northeast region, broken down by city” was translated into an API call that the system could act on:
json
{
“description”: “User requested sales volume for the given date range, here is the API call to get the response”,
“api_call”: “/api/sales_volume”,
“post_body”: {
“start_date”: “2026-01-01”,
“end_date”: “2026-03-31”,
“region”: “northeast”
}
}
The rest of the pipeline was conventional engineering. The system dispatched the call to the right backend — we had integrations with internal reporting portals, Salesforce, and several homegrown services — applied a large language model (LLM)(-generated JSON query to filter and shape the response, and delivered it via email, as a Drive document, or rendered as a chart in the browser.
By mid-2025, the system was generating several hundred reports a month. These reports were consumed by leadership and analysts and circulated to external stakeholders. It had become the default way most teams pulled ad-hoc data.
The contract between the LLM and the rest of the system was a structured JSON object as described in the above example.
json
{
“description”: “User requested sales volume for the given date range, here is the API call to get the response”,
“api_call”: “/api/sales_volume”,
“post_body”: {
“start_date”: “2026-01-01”,
“end_date”: “2026-03-31”,
“region”: “northeast”
}
}
We built it on Claude Sonnet 3.5 in early 2025. We upgraded to 3.7 without incident, and to 4.0 without incident. By the time Sonnet 4.5 shipped, we had grown complacent about the stability and predictability of LLMs in solving what we believed was a simple problem. Model upgrades had become routine, like bumping a minor version of a well-behaved library.
Then we rolled out 4.5. For a meaningful percentage of requests, the model began folding the contents of post_body into the description field. Two failure modes followed.
First, the filter parameters never reached the API. Our system read post_body as the source of truth for the request payload, and that field came back empty. The API call was made without the date range or region filter. Depending on the specific API being called, the backend either returned sales volume for all time or all regions or returned a 500 error.
Second, the model started asking clarifying questions in its response. This was new. Earlier versions always took a best-effort approach to an ambiguous request and returned a structured object. Sonnet 4.5, being more cautious, would sometimes respond with a question instead. Our system had no path for this. It had been built on the assumption that every model invocation would result in an API call. There was no human-in-the-loop component and no state to hold a partially completed request. This caused downstream systems to break in multiple ways.
We rolled back to 4.0. That was harder than it should have been: Between the 4.0 and 4.5 deployments, our team had added new API integrations, all of which were qualified against 4.5. Reverting the model meant requalifying every one of them against 4.0 under time pressure.
Software engineering rests on the ability to bound the effect of a change. When you upgrade a driver or library, you read the release notes to see whether to expect breaking changes. Unit tests circumscribe what could possibly have moved. You can leverage the following property: The system being changed is deterministic enough that its behavior can be predicted, or at least sampled densely enough to give you confidence. The blast radius is bounded by construction.
LLM-backed systems break this assumption. The component that produces your output is not under your control. You cannot diff a model version bump from 4.0 to 4.5. It is a wholesale replacement of the functionality on which your system depends.
This is what we mean by an infinite blast radius: a change whose downstream effects cannot be enumerated in advance because the input space (natural language) and the failure modes (anything the model might do differently) are both unbounded.
The post-mortem revealed that our prompt had always been under-specified. We had told the model to return a JSON object with three fields. We had described what each field was for. We did not explicitly state that the description must be a natural-language string and must not contain serialized representations of other fields.
Earlier versions of the model inferred this constraint from context. Sonnet 4.5, evidently better at being “helpful” in its formatting choices, decided that inquiring for clarification or providing the request body in the description made the response more useful. From the model’s perspective, this was a reasonable interpretation of an ambiguous instruction. However, this violated the assumptions under which our system was built.
The bug was not in the model. The bug was in our assumption that the model would continue to fill in our specification gaps as it always had. Three successful upgrades had trained us to believe those gaps were safe.
Structured output modes and tool-use APIs would have caught this specific failure at the schema level. We weren’t using them for engineering reasons outside the scope of this article. But schemas only constrain syntax, not semantics. A schema cannot specify that a clarifying question shouldn’t appear in a system with no path for clarification, or that a date range should never silently default to all-time. Schemas solve the easier half of the problem.
The discipline that closes this gap is to treat the evaluation suite — not the prompt — as the formal specification of the system. The prompt is an implementation of the spec. The model is an interpreter. The evals are the spec itself, and any model or prompt change is valid if and only if it passes them.
In practice, an eval is a triple: An input, a property the output must satisfy, and a scoring function. For our system, the eval that would have caught the 4.5 regression looks roughly like this:
python
def test_description_contains_no_serialized_payload(response):
desc = response[“description”].lower()
forbidden = [“curl”, “post_body”, “{“, “http://”, “https://”]
assert not any(token in desc for token in forbidden), \
f”description leaked structured content: {response[‘description’]}”
A few hundred such properties, some written by hand for known-important invariants, some generated as regression tests from real production traffic, some scored by an LLM-as-judge for fuzzier qualities like tone, become a gate. Model upgrades and prompt changes should be treated as pull requests that must turn the suite green before they merge.
Evals are expensive to build and maintain. They drift as your product changes. LLM-as-judge scoring introduces its own variance in outcomes. And the suite can only catch failure modes you have thought to specify — you cannot eval your way to safety against a category of failure you have never imagined. We learned this lesson the hard way: Nobody on our team had ever written an assertion that said “the description field should not contain a curl command,” because nobody had thought the model would put one there.
Evals are not a silver bullet. They give you the ability to bound the blast radius of a change in the only way available when the underlying function is a black box: By densely sampling the input-output response you actually care about, and refusing to deploy when that behavior moves.
The engineering community has yet to develop a body of knowledge for writing effective evals. There are no widely accepted standards for what ‘coverage’ means in natural language input spaces. CI/CD systems were not built to gate probabilistic test outcomes. As agents take on more autonomous work — writing code, moving money, scheduling infrastructure changes — the gap between “the model passed our smoke tests” and “we know what this system will do in production” becomes the central engineering problem of the next several years.
The teams that close that gap will be the ones who stop treating evals as a quality-assurance afterthought and start treating them as the actual specification of what their system is.
Vijay Sagar Gullapalli is Founding AI Engineer at Adopt AI and a USPTO-patented inventor.
Sarat Mahavratayajula is a Senior Software Engineer at Sherwin-Williams.
Microsoft used its Build 2026 conference this week to push a clear message: agents are rapidly moving into production throughout enterprise systems, and the winning platform will be the one that gives them reliable context, governance, identity, memory…
When someone on a team corrects an AI agent — better prompts, better feedback, better context — that improvement disappears the moment a colleague opens the same tool. The correction doesn’t transfer, and the next person starts from zero.
The problem compounds in multi-agent workflows, where teams expect agents to share context across users and tasks. Without a shared memory layer, every team member effectively trains a different version of the same agent — and those versions never sync.
That gap shows up in the numbers. According to Asana’s own research, 75% of knowledge workers use AI on the job, but only 5% of companies have reported productivity gains.
“Model providers are getting really, really good at improving reasoning and retry loops, but what they’re not good at is bringing the enterprise work context in a way that human beings can reason about for shared memory,” Asana Chief Product Officer Arnab Bose told VentureBeat.
Asana had been building toward an agentic platform that centers context and shared memory. Its Agentic Work Management platform ensures that if any team member corrects an agent, that correction applies to everyone else on the team.
“That context graph is automatically provided to agents operating inside Asana’s system so you don’t have to have every human member of the team become an expert at prompt engineering or context engineering,” Bose said.
Bose said the shared memory architecture matters beyond Asana’s own product; it’s the design decision enterprises need to make for any multi-agent system.
Shared memory also becomes important when enterprises begin moving from simple single agents to multi-agent workflows that need to share context and behaviors.
The models powering agents are stateless by design, so memory becomes a dedicated layer outside of a context window. While this area of AI innovation is marching towards maturity, the question of what gets stored, who controls it, and how it stays consistent when different agents and users write to the same instance remains largely unsolved.
This is manageable for use cases with only one user. However, in enterprise agentic workflows, the idea is for agents to work with the entire team. Most platforms have agents that still act for individuals, which leads to task repeating and inconsistent versions of reality and spreading mistakes. Agents could then also contradict each other.
Sriharsha Chintalapani, co-founder and CTO of Collate, said in an email to VentureBeat that the lack of shared memory is a major obstacle for multi-agent workflows particularly around consistency.
“Agents are sensitive to the quality of their prompts,” Chintalapani said. “Someone with a strong understanding of the task will generally get more accurate results than someone less experienced. Partly that’s because they’re able to construct more detailed prompts, but also because they’re able to give the agent better feedback. The agent remembers the corrections it’s received and applies that knowledge to successive prompts. The more accurate the feedback, the better the agent will perform for that user. “
He added that organizations should stop treating shared memory solely as a prompt engineering problem and think of building systems that repeat context across every conversation.
Neej Gore, chief data officer at Zeta Global, said in a separate email that shared context becomes a living memory that “compounds intelligence across the enterprise.”
The opportunity may lie in building AI agents that retrieve memory relationally, pulling in relevant context based on what’s being asked — an approach Chintalapani says few organizations outside the largest model providers are equipped to build.
AI agents already proliferate enterprises; it’s just that many of these operate as personal agents doing work specific to individual users. Most prompts start from one person, any files are uploaded by one account, and even for agents living in a company-wide system mostly learn individual user preferences.
Most enterprise AI workflow platforms recognize that memory is important but approach it through different lenses. For example, Microsoft’s Copilot takes an individual-first approach by learning a user’s role within the organization, tone preferences and working patterns, which are then stored as personal memories for the agent to apply across the different Microsoft 365 surfaces.
For engineering and orchestration teams evaluating agentic platforms, the shared memory question is now a procurement criterion — not just a technical nicety. An agent that learns only for the person using it will require ongoing individual upkeep. One connected to a team-wide memory layer builds institutional knowledge automatically.
Agentic AI is moving rapidly from the developer terminal to the corporate world.
On Tuesday, OpenAI announced a major update of its agentic AI platform Codex, introducing domain-specific workflows, a rapid, semi-private web hosting feature within it for enterprises called “Sites,” and an in-place editing tool named “Annotations”.
The release marks a deliberate strategy to transform Codex from a specialized programming assistant into an everyday operating environment for business professionals.
Non-developers—including financial analysts, marketers, operators, and researchers—now constitute approximately 20% of the platform’s 5 million weekly users and are adopting the technology three times faster than traditional engineers, according to research shared by OpenAI with VentureBeat and other outlets.
OpenAI is capitalizing on this shift to position Codex as the premier application for white-collar task automation. The timing of the announcement is highly strategic, arriving precisely as its own primary investor turned business rival Microsoft this week kicks off its annual BUILD developer conference in San Francisco—where a slate of competing enterprise productivity tools is expected—and hot on the heels of Anthropic’s rapid adoption among knowledge-workers via its Claude Cowork and Claude Code platorms.
For business users, the most critical technical upgrade is the elimination of full-document regeneration. Previously, instructing an AI to update a specific chart or spreadsheet calculation often meant the model had to rewrite the entire file, which frequently broke custom formatting or introduced hallucinations.
OpenAI addresses this through Annotations, a localized context-scoping mechanism. As demonstrated in the company’s release materials, the platform maps a document’s underlying data schema.
When a user highlights a specific segment—such as a block of cells in a financial model—Codex isolates those exact data arrays.
If an analyst prompts the system to “Add a chart of revenue, EBITDA, and net income over the selected years,” the model executes the code strictly within that boundary, generating the visualization while leaving the surrounding cell dependencies, styles, and unselected formulas completely untouched.
To further anchor Codex in daily enterprise operations, OpenAI has introduced modular software bundles and a rapid-prototyping hosting environment.
The company is rolling out six role-specific plugins that aggregate 62 popular business applications (including Snowflake, Figma, and Salesforce) and 110 automated skills straight out of the box.
Data Analytics: Unifies cloud environments like Snowflake, Databricks Genie, Hex, and Tableau to translate natural language inquiries into data reports and change-analysis dashboards.
Creative Production: Connects Figma, Canva, Shutterstock, Picsart, and Fal to generate and iterate on ad variations, campaign boards, and e-commerce assets directly from text briefs.
Sales: Integrates pipeline infrastructure across Salesforce, HubSpot, Slack, Outreach, Clay, Rox, and Actively to automate follow-up communications, close plans, and account risk reviews.
Product Design: Bridges Figma and Canva environments to audit live user journeys and transform static wireframes into clickable prototypes.
Public Equity & Investment Banking: Syncs institutional market feeds—including Moody’s, Daloopa, Datasite, FactSet, LSEG, S&P, PitchBook, and Hebbia—to streamline financial modeling, competitive landscaping, and pitch book preparation.
These integrations allow distinct departments—from data analytics and creative production to sales and investment banking—to automate complex, multi-step workflows without requiring IT to build custom API connections.
Concurrently, the new Sites feature introduces an interactive canvas that converts static data inputs or text documents into functional, web-hosted internal applications.
Rolling out in preview for Business and Enterprise tiers, Sites allow cross-functional teams to bypass front-end development.
Financial leaders, for example, can transform a static spreadsheet into an interactive scenario planner shared via a secure workspace URL, allowing executives to tweak assumptions in a live web app rather than clicking through document tabs.
Instead of static decks, Sites promise to keep enterprises updated on their latest metrics and important information in an easily digestible way.
A critical operational distinction in this rollout centers on exactly where these new features can be executed. Codex’s existing infrastructure runs natively across multiple surfaces, including IDE extensions and the terminal command line.
However, the release documentation notes that Sites are rolling out “through the Codex app” and that plugins are managed via a “Codex plugin directory”.
An OpenAI spokesperson confirmed that Plugins and Sites are available int he CLI and desktop app, while Sites are hosted by OpenAI.
These updates operate entirely within OpenAI’s closed, proprietary enterprise licensing model. Unlike open-source frameworks, enterprise clients do not maintain code-level ownership over Codex’s integration nodes.
Instead, system administrators manage deployment through centralized workspace settings, giving them explicit authority to enable or disable hosted “Sites” and restrict underlying application permissions.
These new capabilities deploy seamlessly on top of Codex’s existing commercial framework. Users will continue to access the agent via established baseline subscription tiers—such as the individual “Plus” plan ($20/month) or the high-volume “Pro” plan ($100/month)—or through a separate, seat-free pay-as-you-go model that draws down pre-purchased utility credits.
Enterprise AI agents are stalling — not because of model performance, but because of permissioning. Every agentic workflow eventually hits the same wall: what is this agent allowed to touch, on whose behalf, and how does the system know?
Workday’s answer is to make its existing system of record the governance layer for agents. Gerrit Kazmaier, the company’s president for product and technology, told VentureBeat in an interview that customers often struggle when they cobble together solutions for their agents.
“Sana makes sure the integrity of the approvals and security model is always adhered to,” Kazmaier said. “Frankly, that’s where we see customers struggling when they try to build do-it–yourself AI by just accessing raw data, so the richness of the security model gets lost, and the results become overly broad.”
Workday, which launched Sana in March, expanded its partnership with Google to bring its Sana agent system of record to the Gemini Enterprise — so agents built on Sana are also discoverable there.
Kazmaier said the biggest hurdle they faced was ensuring agent accuracy, especially for HR and finance users.
“Almost right is not acceptable,” Kazmaier said. “Think about paying people correctly, closing the books or managing work schedules reliably.”
Accuracy is harder to evaluate here than in most AI contexts. Policy configurations, role-based security, and organizational hierarchies are deeply interrelated — a small error compounds. And unlike most generative AI outputs, HR and finance queries often lack a correction loop. By the time a paycheck processes incorrectly or an interview is scheduled wrong, the damage is done.
Workday addressed this by building Gemini in as its base reasoning layer, then adding its context engine and business process logic on top. Workday also added verification and classification models that “interrogate” outputs before execution.
Accuracy and identity, it turns out, are the same question: does the system know enough about the agent, the authorizing human, and the current state of the record to act correctly?
Workday’s advantage is that it can infer its customers’ organizational structures from the data they provide. Already, third-party identity providers like Okta verify their information by checking Workday, so its context is the system of record for many enterprises. Kazmaier said the Sana Self-Service Agent uses Gemini as the conversational surface to trigger the workflow. The user is then authenticated and authorized through Workday’s identity and security model. Sana agents will only act on behalf of that user and work within their current permissions.
Audit trails follow the same logic: Gemini retains only interaction logs, while the main audit remains within Workday and its customer.
For many practitioners in the HR and finance space, the permission and governance layer in the agent system of record is key in regulated spaces.
“It has to live in the system of record, that’s not a preference, that’s the only way it works,” said Dan Obendorfer, director of product at Würk, in an email to VentureBeat. “If your permissions are defined somewhere outside of where the data actually lives, you’ve already lost.”
Kadan Stadelmann, chief technology officer and co-founder of Compance.AI, made the same point separately. “Without agent ownership, performance, costs or actions, chaos ensues.”