Apple will launch the iPhone 18 Pro in September, but wait six months to launch the iPhone 18. How will this strategy benefit Apple’s bottom line?
What’s next for the Gemini Agent? Hidden Android 17 code reveals new autonomous skills and task scheduling. But does your phone meet the strict requirements?
For AI systems to keep improving in knowledge work, they need either a reliable mechanism for autonomous self-improvement or human evaluators capable of catching errors and generating high-quality feedback. The industry has invested enormously in the first. It’s giving almost no thought to what’s happening to the second.
I’d argue that we need to treat the human evaluation problem with just as much rigor and investment as we put into building the model capabilities themselves. New grad hiring at major tech companies has dropped by half since 2019. Document review, first-pass research, data cleaning, code review: Models handle these now. The economists tracking this call it displacement. The companies doing it call it efficiency. Neither are focusing on the future problem.
The obvious pushback is reinforcement learning (RL). AlphaZero learned Go, chess, and Shogi at superhuman levels without human data and generated novel strategies in the process. Move 37 in the 2016 match against Lee Sedol, a move professionals said they would never have played, didn’t come from human annotation. It emerged from AI self-play.
What enables this is the stability of the environment. Move 37 is a novel move within the fixed state space of Go. The rules are complete, unambiguous, and permanent. More importantly, the reward signal is perfect: Win or lose, and immediate, with no room for interpretation. The system always knows whether a move was good because the game eventually ends with a clear result.
Knowledge work doesn’t have either of those properties. The rules in any professional domain are dynamic and continuously rewritten by the humans operating in them. New laws get passed. New financial instruments are invented. A legal strategy that worked in 2022 may fail in a jurisdiction that has since changed its interpretation. Whether a medical diagnosis was right may not be known for years. Without a stable environment and an unambiguous reward signal, you cannot close the loop. You need humans in the evaluation chain to continue teaching the model.
The AI systems being built today were trained on the expertise of people who went through exactly that formation. The difference now is that entry-level jobs that develop such expertise were automated first. Which means the next generation of potential experts is not accumulating the kind of judgment that makes a human evaluator worth having in the loop.
History has examples of knowledge dying. Roman concrete. Gothic construction techniques. Mathematical traditions that took centuries to recover. But in every historical case, the cause was external: Plague, conquest, the collapse of the institutions that hosted the knowledge. What’s different here is that no external force is required. Fields could atrophy not from catastrophe but from a thousand individually rational economic decisions, each one sensible in isolation. That’s a new mechanism, and we don’t have much practice recognizing it while it’s happening.
At its logical limit, this isn’t just a pipeline problem. It’s a demand collapse for the expertise itself.
Consider advanced mathematics. It doesn’t atrophy because we stop training mathematicians. It atrophies because organizations stop needing mathematicians for their day-to-day work, the economic incentive to become one disappears, the population of people who can do frontier mathematical reasoning shrinks, and the field’s capacity to generate novel insight quietly collapses. The same logic applies to coding. Our question is not “will AI write code” but “if AI writes all production code, who develops the deep architectural intuition that produces genuinely novel systems design?”
There is a critical difference between a field being automated and a field being understood. We can automate a huge amount of structural engineering today, but the abstract knowledge of why certain approaches work lives in the heads of people who spent years doing it wrong first. If you eliminate the practice, you don’t just lose the practitioners. You lose the capacity to know what you’ve lost.
Advanced mathematics, theoretical computer science, deep legal reasoning, complex systems architecture: When the last person who deeply understands a subfield of algebra retires and no one replaces them because the funding dried up and the career path disappeared, that knowledge isn’t likely to be rediscovered any time soon.
It’s gone. And nobody notices because the models trained on their work still perform well on benchmarks for another decade. I think of this as a hollowing out: The surface capability remains (models can still produce outputs that look expert) while the underlying human capacity to validate, extend, or correct that expertise quietly disappears.
The current approach is rubric-based evaluation. Constitutional AI, reinforcement learning from AI feedback (RLAIF), and structured criteria that let models score models are serious techniques that meaningfully reduce dependence on human evaluators. I’m not dismissing them.
Their limitation is this: A rubric can only capture what the person who wrote it knew to measure. Optimize hard against it and you get a model that’s very good at satisfying the rubric. That’s not the same thing as a model that’s actually right.
Rubrics scale the explicit, articulable part of judgment. The deeper part, the instinct, the felt sense that something is off, doesn’t fit in a rubric. You can’t write it down because you need to experience it first before you know what to write.
This isn’t an argument for slowing development. The capability gains are real. And it’s possible that researchers will find ways to close the evaluation loop without human judgment. Maybe synthetic data pipelines get good enough. Maybe models develop reliable self-correction mechanisms we can’t yet imagine.
But we don’t have those today. And in the meantime, we’re dismantling the human infrastructure that currently fills the gap, not as a deliberate decision but as a byproduct of a thousand rational ones. The responsible version of this transition isn’t to assume the problem will solve itself. It’s to treat the evaluation gap as an open research problem with the same urgency we bring to capability gains.
The thing AI most needs from humans is the thing we’re least focused on preserving. Whether that’s permanently true or temporarily true, the cost of ignoring it is the same.
Ahmad Al-Dahle is CTO of Airbnb.
Google is giving the Samsung Galaxy Z Fold 8 a serious AI feature upgrade that Apple won’t compete with for the iPhone Fold at release.
Samsung has dropped the Galaxy S26 Ultra price, but last year’s equivalent deal was much better. The RAM crisis is continuing to affect smartphone pricing.
New research into difficult-to-target proteins is exploring new ways to treat diseases that current drugs still struggle to address.
When the iPhone 18 series is released, it’s expected to feature a new modem which will improve privacy in one key way.
his week’s Apple headlines: Taking a look back at this week’s news and headlines from across the Apple world, including iPhone 18 Pro pricing, iPhone 17 US success, more MacBook Neo laptops, macOS 27 will look better, Bluey is going to take over your i…
This week’s Android headlines: The Android Show headlines, Googlebook debuts, Android 17 changes, Pixel 11 specs, Acer’s new mid-range tablet, Xiaomi 17 Max confirmed, Apple supports Google in EU AI clash, and more…
The company formerly known as Intercom just did something that no major customer service platform has attempted at scale: it built an AI agent whose sole job is to manage another AI agent.
Fin Operator, announced Thursday at a live event in San Francisco, is a new AI-powered system designed specifically for the back-office teams that configure, monitor, and improve Fin, the company’s customer-facing AI agent. Rather than replacing human support agents — which is what Fin itself does on the front lines — Operator targets the growing army of support operations professionals who spend their days updating knowledge bases, debugging conversation failures, and combing through performance dashboards.
“Fin is an agent for your customers,” Brian Donohue, the company’s VP of Product, told VentureBeat in an exclusive interview ahead of the launch. “Operator is an agent for your support ops team. This is an agent for the back office team who manages Fin and then manages their human agents.”
The announcement arrives at a pivotal moment for the company. Just two days ago, CEO Eoghan McCabe formally renamed the 15-year-old company from Intercom to Fin — an aggressive signal that the AI agent is now the business, not merely a feature of it. Fin recently crossed $100 million in annual recurring revenue and is growing at 3.5x. The broader company generates $400 million in ARR, meaning the AI agent now accounts for roughly a quarter of total revenue and virtually all of its growth.
Fin Operator enters early access for Pro-tier users starting today, with general availability planned for summer 2026.
As companies push their AI agents to handle more conversations — Fin alone now resolves more than two million customer issues each week across 8,000 customers globally, including Anthropic, DoorDash, and Mercury — the operational complexity behind those systems has exploded. Someone has to keep the knowledge base current. Someone has to figure out why the bot entered an infinite loop with a frustrated customer last Tuesday. Someone has to analyze whether the automation rate dropped after a product update.
That “someone” is the support operations team, and according to Donohue, they are drowning.
“Almost every support ops team is already doing data analysis and knowledge management — that’s table stakes today,” Donohue said. “Where teams struggle is the agent builder work. It’s a new skill set, and most don’t have enough time for it. They get their first iteration up and running, and then they get stuck.”
The problem is structural. AI customer agents are not static software. They require constant tuning — a process that looks more like training a new employee than configuring a SaaS tool. Each customer conversation is a potential source of failure, and each failure requires diagnosis, root-cause analysis, a configuration fix, testing, and monitoring. It is tedious, technical, and relentless. Fin Operator aims to collapse that entire loop into a conversational interface.
Donohue described Operator as filling three distinct roles that typically consume the bandwidth of support ops teams: expert data analyst, expert knowledge manager, and expert agent builder.
As a data analyst, Operator can field high-level questions like, “How did my team perform last week?” and generate on-the-fly charts, trend reports, and drill-down analyses across all of the data already stored in Intercom’s platform. The company has loaded Operator with contextual knowledge about customer-specific data attributes to help it interpret workspace-specific metrics accurately.
As a knowledge manager, Operator can ingest a product update — say, a three-page PDF describing a new feature — and autonomously search the company’s entire content library to identify what needs to change. It finds gaps, drafts new articles, suggests edits to existing ones, and presents everything in a diff-style review interface. The underlying search engine is the same semantic search system that Intercom has built and optimized for Fin over more than two years.
“On that knowledge management front, you just have such a time compression of something that would take, certainly hours, sometimes days, into the space of about 10 minutes,” Donohue said.
As an agent builder, Operator introduces what the company calls a “debugger skill.” Support ops teams can paste in a link to a conversation where Fin misbehaved, and Operator will trace every step of Fin’s internal reasoning, identify the root cause — often a piece of guidance that unintentionally creates a loop — propose a rewrite, back-test the change against the original conversation, and then suggest creating a production monitor to catch similar issues going forward.
“This is literally what our professional services team does,” Donohue explained. “You’ve written guidance that is unintentionally causing Fin to repeat itself — this happens a lot. You didn’t realize it, but you never gave it an escape hatch.”
One of the most consequential design decisions in Fin Operator is what the company calls its “proposal system” — a mechanism that functions like a pull request in software engineering.
Every change that Operator recommends — whether it is an edit to a help article, a rewrite of an AI guidance rule, or the creation of a new QA monitor — appears as a proposal with a full diff view. Users can inspect, edit, and approve each change before it takes effect. Nothing goes live without a human clicking “Apply.”
“Right now, we’re taking zero risk on this — Fin cannot make any changes to the system without human approval,” Donohue emphasized. “Nothing goes live until a human clicks apply.”
This is a notable architectural choice. In a market increasingly enamored with fully autonomous AI systems, the company is deliberately keeping a human approval gate in place — at least for now. Donohue acknowledged this will evolve, but said the current moment demands caution: “It’s too big a leap to just let Operator make changes automatically and then tell the team, ‘Hey, let me tell you about what I did.'”
For enterprise buyers evaluating AI tools, this design point matters. It is the difference between an AI system that proposes changes and one that enacts them — a distinction that compliance teams, security officers, and risk managers will scrutinize closely.
In a revealing technical detail, Donohue confirmed that Fin Operator does not use the company’s proprietary Apex models — the same custom AI models that power the customer-facing Fin agent and that the company has promoted as outperforming GPT-5.4 and Claude Sonnet 4.6 in customer service benchmarks.
Instead, Operator runs on Anthropic’s Claude.
“We’re not using our custom models,” Donohue said. “Those are designed to directly answer customer questions, whereas these are closer to what frontier models are best suited for. This is really closer to software engineering.”
The distinction is telling. Fin’s Apex models are optimized for one thing: resolving customer service conversations with minimal hallucination and maximum accuracy. Operator’s tasks — analyzing data, writing code-like configurations, debugging complex reasoning chains — demand a different kind of intelligence. Donohue characterized these capabilities as more akin to software engineering, an area where Anthropic’s Claude models have been deliberately optimized.
The company has not ruled out building custom models for Operator in the future, but Donohue positioned it as a lower priority. What the team has built around Claude, he argued, is the differentiated layer: the proposal system, the debugger skill, the semantic search integration, the data attribution logic, and the charting capabilities that make Operator more than just “Claude inside the app.”
Fin Operator is currently in beta with roughly 200 customers, a number Donohue said has “ramped up pretty fast the last couple of weeks.”
Constantina Samara, VP of Customer Support, Enablement & Trust at Synthesia, said the tool has already changed how her team works: “Previously, improving how Fin handles a conversation often meant reviewing everything yourself — the conversation, the configuration, the content. With Fin Operator, you just ask. It walks you through what happened and makes improving Fin dramatically easier.”
Jordan Thompson, an AI Conversational Analyst at Raylo, reported that he has been using Operator daily and has run head-to-head comparisons between Operator’s analysis and his own manual work. “It’s very accurate,” Thompson said. “It’s just as strong at high-level trend analysis as it is at debugging individual conversations. That’s a real limitation when using an LLM connector on its own — you get conversational depth but nothing on reporting or trends.”
Donohue also shared an internal anecdote from the company’s own knowledge management team. Beth, who leads knowledge operations, told the product team that Operator made her feel like she had “five more people on my team.” Whether internal testimonials carry the same weight as external customer validation is debatable, but Donohue said the knowledge management use case consistently generates the most visceral reactions because the time savings are so stark — collapsing hours or days of content auditing into roughly 10 minutes.
Fin Operator will live inside the company’s Pro add-on tier — a relatively new bundle that already includes advanced analytics features like CX scoring, topic detection, real-time issue detection, and quality assurance monitoring across both AI and human agent conversations.
The pricing model introduces something new for the company: usage-based billing. Intercom has historically relied on outcome-based pricing — charging roughly $0.99 per conversation that Fin resolves without human intervention. Operator’s work does not map cleanly to that model because it produces configuration changes, not customer resolutions.
“This has pushed us to a different model, to go more into that usage model for support ops teams,” Donohue said. “We’ll try to be generous with the usage amounts that come into Pro, but for people who are leaning heavily in, we’ll have the ability to buy more usage blocks.”
The shift is worth watching. Outcome-based pricing was one of the company’s most distinctive market positions — a bet that customers would pay for results rather than seats. Extending that philosophy to internal operations work proved impractical, which suggests that as AI agents take on more diverse roles within an organization, the pricing models that support them will need to become equally diverse.
Fin Operator lands in an increasingly competitive landscape. Zendesk, Salesforce, Sierra, and a constellation of AI-native startups are all building some version of AI-powered support operations tooling. The broader AI automation market is projected to reach $169 billion in 2026, according to Grand View Research, growing at a 31.4% compound annual rate.
But Donohue argued that Operator’s differentiation lies in two areas. First, breadth: Operator works across the full surface area of the company’s configuration system — data, content, procedures, simulations, guidance, and monitoring — rather than addressing a single narrow use case. Second, the fact that it spans both AI and human operations.
“Most critically, where I think we have the most differentiation is because it’s for your human system and your AI system,” Donohue said. “That’s really one of the unique spaces we have — to have a first-class AI agent and a first-class help desk, and Operator works across both.”
The competitive positioning also benefits from timing. The company’s recent corporate rebrand from Intercom to Fin signals a wholesale commitment to AI that legacy players may struggle to match. As CEO McCabe wrote in announcing the name change, the AI agent “is about to be the largest part of our business.” The help desk product continues as Intercom 2, but the parent company now carries the name of its AI agent — a branding move that some industry observers have interpreted as pre-IPO positioning. The Fin API Platform, launched in early April, adds another dimension: the company opened its proprietary Apex models to third-party developers and even offered to license the technology to direct competitors like Decagon and Sierra.
Step back from the product specifics and Fin Operator represents something potentially more consequential than a new dashboard or analytics tool. It is one of the first commercial products to explicitly embody the emerging paradigm of AI agents that manage other AI agents — a two-layer abstraction that is beginning to reshape how companies think about operational software.
Donohue was emphatic on this point. The real paradigm shift, he argued, is not the chat interface replacing buttons and menus. It is that the AI is doing the actual knowledge work — figuring out what should change, why, and how.
“The UX change is secondary, even though it’s most visible,” Donohue said. “The change is that we are identifying and doing the work of support operations. It’s doing the work of what the knowledge manager is doing, so that they just have to approve that. That’s the huge shift.”
The analogy to software engineering is apt. Over the past year, AI coding agents have fundamentally altered the daily workflow of developers, shifting their primary responsibility from writing code to reviewing and guiding the AI that writes it. Donohue sees the same transformation arriving for support operations professionals.
“Software engineers — three months have upended their world, where their primary job now is managing agents who are actually writing the code,” he said. “Similarly now, support ops, your job is to manage an agent who’s managing the agent for your customers.”
Whether this vision pans out at enterprise scale remains to be seen. The company is still launching Operator in beta precisely because it wants to keep refining quality through what Donohue described as a painstaking, conversation-by-conversation debugging process. “We’ve spent three months, conversation by conversation, learning, fixing, learning, fixing, to get it where it’s robust,” he said.
But if the early returns hold, Fin Operator may preview what the next generation of enterprise software looks like: not tools that help humans do work faster, but agents that do the work themselves, subject to human judgment and approval. For customer service leaders already running AI agents in production, the question is no longer just “how good is my bot?” It is now, inevitably, “who is managing it?” And increasingly, the answer is another bot.