News
As AI copilots, autonomous agents, and conversational companions continue their march into the mainstream, the teams tasked with evaluating them are no longer asking: Did the model produce the correct answer?
Increasingly, they are asking whether the system was engaging enough and created enough value for users to return tomorrow, next week or next month.
It is a shift that fundamentally changes what evaluation means.
In the age of AI, "good" is a moving target. What delights one user may frustrate another, and what offers value to one business could be deemed irrelevant by the next.
That’s why success can no longer be measured solely through generic, external benchmarks, telemetry dashboards, or "LLM-as-a-judge" scores.
Real-time signalsModels grow stronger today not by adhering to an external standard, but based on traces and real-time signals from inside the organization. Evaluation, in fact, is becoming a core part of how organizations build and protect their competitive advantage.
Companies are increasingly creating private evaluation systems in-house that measure progress against outcomes that matter to their business, using real workflows, institutional knowledge, and accumulated judgment as the standard.
The new goal is not simply to assess model performance, but to create a learning loop where human expertise continuously improves AI systems and AI systems amplify human expertise in return.
This feedback cycle turns organizational knowledge into a compounding asset. Every interaction generates new training signals, strengthens institutional memory and improves future performance.
In this new world, private evaluation is the mechanism through which firms build, retain and compound their unique intellectual capital.
Tools no longer fit for purposeIn agentic systems, the quality of the experience emerges over long sequences of interactions rather than individual outputs.
Many evaluation benchmarks still rely heavily on turn-level analysis, measuring isolated prompt-response pairs against predefined criteria. Those can identify a model's capabilities in theories – hallucinations, toxicity or syntax errors – but they can’t reliably determine whether a forgotten preference, a broken memory chain or a subtle degradation in user experience caused a user to disengage days or weeks later.
With the growing adoption of consumer AI, preference learning increasingly operates across the entire user journey instead of within isolated prompts. A dashboard can score an individual response, but it cannot fully understand why a user returned three days later, abandoned a workflow midway through a session or learned to fully trust one interaction versus another.
Those signals often live inside the product itself.
AI redefining evaluation: first-party takes the stageThis is why evaluation is moving from a support function on the sidelines to a core product capability.
Development teams are increasingly moving away from external dashboards and generalized scoring systems and building proprietary feedback loops directly into their products. These systems combine behavioral analytics, user retention data, preference learning, reinforcement signals and post-training pipelines tailored to their own applications.
Their reasoning is simple: staying close to the user is the only way to understand what "good" actually means.
The AI companies poised to win are those building closed-loop systems that connect user behavior, offline analysis, reward-model recalibration and online validation. The most advanced AI products already use live behavioral feedback to refine responses, personalize interactions and improve retention.
Static offline benchmarks are giving way to live preference learning. Isolated, single-turn tests are being replaced by trajectory-level behavioral analysis. Evaluation is moving from the support layer to the core operating system of AI products.
The fate of traditional eval vendorsSo where does this leave well-known platforms such as LangSmith, Arize, and Weights & Biases as AI providers absorb more of the stack?
Interestingly, the biggest threat these companies face may not come from Anthropic or OpenAI, but from their own customers.
These vendors are not necessarily being displaced from above. They are increasingly being bypassed and disregarded from below as AI companies realize that evaluation is inseparable from the product itself.
Generic third-party platforms can determine whether an answer resembles benchmark data. External observability vendors can process telemetry and surface analytics. But they are not truly connected to the context that increasingly defines product quality.
Every serious AI company is now discovering the same thing: evaluation is the product. And companies cannot outsource their product.
Defining success is the new critical moatThis shift is accelerating because the rest of the AI stack is becoming increasingly commoditized.
Access to high-performing foundation models is rapidly expanding. Startups and enterprises alike can access powerful APIs with relatively low barriers to entry. Prompt engineering is unlikely to remain a durable differentiator. Basic orchestration layers are becoming standardized. Even generic "LLM-as-a-judge" scoring systems are increasingly available as built-in platform features.
As those layers become commodities, a company's internal understanding of user success becomes a more important differentiator.
The signals that matter for a coding assistant differ from those that matter for a healthcare agent, a tutoring system or an AI companion. Even within the same category, companies may optimize for entirely different outcomes, including engagement, trust, efficiency, emotional resonance or long-term retention.
In a world where models, infrastructure and tooling are increasingly rented, defining success is something that can and must be owned. In fact, it may become the most important piece of intellectual property a company owns.
We've reviewed, rated, and ranked the best business software.
This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today.
The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit
- LinkedIn will hold GPU investment and its compute and storage footprint flat rather than expanding its AI data centers
- The company says it roughly doubled the efficiency of its existing GPUs in six months through accumulated improvements to utilization, model distillation, and workload allocation, not a single breakthrough
- This makes it an exception to the rule under owner Microsoft's umbrella
LinkedIn has revealed it does not plan to spend aggressively on expanding its AI data centers over its next fiscal year - instead keeping GPU investment flat and holding its compute and storage footprint roughly where it is.
The stated reason is not restraint for its own sake: the company says it has found ways to get about twice as much out of its existing GPUs over the past six months and intends to spend that headroom on new features rather than new hardware.
Speaking to Wired, Erran Berger, LinkedIn's engineering CTO, framed the goal as keeping the compute footprint flat or close to it while still shipping more compute-hungry things into production - all while acknowledging this is an unusual position to take publicly right now.
Skipping the norm, both at the industry-level and the parent companyLinkedIn has been wholly owned by Microsoft since its December 2016 acquisition of the company for $26.2 billion.
Despite being part of the software giant, LinkedIn's approach seems vastly different from that of a company that recently saw its shares rally hard as it showcased AI spending finally resulting in measurable returns.
Microsoft also recently closed its own fiscal year, reporting $41 billion in capital expenditure in the most recent quarter alone. It added 31 data centers in that quarter and 88 across the year, and expects to spend more than $50 billion in the current quarter.
This draws an interesting parallel: while one part of Microsoft is deploying capital at a rate few companies in history have matched, a subsidiary with more than a billion users has decided to sit out the year, citing efficiency gains. This makes it a considerably more interesting approach than what would otherwise be a "company shows AI discipline" affair.
While parallels exist, it is prudent to point out that Microsoft's spending is overwhelmingly driven by Azure customer demand, and specifically by capacity it has contracted to supply to OpenAI, rather than by internal product workloads, while LinkedIn's compute is a rounding error against that.
Despite that, it offers an interesting alternative perspective when a large consumer platform can add generative AI features for a year without adding hardware; still, it runs counter to the prevailing assumption that AI features and capacity growth are inseparable.
What makes this possible for LinkedIn?The answer runs back to a decision most coverage treated as an embarrassment at the time.
LinkedIn announced in 2019 that it would migrate its infrastructure onto Azure under a project codenamed Blueshift. In 2022, it quietly shelved that plan. An internal memo at the time cited Azure's own demand pressures and said LinkedIn would focus on scaling its on-premises infrastructure instead; subsequent reporting established that the migration had also run into difficulties because LinkedIn's in-house tooling did not transfer cleanly to Azure.
LinkedIn instead committed to its own data centers in Oregon, Texas, and Virginia.
LinkedIn's CTO for infrastructure, Raghu Hiremagalur, now argues that owning the full stack is precisely what makes this year's plan feasible, because the company can instrument every layer and treat efficiency as a standing investment rather than a one-off cost-cutting exercise.
“I really want to double underscore that for a company of our scale, to say a full year we're going to do this with no incremental storage and compute is no small feat, but it's taken a ton of work to get there," Hiremagalur said.
Reading what, in 2022, was a retreat as a 2026 advantage is self-serving, but it is not obviously wrong. The efficiency work itself is described as an accumulation rather than a breakthrough: better GPU utilization and workload allocation, distilling larger models into smaller ones, and rethinking how work is divided across training, inference, storage, and systems design.
Hiremagalur has also said the cost of serving each query had been climbing steadily while stored data was doubling annually, which he characterized as unsustainable. That is the more revealing framing. The efficiency push reads less like a strategic choice about the AI market and more like a company that looked at its own cost curve and decided it had to bend.
It makes it worth pointing out that LinkedIn has not really solved anything; It is that a platform of this size has publicly said out loud that its compute constraint is deliberate, at a moment when the four largest US hyperscalers have committed to something in the region of $600 billion to $700 billion of capital expenditure for the calendar year between them.
Almost every incentive in the industry currently runs toward announcing capacity rather than efficiency, making it an interesting outlier in a field dominated by daily capex announcements.
The more useful question is whether the approach holds. If it does, the argument that AI product ambition requires proportional growth in hardware gets meaningfully weaker. If it does not, this will read as an efficiency drive that met the hardware demands of a real product roadmap but failed to do so at a time when AI spending is increasingly scrutinized, even as LinkedIn itself recently allowed users to mark what they feel is 'AI slop'.


