{"id":6199,"date":"2026-07-21T19:00:33","date_gmt":"2026-07-21T23:00:33","guid":{"rendered":"https:\/\/workai.tv\/news\/2026\/07\/ai-engineering\/the-biggest-challenge-in-your-stack-evals-evals-evals\/"},"modified":"2026-07-21T19:00:33","modified_gmt":"2026-07-21T23:00:33","slug":"the-biggest-challenge-in-your-stack-evals-evals-evals","status":"publish","type":"post","link":"https:\/\/workai.tv\/news\/2026\/07\/ai-engineering\/the-biggest-challenge-in-your-stack-evals-evals-evals\/","title":{"rendered":"The biggest challenge in your stack? Evals, Evals, Evals"},"content":{"rendered":"<h2>Share with your CTO<\/h2>\n<p>Eight practitioner talks from the AI Engineer Melbourne keynote and World&#8217;s Fair surface a consensus that most enterprise AI teams are building on a cracked foundation: they have no reliable way to measure whether their models are actually working. Across <a href=\"https:\/\/finance.biggo.com\/podcast\/108e1daab4c6f435\" target=\"_blank\" rel=\"noopener nofollow\">sessions covering model selection, agent memory, voice infrastructure, and security<\/a>, the same failure mode recurs. Notion runs a model picker that re-evaluates its default every three to four weeks. Snyk&#8217;s benchmarks show frontier LLMs catching only 75% of vulnerabilities that deterministic tools catch, with an F1 score of 40% for LLM-only detection. Pruna&#8217;s CEO demonstrates that the top-ranked model on public leaderboards loses 40% of its head-to-head battles on specific tasks. The unifying prescription: build your own evaluation infrastructure or cede all strategic leverage to your model vendor.<\/p>\n<h2>What this means for your business<\/h2>\n<p>The Pareto problem is already here. Your engineering team is almost certainly picking models based on general leaderboard rankings or manual spot-checks, neither of which reflects your actual production distribution. Notion&#8217;s &#8220;auto&#8221; model picker, which reroutes 75% of customer workloads every few weeks, is not a curiosity. It is the logical endpoint of treating model selection as a continuous operational discipline rather than a quarterly procurement decision. Teams that cannot run that loop are overpaying, often by 3x or more, for capability they cannot verify they need.<\/p>\n<p>The security dimension makes this urgent in a way that pure cost arguments do not. Snyk&#8217;s data is damning: enterprise vulnerability backlogs grew 108% year-over-year in organizations using AI-assisted code generation, not despite AI, but partly because of it. The generator-validator separation principle, using a different model or deterministic system to check the output of your code-generating model, is architecturally sound and empirically necessary. The same LLM that writes your code has structural incentives, baked into its training, to validate its own output favorably. That is not a solvable prompt engineering problem. It is a system design problem.<\/p>\n<p>The signal worth watching: the gap between what frontier models cost and what deterministic computation costs is widening, not narrowing. Voice agent teams running one million outbound calls per day are using Rust state machines and regex for anything that can be computed, reserving LLM calls strictly for arbitrary interpretation. That discipline, applied to code review, compliance checking, and data validation, is where the next wave of cost reduction lives for enterprise AI stacks. The question worth holding is whether your current architecture has any layer of deterministic verification at all, or whether you have handed every decision to a model you cannot audit.<\/p>\n<h2>Concept deep-dive: Evals infrastructure<\/h2>\n<p>Evals infrastructure is the internal system a team builds to measure whether its AI models are performing correctly on the team&#8217;s own tasks, using the team&#8217;s own data. It exists because public benchmarks are built on generic samples that do not reflect any specific enterprise&#8217;s prompt distribution, user behavior, or quality threshold. Think of it as the difference between a car manufacturer&#8217;s crash-test rating and your own driver&#8217;s accident history. The business connection is direct: without evals, every model upgrade, price change, or provider deprecation forces a guessing game. With evals, model selection becomes a data-driven operational decision rather than a vendor relationship.<\/p>\n<p><em>Based on reporting from <a href=\"https:\/\/finance.biggo.com\/podcast\/108e1daab4c6f435\" target=\"_blank\" rel=\"noopener nofollow\">&#8220;The biggest challenge in your stack? Evals, Evals, Evals&#8221;<\/a>, originally published 2026-07-21 06:41:00.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Share with your CTO Eight practitioner talks from the AI Engineer Melbourne keynote and World&#8217;s Fair surface a consensus that most enterprise AI teams are building on a cracked foundation: they have no reliable way to measure whether their models are actually working. Across sessions covering model selection, agent memory, voice infrastructure, and security, the [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[145],"tags":[],"tmauthors":[],"class_list":["post-6199","post","type-post","status-publish","format-standard","category-ai-engineering"],"_links":{"self":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/posts\/6199","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/comments?post=6199"}],"version-history":[{"count":0,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/posts\/6199\/revisions"}],"wp:attachment":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/media?parent=6199"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/categories?post=6199"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/tags?post=6199"},{"taxonomy":"tmauthors","embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/tmauthors?post=6199"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}