The Weekly
Jul 20–26 | Control planes own AI · Evals over PRDs · Capital traps
The AI moat is moving from models to control planes
The useful AI debate this week was not whether OpenAI, Anthropic, Kimi or Qwen has the best model. It was whether the model is still the strategic centre of the stack. Fireworks says it is already processing more than 40 trillion tokens a day, mostly through customised models. OpenRouter was described as having roughly half its traffic through China-created models. Enterprises are mixing frontier, open, small and domain-specific models because no single provider cleanly fits cost, trust, latency and quality across every workflow.
That points to a more uncomfortable conclusion for founders: generic intelligence is becoming infrastructure faster than many application companies want to admit. If the base model keeps getting cheaper and more interchangeable, the valuable layer is the one that routes the task, chooses the model, applies proprietary data, evaluates the output and learns from production feedback. The control plane decides when to use Claude, when to use Kimi, when to use an open model, and when the task deserves a tuned internal model.
The contrarian read is that “own the model” is the wrong instinct for most companies. The better question is whether you own the workflow intelligence around the model. Open source is not just a low-cost substitute for frontier APIs. In many enterprise settings, it is the mechanism that makes custom intelligence, data control and cost discipline possible.
The operator consequence is immediate. Do not leave model choice to procurement after the roadmap is set. Map your workflows by sensitivity, quality requirement, cost per task and switching risk. Build evals and routing before usage explodes. If your product runs on proprietary work, labels or judgement, the moat is unlikely to be a chatbot wrapper. It is the system that keeps choosing, testing and improving the intelligence behind the workflow.
Evals are replacing PRDs as the real AI product spec
Dianne Penn, Anthropic’s first technical PM, gave the cleanest operating lesson of the week: “evals are the new PRDs.” That line is easy to turn into a slogan, but the underlying shift is more important. AI product work is no longer a linear process where a PM writes the requirements, engineering ships the feature and the market responds. The product surface, the model behaviour and the failure modes co-evolve.
Early Anthropic had around five product engineers and one engineer for the API business. Some early eval sets were only 30 to 40 examples, built around concrete failures like bad JSON or weak schema adherence. Later, some evals reached 99.9% or 100%. The lesson is not that evals need to be grand research artefacts. It is that they create a shared language between product, research and users. They turn vague pain into something the system can repeatedly test.
The provocative read is that a lot of high-level AI strategy is a substitute for doing the gritty work. Reading outputs, collecting failures, writing examples, testing edge cases and deciding what not to build are becoming the core PM skills. Mercor’s Osvald Nitski made the same point from another angle: as coding agents reduce the value of raw execution speed, judgement about business impact and product boundaries becomes more valuable, not less.
For operators, the move is simple but demanding. Replace fuzzy AI specs with failure-mode documents, eval sets and verification plans. Separate “writing the thing” from “checking whether the thing is right.” Make managers use the tools, read transcripts and work with outputs directly. Token spend is only an input. The scarce resource is the speed and quality of organisational learning.
Data is back because AI creates more state, not less
The old enterprise software story treated data infrastructure as a mature, slightly boring layer beneath the action. AI breaks that story. CJ Desai’s framing of MongoDB as the “memory layer” for AI is self-interested, but the logic is hard to dismiss. Agents, inference, voice, image, video and long-running workflows all create state. That state has to be stored, retrieved, secured, audited and governed.
The more AI is used, the heavier the infrastructure burden becomes. Desai described enterprises mixing open, closed, small, large, horizontal and domain-specific models. He also cited a Texas hyperscaler customer that wanted more cloud capacity and was told there was none, plus a large US telecom forced to sign with a second hyperscaler because regional capacity was unavailable. “Cloud-first” is no longer a one-way march if capacity, sovereignty and resilience are back on the table.
This is the less fashionable side of AI adoption. Demos make agents look weightless. Production makes them stateful. Desai said one enterprise agent architecture still had 55 boxes, even if that was about half as many as it might have had a year earlier. Enterprise buyers are not just asking whether an agent can act. They are asking where the data lives, who can inspect it, whether outcomes are deterministic enough, and what happens when the system gets audited.
The tactical implication is to treat observability, auditability and data ownership as product features, not admin chores. If you are building AI into a product, map what new data the system creates, where it sits, who owns it and how it can be deleted or inspected. If you sell into enterprise, assume mixed-model, multi-cloud and on-prem reality. The winning infrastructure layer will be the one that survives heterogeneity, not the one that assumes a clean model monoculture.
Constraints are not friction. They are the operating system
DHH’s Basecamp story and Senra Systems’ manufacturing story look unrelated until you strip them back. Both are arguments for constraint as a quality mechanism. Basecamp 1 was built on about 10 hours a week and 380 hours total. The old 2004 version stopped being sold in 2010, yet still has paying customers and still makes millions in profit. Senra’s Chris Venner, meanwhile, argues that process should lag scale, not lead it, and that the repeatability of a manufacturing model is only really proven after three sites.
The shared insight is that speed without selection creates waste. AI makes software easier to produce, which means the default failure mode becomes overbuilding. Capital, headcount and tools remove the natural friction that once forced taste. In hardware and manufacturing, the same pattern shows up as premature process, handoffs and coordination drag. Venner’s claim that a six-person Senra team can do what once took roughly 175 people at SpaceX is not just an AI productivity story. It is a warning that coordination layers are about to be compressed.
The contrarian read is that more process can be anti-scale, more features can weaken adoption, and more capacity can make teams less disciplined. DHH’s “less software” strategy is not nostalgia. It is a retention strategy: fewer things to learn, fewer things to explain, fewer ways to confuse the customer. Venner’s “law of threes” is the industrial version of the same point: one site can be craft, two can be promise, three starts to prove a system.
Operators should turn refusal into an operating habit. Keep a visible kill list of features that are good but make the product harder to adopt or defend. Ask whether each process exists for accountability or managerial reassurance. Look for hidden “cable harness” problems in the business: unglamorous but essential work that remains manual, fragmented and fragile. Constraints are not the enemy of growth. They are often what preserves quality as growth arrives.
Ownership, not income, is the political and commercial fault line
David Friedberg’s strongest argument was not the familiar one about inequality. It was about ownership. His claim is that America’s real divide is between people who can compound capital and people whose labour income gets absorbed by housing, healthcare, education and taxes before they can build a balance sheet. In that world, wages matter, but they are not enough.
His Social Security example is deliberately provocative: he argues the trust fund would be about $37 trillion larger if contributions since 1982 had been invested in the S&P 500 rather than Treasuries. The precise policy path is debatable, but the framing is useful. Retirement policy, compensation policy and consumer finance are really ownership policy. Friedberg’s proposed national KPI, converting 2% of Americans from labour to capital each year, is a sharper metric than another abstract debate about prosperity.
The contrarian read is that the loudest political energy may be a symptom of asset poverty. People without capital experience the economy as a series of price shocks, not as a compounding machine. That changes what they buy, what they resent, what they believe and which promises feel credible. The commercial implication is that affordability pressure is not a mood. It is a structural condition.
If you build in fintech, payroll, benefits, housing, education or compensation, design for accumulation rather than transaction volume alone. Help users save, own, hedge, compound or participate in upside. In GTM, sell cost avoidance and balance-sheet improvement before abstract convenience. Inside companies, treat equity, retirement contribution and long-term ownership as part of the talent system, not as benefits-table decoration.
Market structure quietly decides what a business becomes
Matthew Smith’s discussion of private capital was nominally about natural gas and private credit, but the reusable lesson was market structure. The system moved from post-1933 separation, to the pre-GFC merged-bank era, to a post-2010 structure where commercial banks are more tightly guarded and private capital fills much of the gap. Private capital grew from roughly $2 trillion pre-GFC to $14 trillion to $15 trillion today. Private credit grew from around $500 billion to about $2 trillion.
The lesson is that the liability side shapes behaviour. Once a firm raises capital at scale, the machine starts to demand deployment. Narrow products become easier to sell. Distribution starts to corrupt underwriting. Fee-related earnings multiples moving from roughly 10-15x to 25-30x plus rewarded industrialised fundraising, but also created incentives to keep the factory fed.
“There is no such thing as semi-liquid” is the line worth carrying beyond asset management. A product built on illiquid assets but sold with liquidity language has a truth problem. The same pattern appears in AI compute commitments, enterprise contracts and fintech balance sheets. Apparent growth can be a symptom of bad product design if the promises made to customers do not match the underlying asset or operating capacity.
Operators should audit their own version of liability mismatch. What have you promised customers, employees, vendors or investors that depends on normal conditions staying normal? What happens if customers want money, compute, service capacity or commitments back faster than expected? Do not optimise for inflows if the product cannot support the outflows. Growth is only healthy when the plumbing underneath it matches the story being sold.
Enterprise AI still needs deployment muscle
Mercor’s story is a useful corrective to the clean software narrative. It describes two modes, “talent only” and “managed service,” and says environments are its fastest-growing data type. Osvald Nitski’s point is that services are a knowledge dissemination problem. A small number of teams know how to deploy agents, run evals and build AI-first workflows. Most enterprises do not.
That means the early enterprise AI market will not behave like pure self-serve SaaS. Buyers need trust structures, data controls, model provenance, workflow mapping and implementation help. They are not deciding between “use AI” and “do not use AI.” They are deciding which workflows are safe enough, valuable enough and legible enough to move first. HR and procurement may be easier. Client advice, legal memos, underwriting or core product work will face a much higher bar.
The contrarian read is that “services are just bad product” is too simple for this phase of the market. In early AI, services can be the bridge that teaches the buyer what is possible and teaches the vendor what should become product. The danger is letting bespoke work become the business instead of a learning wedge.
For founders and GTM teams, deployment is not an after-sales problem. Build a trust memo before the enterprise buyer asks for it: data flow, hosting, access control, auditability, model provenance and human review. Use services where they accelerate learning, but aggressively productise the repeatable parts. The near-term winner may be the company that can deploy working systems, not the one with the best demo.