The Open-Weight Unbundling
with Jake Saper, General Partner @ Emergence Capital
Jake Saper is a General Partner at Emergence Capital, where he’s spent 12 years backing companies at the frontier of enterprise software, including early bets on Zoom, Veeva, Ironclad, and Physical Intelligence. In 2024, Jake coined the term AI Native Services in his essay The Death of Deloitte, arguing that AI would fundamentally reshape professional services.
Emergence was also an early bettor on open-source AI infrastructure, leading investments in Together AI, which just raised $800M at an $8.3B valuation, and in OpenAI’s $4B deployment spinout, DeployCo.
Tune in to hear our in-depth discussion with Jake on why the open-weight surge is unbundling the model labs and the coming reckoning over where value actually accrues in Vertical AI.
Today’s Episode
Last week, Chinese model lab Moonshot AI released Kimi K3 — a 2.8T parameter open-weight large language model. Kimi K3 is the first open-weight model to crack the frontier to this extent — scoring first in a popular front-end test (ahead of both Fable 5 and GPT-5.6 Sol) and second only to Fable 5 on another measuring long-horizon knowledge work. And while not “free,” as open-source1 might imply, K3 achieves the above at roughly half the cost per task (and that’s on a relatively favorable Opus 4.8 comparison). K3’s full weights go public on July 27.
K3 is just the latest in a cascade of open-weight models. It is estimated that open models are used by 80% of startups and represent 25-50% of volume across major platforms like OpenRouter and Vercel. Together AI’s annual bookings crossed $1.15B as customers seek to cut ballooning inference costs. The gap between frontier and open-weight performance — once measured in years — has compressed to 4–6 months, and appears to be shrinking.
Jake and Emergence called this shift early, backing Together AI in 2023 on the conviction that enterprises would want the flexibility, security, and cost advantages of open-source models. In July 2026, the question is no longer whether open weights will be in demand, or matter — it’s what happens when the model becomes less of a moat.
The Tokenmaxxing Hangover
Tokenmaxxing describes a phase of eager, first-wave LLM adoption in which enterprises encouraged pure token usage volume from its employees, presumably to accelerate adoption, value, or just generally prove they were AI-pilled and braced for the future. Some of the results are unsurprising: Uber blew through its annual AI budget in a quarter, Amazon shut down its leaderboard because employees were setting up agents in pointless loops just to climb the ranks; founders bragged about monthly token spends in the millions.
Jake uses a fitting term for the aftermath, which is no doubt a catalyst for the open-weight surge: the tokenmaxxing hangover.
There was an era three months ago when there was a lot of pride that companies would take in how much their token spend was going up as an indicator of — turns out — a very poor indicator of value created by AI. The hangover treatment has been to more heavily scrutinize what those tokens have been spent on and to find effectively cheaper tokens to get the same or better outputs.
Companies that were routing everything through frontier APIs at premium prices are discovering that open-source models — used judiciously, in the right use cases — can deliver comparable quality at a fraction of the cost. This is all pushing the ecosystem toward a multi-model paradigm: frontier closed-source models for the highest-intensity tasks (a small portion of total runs) and significantly cheaper open-weight models for more quotidian tasks (the majority of workloads).
Three forces are accelerating this more diverse future of model use:
Open-weight model quality has caught up. Kimi K3 is months behind the leading closed-source labs, not years. And that trend appears to have accelerated. Few deny that foreign open-weight models — less exposed to US IP law — are distilled from frontier models, likely putting a natural limit on how small the gap can get for holistic performance. But the capability difference simply isn’t as noticeable as it once was, and in certain domains, it’s arguably non-existent.
The US federal government has shown a willingness to intercede. Both to gate powerful frontier models (as they did with Mythos) and, potentially, to crack down on Chinese open-source models almost certainly distilled from US models.2 Regulations incentivize companies to further control their AI infrastructure.
Enterprises are beginning to seek their own AI moats. For example — as we’ve discussed at length here at The Verticalist — capturing, protecting, and embedding proprietary data into their stacks. Beyond this, there’s the possibility of edging into the model layer: fine-tuning an LLM with open weights to their unique purpose. That makes open source not just cheaper, but strategically essential for anyone trying to build a defensible data asset.
A quick word from our partner on the Verticals podcast, Parafin. They just launched Spend, a business credit card purpose-built for SMBs that vertical platforms can deploy in weeks, adding to lending and insurance products already powering DoorDash, Gusto, and TikTok Shop. Named to the 2026 Forbes Fintech 50, with a fresh credit facility led by Goldman Sachs, Parafin is built for founders like you. Explore a custom program for your platform today →
The Labs Respond
The frontier labs seem to be doing their best to couch panic over the open weight news in grave national security concerns. A tweet from Dean Ball, OpenAI’s “head of strategic futures” is making the rounds branding open weights as “inherently decelerationist,” likely to result in “AI communism” that he characterizes as a “dystopian hellscape.” The incentive for frontier labs to support government intervention to suppress direct competition requires no explanation. The deeper response, however, lies in their approach of new layers beyond the model.
OpenAI hired Jason Boehmig, the co-founder and former CEO of Ironclad, to lead its new legal vertical in June, a direct move into application-layer territory. Harvey and Legora have been among its largest customers. Both OpenAI and Anthropic are hiring vertical-specific leaders across industries. And the app layer hasn’t been the labs’ only expansion salvo. OpenAI’s DeployCo raised $4B backed by Emergence, TPG, SoftBank, and Goldman Sachs to embed forward-deployed engineers inside enterprises. AI implementation services are perhaps an even more peripheral adjacency to their model-layer competencies that would create TAM breathing room as infrastructure is squeezed.
While their call for open-weight regulation is louder, the layer expansion feels even more telling about how seriously the labs are taking the potential erosion of their core moats. How successful these new initiatives will be is a separate question we will be writing more on soon. But as we wrote about Claude for Science last week (and OpenAI’s legal product is not dissimilar), the vision is nascent — today, a core model plus a set of connectors. As we discuss with Jake, that’s not what made vertical software companies defensible. While Vertical AI undoubtedly moves faster than its SaaS forebears did, it’s not just a question of product development — those with the strongest ostensible moats (e.g. Veeva, Fair Isaac, Epic) have gone deep on unique workflows and parsing of first-party vertical data for years.
The labs’ dilemma is that their core business is under commodity pressure from open weights. Therefore, moving up the stack to applications and / or services looks like a natural escape hatch — and one they may be uniquely positioned for, if only due to their vast balance sheets. But that means competing with their own API customers in a very existential way. It’s a self-reinforcing loop: the labs go vertical, which makes their customers warier, which accelerates the shift to open source, which puts more pressure on the labs’ core margins, which pushes them further into vertical plays... and so on.
As they look to their own long-term defensibility, even enterprise operating businesses are catching on. As Jake put it, if you’re a pharma company and you give a model lab all your proprietary molecular research, you’re handing your most valuable IP to an entity that’s now hiring vertical leaders and building competing products. “I think for good reason, people are increasingly wary. And that’s another reason why open source is being utilized — so people can keep their data within their house.”
As we covered in The Dispatcher Problem, the big question is who captures value as the marginal cost of intelligence falls toward zero. The recent step-changes in open-weight capabilities and the shrinking frontier gap make that question more urgent. When the model is no longer a meaningful differentiator, value migrates to whoever controls the workflow, the data, and the customer relationship. Perhaps the value also flows to the third parties facilitating all these value-chain shifts (e.g. integrators and other AI services).
We’ve said it before and we’ll say it again — we believe the immutable primitives of app-layer defensibility are workflow depth and data gravity. With more margin to go around, cheaper tokens only make the Vertical AI opportunity more compelling. And in a world of model-layer commoditization, deriving defensibility from the app layer has never been more important. The model labs themselves seem to agree.
Controlling Your Own Destiny
The best Vertical AI companies aren’t waiting for the labs to sort themselves out. They’re already taking steps to cement their app-layer moats — and even to take incrementally more control over their AI infrastructure.
Harvey is a visible example. The legal AI company is training open-source models to encode law firm workflows while simultaneously running a multi-model stack across Anthropic, Google, and OpenAI. The goal is to ensure that no single provider has leverage over the business. As Jake put it, “You’re seeing a lot of the leading agentic software companies focus an increasing amount of their spend on fine-tuning their own open-source models versus relying on the closed-source labs.”
What makes this possible is the convergence we’ve been describing: open-source models are good enough for the majority of workloads, and fine-tuning them with proprietary data is feasible only when the weights are open (unless you want to fund ground-up LLM training yourself). Making proprietary evals3 more valuable than they’ve ever been.
Jake made a related point worth noting: the companies that own the full outcome — AI-Native Services, or AINS4 — may be in a unique position to gather that eval data. A fund administrator like Hanover Park (a past guest on Verticals) sees the complete lifecycle: the AI output, the human review, the client feedback, the audit result. A software vendor selling into the same fund admin sees a truncated slice.5 That’s a separate argument from the open-weight thesis, but it points in the same direction: value is migrating from model access to the proprietary data and workflows built on top of it.
The Takeaway for Vertical Founders
The open-weight unbundling is a control story.
A year ago, Vertical AI companies were largely dependent on one or two frontier labs for model access, fine-tuning, and the roadmap. That dependency is unwinding. Open-weight models today are good enough for many tasks. The cost advantage can run as high as 10x. And the frontier labs may be accelerating the process by making dependency more dangerous, as they seek to compete with their own API customers.
From what we’ve learned so far, the founders navigating the unbundling well are doing three things:
Going Multi-Model: running multi-LLM architectures that route by complexity and cost and function. Knowing what works, for what, and why is an edge.
Exploring Fine-Tuning: supercharging open-weight models with proprietary data, turning customer feedback into a compounding asset beyond simple reinforcement learning. The app layer strikes back.
Depthmaxxing: Maximizing channel checks, customer data depth / breadth, champion relationships, physical proximity (via forward-deployed X), ICP granularity, and general user obsession. Businesses are both greedy and fearful around AI today. They want partners who are on their side and will help them take advantage — but only businesses that go all-in on understanding them will earn their trust (and thus the right) to solve their big AI problems. Founders securing airtight footholds and relationships are finding they can offer more (and more upmarket) than they expected.
As the cost of intelligence falls and the weights open up, the bulk of LLM use recedes into infrastructure. What remains is the vertical knowledge, the proprietary data, and the willingness to stand behind the outcome. As we’ve argued in many past essays, what makes a Vertical AI business durable hasn’t changed — open weights just make the ultimate prize dearer and clearer.
See you next week.
Key Moments from this Episode
00:00 — Why the AI moat is changing
05:40 — Open source catches up to frontier models
10:58 — Why enterprises are embracing open source AI
16:20 — Can OpenAI win the application layer?
24:01 — Why services are becoming software
31:28 — The race to build proprietary AI data moats
36:38 — What AI-native services get right
40:36 — Why brand still matters in AI
45:49 — The AI founder playbook for the next decade
48:37 — The biggest AI opportunities still ahead
Subscribe to Verticals to get new episodes every week, available wherever you watch or listen:
A note on terms: this piece uses "open source" and "open weight" loosely, following general parlance. But they are not the same thing. Kimi K3, DeepSeek, and Qwen are open weight — you can download the parameters and run them. The training data, the data pipeline, and the training code, however, are still proprietary. True open-source models (in the sense the OSI definition requires) release enough for a competent third party to fully reproduce the model: weights, training code, etc. Very few “frontier” models would qualify (e.g. AI2's OLMo). The distinction matters for at least two reasons: (1) weights, not reproducibility, commoditizes inference pricing; and (2) anyone claiming releases like Kimi K3 represent open science rather than a competitive strategy are over-simplifying — even in cases where such strategy is geopolitical as much as it is commercial in nature.
Although, it’s important to note that even OpenAI has admitted distilling can’t explain the extent of K3’s outperformance. Not to mention, it’s unlikely US model labs haven’t dabbled in distillation from time to time.
“Evals” are how you measure whether an AI model actually performs its task: structured tests that score outputs against the ground truth of messy, domain-specific, real-world edge cases. In Vertical AI, evals are the bedrock of your moat, across workflows and data — measuring your success in encoding human / expert / tribal judgment into scalable, effective Vertical AI.
Coined by Jake himself!



