# Will OpenRouter Sell for $10BN to Stripe?

Why Chinese Open Models Are Beating America—and What Happens Next · Why Enterprises Are More Fearful of Anthropic and OpenAI Than China · Is the Routing Layer Becoming a Commodity with Alex Atallah

20VC · Aug 10, 2026 · 60 min · 10,826 words
Speakers: Alex Atallah, Harry Stebbings
Source: https://www.996.fm/episodes/20vc--ep-ca4673c2/

## Cold open

**Alex Atallah** [0:00]:

This going be like the biggest, biggest, market in tech ever. A lot of companies are making routers because it's fashionable. The model labs have several incentives to go after you eventually. In July, we launched 70 models. about one model every hours. America is very, very, behind still. But GLM 5.2 was a really big, big step for open weight models.

**Harry Stebbings** [0:25]:

There are reports that you are selling to Stripe for $10 Is that to happen? This is is VC

## Intro

**Harry Stebbings** [0:32]:

with me, Harry Stebbings. Now, the only thing that I really care about anymore is providing the best, most relevant interviews at the right time to you. So today we have Alex Atallah, co-founder and CEO of OpenRouter, the gateway to the world of LLMs. They reportedly have had offers from Stripe for billion. They've raised at a valuation of over a They are the market leader, and this interview could not come at a more prescient time. It'll be very interesting to see whether the company chooses to stay private or sell to Stripe. We shall see. But this interview was recorded before, so we will check back in a couple of weeks. This was an incredible show and it was awesome to have Alex in the studio. But before we dive into the show today,

## Sponsor read

**Harry Stebbings** [1:16]:

founders face a different set of challenges at every stage of growth. For Sid Shait, co-founder and CEO of Matrix, Morgan delivered the guidance and expertise to help navigate what came next. He credits Morgan's high touch approach with supporting Dmatrix Matrix as it grew and expanded internationally. Whether you're in the early days or expanding into new markets, JPMorgan startups navigate complexity with real confidence, offering personalized guidance and deep sector expertise. Find out how JPMorgan helps founders at jpmorgan.com forward slash grow without limits. JPMorgan is the bank of the innovation economy. While JP supports growth, Corgi protects it. My word, what an arresting first line. Get your ass covered with Corgi insurance, and I'll tell you why. If you're running a business right now, you already know this pain all too well. Getting insurance, it's really slow, it's confusing, and my word, it's full of paperwork. Well, that's exactly why Corgi is here to change the game. Corgi is the first and only insurance carrier designed specifically for tech companies, allowing you to get covered in minutes instead of days. Corgi provides essential coverages for all growth stages such as D and O, E and O liability, cyber, commercial, general liability, and more. Get your ass covered. I love the way we say ass with Corgi Insurance alongside thousands of other startups at corgi.com/20vc today. That's corgi.com/20vc. You won't regret it. While Corgi covers risk, Flex gives you room to move. Business owners run their whole financial life on Flex. one platform from business revenue to their personal spend. Float every purchase for sixty days. Tap capital that grows with your revenue and pay vendors in in 170 countries across 32 currencies, plus the whole back office, bills, expenses, accounting, all in one place. So you spend less time reconciling and more time growing. That's why thousands of owners use Flex, named one of Fast Company's most innovative companies of 2026. Visit flex.one, that's dot, and use the code 20 v c.

**Alex Atallah** [3:36]:

You have now arrived at your destination.

## Conversation

**Harry Stebbings** [3:39]:

Alex, I am so excited for this, dude. I have wanted to make this one happen for a while. I've heard so many things from Matt at Menlo. I've stalked the shit out of you speaking to Anjani, even your roommate before this show. So thank you for joining me, dude.

**Alex Atallah** [3:52]:

Thank you. It's great to be here.

**Harry Stebbings** [3:54]:

Now I want start with a little bit pre OpenRouter and start on OpenSea. It was a pretty incredible journey. What did you take with you to OpenRouter having seen all that you saw with OpenSea?

**Alex Atallah** [4:07]:

Yeah. So OpenSea started as the first NFT marketplace. Similar to OpenRouter, it was very small for a long time. Like we kept the team very small until the series A roughly, or you know, a little bit afterwards. And this was before AI. Right after NFTs started blowing up in in 2020, October 2020, were like, oh my goodness. Like, we are understaffed. The servers are melting. Like our search index was exploding. We had a couple big outages. It was tough to like keep the site up. and it was like, oh my God, we're to become like Twitter fail whale, but like applied to crypto. My biggest goal was to have us not be the Twitter fail whale for crypto. And it took a little bit to like create the team, get platform and infrastructure under control, like, make sure we we could predictably scale. In other words, like do load testing to like help the site sustain 10x load even when we weren't seeing that load. Because with crypto, you just don't know. know like these moments where we would get these incredible traffic spikes and be very dependent on the content and the community. So I built like a lot of infrastructure and scaling responsibilities then that I took to OpenRouter and spent a lot of time like thinking about, okay, how do we make something that is to basically be always up and that people can really count on from an infrastructure point of view, even when there are huge surges in tumultuous markets, Which has been very helpful for AI, of course, because like all companies, especially Anthropic, have seen like unpredictable growth, and we have as well. And we've we've we've had like a couple bumps, but overall, it's been like significantly better. And And like, OpenSea just kind of like drilled that into me in a way where I could, like, take it productively to OpenRouter.

**Harry Stebbings** [6:01]:

Can I ask you, when you, go back to the founding thesis of the company, what has happened in the ecosystem in the model landscape that you did not expect to happen?

**Alex Atallah** [6:10]:

Well, one thing that we did not expect was that an ecosystem of companies would emerge to host and serve the open weight models. Like early on, it wasn't clear that that market wasn't going to be a monopoly, where like just the three hyperscalers serve all the open weight models and startups don't, you know, they're are they're really far behind. In reality, you know, how often do you hear people running GLM on a hyperscaler? Never. like they're using the the inference providers like Fireworks and Together, and there's like, big list that we see doing the best job of hosting all the the open weight models. In the early days, we had, I think we called it provider provider one and provider fallback. We didn't like show which providers were actually doing the hosting because we weren't really a marketplace. We were kind of an an exploration tool for like, finding and discovering new LLMs. And we wanted to build like, a marketplace of model labs, but like, the inference provider layer, we weren't sure would actually be a marketplace. And it turned out that those companies were doing a way better job than the hyperscalers, were way faster to host the models and figure out these edge cases to hosting them. And uptime was just to be a a constant problem. It wasn't going to like, magically get solved by the supply side of the market.

**Harry Stebbings** [7:31]:

A lot of people suggest that that inference provider commoditizable element or layer that will be removed or see margin reduction competed out over time. What would you say to that theory?

**Alex Atallah** [7:42]:

Right now, we're in a massively supply constrained market, market, and and it's it's likely likely to be supply constrained for a while where all the inference providers are short, pretty much constantly short. And you're like, okay, so GPUs are are really, really, beneficial. And like, why doesn't Google or Amazon or Azure run around and like buy up all the GPUs and take all these inference providers out of business? Well, the people making the GPUs don't want that. Like one of NVIDIA's top priorities is not having customer concentration. They want lots of customers to all have like separate allocations of GPUs. They want the heterogeneity of the market. They want like competition on the compute layer. And this is good for the ecosystem. Like, users also want this. It's good for Nvidia it's good for end users as well. It like, allows these inference providers to kind of like, come up with new innovations on like, how to serve the models better. Even a single model like Kimi K3, like Moonshot just posted a benchmark showing all the inference providers and how well they're serving Kimi K3. And they're pretty different numbers for benchmarks that are really static, are well known. We post this continuously all the time. We always are like benchmarking all of the models on all of the inference providers, all the open weight providers, and finding really different results constantly. The results change over time. These models are like very they're very emotional. They're they're very non-deterministic.

**Harry Stebbings** [9:13]:

I had Lynn on the show from fireworks and she said that you know, I said about Gavin Baker and a token is a token is what he said. And she kind of corrected me that a token is not a token, actually, because one provider can make a token go so much further than another token. It's like, how do you get to the store where you can drive around the whole block, or you can drive straight to the store. Tokens can be made more efficient and go further and that's the job of the provider.

**Alex Atallah** [9:38]:

Yeah, I agree with that. I think that in some ways we are providing a service to help people discover providers. Ultimately, when when one provider is making a token go further, we spend an enormous amount of time on our router, central router tech, so that that provider immediately gets more traffic. As soon as we detect that like there's a quality improvement or a speed up or a price reduction happening, immediately starts getting more traffic. This stuff happens like seven, every five minutes. There are big changes for the big models. and so it like actually does make the experience better.

**Harry Stebbings** [10:13]:

You can only invest in one inference provider. Which one do you invest in?

**Alex Atallah** [10:18]:

I probably have to stay, you know, stay neutral on this. I do I really like the inference providers that are doing like, custom hardware and very, very, like, low level optimizations. I like providers that are also trying to figure out how to make customization easier. Today, you fine tune models and you create this, like, new fully independent model from the base model. Many inference providers are of like creating these LORAs or some some call them cartridges are much more portable potentially between models. And And we might see a future where when you do a fine tune and you to like, change the base model layer, it only costs like maybe a few dollars, maybe a few dozen dollars to change it.

**Harry Stebbings** [11:01]:

It's okay. I understood that fireworks is your favorite. It's okay. I get it. Mine too. My question is, when Lynn was on the show, she was like, oh, you don't to rent your intelligence, you to own it. And we're to see companies have specialized models, which is trained on their own data and proprietary to them. In a world of every company having specialized models that's really tuned to them and their preferences, is that good for an OpenRouter business or not? Oh, definitely. I mean Why? Because you'd stick on one model, which is yours, proprietary, trained on yours, and not be open to the diaspora of models that is available.

**Alex Atallah** [11:39]:

No, I disagree. I I think our mission from the very beginning has been to increase neurodiversity in AI for the whole ecosystem. And we really believe that like, a multimodal future is inevitable. When you start, let's say there's one model that, like, hypothetically, let's say you're right. Let's say there's one model that fulfills all of your desires, either within your company or, like, as a consumer. More and more people start using that model. And then someone decides, you know what? I'm going to, like, create a neurodivergent model. I'm to create a model that's like, a little bit different, that, like, talks a little differently, that has ideas that the the first model like, could never have come up with because it's like completely different data that's being used to train it, then it of creates inevitable demand to use both models. creativity is not verifiable. Like you can't really put an easy number on on creative ideas. And when you use two models together, you're more likely to get creative ideas than if you just use one. It's just fact if if that other model was trained in a different way on a different set, or has like, made a big update, consolidation on one model just seems like it just doesn't make any sense to me.

**Harry Stebbings** [12:47]:

Totally get you. So you'll have companies which have like a core workflow or their core, which is their own specialized model, and then they'll use a plethora of other models and they'll use OpenRouter for those other model selection.

**Alex Atallah** [12:58]:

Yes. And I think that when companies make like, to get back to your question, when they make their own model trained on their own data, the ecosystem around you is all doing the same thing. You have to like play out the game theory for these things a little bit. Like if everybody is doing this as well, and all the model apps are creating new models constantly using new data that they've acquired, that they've bought from other companies, that's all like potentially data that's valuable to you. What is in your best interest? It's to go and try out those other models and like, see if you can be more productive with them, if you can like merge them together to get better state of the art performance, if you can reduce your costs using these other models. Whether your goal is to improve your margins or grow your company, you are incentivized to go use what the ecosystem creates. So the the model that you made, you're gonna have to continuously improve it to keep up, and it's never to win the whole market. So it's going be a massive market. This is to be like the biggest, biggest, market in tech ever, and biggest market probably in human history. No one's to win all of it. You're not to build a model that wins all of it. So you might as well build a model that is is known to specialize in something very useful and that's very important to your company and your business and be known for that specialty. And I think a lot of enterprises are to move that direction, make their own models, make their own branded intelligence. Your brand is a big part of your moat, and that model will be be a way your brand carries around.

**Harry Stebbings** [14:26]:

You mentioned the immense time that you spend on the routing technology that you have. A lot of people are thinking that we're seeing the commoditization of the routing technology. You're seeing ramp release products like this. I mentioned earlier of Merge, company we invested in has released that product. Several are releasing kind of routing technology similar or claiming to be similar. Are we seeing the commoditization of this layer?

**Alex Atallah** [14:48]:

I think a lot yeah, a lot of companies are making routers because it's fashionable. You know, they're seeing growth happen here or they're making gateways at least. There's two issues with First, it immediately puts you in the mindset of copying instead of like, you know, winning something. You're playing to play, you're playing to exist rather than playing to win. And maybe you're just trying to like play to serve your your existing customer base and you you want see some AI growth happen. I think you know, it immediately kind of like puts that gateway many, many, months behind the companies that are fully focused on it. Like I am a 100% focused on building the best router and gateway and LLM marketplace. And it shows in our product and you know, the the benchmarks that we create internally and how we see ourselves compared to the competition. This is not a side quest for us like it it may be for some other companies. The other problem is that it reduces the leverage of all of your users. Like I really deeply believe in giving users and developers more leverage. Fundamentally, giving them access to more models is about giving them more leverage over all the innovations that happen in AI. You want to be able to like, access them all. You to reduce your dependency on any individual one. If you you you build on top of a router or a gateway that doesn't give you access to the full market or full flexibility or full customizability, it doesn't give you the full leverage of the whole ecosystem, then you're of like being cut out. You're cutting out all employees at your company of things that they they need. And so like, OpenRouter is fundamentally about giving people more choice because that gives them more leverage.

**Harry Stebbings** [16:30]:

You do that at a price at 5.5% take.

**Alex Atallah** [16:34]:

That that was sort of our pay-go plan. We then added an enterprise plan with like, a totally different pricing model. and it's been very successful so far. It's kind of based on like, committed spend and then, you know, no fees on that committed spend.

**Harry Stebbings** [16:50]:

Because that was going be my question. Ultimately, companies will love love it small. and then as you scale, you're like, shit, this is really freaking expensive. I'll just build my own routing tech now because it's become such a significant part of my cost base, actually.

**Alex Atallah** [17:04]:

I mean, I of figured like some of those companies, you know, just haven't realized we have like an enterprise plan and some of it is like our fault for not having like a better, I think, more detailed pricing model. We're soon to introduce like a kind of a business self self-serve plan that also just makes it make a lot more sense. And if you if you have your own inference, like you bring your own inference to OpenRouter, if you bring your own keys, that fee goes away. For inference that we are providing you, like when you go into OpenRouter's capacity and you're not on our enterprise plan, that's when that fee comes in. Otherwise, like, you you know, we need to be able to like predict demand a little bit. So that's why we do these committed spend.

**Harry Stebbings** [17:43]:

What will be the main revenue line of OpenRouter in three years time?

**Alex Atallah** [17:47]:

I think it's going depend on on the economy in so many ways. Like if the overall AI market keeps growing the way it's been growing over the next four years with like 10 to 15x x every year or potentially more, it's a lot of growth. You know, I think under that world, I would expect people to continue to underestimate how much inference they're going to need, and thus our our revenue is to be dominated by, you know, the same things that dominate it today, which is like us helping people with with an unplanned inference capacity, both enterprises and startups. That's what OpenRouter is best at, like, when you need to try models that you weren't expecting, you need to try when you're like using more inference than you thought you were going to use on particular models. Like, we make, we make sure that that is not going to be an issue for your company by providing the best failover and best uptime. And this is really, really, a good thing to do when the market is like continuously underestimating its inference needs and growing at this rate. If this growth rate continues over the next four years, it's going to be a a wild amount of growth and the economy has some limits to it. I can see like major SMB SaaS, like growing for us. You know, if growth like does not keep going going 10x, 15x per year.

**Harry Stebbings** [19:12]:

We've seen token prices fall 90% give or take in in like 18 months. Is the reduction of token prices helpful or hurtful to your business? Because obviously you have a take on spend. If they come down and spend is more efficient, seemingly it's bad for your You have a shrinking pie to to take from.

**Alex Atallah** [19:31]:

Well, a lot of people talk about the Jevons paradox, that when prices go down by 10x x, the usage increases by more than 10x. But like, no one has really done a great job modeling it. We do have a lot of spot stories that confirm it. For example, GPT 5.6 Luna on OpenRouter. OpenAI cut prices by 5x x, and then in coordination with us by another 2x. So in total, the price of Luna has dropped 10x x on OpenRouter over the last two weeks. Guess how much usage has grown? 13x. x. So it's close to perfect Jevons paradox story where you drop prices 10x x and usage grows by more than 10x, just a bit more. And the also the the usage is pretty stable. Like it, like grew it like flattened out at 13x x, And then, you know, it's been of like growing at the same rate that it was growing before it hit the the 13x x multiple. So that's pretty interesting. And it's a pretty low variable. Like there are few other confounding variables in the story. And it was also done in the middle of DeepSeek launching and having a really, really, good price and GLM having a really good price. Like now Luna is being used more than GLM on OpenRouter. GLM used to be like one of the three four models by token volume, and now Luna is past it. This is the first time OpenAI has had a model on our platform in the top three to five models by token volume in an extremely long time. So it it was a really big and interesting move.

**Harry Stebbings** [21:09]:

How reflective of the market are your token volumes? Because it's about I may get this wrong, about maybe one and a half, 2% of say, token volumes. And so how reflective are they? Because a lot of people, I say, oh, the top five models when I look at OpenRouter are all Chinese, what does that mean? They'll go, oh, well, Harry, no offense to OpenRouter, but it's not reflective of the market. And most people who use Frontier, it doesn't go through that. they use Frontier APIs, and so it's not counted. To what extent are your rankings reflective of true token usage?

**Alex Atallah** [21:41]:

Yeah, it's a really good question. You know, we we try to estimate how they're off by, you know, just surveying people sometimes or looking at like the surveys other people have done. We definitely have a bias to people who believe our thesis, which is that the future is multimodal and companies who want multiple models. And there are still companies out there. I basically rarely, very rarely run into them now, but there still companies out there that are just like, oh yeah, we're an OpenAI shop. You know, we only do OpenAI models. And so we're not to see any of those companies. And I think those companies are primarily focused on, you know, the hyperscalers, OpenAI, Anthropic, and Gemini. So we do probably like, undercount the frontier models. But I think like, over time, our thesis is becoming more and more common to see in other companies in the moment that like, they're like, oh, yeah, Like, we need to use other models. Then our data becomes more representative. And as we scale up, the data becomes more representative in general. So my hope is that like, that it just becomes better and better data over time.

**Harry Stebbings** [22:48]:

Alex Karp said on CNBC in his rather wonderfully energetic way that companies are terrified of working with frontier model providers.

**Alex Atallah** [22:55]:

Do

**Harry Stebbings** [22:55]:

Do you think they are?

**Alex Atallah** [22:57]:

I haven't seen what he talked about there when I talked to our customers, but there was definitely like a little, there was some skittishness that the, particularly when when Claude design came out around Figma. And that part I did see. And I do think that there are like real concerns. Like Figma is very different, but if if like a a startup is only building a like go to market wrapper around intelligence, like, hey, we are we're a company that of like, brings AI to this market and does so by doing the right integrations and customizing the system prompt. You're to be fine if the model labs don't care about that market, which there will be many markets like that. But the model labs have several incentives to go after you eventually. One is getting multiple teams within companies they do care about to be dependent on them. This is my theory behind why, like Claude design was strategic. While it's not like a massive amount of revenue for Anthropic, like not probably not a significant amount of revenue, it does get the design team to really care about Anthropic models. And so the companies that like they want, they now have another team that really wants to stick to Anthropic. So that team strategy can make you compete with the model labs. And so I think like companies like that, that, find themselves like, oh, we're like building a product for a team that has now become strategic for the model labs for like companies they actually care about. That's where I see probably the most near term threat.

**Harry Stebbings** [24:33]:

Do you think Claude design will have a meaningful impact on the Figma business? I speak to many founders today who are bluntly switching from Figma to Claude design, and it's cannibalizing their Figma usage. Do you think that will happen?

**Alex Atallah** [24:46]:

So I I I saw a lot of designers try out Claude Design, including our own. But so far, I haven't heard of the repeat story. I don't know. Honestly, like, I have not talked to very many designers about this. I certainly haven't heard a lot of chatter about Claude design. And if you just look at the numbers for Figma, they're quite good. Like they had a very, very incredible earnings.

**Harry Stebbings** [25:11]:

This is why you don't to be public, dude. You see like, great numbers. Figma, down. I'm like, poor Dylan. Like, give the fuck. What?

**Alex Atallah** [25:21]:

that was crazy.

**Harry Stebbings** [25:22]:

Do you what I mean? Really? Come on. We were talking about the different models that we have on offer and whether companies are willing to work with frontier models. The rate of model development feels immense. Do you think we see the same rate of model development continue over the next year, two years, three years,

**Alex Atallah** [25:41]:

Frontier model development or general model?

**Harry Stebbings** [25:44]:

General model. Both frontier and open.

**Alex Atallah** [25:46]:

Yeah.

**Harry Stebbings** [25:46]:

Just because, I mean, every single day there's there's two, three, four new models.

**Alex Atallah** [25:50]:

In July, we launched 70 models. about one model every hours. There's some agent labs starting too that are all to of like probably make models eventually. Like Jeff Dean is starting an agent lab right now from Google. The companies that are known for making agents have an incentive to create their own model, a very clear incentive to create their own models and distribute it through the agent. And we haven't even seen the start of that. sorry. We've seen the start of seen it really pick up. Like Cursor has a model. Does Lovable have a model yet? I don't think so.

**Harry Stebbings** [26:26]:

Not Not publicly.

**Alex Atallah** [26:27]:

Yeah. So the agent labs are to, I think, develop models. This pressure from both the GPU makers like Nvidia to like, create more competition in the space and create more diversity in the space, plus us, plus investors who just want to try new things that all could improve intelligence in some neurodivergent way. I think those are those are strong incentives. I think that there are enough to to incentivize more founders to make Neolabs. And if like, if if American open weight models pick up in steam, then it gives these Neolabs a base to train on that's not Chinese, which will then probably create more American Neolabs.

**Harry Stebbings** [27:13]:

Do you think we should be concerned by the rate and quality of Chinese open models?

**Alex Atallah** [27:19]:

We should. We're behind. America is very, very, behind still. I think things are picking up. I think, you know, we we have poolside, we have thinking machines. We have RC.

**Harry Stebbings** [27:31]:

Do you feel a sense of responsibility for that? And what I mean by that is like, you know, you are a routing business and you could route a company to a Chinese model that, who knows, people are worried about backdoors, CCP involvement. You could be the the deliverer of that to those models. Do Do you feel a sense of responsibility for that?

**Alex Atallah** [27:52]:

So we we do feel a responsibility to have safe access for all of these models. Like customer trust is like our paramount goal. If one of these models is unsafe to use, and generally considered unsafe, we pull it from the platform. If there's like a way to use it in an unsafe way, I mean, there's a way to use all the models in an unsafe way. then we believe in using technology to make it safe and to work with the model labs themselves to figure out how they're doing it on their side so that we can be state of the art or better. We spend an enormous amount of time making sure that our practices match what the best things that we're seeing coming out of the the labs or better because we're we're a way of like exploring all the models and finding them for the first time. We're a good focal point for deploying safety measures across your whole company. For example, we have prompt injection protection. You can just turn it on and immediately flag prompts that look like prompt injection that's trying to happen. We have PII redaction. We have we have like a couple different things that you can automatically just turn on with a click and get an added safety layer on top of all of your inference. And we build that so that enterprises feel like they can safely deploy new models and that their employees can try them out. I I I think of the models a little bit like the Internet. You You can't just like, ban the Internet at your company because there are there's some bad things on the Internet. You can create guardrails, and you should. You need to use AI to build the best possible guardrails that you can.

**Harry Stebbings** [29:29]:

I'm with you. But do you think you actually know what's going on within Moonshot or Alibaba with Kwan? These are incredibly secretive organizations in the depths of China.

**Alex Atallah** [29:40]:

Can't pretend I know what's going on inside of them. As a US company, like, we're to follow, the best practices of what happens in the US to make sure that we're not doing something irresponsible.

**Harry Stebbings** [29:50]:

What do you think US companies are more nervous of? Frontier models or Chinese models?

**Alex Atallah** [29:55]:

I think they're they're more nervous about frontier models usually, part because there's just much more confusion around the data policy, about what's actually happening to the props that I'm sending and where they're being stored and how they're being looked at. And you can't run them on your own machine or in a provider of your choice. And so that just immediately creates all of this uncertainty in a lot of enterprises. And it's uncertainty that they can also pattern match. It's very similar to like, you know, running on their own infra versus running in their VPC and knowing like, who can see the data.

**Harry Stebbings** [30:31]:

How extraordinary is that though? Like they're more nervous of like US companies headquartered in Silicon Valley where you can see and touch and feel the headquarters and the leaders. It's just like, what a strange world to be in.

**Alex Atallah** [30:44]:

Yeah, it it is very strange, especially with the frontier models having the biggest cyber posture right now. What

**Harry Stebbings** [30:51]:

What do you make of every company kind of posturing, we hacked someone? Firstly had OpenAI, then you had Anthropic, and then you had Zuck coming out. I don't to miss the party. We did too.

**Alex Atallah** [31:00]:

Yeah. Well, I think they have to talk about it. The right thing to do is to reveal when there's been a cyber incident involving your model. Covering it up doesn't work it's not gonna work in the long term. And it certainly looks like they're all bragging about it. But really, if you were in their position and something happened with one of the models and you had to make the choice about whether to publish it or not, I think the right thing to do is to publish it regardless of what how people are spin it. So I doubt that they're, you know, they're actually thinking of the felony bench or whatever it's called.

**Harry Stebbings** [31:37]:

How significant was the latest Kimi model, which got so much attention? Was it as significant as everyone thought?

**Alex Atallah** [31:44]:

It's quite good. It's It's not cyber capable in the the same way the frontier models are. And long horizon tasks, I think it's still a bit behind the frontier models. But GLM 5.2 was a really big, big step for open weight models. Kimmy was of like moonshot getting up to that step. That's a little bit how I see it. And Kimmy's also a very good writer. Like the voice and tone are both pretty good. Whereas like some of the frontier models like have like voice degradation that happens when they get better at coding, especially. And I was like, oh my God, Like, I can't read this output anymore. The output sounds like three of the four arguments you made are right. and one is a turning point. Know? And Here's the rub. Like, I it's just it sometimes just just impossible to read what what they're saying. And this stuff is fixable. But Kimmy, I think, has always had pretty interesting writing.

**Harry Stebbings** [32:47]:

In months, will the chasm between US open source and Chinese open source be bigger or smaller than it is today? my fear is that it will be bigger because when you have DeepSeek, it becomes a national champion in China. And I mean, Xi Jinping is going, this is our AI horse. I will concentrate all of my money and efforts behind this, and I will supplement this ecosystem to the end. This is the winner. And then when you see another moonshot come out, suddenly all regulation gets moved aside, all policy gets pushed aside, all funding becomes available, Everything is allowed. You are free to run. And these guys are unabridged in their ability to do whatever they want to get to the end goal. Whereas OpenAI and Anthropic and all and all the other providers in the US, especially Source. fuck. you got to try raising billions of dollars for a US source model. Bit tough, actually. Not impossible at all, but tougher. Business model questionable. AI research is super expensive and you're competing against OpenAI Anthropic. I think the comparative landscapes they sit in mean that the Chinese open source providers are just inherently advantaged, sadly.

**Alex Atallah** [33:55]:

And they have very, very, good researchers. and I think Americans underestimate that a lot. I do think they're to be concerned about the cyber posture of their models. and they do seem very concerned about like, censoring the models models and and censoring censoring the information that the models can provide to people. So while while while today people complain about American models censoring more due to cyber, I'm not sure that's always to hold. And as like the Chinese models like grow in importance for China, what are they going do? Are they going like drop the great firewall? Are they to like, give up on on putting the firewall around the models? Like, I don't know that much about China, but it does seem like kind of strange that they don't seem to care more that the models are. I've never seen anyone do a profile of like, what you can do with DeepSeek that you can't do with the Internet in China that's available to you within the border. like what information you can access. I've never seen anyone of like do a real deep dive. Like how far past the firewall does DeepSeek go? If the firewall matters to China, if it's to matter in years, like something's gonna change.

**Harry Stebbings** [35:04]:

Well, what's interesting is like, obviously the abilities of the Chinese models outside of China is immense. The abilities of the Chinese models inside China is actually relatively limited.

**Alex Atallah** [35:12]:

Oh, the the guardrails.

**Harry Stebbings** [35:13]:

The guardrails. are incredibly stringent and prohibitive. prohibitive. It's ironic that they are incredibly superior to us. Shit domestically. Terrible. Interesting. I literally just had my dear friend Jason Lampkin, who runs SaaS, to come back and be like, couldn't figure out what time Starbucks opened on DeepSeek. Wasn't wasn't on offer. Would say not allowed. Wild. Very Very basic rudimentary requests. We're speaking about all of these different models. And the thing I think is like, what about loyalty? And you have this incredible seat in the ecosystem where you can see everything. Do we see any developer loyalty today with models?

**Alex Atallah** [35:49]:

Honestly, we do see some. You know, we we try to make switching costs close to zero so that when new models come out, people can try them out really easily. But we also measure retention and churn from all the models. We share this with this data with model labs too when they ask for it, so they could know like, oh, you know, for my model that just came out, which models drove traffic to it? And for those users, Like, when they leave, which models are they leaving to? And we'll like make this more and more available to the to the world soon. And we do notice in the churn data, there are developers who of like, continuously stick to models even when there are better models out there, better models for their use cases. I think it's a combination of like a couple probably root factors. One is my app works and I don't to break it. You know, If the support bot starts saying something weird that I didn't expect, why add more headache? I've already done all this optimization and like, I've already put all these guardrails around it. Another is new models are not necessarily going to make your pricing better. In fact, in general, what happens is that the current models, like price goes down over time. And especially when new advancements in the labs happen, you'll you'll you'll see like intelligence jump, but like the price curve, like also jumps and then we'll start going down over time. So it's not necessarily the most like price effective thing to do to like, shift over to the the newest model, even for open weights. The third reason it's just fundamental, like trust in the outputs. Like I'm using a model to do my work and I like the way it talks, I probably have like some eval, like a personal eval. A lot of people have these personal evals that are just these random tests that they give the models. If the random test doesn't look really good on the new model, it'll just be like, good. I liked Kimi K 2.6 anyway.

**Harry Stebbings** [37:45]:

People thought before that memory would be the retentive mechanism. Well, OpenAI has all of my previous prompts. It knows that I live in London, I do podcasting, podcasting, and and that, and that will make it a better model for me moving forward. Is memory no longer a retentive mechanism?

**Alex Atallah** [38:03]:

Memory is really interesting. I I've always thought it like is a retentive mechanism. and the question is where it lives. Is going live with the model? Is it to live with the inference provider? Is it going to live with the app? Is it to live with the infrastructure provider, the the router? My guess is that all of those layers are going to try to own memory in different ways, and there are to be advantages to sticking your memory in each layer. You know, if you stick it with the app, then the memory has like the most app related context and is model agnostic. If you stick it with the model, the memory might perform the best on personalized benchmarks and perhaps have the best ultimate intelligence. And I think the model labs are going work on memory. And then the ultimate thing might be like, is there a good combination? Can I use memory in the model and memory at the infrastructure layer or the app layer at the same time? Like, is is that to confuse the model? We don't know yet. I I do think that like, it's impossible for one layer to capture all valuable memory because the apps own so much important context that the model labs don't have. And the the model labs, in order to get this to work, they'll have to incentivize the apps to like, give them that context. Speaking

**Harry Stebbings** [39:17]:

of Apps and the models there, Claude code, cursor, bundle, model and harness. Is the router absorbed into the agent framework before it ever has the chance to be independent when you have the agent and the harness together?

**Alex Atallah** [39:32]:

The harnesses are pretty interesting because in our early days, one of our early bets was that most apps were underestimating the desire for users to choose the model. Most apps in the very early days, like 2023 and 2024, it wasn't even clear which model was being used under the hood. They were like, people are not to care about that. They They just want AI. And one of our like, strong convictions then was that, no, people are going to want to use particular models. They're going to care about who they're talking to. It's like, I to know which employees I'm talking to when I'm trying to solve a problem, and models will be kind of like that. And that has played out, you know, like in Notion, you can like choose the model that you you talk to, even though you would think an app like that might want to obscure it completely. A similar thing happened with harnesses where, particularly with developers, they started to build an affinity to different harnesses, and that's because it's like it's a user experience. So I think that is my favorite argument for why harnesses are to stick around. Not that like they're being bundled with the models, because in fact, like, as models get better, they get more resourceful, and the the junk that gets thrown in the system prompt just becomes a handicap. Anthropic, I think, published like, a good article about this where they showed that like, oh, we got like, we got rid of stuff from the system prompt and suddenly fewer contradictions showed up later on with user prompts and the model performed better. And we and we're seeing a lot of the harnesses right now are like deleting code in order to perform better with the latest frontier models. That I don't think means that harnesses are bad. In fact, I think we'll see more harnesses come up in the future because it's a way of building a user experience on top of models. It's a way for developers who are not model labs to own a user relationship, and that is just going to be incredibly valuable for the economy to have that layer.

**Harry Stebbings** [41:28]:

I'm gonna get killed for this. What's the difference between a harness and an app? Feels like this word wank of like, everyone's talking about harnesses and harness. and I'm like, is that not an app? Hello?

**Alex Atallah** [41:40]:

The nice thing about the harnesses compared to the apps is that they're more composable. I can have a harness call another harness. I can have a harness spin up another harness in a sandbox in the cloud.

**Harry Stebbings** [41:51]:

I don't know what APIs did for apps.

**Alex Atallah** [41:53]:

Yes, but it's much more reliable and deterministic and sort of easy for users to grok with a harness because the harnesses are Unix based. They all have and and the models are so well trained on Unix on bash commands. Whereas, like, you know, if I'm like telling a harness to go orchestrate an app in the cloud, it's to be like, oh, boy, does this app like, how do you log into this this app? Is it like, do I need your password? Do I need to fire up a virtual browser? It's to be pretty slow. I'll figure it out. Okay. I I like fired up a browser and like, now I need your password and I'm gonna, like try to find the input where to put it in. And And there's probably an API in this app somewhere. I need to like look up the docs to figure it out. And okay, now I've got the API. there's so many like unknown unknowns when you're composing around an app. Very, Very, very, very, few unknown unknowns when you're composing around a harness. So I I think it just gives developers more flexibility and flexibility that they can inspect. Like API calls, you're just seeing a whole bunch of code flying around the screen. A harness. Oh, I can like jump into the harness and like, look at what's going on and talk in English about it. So it's much more user friendly.

**Harry Stebbings** [43:06]:

We've seen Meta and Muse really be a focus for Zuck. We've seen Alex Wang front and center much more. Were you impressed by what Meta delivered with Muse?

**Alex Atallah** [43:15]:

They've been doing a good Yeah. I mean, it takes a while to set up a whole new model lab from scratch. and, yeah, I'm sure a lot of, like, organizational debt to deal with. Like Do you think they will be a serious challenger? I do. I think they they they have the resources. I think there's some competitive things they can do around the model that helps people in ways that the model labs are not as interested in doing. like just having like a social network and like a focus on people. You know, it's it's like something for the brand that maybe Grok and like SpaceX AI have it too, but they do need to find their niche. Like I think people don't quite know what to do with Muse Spark yet, like when to use it or when to go for it or what it's like, like, core advantages. They just released a coding harness. They're trying to be like a generally capable model right now. I expect that in the future, they're going to be like, look, we are way better at this thing. And that's that's to be a really important moment for them.

**Harry Stebbings** [44:17]:

I Fascinating. I have to I was impressed by it, actually. Do you know what I use now? Maybe plug in one of our our mutual friends, but Anastasios and Arena. And it's so weird. I'll put my prompt in Arena. And then obviously it comes back with a load of different model options.

**Alex Atallah** [44:30]:

Yeah.

**Harry Stebbings** [44:31]:

And, you know, I come back with, I used one the other day, Pergamom? Yeah. It was like Kimmy and Pergamom, and they offer you four different options. And it takes me to models that I would never have used before. And actually muses come up a couple of times for being pretty impressive. But I love that in terms of this like discovery mechanism to models that I would never have used. I would never to a Kimi, honestly, dude. just go fucking chat ChatGPT. GPT. really It's interesting. It basically goes to the point of the model layer just becoming a utility layer.

**Alex Atallah** [44:58]:

What do you mean by that?

**Harry Stebbings** [44:59]:

Well, actually, I have no loyalty to them. I have no affiliation with brand. I go to arena and I to see what you got for me. Show me the results. I don't care if it's Kimmy or Muse or Claude or Sonnet or do you know what I mean? And actually, I just to see the options you got and I'll pick the best from there. I'd rather run four in parallel. Do you buy this whole we're gonna have one frontier model run four open models. And the frontier model might be be 160 IQ points and the open models might be be 120 IQ points, but that will be a model infrastructure or structure that we'll work with.

**Alex Atallah** [45:34]:

Totally think that that is a great architecture that everybody needs to explore. We've been helping lots of developers do this. You have subagents. agents. We have a sub agent server tool that we like tune to be really, really, good at using models generally. And then you have an an orchestrator model that calls out to the sub agents when it wants particular tasks to get done. And these sub agents are just very, very, low cost, and they're focused on deterministic tasks. This is what open weight models are generally really good at compared to frontier models. When you have a deterministic task where you know the shape of the output, you know the type of problem that you're working on, and it's a type of problem that has been solved, like classifying some text, for example, then you should definitely use like, a low cost model from OpenRouter and and then have the orchestrator model read the results and then go and continue working on the like, unknown non-deterministic task that it was set out to do.

**Harry Stebbings** [46:33]:

I to create a open American ecosystem. Yeah, More amazing open American models. And I make you head of this program. What would you do to encourage, incentivize the open US ecosystem to compete more vociferously with the Chinese?

**Alex Atallah** [46:51]:

I think I would spend time talking to the current American labs a little bit more to figure out what distilling the Chinese models looks like for them and how effective it is. You can probably get pretty far distilling the Chinese models. The nice thing about the open weight models and the Chinese models that they allow distillation they like most of them. And that means that you can like take the outputs of these models to do reinforcement learning on top of the model that you're building. This is just like a very important and common practice in AI that all labs do. And also when you distill, you see the output. So you can inspect inspect them to make sure that they're aligned. So if there's anything about the open weight models that you're worried about not being aligned with, like, the voice or constitution of the model you're creating, you have a much better shot at catching it when you're doing these RL rollouts. The other thing I would try to figure out is the compute question. Like, compute is just a huge advantage that I think we still have relative to China, and these Neolabs need a shot. And there like, there needs to be an easier way to get compute to the right talent in all countries. But especially if we're trying to create like a competitive American Neolab system. NVIDIA has been doing a good job of this, but there's Google, there's TPUs, there's Trainium from Amazon. I would work with all of the hardware companies and also the Neo chips to help with compute.

**Harry Stebbings** [48:30]:

I don't think we will have that compute advantage for long. I think you see DeepSeek and ByteDance both aggressively pursuing their own chips now. The export controls mean that they have to And this is like the number one problem for Xi Jinping in his race to win the AI war. If they build a bridge in four weeks, I think they'll manage a chip in six months.

**Alex Atallah** [48:49]:

Yeah. Like, staying ahead on the on the chip war is critical for America. Is distillation wrong? I mean, distillation is it's a technique to build models. Like

**Harry Stebbings** [49:00]:

But people like, view it with cynicism and shade. Well, and then just distilled models.

**Alex Atallah** [49:05]:

It's a technique to build models. The closed weight model labs distill models too. like like Sonnet is a partially distilled version of Opus. And like, this is how you like, make smaller models out of bigger models. It's an important way to just just just teach your model new things when when you find like something useful in the ecosystem. We do think that labs have a right to say it's not allowed in their terms of service. A company can cut off access to someone who is trying to build a competitive model. If you're just trying to build like, a smaller model that's really focused on doing one specific thing that's not competitive, most of the frontier labs don't prohibit that to my knowledge. But there are to be markets for companies that allow it and companies that don't. And we we make sure that we help, both companies uphold their terms of service.

**Harry Stebbings** [49:54]:

I have to ask you one question before we do a quick fire round. I'm to get killed if I don't ask it. There are reports that you are selling to Stripe for billion. Is that to happen?

**Alex Atallah** [50:06]:

I can't can't comment. But whatever happens, we're we're to execute on the vision. What we're doing is critical for the ecosystem, and we believe for like safe access to AI where one monopoly doesn't take over, where we have like a vibrant ecosystem of of models that everyone can explore. And And when new providers and new server tools and new inference adjacent tech comes online, there's a really easy way to discover it and connect it with all of your with all of your existing AI.

**Harry Stebbings** [50:36]:

I was thinking thinking these situations? my response would be like, well, I own 22 of the company, billion, Ooh. Now I'm a venture capitalist. But, is it hard not to think like that?

**Alex Atallah** [50:48]:

I don't really think about it. Do you know? I don't spend a lot personally. What what I think about when I what I do with like, personal capital, I I really want to help people work on problems that are not that just don't lend themselves very well to venture capital. They're sort of falling in this gray area of problems that like, people need to solve, but are really tough to fund because they don't come with a business model attached. And I think they're like very cool things to do now in in the nonprofit space because you can use AI to review way more data than you ever could before. I'm not quite ready to talk about it publicly yet, but I do to do something that like, helps researchers work on those problems and like, get grants to do it.

**Harry Stebbings** [51:35]:

One really cool example, think of this is David Fialkow, who's one of the founders of General Catalyst, who basically finds incredible stories that won't get funded for movies and funds them to shine a light on them because he thinks they're very important. So like the Dissident, which, you know, obviously told the story of Khashoggi and Khashoggi being you know, And then, you know, Icarus, which is the story of the Russian doping. And like, these were films that would not get funded had it not been for his funding because they were politically sensitive, charged. And he's like, I'm going to enable the stories of these forbidden tales.

**Alex Atallah** [52:11]:

Yeah, it's kind yeah. of like that. I love I love that stuff.

**Harry Stebbings** [52:14]:

He's great. He's fucking awesome. Anyway, are you ready for a quick fire round?

**Alex Atallah** [52:19]:

Sure.

**Harry Stebbings** [52:19]:

Okay. What is the most underrated model on OpenRouter today?

**Alex Atallah** [52:23]:

Ooh, good one. I mean, first, I like, like, poolside's models are great. I'd probably like my fire round answer. Like new American lab building interesting coding models that are small, highly effective, and they're building a lot of useful tools for accessing them. Good team.

**Harry Stebbings** [52:41]:

70% of of NeoLabs will die in the next three years. Agree or disagree?

**Alex Atallah** [52:46]:

Disagree. 70 seems very high. Of NeoLabs, there aren't that many NeoLabs. If like getting acquired by one of the model labs counts as die, I do think there'll probably be some like potential consolidation. But if you if you include the consolidation, I'd say 50.

**Harry Stebbings** [53:04]:

Do you think Dario should be less negative and more positive as a voice in AI?

**Alex Atallah** [53:09]:

I think it's important to have somebody who is very paranoid about the future and and how things are to shake up. And I personally appreciate anthropic paranoia. Obviously, there are areas where I I want other model labs to not feel like they're just being pushed pushed off the table. But I'm a big believer in in neurodiversity. and like, Anthropic is a part of the neurodiversity map that really matters. And if if no one is being extremely paranoid, then no one is offering that voice. And so I appreciate that they're doing it.

**Harry Stebbings** [53:45]:

What's the craziest thing that you see in your seat on top of everyone's usage that you don't think people talk about enough?

**Alex Atallah** [53:53]:

I mean, a lot of companies are obviously worried about cost management and freaking out about the amount of inference they're spending, and they don't know how to think about it. It's like a whole new way of like, doing business and thinking about your your OpEx. Like, the old way of thinking about how how much you give your employees, you like, give them a salary and you kind of forget about it. it. someone knows what what everyone's making, but like, it's a static number that like, gets readjusted on on a quarterly basis maybe after performance reviews. Really, your your employees all cost totally dynamic, different amounts now. I think a lot of companies are putting it on them to do routing. And I think in the future there's a good chance that it will like, get pushed downwards to the employee level. Your employees should like, figure out which tools and models to use that are best for their tasks. And then we should figure out how much you're costing like, due to the choices that you make as an employee. Your your cost as an employee is is to be a dynamic number, and it's dependent be be dependent dependent on how much that employee is like, effectively using expensive and cheap models to do their job. I advise companies to of like still do their normal management work, like have their managers kind of assess how effective and productive employees are, but also line it up with how much their employees cost, and then kind of come up with, a quadrant of celebration. Like these employees are like doing a good job and they're pretty price effective or cost effective. And then a quadrant of concern. these employees are kind maybe doing a so-so job and whoa, they are not cost effective at all. Their AI psychosis is off the charts. And then you address the quadrant of concern. So I don't think people talk about like, basically how you think of like, employee cost in the age of AI and that it's it's really it should be a dynamic number and not a a static thing that like, only a few people know about and it's gone.

**Harry Stebbings** [55:46]:

Wonderful. But can you imagine going to someone, oh, I'm sorry, you were worth $100 last month. Now you're worth I I think it would make planning

**Alex Atallah** [55:54]:

personal. They are in control of how much they cost. That's the great thing. Like all employees are in control of how much they cost and and can like influence that. Now you get to think like, okay, how good am I as an employee and how efficient am I being as well?

**Harry Stebbings** [56:10]:

Final one. When you look at the landscape today, there are so many things to be excited about. What are you singly most excited about?

**Alex Atallah** [56:17]:

Two things come to mind. One is rare disease research, which I think is one of those things that has been intelligence bottlenecked or or really just the inference bottlenecked. Like, it involves like, trying out lots of ideas and seeing if they work. The other is crowdsourcing productive urban life improvements. For example, like, imagine if someone was curious about finding every lead pipe in America or every lead pipe in the UK and like had an approach to it. But they really need like to make it mature and stress test it. Like now you can use AI to do that, and we just might solve some weird problems that everyone's just of given up on because you need like, a crazy idea to come from somewhere. Brilliant ideas are sort of evenly distributed all over the world. They can come from anywhere. And now you just give them leverage to actually work. So I'm excited about sort of very like, broad of urban urban or or rural quality of life improvements that we'll be able to make.

**Harry Stebbings** [57:21]:

Alex, dude, I've wanted to do this one for a while. I'm so glad we could do it in person as well. I was worried that we were to have to do it remote. It is so much nicer to do it in person. You've been fantastic. So thank you so much for doing it with me.

**Alex Atallah** [57:32]:

Likewise. This was great.

**Harry Stebbings** [57:35]:

But before we leave you today,

## Sponsor read

**Harry Stebbings** [57:37]:

founders face a different set of challenges at every stage of growth. For Sid Sheit, co-founder and CEO of dmatrix, JP delivered the guidance and expertise to help navigate what came next. He credits Morgan's high touch approach with supporting dmatrix as it grew and expanded internationally. Whether you're in the early days or expanding into new markets, Morgan helps startups navigate complexity with real confidence, offering personalized guidance and deep sector expertise. Find out how JPMorgan helps founders at jpmorgan.com forward slash grow without limits. Morgan is the bank of the innovation economy. While JPMorgan supports growth, Corgi protects it. My word, what an arresting first line. Get your ass covered with Corgi insurance, and I'll tell you why. If you're running a business right now, you already know this pain all too well. Getting insurance, it's really slow, it's confusing, and my word, it's full of paperwork. Well, that's exactly why Corgi is here to change the game. Corgi is the first and only insurance carrier designed specifically for tech companies, allowing you to get covered in minutes instead of days. Corgi provides essential coverages for all growth stages such as DNO, E&O and E&O liability, cyber, commercial, general liability, and more. Get your ass covered. I love the way we say ass with Corgi Insurance alongside thousands of other startups at corgi.com/20vc today. That's corgi.com com forward slash two zero 20VC. You You won't regret it. While Corgi covers risk, Flex gives you room to move. Business owners run their whole financial life on Flex. One platform from business revenue to their personal spend. Float every purchase for days. Tap capital that grows with your revenue and pay vendors in in 170 countries across 32 currencies, plus the whole back office, bills, expenses, accounting, all in one place. So you spend less time reconciling and more time growing. That's why thousands of owners use Flex, named one of Fast Company's most innovative companies of 2026. Visit flex.one, that's dot, and use the code 20VC.
