Ship Happens

Building Search for AI, Not Humans | Saahil Jain, CTO of You.com

Episode Summary

Search has always been built for people—but what happens when the user is an AI agent? In this episode of Ship Happens, host Per Krogslund sits down with Saahil Jain, CTO of You.com, to explore why the next generation of search needs to be designed for machines, not humans.

Episode Notes

As AI agents become increasingly capable of researching, reasoning, and completing complex tasks, traditional search engines aren't equipped to meet their needs.

Saahil shares how You.com evolved from a consumer search engine into enterprise AI infrastructure powering more than one billion monthly queries. Together, they discuss why search built around human behavior breaks down for AI agents, how enterprises are moving beyond AI experimentation into operational AI, and what it takes to build reliable systems that balance speed, accuracy, and cost.

From evaluation frameworks and multi-model architectures to continuous monitoring and the future of enterprise AI, this conversation offers an inside look at the infrastructure powering the next generation of intelligent software.

 

 

What You'll Learn

 

 

Episode Timestamps

00:00 The Agentic Search Boom

00:43 Meet You.com’s CTO

01:39 Building Search for AI Agents

02:45 From Consumer Search to Enterprise AI

03:59 Why AI Agents Can’t Just Use Google

06:55 Rethinking Search for Enterprise AI

07:39 Giving Agents Control Over Search

09:05 Where Companies Are on the AI Adoption Curve

11:42 AI-Native vs. Traditional Companies

13:11 The Real Metrics Behind AI Operations

14:46 Why AI Evals Need to Go Beyond Benchmarks

17:13 Building for a Multi-Model World

19:42 Shipping AI: Gates, Drift, and Monitoring

21:17 Rethinking Org Design for the AI Era

22:53 What Actually Wins in AI Products

25:33 Avoiding AI’s Biggest Engineering Trap

27:58 Where the Next AI Moats Will Come From

30:55 Closing Thoughts

 

Key Takeaways

 

Notable Quotes

"The future of search isn't just helping people find answers—it's helping AI agents get the right ones."

"Shipping AI isn't the finish line. Operating it reliably is where the real work begins."

"As models evolve, your evaluation strategy has to evolve with them."
 

Links & Resources

You.com
Docker.com

 

Episode Transcription


 

Saahil Jain: [00:00:00] We predict that in the coming year or two, the amount of agentic searches will continue to exponentially increase. I think there's already a statistic from Cloudflare that shows that more agents request HTML pages than humans, and we expect this to only further increase over time. Google is the search engine for humans.

Why is it not the case that Google is the search engine for agents? AI natives and companies that in some ways have less to lose can move a lot faster than companies that have a lot of obligations, like banks, where they need to abide by regulation, they're more conservative

Per Krogslund: Welcome to another episode of Ship Happens. My name's Per, and in this, this podcast, I get to sit down with really smart people and talk about what's interesting and exciting in the world of AI, but also what's challenging and what's really happening in the space of customers and companies and so on. So today, I have [00:01:00] the absolute pleasure to have Sahil Jain joining me.

He is the CTO of You.com. We're gonna talk about really going from AI adoption to operati- operation, and we're also gonna talk about what You.com does in a second. I'll let Sahil do the talking about that so we get the full understanding of what you're doing. First of all, Sahil, welcome to the show. It's lovely to have you here.

And I think for people in the industry, You.com has been known as this AI search engine. That's what most people's perception of You.com is. What are they missing? What else do you do? It's not the whole story. 

Saahil Jain: Yeah. First of all, thank you for having me on. Very excited to, to chat about various topics related to AI and search.

And in terms of You.com, if we just zoom out for a second, what we're focused on is building the AI search infrastructure for agents. And y- I think you're right that we've been known as many different things over time, but our current focus really is building out this AI search infrastructure that can help power the next generation of AI agents.

And there's many reasons which we view this [00:02:00] and why we focus on this. We actually started in 2020, and when we started in 2020, we launched as the consumer search engine, so we were a search engine for humans because agents didn't yet exist in their current form. And then what we realized over time, as we iterated on our product, is that the most important aspect of what we built, our search engine, could be used to power agents, and that actually is an even more powerful user.

We predict that in the coming year or two, the amount of agentic searches will continue to exponentially increase. I think there's already a statistic from Cloudflare that shows that more agents request HTML pages than humans, and we expect this to only further increase over time. So we're building for that future essentially, of how do we build search that can make sure that agents are well-grounded in the external world.

Per Krogslund: So does that also mean that you went from being, you can say, a consumer-facing product where you had all the, like you say, a classic like B2C business, but now you, I guess you're engaging more with enterprises. So how is the difference in dynamic here in the, the way you're doing this? 

Saahil Jain: Exactly, [00:03:00] yeah. So we did make that transition.

We started as more of a consumer-facing product, and then we ended up transitioning towards an enterprise-focused business. And right now, we essentially power over one billion queries a month for enterprises. And in order to do that, we have to follow certain types of governance, regulation, et cetera, that is really important for enterprises that is a bit different when it comes to consumer.

So we've definitely had to change our identity, the way we approach technology as we've evolved our company. And I think what we realized is, by building search for AI, we're able to really align our incentives properly with the customers that we support as opposed to... And with the-- we also were trying to do that with consumers, but obviously the consumer landscape when it comes to search is a very different beast than this, this new world in which agents are much more willing to try different types of search engines, use different types of search APIs, and we think it's a much more competitive market than consumer Search was.

But yeah, so we've had to evolve and we can talk a little bit more about that as well if you have any questions. 

Per Krogslund: [00:04:00] I have a, I have a follow-up question, which I'm guessing a lot of people have, like why... If you can get an agent to do anything, why are you just not pointing your agent at google.com instead of using you.com?

Saahil Jain: Yeah, that's a great question. 

Per Krogslund: Search has been a, search has been a commodity for the last two decades, and now we're entering this era where we're having specialized services do it. 

Saahil Jain: Yeah, so there's a couple of reasons, and this is something that comes up a lot is Google is the search engine for humans.

Why is it not the case that Google is the search engine for agents? And there's a couple reasons why, one of which is Google's search engine is highly optimized for the consumer experience, and right now they even bundle it very closely with Gemini. It's their ch- their, their crown asset. It's one of the things that allows Google to distinguish itself.

Their search index is this proprietary piece of technology that they've built and they've created a moat around. And for them to expose that directly actually is at odds with their core business. So what this essentially means is other businesses, when they wanna use Google, are forced to essentially use what we- what's called SERP APIs, [00:05:00] which are APIs that essentially crawl Google.

And there's a couple issues with this approach, one of which is reliability. So when you're trying to use Google in this way, reliability can fluctuate because it's a cat and mouse game where Google is trying to block some of these scrapers, they're trying to scrape. There was actually an interesting court case that came out regarding SERP API versus Google, in which SERP API actually seemed to win the court case in terms of scraping.

But regardless, there's a lot of reliability questions over there, and the second one is zero data retention. So when you're using a SERP API, your queries are going to Google, and for many enterprises, they need to make sure that their queries and their customer data never leaves the cust- the vendor that they're using.

And if you're using Google, you can't make that guarantee. And then the third piece is we're actually thinking about building search differently for agents. So humans have very specific quirks compared to agents that you can optimize search differently for agents. One example of this is how we process information.

We tend to look very closely at the first couple results, and [00:06:00] anything below the fifth result or on the second page of Google, you just don't look at. Agents, on the other hand, can process large amounts of information more efficiently, and with token prices decreasing, it's actually a very different form of ingesting information than humans.

So we can rebuild search from the ground up for this new customer or this new user. So those are a couple reasons why we think that there's gonna be this new ecosystem of search APIs and this renaissance in search for agents compared to in the consumer place where I think humans became very used to using Google.

And Google is a great product, by the way, and they've built it. It's an amazing product that works phenomenally well for humans. 

Per Krogslund: Yeah. And you bring up a really interesting point here in the sense that you've- By first being a consumer business for humans and now going into this enterprise business focusing on agents, you've tried to have both of these customers and seen the difference in behavior.

So what you're describing here that, yeah, agents doesn't really care about the fold, whereas humans very quickly get impatient. 

Saahil Jain: Yeah, so tool use. So I think what we think about is search... What we're building essentially is a [00:07:00] tool for AI, and there's gonna be other tools as well, and the hi- the higher level term that we use to describe this is what we call augmented intelligence.

We believe there's two types of horizontal AI companies. There's what we call core intelligence companies. These are the companies building the foundation models like OpenAI, Anthropic, Gemini, and there's companies that are building what we call the tools and services that these foundation models depend on.

So you can think of these as, like, the roads that these self-driving cars are using to go further. So we're building the roads essentially that allow these AI models to do more, and search is that core tool or road that we're focused on. 

Per Krogslund: So in, in just... Y- you say you're gonna launch more things, which is very interesting, uh, but also in, if...

A question I had around the search is, are you seeing a need for adding additional parameters that a traditional, like, human wouldn't care about? Where you say, "Oh, this actually really matters to an agent to, to increase the performance of doing their job with search." Are there things that machines care about that people generally don't?

I know people generally don't care about search in settings, but are there things, like, you don't even see as a setting on Google, for instance? 

Saahil Jain: Yeah, [00:08:00] absolutely, and I'll give you one example of this. So w- in our research APIs, one very basic parameter we expose is something called research effort, and what this essentially means is it allows the agent to pick the amount of effort that needs to be put into a query, and that basically changes the cost, accuracy, latency, dynamics.

And this is something that a human necessarily wouldn't do because when you're using search, you're being monetized in different ways, via advertising, for example. But when agents are using search, they're actually in some ways becoming more and more capable of directing their own budgets. So we actually View this in some ways as, you know, and this is one of the innovations that we worked on, was providing budgets to our agents, which other agents use.

You can also view some of our tools as sub-agents, so search sub-agents, like a research sub-agent. And in that case, the main agent can s- configure the amount of effort that the sub-agent should spend, and that's a very interesting phenomenon that's happening, and I think the better the agents get, the more and more knobs that we're gonna continue to expose to these [00:09:00] agents so that they can really optimize what they're getting out of search and research.

Per Krogslund: Cool. That's super interesting, but I also wanna, wanna pivot a little bit here on, on topic because I also wanna talk to you about AI adoption, and also AI operations really, which is a little bit further down the maturity curve for most companies. So where does you.com f- typically fit into, you could say, an organization's adoption of AI?

So we've seen the past couple years, like saying, everyone is saying they're doing AI, which often is like handing out a bunch of Claude and OpenAI licenses and then see what happens. Um, but I guess companies who come to you and say, "We need a search agent AI for our agents," they're a little bit further down the adoption curve.

So I'm assuming here that you are probably working with companies who are a little bit more mature in their approach to AI than what you've seen with just have some AI licenses and go to town kind of approach. 

Saahil Jain: Yeah, that's a really good question. Yeah, so we generally work, we work with a range of companies, ranging from your traditional companies that are now innovating really fast to your AI natives.[00:10:00]

So we have a spectrum of companies across industries, and on, on one side, on the AI native side, over there what you're seeing is these companies are building out AI experiences in which they're using a lot of these reasoning models, and one of the core limitations they face is that these reasoning models don't have access to the external world, and if they did, they would be able to achieve and do so much more for their customers, and that's where we come in.

We essentially play the role of this tool that allows them to do more, and in that sense, we are essentially part of the critical tool stack for companies that are hitting a certain point and they wanna go further. One example of a company that we work with is Factory. So Factory AI is building out these droids that allow developers to do more, and they're model agnostic, so they're using a lot of different models behind the scenes, and because of that, they need an independent search provider that can power a lot of those models, and that's where we partner.

So that's an example of an AI native company. And then on the other side, we partner with more traditional companies like some banks, retail, et cetera, [00:11:00] where they're actually now starting to build their own agents and they need search APIs as well. What we find is that when people are first developing, they'll often just use the generic off-the-shelf implementations from OpenAI, Anthropic, but then once they hit a certain point in which they need to further improve performance and reduce cost, they then start- Really examining each particular choice and start optimizing it, and that's usually where we come in.

And we also help provide evaluations where we show that by using our search API with some of those models, whether it's a frontier model like Anthropic or OpenAI, we can improve performance. And we partner very deeply with the foundation labs as well as the open source community to make sure that our tools work really well across models.

Per Krogslund: Yeah. When you both have a chance to, you can say, interact with AI-native companies and these more traditional companies, how do you see the difference in their approach to adopting and, you could say, operationalizing AI? Do you see changes in behavior of, you can say, these more like legacy companies, if we can call them that, versus AI-native companies?

What's the biggest difference? 

Saahil Jain: Yeah. There's a couple differences. I would say [00:12:00] AI-native companies obviously can move a lot faster, so they're often using many more models. They have, they have, they're probably using every single one of the frontier models, plus a variety of the open source models, and they're very quick to add new sub-processors.

On the other hand, when we talk to the traditional companies, they've often, they're often very tied to a core infrastructure provider, whether it's Google with Gemini or AWS or what- whatever they have, and adding new sub-processors takes a lot of time. So you really have to work and build faith and build trust in your product over the course of many months, if not sometimes even a year, with a lot of these traditional companies.

So that, that's one of the main differences in terms of velocity, AI natives and companies that In some ways have less to lose, can move a lot faster than companies that have a lot of obligations, like banks, where they need to abide by regulation, they're more conservative. But what we're seeing regardless is a lot of these more traditional companies are starting to innovate very fast, so they're building their own technology.

Almost every one of these [00:13:00] traditional companies now are not just using off-the-shelf solutions, they're building their own agents, and as they're building their own agents, we then work with them to see if we can embed some of our tools with their agents. 

Per Krogslund: Yeah. And so you're also seeing this maturity go up over the last six months, and I know even a lot of the, like you say, legacy of traditional companies are really talking about operationalizing, right?

It's not about adoption, it's not about everyone x-x AI now. It's not a thing that you need to push adoption, right? But now it's about operationalizing, and how do you see that, like, happening and out... When companies talk about operationalizing AI, how do you see that play out? What do they care about moving from adoption to operationalizing?

Saahil Jain: Yeah. So I would say when companies are operationalizing AI, there's two themes that become more important. One is cost, and the other one is accuracy/predictability. So when you're in the prototyping staging, stages, and this is what we generally saw with the AI industry, nobody was really optimizing cost at all.

People were okay with just spending a ton. What we've now seen in the last couple of months [00:14:00] is there's been a huge shift in the narrative from token maxing to being efficient in terms of your usage. So this is where we're seeing costs become a much bigger part of the conversation when it comes to AI adoption.

People are now, and this is also why you see a lot of routing companies that have had tremendous growth, a lot of inference companies as well, that are making a lot of money and growing extremely fast because of this desire to reduce cost. If you can use a cheaper model for some workloads, people are now doing that, whereas they weren't before.

And the second one is accuracy. When you start productionizing, it's not enough to just have okay accuracy on some benchmarks or some what we call hero queries. You actually need your system to work predictably on a real-world distribution, and that requires a very different level of maturity compared to the prototyping stage.

Per Krogslund: Let's go into that a little bit, how you do that, like, around, you say benchmarking, evals, and so on. How do you see it from your side? I guess you have experience in this area as well. How do you approach this, like ensuring that you have this accuracy that you wanna offer your [00:15:00] customers? 

Saahil Jain: Yeah, it's a great question.

So we actually, we think very deeply about AI evals. We even have a team, an AI evals team, whose main charter is to innovate and continue to benchmark our various tools and offerings, and really understand how we can continue to support and improve our customers. In, in terms of accuracy, we also do some research on evaluation, and there's-- it-- we're in a really interesting state right now when it comes to eval, and this is, uh, I'll give you an example of this.

So we actually worked on some research which was published and received the best paper award at the AAAI conference. And in that research, what we looked at was, uh, understanding how, basically understanding the consistency of how agents perform in some of these benchmarks. It's become common practice in benchmarks to basically pr-provide just one number over a benchmark.

So there's a hundred questions on a benchmark or a thousand. You provide some number, and then you report your number versus all the other numbers in the industry. And oftentimes, you know, because people are moving fast, you're not really providing things like statistical significance and [00:16:00] all these, you know, statistical tools that are more common in peer-reviewed science.

And that's okay because at the end of the day, this is not a peer publishing benchmark. This is not a peer-reviewed article, and most people who are viewing benchmarks don't v-don't view them that way anyways. They just view them as more directional. Are you in the consideration set? That type of stuff. But I think it's still interesting to think about how can we do better, and it does, it, it is complicated.

So you know what, one thing we found, for example, is it's not enough to just look at even something like variance. It's important to understand why you have variance on a benchmark. So even if you were to support, variance can come in two ways. It can come from what we call agent inconsistency or task difficulty.

So what that means is if you have different tasks on your benchmark, some of which are more difficult than others, you expect there to be some variance, which is expected. But then there's also variance that comes from agent inconsistency, which is the agent on the same question doing different, having very different results.

So in general, you want your agents to have low agent inconsistency, and this is something you can measure with tools like intracorrelation coefficients, and we [00:17:00] propose using these type of statistical measures to better evaluate agents. So there's a lot of interesting research that can be done in terms of really understanding accuracy, like one level deeper than just that initial benchmark.

And that's what we're also thinking very deeply about. 

Per Krogslund: So for someone who want, maybe wants to make that migration from being, okay, they've been using Claude or OpenAI for everything. There's just been... Yeah, they've basically been token maxing, but now they wanna move into the space where they start taking ownership of their usage and maybe say, "Okay, we wanna have maybe a spectrum of models, and we wanna try to find the right use cases."

How would they approach that? Like, how would... Is it just start doing things and then measuring, getting different evals of how the model would respond to a given task? Or like, how would you approach it? Maybe you have a little bit more concrete and practical experience in this space than most companies.

So just curious how a company who are basically starting from one model for everything and maybe wanna move into this more diversified sh- would approach that. 

Saahil Jain: Yeah. I think the first piece is to just understand what the bottleneck is, and the reason why. So is it cost? Is it accuracy? Is it a growth area?

If [00:18:00] using one model for some task is working fine, and it's not a cost center, and it's not necessarily something that's a bottleneck, that's okay. You don't always need to optimize every system. If something's working okay, that's fine. But let's say the system is now important, it's growing, cost is relevant, you wanna improve margins, then you need to start optimizing your system, and maybe going multi-model, multi-tool is the best approach.

And when you get to that point, a prerequisite usually is to have some set of benchmarks and eval sets. This is often the main blocker. Yeah, I think evaluation is one of those almost trite terms now that everybody's talking about. But what we see is that most businesses still are Struggling to properly evaluate their agents, and I think that's something where there's need for hu- I, I think there's need for lots of education in, across the industry.

And over there, and this is why we have an AI evals team that often helps companies evaluate their agents and understand should they split apart their one model into multiple models and then use our tool across those [00:19:00] models. Those are the types of nuanced questions that we get into with customers. But it usually starts, step one is just having some eval set, so you can understand what the quality and cost implications would be if you were to switch one model to many.

And when you get to that point, then it becomes a question of understanding the query distribution, understanding what types of models may be good fits, and understanding also kind of the expectations of the user. Are they very sensitive? Is it a complicated task? That's the other piece, is what is the actual difficulty of the task?

If it's a very simple one, you don't need to be using Fable Vive. You could get away with using a much cheaper distilled open source model. So a lot of it comes down to just understanding the actual query distribution and having your evals simulate that. 

Per Krogslund: Yeah. So would you say, maybe from a sense of met- mental model here of how people or companies would approach these metrics, would you say this is more a way of thinking about your measurement like uptime, where you have a consistently graph over time of saying, "This is our overall performance here"?

Or is it more like a gate, like saying, "You do this. Okay, we [00:20:00] deploy this thing because we know the evals are good for whatever task we give it here"? Or is it really this is a thing we need to care about over time because things can be sliding because either the, maybe the model updates, so they change the model, or the dataset changes, or just the context gets bigger or whatever.

So it's more like how companies think about this going forward. This is, is this a gate, or is this like a progress over time? 

Saahil Jain: Yeah, that's, it's a good question. I would say it's ideally both. So you need some gate in order to feel comfortable introducing a change, and over there, there's a ton of practices from traditional software that we've mastered over the last 20 years regarding AB testing, rollouts, that type of stuff, which are very useful.

So that's kind of the gating stage. You have some offline set. You do well on it. You then roll it out to your population. And over there you can use common statistical measures, AB tests, all that type of stuff to make sure you roll it out. But that's not enough because your users' behaviors are changing.

The models may change behind the scenes if you're using proprietary model, so you then also need to be measuring how you're doing on your new query distribution. So one [00:21:00] approach here is if you have queries that are, you can use for analysis, so users opted in or given permission, you can then use those queries to basically, you can sample from them and then create new benchmarks and see how you're doing on your new query distribution.

So I would say you need to gate, and then you also need to actively measure to make sure that you're not drifting- 

Per Krogslund: Yeah So you feel like, so again, from your company's point of view, in a sense that we've had, I think we've had, like, 10 years of reasonable stability in the IT industry, like with cloud, Kubernetes 1, all that.

Like, it was a very, it was a very predictable method of how you built products. Like, the products were different, but, like, it was very predictable how the SDLC ran across a company, and so also, like, how the roles were divided between the different divisions. It was pretty... It's been pretty static for 10 years.

But now it's all changing, and so how do you see it from your... Have you organized yourself in a different way than you would, say, like a normal company? I know you have an evals team, which is clearly different, but when you think about this, maybe also the AI-native companies, are they organized in a, like, fundamentally different way than a normal engineering org would be?

Yeah, 

Saahil Jain: it's a good question. I would say right now, [00:22:00] I don't know if I've seen any systematic changes in how organizations are set up, but what I'll say is some trends are, I think the sizes of the orgs are different. So I think what you'll see is engineers owning a lot more, and teams becoming smaller. So I think it's becoming clear that companies can do more with fewer people because of a lot of the leverage that AI gives to people.

That's one change that I think is quite systematic across AI-native organizations as well as traditional organizations. Small, nimble teams can often outperform larger, more complicated teams because of the fact that each engineer can go end-to-end much faster. We have some horizontal teams, like security, for example.

Security is important across all the products, so that's a key horizontal. And then platform also is a horizontal. When you talk about stability, reliability, uptime across service areas, DevOps processes, CI/CD, those are horizontal pieces. 

Per Krogslund: What is the key differentiator, then, if people are really succeeding with AI?

Is it leadership? Is it they [00:23:00] have a vision or an idea? What is the difference between people who are really failing here and people who are really succeeding? And maybe it's too early to say. We are very early in the adoption space, but, like, where do you see, like, where do you see the main difference in what is failing and what is really succeeding?

Saahil Jain: Yeah, this is a, this is a, yeah, this is a really good question. The main a- the, there's many different ways to answer this, and I think many ways in which people are succeeding. But one, one general note in this AI world is it's really important to be intentional about the product you're building, and understand how it is going to play into a world in which models become more and more intelligent.

So a lot of companies that built a lot of technology that in some ways relies on the core deficiencies of current models, you know, were heavily displaced as the models became better and their va- their value proposition became less clear. An example of this is there was a range of companies that existed that basically prompt engineered these early terrible models and were able to build experiences by building a lot of scaffolding on top of these.

But then as the models got a lot better, the need for that [00:24:00] scaffolding disappeared, and with that, the value proposition of those companies also disappeared. So the key to me is really thinking strategically and ensuring that your product or your company will benefit as the models become a lot better.

If your company will be threatened by the next iteration of a frontier model, then it's a very uncomfortable position to be, because those models have continued to become better at an exponential rate. And I think we should assume that these models will continue to become more intelligent. And in many cases, intelligence is no longer the bottle- the bottleneck on many applications.

Um, and you have to, you have to understand which category you lie in and not get a bit confused, because I think I've seen a lot of companies who are not foundation labs try to build technology that Will eventually be solved by a foundation lab and iteration of the model, and that, that is a little bit of a hard truth.

But I think that it's better to accept that and then build products that, around that ecosystem and assume and make those models go even further. So that's why I described in my head, there's this idea of this... The way I view it actually is [00:25:00] inverse, the inverse of what we saw with self-driving. We actually had roads long before we had these self-driving cars, and the self-driving cars have been now operating on these roads, and you see them in Silicon Valley and Phoenix and whatnot.

In the case of AI, we had these super, these very intelligent models come around before there were many roads that allowed them to connect, and one of those roads is search, for example. And that's just one example, but that's, that's a vertical focus. But then there's also complete-- there's other focuses, too, like new product surface areas that I can't even imagine right now that are coming about, and new form factors and ways to interact with the world.

Per Krogslund: So looking a little bit back on, you can say, adoption among these companies that you interact with, in this whole phase of really adopting and maybe operationalizing AI, like, where do you see companies typically waste the most time? Like, what is the has- hardest hurdle to get over? And then maybe an extension of that, like, what kind of problems naturally just disappear after six months of where you have this, like, growing pains, and then what typically stops being a problem after six months?

Saahil Jain: Yeah, I would say, so in [00:26:00] terms of wasting time, I can just speak from our own experience, and I wouldn't call this wasting time, I would call it learning and evolving with the industry. Um, but I'll give you an example. So we spent a lot of time working on very complicated scaffolding around models and complicated orchestration.

So our advanced research and insights agent, when we launched it, was actually better than OpenAI's deep research on, on various benchmarks. We would calculate win rates, this type of stuff, so we launched something quite early that was very performant, and we put a lot of effort into coming up with this really interesting orchestration system that was able to paralyze a lot of the tasks, had a lot of ethics, plus model usage in, in very interesting ways.

And then what we realized is over time, as the models got better, we were able to improve the system by removing more and more components from it. And what that basically meant is we saw that these models are becoming much better at long-horizon tasks, and then focusing on the right tools with those long-horizon models or those models that are capable of doing long-horizon tasks, was a much better use of our energy than trying to [00:27:00] constantly use manual heuristics to account for some of the existing gaps.

Oftentimes, companies need to do that because you are deploying something to production today, not tomorrow, so you need to come up with it. But what I would say is I think it's probably worth putting Minimal effort into some of those pieces that we think are going to essentially be solved in six months by model improvements themselves.

So I think that's kind of one area in which I think at least I've personally probably spent too much time, and that didn't really stand up to the test of time. So I think it's important to think about, okay, what are we working on that is going to be useful six months, a year from now? What are we working on that's probably going to be less valuable six months, a year from now?

You still may need to do it right now, and that's okay, but it helps frame the investments there. So maybe it's okay to have a little bit more tech debt in something that you're working on if you think it's not going to exist in six months. Whereas the piece that you're working on that's more foundational, you want to focus and make sure that is much more stable and sustainable and foundational.

Per Krogslund: Amazing. I'll just wrap up here a little bit by asking you [00:28:00] to look a little bit into the future. The conversation has been a lot about models and optimizing for the right model. We have model routing as a very new concept that a lot of people are adopting now. We are probably going to assume that the model is going to become a commodity in this sense in a lot of ways.

Of course, there's going to be model innovation going forward. There's going to continue being capabilities, but anyone can get access to a lot of different models. It's not something that is blocked, right? So whenever everyone has access to the same technology, like, where do you think the proprietary value lives in five years?

What is it that really sets things apart in the space? Is it the model? So is it the access to data, or is it really how well you do evaluations? What do you think it is? And five years is a long time in this time, I know, so. 

Saahil Jain: Yeah, five years, I have no... Who knows what's going to happen in five years. I think it's, I think it's very hard to predict the future right now.

But what I will say is it depends on who you are. So if you're a foundation lab, I think the core IP over there is the actual model itself, um, in which case you want that model to be performing well in benchmarking. It's [00:29:00] versatile, self-improving. I think that's another thing that's becoming more important, is this idea of recursive self-improvement of models.

So that's if you're a model company. If you're not a model company- It depends a bit on what type of company you are. There's many companies that I think are much more partnership focused. So these are companies that have amassed lots of partnerships that are sticky over time. I view maybe Amazon even in this category, where they have all these partnerships with buyers and sellers, and warehouses, and distribution.

That's, that I think is still gonna be quite important even five years from now, as it is today. But then there's other kind of companies like software companies or pure software companies where I do think a lot of the advantage will probably be the data. Most of the data is actually not public in the world.

There's a lot of data that's private, and I think that many companies that are able to properly collect, enrich, and use their private data will be in a very good state five years from now, and those foundation models will be able to properly leverage and [00:30:00] accelerate those companies with their unique private data as a moat.

So that's another, I guess, prediction, is a lot of the companies that have unique data assets will continue to be relevant and important. And then another piece is, at the end of the day, it's, uh, it just comes down to making customers happy, whether it's consumers or whether it's other businesses. And I think if you're able to produce businesses that provide value to people, and customers love the product, love working with you, that's very useful, and that, that will always be valued in the market.

And I think that what we may see is more companies, but smaller companies. So I think we'll see the rise of many more small companies run by a couple people that are, maybe look more like service businesses, but it's okay because they don't need to have crazy amounts of revenue on like this venture scale.

So I think we'll see more and more of that with AI. We'll see a lot of companies that are actually, they look like service-based businesses, but their margins may not be as good, but that's okay because they're much smaller and potentially bootstrapped. 

Per Krogslund: And I think that was my last question for this time.

Thank you so much for your time here, and thank [00:31:00] you for answering all my questions and delivering your insights. And, and I need to do an outro, which is, this is the podcast Ship Happens, and it's sponsored by Docker. And again, thank you so much for your time, Sahil. It was a pleasure talking to you, and see everyone next time.