Episode Transcript
Available transcripts are automatically generated. Complete accuracy is not guaranteed.
Speaker 1 (00:16):
Welcome to tech Stuff. I'm os Voloscian. There are many
ways to think about AI. The one that recently struck
me quite a lot comes from Demis Hassebus, founder of
Google's Deep Mind, who said, if you stop to think
about it, we've essentially found a way to make sand think.
He was talking about silicon, the element that is melted
and sliced into the chips that power AI, and today
(00:38):
I want to talk about those chips. Our guest today
is Steve Vassalo, a venture capitalist who was the first
backer of Cerebras, a chip company that went public earlier
this year in the largest ever IPO for a semiconductor company. Steve,
Welcome to tech Stuff.
Speaker 2 (00:54):
Thanks Arles. It's great to be here, great.
Speaker 1 (00:56):
To see I'm curious what's your reaction to that, To
that framing of demis making sand think.
Speaker 3 (01:02):
I mean, it's such a great reminder of what it
is that we're doing and how fundamental it is, both
from a physics perspective and all the hard work that
is required to take a piece of sand and turn
it into ultimately intelligence. And then there's also something that
feels sort of I don't know, almost I would say
sort of you know, just sort of scary and fun
in a humanity sense, like, Wow, what are we building
(01:24):
right now? What are we doing to be able to
sort of turn this, you know, very basic material that
surrounds us everywhere into something that is super intelligent.
Speaker 1 (01:35):
I want to ask you all about the story of
Cerebrest and and also frankly, what wayfer scale means. But
before we get there, can you just zoom out a
little bit and give us the wider context of what's
going on with chips right now? Because I think we
will see the headlines. I mean, I'm going to read
a few, but the New York Times recently had it's
not just in Nvidia. The AI boom has ignited Asia's
(01:56):
chip companies. Dull Street Journal had a couple of months
ago inside Apple's push to build an all American Chip
axios in July, Chip craze for the layman, what's going
on here?
Speaker 2 (02:07):
Yeah? I love that question.
Speaker 3 (02:08):
So the way I think about it is really sort
of starts with what I'll call workloads, And the workloads are,
you know, what is the work of computing, And if
you go back to the nineteen seventies and eighties, the
workloads at that time were very serialized. They were kind
of diverse if you think about the invention of things
like the spreadsheet, or you know, as a kid, I
(02:30):
wrote my college essays on word Star and an eighty
eighty six IBM computer.
Speaker 1 (02:35):
Diverse in a sense that different computers for different applications,
or diverse in what sense.
Speaker 3 (02:40):
Diverse in that you had one computing platform, which turned
out to be Intel's X eighty six platform that turned
out to be a great general purpose computing platform for
a wide variety of applications. And then sometime in the
kind of mid nineteen eighties, we began to see a
new workload emerge. And that workload really came from two
areas as one was from gaming and the rendering of
(03:02):
pixels and a need for paralyzation of the compute. And
then also in believe it or not, the CAD tools,
so you know, building new products, whether its futures exactly.
Speaker 2 (03:13):
So computer aided.
Speaker 3 (03:14):
Design pushed the perimeter of what was possible on traditional
computing CPU general purpose systems. So that was really the
dawn of the GPU, and really that was Nvidia. And
by the way, at the time of Nvidia's founding, there
were like we're seeing today thirty other companies working on
graphics processing units. And I say graphics processing units. We
many people, you know, see the term GPU. They don't
(03:36):
know or remember that, like graphics is the g It
was purpose built for that workload. And then the next
real shift was a few decades later, but was really
a constraint a workload driven by mobile devices, and the
constraints there are totally different, right. This is form factor,
this is power consumption, battery life of course related to that,
(03:57):
and so you've got a totally different workload. And then
what we began to see this is now zooming forward
to kind of the twenty fifteen timeframe, really starting in
the twenty twelve timeframes, but the workloads began to really
build in the mid twenty tens. And what you saw was,
oh wait a minute, statistical inference over large data sets.
Our portfolio companies were beginning to hire data science teams
(04:18):
business models that weren't possible before you had sort of
a real awareness of kind of the data underlying your This.
Speaker 1 (04:26):
Was the era when phrases like big data es out
computing mark andresen saying software eats the world. Is this
the area referring to sort.
Speaker 3 (04:34):
Of about five years later than software into the world,
but kind of roughly contemporaneous with that. But really what
you're seeing big data is absolutely right and in a
sense that, wow, there are signals inside of our businesses
that if we're more attuned to we could make better decisions.
And what was interesting at that time, and this kind
of ties back to the sort of the GPU wave
(04:55):
was those kinds of workloads. AI workloads are very data
io intensive, so that just means the data is moving
through the system with with a lot of bandwidth, memory
bandwidth and with high high frequency.
Speaker 1 (05:08):
Back and forth, back and forth. We had somebody who
I'm sure you know, Nick McEwen from course on the
show recently, who was at Cisco in the days of
building the kind of you know dominat net Round.
Speaker 3 (05:18):
He was building networks, networking systems in that company was
acquired by Cisco.
Speaker 2 (05:21):
Exactly.
Speaker 1 (05:22):
He was saying, today's data center, a single data center,
has more back and forth connections than the whole Internet
did in the nineties.
Speaker 2 (05:29):
Exactly, exactly. In fact, was fun.
Speaker 3 (05:31):
There was an announcement yesterday and videos new vera Rubin
and you looked at I mean I sort of laughed
when I looked at the photo of the launch because
you're staring at this thing and it was just basically
this medusa of basically cables and connectors behind all of that,
you know, literally probably miles of AMFONYL cables.
Speaker 2 (05:52):
Was of course some graphics processing.
Speaker 3 (05:53):
Units, but what they're having to do is connect all
that stuff back together. And so this ties to the
point around wafer scale, which is, if you are going
to build something starting from scratch, not using platforms that
already exist, not relying on CPUs or GPUs or ARM architectures,
and you were to say, what would you build that's
really designed for these kinds of workloads?
Speaker 2 (06:15):
You build what we built at Cerebras.
Speaker 3 (06:17):
And that's the journey that we started in twenty fifteen,
and really it's born of this notion of big, big
workloads deserve custom purpose built silicon, and that has been
really the story of every silicon wave over the last
five decades.
Speaker 1 (06:31):
And talk about this concept of wafers scale because I can.
I mean, if you look at the Cerebras chip, it's
fifty times bigger than an Nvidia GPU. And as you
mentioned with with Jensen holding up with verrub and it
is this kind of like, you know, sort of different
silicon chips connected together with cables and stuff, whereas your
(06:52):
chip is is kind of a single huge piece of silicon.
Why do they do that and why do you keep
it as a single piece.
Speaker 3 (06:58):
The way to think about sort of silicon, MANUF is
it's grown as you started this conversation with sand, you know,
it's grown from sand into single crystalline columns, and then
those columns are sliced very thinly that is then cut
to basically turn it into a nine inch roughly nine
inch square, and that nine inch square has trillions of
transistors on it. And the way Nvidia or ARM or
(07:20):
others build their chips is they might start with a
twelve in wafer, but then they cut it. They cut
that wafer into many smaller components that are, as you said,
sort of fifty or fifty eight times smaller than ours,
and then they cut it and then in many cases
they're actually literally putting them back together. And you ask
a question, which is why cut it? That's exactly the
question that we asked ourselves, which is you shouldn't cut it.
Speaker 2 (07:43):
You should keep it all together.
Speaker 3 (07:44):
Because as soon as you as soon as you take
communications from a chip off of the chip. So once
once a bit needs to move from one chip to another,
you basically have other components involved, and everything you add
as you would imagine with any sytem them. As soon
as you add another component between two pieces of silicon,
you've slowed it down, You've created new constraints. And so
(08:07):
the big idea here is don't do any of that.
I'm sure you're familiar. I mean, Elon, you know he's
working on other stuff in this domain, but you know
he has this, this sort of rule of thumb. When
you work at SpaceX or one of his companies, you
hear this all the time, which is the best part
is no part. And the point of that is you'd
like to reduce all of your components down to their
(08:28):
simplest sort of platonic ideal form.
Speaker 2 (08:30):
And that is exactly what.
Speaker 3 (08:31):
We decided to do with Cerebus, which is, instead of
taking a twelve inch column of silicon, slicing it into
hundreds of smaller pieces, then putting those on a motherboard
and reconnecting them.
Speaker 2 (08:42):
With copper wires or copper traces.
Speaker 3 (08:45):
Let's keep it all on one piece so that the
data can move as fast as possible throughout that single
piece of silicon, and we'll get a thousand x improvement
in performance. And that's from a memory bandwidth perspective. That's
exactly what we delivered, and so's I were now call
it roughly fifteen in some cases even more than that
times faster than any other inference platform on the planet.
Speaker 1 (09:07):
What do you think the biggest misconception is about the
quote unquote chip craze today.
Speaker 3 (09:13):
I would say the biggest misconception is how hard it
is to build a new platform to yield those systems,
and the yield it's basically for everyone you make, how
many of those can you use? And that that is
sort of one of those fundamental drivers in the semiconductor industry.
But in the case of large semiconductor systems powering it
so you know, most of these systems are now taking
(09:35):
tens of kilowatts. Back when I was building startups a
couple decades ago, you know, a rack was somewhere on
the other of two to three kilowatts. We now have
racks that are sixty kilowatts, and so you're trying to
power this system now. Of course, once you're powering it.
Guess what else you have to do. You have to
cool it. You have to get that all the all
the heat out of that system. And then of course,
once if that's that's now working, now you're going to
(09:57):
program this thing that has you know, two and a
half trillion trans and then get you know, modern models
to work on that and get them up and running
in a matter of days or weeks instead of months
or years. And so you're compounding many, many hard problems.
And in most businesses, when you compound hard problems, what
comes out the other end often doesn't work. And so,
(10:18):
you know, the single greatest challenge really is can you
actually make this thing at scale?
Speaker 4 (10:23):
Yeah?
Speaker 1 (10:23):
I think I think you said somewhere, you know, I
sometimes joke that Cerebras was five startups in one. You've
sold wafer scale, great nive power system, cool it, package it,
deploy code onto it, integrated with an existing frameworks, get
into data centers. Each one of those challenges was quite
literally its own company. So I mean, let's go back
to twenty fifteen. I mean, you had this insight, you know,
(10:43):
every new wave of computing demands and new type of
silicon right and We talked about the computing of the eighties,
one type of chips and the GPUs for gaming, which
ended up by sort of accident essentially powering the air revolution,
the arm chips that power cell phones, and then these
these new chips that cerebreast design for data centers. Essentially
(11:04):
exactly who were you in twenty fifteen, and we've got
the kind of we've got the global context, the industry context,
what's the what was the Steve context? And what made you?
I mean nowadays twenty twenty six. When we think about
venture capital and investing, you know, we often think about
like momentum based, like how do I get into anthropic
before it goes public right? Versus like how do I
(11:26):
make a ten year plus bet on something which which
you know is not only requires all these macro things
to break my way, but also has like just layer
upon layer, up on layer of complexity. And you know,
these these all these different elements we talked about in
terms of cooling and packaging and integration. What what made
you want to do this? And how much of your career,
(11:49):
shall we say, or your capital, your financial capital, your
relation capital did you stake on this?
Speaker 3 (11:54):
I love that, Yeah, so I mean the quick bounce
on me is you know, I'm trained as a roboticist.
I lived for the first ten years of my career
at the intersection of electrical and mechanical systems. I worked
at a company called Ideo, designing products.
Speaker 2 (12:10):
Some of those.
Speaker 3 (12:10):
Products brought those two worlds together as well as sort
of the design world, sort of really answering the question
of once you've solved those hard technical problems at inter
section of electrical and mechanical systems, does anyone actually care?
Do people love it and want it? So I spent
really a decade building products. Came to Foundation actually as
an entrepreneur in residence on the path to starting another company.
And this is now in the spring of two thousand
(12:31):
and seven. And so I'm a builder by nature, if
I would say, sort of more of an accidental venture
capitalist than one who sort of, you know, thought to
do this from when I was a teenager. I didn't
I literally didn't know the term venture capital until I
moved to Silicon Valley back in the mid nineteen nineties.
And the specific context and I love the question around
(12:53):
sort of what was required to get an investment like this.
Speaker 2 (12:57):
Approved as it were, more than a decade ago.
Speaker 3 (13:00):
The answer to that question is is really sort of
central to how we work at Foundation Capital, which is
we go really deep in opportunity space, and we're not
trying to be everything to everyone, and when we go deep,
we try to kind of build what we call our
points of view, a sense of what would need to
be true for us to be excited about a new
investment in this area. What are the attributes of the
(13:21):
founders that would make them successful or not, What is
the timing of the market opportunity, because of course many
venture capital investments, you know, have great ideas, but are
maybe a decade too early or in some cases multiple
dedcades too early. And so for me, I had been
studying this space, and I'd actually met Andrew Feldman and
his co founder Gary Back no joke, in October of
(13:42):
two thousand and seven, seven years before we really re
engaged around the opportunity at Cerebras, and they were actually
at that time just starting their prior company, a company
called c Micro, which was kind of also in the
kind of data center think sort of warehouse scale computing concept,
and we didn't actually invest in that company. I hit
it off with them and they were acquired by AMD
(14:05):
maybe five years later, and a couple of years into that,
I reached back out to Andrew and I could tell immediately,
I mean, I had this intuition going into the conversation
that he was not going to stay long for AMD.
He had another startup in him. And so we basically,
over the course of the next two years riffed on
a whole bunch of ideas that at that time were
kind of around that big data concept that you and
(14:26):
I talked about earlier, and we were looking at a
whole bunch of companies were using Andrew as a sounding
board for the opportunity space. And then he and Gary
and then Sean and then Michael and JP started to
basically converge on this idea around Hey, wait a minute,
there's this new workload, and this new workload we got
to pay attention to, and this new workload while as
(14:46):
you said, sort of NVIDIA is a bit of a
happy accident that it happens to be better than CPU's,
but it's not what you build. And so the you know,
these five founders basically set out to go build something
that was purpose built, and I had this prepared mind
I'd been working with Andrew on kind of a range
of concepts for two years. So when he finally said
this is now kind of March timeframe of twenty sixteen,
(15:06):
it's like, I think we're ready. I told him, I said, look,
we want to be your first term sheet and so
we It was no joke on April first, so April
Fool's Day, I said, Andrew, you know, we presented it
with an offer, and over the course of the next
few weeks we basically converged on that initial financing, which
we which we did in partnership with Eric at Benchmark
and Pierre and Leora at Eclipse, And yeah, that was
(15:28):
the beginning of a cerebras back in May of twenty sixteen.
Speaker 1 (15:33):
What was the closest moment between twenty sixteen and ringing
the bet On in the New York Sulck Exchange to failure?
Speaker 3 (15:39):
Oh many, We had many, so many white knuckle moments.
We didn't ship the first product for three and a
half years from that May twenty sixteen. I'd say there
were lots of challenges around getting that first semiconductor system
to work. You know, we'd spent multiple millions building the
first one and no joke. The very first system we
plugged it in and it blew up. I mean we
(16:02):
joke about like we call those thermal events because you
never want to tell your landlord you've.
Speaker 2 (16:05):
Had a fire. But we had an event.
Speaker 3 (16:08):
We had a small thermal event around our first system.
And what you're trusting in those cases is, wow, that
was very expensive. That one system you could you could
as certain cost you know, millions of dollars with all
the R and D that was invested in it. And
then you've got to just trust that we're going to
come back, you know, the next day and and dig
through what went wrong and figure it out.
Speaker 2 (16:28):
And and we did. And that team is just.
Speaker 3 (16:30):
Born of so much grit around sort of getting a
system that again there were no there were no playbooks
for this. TSMC had never done anything like what we
were asking them to do. And then the other piece
I would say that was a very scary moment was
when we realized, you know, we were initially actually going
to build an air cold system because every other system
in the data center, you know, including all the GPU systems,
(16:51):
were all air cooled, and we realized to make our
system work, we would need to do we'd need to
water cool it, and there was lots of questions around
how that would work. APR kind of head of mechanical
engineering and system design, had never done this.
Speaker 2 (17:04):
He kind of admended that to me later. And you know, keap.
Speaker 3 (17:06):
Dissipation at that scale is a really hard problem. And
so we we built a lot of prototypes. We actually
ended up having to bring in several consultants and I
spent you know, probably you know it was not hundreds
of hours, but you know up in that range working
with the team on how do we actually solve this
really really hard problems because you'd come out of the
lab some days thinking, guys, we might just be trying
(17:27):
to beat the laws of physics, which is never never
a good idea.
Speaker 1 (17:31):
So you, in a sense you had all of these
pod engineering problems and now engineering problems. Some of them
are solved, but no doubt are new ones. But also
of course you have you know, you have competition within Vidia.
Is that is that essentially? I mean, how does the
business look today? Is it? Is it a pitch battle
of Nvidia?
Speaker 4 (17:48):
Like?
Speaker 1 (17:48):
Is that the central kind of business challenge?
Speaker 2 (17:51):
Yeah?
Speaker 3 (17:51):
I mean they are the eight hundred bound guerilla no question,
and we are the insurgent and they've built you know,
they've built an extraordinary category. When when Nvidia moved into
the data center space, which was kind of in the
mid twenty tens and it's now you know, it's more
than half of their business, and so they are definitely
you know there, they are the player that we most
(18:12):
directly get compared with, which I'm fine with. And they
have a different architecture and you know, we fundamentally believe
we're doing we're doing things differently for again back to
the sort of you know, initial premise here, because we
think this workload deserves its own silicon.
Speaker 1 (18:27):
So so displacing them with clients, I guess is one
challenge on the on the on the demand side, shall
we say, what about on the supply side? I mean
you mentioned that TSMC actually you know, fabricate these these
chips for you guys, Are you like literally fighting of
a factory flaw space in Taiwan with Nvidia to to
to see which orders for chips they'll fulfill faster? I mean,
(18:48):
how do you how do you navigate that relationship?
Speaker 3 (18:51):
Yeah, So TSMC has been an extraordinary partner to us,
and they are to really everyone in the industry. And
what I appreciate most about them, and I think what
they are uniquely well suited to doing is working on
new ideas even before the market is sort of clearly
there for them. And they have been a great partner
to us in that regard, because again, if you're looking
(19:11):
at Cerebras back in twenty sixteen and your TSMC and
you've got you know, customers like Video or Apple and
many many others of scale that are you know, would
dwarf anything that we could ask of them. For certainly
several years, they had no business in working with us,
and yet they know that and they've proven this out
time and time again that by working with startups on
(19:33):
hard problems, they get better. And so you really want
you want partners like that. You want first customers that
are like that, that help you make your systems better,
that help you debug them, help you find problems that
you wouldn't have and resolve those problems you you wouldn't
have otherwise.
Speaker 2 (19:48):
And so they've been a great partner to us.
Speaker 3 (19:50):
But yeah, I think everyone in the world is asking
more of TSMC. They're having to make you know, investments
and you know, as is ASML and that the companies
that make their machines are also in high demand. And
I think the entire world is basically looking at this
market opportunity in front of us and saying, well, what
is the market for intelligence?
Speaker 2 (20:10):
It's unbounded, right like you you.
Speaker 3 (20:12):
Know, these are trillion dollar markets today that could be
one hundred trillion dollar markets in the future. And so
you know, working on for them, working on the fundamental
technology underlying that wave that intelligence.
Speaker 2 (20:24):
They're in high demand, no question. But they've been great,
great partners.
Speaker 4 (20:27):
To us.
Speaker 1 (20:48):
On the demands side. As I will send it. In
twenty twenty five, between eighteen ninety percent of the revenues
came from setting chips to to UAE based entities g
foot two and the Muhammad Benzai a University of Artificial Intelligence.
I guess, first question what do they use those chips for?
And second question is how do you How are you
(21:10):
diversifying the business and are you focus on winning market
share from in video or just maintaining your position As
you put it, the kind of the market for intelligence grows.
Speaker 3 (21:21):
Yeah, So I mean we got started really selling systems
to great partners early in the life cycle of a company.
We had partners with National Labs, we had pharmaceutical companies,
we had energy companies working with US buying systems, and
that was primarily in the training days of.
Speaker 2 (21:40):
Three versus History. In twenty twenty four, we launched our
inference offering.
Speaker 3 (21:45):
We built it, we kind of had a board meeting,
realized that as everyone was seeing, inference workloads were growing
very very quickly, beginning to sort of really dwarf the
training workloads, and so we ourselves realized that our architecture,
our way for scale engine was going to be really
appropriate for that. So we basically rotated towards an inference offering,
(22:07):
and as part of that, we developed a partnership with
G forty two, which is the entity you just described,
and they have been purchasing our systems and working with
our cloud infrastructure for their customers. So think of them
as sort of the AWS for the Middle East, and
they have been able to really work closely with their
own customers as well as some of our customers. In fact,
(22:30):
we have actually put a number of our customer workloads
back onto those systems that we that we put in
place now a couple of years ago.
Speaker 2 (22:37):
And then your question around diversification.
Speaker 3 (22:39):
So earlier this year we announced our deal with open ai,
which is more than a twenty billion dollar multi year
deal to build seven hundred and fifty megawatts of compute
and we were after announcing that we were only six
weeks to launching our first model with them. We love
working with them. I mean it's so great to see
too super high performance teams focused on, you know, making
(23:02):
the fastest inference on the planet, which is which is
what we launched and what we have continued to prove
out with you know, lots more to come there. So
open aye is probably the single largest sort of counterweight
to the question of you know, kind of customer concentration.
And then back in March of this year, we also
announced our partnership with AWS. So I think what you'll
see is over the coming quarters is you know, us
(23:23):
scaling those deployments with those customers. And then there's many
many more that are you know, kind of wanting to
work with us. And I think back in the early
two thousands, there are lots of companies building search products,
and Google happened to have the best one. It was
also one that was just ruthlessly focused on speed. In fact,
you remember the early days of Google, they would tell
you how many milliseconds it took to return the you know,
(23:43):
to return the query. And they did that because a
it was it was important to users. But if they
if they weren't fast, you know, their users were going
to go do other things. And so performance and inference
is absolutely fundamental. And so we while while we might
only be i'll call it sort of ten or fifteen
percent of the market today in terms of who wants
to have high performance, we think one hundred percent of
the market towards most towards high performance over time.
Speaker 1 (24:06):
Obviously G forty two's connections with with with China cause
some consternation in the US, and I believe that was
part of the reason why this rebrass IPO was originally delayed.
So I'm curious for your kind of take on on
on the geopolitics of all this. And also, I mean,
you are a you know, a sort of roboticist and
(24:27):
an IDEO guy based in Silicon Valley with like a
real passion of engineering, and all of a sudden you're
you know, on the board and the lead investor in
this in this company that's in the middle of you know,
politics and geopolitics, how do you navigate that?
Speaker 3 (24:41):
Yeah, so we I mean we see these you know,
these markets. Market again for intelligence is infinite, is unbounded,
and it's coming from everywhere in the world.
Speaker 2 (24:51):
And and I.
Speaker 3 (24:52):
Would also assert that the you know, the geopolitics of
the last four or five years have only made those
international parties more to get access to alternative systems, alternative
technologies than just the sort of you know, you know,
kind of think of it as the traditional hyper scalers
or Nvidia for that matter. I mean, we have prospective
customers in Europe and they're desperately seeking any alternative to
(25:17):
purchasing more in video gear. And so there's I think
a really sort of a pull around the world for
alternatives that are coming from parties who don't feel like
they've they've got eighty.
Speaker 2 (25:29):
Or eighty five percent market share.
Speaker 1 (25:31):
Why because if fear that they're basically not going to
get they're not going to be a big enough client
to get preferential.
Speaker 2 (25:37):
Get the allocations they are.
Speaker 3 (25:39):
Yeah, there's you know, these are you know, you're you're
you're seeing companies have to and and sovereign governments have
to come to here in Silicon Valley and I mean
Jensen's made more than a few trips to the Middle
East over the last few years. But you know before
that they were all coming here to kiss the ring
and get allocation.
Speaker 1 (25:56):
Well knows you about the open a ideal. Obviously huge
announcement a few guys, and then the market was sort
of questioning, oh, is it going to be more expensive
than they thought to actually service this deal which is
not a not a problem, which is you know, confined
to cerebras. I know, Google it getting beaten up for
the kind of cost of deployment of that cloud computing.
But how do you think about that?
Speaker 3 (26:19):
Yeah, So we it was interesting because when you know,
we did the IPO roadshow, we shared all of that
in conversations as well as in the S one, and
you know, there was this question mostly in my sense
was it was folks either not reading or perhaps misunderstanding,
or just looking for ways to sort of throw bricks.
Speaker 2 (26:36):
At young companies.
Speaker 3 (26:38):
So we basically had a home run up and down
the P and L beat really every metric. And the
question that folks were raising was around gross margin, which
was off by I think ten or fifteen percent, and
that was entirely because we were renting back capacity over
the next few quarters from G forty two, which is
the thing we talked about earlier. So we had you know,
they had deployed systems, we had customers who wanted more,
(26:59):
and so we ultimately, I mean, we put those systems
in place, and then we've developed an agreement with them
which was, hey, can we rent back that capacity for
these customers, which they were they were happy to do,
and you know, we had to take a gross margin
hit in the in the near term for that. So
this was not a surprise for any of our institutional investors.
But you know, folks are looking for, you know, whatever
might be a crack in the story for your first
earnings call.
Speaker 2 (27:19):
But yeah, this was this was a big nothing burger
in my mind.
Speaker 1 (27:23):
And what about I mean, how do you look at
the competitive dynamics beyond Nvidia obviously, Microsoft, Google, Amazon, Meta
all now rushing into the space of chip design and
and and you know for their specialized use cases in
many case which I mentioned on how computeeing. How do
you look at what they're doing, are you worried about it?
(27:44):
And how do you maintain the edge?
Speaker 3 (27:47):
Yeah, so I think broadly the way you think about
these things is you know, large markets like this, and
as we said, these are the mother of all markets,
you're never going to have them all to yourself. You know.
If you're right about these things, they're going to attract competition.
That's just basics and you know economics. And so what
we observe is every one of the hyperscalers has already
(28:09):
developed or is it you know, in development on their
own silicon. I mean, obviously Google's had their TPUs in place.
They're on their fourth or fifth generation now, the Trainium
and Infernia platform at AWS, and again we're in partnership
with AWS. Fact you know, we've we've shared some of
the ways in which we're working together on a very
novel disaggregated solution where we can kind of get the
(28:30):
best of their hardware and the best of our hardware,
which is where I think it will be really interesting
things to sort of.
Speaker 2 (28:35):
See that play forward.
Speaker 3 (28:36):
Facebook of course, has been talking about their own platform
to build their own silicon, which you know they've been
doing in partnership with Broadcom for some time. And so
you know, we see this as any large market is
going to attract multiple players, and what I actually really
appreciate about it is everyone's at least looking at their
workloads and saying, what is it that we should build
here that's well suited to us, And what we're building
(28:58):
at THREEBRUS is very well suited these inference workloads. It's
also well suited to training, and it's you know, it's
something that took as I said, you know, it took
us multiple years in three generations before it really felt
like we had a system that was really singing. And
most of the other insurgents, you know, are still telling
you what's about to ship you know whatever in the
fall or in Q one of.
Speaker 2 (29:18):
Twenty twenty seven.
Speaker 3 (29:18):
But we, you know, we have systems out in the wild,
and we're scaling dramatically, and we're also working on our
next generation systems, which are which are super exciting.
Speaker 2 (29:25):
So that's how I think about it.
Speaker 3 (29:26):
I think there's going to be many many more, could
be more upstarts. You know this, this this market is massive,
and so we should assume there's gonna be many, many winners.
Speaker 1 (29:36):
So you have this claim now, which which all vcs
would love to be able to make. That you essentially
soar around the corner, right, and so that must be
a fun moment for you. And I'm curious about two things.
One is, you know, ten years ago you saw this
(29:57):
kind of very strong demand see all for you big
data processing what became AI workloads and invested in and
built Cerebrass entergy for thease opportunity. What are you seeing
now ten years into the future in terms of demand
signals that you're investing in, either through your fund or
(30:18):
as part of Cerebress.
Speaker 3 (30:20):
Well, when I think about what today we're excited about,
I mean what you see in our industry. I think
in venture capital and tech more broadly, and you've seen
this as much.
Speaker 2 (30:29):
As I have.
Speaker 3 (30:29):
Is every wave, every platform, new platform technology, whether you
know we were talking about sort of the X eighty
six platform or GPUs or mobile and now AI. There's
a sort of quite roughly a decade of building out
the infrastructure and then there's multiple decades of building out
the application layer on top of it. And so there's
(30:50):
no question and we're spending you know, we more than
one hundred companies broadly, and that you describe as sort
of AIRAI native, you know, across our portfolio where a
thirty one year old venture firm we've been investing in
AI really sort of the sort of fundamentals of it
since two thousand and eight two thousand and nine time frame,
and so we look at this next waves as involving
(31:11):
a lot of innovation at the application layer.
Speaker 2 (31:13):
Now what's tricky about this, as I'm.
Speaker 3 (31:15):
Sure you're aware of as well, is well which parts
of those application layer are going to get completely commoditized
by perhaps even the fundamental foundation models that you're building
on top of. And so we look for the sort
of thin, sort of proprietary data loops, you know, ways
in which there's a data flywheel where with every use
of the product, you know, your insight around your customer
(31:36):
and your value proposition gets deeper and richer. And we
look for sort of really interesting ways in which things
that were today delivered as a service, oftentimes literally as
maybe it's a software service, maybe it's a human service,
maybe it's labor, and we look for ways that that
is going to be automated over time. And so we've
made you know, a number of really interesting investments in
this area. And then when it comes to sort of
(31:57):
AI infrastructure, I mean we all will have to kind
of like point to the you know, the the analogies
to how are you know natural systems work, whether it's
a human brain or not. But you know, these systems
that we're building are far less sample efficient than the
way natural systems work.
Speaker 4 (32:15):
You know.
Speaker 3 (32:15):
You know, if you've got a twenty whatt brain operating
at two hundred hertz and it can do, you know,
sample efficiency is much much higher than what is used
to train a very intelligent next generation frontier model.
Speaker 1 (32:29):
Some of thefsionist means that you can learn from fewer examples.
Speaker 2 (32:31):
Yeah, exactly.
Speaker 3 (32:32):
Yeah, and so I think there's going to be some
really interesting work in this area that then allows us
to build models that are more efficient both from an.
Speaker 2 (32:41):
Energy perspective but from a data perspective.
Speaker 3 (32:43):
So yeah, I'm very excited about, you know, some of
the work that's being done in this domain.
Speaker 1 (32:48):
Steve has outa thank you, thanks so much, as it's
just so great to talk with you for tech stuff.
(33:12):
I'm Osvoloshin. This episode was produced by Eliza Dennis. It
was executive produced by me and Julian Nutta for Kaleidoscope
and Katrian norvelve iHeart Podcasts. Jack Qinsley mixed this episode
and Kyle Murdoch wrote Olph theme song,