Episode Transcript
Available transcripts are automatically generated. Complete accuracy is not guaranteed.
SPEAKER_02 (00:00):
Today we're asking
one simple question.
Can your platform actuallyhandle AI or will your GPUs melt
first?
This is Confluent Developer.
SPEAKER_00 (00:12):
We're constantly
trying to solve and innovate for
these problems right now.
Okay, how do we track thatvariability across the entire,
say, GPU data center?
The problems that we had 10 or20 years ago are the same
problems we have now, just in adifferent context.
SPEAKER_02 (00:27):
Hello everyone.
I'm Adi Polak, and you'relistening to Confluent Developer
Podcast, where we uncovered thehuman stories behind complex
software systems.
In this episode, I'm discussingwith Brian Oliver, a platform
architect and author, and one ofthe voices shaping how
large-scale AI systems areactually being delivered in
(00:51):
production.
From teaching swimming lessonsand building gaming in his teens
to writing Lua add-ons for Worldof Warcraft, to now scaling
GPU's platforms and influencingthe future of cloud native AI.
Brian's journey is all aboutperformance, resilience in the
(01:14):
pursuit of predictability in aworld that is anything but
Brian.
SPEAKER_00 (01:28):
Hi, how are you?
SPEAKER_02 (01:30):
I'm good.
I'm I'm so excited to that wefound the time to actually talk.
I know you've been doing a lotof things recently and you've
been super busy, you know,working on architectures and
GPUs and platform engineeringand so many good things.
So what what keeps you busythese days?
SPEAKER_00 (01:48):
Uh yeah, so really
deeply ingrained in large-scale
GPU work right now.
Um giving a talk at QCon in NewYork, QCon AI in New York about
some of that, like doing chaosengineering and large-scale
GPUs.
Um also, you know, book writingand giving talks and all this
other stuff.
It's uh yeah, it keeps me prettybusy.
(02:08):
For sure.
SPEAKER_02 (02:09):
Right.
Book writing is is a labor oflove.
SPEAKER_00 (02:13):
It really is.
Yeah, we just finished our firstone, um Effective Platform
Engineering with Manning.
Uh it arrived on my door a fewdays ago, which was pretty cool.
And then uh my second book withO'Reilly is uh we the title may
be changing actually, butcurrently it's um delivery
systems, but we are reframing itto have some AI stuff in it.
(02:34):
Um so it might be something likedelivery systems in the age of
AI or something like that.
So but that'll be in about ayear from now.
unknown (02:41):
Yeah.
SPEAKER_02 (02:42):
That's really
interesting.
So it's our do you do you seelike a profound change to our
delivery systems with AI today?
SPEAKER_00 (02:50):
Yes.
Um this is maybe also aprecursor to the ThoughtWorks
radar coming out pretty soon.
Um I'm on the committee thatwrites that.
And one of the central themes ofthe most recent radar session in
Bucharest was um we talked a lotabout AI scheduling, which is a
deep problem right now.
Um, one of the things going onin this space is you have these
(03:13):
really massive data centers withyou know thousands of GPUs and
they all cost millions ofdollars.
Like a I think a GB200 rack thatNvidia came out with recently
costs like three million dollarsfor one rack.
Um and they're allnetwork-linked GPU to GPU, and
they're they're next to eachother.
Um and so one of the things thatKubernetes and other scheduling
(03:36):
systems wasn't super aware of islike how to deploy jobs in a
contiguous way, where it's likethat because the jobs are split
up across tons of GPUs, so youwant the job to hit like a block
of GPUs that's close to eachother to reduce latency.
Um so we call this topologyaware scheduling, something
that's existed in computing forages, but now it's a GPU
(03:57):
problem.
And so all of our deliverysystems are starting to adapt to
this new requirement.
Um it's really interesting.
You see a lot of projects in theCNCF and in SLERM that are
tackling this problem.
SPEAKER_02 (04:10):
That's amazing.
I'm guessing I can read moreabout it in your upcoming book,
right?
SPEAKER_00 (04:14):
Yeah, the O'Reilly
book, we're gonna talk about it
a lot.
Um, in the more short term,we're gonna have some things on
the ThoughtWorks radar for thatas well, in the themes as well
as some of the blips.
Um yeah, definitely check thatout.
SPEAKER_02 (04:27):
Yeah, I'm definitely
subscribed to that, and every
time it comes out, it's uh italways becomes a huge
conversation uh with my team andin the company about all the new
innovations in uh what is uh youknow uh what has proven itself
to uh to work well inproduction, what is still
experimental, and uh I'm superexcited for that.
SPEAKER_00 (04:48):
Yeah, it's there's
pretty much this whole radar
edition is all about AI stuff.
It just dominated all of theconversations, so it's gonna be
a good one.
Yeah.
SPEAKER_02 (04:56):
Yeah.
And hopefully it becomes morepractical.
Um so solutions are use casesare getting into production and
uh being materialized there aswell.
SPEAKER_00 (05:06):
Yeah, yeah, that was
a lot of a like it seems like in
the past sessions we it was lotsof like, oh, experiment with
this technology, try that.
Now it's like a big, yeah,bigger focus on like all the
production stuff and likescaling.
Um, and those are all becomingmore interesting topics now.
Yeah.
SPEAKER_02 (05:23):
Exciting times,
exciting times.
For sure.
Um, so I don't know, you know,uh we have a lot of people
listening in, and a lot ofpeople are curious about you and
and your background and and morethings.
So maybe you can take us back intime to your first ever job.
SPEAKER_00 (05:40):
Oh boy.
Um my first job was as ateenager.
I I worked at a swimmingacademy.
Um, and I started off withhelping them build pools because
they were expanding um theirscaling.
Um, and then I eventually got myI didn't like that part, so I
got the the certifications to bea lifeguard and a swim
(06:01):
instructor there.
Um and so I started off withteaching kids how to swim and
eventually moved into liketeaching like um triathletes and
stuff, um, like working on theirperformance tuning, and and I
really enjoyed that.
And then as I got older, I'm inlike my late teens, I worked at
a land gaming arena.
(06:22):
Um, and that was a a reallylow-paying but really fun job.
We we got into some pro gamingthrough that because we had
teams that were hosted there.
So uh I never made it onto uhlike a starting squad, but I was
an alternate on some likebattlefield pro teams and call
of duty.
So I got to participate in somesome like game battles
(06:42):
tournaments, and it was really,really fun.
It was uh it was a cool job.
I enjoyed it.
SPEAKER_02 (06:48):
Sound exciting.
So the love from like scalingpools and uh swimming lessons to
scelling GPUs and performancetuning.
Uh it's kind of like a themethere.
SPEAKER_00 (06:59):
I know, right?
Yeah, it it's funny.
I actually wasn't even intocomputer science or mathematics.
Uh I but I grew up playing Worldof Warcraft from like age, I
don't know, 13 to in well intomy adult years.
I don't play anymore, but I Iwould write um add-ons for the
game in Lua.
(07:20):
So I was experiencing I got someexperience programming early on,
but I never like considered itas a career or anything.
I was gonna be an English majorin college, and I ended up uh uh
switching to computer sciencebecause I took this symbolic
logic course, which was like aphilosophy course.
Um, but it's basically justdiscrete math but more abstract,
which is the sort of core mathbehind computer science, and so
(07:42):
I ended up switching because Ienjoyed it so much.
Yeah.
SPEAKER_02 (07:45):
It's fascinating.
It's always interesting how youknow discrete math and some uh
you know aspects of mathematicscan be it's more philosophy
driven, and uh, and then webring it into the computer
science world, into the actualuh zero ones and uh soon, who
knows, quantum and the rest ofthe things.
SPEAKER_00 (08:02):
So yeah, it was
really fascinating.
Like it in the philosophyaspect, it's lots of like set
theory, and but it it ends upbeing very similar to
manipulations you do in discretemath and techniques you use, um,
which is really just it reallydrove me towards computer
science and math at the end ofthe day.
I went from wanting to gradeEnglish papers and teach British
(08:23):
lit to to yeah, programming GPUsis quite the switch.
SPEAKER_02 (08:29):
I can imagine.
And and and right on time.
I think when I look into thefuture, when I hear you know,
people like um Nvidia CEO and soforth and so on, I think GPUs
might be dominate like the thefuture cloud.
SPEAKER_00 (08:43):
Yeah, yeah, I think
so.
Um we're gonna see a lot of Imean we already do, we see a lot
of like large-scale cheaper GPUum scaling happening for the
inference requirements.
But what's interesting is thethe model sizes are changing the
definition of what small andlarge mean.
(09:03):
Um like the what small languagemodel means now is kind of a
moving target to the point whereit's like you have these um GPU
racks that I was talking aboutearlier, the$3 million one, and
that thing has 13 terabytes ofmemory, unified memory, and it
acts as one GPU.
That's now the new large in alot of ways.
(09:27):
Um, what does that mean mean forsmall language models?
Um and the size is kind of justthis like moving target.
So we're starting to see likethose scaled requirements I was
talking about earlier withscheduling now are even being
applied to small language modelsbecause they still need to be
spread across multiple GPUs.
Um so we're starting to seethese distributed complex needs
(09:49):
um occur even in everydaycomputing and companies that
don't have access to large-scaleGPU infrastructure.
Yeah.
SPEAKER_02 (09:57):
Interesting.
So we we didn't yet I think whenwe started, it was like a small
language model should fit on onemachine.
So essentially beginning whatI've you know experienced was it
wasn't uh a distributed topologybehind that.
So now you know we're seeingthat we're growing past that.
But uh a little bit, yeah.
SPEAKER_00 (10:16):
Like you you have
like uh what is I think Quen 72B
or it's a I can't rememberexactly, but there's there's
some models that are like 72million parameters, and you
would now almost think of thoseas small when you compare to
who's using 13 terabyte um GPUracks and even spreading across
multiple of them.
(10:36):
Um maybe that's just consideredmassive scale and we we need a
new definition.
Um but yeah, it's reallyinteresting.
It's definitely a moving target.
SPEAKER_02 (10:45):
Yeah, and it
definitely pushes you know
platform engineers to to thinkbeyond uh some of the practices
we had so far.
Um this is why I'm I'm I'm superexcited for your books.
So uh I'm gonna order both ofthem and uh and read through
them and have some notes.
Um I'm gonna I'm curious, you'vebeen solving a lot of tech
(11:06):
challenges in in your career,and uh of course you're at the
forefront of what's happeningnow in infrastructure and
platforms.
So maybe you want to share somechallenges that you're working
on and perhaps some learnings oror things that you know uh came
out of them.
SPEAKER_01 (11:22):
Now a quick word
from our sponsor.
Confluent developer the podcastis brought to you by Confluent
Developer the website, which haseverything you need as a
developer of data streamingsystems.
And it's completely free.
We've got curriculum, hands-onexercises, executable tutorials,
the online data streamingengineer certification, also
free.
A way to find a meetup near you,those are free.
(11:45):
Everything is there.
I really want you to besuccessful in your journey as a
data streaming engineer, andthis is the site that has what
you need.
Check it out atdeveloper.confluent.io.
That's developer.confluent.io.
Now back to the show.
SPEAKER_00 (12:02):
Sure.
Um, did you want were youinterested in challenges I had
many years ago or more recent?
SPEAKER_02 (12:08):
I think the recent
one are super fascinating.
Um, but you know, I'll let youdecide.
Cool.
SPEAKER_00 (12:14):
Yeah.
Um yeah, I think the one of theinteresting challenges now is um
not just the topology awarenessum like we were talking about
earlier with trying to reducelatency for deploying AI
workloads.
Um there's this new kind ofnewer concept.
There was a paper presented atSupercomputing Conference 2024
(12:36):
here in Atlanta on umvariability aware scheduling.
And I'd never thought about thisbefore, but when you have, say,
a hundred GPUs and they're allidentical, um, the difference in
their performance is um muchlarger than 100 identical CPUs.
It's much less predictable, infact.
(12:57):
Um GPUs are just not as stableof hardware, not as predictable.
So, like if you have a datacenter full of them and they're
all identical, some are going tobe performing better than others
by a pretty significantpercentage in some cases.
And this could be related tosome are being cooled better
than others, some are maybe notgetting the amount of power
(13:17):
they're supposed to get.
Like, there's all thesedifferent factors into why
they're not quite aspredictable.
And so one thing we're startingto think about is okay, how do
we track that variability acrossthe entire, say, GPU data center
and figure out, okay, this blockover here has a few that aren't
performing as well.
Maybe the cooling's not doinggreat over there, but it's it's
(13:40):
still working fine.
We could still run things overthere.
So, how do we maybe like say,okay, if we're doing like a
multi-tenant architecture datacenter, send like, okay, we have
this really premium customer orreally critical job we need to
run.
We're not gonna put them in ablock of GPUs that has maybe
some that aren't performing aswell.
Um, and then there's alsodifferent types of training
(14:04):
jobs, even um where it's likesome don't really care about
those performance differences,whereas others do.
So it's like you are alsostarting to maybe categorize the
types of training jobs you'rerunning and then scheduling it
to certain areas of your datacenter and others depending on
the needs.
Um, so it's the complexity justkeeps going up and up.
(14:24):
And so we're we're constantlytrying to solve and innovate for
for these problems right now,and then new techniques are
coming out, like new papers arecoming out every day where it's
like our team chat is like, readthis paper, because this is
really like it, this is almost adaily occurrence now, um, where
we're like, oh, cool.
Oh, yeah.
SPEAKER_02 (14:42):
Yeah, it's
fascinating.
So essentially, you know, I canbuy the same hardware, I'm
paying you know good money forit.
And uh because my cooling systemor electricity system is not uh
top-notch across, I'm guessing.
Um, this is when I'll startsaying seeing like different
performance coming out of uhthese GPU, and I'm guessing, you
(15:02):
know, even if you put all theGPUs in one rack, just the
proximity of one GPU to theother and the the heat that
comes out of that could be uhimpacting the performance.
Um and you mentioned that thereare some algorithms that are uh
okay with this uh you knowperformance differences and and
(15:23):
variability, and some of themare are not.
I'm curious, like what's uh youknow, what's in the algorithm
makes the difference?
Um and how could people navigatethat perhaps?
SPEAKER_00 (15:34):
It's not so much the
the algorithm, even it's um it's
like you need to have the thedata.
Um so like if you you're alwaysrunning AI workloads inside of
say, I'm using like one datacenter as kind of an abstract
way of talking about thisproblem, but this is across
hundreds.
Um it's you need to collecttelemetry on all of that
(15:57):
hardware and then sort ofcentralize that into some sort
of monitoring system.
So then your scheduling systemcan be aware of like each one um
individually so that it can thenmake intelligent decisions.
Um then you have technologieslike Slurm in the sort of old
school.
I would say old school becauseit's actually becoming quite
(16:18):
modernized by cloud nativerequirements.
Um but Slurm's been around for along time, since I think 2004,
that is aware of all of thesedifferent types of nodes that
are available to it.
Um and Slurm actually runs likea daemon on every single GPU
server in your data center.
So it's collecting informationon that node, it knows all of
(16:41):
its different um aspects, typesof GPU, all that.
And then it can make sort ofgood decisions on those basic
hardware things, but it's notaware of the those variability
differences.
Like you you have to implementthat yourself.
Um, so we're starting to look atalso CNCF-based projects like uh
Q or DRA and Kubernetes um APIadvancements.
(17:04):
Um there's also like the Kaischeduler, um Skypilot, which is
uh Kates or non-Kates.
Um all of these can befine-tuned to be aware of the
topology.
Um, but the it takes investmentand time from your team in order
to take advantage of that.
Yeah.
SPEAKER_02 (17:21):
Yeah.
And is there a way to improvelike a specific Jupyter
performance?
Let's say now you know we youdon't send any workload to it,
put it on idle for X amount oftime.
Is that something you've beenlooking into?
SPEAKER_00 (17:34):
Not as much, but we
we are trying, we are looking
into the talk I'm giving at QConis about um maybe not driving
more performance out of currentones.
A lot of that has to do withlike cycling them out and
scheduling them for work ormaintenance from, say, data
center team, but um getting morepredictability out of the
performance by maybe doing chaosexperiments um on your GPU
(17:58):
infrastructure so you start toknow ahead of time where the
issues are instead of findingout while you're currently
running workloads.
Um we're starting to think aboutlike, okay, how do we apply the
concept of chaos engineeringfrom the API world and start to
move it into this GPU world?
Um and it's very different.
It's hard to put something likethat together because you need
(18:20):
access to really expensiveinfrastructure in order to
really make it work.
Um But we're we're starting towork with um Chronosphere on an
open source project to try andget something that that's
functional for doing that.
Yeah.
SPEAKER_02 (18:34):
Yeah, it's it's
fascinating.
I wonder if you can walk methrough like what would an
experiment look like.
SPEAKER_00 (18:40):
Um yeah.
For sure.
Um so you you might have youryour typical ones um where
you're like say blackholingcertain nodes.
Um what black holing is is it'sthe concept of like things that
can go in, but they can't comeout of, say, a node.
Um you might do like take nodesdown, like a chaos monkey kind
(19:01):
of thing, like all the normalstuff.
Um but where things getinteresting is you also want to
do things specifically to theGPU.
And one node might have, um thisis quite normal, like one server
might have eight GPUs on it.
And the way you shard up areally large training job is you
like say you're writing somePyTorch code, you might inform
(19:24):
your training job or yourscheduler, okay, you're going to
be split up across, say, 16GPUs, and this is going to be on
two servers, so sets of eight.
Um well, the way that that ummodel that's being trained gets
divided up, if one GPU from theset of eight goes down, you then
(19:45):
need to stop the job and switchover to another server that has
a full set of eight.
Like you can't just continue onwith the odd number.
You either have to tweak yourjob and your parameters or move
over to another node that has afull set of eight going.
So in the chaos world, whatwe're thinking about is like,
okay, how do we go into theserver and maybe bring down one
(20:06):
of the eight and see how ourscheduling processes respond?
Um we could also maybe createnoisy neighbor type things where
it's like there's somethinggoing on with one of those GPUs.
Maybe it's being consumed byanother job or another daemon or
maybe an agent collector.
Um and then you can get reallylow-level and like eBPF or um
(20:27):
NVIDIA has um DCGM, which isthis um sort of low-level GPU
management tool where you caninteract and actually send
commands to directly to GPUs,and you can send it faults.
So you can pretend to createlike heat or power faults
through DCGM.
Um so we're starting to thinkabout things like that as well.
unknown (20:50):
Yeah.
SPEAKER_02 (20:50):
That's cool.
So Nvidia actually thought aboutit and it's like, hey, we know
it's gonna be stress test, anduh people are gonna build
different experiments, and uh,we're enabling people to uh to
send faults to our GPUs as well.
Um kind of, yeah.
SPEAKER_00 (21:03):
They they did it for
themselves really, because they
they need to test their GPUs umwhen they're building them.
So they built the tool, I think,more for that.
Um the DCGM has two projects, umthe exporter, which was created
for the Kubernetes world, whichjust collects metrics on the
GPUs running, and then the DCGMlike CLI tool, which actually
(21:25):
like you can run commandsagainst.
And they really made it forthemselves to test and work on
GPUs and verify they're good.
But it's just a CLI process youcan build it into like a job or
API, that kind of thing.
Yeah.
SPEAKER_02 (21:39):
Yeah, it's amazing
how in software, you know,
sometimes you build one thing,it's just for us to test and
validate that everything works,and then we expose it to the
world, and the world's like,hmm.
I have some other use cases justfor that.
SPEAKER_00 (21:51):
Exactly.
Yeah, exactly that.
SPEAKER_02 (21:55):
Very cool.
So I I'm curious, like you know,it's um I didn't realize that.
The beginning to be honest, likeGPUs is going to have a
different performance becauseI've been um you know of a
working a lot with more CPUs,and GPUs are kind of like an
additional to uh to the existingCPUs that I I worked with.
And that that's really asurprise to see that there is uh
(22:16):
variability uh in in the samemachines.
I'm curious if there were youknow more things that kind of
like are were less intuitive toyou that you learned through
working on these projects.
SPEAKER_00 (22:29):
Um mainly it was
around the the differences
between machine learning, likewhat a machine learning engineer
wants and needs and anoperations engineer wants and
needs, because they're the thecurrent like AI workload and
supercomputing communities umhave really been using this
(22:49):
Slurm project for a long time.
And um the Kubernetes communityis really trying to move into
supporting this space becausethe the AI operations engineers
want to use Kubernetes and theMLE engineers just want their
jobs to work.
And that's been Slurm for a longtime for them.
It's just like they can useSlurm APIs and quickly write
(23:10):
their jobs and they don't haveto care, um, which is different,
but also similar to like theproblems we've heard from like
when DevOps came around, whereit's like devs just want their
code to work, they just want towrite code, and operations
engineers want it to bepredictable.
And Kubernetes kind of solvedthat that problem of helping
them meet in the middle.
Now we're trying to do the samething with AI workloads and get
(23:33):
to the point where our MLEengineers can, um MLE is the
acronym we use for machinelearning engineer, um, can
create jobs that are deployed toKubernetes, and it's that same,
like they don't have to carewhere it's running.
Um I was at the uh Kubernetes umcontributor summit in in Paris
(23:56):
last year.
And Tim Hawken got on stage andhe was like, we have to start
like all focusing on AI support,or our project is going to get
left behind for something else.
And from and that conference wasso focused, that KubeCon was so
focused on AI.
Um and since that moment, itreally has exploded in the CNCF
(24:17):
space.
SPEAKER_02 (24:18):
Yeah, it's um it's
fascinating how requirements get
changed, but the kind of thenature of the two different
types of engineers that worktogether in a company, it's
like, you know, yeah, MLengineer wants to get things
done, platform engineer want toget real real ability and uh
predictability for for theplatform.
(24:40):
Um so beyond like we have theGPUs, it's kind of a game
changer and they bring new umchallenges, we'll put it that
way.
If we do technical problems orchallenges that uh we got to
solve.
What about like working withEmily's are requirements?
You mentioned PyTorch.
Um are there other requirementsthat they're now focused on
(25:02):
these days?
SPEAKER_00 (25:03):
Yeah, there is
PyTorch is just you know one
project of many that uh machinelearning engineers are using.
Um so we the the approach we'retrying to focus on is um helping
them sort of containerize theirjobs and workloads so that they
can be cloud native.
Um what's interesting is youeven have some different
container runtimes in thisspace.
(25:25):
Like um there's nroot, which isum one of NVIDIA's sort of
container runtime.
Um you also have like the latestand greatest GPU architectures
um have some differences betweenmaybe traditional um server and
node architectures that requirechanges to these upstream
projects in order to supportthem.
(25:46):
Um so it's re you start to seelike certain projects need to
sort of adapt um to theseever-changing like the um the
GB200 um has some differences inits hardware architecture that
require um you to think a littlebit differently about how you do
AI and how you do training inorder to make use of that really
(26:08):
large-scale um hardware.
Um, because it's technicallylike one GB200 is, I think, like
72 GPUs, but the way they'vedesigned the architecture is it
acts as one.
Um so the way you write yourjobs is a little bit different
in that context.
Um I've never written one myselffor GB200, so I just know kind
(26:31):
of how that works.
Um but you know, getting accessto that kind of thing is is very
expensive.
So it's all kind of theoreticalfor some of us.
Yeah.
SPEAKER_02 (26:39):
I can imagine.
And it's probably like the Emilyneeds to change some of the way,
or the researcher need to changesome of the way they're
developing.
Or is it like a platformsupposed to be kind of like um
catch them all, uh figure out uhyou know which hardware we're
running on right now and try todo the adaptations?
SPEAKER_00 (27:00):
Yeah, they
definitely need to be aware.
Um even you know, the MLEs thataren't working in that bleeding
edge, really large scale.
Like let's say they're usinglike, you know, the one of the
standard GPUs you see in cloudsis like the um, I think it's the
Tesla T4 or L4, um, or maybelike some H100 um GPUs, which
(27:22):
are quite large.
I think they support like 80gigabytes of VRAM per.
Um the MLE needs to be aware oflike what node server, like GPU
server they're targeting and howmany GPUs it has, how much
virtual memory it has.
Those are all things they'reprobably always going to have to
care about.
Um but the things we're tryingto abstract are how that
(27:44):
workload gets scheduled to thatinfrastructure.
They shouldn't care if it's umSlonar or Kubernetes or
whatever.
Like they should be able to justwrite their training workload or
their inference workload anddeploy it.
And so that's the abstractionwe're trying to get to.
SPEAKER_02 (28:00):
Very cool.
Very cool.
Um one or two learnings that youwant to share today from all
these great projects that you'reworking on.
SPEAKER_00 (28:10):
Um or two learning.
I think the it's funny to me howthe same.
I guess one of the things Ilearned is the problems that we
had 10 or 20 years ago are thesame problems we have now, just
in a different context.
So I'm learning like theproblems repeat themselves just
in different ways.
Like the, you know, we weretrying to bridge the gap between
(28:31):
developers and operations 15years ago when Kubernetes solved
that, and now we're we're kindof dealing with the same
problem.
So it's like I'm I think whatI'm learning is I'm starting to
recognize these problem patternsum in computing and realize
like, oh, this is similar to aproblem we solved 10 years ago,
it's just a different kind ofhardware or context or whatever.
(28:51):
So I'm learning to mayberecognize those things, um, and
that helps me like think abouthow to solve them.
unknown (28:56):
Yeah.
SPEAKER_02 (28:58):
Yeah, uh software is
software is software.
It's uh great advice.
Always stick to fundamentals,you know, know the core, know
the basics.
SPEAKER_00 (29:06):
Exactly.
Yeah.
SPEAKER_02 (29:08):
Awesome.
Brian, thank you so much forsharing your expertise and
experience here today.
I am super excited to read yourbooks.
And I know you also you alsogive a workshop, right?
SPEAKER_00 (29:19):
Yeah, so we um me
and one of the other co-authors
on the platform engineering bookjust announced a workshop on
platformengineering.org.
It'll be the platform architectcertification for that website.
Um, and then I'm teaching aworkshop on platforms with AI at
KuCon San Francisco coming up inNovember.
SPEAKER_02 (29:42):
Very cool.
So if anyone is listening athome and want to catch up and
want to become certified andknowledgeable in all this space
and build a career there, theyshould definitely uh go to this
website and uh check out yourcertificate and uh get your
books.
Thank you so much.
Thank you.
It was a pleasure.