All Episodes

February 2, 2026 30 mins

Tim Berglund talks to Richie Artoul (WarpStream/Confluent) about his career in data infrastructure. Richie’s first job: working at Howie’s Game Shack, a walk‑in LAN gaming cafe. His challenge: working at Datadog on a new log storage system.

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
Listen
Watch
Mark as Played
Transcript

Episode Transcript

Available transcripts are automatically generated. Complete accuracy is not guaranteed.
SPEAKER_00 (00:00):
Today, from walk-in land parties to discless data
infrastructure pioneer, this isConfluent Developer.

SPEAKER_01 (00:07):
And as we were getting ready to kind of like
migrate our first product, werealized it just like kind of
didn't work.
Someone's day is going to beruined every day at that level
of nights.
When it was not working and wewere grinding on it trying to
make it work, I was like, man, Imight have to eat some crow.
Like, this is not working.

SPEAKER_00 (00:21):
Hey there, everybody.
I'm Tim Berglund, and welcome toConfluent Developer, the podcast
where we explore the journeys ofsoftware developers tackling
really hard problems.
In this episode, I'minterviewing Richie Artoul,
who's famous for being thefounder of the pioneering
Diskless Kafka implementationwarp stream.
We talk about his first job at agaming cafe and some really

(00:44):
pivotal work he did at Datadogredesigning their log storage
engine and indexing system.
Richie tells a great story.
I always really enjoy talking tohim, so let's get to it.
Welcome to another episode ofthe Confluent Developer Podcast.

(01:05):
I am your host today, TimBergman, and I'm joined by
Richie Artool.
Richie, welcome to the show.

SPEAKER_01 (01:11):
Hey Tim, thanks for having me on.

SPEAKER_00 (01:13):
You got what do you uh what do you do right now?
What's your uh current job,current title?
What do you what are you allabout?

SPEAKER_01 (01:20):
Uh my current job is director of engineering uh for
warp stream at Confluent.
Uh so I lead uh essentially whatis like the Warp Stream group at
Confluent.
Um so we're like a product lineum um so we have kind of like
our uh our own software and umproduct yeah um and it is a

(01:44):
fascinating one.

SPEAKER_00 (01:45):
Maybe we'll talk about it today.
Maybe we won't.
What was before you did that,long before that, your first
job?
What was your first job ever?

SPEAKER_01 (01:54):
Um you want like my real first job, right?

SPEAKER_00 (01:57):
Real first job.
Like I don't want your first,you know, uh interned at Google
writing that's like your actualOkay, cool.
Yeah.

SPEAKER_01 (02:05):
Um yeah, so I I mean I actually didn't study like
computer science in college orwhatever, so my my first job out
of college also wasn't tech, butokay.
Um my first real job was workingat a place called uh Howie's
Game Shack, uh which I think Iwas 16, and it was like uh I
forget what they're called, likeone of those places you can go,
like an internet cafe.
Yeah, but it was all games, sothey had like 50 Xboxes and like

(02:28):
a hundred PCs, uh and peoplewould go and play like Dota,
like on Warcraft.
I don't know if you ever heardof that game, and uh
Counter-Strike and World ofWarcraft there.
Um, and I think this is whenlike the iPhone had just come
out and I like really wanted onebecause I just wanted to have
access to the internet all thetime.

SPEAKER_00 (02:48):
Sure.

SPEAKER_01 (02:48):
I was a nerd.
Right.
Uh and I was like, okay, so Ineed a job.
Uh so I got a summer job there.
Um and I think I made just aboutenough money to pay for an
iPhone.
So that is amazing.

SPEAKER_00 (02:57):
So it was like a retail land party was the the
the business.

SPEAKER_01 (03:00):
Yeah, that's exactly what it is.
Yeah, it was an interesting,interesting group of people.
That is wrong.
Yes.

SPEAKER_00 (03:06):
I mean look, it's it's kind of our people.
Yeah, yeah.
We can say that, and yes, it isa very interesting group of
people.
Yeah, yeah.

SPEAKER_01 (03:14):
Uh well, especially just because it was a lot of
like um, you know, high schoolkids and college kids, and I was
like 16.
So uh it was fun.
It was a good it was a goodfirst job.

SPEAKER_00 (03:23):
That that is an amazing first job.
I love that.
Uh and I and I want to go there.
I I mean I don't know if I seeuh even a business that exists
anymore, but just sounds like afun place to hang out.

unknown (03:32):
Yeah.

SPEAKER_00 (03:33):
Well, I you mentioned that you weren't a
computer science major.
Uh this isn't scripted, but whatuh what was your major?
What'd you what'd you do incollege?

SPEAKER_01 (03:40):
Uh I studied biochemistry in pharmacology.
Um so uh you know, I don't know.
My dad's a doctor.
I was told to be a doctor from ayoung age.
Uh that was that was the plan.
Um and then uh in college I kindof realized like I didn't like
hospitals, so I was like, Iprobably shouldn't be a doctor.

(04:02):
Um but um you know I was youknow kind of like top of my
class, pre-med, that sort ofthing.
Uh and then just decided to nottake the MCAT.
Um so I uh just ended up workingsome kind of weird um office job
doing like you know, editingword documents and type of

(04:24):
thing.
Finding finding your way, yeah.
Yeah.
Uh and then obviously hatedthat.
Um and I think like about 10months into that job, I just was
kind of like losing my mind andjust kind of like rage quit and
went to a coding boot camp.
Um and so that's kind of how Igot into tech.

SPEAKER_00 (04:42):
I love it.
I I did not know that part ofyour story.
Um that uh that is fantastic.
Uh and you know, well, I thinkwe're all kind of glad you did.
Thanks.
Um so the big question of theshow What's the most interesting
problem you've ever solved?
Uh I mean I I gotta tell you,I'm expecting you to say Warp
Stream, but I I I don't maybeyou're not gonna.

(05:04):
So you get to you get tosurprise me.
What is it?

SPEAKER_01 (05:07):
Yeah, I I think uh yeah, I mean it's lots of stuff
obviously I could talk about uhrelated to Warp Stream.

SPEAKER_00 (05:15):
Uh we might we might have you on the show again, just
so you know.
Yeah, exactly.

SPEAKER_01 (05:19):
Um but one of the things I thought uh might be
kind of interesting was um I wasactually my job before Warp
Stream.
Yeah.
Um which I was I was working atDatadog.
Um and you know, we basicbasically um you know Datadog
ingests and stores all your logsand events and and time series
data.
Um and I was part of the teamthat was building the new

(05:41):
storage system, basically, toreplace the existing one um for
a lot of that data.
Um and you know, building thatsystem was like super fun and we
we we did a bunch of cool stuff,but like the the part that was
actually like really hard andkind of a grind was like, okay,
how do we get this actually outinto production now for like a
hundred percent of products andcustomers?

(06:02):
You know, Datadog has like tensof thousands of customers and
dozens of products that are allpowered by like one database.

SPEAKER_00 (06:09):
Um you gotta imagine Let me actually I want to ask
you about that.
Like, what what was wrong withthe world of Datadog that they
wanted to build the new system?
What what did it and and I'lltake like the business
perspective, I'll take thedeveloper perspective, hands-on.
Like, what was it?

SPEAKER_01 (06:26):
Yeah, so the the business perspective is I think
super easy.
Uh one, it was like just old,like insanely expensive, what
they were doing.
There you go.
Um, which is for all the samereasons, really, that um I'll
slip warp stream in.
Uh, which is that uh, you know,running kind of open source
Kafka yourself can be reallyexpensive.
Networking, storage,auto-scaling, all that type of
stuff.
So a lot of the same impetus forWarp Stream, uh, but this was

(06:48):
before before Warp Stream.
Um, but similar problems.
Uh and so that was kind of likethe cost was an obvious one, but
actually a lot of it too waslike there was a lot of stuff
that customers wanted in termsof like features uh that we just
couldn't deliver with theexisting system.
Um so like one of those, um Idon't know how obvious this will
be to someone who's not like adatadog user, but like if you

(07:09):
imagine a logging UI, right?
Like there's a little bar at thetop, and you're like you filter
on things.
And in the older versions ofDatadog, you used to have to say
beforehand, these are fieldsthat are interesting to me and I
would like to be able to filteron.

SPEAKER_00 (07:23):
Before ingest.

SPEAKER_01 (07:24):
Uh correct.
Yeah.
Oh, okay.
Good.

SPEAKER_00 (07:27):
As long as you know everything ahead of time, it's
fine.

SPEAKER_01 (07:28):
Yeah, which is great.
So it's like you'd be in anincident and you're like, oh, I
need to filter on this field,and we're like, uh no, you
didn't tell us beforehand.
And then you could startindexing it in the moment, but
it would only apply to new datacoming in, you know, several
minutes later and not toanything in the past.
So that's obviously extremelyfrustrating.
Yeah.
Um but if you think about it,like about without going to too
many details into how the oldsystem works, if you imagine

(07:50):
like a traditional searchsystem, usually there's like an
index, you like a schema for theindex, right?
You're like, well, these are thefields that are important.
But like every customer's logslook different, so that doesn't
really work for like uh anobservability use case where the
customer doesn't like give youthe schema.
Um so that was just like oneexample of a feature um that we
couldn't really implement in theold system that the new one
could.
Um and the other was like likeuh just query performance really

(08:14):
like you know, the old systemworked fine if you just need to
query a couple hours of data,but what happened is someone
wants to query like three monthsof data.
Nope.
Um and they're a relatively lowvolume customer, so they they
were on like one shard overhere.
And a shard is like physicallytied to like you know a machine
with a fixed number of cores.
And so if we only had your dataon this machine, but you wanted

(08:36):
to run some massive query, therewas no way for us to parallelize
it.
Um beyond what we had decidedthree months ago how many
machines yet you deserved,basically.
Gotcha.
Uh so kind of um that sort ofthing.

SPEAKER_00 (08:48):
So which is a very much first generation
distributed data architecturekind of choice.
That's common.

SPEAKER_01 (08:54):
Yeah, yeah.
Um so those are the kind of thereasons that we were building
something new.
We wanted something cheaper,something more cloud native,
easier to scale, easier tomanage, but also something that
would allow us to kind of buildsome of these new features and
something where like, hey, evenif you're a tiny customer, if I
want to throw a thousandcomputers at your query for a
second, you know, I can.

SPEAKER_00 (09:11):
Okay, so thank you.
That's that's that's the whatwas wrong, and that makes a lot
of sense.
And you were starting to diveinto like a particular storage
problem.
Uh keep going.
Now a quick word from oursponsor.
Confluent Developer the Podcastis brought to you by Confluent
Developer, the website, whichhas everything you need as a
developer of data streamingsystems.

(09:31):
And it's completely free.
We've got curriculum, hands-onexercises, executable tutorials,
the online data streamingengineer certification, also
free.
A way to find a meetup near you,those are free.
Everything is there.
I really want you to besuccessful in your journey as a
data streaming engineer, andthis is the site that has what
you need.

(09:51):
Check it out atdeveloper.confluent.io.
That's developer.confluent.io.
Now back to the show.

SPEAKER_01 (10:00):
Oh, yeah, it was just the general problem I would
say of like getting out toproduction.
Because it like, you know, abouta year in, we were like, okay,
we've built a pretty goodsystem.
Like it's not perfect, but it'sdecent.
Like it can do some pretty coolstuff.
Uh, but like the reality of likehot swapping the database for
like 30,000 customers and 13distinct products with different

(10:21):
features, different querypatterns, like the reality of
doing that um without breakingeverything was was hard.
Um and um so there were kind oflike a number, and we what we
ended up doing is we justtackled it product by product.
So we pick like the simplestproduct we could, and we'd be
like, okay, like let's justfocus on this one.

(10:42):
And and because you gottaimagine too, if you if you mess
up, like let's say filteringbecause Nog language query
language is super weird becauselike so imagine you're like um
duration larger than 200 is likea filter you put into your logs,
right?
Okay.
Well, some customers may beemitting the duration field as a
string, and some other peoplemay be emitting it as a float,

(11:05):
and someone else might beimplementing it as an integer,
and some customers may be mixingand matching those from like the
same service.
And so the query semantics getfunky.
And if you mess up the querysemantics, someone might get
paged in the middle of the nightfor something they shouldn't
have been.
Um, or worse.

SPEAKER_00 (11:20):
Or not.

SPEAKER_01 (11:20):
Yeah, the opposite, yeah, exactly.
Um so you have to be verycareful and do all this kind of
query shadowing stuff.
Uh and I I remember one of thethings that like really killed
me was um you know, we were kindof getting ready to migrate our
first project, and thisrequirement that I like had been
vaguely familiar with but didn'treally understand the the the
real implications of, which isthat in Datadog, they'll only

(11:43):
they they had the storage systembasically ensures that you only
ingest each piece of data once.
Like every log, everything thatwe ingest, like a JSON record,
has an ID.
Okay.
And the storage system iscapable of making sure that it
only ingests it one time.
Okay.
Um, which is also superimportant because like you can
imagine a lot of people havemonitors that are like, if this

(12:03):
happens more than three times infive minutes, page me.
And it's like, well, if someintermediary service in
Didadog's processing pipelinerestarts and replays a couple of
messages, and then those end upin storage twice instead of
once, that will like ruinsomeone's entire day, basically,
right?
Okay.
Okay.
And so it's extremely importantthat like if you send us

(12:23):
something once, we store itonce, and it shows up once in
the query.
Um, and that, you know, if youimagine kind of like a database
uh that's on SSDs and sharded,uh it's relatively
straightforward to implementthat feature.
It's not easy, but like you cando it.
You're like, okay, well, thisrecord goes to this node on this
shard, and that thing can kindof check because it's you know
it's a KV store or whatever itis.

(12:45):
Um but when you've built like akind of columnar store on top of
object storage, um like hey,does this ID already exist in
the system?
Becomes like a very hard problemum to answer.
Um and so we we built a thing todo this, and as we were getting
ready to kind of like migrateour first product, we realized

(13:06):
it just like kind of didn'twork.
Like it worked 99.9% of thetime, but that wasn't like good
enough.

SPEAKER_00 (13:12):
There were some cases that that were there,
yeah.

SPEAKER_01 (13:15):
Yeah, I think at one point we had it up to like we
measured it.
We're like, okay, we're at likefive or six nines of
deduplication, which like seemslike a lot, but is still like
not enough.
Sure, sure.

SPEAKER_00 (13:25):
When you're looking at volumes like you're
ingesting, there's still asignificant number of
duplicates.

SPEAKER_01 (13:29):
Yeah, like someone's day is going to be ruined every
day, uh, if we at that at thatlevel of nines.
Um and so it was like me andanother engineer, we were
basically like all of the workwe did building the system is
for nothing until we solve thislike to like, you know, yeah,
29s or whatever, basically.

(13:49):
Um and so it took us about threemonths of basically just
grinding on that feature um toget it working.
Um yeah.

SPEAKER_00 (13:59):
How how did you find that you know you started by
writing something that I'm gonnaguess you believed did
deduplication and then you findout you're wrong.
How did you find out you werewrong at the beginning of that
three months?

SPEAKER_01 (14:16):
Yeah, I'm I'm trying to remember how we noticed at
first.
Um well in the beginning itdidn't work well enough that
like people would like noticeand complain.
Like our internal c like thedatadoc things, like not before
we migrated customers, we'dmigrate datadog internal
products.
Okay.
And there were like parts of theproduct that would just like
break, basically, if thishappened.

(14:37):
Okay.
Um and then when we realized theproblem was more complicated
than we had kind of anticipated,I think what we ended up doing
was writing a bunch of toolingto detect it.
Um so we had a couple of tricks.
Um, one, we could detect likefor certain types of queries, it
was easy to detect that like,hey, there's two duplicate
events in this result.
So we did that.

(14:58):
And then um we had thiscompaction system that was
constantly merging data in thebackground.
Um and so we and what that wouldthe that system was designed to
bring similar data closertogether, basically, to make
queries faster.
And so we added a bunch of codein there too to detect um

(15:19):
essentially like, hey, whileyou're bringing similar data
closer together, kind of lookaround yourself to see, hey, are
there duplicate kind of do likea sliding window scan during the
compaction to see if you haveduplicate events in there.
Okay.
And and I think we then builtsome other tools that would
essentially go and download abunch of files, analyze them,
look for duplicate events, andthen emit logs and stuff.

(15:40):
Um so we did we built a bunch oftooling basically to kind of
passively scan in the backgroundto see if we messed anything up.
Um and then we would go andmanually cause this to happen to
see if our tooling detected it.
Right.
Um so we we just had to build abunch of tooling.
Um and we probably should havedone that.
Um this probably should havebeen the first thing we did.
I think we just were kind oflike in a rush and didn't build

(16:02):
enough tooling uh in thebeginning.

SPEAKER_00 (16:04):
Uh you usually usually believe you're doing it
right.
You know, you think you have asolution that works until you
bump up against reality.

SPEAKER_01 (16:12):
Yeah, exactly.

SPEAKER_00 (16:13):
Uh what was that tooling the final answer uh that
these these tools run and anddetect duplicates, or did that
cause then you to go back intothe deduplicate?
I mean, did you did youeffectively solve deduplication
on the input?

SPEAKER_01 (16:31):
The tooling gave us like a benchmark, basically
like, okay, until these go away,like we're not done.
And even then we may still notbe done, but we'll be like much
closer to done.
Uh because I mean it would likeonce we started building smarter
tooling, we could there weretons of operations we could do,
and we'd be like, oh shit, itfired because you know we scaled
this up, scaled that down,killed this node at the wrong
time, or we would inject faultsand see that was happening.

(16:52):
So I would say that's actuallywhen the work began, um, is when
we started having those tools.
And then it was like literallylike, okay, um well, and part of
what made this hard too is itwasn't just localized to our
system.
Like our system, in order forour thing to function correctly,
it depended on services sittingupstream of us.
Because essentially what thesystem did was like, for

(17:12):
example, if two duplicate eventswere ever sent to um we had
these things called shards, anda shard was essentially a group
of Kafka partitions.
So if anyone ever sent um yousent an event with this ID to
this shard, and then later yousent a duplicate event to
another shard, our deduplicationsystem just did not work.
So all of our upstream thingsalso had to be functioning as

(17:34):
well, which ended up being coolbecause eventually we got our
thing working perfectly, andlike a couple months later it
fired, and it turned out we hadcaught a bug in something
upstream of us, basically.
Um so our thing, it was ourscanners, because they
essentially just ran passivelyin production, were able to
catch issues that arose evenoutside of our own system, um,

(17:54):
which was cool.
Uh, but that's what kind of whatmade the problem so hard in the
beginning is like there's somany moving pieces, and you
like, you know, it's notworking, but where did it go
wrong?
That's very hard, right?
Um so I think we we um someoneended up writing a uh uh TLA
verification of the protocol wewere trying to implement just to
make sure, like, is there justlike a logical fallacy in the

(18:17):
algorithm we're implementing?

SPEAKER_00 (18:18):
And if uh anybody listening doesn't know TN TLA,
tell us briefly about that.

SPEAKER_01 (18:22):
Uh man, I'm not the right person to explain these
things, but it's a it's a uhdon't quote me on my
explanation, but it's a it'sessentially a um uh under the
category of formal methods, andwhat it does is allow you to
express distributed systems andconcurrent algorithms in like a
formal language, um, and then uhessentially run it through a

(18:45):
simulator.
Um, and the simulator willessentially like explore the
solution space and tell you ifany of the like invariants you
said must always be true or everviolated.

SPEAKER_00 (18:56):
Right.
And it's pretty common, and I amblanking on what TLA stands for.
We'll look it up, we'll put itin the show notes.

SPEAKER_01 (19:01):
Yeah.

SPEAKER_00 (19:01):
Uh but I I think it's common experience when you
know you've built your system,you're like, hey, this works,
let's do TLA.
Oh, it doesn't work.
It's it's it's usually humbling,right?

SPEAKER_01 (19:12):
Yeah, and the the thing that's really cool about
TLA plus, I haven't done it aton.
I've done it like two or threetimes in my career.
Um, but it's like you know whatpeople say like writing, like
you don't underst you don'tunderstand something or you
haven't thought somethingthrough until you've written it
down because writing isthinking.
Yeah, I think TLA plus takesthat to like the logical
extreme.

(19:33):
There you go.
For computer algorithms, it'slike you don't really understand
a distributed algorithm untilyou've implemented it in TLA
plus, and then you really knowlike it's almost like I often
will find bugs, not because TLAplus like the simulator found it
for me.
It's because while in theprocess of trying to translate
my thought into a TLA plus spec,you're just like, oh, that's

(19:53):
that's just like wrong.
Like you'll see it while you'retrying to implement it.

SPEAKER_00 (19:57):
The discipline of translating it into the formal
specification forces you togrind through.
Yep, just like writing or or uheven explaining verbally, you
know, if you if you can't dothat in a simple way, um it's
yeah, I totally agree.
That's uh without had havingused TLA to verify any
distributed algorithms I've wI've written, the whole the

(20:19):
whole concept uh that makessense.
I mean you were kind of you youhad uh deduplication solved and
you said you had caught a bug.
It had been running okay for afew months, and your monitors
caught a bug in some otherupstream service, which is
really cool.
What was um this could rabbithole, and we don't need to do

(20:39):
that.
So kind of in brief, the youdescribed the the problem at the
beginning.
When you you write something tostorage, you need to make sure
you're only writing it once.
And uh uh there are two things,I suppose, two categories of
things that could be hard aboutthat.
One is the definition ofsomething, like the identity
identity of the thing.

(21:00):
Uh I'm gonna guess that wassolved.
There were some unique IDssomewhere, that was all okay.
Uh, and the hard part is raceconditions on the writing or the
reading.
Um what was the storage?
Uh was it a blob store?
What was it disks?
Uh was it something you guys hadbuilt?

SPEAKER_01 (21:19):
So um this was another thing that was
interesting, which was like wedecided, okay, so we'd built
this completely statelessingestion thing, like you know,
our compaction servicestateless, our ingestion service
was stateless, our queryservices were stateless.
Go and delete nuke or whatever.
Yeah.

SPEAKER_00 (21:36):
In your architectural preferences here.
Yeah.

SPEAKER_01 (21:39):
I I like to sleep at night, you know?
Yeah.
Um and we um so we built thiscompletely stateless system, and
you could go delete anything atany time and you would never
lose any data.
And then when we were designingthis deduplication system, we
were like, we can make this workwith no disks, too, and no state
and any of this stuff.
And I got a lot Not just me.

(22:00):
We got a lot of flack for thatbecause I think there were some
other people at the company thatwanted us to just like
essentially stick Rocksdb onsome nodes and do the
deduplication there.
As as one does.
And have it be pretty stateful.
And that would have made theproblem very maybe not very
easy, but let's say much easier.
Um but we really didn't want todo that because we were like,
man, we got like we literallybuilt this whole thing, and now

(22:23):
we're at the final step, and wereally don't want to just shove
disks in here at the very lastmoment.
And then when it was not workingand we were grinding on it,
trying to make it work, I waslike, man, I might have to eat
some crow.
Like this is not working.
It should be able to work, butit's not working.
And what was interesting is thatwe actually were using the
system we built to store if ifyou think about it, like the
deduplication thing essentiallyboiled down to um it was

(22:49):
essentially essentially anotherdatabase um that needed to track
IDs and be able to commitatomically with the other data
file ingestion.
Yeah.
And we were like, well, if westart creating files that just
have IDs, we can organize themthe way we want so we can query
them quickly.
But then we'll have a lot offiles and we'll need to compact
them, and then when the filesget old, we'll need to expire

(23:09):
them.
And we're like, we're gonna haveto rebuild half of this
database.
And so what we ended up doingwas storing them in the system
itself that they were alsoserving as the deduplication
layer.
So they were essentially specialtables in Husky, and they would
go through their own compaction,which sounds like meta and weird
and like it wouldn't work, butit did end up working and saving
us a lot of work.
Um, because then it's like,well, you get compaction for

(23:31):
free, and you get data experienfor free, and you get a file
format for free.
Um so that's how we did it.
And the the IDs did have theywere very specifically shaped,
like they had time in them andsome like tenant information and
whatever, which allowed us tocompress them like crazy in
memory.
Because what what we didn't endup any doing any discs, but we

(23:51):
did end up doing is like havingessentially a a hot set that was
in memory that we're that wecould verify against very
quickly.
Okay.
Um and you know, we wouldessentially page them in and out
of memory um very quickly.
Um sorry, I feel like I forgotwhat the original question you
were asking was.

SPEAKER_00 (24:09):
Um what what was hard about the like I basically
like uh I think the most recentquestion was were there disks or
not?
And you're saying ultimately no?

SPEAKER_01 (24:20):
No, there were no.
We did managed to make it workwith no disks at all, um built
on top of the system that it wassupposed to be essentially
deduplicating for.
Um which ended up being um andit was nice because it ended up
being a really simple once weironed out the kinks, it was
very stable, and you just kindof like you could add new nodes
and they could start processingdata, and they would load

(24:42):
whatever they needed fordeduplication into memory.
And if they died, work would gettransferred over and you didn't
really have to think about ittoo much.
But there was definitely awindow where I was like, this
may not work, like maybe we domore than we can chew.
Um and we just kind ofeventually we ground through it.
But I I was definitely worriedfor that.
Was probably the most in mycareer of like, mm, this may
never work type of thing.

(25:02):
Um because we were working, youknow, it was like three months
of like probably six to sevendays a week, 10 to 12 hour days,
just trying to get this thing.
Because it was blockingeverything, you know what I
mean?
Like until this was until likethis graph we had essentially
went to zero.
Uh you couldn't get any valuefrom any of the system we'd
built, yeah.

SPEAKER_00 (25:19):
Right, right.
Of course.
That makes sense.
And that value after you shippedit, um, you know, beforehand,
you had to define indexes inadvance.
Uh it was really expensive, itwas slow, uh, those things got
better.
I mean, that is that that'syeah, yeah, way better.

SPEAKER_01 (25:36):
So, like, you know, product owners could be like,
you know, I would basically tellthem the difference between one
day of data retention and twoweeks of data retention is
almost nothing.
Um, and so cust uh productscould have higher retention by
default.
Um, you didn't have to pre-indexyour field.
So if you're in an incident, yousaw something weird, you needed
to go look at some historicaldata during the incident, you

(25:58):
just could.
Uh, and that started working.
We were able to add morepowerful auto-complete.
Uh and obviously the the companysaved just you know a ton of
money not running this superexpensive um you know, kind of
SSD-based system for what isessentially the lowest value per
byte data on the planet, right?
Which is like you know, randommetrics and logs coming out of

(26:19):
your software, right?

SPEAKER_00 (26:20):
Yep, the exhaust.
And uh what year was this?

SPEAKER_01 (26:24):
Uh that's a good question.
What year uh was how old is WarpStream?
Warp Stream's like two and ahalf years old.
Uh so I guess this would havebeen like five years ago.

SPEAKER_00 (26:35):
And and yeah, like Warp Stream's, I guess that's
really the the thing.
Warp stream's two and a half.
Um and now this is late 2025when we're recording this.
Um Diskless whatever is the coolthing.
You know, if if there's a kindof data infrastructure, someone
is working, and probably two orthree someones are working on a
diskless version of it, yeah.

(26:56):
Um search various kinds ofdatabases, uh certainly,
certainly uh distributed logs.
Um that wasn't true then.
That was a you know, you try todo this all diskless.
That that was a much more muchriskier choice for you and
direction for you.
And it's it's funny, your storykind of pivots on that that all

(27:21):
is lost moment of oh crap, maybeI did the wrong thing and we
have to go back to doing thisthe lame way.
Um and you didn't, you pulled itout, but that's uh these days,
like if that's happening in in2025, early 2026, oh you want to
do some discourse, big deal.
You know, everybody does that.

(27:42):
It's hard, but it's it'spopular.
It just wasn't then, and I Ijust I'm pointing out that that
was uh um an admirable andslightly revolutionary approach.

SPEAKER_01 (27:52):
Well that and that was I think less uh I don't know
what word to use, unhingedmaybe, than uh than when
Warfream actually, because likeI think at that time, you know,
we were kind of following in thefootsteps of like like you know,
the Snowflake paper waspublished.
People knew that buildinganalytical databases and
columnar stores on top of objectstorage was possible.

(28:12):
Um people thought maybe, oh,it'll have to be like a cold
store.
Because that's how the projectstarted.
People were like, well, you'regonna build like the cold tier
or some special really slowversion of the logs product
that's that's cheaper orwhatever.
And then like six months intothe project, like our
engineering, one of ourengineering leaders was like,
no, you need to make this workfor every product and to get rid

(28:33):
of the old system to power allthe real-time data as well.
And I was like, Ah man, that'sgonna be hard.
Um, because there's you know,we'd already set some latency
expectations with the the theold product.
Um and I think warp stream waseven more extreme that like
people, you know, the the waypeople used to describe it was
an aggressive architecturaldecision or whatever.

SPEAKER_00 (28:53):
I've probably used those words myself.

SPEAKER_01 (28:55):
Yeah, because it's you think of it being as a much
more you can kind of imagine,okay, like if a query takes a
second to run to answer a human,that seems more reasonable than
like a computer waiting 500milliseconds to make sure
something is durable orwhatever.
Right.
Um but I don't know.
I think if you're working onstuff that's like really hard or
bleeding edge, at some point youwill find yourself questioning,

(29:17):
like you will have a at somepoint you'll have a come to
Jesus moment where you're like,I don't know if this is ever
going to work.
Exactly.

SPEAKER_00 (29:24):
Am I the stupid one?

unknown (29:26):
Yeah.

SPEAKER_01 (29:27):
Like I I've had that moment in pretty much every
major product I've project I'veever worked on.
Yeah.
Just because it's like if you'redoing something ambitious, you
will eventually reach a pointwhere you're like, I don't know
if it's gonna work.
Like it the the napkin math saidit would, but I'm not getting
the results yet.

SPEAKER_00 (29:41):
You know, uh as I like to say, it's hard to make
things, and if you're notstruggling a little bit while
you're making them, you'reyou're you know, maybe not
making something that is is allthat interesting.
Yeah, I agree with that.
My guest today has been RichieArtoul.
Richie, thanks so much for beinga part of the Confluent
Developer Podcast.

SPEAKER_01 (30:00):
Thanks for having me, too.
Advertise With Us

Popular Podcasts

Betrayal Weekly

Betrayal Weekly

Betrayal Weekly is back for a new season. Every Thursday, Betrayal Weekly shares first-hand accounts of broken trust, shocking deceptions, and the trail of destruction they leave behind. Hosted by Andrea Gunning, this weekly ongoing series digs into real-life stories of betrayal and the aftermath. From stories of double lives to dark discoveries, these are cautionary tales and accounts of resilience against all odds. From the producers of the critically acclaimed Betrayal series, Betrayal Weekly drops new episodes every Thursday. If you would like to share your story, you can reach out to the Betrayal Team by emailing them at betrayalpod@gmail.com and follow us on Instagram at @betrayalpod and @glasspodcasts. Please join our Substack for additional exclusive content, curated book recommendations, and community discussions. Sign up FREE by clicking this link Beyond Betrayal Substack. Join our community dedicated to truth, resilience, and healing. Your voice matters! Be a part of our Betrayal journey on Substack.

Stuff You Should Know

Stuff You Should Know

If you've ever wanted to know about champagne, satanism, the Stonewall Uprising, chaos theory, LSD, El Nino, true crime and Rosa Parks, then look no further. Josh and Chuck have you covered.

Dateline NBC

Dateline NBC

Current and classic episodes, featuring compelling true-crime mysteries, powerful documentaries and in-depth investigations. Follow now to get the latest episodes of Dateline NBC completely free, or subscribe to Dateline Premium for ad-free listening and exclusive bonus content: DatelinePremium.com

Music, radio and podcasts, all free. Listen online or download the iHeart App.

Connect

© 2026 iHeartMedia, Inc.

  • Help
  • Privacy Policy
  • Terms of Use
  • AdChoicesAd Choices