Episode Transcript
Available transcripts are automatically generated. Complete accuracy is not guaranteed.
Narrator (00:01):
Welcome to the
Practical AI Podcast, where we
break down the real worldapplications of artificial
intelligence and how it'sshaping the way we live, work,
and create. Our goal is to helpmake AI technology practical,
productive, and accessible toeveryone. Whether you're a
developer, business leader, orjust curious about the tech
(00:22):
behind the buzz, you're in theright place. Be sure to connect
with us on LinkedIn, X, or BlueSky to stay up to date with
episode drops, behind the scenescontent, and AI insights. You
can learn more atpracticalai.fm.
Now onto the show.
Dan (00:41):
Welcome to another episode
of the Practical AI Podcast.
This is Daniel Whitenack. I amthe CEO at Prediction Guard, and
I'm joined as always by mycohost, Chris Benson, who is a
principal AI and autonomyresearch engineer. How are you
doing, Chris?
Chris (00:56):
Hey. I'm doing great.
Can't wait to get into today's
conversation. It's gonna be fun.
Dan (01:01):
Yes. Yes. For for an audio
podcast, we're gonna talk about
a lot of interesting visualthings. Maybe before we get
started, just just a littleteaser. Practical AI is posting
some videos on YouTube now.
So if you do consume podcaststhat way, you might go check us
(01:22):
out on our on our YouTube page.But speaking of images, videos,
and more specifically, imagegeneration, really excited to
have with us today DustinPodell, who is cofounder and
researcher at Black Forest Labs.Welcome, Dustin.
Dustin (01:38):
Yeah. Thanks for having
me, guys. It's, really great to
be here.
Dan (01:41):
Yeah. And and I know that
Black Forest Labs does more than
just kind of, raw imagegeneration. There's a lot of
workflow related things,hardware optimization, all sorts
of cool stuff you're involvedwith. But as we get into some of
that, I'm wondering if you canjust help our audience with a
bit of a state of imagegeneration methods and
(02:04):
workflows, for the industry.We've talked on the show before
about diffusion models, andwe'll link some of those
episodes in the show notesmaybe.
But a lot has happened. Right?There's a lot of people working
on a lot of interesting things.And, I'd love to kinda
understand, like, over the lastyear, what are some of those
main main points that might begood for people to orient
(02:27):
themselves to where things areat now?
Dustin (02:29):
Yeah. Yeah. No. It's a
it's a good question. I mean, if
I'm allowed to take it even alittle bit further back, mean,
state of
Chris (02:35):
Absolutely.
Dustin (02:35):
Yeah. Yeah. The state of
the state of image gen, video
gen, generative models in in asa whole has kind of gone crazy,
so to speak, in the last, like,three or four years.
Chris (02:47):
Yeah.
Dustin (02:48):
So where are we now?
Where we came from about four
years ago, where I first kind ofentered into the scene, so to
speak, is we were at Models thatwere essentially just doing
little blobs of color that werekinda related a little bit to
where you were with the prompt,you know, okay, oh, a lighthouse
(03:08):
on the beach or this and that,and you would get something
that, okay, that vaguely lookslike it, and you would chose
someone in the beach. Yeah, Icould see it, I guess, and maybe
it would interest a few nerdypeople, and then now today,
we're at the point where, ifyou've probably been seeing
plenty of this stuff online,where we're seeing whole short
(03:29):
films made entirely with AIgeneration where certain scenes
are almost entirelyindistinguishable from from
reality, so to speak. So I wouldsay we we've we've come quite
far, but I will also say thecore of the technology hasn't
actually changed that much. It'sbeen a pretty nice, like,
forward progress.
I don't wanna, like, diveimmediately into anything
(03:50):
technical here, but it's it'scertainly I mean, for anyone
who's been paying attention, I'msure or or or anyone who really
hasn't been paying attention,this probably came a bit out of
out of nowhere, so, you know.
Chris (03:59):
Yeah. Was gonna say, I
guess, you know, with the
broadly attention by the generalpublic is so much on kind of
more of the LLM generative worldin terms of, you know, that that
stuff, and everyone's finally onapps regardless of whether
they're technical or not. And soI think a lot of people kind of
(04:19):
miss the tremendous advancementsyou guys are making on that
side. And so, like, you know,could could you could you take a
quick moment, and now thatyou've kinda done the highest
level, maybe step through someof the things that people may
remember, like we talked aboutstable diffusion and things, a
couple of those things that havekinda led up to what we're gonna
dive into today with some morespecifics. And and as Daniel
(04:42):
mentioned, we can we can offersome links for past
conversations if people wannadive into those specifically,
but that would that would reallybe interesting to kinda hear
distinct points on that timeline
Dustin (05:39):
in 2017, some
researchers at Google figured
out this very cool thing calledthe transformer, which I'm sure
people have have heard many,many times about right now,
which These are what you wouldcall an autoregressive language
(06:40):
model. And so what this is isit's predicting data essentially
one piece at a time. So itpredicts one piece of data, and
then it looks,
Chris (07:33):
I love, yeah, I feel like
I'm watching a thriller where
you're about to reveal the nextthing. It's good.
Dustin (07:40):
Yeah, so what we do with
the Fusion is, I'm just gonna
tell practically what we do, andthen we can kinda break down.
You could ask me questions.
Chris (07:46):
That's fine.
Dustin (07:48):
So what we do with the
Fusion is we take a whole
continuous medium, and what Imean by continuous medium is the
most famous one is images. Theworld is continuous. There's not
you know, it's not that there's,you know, one you know, like a
letter is discrete. There's a T,there's an A, a B, a C, a D, an
E, and an F. A color, shape, thestructure of the world, the
(08:09):
color, the light, all of this,these are continuous natural
mediums.
And so what
Chris (08:17):
we
Dustin (08:17):
do is we take one of
these continuous natural mediums
that that we've now encoded intodigital space, so we use RGB.
So, you know, okay, 256 colorswe typically use for red, green,
and blue, and then we encodethis into an image. And then
what we do is instead of tryingto predict so an autoregressive
view might try to predict likepixel by pixel by pixel, instead
(08:38):
what we do is we essentially tryto remove information, and the
way we remove that informationis by adding noise. So you
imagine you like add a littlebit of grain on top of an image,
you can still kinda see theimage, maybe it's a little
grainy, maybe it's a littleblurry, oh, you can't quite see
it, and you add a little bitmore, maybe you see some of the
shapes, maybe, okay, there'slike, is that a dog? Yeah, you
kinda see it, the color startsto fade and then you keep adding
(09:00):
and eventually you just get,okay, this is just noise.
I don't see anything in thisimage at all. And with the
diffusion model, what we'retrying to do is we're
essentially adding a bit ofnoise and then we're saying, now
try to predict what the cleanimage is essentially for this.
So, you know, okay, so we add alittle bit of noise, predict the
clean, it just has to remove alittle bit of noise. You add a
(09:21):
lot of noise, and now what ithas to do is there's so little
information, it has to actuallylike come up with essentially
like how to infill thisproperly. And then what you do
is then in in what we callinference, which is when you
actually run the model, you askfor an image.
What you would do is then youstart a fully noisy image and
you say, give me a dog, youknow, wearing a top hat on the
(09:42):
beach. And then essentially whatit does is it tries to remove a
little bit of that noise to getcloser to this idea of a dog in
a top hat at a beach. And youmight get some structure, it
tries to figure it out, it'svery coarse, it's very just
trying to figure out the shapeand then it has a little bit of
shape to latch onto and then itdoes the next step where it
removes a little bit more ofthat. Oh, and it starts to come
into view. It starts to youstart to get a little bit more
in focus, and you do this do dodo do do do do and what you see
(10:05):
is essentially you completelyremove the noise of this image
and oh my goodness, now you havethis this completely generated
image of a dog on the beach witha top hat.
For a long time, this was, Idon't wanna say not good. I
mean, all models have gotten alot better, but when it first
started out, these looked likeblobs of color. It was like
these little, like you couldbarely see it. But
fundamentally, what I justdescribed is still the process
(10:27):
that's been happening over thelast four years from these old
image generator models or anyonethat maybe tried like DALL mini,
I mean, the old stable diffusionmodels that we worked on, you
know, I don't know, many of thedifferent models are out there
now to these modern video modelsthat are producing whole ten
second sequences, fundamentally,it's doing the exact same
process of removing thisinformation slowly with this
(10:50):
noise, and then just slowlygoing the other direction,
essentially creating info toinfill if this makes sense.
Pause there
Dan (10:58):
for a second just to kinda
like that.
Chris (11:00):
Let me tell you, because
as I'm listening, very, very
good explanation, probably thebest one that I've heard, so you
are right on target if you justcarry on what you're doing,
because I'm I'm enjoying this.
Dan (11:13):
And and you mentioned that
when this started, right, you
sort of ended with these blobs,etcetera. I think there's a lot
of people that have seen veryimpressive things like the, like
you were saying, like wholeparts of maybe movies or
commercials or something
Dustin (11:30):
Mhmm.
Dan (11:30):
Being generated in this
way. What what is the current, I
guess, state of the art in termsof, whether it be your your
models or others? And I knowthere's more to talk through as
you get in. Like, there's morethan just raw image generation.
There's how this fits intoworkflows and and doing certain
tasks.
But what is kind of the state ofthe art in terms of what we're
(11:51):
able to generate? Where are thewhere are the bounds currently
and in terms of quality and thatsort of thing?
Dustin (11:56):
Sure. So I'll I'll tell
you kinda like in the last year,
I'll I'll split it up into,let's call it three different
categories, but they kinda blenda little bit, which is these
continuous mediums I talkedabout. So I gave the example of
a continuous medium with image.Video is obviously an extension
of that, adding the the timedimension. But then there's also
audio as well, which now a lotof these video models generate
(12:16):
audio, but then there's alsothings like, I don't know, maybe
you've heard the the like, Sunoor these music models.
Some of these are also usingthis exact same technique of
removing the information withnoise, bringing it back in, it's
just a different medium thatthey're essentially applying to.
I'm happy to talk about ourmodels all day. I love talking
about our models. But I alsowanna be fair to just anyone
(12:37):
listening and honest to kind ofthe state of the world on if you
wanna go out and try some nicethings out there as well. I will
say that before I kinda likedive a little bit into what I
think is maybe the best one torecommend, I will make like one
statement of that like bestquote unquote here is a little
(12:57):
bit hard to define sometimes,and this is something we're
always like wrestling with islike what makes a best model for
someone.
Because, you know, obviously youwhen someone uses these models,
they're typically prompting orputting in now you can reference
images or videos or other sortsof things and kind of treat it a
little bit like a more like alike a creative companion. And
(13:18):
then what you expect out of thiscan be very different for
different types of people. Thatsaid, I would say that if you
wanna see kind of the quoteunquote state of the art right
now, For what I would say islike text to video at the
moment, it would be Seedance. SoSeedance has done some extremely
impressive work with theSeedance two model over the over
(13:40):
the last year, so they're doingI think it's now four k
generations up to fifteenseconds, which is, I mean, very
nice quality. It's focused verymuch on like cinematics.
I will throw a shout out toanyone that, if anyone from the
Sora team ever listens to this,I think Sora two is quite a nice
model. It didn't rank so high onsome of these like leaderboards
(14:02):
that we In the AI community,there's a I'm sure you guys talk
about this plenty, likeleaderboards, rankings, where
models place. I think in the LLMside of things, this is much
better covered where there'svery clear metrics of how well
does it program, how well doesit do math. On the more creative
side of things with videomodels, image models, audio
models, we typically end upfalling down to this single
(14:22):
preference type benchmark, whichis just like a general
preference. Everyone in theworld goes and can can vote on,
you know, on these leaderhordes.
There's a couple different nicecompanies and sites that that do
this. But I I feel like thiskind of removes some of the the
specificity that the or thegranularity. That's the word I
was looking for.
Sponsor (14:43):
Listen. I've been to an
incredible amount of AI events,
many of which are good, but manyof which are not practical. You
know I love practicality. We'reon practical AI. And some are
just hype focused.
Some are sales focused. That'swhy I'm always eager to share
about an event that I trulythink is practical and useful
(15:04):
for people. That's what Idiscovered last year at the
Midwest AI Summit. And they'regonna have another Midwest AI
Summit October 15 inIndianapolis this year, 2026.
One of the reasons why I lovethis event was there was an
actual AI engineering loungewhere you could sit down and
talk through your use cases withactual experienced AI engineers
(15:26):
and practitioners to reallybrainstorm and come back from
the event with actual solutionsand practicality rather than
just a bunch of content andslides.
But there were also amazingkeynote speakers, speakers from,
even that had been on thepodcast before, like Rajeev
Shah, was at last year's event.I would really recommend that
(15:50):
you go to midwestaisummit.com.And, for our listeners, you can
actually get 20% off with thecode Practical AI 20. So go to
midwestaisummit.com. Don't missyour chance to attend this event
and get 20% off with the codePractical AI 20.
Chris (16:12):
I was gonna ask you, are
the models that you're
describing, are they kinda thetraditional diffusion models? Or
you had mentioned in passing afew minutes ago about flow
matching, and so it's just in myhead, I'm trying to kinda
categorize how do the ones thatyou're talking about right now,
how do they fit in? And and,like, and how what is to go back
(16:33):
and pull that up, as you weretalking about diffusion and you
made the reference to flowmatching
Dustin (16:38):
Mhmm.
Chris (16:38):
How do those what is that
what is that transition and how
do those models that you'retalking about now fit into that?
Dustin (16:44):
So I'll I'll I'll first
keep it very simple and say all
of these models are doing thesame process of what I described
earlier of adding noise andtraining the model to then
remove this noise. Thedifference kinda This is where
it gets a bit more technical inhow we actually approach this.
So we used to do more of thisprocess called diffusion, which
(17:08):
I I don't I don't I don'thonestly know if I wanna get
into how how deep and technicalthis is, but essentially what
we've done is we've we'vecleaned this process up to what
we call flow matching, and flowmatching is essentially this
very simple process of stilldoing the same thing, still
training the model to removethis noise, but fundamentally
what it's learning under thesurface is anyone out there who
(17:30):
has ever seen a velocity map ora flow map, what this
essentially looks like is if youcan envision a I almost wanna I
Dan (17:38):
don't know, Mark, you got
Chris (17:39):
the whiteboard behind it.
Don't know if want the
whiteboard. You're gonna
Dustin (17:41):
have to
Chris (17:41):
For the audio folks only.
They're not gonna see it, but
that
Dustin (17:44):
was great. Probably
about
Chris (17:46):
to turn to the
whiteboard. Yeah, it was great.
Dustin, if you're on video, yousaw it, but Dustin was literally
about to turn back to thewhiteboard behind him, which I
love. If I was in the room, I'dbe like, Go, man. Take me there.
Unfortunately, a good bit of theaudience won't be able to see
us. You're going to have todescribe it.
Dustin (18:02):
Yeah, yeah, no Let's
keep it purely in the audio
space. So I'll draw a picturewith my words as best as I can.
I spent enough time promptinganyway, so hopefully I can do
this.
Chris (18:13):
All good.
Dustin (18:14):
But yeah, essentially
the way I like to think about it
is if you imagine like alandscape, like I don't know,
you can imagine your town or theregion you're in, and imagine
you're looking at it from likethe sky and you see like your
house, maybe right in the centerof this map. And what you might
(18:35):
wanna do Now imagine, okay,there's wind going all over the
area, and the thing that youwanna do is you wanna be able to
throw like a paper airplane fromanywhere in this in this
landscape, your city, and youwant the wind to carry it so it
lands on your house. Andfunctionally, what we're trying
to train here is essentiallythat in a much, much grander
(18:56):
hyperdimensional space, whereinstead of it being your house
and instead of it being a windand a paper airplane, what it is
is your house in the scenario iswhat we would call the manifold
of real images. It's essentiallythe place in I'm trying not to
use too many stop me if I usetoo many words here. Wanna
Narrator (19:16):
say biggest.
Chris (19:17):
I'll I'll ask you, but
you're doing fine. Keep going.
We're fine on technical. We'lljust gonna explain it as we go.
Dustin (19:21):
Yeah. It's it's the
place in latent space that would
be where the actual real imagesimages are. So another way I'll
I'll I'll try to try to describein the simple terms, okay, with
your your town here, you're intwo d space. You're where you
are above an x y and the paperairplane has to fly in this x y
coordinate, you land and thenyou have the x y. In this space,
(19:44):
instead of it just being twocoordinates, it's enough to have
all of the colors of the entireimage, so that's why I say it's
hyperdimensional.
It's the center is whereliterally real images are. And
so what we're doing is insteadof it being wind and the rest of
your town, imagine the rest ofyour town here is every image
that isn't real. And what thatmeans is noise. If you think
(20:04):
about noise as a real image, youcan go generate a bunch of noise
and save it to your computer,but that noise would exist
somewhere in this space of allpossible images. They're not
real images.
They're they're noise, but theyyou can you can save them. You
know, they're they're they'reRGB values. You have them. And
so what we're doing with thisthis essentially this flow
matching is we're training themodel like, when we say we're
(20:25):
we're training it to removenoise, what's really happening
under the surface is we'retraining this flow map so that
we can land anywhere in thisfield of noise. And then these
flows, these winds in ourscenario, when you take a step,
will take you closer to yourhouse or to the manifold of real
images.
And that's what like as you getcloser, it starts to look more
(20:47):
and more like a real image. It'sit gets blurry. It has this. And
then when you actually finallyland on this this manifold,
boom, you have hopefully a realimage, but you'll have some
image. You know?
And then the better the model istrained, the better you have
mapped this flow to actuallytake you to where you want to
end up.
Dan (21:03):
And and I'm assuming yeah.
No. I and I'm assuming because
you are kind of mapping this,from from where you start to to
where you end up with this realimage, maybe in a way that's,
just to be crude, I guess, lessrandom. I mean, at at the end of
the day, all AI and machinelearning, like, it's sort of
(21:24):
that training process is verymuch trial and error, but we
have a lot of optimizationsaround it. Right?
So I'm I'm assuming that thisthen allows you maybe to, I
guess, advantage wise shortenthe the training or make that
more efficient? Or am I am Imisconstruing that in some way?
Dustin (21:44):
Yeah. Yeah. I mean,
definitely, like, one of the
things that's taken us a lotfurther over the last three to
four years is just figuring outa lot of optimizations both in
the actual training itself andthen also in just, you know,
architectures of the actualmodels have improved. But I will
say the like fundamentalunderlying process of just
(22:05):
removing this information andfinding a path back to it is
still the same process. It'sit's just like the car you know,
you've been having cars go downthe road since the nineteen
fifties, and we've made betterengines and better safety and
all this, but you're stilldriving down the road.
Like, that that's yeah.
Dan (22:20):
Yeah. That that makes
sense. And now, definitely, I I
think that's a good setting inthe sense that we've got to a
point where, you know, your thethe models coming out of Black
Forest Labs, the models comingout from other places, the
models that are even, like, nowin my text messaging. Right? I
get an image.
I can immediately remix it withan image model. Right? In in
(22:42):
some way. So these things arebecoming more embedded in our
lives, but I don't know ifeveryone in the audience, some
some might have been along forthis ride where I was like, oh,
cool. I can generate an image ofa astronaut riding a horse on,
you know, you know, wherever.
And that doesn't seem thatpractical to folks. So I'm
(23:04):
wondering if you could now kindof given that that foundation
that we have and we know sort ofwhere we're oriented in the
state of, I love how you put iton your website, Visual
Intelligence, which I think is,I I I love that statement
because it gets to more likelanguage, although it it's being
used a lot in terms of agentsand intelligence like you
(23:26):
mentioned, it is it is very mucha subset of the information that
we process as as humans. Right?There's this visual element,
there's the audio, etcetera. So,now that we have that
foundation, could you help theaudience understand some of, I
guess, the practicalities andthe outworkings of the, like,
okay, we can do this now.
(23:47):
So what? So how does that howdoes that help people in the
real world other in ways otherthan maybe just pure creativity?
There's certainly like the the,cinema, like you were saying,
that side of things. Noteveryone's gonna be generating
movies maybe, or maybe morepeople will, I I guess. But,
but, yeah, I I think you'reunderstanding what I'm saying.
(24:08):
Like, where where is this gonnaimpact where is it impacting me
now? Where is it practicallygoing in terms of the
application?
Dustin (24:15):
Yeah. Yeah. Abs I'm very
happy to talk about this. This
is a I think we're goingthrough, like, a very nice
transition period right nowwhere we finally get to leverage
these for some very usefulthings. And if I'm allowed to
take a little bit of a tangentand kinda Please.
Lead up to it again.
Chris (24:28):
I yeah. Wherever you
wanna go is good. We're all
good.
Dustin (24:30):
Yeah. Yeah. So, I mean,
I guess I'll take it back to,
like, a little bit of a timelineof things. So okay. So so early
on, we had these nice modelswhere most people recognize them
for you put in a prompt, you getan image out, you put in a
prompt, you get a video, maybeyou get a song.
It's just this one way okay, youmake something, this appeals to
maybe creators or ad agencies ormovie, know, now
(24:52):
cinematographers. And and Imean, we love this stuff. We
love the creative side of thisand and and what it allows
because it fundamentally to me,this is like a nice potential
communication tool. But then Ifeel like something changed. I
don't wanna say changed, butthere was definitely a timeline
moment when we started to moveinto editing.
(25:15):
So was it two years ago now? Ayear ago? I don't know, around a
year, or take, we released ourfirst in context editing model
called FluxContext. And on thesurface, would look at this and
go, okay, well, is this is animage editing model. This is
something you could take aphoto, you can clean it up, you
can add a hat, you can do sillythings, you know, you can
(25:37):
whatever you wanna do with it.
It's a general editing modelthat's supposed to supposed to
do all these interrelationalthings. And on the surface,
that's very cool. It's it seemslike another creative thing. But
if you think about, like, what'sactually going on under the
surface for the model to be ableto do this, it has to understand
a significant amount ofrelationships in the world and
what it means for these likelike how these relationships
actually interact together. Soif I say take a picture of a I
(26:00):
don't know, we have like a waterglass here on the table, and I
take a picture of this waterglass and I say to the model,
knock the water glass over andshow me what happens.
The model has to understand somepart of the actual world. Like,
it has to actually model theworld in some way to know, okay,
it spills over, maybe somethinggets wet, maybe x y z happens.
(26:23):
And essentially what we're we'retrying to do is train a model
that could do all of these typesof relationships. So so the
model is learning not just likethis one one thing, but just how
the world works fundamentally soyou can do this editing. Now we
can take this a step further andwe can look at all the video
models that are coming out rightnow and you can do a very
similar thing.
You can take an init image andsay, okay, this person now grabs
(26:46):
a fire extinguisher and puts outa fire and it has to actually
understand these relationshipsto do that. So, well, maybe this
wasn't our original goal wayback in the day. We were trying
to make very cool models andfigure out how to I I think in
some sense we wanted to modelthe world, but maybe we weren't
thinking this far ahead.Inherently through this whole
process, we've learned to buildthese models that are developing
(27:09):
this, and I'm really trying toavoid the use of the term world
model here.
Chris (27:13):
Because I'm gonna go
there if you don't. I'm just
telling you.
Dustin (27:17):
Yeah, you might notice
me trying to skirt this, because
I think it's a little overusedthese days, we can definitely
talk about it.
Chris (27:25):
But fundamentally, this
is too many people, which is
where I was going.
Dustin (27:28):
Yeah. Fundamentally,
this is what these models are
doing when you train them atscale and you and you really
train them to be general androbust is they need to be able
to simulate parts of the worldto get an output. And on on one
end, you can take that and go,okay. Well, this is nice for
creativity because you can makea film scene and it looks nice.
But on the other side, this iswhy we're starting to move in
this in this era, you know, thisthis area we're we're calling
(27:50):
visual intelligence, and nowwe're starting to see how we can
actually leverage what we'recalling not just us calling it,
this is this is a general fieldterm, but the representation
inside of this model of theworld to actually go do
practical things.
Now this isn't to say we'regonna drop the creativity side
of things, we're stilldefinitely pushing on the side.
This is something that we're allstill very fond of, but looking
(28:11):
forward, if our models areunderstanding these
relationships and the physics ofthe world like this, well, this
is a great place to put it insay robotics and actually, okay,
a robot has this understanding,this model of the world inside
of it to go act and take actionswith confidence in the world. So
(28:31):
I'll leave it there as kind oftrying to build up to it, but
Chris (28:34):
Let me ask. I'm just
gonna go there because that's
kinda you're in the area that Ispend all my time, which is, you
know, embodied intelligence,robotics, you know, UXVs and
stuff like that. That's myworld. And so in even within our
space, the notion of worldmodels has there's a lot of
interpretation. If you get intoa meeting with 20 people,
(28:56):
there's 20 differentdefinitions, and we have to
start off by sorting all thatout, you know, in terms of how
we're communicating.
As you add in this notion thereand of this kind of contextual
understanding, you know, that ithas a representation. I am
curious before we move on. Doyou how do you do you and you've
mentioned now robotics. Do yousee it as the same thing, or is
(29:19):
it kind of yet another variationof a world model, you know,
using the word and stuff? Like,how how closely do would you
would you believe those two tobe related, you know, as you're
talking about that?
Just because both are bigtopics, you know, and Sorry.
Dustin (29:33):
Sorry. Can I can I just
ask for clarification? You asked
them, like, how close do Ibelieve, like, our models are
Chris (29:37):
kind of, like,
Dustin (29:38):
a world model or a
Chris (29:39):
Well, like, when you say
world model, I I'm just kind of
trying to to clarify the samething I did when there was 20
people in the room and you'reasking what they mean by it.
Mhmm. And you as you mentionedrobotics and stuff, we all need
this representation of the worldout there so that we can do
better about about acknowledgingthe context of what we're
working on, whether it'srobotics or in visual
(30:00):
intelligence, presumably. Arethey I'm just curious. In your
mind, do you think that they arevery closely related, or are
they kinda distinct ideas ofwhat a world model is?
What what is your take on that?
Dustin (30:10):
I would say I would say
that they're pretty related. I
mean, I I to to me, it's, likefundamentally under the surface
of what we're doing is we'retrying to teach the model as
much as we can about the worldthat we exist in and then asking
for it to utilize that in someway. And up to this point, it's
basically been mostly throughcreative mediums, but like that
(30:30):
representation that we'retalking about that, and again,
I'm trying to avoid this termbecause we, I don't know, it's a
little overused, term we're alittle bit, but it is. I mean,
this is what it is. It ismodeling, it's creating a
representation of the world welive in, and then it is using
essentially that intelligence ithas to be able to act in it.
So the idea is that if itunderstands enough of these
(30:50):
relationships and and theactual, like, physical
properties of it, it's it's areally good foundation to, you
know, build and train roboticson top of.
Chris (30:58):
Gotcha.
Dan (30:58):
And I
Dustin (30:59):
hope that hope that
answers that.
Chris (31:00):
Don't know that That did.
No. That was good. Yeah. Thank
you for thank you for putting upwith with that.
I was just curious.
Dustin (31:05):
No. Not not at
Chris (31:06):
It's not unusual to to to
navigate that. So go ahead,
Daniel. Sorry about that.
Dan (31:11):
Yeah. I I from the from the
non robotics person in in the
room, that that was reallyinteresting. I appreciate you
going into that. And I'm and I'mwondering there's that kind of
outworking of some of this whereyou are now kind of
understanding the context that'sin this model, maybe how it's
representing the world. There'salso things that I've seen just
(31:33):
crossing my paths, whether it bekind of in ecommerce or, like I
say, my my phone, my textmessaging.
So on on the web where thesemodels are being more and more
integrated into workflows,whether that be kind of like try
on these clothes or glasses orwhatever, or it's a creative
(31:53):
tools in, like, visual editingplatforms. Are you seeing that
with with your all's models interms of kind of what's the
state of and I know, I wanna getinto your model families here in
a second, and they do differentthings. Right? And many of them,
I I know there's a big, somethat are that are open weight,
(32:16):
so you might not know all theways that they're being used.
Right?
But from from at least thosepartners that you're working
with, in terms of of today, whatare some of those creative and
maybe more most practical usesof these models that you see out
there beyond just kind of the,fun image generation kinda side
of things.
Dustin (32:36):
Yeah. Yeah. No.
Absolutely. I would say, again,
I'll come back to the the momentwe started to get context into
the model that wasn't just textmade a huge like, it was it was
a fundamental change in whatthese models could do and how
people actually worked withthem.
Now our first model, Flexcontext, only could take one
image reference, it was mostlyused as an editing model, but
(32:57):
you could reference a picture ofa product and then tell it to
generate like a nice productphotography set, and this was
very nice. Then moving on, weintroduced the Flux two family
and then the Klein speediersmaller variant that now could
take many references. And thennow we're looking at mean,
people and this kinda come backscomes back again to what I was
(33:19):
just saying of having thisrepresentation of how things
relate, and that the better weactually build this
representation, the moreinteresting ways you can
essentially tie in all of thesedifferent things. So I mean,
I'll throw out one of the mostcommonly used, or I don't wanna
say commonly used, but like kindof obvious cases is okay, like
clothing drive. I'm like, thisis a very, you know, here's a
(33:41):
picture of me, here's a pictureof some clothes.
Could you please, you know, showme what this looks like? I've
seen people like I actually didsome home decoration earlier
this year just trying to see,like, what different couches and
furniture look like in my place.One of the most interesting
uses, I'll say, that kind ofstood out to me that I saw at a
hackathon was and I I don't knowhow much this could be used for
(34:06):
like real planning purposes, butI just thought it was
interesting, is someone wastaking pictures of fire exits in
a building and then generatingwhat it would look like if a
crowd was trying to leavethrough this fire exit in an
emergency, so that theyBrilliant. Could actually like
Yeah. So they could actuallygauge, like, what would this
look under, like emergency,like, where is the crowd?
(34:27):
Like, I don't know. Obviously,there there's, you know, there's
a generated component, so youyou have to take it with a
little bit of grain of salt, butyou still could get a general
idea of this is what this wouldlook like under this scenario. I
thought
Chris (34:36):
that extremely derailing
us, you just sparked a whole
bunch of ideas on that in myhead. So, like, I'm I'm just
like, oh, this great stuff. Keepgoing. Sorry.
Sponsor (34:46):
If you've been
listening to the show over the
past few months, you realizejust how transformative AgenTic
AI is, whether that's ClaudeCode or Hermes Agent or custom
built software that you'redeploying for operational
efficiencies or as new productsto your customers. Regardless of
your maturity now, this is theworld that we're headed towards,
(35:09):
this agentic AI world. Andthere's a lot of security and
governance teams that aren'tletting these agents go into
production because of risksrelated to agency and autonomy
and how do you take care ofthings like prompt injections or
insecure tool usage? There's alot to take care of, and that's
(35:29):
why I'm personally spending mytime outside of the show working
with an amazing team of AIengineers to build Prediction
Guard. Prediction Guard is an AIcontrol plane that you run-in
your own infrastructure behindyour firewall.
Developers can build on top ofthis control plane using
everything that they wanna use,OpenAI and Anthropic compatible
(35:50):
APIs, MCP servers, frameworkslike LangChain, but all of this
is plugged into a built ingovernance harness that enforces
your organization's AI policies,and all of that telemetry goes
back to your monitoring andalerting systems. I'd encourage
you to check out what we'redoing at
(36:13):
You can schedule a demo with meand the team and I'd love to get
your feedback on what we'redoing. So visit us at
That'spredictionguard.com/practicalai.
Dan (36:28):
These are all so you
mentioned the the flux, family
of models. So that's BlackForest family of models or at
least some of the models thatthat you've worked on. Could you
could you just kind of give us aconcrete you mentioned a couple
by name, but, like, give us atour of the the family of models
and then maybe where, I knowthat there's some on Hugging
(36:49):
Face. People can find them.People can can look at them, but
maybe just give us a little bitof a tour of the model family
and then, anything that, thatyou wanna share about some of
the some of the distinctionsbetween them, the different ones
of them?
Dustin (37:03):
Yeah. Yeah. I mean,
we're I I would say
fundamentally, like, as a core,we're still like a I'll say
we're still like a research labthat wants to push the
boundaries, so with each one ofour model releases, we want to
try to level up some capabilityof the model, not just make a
general improvement, but reallysee some new versioning with it.
(37:24):
So I mean, can take back throughThere's not that many models, so
can take back through a littlebit of history. Was a little
That was around two years agonow, we released our first
family models, which was theFlux series, and this was the
first set of models that wecreated after we formed the
company after a very fun butintense, I think it was around
four or five month sprint toreally like build our chops here
(37:47):
and get something out.
And with that, we came up withthis, at the time, this
essentially distinction betweenthe models where we released the
Flux1. So we had the Flux1 Promodel, which was on our API. We
had the Flux1 dev model, whichwas this commercially licensable
model, but the weights were openfor people out in the world to
(38:08):
use. And then we had the Flux1Chanel model, which was a high
speed, like step distilled modelthat was totally I don't know if
it was MIT or Apache, but itopen to use for whatever
purposes you want. Then fromthere, we worked on our tools
(38:31):
series of models where werealized like we needed a lot
more control with the models,and this is where alongside this
work started the context projectto, okay, these are people who
want a lot more control of themodels.
These models are learning a lotof relationships. How can we
leverage this? Which led up tothe Flux context release, which
I talked about. This was our bigedit release. Everything I
mentioned up to this point wasstill kind of our like
(38:51):
historical models.
Now we're getting into our moreup to date models, although
they're getting a little bitolder now, but we have a, I
don't know, something I'mexcited coming, I don't wanna
get dates soon, hopefully laterthis summer, that'll be very
exciting to talk about when itcomes out, but
Chris (39:07):
Looking forward to it.
Yeah. We'll have to get you back
on.
Dustin (39:09):
Definitely. Yeah. Happy
to happen to come back. But then
we got to our Flux two series ofmodels, and Flux two was a very
big upgrade for us, where wereally pushed the capabilities
not just T2I, but also onediting, not just on single
image, but this is where we alsointroduced like the multi image,
the kind of omni edit where youcould put many people, many
different items, have all sortsof relationships. This is where
we started to see, you know, notjust these kind of more standard
(39:32):
advertising use cases, whichstill were like the most common
and we really wanna supportthis, but like some more
interesting how people kind ofbuild, you know, like things
like this fire exit thing hereor other types of stuff.
And then right after that, wecame to our Klein series of
models, which was essentiallylike a size distillation where
we really wanted to pack as muchperformance as we could into a
(39:55):
small model for both use on ourAPI, but also just for open
weight release. We know thatthere's many people out in the
world who use our modelslocally, and we really wanted to
make sure they had somethingpowerful they could use. So it
was a text image and an editingmodel. It was a was a very fast
model, and it was it was prettysmall in comparison to you know,
it was even smaller than our ourFlux one series. And then we've
(40:17):
done further a further speed upon that with our our Klein KV
model, which introduced, Ibelieve for the first time KV
caching, which is this I'll justsay it's an optimization
technique that's very commonwith the language model world
that we were able to bring intothe editing world to get a very,
very big speed up on localediting.
So people who wanted to actuallyuse these models locally, but
(40:38):
also, you know, we also servethis as well, you know. And now
now since then, we've done acouple blog releases, one fairly
fun research, you know, researchblog on on something called self
flow, which I'm reppingpartially because I am on there.
I'm not not the lead author. Ourlead author is incredible. He
list.
She's amazing. But and then nowwe're we're kind of all pushing
(41:01):
forward to our our next bigrelease, which is hopefully
gonna be I I'm quite excitedabout
Dan (41:06):
That that's awesome. And
just to follow-up on that, I
know some people out there, ourlisteners, are always trying
things on, you know, on theirlaptop or or wherever they're
they're they're pulling downthings. I know for quite some
time when I tried to access someof these models and run them
myself, it was either verydifficult to find find, you
(41:30):
know, with my limited resources,the the right the right kind of
configuration to run this, butalso it was sometimes incredibly
slow. So what to, just give asense of some of those, you
know, more hardware optimizedmodels that you mentioned, I
think the the Klein and othermodels. How, maybe my question
(41:50):
is, can I reasonably run one ofthese models on my laptop now
and create a a great image?
Like, what's what's requiredhere? Certainly, I'm not gonna
serve my production web app inan enterprise environment off of
my my laptop, but just maybegive us a little bit of a sense
because that has changed on theLLM side, right, where that has
continually updated andobviously the smaller models are
(42:13):
not up to the same outputquality as the larger models,
but you can run a small model,you know, even on a CPU now, at
least for some tasks that, is ispretty reasonable. So how how
has that progression happened onthe visual intelligence or image
generation side?
Dustin (42:32):
Yeah. I I would say,
I'll I'll first state that the
LLM world definitely has a lotmore people working on it, so
they they definitely have a lotmore of this up and down. That
said, with our Klein series, Ihaven't run this personally.
Maybe I'm not the best person inthe world to speak on this, but
I'm 99% certain you could runthis on, say, like a modern M
series Mac, like a MacBook Pro.As for speed, I don't have any
(42:56):
numbers I can quote because Ihaven't tested this myself, but
this is always like a big kindof trade off in this world of
how, you know, even on thelanguage side, you know, you see
the scaling of a, you know, alot of the new big open releases
that people are excited aboutare getting into the hundreds of
billions of parameters, and thisis why we did the Klein series,
(43:17):
for example, because Flux twowas a 32E model, it was quite
chunky, you could run it on ahigher end local computer, but
it was slower to run locally,you needed some more powerful
computing, this is where likethe Klein series popped in, but
it's always a trade off becausewe wanna push performance and
make something that we're reallyproud of, and then we wanna try
(43:38):
to bring that again back intosmaller, faster scale.
And this is something we'vecontinually done with our
Fluxone series, we had Schnell,with Flux2, we had Klein. I
don't wanna promise anythinggoing forward, but we definitely
wanna keep bringing this biggerpower that we generate into
these smaller models that peoplecan hopefully run locally.
Chris (44:00):
Absolutely. Super cool.
I'm excited to try some of
those. As we are starting towind up, you guys have done some
really, really cool, work inthis space. One of the things we
always like to to ask as we'reclosing out these episodes are
kinda like we wanna get aglimpse into your head, into
(44:21):
your thinking about what's tocome.
And and kinda to to frame it alittle bit as you're thinking
about, like, you know, you'reyou're out of the the day's work
and you're, you know, enjoyingyour evening and your mind
wanders and you're kindathinking about, like, down the
road. What would you know, whatyou want, what where your
passion is is taking you andwhat you'd like to see. Can you
(44:44):
share can you share a little bitabout what the future that you
would like to to create withVisual Intelligence in terms of
what's next that you guyshaven't addressed? I'm not
asking for a product release somuch because I know those those
need to stay in, but kind ofjust like what's what what, you
know, where what is youraspiration? What's your dream in
terms of where these kind ofcapabilities might ultimately
(45:06):
lead?
And and love to love to get youryour your kind of your dream
context, if you will, as wefinish up.
Dustin (45:15):
Sure. If I'm allowed to
to give two split answers here,
there's a
Chris (45:21):
Totally fine.
Dustin (45:21):
There's kind of two
areas that constantly sit on the
front of my mind. At some point,they'll hopefully become one,
but I imagine the next and thisis no glimpse into saying what
we're doing, but this is stuffI'm personally excited about.
Chris (45:33):
Fair enough.
Dustin (45:33):
Well, I would say the
two areas that I'm the most
excited about is, one, is longcontext truly multimodal models.
So nowadays, obviously, we'reseeing lots of people work with
agents constantly, And on ourfront where we're having these
more like continuous generativemodels, we're starting to
(45:54):
introduce more and morecontextually usage with these
references. I'm looking forwardto when this kind of bridges a
little bit more, And we havemodels that not just can like,
okay, you have, say the agentcalls the generative model, but
when this kind of becomes moreor less a model that not just
can, you know, do work in thetext and language and agent, but
(46:17):
actually maybe can thinkvisually as well and generate
audio for you and has the fullcontext of say all the stuff
you've done over the last fewweeks. You know, you don't need
to say give it the reference,give it the right prompt. It
just like already has thiscontext and can reference this
as needed in this, like,continuous space.
This excites me a lot. The otherside, I'll say, is the real time
(46:38):
stuff that's you know, we'reseeing a little bit of it now
out there. I think that this is,like, very early, and I think
it's gonna be very exciting forreal time video, audio, duplex
interactions, know, being ableto I don't know. You see a
little bit of these, like,interactive, you know, stuff
with, like, Genie where, youknow, play the game, but also on
(46:58):
the side of, you can bridge thisback into robotics where it
needs to take in the real worldin real time and make decisions.
Chris (47:04):
Yeah. I was gonna say
that sounds really familiar in
terms of interest on that sideof things. Mhmm. So, yeah,
really cool. Great conversation,and your lead in to to what what
Black Forest Labs is doing wasalso really good contextually in
terms of kind of explaining.
So hope our audience got a lotout of that. Dustin, thank you
(47:26):
very, very much for coming onthe show. Great conversation.
And as you've as you've hinted,there are things to come, and
I'm looking forward to to havingour next conversation as things
move forward a little bit.
Dustin (47:38):
Absolutely. Thank you so
much for having me, guys.
Narrator (47:44):
All right. That's our
show for this week. If you
haven't checked out our website,head to practicalai.fm, and be
sure to connect with us onLinkedIn, X, or Blue Sky. You'll
see us posting insights relatedto the latest AI developments,
and we would love for you tojoin the conversation. Thanks to
our partner Prediction Guard forproviding operational support
(48:04):
for the show.
Check them out atpredictionguard.com. Also,
thanks to Breakmaster Cylinderfor the beats and to you for
listening. That's all for now,but you'll hear from us again
next week.