Episode Transcript
Available transcripts are automatically generated. Complete accuracy is not guaranteed.
SPEAKER_00 (00:19):
Hello and welcome to
Full Tech Ahead.
I am your host, Amanda Razzani,and with me today, I'm so
excited to have MatthewShakstead.
He is the CEO of Parallel Works.
How are you doing today?
SPEAKER_01 (00:32):
Hey, Amanda, doing
great.
Thanks for having me.
SPEAKER_00 (00:34):
Thank you for coming
on the show.
Can you share a little bit aboutParallel Works?
What services do you provide?
SPEAKER_01 (00:41):
Sure.
Yeah.
We are a software company thatmakes a product called Activate
that I call a computing controlplane.
So it's a piece of software thatorganizations deploy, usually in
their own cloud or on-premboundaries.
They connect all their variouscomputing resources together
into it.
(01:02):
And it becomes the single paneof glass that the end users
access the computing through.
And we do a bunch of more adminand operational capabilities as
well for IT organizations andsuch.
So can talk about that.
SPEAKER_00 (01:16):
Great.
So adding that uh visibility allin one place.
SPEAKER_01 (01:21):
Correct.
Yes, that is definitely a bigaspect of it.
SPEAKER_00 (01:24):
Well, that brings us
in our topic for the day, which
is the incorporation of AI.
Of course, it's been swift andfurious.
And one of the problems businessleaders face is how many tokens
are being used with this AI use.
And we know that the cost of AIuse is increasing.
The amount of tokens needed forthe AI is increasing.
(01:48):
And this is presenting a littlebit of a challenge to companies.
So to start off, from yourexperience, where is the biggest
problem that business leadersface right now?
SPEAKER_01 (02:03):
Yeah, great
question.
Uh, and our company is kind ofright in the middle of this, as
many of the organizations thatwe work with start bringing in
frontier models from publicproviders, or they're building
on-prem private infrastructureand deploying models there.
There's different sets ofchallenges in both of those.
(02:24):
But I'd say the biggestchallenge, and I think
everybody's probably seen it inthe news, is you know, you open
up frontier models and API keysto end users within an
organization, you know, tens,hundreds, even thousands of
people now starting to use theseenterprise approved boundaries,
(02:46):
you know, for their models, andthe costs start skyrocketing.
And unless that is kind ofthought about in the very
beginning, in terms of how areyou going to control and monitor
token usage and capability, itit really becomes a big problem.
Like we've seen the Uber storyrecently, where you burn through
the entire budget in a fewmonths that was set because uh a
(03:09):
lot of these things start assomewhat unrestricted.
And and you hit it right.
The the cost of these frontiermodels when you're going to
public providers is expensiveand getting more and more
expensive kind of by the day.
So that's kind of the world thatI find myself in now is one of
governance around these AImodels and using them
(03:30):
effectively and giving theoperators of the models and the
administrators a way to kind ofmonitor usage for one, but
actually go towards cost controland basically allocation, which
I'll kind of talk about as wellin the in the worlds that we
came from previously.
Yeah.
SPEAKER_00 (03:49):
Yes, absolutely.
And there are so many AI toolsbeing used, it's hard for
departments to realize, okay,who's using what tools at what
level, what's really needed,what isn't, what's maximizing
the best end result, what reallyisn't efficient and not giving a
(04:12):
ROI.
So what advice do you have forbusiness leaders to tackle this?
SPEAKER_01 (04:19):
Well, I I'd say I'm
I'm seeing it kind of as a trend
that I've seen take place fromkind of on-prem high performance
computing resources where youknow you bring these systems in
to an internal boundary and youknow, a data center that you
manage, maybe, and then you openit up for pioneer pioneer usage
(04:39):
where hey, everybody come in anduse this thing, and then
eventually you need to start uhallocating out time for units
out, cost tokens, you know, kindof similar units, I'd say, uh,
because the resources start toget constrained within some set
budget.
And you know, I think what youknow, maybe 10 years ago now,
(05:02):
there's a big movement intocloud resources, at least in the
the world I'm from, which ishigh performance computing
resources in the cloud, andstarted moving over with kind of
constrained budgets as well.
And you have, you know, it was asimilar challenge.
I'd say, how do we keep controlof our ballooning cloud costs?
Which starts with tracking ofwhat's actually being used
(05:25):
across resources, taggingpolicies, and then the costs
start reaching a point whereeconomics would actually let
them bring certain aspects oftheir operations back on-prem.
And you go back to these cyclesof hey, we're actually going to
purchase a data center and thewhole operational team to get
this thing running, uh, becausewe can run it at a base load of
(05:46):
90%.
And I think that probably is asimilar trend we're going to see
in these AI models whereorganizations can get started
very quickly running in thesefrontier provider models, you
know, the public endpoints, ifyou will.
And you can put cost aroundthem, and then you'll reach a
point where, hey, we're athundreds of thousands of people
(06:08):
leveraging these things now,using it, you know, 90% baseload
or whatever that that metric isthat makes sense for a given
organization.
And hey, let's bring it on-prem.
And we're gonna, you've seen theopen weight models get better
and better month over month, andthat trend will continue, and
there will be certain sets ofthings running, you know, on
(06:28):
dedicated infrastructureon-prem, which change you know
the economic story a little bit.
You still want to be able tohave control and you know usage
tracking on there, but you know,it's kind of a similar
transition I thought that weI've seen with like cloud HPC in
the last you know 10 years orso.
So uh so back to your pointthough, what what to keep in
mind?
I'd say from the very beginning,having strong visibility in
(06:52):
where the tokens are going.
You know, you connect in a cloudAPI endpoint or you know,
bedrock or pick your provider,make sure from the very
beginning you have visibility inwho's doing what, because that's
going to start growing veryquickly.
If you can put guardrails aroundthese things, uh either from the
provider themselves or from likea control plane, what my company
(07:15):
does, so you can actually plugin all the different models you
want and put the same type ofcost control guardrails around
them, and then start watchingthe behavior of it and you know,
have in mind that hey, maybe ina year it does make sense to
actually bring some dedicatedresources on-prem if you're not
already doing that.
So and kind of manage it thatway.
So that's my my tip kind of havevisibility and guardrails set up
(07:38):
from the very beginning.
Otherwise, you get into kind ofthis untethered access, which as
we've seen, makes really goodnews stories lately.
SPEAKER_00 (07:46):
So yeah, absolutely.
It's interesting how quickly wesaw the shift to cloud, you
know, migrate everything tocloud, and now how it looks like
okay, actually, let's bring someof this back on premises.
Uh, there's this maybe there's abalance that's needed, not
everything needs to be out goingto cloud.
SPEAKER_01 (08:04):
Yeah, I think
there's a balance for sure.
And I'm, you know, my softwareenables organizations to create
hybrid computing environments,so a mix of on-prem and cloud.
So I'm I have a have a vestedinterest in that positioning,
but I really do believe thatthere's a place for both.
And as an organization kind ofmatures on their computing
(08:26):
journey, just economically andkind of baseload, maybe
security-wise, sovereignty-wise,there's reasons to put
infrastructure inside of yourown boundary, you know, on-prem
and operate it that way with allthe costs assumed.
And then there's really greatplaces for cloud where, you
know, they're typically gettingthe latest generation technology
on the floor faster thananybody, faster than most
(08:47):
organizations can get itthemselves.
You're able to really leveragethat uh effectively.
And you know, it kind of becomesa seamless transition to kind of
move that technology from thecloud back to on-prem when you
can, and then, you know, it's acontinuous cycle that way.
So that's how I I've been seeingit that way.
There's places for both forsure.
SPEAKER_00 (09:05):
Right.
And when you were talking aboutputting in these expenditure
guardrails, is are you sayingthat there's a way that uh
company leaders could, you know,say somebody's using these
different AI tools, and oncethey have basically gone to the
threshold of the allottedbudget, they just suddenly
wouldn't be able to use thetools or move forward without
(09:25):
getting some sort of approval?
SPEAKER_01 (09:27):
Yep, that's exactly
right.
You summed it up.
So a system like ours, andthere's others out there too,
I'll say.
But uh we we're different inthat we bring a lot of other
infrastructure types into thesame common allocation and kind
of cost guardrail frameworkbeyond just LLMs, I'd say.
(09:49):
So we're we're touchingKubernetes clusters, you know,
virtualized environments, batchschedulers, cloud resources.
And we rolled out basically anAI LLM gateway because our
customers are asking us, hey,how do we start actually cost
controlling those as well?
So, what it allows you to do isbasically from the admin
(10:09):
perspective, plug in thedifferent models you want your
teams or your end users to use.
So you may just you may havejust purchased like an OpenShift
cluster and you're serving upyour own open weights or you
know, private weights that havebeen fine-tuned, and you want to
make those available to yourusers.
You want to bring in bedrock orAzure Foundry resources, you
(10:31):
want to bring in your enterprisecloud account, whatever it is.
You can plug those all into ourplatform from an admin
perspective, and then startproviding access to specific
groups of who you want to beable to access to particular
models, or users can even servethem up.
We deliver all those modelsthrough a single OpenAI
(10:53):
compatible endpoint.
So you plug that into your codeassist tool or your chat or
wherever you want it, it listsall the models you have access
to and the given token budgetthat you've been allotted.
And so users are using theirmodels, and as soon as it
basically is out of budget, itwill deny the request until you
(11:15):
get more allotted.
So that's exactly what'shappening.
Yep.
SPEAKER_00 (11:18):
So that brings me to
my next question is I think that
the end result of this wouldobviously be some really good
company insight and realizingwhat tools are needed and at
what level.
But in the beginning, is theremaybe a little bit of chaos as
people that are used to usingcertain tools suddenly are
stopped in the middle of aproject because they can't move
further.
Um, and how do you advisebusiness leaders deal with that
(11:40):
process?
You know, that what I thinkwould be a first initial chaos
before, you know, it irons out.
SPEAKER_01 (11:47):
Yeah, there's even a
chaos before that, too, at least
that I've been seeing where alot of end users, like the
practitioners, people doingwork, uh, maybe start with even
like a personal account, youknow, and they're they're on
their own, you know, co-pilotaccount or whatever, and they're
using that.
And then suddenly theirorganization says, Hey, you need
to use this enterprise mandated,you know, set of tools.
(12:09):
And hey, you're moving to Cloudor Pick Your Provider, OpenAI,
Enterprise, et cetera.
So that's like the first levelof transition.
And then, hey, you're beingallocated the equivalent of
$10,000 a quarter or$10,000 amonth or whatever for your work,
use it sparingly.
(12:30):
Or, and that, and that's where Ithink end users do need to start
actually thinking about how arethey using their models and
actually maybe having processthese that can use less
expensive models for certainthings, or if you're bringing
this all into a hybridinfrastructure, hey, we're gonna
send certain sets of tasks toour on-prem dedicated
infrastructure where the tokenrates, what I'm being charged,
(12:50):
are you know maybe six X lessthan some of the public
providers.
Move these tasks over there.
So you can start to be aware ofthat.
I've seen usually users notunless they have a economic
incentive to do so.
So, oh, they see what they'reactually spending, they know
they have this work to be donein the next month.
(13:11):
So I got I need to make thatlast by maybe using a, you know,
our private internal model, forexample.
Yeah, I guess that's my mycomments there.
There is definitely chaos, but Ithink it just brings awareness
to the actual end users, howthey're actually interacting
with the models, I would saymore so, which is similar to
what cloud, I think, was it'slike here, you know, you can
(13:33):
spin up whatever you want in thecloud, and then suddenly, oh, we
have a$50 million budget a year.
Uh something needs to change.
SPEAKER_00 (13:39):
Right.
unknown (13:40):
Yeah.
SPEAKER_00 (13:41):
Do you think though
that there's gonna be, because
I'm hearing a lot about theincreased token costs and the
amount needed for these tools,that some of them were initially
free, everybody started usingthem heavily.
Now suddenly they're costingmoney.
Um, other ones are costing more.
Do you think there's gonna besome pushback, especially when
companies do put in theseregulations and start cutting
(14:03):
the fat, if you will?
Do you think that maybe we'llsee prices go back down as they
say, oh shoot, nobody wants touse these tools anymore because
how expensive they are, they'recutting back.
And we'll see those prices fallback down.
Because I think right now it'ssort of a money grab.
Oh, they need these tools, we'regonna make them very expensive
now.
SPEAKER_01 (14:22):
Yeah, or or
actually, I mean, subscription
models before you reach theseenterprise plans for the
frontier lab providers and suchthat are being subsidized,
essentially, right?
So you're getting if you were torun the same processes with like
an API key and pay like the pertoken input output rates, you're
getting way, way more valueuntil you transition to these
(14:44):
like enterprise API keys.
And then suddenly it's like, oh,we're spending, yeah.
You've seen it in the news, youknow, we just blew our entire
budget.
I think that will definitely bea drive, you know.
I I want to say it will be astrong driver for organizations
to seriously think about thistransition from moving from just
frontier models using the publicproviders to on-prem
(15:06):
infrastructure, where you have amuch more predictable cost
control, like you saw with youknow on-prem HPC and everything
else.
And as the open weights getbetter and better, which I think
we're starting to see, but youknow, GLM 5.2 coming out
recently.
That's that was like big newswhere it's really catching up in
(15:26):
terms of capability.
You can start running openweight models for certain
classes of things with much morepredictable cost.
And then that's going to be adirect driver to the frontier
models because, like, why issomeone gonna run that if I can
run similar things in on-premeinfrastructure that I control
three month or for three-yearcadence, right?
Refresh cadence, uh, or you gorent it from a neo cloud
(15:48):
provider or a hyperscaler, butyou rent the infrastructure and
host it yourself.
But at least you have very clearcost control.
You know, you know what, youknow what the max consumption is
gonna be.
And yeah, you still need tocontrol the token usage because
there's contention.
But yeah, so I think it's a it'sdefinitely gonna be a big
challenge as they want toincrease those rates to recap
all recoup all the massiveinvestments they did, but you
(16:10):
have you know open weight modelscoming in with you know on-prem
infrastructure that could standup at least some portion of what
these things are doing,especially for those like base
loads, where hey, we can runthis thing, you know, 80%
utilization reliably, and thismodel's good enough for what it
needs to do.
That will get that will getchallenged.
And I think you you kind ofmentioned it the you know,
(16:33):
you're hearing agentic workloadsand really people using 10x of
what we're at right now becausethey have you know 100 agents
doing XYZ things at kind of alltimes.
Now I was just at someconference, they're like, Oh,
yeah, the the compute need'sgonna go up like a thousand
times for every individual.
You know, how's that gonnaactually happen?
SPEAKER_00 (16:50):
So Yeah, and how
yeah, how is that sustainable?
SPEAKER_01 (16:54):
Sustainable from a
yeah, cost model and uh, you
know, it it definitely willchange, I think, the the shape
of the compute, you know, movingforward.
Uh, it starts unfolding, whichis not you know that far away, I
think.
SPEAKER_00 (17:07):
Well, if there was
one key takeaway you could leave
our audience with today, whatwould that be?
SPEAKER_01 (17:12):
I I'm going back to
the the beginning point I had
that you know, as you're rollingout these models, whether
they're Frontier Lab, onhyperscale providers, Azure, you
know, AWS, vertex, whatever, uh,whether you're bringing them
on-prem, starting from the verybeginning, putting in a system
for visibility across all thedifferent models and guardrails,
(17:34):
if you can, you know, putting ina system for guardrails so that
you do have the option to turnoff the faucet if you need to.
That's my biggest, I'd say, tipright now from what I've I've
been seeing.
SPEAKER_00 (17:44):
So okay, great.
Well, thank you so much forcoming on the show and sharing
your insights with us today.
SPEAKER_01 (17:50):
Sure.
Thanks, Amanda.
Good questions.
SPEAKER_00 (17:52):
And thank you to our
audience.
If you have any questions orcomments, please leave them
below and I'll try to respondback as soon as possible.
Have a wonderful week.