All Episodes

October 2, 2015 12 mins

The multi-armed bandit problem is named with reference to slot machines (one armed bandits). Given the chance to play from a pool of slot machines, all with unknown payout frequencies, how can you maximize your reward? If you knew in advance which machine was best, you would play exclusively that machine. Any strategy less than this will, on average, earn less payout, and the difference can be called the "regret".

You can try each slot machine to learn about it, which we refer to as exploration. When you've spent enough time to be convinced you've identified the best machine, you can then double down and exploit that knowledge. But how do you best balance exploration and exploitation to minimize the regret of your play?

This mini-episode explores a few examples including restaurant selection and A/B testing to discuss the nature of this problem. In the end we touch briefly on Thompson sampling as a solution.

Mark as Played

Advertise With Us

Popular Podcasts

Dateline NBC
Monster: BTK

Monster: BTK

'Monster: BTK', the newest installment in the 'Monster' franchise, reveals the true story of the Wichita, Kansas serial killer who murdered at least 10 people between 1974 and 1991. Known by the moniker, BTK – Bind Torture Kill, his notoriety was bolstered by the taunting letters he sent to police, and the chilling phone calls he made to media outlets. BTK's identity was finally revealed in 2005 to the shock of his family, his community, and the world. He was the serial killer next door. From Tenderfoot TV & iHeartPodcasts, this is 'Monster: BTK'.

Stuff You Should Know

Stuff You Should Know

If you've ever wanted to know about champagne, satanism, the Stonewall Uprising, chaos theory, LSD, El Nino, true crime and Rosa Parks, then look no further. Josh and Chuck have you covered.

Music, radio and podcasts, all free. Listen online or download the iHeart App.

Connect

© 2025 iHeartMedia, Inc.