What are they gonna play?
A statistically honest guess at the next Phish setlist — built from all 2,029 shows and 967 songs since 1983, refreshed daily from phish.net. It knows the current rotation, who's overdue, what opens, what closes, and what follows the big jam.
Every default here is the exact setting this tool's accuracy was tested at, so nothing needs changing — it's for poking at how the prediction shifts, or testing the model against a night that already happened.
1 · Which show are we predicting?
2 · Which history counts?the window the odds are measured over
5 years is the sweet spot — it's the setting every accuracy number on this page was tested at. Recent shows always count more than old ones.
3 · How much do recent shows count?
4 · Special cases
5 · Display onlychanges a label, never the odds
Predicted setlist
Every song gets a probability for the next show, measured from 2,029 setlists going back to 1983.
Realistic rolls the dice on each song at its own probability and lays the result out as a show — one plausible night, not the likeliest one.
Most likely instead ranks every song and shows the arithmetic behind it.
~150 songs in rotation, ~18 played a night.
P(played) = how often it's been played × how overdue it is → corrected against reality
How often — the share of shows in your window that included the song. A song woven in and out of one night counts once. Recent shows count more than old ones (adjustable half-life), so the number follows where a song is now rather than its career average.
How overdue — measured from the full history rather than assumed. A song is suppressed the night after it's played (0.45×), peaks about 4–8 shows later (1.12×), and drops well below its own average once it's been gone 20+ shows (0.22×) — long-dormant songs come back less often than raw frequency implies, not more. This term does most of the predictive work.
Overdue against its own habit — the two corrections above work on the calendar, which can't tell a song that plays every four shows from one that plays every twenty. So a further correction is fit on the ratio of a song's current gap to its own average gap. It matters a lot: a staple sitting past twice its usual gap is bumped 2.1×. Before this was added, exactly the songs you'd swear were due — big, frequently played, conspicuously absent — were predicted at 25% and actually turned up 39% of the time. That gap is now closed.
Corrected against reality — raw scores are checked against what actually happened in the ~120 shows before your target date, at your current settings, and bent to match. Nothing is held out: a themed run (like the 2026 90s-throwback MSG run) picks its setlist by a rule this model can't see, but it's a handful of shows out of thousands — it washes out in the average rather than skewing it. We checked: leaving it in versus pulling it out moves any song's number by at most 1.5 points. Not worth a special case. The correction refits whenever you change a setting.
Inseparable pairs — Mike's Song → Weekapaug Groove (same set, songs in between), The Horse → Silent in the Morning (back to back), Tweezer → Tweezer Reprise (later in the night). Nine such pairs are mined from history and drawn as single units, so the partner arrives at its true conditional rate instead of doubling the pair's odds.
Songs that share a set — thirteen further pairs turn up in the same set well beyond chance: Tweezer + Piper (2.6×), Also Sprach + Piper (2.6×), Slave + Light (2.7×), Funky Bitch + The Moma Dance (3.2×), Harry Hood + Twist (2.2×). They nudge set assignment rather than forcing it.
Where a song lives — each song carries its own history across seven slots (set 1 opener/song/closer, set 2 opener/song/closer, encore). Songs that only ever land as closers or encores are barred from mid-set outright: Tweezer Reprise is 92% encore-or-set-2-closer, and it will never appear in the middle of a set here. Slots are filled by weighted draw, not by the single best fit — real encores draw on 193 different songs, and the five most common cover only 21% of them.
Set length in minutes — sets fill to time (~73 min, ~67 min, ~15 min encore), not to a song count, so a night that draws Tweezer, Down with Disease and Ruby Waves simply fits fewer songs.
Cool-downs — the breather that follows a big set-2 jam is drawn from a curated list of ballads and low-energy songs (Waste, Shade, Lonely Trip, Lifeboy, Leaves, Dirt, Strange Design, Velvet Sea and about fifty more), then ranked by how often each actually follows a 10+ minute jam. This is the one hand-tuned piece of the model, and deliberately so: tempo and mood appear in no available dataset, and the structural signals genuinely cannot tell a ballad from a short fast song — NICU sits 98.6% mid-set, identical to Lifeboy, and Poor Heart is 100% mid-set but is a bluegrass sprint. A measured proxy would keep producing wrong answers, so it isn't used here.
Position in a run — opening, middle and closing nights get measurably different setlists, and it's filled in for you automatically whenever your show date matches something real (already happened or on the announced schedule), rather than asked as a guess. Closing nights run longer (21.4 songs vs 18.0 on openers) and favour Wilson, AC/DC Bag, You Enjoy Myself, Buried Alive and Gotta Jibboo (about 2×); opening nights favour Stash, Blaze On, Bouncing Around the Room and Sample in a Jar.
No repeats inside a run — when the show you're predicting is night 2 or later of a stand at one venue, everything already played earlier in that run is all but ruled out. This is the strongest single pattern in the whole dataset: measured across 2009 onward, just 0.3% of a night's songs had already appeared earlier in the same run, against 17.8% for any three preceding shows. The Baker's Dozen ran thirteen nights and 237 songs with exactly one repeat. Adding this lifted backtested top-20 accuracy from 4.76 to 5.14 songs overall, and to 5.56 on run nights specifically. Ruling songs out doesn't shorten the night, it concentrates it — a run night actually runs longer than a one-off (18.1 songs against 16.7) — so the probability taken off the already-played songs is handed back to everything still eligible. Deep into a run that's a large boost for whatever's left; on night 2 it barely moves.
Bustouts — a return after a long absence, with the threshold set in Advanced settings. At 100+ shows this happens roughly every other show; 68% of real shows have none. At most one appears per night here. Songs gone so long they have no plays in your window share a measured 0.25/show budget, split across career-size tiers in their real proportions — so a song with one or two career plays can return, but almost never does.
Date-locked songs — anything played almost exclusively on one calendar date is excluded unless you set the show date to match. Exactly one song qualifies: Auld Lang Syne, 28 of its 29 outings on December 31.
Tested against 143 shows the model never saw (2023 through January 2026, themed runs excluded): the top 20 songs by probability contain about 5 of the ~18 played. That is far better than chance and, as far as we can tell, better than anything else public — phish.net's Trey's Notebook, which ranks by plays in the last year, gets about 3. Bolting our overdue-ness term onto their one-year count lifts them to 4.8, which says plainly that the overdue term is where nearly all the edge lives; window length barely matters.
But five of eighteen is the honest ceiling here, and it's worth being blunt about why: Phish setlists are close to irreducibly random. Roughly 150 songs are in live rotation, the band deliberately avoids patterns, and a good chunk of any night is songs nobody would have ranked highly. No amount of modelling fixes that. Treat a generated setlist as a plausible night, not a forecast.
Known weak spots, stated plainly: song selection rests on much firmer ground than song ordering — slot tendencies are real but weak. We checked openers specifically, ranked by each song's odds of showing up at all times its own history of opening a set, against the same 143-show window above (139 of those had one clearly recorded opener): the real set 1 opener landed in our top 5 24.5% of the time, top 3 15.1%. Worse than song selection overall — ordering is genuinely the harder problem, and we're not going to pretend otherwise. Bonded pairs still nudge their partners' rates upward slightly. The cool-down pattern is genuine but mild (18% vs a 14% base rate), so it fires on only about 10% of generated nights. Nothing here can anticipate a themed show, a guest sit-in, a debut, or a one-off gag, and those are a real share of any tour.
Setlists come from the phish.net API — the work of the volunteer-run, non-profit Mockingbird Foundation, and worth supporting. Song lengths come from phish.in (median of every circulating version since 2009). Neither project is affiliated with this tool, and any errors here are ours.
Every song, scored for the next show. Same engine as the Predicted Setlist tab, shown per-song with the full math — plus who's most likely to open, close, and encore, and the bustout watch. Sort any column, search any song.
Most likely by slot — who opens, who closes, who encores
Each percentage is the chance of that exact slot — the song being played AND landing there — so the numbers run smaller than plain play odds. It's a song's play probability times its own historical share of that slot: the same math the Phantasy Tour picks are built on. Click any song for its history.
Top 15 by chance of being played at the next show
Formula: recency-weighted frequency (distinct shows played ÷ shows in window, weighted by your recency setting) × gap multiplier (empirical suppression/boost by shows-since-last-played) → calibrated against actual play rates. The calibration is dynamic: it refits to your current window and recency settings on the ~120 shows before your "as of" date (excluding the themed 2026 Mexico/Sphere/MSG-90s runs), so changing settings re-anchors the probabilities honestly instead of reusing a stale correction.
Bustout watch
| Song | Prob. | Wtd. freq | Gap now | Avg gap | Gap vs avg | Gap mult | Length | Jam follow | Usual slot | Last played |
|---|
Top 15 songs by frequency in this window
Frequency = distinct shows played ÷ shows in window (recency weighting applies). Debut-clamped for songs newer than the window. This view is plain history — no prediction, no calibration, just what actually happened in the window you set at the top of the page.
| Song | Shows played | Shows in window | Frequency | Once every | Debut | Last played | Gap |
|---|
The official call
This is the committed prediction: a consensus of 500 generator runs at default settings — the songs that turned up most often, assembled under the same structural rules every drawn night obeys — snapshotted by the daily build before the show and never re-rolled. Grading happens on the first build after the show, and a graded row is never touched again — that's what makes the table below resistant to after-the-fact fiddling.
Show by show
| Date | Venue | Official call | Top-20 | Of setlist | Opener |
|---|
Click any graded row to see the actual songs — what we called and hit, what we called and missed, and what they played that we never had. Of setlist is recall (share of the played setlist we had in our top 20); Top-20 is precision out of 20 picks. The two move in opposite directions when a show runs long, so both are shown rather than one headline number.
Live rows are graded against the exact prediction committed before the show. Backfill rows (greyed) were computed retroactively over past shows by the identical method — same code, same default settings, stats cut to the day before each show — but they can't be verified the way a live row can, so they're counted separately and will never mix with the live average.
Playing a setlist prediction game? This turns the probabilities into the highest-expected-value entry — FantasyPhish's 13 picks for a single show, or Phantasy Tour's 10-song Tour Tournament team. Copy the picks straight into your entry form.
Optimal picks
Expected points = P(song lands in that exact slot) × the slot's point value + P(it shows up somewhere else) × 1 consolation point
"Chance it lands there" is the song's overall probability of being played multiplied by its own historical share of that slot. Example: a song with a 30% chance of being played that has opened set 1 in 25% of its appearances has a 30% × 25% = 7.5% chance at the Set I Opener slot. For a 9-point slot that's 9 × 0.075 = 0.68 expected points, plus the consolation point for the remaining ~22.5% of the time it's played somewhere else — about 0.9 expected points total. (FantasyPhish has no consolation point, so it's just points × chance.)
How the team is picked — every song is scored against every position, then the highest-value combination is assembled greedily, each song used at most once. That's why the opener and encore picks aren't simply the most likely songs of the night: a slightly less likely song that always opens beats a more likely song that rarely does.
Tour Tournament odds compound over the tour: a 10%-per-show song has a 1 − 0.920 ≈ 88% chance of landing at least once in 20 shows. The One Timer slot instead wants exactly one appearance — a second outing wipes it out — so it favors mid-probability songs, not locks.
Known weak spot, stated plainly: the "overdue" discount uses a curve measured across all songs, which only mildly suppresses a song played 2 shows ago. For strict-cycle songs — Tweezer Reprise being the canonical case, since it only appears when Tweezer does — that global curve under-discounts a recent appearance. If Reprise (or a song like it) shows up as a slot pick right after it was just played, apply your own judgment before locking it in.