Methodology
Every number on this site comes from official lottery prize data plus arithmetic you can check. This page is the arithmetic.
What the lotteries publish
For each game a state lottery publishes the prize levels, the printed odds of hitting each level, how many prizes of each level the game started with, and how many are still unclaimed. Some states publish tickets printed instead of per-tier odds; where that happens we derive the odds, and we say so below.
The formulas
For a prize tier t with published odds 1 in O, N prizes at launch and R unclaimed:
- Tickets printed for that tier = O × N. Every tier implies the same print run, so we average across tiers to estimate the game's ticket count.
- Tickets sold = O × (N − R): if a tier has given up prizes, that many tickets must have been sold.
- Tickets remaining = tickets printed − tickets sold, averaged across tiers.
- Odds now for a tier = tickets remaining ÷ R.
- Any-prize odds now = tickets remaining ÷ (sum of R across all tiers).
- Current payout = (sum over tiers of R × prize amount) ÷ tickets remaining ÷ ticket price. Expected value per ticket is the same figure before dividing by price.
- Cash odds repeat the calculation ignoring free-ticket tiers. Profit odds ignore every tier worth the ticket price or less. $600+ odds count only tiers at or above the claim-at-the-lottery threshold.
Annuity prizes are converted to their nominal total: "$1,000 a week for life" is counted as $1,000,000, and "$50,000 a year for 20 years" as $1,000,000, matching how the lotteries themselves advertise them. Free-ticket prizes are counted at the ticket price.
Known weaknesses
The tickets-remaining estimate is an unweighted average across tiers, so a top-prize tier with five prizes counts as much as a $5 tier with two million. When a game is nearly sold out the small tiers dominate reality but the average lags. We keep this method because it is the one behind every historical number on the site, and changing it would silently rewrite years of comparisons; a weighted version is planned as a separate column.
Texas does not publish per-tier odds, so we derive them from the stated ticket count divided by prizes in that tier. Michigan publishes only the top prize levels, so its lower tiers are derived from overall odds. New Jersey publishes tickets printed rather than tier odds. In each case the figures come out very close to the states' own published overall odds, but they are derived, not quoted.
Finally, remaining-prize counts are only as fresh as the lottery's own publishing. A prize claimed today may not appear for a day or two. Where a state's data stops updating, we show a notice on that state's pages rather than presenting old figures as current.
Why some games show a payout above 100%
About one game in seventy on this site shows a current payout over 100%. That is not a typo and not, by itself, an error: it says the prizes still unclaimed are worth more than the tickets still unsold would cost to buy. It happens late in a game's life, when most tickets have sold but the largest prizes have not been hit. It is also exactly where the averaging weakness above bites hardest, because the estimate of tickets remaining is thinnest when a game is nearly sold out. We publish the number as the model produces it, mark it on the tables, and say plainly that its precision is low. What it reliably tells you is that the game is unusually rich in unclaimed prizes. What it does not tell you is that buying it is profitable: the remaining tickets are scattered across thousands of retailers and you cannot buy them all.
Draw games: frequency, significance and expected value
Scratcher odds change as prizes are claimed. Draw-game odds never change at all, so the analysis is a different job: not "what is this ticket worth now" but "does anything in the history mean anything". Everything below is plain arithmetic you can reproduce.
Matrix eras — why counts start where they do
A frequency count is meaningless across a change in the pool. Ball 68 cannot be cold over a period when it did not exist. So every count is scoped to the span during which that pool's matrix was unchanged, and the era table is checked against the data two ways: every draw must fit its declared matrix, and every era must contain a draw that actually reaches its declared ceiling. A typo in either direction fails one of those checks and the build stops.
The two pools are scoped separately, because they do not move together. Mega Millions has drawn five numbers from 70 since October 2017, but the Mega Ball fell from 25 to 24 in April 2025. Scoping the main numbers to the newest game version would throw away several hundred perfectly comparable draws; scoping each pool to its own span keeps every draw that is genuinely comparable and no more.
Is a number hot?
Each number appears in a draw with probability p = k/m (k balls drawn from a pool of m),
independently across draws, so over n draws its count is exactly Binomial(n, p):
expected = n · p
sd = sqrt(n · p · (1 − p))
z = (observed − expected) / sd
A z-score on its own is not a finding, because the hottest of m numbers is supposed to look
extreme. The number to compare it against is the expected maximum of m standard normals, which we
compute by direct quadrature of E[M] = ∫ x · m · φ(x) · Φ(x)^(m−1) dx. For a 69-number
pool that is about 2.37 standard deviations. So a leader sitting 2.3 standard deviations high is not
a hot number; it is an ordinary Tuesday.
For the pool as a whole we run a goodness-of-fit test against "every number equally likely". The
textbook statistic divides each squared residual by the expected count; here each cell is Binomial
rather than a multinomial cell, so the variance is n·p·(1−p) and the (1−p)
belongs in the denominator:
chi-square = Σ (observed − expected)² / (expected · (1 − p)) on m − 1 degrees of freedom One approximation is worth naming: the m counts are weakly negatively correlated, because every draw contributes exactly k of them. The correlation is of order 1/(m−1) and ignoring it moves these numbers by less than their third digit.
Droughts
Gaps between appearances are geometric, so the chance a given number goes at least g
draws without appearing is (1 − p)^g. We publish two different figures from that, because
quoting only one is how "overdue" lists mislead: the chance for one named number right now,
which is small, and the number of times a drought that long has occurred somewhere in the span, which
is usually large. A number is never due.
Expected value
Prize-tier odds are exact combinatorics — C(k,j)·C(m−k,k−j)/C(m,k) for matching j of k
main numbers, times the bonus-ball probability — and reproduce each operator's published table to
eight significant figures. The fixed-prize expectation is the sum of probability × prize over every
tier below the jackpot, including any built-in multiplier: Mega Millions has applied a random 2× to
10× to every non-jackpot prize since April 2025, averaging 3.0×, and leaving it out understates that
game's return by two thirds.
The jackpot term is the part nobody else publishes, because it needs the sharing correction:
EV_jackpot = P(jackpot) · cash_value · (1 − tax) · E[1 / winners]
E[1 / winners] = (1 − e^−λ) / λ for λ = tickets_sold · P(jackpot) Your own ticket is the "1"; every other ticket sold is an independent chance to share. Ticket sales are the one genuinely uncertain input — we interpolate between published sales at known jackpot levels, and label the figure as an estimate wherever it appears. Cash value is taken at 50% of the advertised annuity, which moves with interest rates.
Two sources, always
No draw is published on one source's word. Every game is cross-checked against several independent official feeds; a draw that one feed omits is published from the others and the omission recorded, and a draw two feeds contradict is settled only by a feed closer to the drawing itself. If nothing can settle it, the game does not publish. Every discrepancy we have found is listed on the sources page — including the six Powerball draws missing from New York's open-data feed.
Update cadence
We re-read all 9 states every day and store a snapshot of every prize tier, which is what makes the history charts and claim logs possible. Snapshots are never overwritten.