A capper posts 8-2. You want to tail. But here is the uncomfortable truth: 8 picks cannot prove anything. We simulated 200,000 bettors to find out exactly how many picks it takes before a record stops being luck — and then audited our own grades against that standard.

How many picks before you can trust a capper — sample size infographic

The data: every figure below is computed from the CappersTracked dataset — 1,460 tracked handicappers, 170,170 graded picks, 372 currently graded (protected variant). Snapshot: 2026-07-29. Simulations are Monte Carlo runs with a fixed seed — every chart in this series is reproducible.

1. The Coin Test

Start with the worst-case assumption: the capper is just flipping a coin against a standard -110 line. A fair coin at -110 loses 4.5% of everything it stakes — that is the vig. So the honest question is not "did they win?", it's "how many picks do they need to prove they are doing better than a coin?"

Statisticians answer with a confidence interval. With n picks, the band around the observed ROI shrinks at the rate 1/√n:

95% CI width = 1.96 × σn

The more picks you have, the narrower the band — and once the band stops touching 0%, the edge is statistically real. Here is how fast the band closes on a fair-coin bettor:

±34%
95% band at 30 picks
±19%
95% band at 100 picks
±11%
95% band at 300 picks
±8%
95% band at 500 picks
±6%
95% band at 1,000 picks

Even at 1,000 picks the best a coin can do is pin its true edge to within ±5.9% — which is why "edge detected" at 100 picks is almost always a regression to the mean, not a discovery.

Fig 1 — At 30 picks the 95% band is ±34.2%. At 1,000 picks it is ±5.9% — and it only stops touching 0% once the edge is statistically real.
Fig 1 — At 30 picks the 95% band is ±34.2%. At 1,000 picks it is ±5.9% — and it only stops touching 0% once the edge is statistically real.

2. How Many Picks "Proves" an Edge?

We computed the pick count where the confidence band finally separates from 0% for four edges — the raw-ROI thresholds behind our A+, A, B+ and B grades:

Edge being proven Picks needed (95% confidence) At 3 picks/day
A+ threshold(raw ROI 8%)537 picks0.5 years
A threshold(raw ROI 5%)1,386 picks1.3 years
B+ threshold(raw ROI 3%)3,865 picks3.5 years
B threshold(raw ROI 1%)34,885 picks~32 years
Table 1: Prove an 8% edge in ~537 picks — a 5% edge takes ~1,386, a 1% edge would take a lifetime.

We verified the 5% number with a Monte Carlo run: 1,386 random bettors with a true 5% edge, 1,386 picks each — 97.3% landed profitable with a confidence band that excludes zero, right where the math says it should.

Now flip the question: how often can a coin fake an A+? At -110, posting a +8% ROI means winning about 57% of your bets. A fair coin needs that exact record just to look like the capper you are about to tail:

Picks Record needed for a +8% ROI A fair coin posts it…
106 of 10~1 in 3
2012 of 20~1 in 4
5029 of 50~1 in 6
10057 of 100~1 in 10
200114 of 200~1 in 36
300170 of 300~1 in 83
Table 2: The same math that "proves" a small edge is the math that lets a coin impersonate a star. Only volume separates the two.
Fig 2 — Proving an 8% edge needs 537 picks; a 5% edge 1,386; a 1% edge would take a lifetime. Log scale so both fit on one chart.
Fig 2 — Proving an 8% edge needs 537 picks; a 5% edge 1,386; a 1% edge would take a lifetime. Log scale so both fit on one chart.

3. What Our Own Grades Sit On

Now the uncomfortable part — our own data. Here is the distribution of graded cappers by pick count, and the median pick count behind each grade:

Fig 3 — 41% of all grades sit on fewer than 100 picks. Volume — not the letter — is the real evidence.
Fig 3 — 41% of all grades sit on fewer than 100 picks. Volume — not the letter — is the real evidence.
Fig 4 — The median A+ is backed by more picks than any other grade. Thin samples and high grades are the exception, not the rule.
Fig 4 — The median A+ is backed by more picks than any other grade. Thin samples and high grades are the exception, not the rule.

The honest read: only 5 of 23 A+ cappers have 300+ picks — the volume where an edge even starts to separate from coin-flipping. Most grades are estimates, not proof. That is exactly why we built Bayesian shrinkage (see the grading methodology): it tells you a capper's most likely skill, and it stays humble when the sample is thin.

4. So When Can You Trust a Capper?

  • Under 30 picks: A grade is a courtesy. Anyone can flip 8-2. Treat an A here as "promising", not "proven".
  • 30–100 picks: The shrinkage estimate is meaningful, but the confidence band is still ±20%+. Variance can flip a grade in a week.
  • 100–300 picks: Trends start to separate from noise. This is where a grade begins to mean something.
  • 300+ picks: A sustained edge here is statistically defensible — and rare. Only 5 of our 372 graded cappers hold an A+ at this volume.

The takeaway is not "ignore small-sample cappers" — it's know what each record can and cannot prove. Volume is the only honest currency in this game. A coin starts 8-2 about 6% of the time; a 56% winner over 800 picks is statistically real — and almost nothing in between can be told apart. Our grades already blend in sample size via shrinkage; this page shows you how thin the evidence really is behind a shiny new A+.

The call is yours — every profile shows the pick count and record behind the grade, so you can judge the evidence on the capper leaderboard.