← Back to waitlist

The Standard

TL;DR

We publish the bar before any model races it, measure in public, and only sell picks from a model that has cleared it.

To enter a paid pond a model needs all four:

100 decided picks, each with a verified closing price.

Mean closing line value of at least +1.5 points.

At least 55% of picks priced better than the closing line.

Picks spread over at least 30 distinct days.

Miss one and nothing is sold. Clear all four and a model can still be publicly evicted if it stops clearing.

Nothing has cleared. Our first model lost money, 50-50-2 at -9.7%, and that record is a file you can download. Watching is free and always will be.

This page explains how Pixel Forge Plays decides which betting models are worth paying for, and what happens to the ones that are not.

The standard below is published before any model races against it, and it does not change during a race. Models move against the standard constantly. The standard itself does not move.

This page lives in a public repository and its full edit history is linked at the bottom. If we ever changed a number in the middle of a race, the commit log would show it, with the date. You do not have to take our word for any of this.


What this is

Pixel Forge Plays is a public lab. We build betting models, run them in the open, and score every one of them against a standard we wrote down in advance.

We do not know yet whether we have an edge. That is the honest answer, and it is the reason the lab exists rather than a sales page.

Watching is free and always will be. The only thing we sell is the picks of a model that has cleared the standard below.


Ponds

A pond is one sport and one market. MLB totals is a pond. MLB moneyline is a different pond. A model that prices run totals is not competing with a model that prices who wins, so they do not race each other.

Every pond carries the identical standard. We do not lower the bar for a quiet market.

Most ponds are empty, and empty is the expected state rather than a stage on the way somewhere. The standard is deliberately hard to clear. If it were easy, clearing it would not tell you anything. A pond that sits empty for an entire season is the system working, not the system stalling.


The entry gate

A model enters a pond by clearing four conditions at once. Falling short on any one of them is not clearing.

1. Sample. At least 100 decided picks.

2. Beat the close on average. Mean closing line value of at least +1.5 points.

3. Beat the close often. At least 55% of picks priced better than the closing line.

4. Spread across time. At least 30 distinct days. A model cannot qualify on three hot afternoons.

Condition one means one hundred decided picks each carrying a verified closing price. A pick we could not measure does not count toward the hundred, and it does not count toward the other three either. There is no version of this where a model sells picks on a record we could not fully measure.

The fourth condition exists because picks made on the same day are not independent evidence. Twenty picks on one slate tell you far less than twenty picks across twenty days, and a standard that ignored this would be easy to game by accident.

When the gate numbers get set

The bars above are not final yet. We are still measuring how much closing line value varies on MLB totals, and right now we have too few observations to know. The numbers we would use today are borrowed from moneyline, where we have more data.

So here is the decision, written down before it happens.

On August 20, 2026 at 12:00 UTC we count our MLB totals evaluations that meet all of these: decided as a win, loss or push; edge of 13 percent or higher; either posted publicly or tracked in the shadow book without being removed by a published rule; and carrying a recorded pre-game closing price. Runline and tennis are excluded. Two evaluations are excluded because their posted and closing prices refer to different lines: Red Sox versus Blue Jays UNDER 9, and Orioles versus Braves OVER 6.5, both from before August 5.

If that count is twenty or higher, we set the gate numbers from what we measured.

If it is under twenty, the numbers above publish unchanged and do not move afterward. Either way we will state the count on this page.

Twenty is not arbitrary. Below it, a measurement of variance is too imprecise to tell a forgiving bar from a punishing one, and setting a permanent number off it would be false precision dressed as rigour.

The August 20 count. On August 20, 2026 at 12:00 UTC we ran the count this page committed to. Twelve qualifying picks carried a recorded closing price. The trigger required twenty. It did not fire, so the four bars above stand exactly as first published on August 11. Other numbers on this page have changed since: we corrected our coverage figures on August 17 and said so above.

We also ran the readings we could have taken instead, because a count is only as honest as the alternatives you did not pick. Counting five further picks that carry a captured closing-line value but not the pre-game price field gives seventeen. Excluding the picks our edge cap suppressed gives four. Every reading we could defend is under twenty.


How we measure the close

This is the part everything else rests on, so here is exactly what we do, including the parts that are imperfect.

The published price. When a pick goes out we record the price we published it at. That is the best number we could find across the books we track, which includes DraftKings, FanDuel, BetMGM, BetRivers, PointsBet, William Hill and Unibet. It is a price you could have taken at that moment.

The close. We then try to record the price again within 90 minutes of first pitch. The 90 minute window matters: a price taken four hours out is not a closing line, it is just an earlier price.

There is a gap between that sentence and our code, and we found it on September 2, 2026, three days before this pond opens.

This page says ninety minutes. Our code looks for a close up to 120 minutes before first pitch. We widened that window on August 12, 2026 and never updated this page to match. Counting every close we can date, 16.8 percent were recorded between ninety and 120 minutes out. The median is 45.8 minutes before first pitch and the furthest out is 119.6 minutes. None is four hours out.

We are not changing the ninety to 120. Widening a published number so that our own records fit inside it is the thing this page exists to prevent. Ninety is what we are aiming at, 120 is what the machine does, and that difference now sits on the page instead of only in our code.

We also cannot date most of our closes. Of 1304 recorded closes, 673 carry no timestamp for the moment we took them, including every close recorded before July 22, 2026 and every run line close. For those we know a close was recorded and we do not know when. Our counts do not filter on this, and none of the numbers on this page has been changed by it.

The two are not always from the same book. Our closing prices come from DraftKings and FanDuel, because those are the closes we can obtain reliably. If a pick was published at the best price across a wider set of books, the comparison is between that price and a DraftKings or FanDuel close. We are telling you this rather than presenting one clean number, because the difference is real and you should know it is there.

The same check found a second gap.

This page lists the books we quote. Pinnacle is not one of them, because most US bettors cannot use it. Our code was picking the best price across a wider list that did include Pinnacle. Of 131 picks we have published, 42 carry a Pinnacle price, which is 32.1 percent of them. That price was real and it was the best one available, but it is not a price we can tell you that you could have taken.

We fixed the selection on September 2, 2026. From that commit forward, the price we publish comes only from the books listed on this page. We have not changed any pick already published. Those rows say what they said when they went out.

One consequence worth stating plainly. When we publish a price from one book and record the close from another, the difference between them is not a number you could have captured yourself. Of the 23 published picks with a recorded close, 20 compare a price at one book against a close at another.

We do not always succeed. Some picks never get a close recorded at all. So we publish our capture rate, which is the share of picks where we successfully recorded one. That number is the foundation under every other number on this page. A grading system whose measurement is invisible would be the exact thing this page exists to refuse, and the capture rate is how you check ours.

Our capture rate, and the bug that made it look worse than it is. As of August 13, 2026, we record a closing price on 143 of 170 attempts for MLB totals and 167 of 178 for MLB moneyline. Over the full life of the project those figures are 143 of 265 and 335 of 360.

The gap between those two pairs is not improvement. It is a bug. Until July 22, 2026, our code asked our data provider for moneyline prices every time, including when the bet was a total. Every over/under lookup came back empty. Ninety-five attempts, nothing recorded, for three weeks, while moneyline ran above ninety percent on the same connection the entire time. Commit e8356e4 fixed it. We did not notice for three weeks, and nothing in our alerting caught it.

We report both eras rather than the one that flatters us, and the commit is in the public log this page links to, so the boundary is checkable and is not a date we picked after seeing the numbers.

And the number behind the number. The rates above answer a narrow question: when we looked for a close, did we find one. They do not answer the wider one: what share of the picks we score ever got looked at. For MLB totals that is currently seven of twenty-five. For moneyline, six of eleven. The bug fixed the looking. It did not fix the coverage, and coverage is the harder problem. We publish both, because publishing only the first would be publishing the flattering half.

Corrected on August 17, 2026. This page previously said eleven of twenty-five for totals and six of ten for moneyline. Both were wrong and both were wrong in our favour. The totals figure overstated our coverage by sixteen points. The moneyline denominator was eleven, not ten. We found this by rebuilding every number on this page from the logs, and we are publishing the method below so anyone can rebuild them too.

How we count. An attempt is one snapshot request written to our logs for a specific pick. A recorded close is a snapshot that came back with a price. Totals are selections of the form OVER or UNDER followed by a number. The two eras are split at commit e8356e4, July 22, 2026.

The capture rates in this section were counted through August 13, 2026, the day it was published, and are frozen at that date. The coverage figures in the paragraph above were recounted on August 17, 2026.

These are event counts, not unique picks. A pick that gets snapshotted twice counts twice. Counted as unique picks instead, the post-fix totals figure is 130 of 146 rather than 143 of 170. We are showing the event count because that is what was published, and we are telling you the difference rather than quietly switching to whichever number reads better.


Why the field size matters

Race enough models and one of them clears any bar by luck alone. That is arithmetic, not cynicism.

So the size of the field is fixed and published before a season starts, and the bar is set with that number already accounted for. A larger field means a higher bar for everyone in it.

Seasons run per pond, not across the lab. Sports do not share a calendar, so each pond has its own opening day, its own field, and its own bar. A pond's field is fixed and published before that pond opens, not before some lab-wide start date. This is what the rule always meant: the bar depends on how many models race each other, and models only race each other inside a pond.

No model may enter a pond after that pond has opened. A model that appears in week three races that pond's next season. Letting one in mid-race would make the published bar a fiction, and we would rather lose a good model for a season than publish a number that is not true.

The Season 1 field for MLB totals

The field is one model: v7. It has been posting publicly since June, and it is the only model whose picks are scored against this pond's threshold.

We also run a second model in shadow. Its picks appear in our public shadow ledger, labelled with its own model id, so you can find them there without asking us. It is not in the Season 1 field, and no model may enter a pond after that pond opens. That is our rule and we are applying it to our own model first.

The reason is worth stating because it does not flatter us. The second model prices on a recalibrated scale, and edge numbers are not comparable across models: a five on one scale is not a five on the other. Applied as written, our published threshold of thirteen percent sits above anything that model has produced or, on our own historical record, could have produced. We could have published a second threshold in its units and put it in the field. We are not going to. A second scoring number published after the first one is a change to the bar, whatever we call it.

A field of one is the weakest version of this. One model measured against a fixed number is not a race between models, and we would rather say so than dress it up. What it is, is a bar published before anything ran, and a model being measured against it where you can watch.

More than one model can fire on the same game. When that happens, only the model already racing in the pond is posted. The other model's selection is still recorded, and the gate scores recorded rows, not what we posted.

Today that is a rule we hold ourselves to, not a machine. The second model is blocked from posting anywhere in the code, so the situation cannot arise yet. If a second model ever enters a pond, this is the rule it enters under.

The two models can also disagree about which side of a game to take. The second model is a recalibration of the first, so it can be less confident on the same game, and less confidence is sometimes enough to move the value from the over to the under. Where that happens, both rows appear in the shadow ledger on opposite sides of the same game.

When a season ends

A pond's season runs from the day that pond opens until its sport's next opening day. For MLB totals, Season 1 runs from September 5, 2026 to MLB's Opening Day in 2027.

A season does one job: a model may only join a pond at a season boundary. We considered ending Season 1 with the regular season on September 27. We are not, because a twenty two day season would reopen our field three weeks after we published it, and that would make the rule we just applied to our own model cosmetic.

The cost is ours. Our second model waits roughly seven months for its chance to race. We are writing this down before that becomes inconvenient rather than after.


Staying in

Clearing the gate is not permanent. A model in a pond is re-evaluated every 50 decided picks against a standing condition: rolling closing line value must stay positive.

A model that fails the standing condition is evicted from the pond. We stop selling it that day.

Eviction is not hidden. We post it, with the numbers that caused it, and the model goes on a public list that only ever grows.

Suspension and eviction are the standard being applied, not the standard being changed. The rules stay put. Models move against them.


When our own measurement breaks

If we fail to capture a closing price for more than ten percent of a racing model's picks in an evaluation window, we suspend that model until the measurement is fixed.

This rule governs models already in a pond. Entry is governed separately and more strictly: a model does not enter a paid pond until it has one hundred decided picks with a verified closing price on every one of them. There is no version of this where a model sells picks on a record we could not fully measure.

This rule is aimed at us, not at the model. A model cannot be graded on picks we failed to measure, and a standard that quietly scored the ones that happened to work would be worthless.

We would rather suspend a model that might be fine than publish a number built on the picks that were easy to check.

We have already failed this test once. The three weeks described above are exactly the failure this rule describes. The rule did not exist when it happened, and even now nothing automated enforces it. That is not an excuse, it is the point of the section below on what is a rule and what is a machine.


What is a rule and what is a machine

Every rule above is a rule we hold ourselves to. Most of them are not yet automated. There is no piece of software that admits a model to a pond, counts down to the next evaluation, or evicts anything on its own. We do that, on the schedule written here, using the numbers published on the public receipts.

We are telling you this because the alternative is implying a machine that does not exist, and you would eventually find out.

What makes the rules checkable is not automation. It is that the inputs are public. Every pick, every price, every close and the capture rate are on the public receipts. If we ever quietly skipped an evaluation or kept selling a model that should have been evicted, the published numbers would not match the published rules, and anyone could see it.

As pieces get automated we will say so here, and the commit log will show when.


What closing line value is

When we publish a pick, we record the price. When the game starts, we record where that price ended up.

If we published a team at +150 and it closed at +120, we got a better number than the market settled on. That is positive closing line value, and it is true whether the bet won or lost.

We grade models on this rather than on win rate because win rate on its own means nothing. A team at -300 is supposed to win 75% of the time. Going 7-3 on heavy favorites is not skill, it is arithmetic working normally.


What closing line value is not

It is not a promise that you make money.

A model can beat the close for a month and still lose money that month. Over a long enough run those two things converge, but "long enough" is longer than most people expect, and anyone who tells you otherwise is selling something.

Clearing this standard is a statement about process, not about outcome. It says a model has been pricing games better than the market settled on. It does not say its next ten picks win, and we will not present it as though it does.

It is also not one number. We publish the price we recorded when a pick was posted, which is the price you could have taken at that moment. It is not the price you get if you act three hours later.

We report both, and we do not average away the difference.


What we score

Every evaluation a model makes at or above the pond's edge threshold, once the game is decided. That includes picks we published and picks we scored but never published.

Our models evaluate far more games than they publish. All of them are logged in advance, timestamped, and scored against the same closes.

This is stricter than scoring only the published picks, not looser. If we graded ourselves solely on what we chose to publish, we would be grading ourselves on a set we selected. Scoring the unpublished ones too makes that impossible.

One condition applies only to published picks: the price a subscriber could actually have taken. Unpublished evaluations have no subscriber price, so that condition is measured only where it means something.

The pond's edge threshold. The pond's edge threshold for MLB totals is thirteen percent. It is the same floor we use to decide whether a pick is worth publishing. We chose thirteen percent on June 15, 2026 as the public posting bar, and we are writing it here on August 20, 2026 as the scoring bar, before the pond opens. That commit raised the floor from nine percent. We moved it up once, in June, before any of this was a gate, and we are not moving it now.

Evaluations below thirteen percent are logged. They do not count toward a gate.

We are naming it now because we had not named it before, and a scoring set that is not published is not a standard. We are not lowering it. Here is what lowering it would do, measured on this model's own MLB totals record since we fixed our closing line lookup on July 22, 2026.

These figures are as of August 20, 2026. They are a snapshot of a measurement that is still accumulating, and we will restate them rather than quietly let them drift.

Floor 13 percent. Decided 26. Measured 12. Mean CLV +1.72. Beat the close 66.7 percent. Distinct measured days 8.

Floor 12 percent. Decided 36. Measured 21. Mean CLV +1.15. Beat the close 61.9 percent. Distinct measured days 13.

Floor 9 percent. Decided 89. Measured 53. Mean CLV +0.71. Beat the close 52.8 percent. Distinct measured days 20.

Floor 5 percent. Decided 186. Measured 114. Mean CLV +0.97. Beat the close 57.0 percent. Distinct measured days 27.

Floor 3 percent. Decided 241. Measured 148. Mean CLV +0.83. Beat the close 55.4 percent. Distinct measured days 28.

At thirteen percent we have twelve measured picks and a mean closing line value of plus 1.72. At three percent we have a hundred and forty eight and a mean of plus 0.83, which is under our own bar of plus 1.5. Lowering the floor buys sample and costs quality. No floor on that table clears all four of our entry conditions.

We could have chosen three percent, called it a principle, and shown you a hundred and forty eight picks instead of twelve. We are keeping thirteen. The cost of keeping it is that this model will not reach a hundred measured picks in the 2026 baseball season. Season one is a measurement period, not a race anything can win this year.

One number we are publishing that does not count. At a three percent floor, this model has a hundred and forty eight measured picks at a mean closing line value of plus 0.83, with 55.4 percent priced better than the close. That is not the gate. It does not count toward one, and no model enters a paid pond on it. It is the largest sample we have of what this model actually does, and it is under our bar. We would rather you see it than not.

One rule removes a pick from scoring, and it is the next section. Three other things narrow what a pond counts, and none of them is a scoring rule: a pond counts only its own sport and market; our edge cap suppresses publication but the pick is still scored; and where we have found a specific data error we name the pick, as we did with two evaluations in the August 20 count. We will not add a scoring removal mid-season.


The one rule that removes picks

When a model's edge on a game comes mainly from a large starting pitcher mismatch, we do not publish that pick and it does not count toward a gate.

We measured that slice before writing this rule. As of August 10, 2026: twenty-three decided picks at roughly minus 43 percent, against plus 0.8 percent for the forty-six other decided picks at the same edge threshold, on a higher average edge. The model's most confident pattern is also its worst one.

Both groups are in the shadow ledger this page links to. The parked group is every decided pick where this rule fired. The comparison group is every other decided pick at or above the same edge threshold, with no window and no trimming.

An earlier version of this paragraph said plus 3.5 percent. That figure came from the most recent forty-three of those forty-six picks rather than all of them, which flattered the comparison. We changed it before this page was announced.

We would rather say that in public than quietly let it into a paid pond.

Two things about how this rule is applied. It is applied by what is recorded on the pick, not by what we meant at the time. Picks published before this rule existed stay in the scoring set. Eighteen picks are in that category. We published them, and a pick cannot be unpublished.

And the rule is frozen for the season. If we ever change what it excludes, that starts a new season rather than restating an old one. Otherwise the denominator moves under the numbers, and the commit log at the bottom of this page would be worth nothing.


A note on the order we did this in

We decided what counts toward a gate before we measured how well we capture closing lines on that set.

The definition we settled on produces a better capture number than the narrower one we had been using, roughly a third rather than roughly a quarter. We are telling you the order because otherwise it looks like we picked the definition that flattered us.

The commit log for this page shows when each was written. That is the point of publishing it here rather than announcing it later.


What we do not claim

We do not claim our models win.

We do not claim past results predict future ones.

We do not delete anything. Losing picks stay up. Bad months stay up. Our first model lost money from June 10 to July 12, 2026, 50-50-2 at -9.7%. That record was finalised on July 14, 2026. Only the note on the freeze file has changed since. The file is public.

When we get something wrong in our own numbers, we post the correction rather than quietly fixing it.


Season 1

The first pond to open is MLB totals. It opens on September 5, 2026. Its field is published above, before that date. The gate numbers on this page are the numbers that field races against.

MLB moneyline is shown and stays empty. An empty pond is not a gap in the display. It is what the standard looks like when nothing has cleared it.

Other ponds open on their own sports' calendars. We said an NFL pond would open when the NFL season does. As of August 24, 2026, it is not opening. We are collecting NFL market data and we will keep collecting it, but what we have is an overlay on the bookmaker's own posted line rather than an independent estimate, and we have no way to grade an NFL pick automatically. A pond needs a model that disagrees with the market and a way to settle the result. We have neither for NFL yet. We would rather say that now than open a pond we cannot score. A pond with a future opening date is not an empty pond. It is a scheduled one.

We would rather run one pond we can measure than four we cannot.

The first verdict will take months, not weeks. The standard requires a real sample spread across real time, and there is no honest way to shorten that. It is entirely possible that nothing clears in Season 1. That outcome would be a result, and we would publish it the same way we would publish a success.

The eviction list ships empty. It will not stay that way forever.


Standard last changed: August 24, 2026. Full history: commit log. Shadow ledger: shadow-ledger.json.

Gamble responsibly. 21+. 1-800-GAMBLER, ncpgambling.org