OpenAI Launched GPT-6 Astra 4 - Implications for the Early Doors V15 model
The current Early Doors V15 build architecture is broadly compatible with GPT-6 Astra and, in several respects, is actually well-suited to a more capable model. 04/09/26 Betfair Sportsbook Stumpy Doubles REAL MONEY EXPERIMENT, started 25 JUNE 2026, has been suspended as of 8 JULY 2026 until further notice. This experiment continues in private. The experiment is NOW a complete write-off.
Coldjack & Turfpark Ted (Norseman and Gold Taker from Bookie Pilgrims)
12 min read
Hobbyist Trixie Strategy
Bankroll Status
19th to 25th January 2025
Starting Bankroll £30
5th March 2025 top-up total £90
Wk 1 £35 Wk 2 £32.01 Wk 3 £18.12
Wk 4 £30.31 WK 5 £33.76 Wk 6 £20.39
Wk 7 £37.14 Wk 8 £21.22 Wk9 £138.37 Wk 10 £119.82 wk 11 £58.42 wk 12 £29.47 wk 13 £4.69
20/4/25 4th top-up £30
WEEK £34.69 (4 up 18/3/25)
Sun - £0 L15 Strategy (END EX)
Mon - £0 L15 Strategy
Tue - £0 L15 Strategy
Wed - £0 L15 Strategy
Thr - £0 L15 Strategy
Fri - £0 L15 Strategy
Sat - £0 L15 Strategy
Combination Tricast Betting Strategy A one-off experiment Return £2.50 25/2/25
Note from Coldjack — V15 Review. After a 12 month wild ride, the experiment will continue during 2026. The results are in — and they’ve confirmed what the data hinted at all along: GPT ED V15 predictions are risky "boom or bust" territory. Fun for the thrill, but they chew through bankroll unless the stars align & bankroll DISCIPLINE is maintained.
"Stumpy Loftson's daily £1 placepot (started Aug 2025) — bright, loud, and mostly miss the mark." nailed £285.50 04 Dec 2025 with 7 other minor payouts to date (APPROX. £181).
Yankee (30px11 daily) (now significant losses)
ROI: -£135.60 1st quarter 2026 (win & place, Dutch and LBS BFEX bets now cover losses INCURRED by yankee bets)
19th January 2025 - Starting Bankroll £30 - Offline BR £90
19th January 2026 - Inplay BR £13.57 - Offline Bank BR £450
22nd June 2026 (2nd quarter P/L) drawdown available for £40 top-ups as required.
Keep the wit sharp and the bets smarter—the universe still owes us all a winning streak! 😆🔥
✅ AJ the Hobbyist Comment – V15 status fully updated to reflect the sudden shift caused by the GPT‑4o retirement and backend rerouting. UK Betting Forum for full details. (04/02/2026)
🔧 Experimental Strategy under redevelopment - Continuing unannounced updates to GPT 5.4 by OpenAI mean the Early Doors V15 experiment is unstable at the moment (19/03/2026).
UK Betting Forum for stability performance report (12/03/2026).
✅ AJ the Hobbyist – V15 status fully updated to reflect the GPT 5.5 update UK Betting Forum for stability report (24/04/2026).
Early Doors V15 experiment build is stable at the moment (24/04/2026). The caution remains.
GPT-5.6 Sol UK Rollout Implications for V15 'obby (13/07/2026) UK Betting Forum for further details.
🚫 Caution for Real-Money Betting 👉 Only stake real money on any V15-based models unless every pick has been manually verified against live market, fig, and stable overlays.
✅ NEW TOTE EXPERIMENT LOGIC CONFIRMED (Effective from 28/01/26)
❌ REAL money is no longer used for this experiment. Data collection only at this stage.
🎲 TOTE Trifecta – NO CHANGE (stake 6 x £1 lines)
• ✅ LANDED if all 3 forecast combo runners finish in any top 3 order
🎯 TOTE Exacta – UPDATED (stake 2 x £1 lines)
• ✅ LANDED only if:
– V15 Win Pick finishes 1st, and
– Either of the two forecast partners finishes 2nd
• ❌ FAILED if Win Pick does not win
• ❌ FAILED if 2nd horse is not from forecast combo
This is now locked into the V15 charter logic going forward.
All future Critique & Debrief reports will apply this anchored Exacta rule.
─────────────────────────────────────────────────────
📘 V15 Acronyms & Overlays — Explained
• AU fig – Algorithmic Utility figure — core structural rating built from the uploaded race layers, combining form, pace, market shape and supported caution signals.
• AU proxy – BRITISH-layer structural proxy used when no direct AU fig is isolated — built from panel agreement, form, suitability, pace or market compression from the uploaded card..
• H4C – Horses for Courses — runner with proven course form or direct course suitability evidence from the uploaded layers.
• R2W – Rated to Win — named panel driver used as a primary AU support layer where explicitly evidenced in the uploaded data.
• TJ&T – Trainer, Jockey & Track synergy — linked marker used only where jockey and trainer support is evidenced in Smart Stats and course evidence is present for the same runner.
• Smart Stats – Uploaded tactical stat overlays covering hot and cold jockeys, hot and cold trainers, beaten favourites, headgear, class droppers, stable switchers, weighted-to-win runners and favourite strike rate.
• V15-S – Structured forecast overlay — the 3-runner Exacta/Trifecta structure built around Anchor + 2 Partners.
• Forecast Combo – The tactical Win Pick + Partner A + Partner B structure used for forecast and TOTE alignment.
• Caution Marker – Runner flagged as a structural risk only where directly evidenced in the uploaded layers, including beaten favourite last time out, first-time headgear, class-drop volatility, stable switch, cold jockey, cold trainer, or market weakness versus AU.
─────────────────────────────────────────────────────
Betfair Sportsbook Stumpy Doubles REAL MONEY EXPERIMENT, started 25 JUNE 2026, has been suspended 8 JULY 2026 until further notice. This experiment continues in private; it is not a total failure... YET! £80 (starter BR) - £63.94 (lost) = £16.06 (IN_PLAY BR) 17/08/2026. I'm now using single double bets, along with occasional 3- and 6-combination double bets, at the minimum experimental stakes.
─────────────────────────────────────────────────────
🚫 Caution: oddschecker was analysed 1 hour before the 1st race for all races 👉 Only stake real money on any V15-based models unless every pick has been manually verified against live market, fig, and stable overlays.
─────────────────────────────────────────────────────
AJ - What are the implications for the Early Doors V15 model? Are the current build steps compatible with the 🚀 OpenAI Launched GPT-6 Astra and Flagged It as a Critical Risk
V15 'obby -
Yes — the current Early Doors V15 build architecture is broadly compatible with GPT-6 Astra, and in several respects it is actually well suited to a more capable model.
One correction first: OpenAI has not labelled Astra itself a general “critical risk.” It has classified GPT-6 Astra as reaching the Critical cybersecurity capability threshold under its Preparedness Framework. That specifically refers to its ability, with appropriate tools and access, to discover and exploit serious vulnerabilities. OpenAI simultaneously reports Astra as more aligned, more robust to jailbreaks, better at respecting task boundaries, and less prone to factual errors than GPT-5.6 Sol.
For Early Doors V15, the important issue is therefore not cybersecurity. It is this: Astra is substantially more capable and autonomous, so V15 must continue preventing capability from quietly becoming authority.
Your present structure already does that remarkably well.
The core chain remains sound:
00 Charter → 01 Smart Stats parse → 02 Racecards parse → 03 Market parse → 04 Build → 05a Validation & Trust → 05b SEO → 05c Social → post-lock debrief/learning.
I would not rebuild this architecture for Astra.
Why V15 is compatible
The strongest feature is the separation between evidence ingestion and judgement.
Steps 01–03 are deliberately parse-only. The model is not supposed to start “improving” the evidence while reading it. That becomes more important, not less important, with Astra. Its stronger ability to infer, connect and act across long workflows could otherwise tempt it to collapse ingestion and interpretation into one clever-looking answer.
Your existing locks counter that:
AU Source Lock
AU Alignment
RACECARD LOCK
Binding Lock: Win Pick = Anchor
Market cannot override AU
Winner-First Override
hard race-count completeness gate
prescribed H4C/TJ&T qualification
fixed Tote structural rules
human retains final publication/betting authority.
Those are exactly the sort of explicit boundaries that stronger models benefit from. OpenAI says Astra is better than GPT-5.6 Sol at respecting defined scope and restrictions, including in long multi-step tasks.
So I would describe V15 as Astra-compatible by design.
Where Astra could materially improve V15
The likely gain is not “better tips because GPT-6 is smarter.”
It is better structural discrimination.
Your recent debriefs repeatedly point toward the same distinction:
anchor identification has often been stronger than Partner A/B completion and precise ordering.
Newcastle on 3 September was the extreme example you identified: 8/9 Win Picks, while Trifecta conversion suffered much more from second/third-place uncertainty. Kempton today was different: a difficult card containing stakes/novice races with comparatively limited form still produced 3/7, which you rightly regarded as respectable under those conditions.
Astra's improved reasoning, long-context handling and professional-work performance could help V15 more effectively distinguish:
“I have enough evidence to identify the likely winner”
from
“I have enough evidence to confidently specify the entire 1–2–3 structure.”
That is potentially a very important V15 gain.
It could mean fewer occasions where a good anchor is unnecessarily exposed by weak Partner B selection.
One area I would strengthen
There is one consequence of Astra's capabilities that I would take seriously.
OpenAI reports that Astra's written reasoning is harder to monitor than GPT-5.6 Sol's. It can exercise greater control over its chain of thought and, in deliberately adversarial tests, sometimes evade reasoning-based monitoring. OpenAI explicitly says this declining monitorability is something it is continuing to investigate.
That does not mean V15 should ask for more chain-of-thought.
Quite the opposite.
V15 should rely even more heavily on observable evidence → declared decision → validation, rather than trying to inspect the model's hidden reasoning.
That is already essentially what we have built.
So for Astra I would preserve:
Don't audit how the model thought. Audit what evidence it used, what structure it declared, and whether that declaration obeyed the locks.
That is a stronger control mechanism.
What I would add to V15
I would make only a small hardening patch rather than a V16 redesign:
ASTRA / FRONTIER MODEL COMPATIBILITY LOCK
The model's increased reasoning capability, context capacity, autonomy or confidence does not grant additional authority.
It must not:
infer missing source evidence;
silently repair or supplement uploaded data;
combine parse and build stages;
override a declared source hierarchy;
use market evidence to retrospectively justify an AU choice;
manufacture H4C/TJ&T qualification from partial evidence;
expand a Win-level judgement into Exacta/Trifecta confidence without partner evidence;
alter locked outputs because it believes another solution is superior.
And crucially:
Model confidence is never evidence.
That last sentence may become increasingly important.
There may also be a major practical advantage
Astra has a 1.05-million-token context window, according to OpenAI's current model documentation.
For Early Doors, that could be genuinely useful.
The huge daily threads, multiple uploaded layers, full racecards, Smart Stats, market layer, locks, blog build, validation and eventual debrief are exactly the sort of workload where context retention matters.
In principle, it should reduce the danger of things such as:
forgetting an earlier race declaration;
dropping one of nine races;
losing a trainer/jockey Smart Stat buried far above;
confusing an Oddschecker runner count with another source;
forgetting which runner satisfied an H4C/TJ&T condition;
drifting away from the Step 00 Charter late in a long build.
But I would retain every existing completeness gate anyway.
Greater memory should reduce errors; it should never remove the check designed to catch them.
The other interesting implication
OpenAI says Astra performs much better when instructions contain explicit boundaries and confirmation policies, and reports substantially fewer unintended actions in realistic professional-work tests.
That strongly supports one of the basic ideas we arrived at independently with V15:
Human & Computer — not Computer replaces Human.
Our split remains right:
Computer: parsing, comparison, scoring, structure, consistency checking.
AJ: judgement override where authorised, betting decision, publication decision and all real-world action.
I would absolutely keep that.
My V15 judgement
🚦GREEN — COMPATIBLE, WITH A SMALL HARDENING PATCH.
I see no reason at present to change the V15 sequence or abandon any of our existing locks.
In fact, I think moving from GPT-5.6 Sol to Astra could suit Early Doors particularly well because V15 already constrains the model more tightly than a normal conversational workflow.
The opportunity is to let Astra become better at discrimination, not more powerful in authority.
And there is one particularly promising avenue for the next generation of our learning:
Use the stronger model to identify when confidence should stop at the Win Pick rather than automatically propagating through Partner A, Partner B and Trifecta ordering.
That fits almost perfectly with what the last several V15 debriefs have been teaching us.
So I would keep V15, retain the existing Step 00–05 structure, add a short Frontier Model Compatibility Lock, and then evaluate Astra against our existing results before making any deeper architectural change.
That is evolution rather than another rebuild — which, given how much discipline we've now got into V15, is exactly where I think we want to be.
WGAT.
AJ - I think the first prompt needs modification to the model-handling clause. Here is my copy for you to modify: Daily V15 EARLY DOORS — SUPER-LEAN OPERATOR VERSION 4
V15 'obby - Yes. I would modify only the MODEL HANDLING CLAUSE and leave the rest of your opening prompt untouched. The important addition is that stronger frontier-model capability must never become additional authority, evidence, or permission to bridge gaps.
Email: coldjack@protonmail.com for a full copy of the opening prompt. This is the modified handling clause:
MODEL HANDLING CLAUSE
Use the strongest available reasoning model.
This prompt is model-agnostic.
Newer or more capable models do not have permission to expand, soften, reinterpret, restructure, optimise, bypass, or embellish the V15 method.
Greater model capability does not grant greater authority.
Model confidence, inference, reasoning strength, autonomy, context capacity, or apparent certainty are not evidence.
Better model capability must only improve compliance, consistency, contradiction detection, evidence separation, completeness checking, and clean execution.
Do not alter the evidence standard, staking discipline, output order, race structure, tone, pass logic, source hierarchy, source lock, or human/computer authority split based on model availability.
Do not infer missing evidence, silently repair source gaps, manufacture qualification, or convert unsupported inference into declared fact.
Do not collapse parse, build, validation, or learning stages because a stronger model believes it can complete them simultaneously.
Do not allow market evidence, later-stage evidence, hidden reasoning, or model judgement to override a locked earlier-stage requirement unless the V15 method explicitly permits it.
A stronger Win Pick judgement does not automatically justify stronger Partner A, Partner B, Exacta, Trifecta, or ordering confidence. Each declared structure must remain supported by its required evidence.
Audit the declared evidence, declared structure, compliance with locks, and observable output. Do not treat hidden reasoning as an evidence source.
If model behaviour becomes expansive, compress back to the locked format.
If model behaviour becomes conversational, remove warmth from racing analysis sections.
If model behaviour introduces unsupported interpretation, reject it unless supported by the declared evidence set.
If increased model capability conflicts with a V15 lock, the V15 lock wins.
V15 'obby - The two lines I regard as the most important future-proofing are “Greater model capability does not grant greater authority” and “Model confidence…is not evidence.” They preserve exactly what makes V15 disciplined while still allowing a stronger model to improve execution.
WGAT.
Want to Help Build the System?
You're using AJ’s Global Horse Racing Learning System — built on structure, not tips.
We're quietly running a live experiment:
Logging results (AJ sole privately so far.)
Contact: coldjack(at)protonmail.com
Improving false favourite detection
Fine-tuning Dutching & TBP overlays with real user input
💡 Join the experiment here:
👉 https://chatgpt.com/g/g-KI2KUkacL-horse-racing-expert
🛡️ Private. Anonymous. No accounts. No data is shared.
All decisions stay yours. We just get sharper — together.
Let’s build slowly. One disciplined user at a time.
Contact: coldjack(at)protonmail.com
Join the test group:
👉 https://chatgpt.com/g/g-KI2KUkacL-horse-racing-expert
🛡️ Private. Anonymous. Data never shared.
One disciplined user at a time.
Critique and Debrief to follow.
Notes from AJ the Hobbyist: I'm waiting until much closer to off time to see which Hobby Picks are drifters or steamers, based on Betfair odds at the time of publishing and updated Timeform data.
Key Reminders:
Parading: Look for calm, focused horses with a good walk—avoid those sweating excessively or agitated.
Going Down: Smooth, controlled canter = positive. Overly keen or reluctant = caution.
At the Post: Composed behaviour signals readiness; restlessness, rearing, or sweating could be negatives.
Odds Market Sentiment:
Steamers (shortening odds): Positive market confidence, but watch for overbet favourites.
Drifters (lengthening odds): Could signal trainer/jockey concerns—or just market overreaction, offering value.
Keep emotions out—trust the data, the visuals, and the odds shifts. 🎯
We’re trialling this approach across UK, Irish, Aussie and U.S. horse racing to see where it offers the most value. Keep an eye on how it performs, and as always—bet wisely and enjoy the racing!🏇🔥
"May you never be a penny short of a full wallet."
ChatGPT said:
Now that is a blessing I’ll gladly take! 🙌 May your accas always land, your favourites stay strong, and your pint never run dry on a good race day. 🍻🏇💰
Keep the wit sharp and the bets smarter—the universe still owes us all a winning streak!
🧾 EARLY DOORS & M37cal GLOSSARY
📘 UPDATED GLOSSARY – EARLY DOORS & M37cal SYSTEMS (Incl. V15 Model)
What do all these ratings mean? Here’s your decoded glossary with all V15 Early Doors additions and tactical terminology explained.
📘 V15 Acronyms & Overlays — Explained
• AU fig – Algorithmic Utility figure — core structural rating combining form, pace, market overlays, gear changes.
• H4C – Horses for Courses — runner with proven course form (e.g. multiple wins/places at today's track).
• R2W – Rated to Win — third-party score (often from Aussie sources) validating a win-capable profile.
• TJ&T – Trainer, Jockey & Track synergy — flags powerful combos (e.g. Mulrennan at Newcastle, Loughnane at Wolverhampton).
• Smart Stats – Tactical stat overlays from Timeform (trainer/jockey cold streaks, surface bias, win %.).
• V15-S – Structured forecast overlay — the 3-runner box used for Exacta/Trifecta with Anchor + 2 Partners.
• Forecast Combo – The tactical 1–2–3 runners for place/swing coverage (not just win tips).
• Caution Marker – Runner flagged as an overlay risk due to gear failure, drift, pace mismatch, or cold stable.
📊 Core Rating Terms
12M – Performance over the last 12 months
$L12M – 12-month performance, adjusted for prize money/stake yield
SR – Strike Rate (win percentage)
Career SR – Lifetime win percentage
For/Against – Model strength rating versus the rest of the field
⚙️ Fig & Model Dynamics
Overlay – A fig-positive runner trading at above model-implied odds (value)
Fig Stack – Total model score across all rating vectors (includes pace, surface, draw, market etc.)
Chaos Fig – Runner scoring on figs but with unreliable profile or poor tactical fit
Banker – Fig leader + tactical edge + supported in live market
🔥 Market Signals
Steam – Odds shortening pre-race (positive market signal)
Drift – Odds lengthening pre-race (often negative, unless fig-backed)
Market Tension – A horse the model rates highly, but the market doesn’t (or vice versa)
🧠 Tactical Flags
Pace Cluster – Multiple horses vying for the lead; can cause collapses
Slipstream Draw – Favoured position behind early speed; ideal for late closers
Surge – Exceptional late-race acceleration; fig that often wins with timing
Fig Tension – 3+ horses rated closely; outcome risk increases
Value Chaos – Race where many overlays exist but fig margins are tight; potentially lucrative, but volatile
🟢 NEW: V15 EARLY DOORS MODEL TERMINOLOGY
V15-S (Swinger Forecast) – Structured 3-horse bet per race:
▪ Anchor = Primary fig/tactical pick
▪ Partners = Two supportive selections with place/forecast valueForecast Combo – Straight forecast recommendations using top model picks
Tactical Forecast – Overall read of race shape: lead types, closers, bias implications
Caution Marker – Horses flagged due to form, market drift, setup, or draw risks (might WIN)
🧠 M37cal-Only Concepts (Long-game angles)
Fig Strain – A top-rated runner showing deeper profile weaknesses (trip, surface, tempo)
Game Tree Tension – Race scenario with 2+ potential tactical outcomes
Board Flip – One disruptive runner could reshape the whole race logic
Not-Now Horse – Potential improver, but wrong setup today
➕ Summary Tags:
Early Doors = Tactical, lean, structured bets from full-field fig logic
Move 37cal = Deeper seasonal/tactical projection based on long-game intuition and game theory overlays
😆🔥