Accuracy Centre
Every forecasting product should publish how it forecasts, how often it has been right, and what it is bad at. This page does that — including when the honest answer is "we don't have enough data yet."
LEXUN's rebuilt platform (runway model v1.0.0) launched in July 2026. Calibration statistics are computed only from forecasts whose outcomes have been reported and resolved against criteria frozen at prediction time. Until at least 10 resolved outcomes per domain exist, no accuracy percentage will appear here — a smaller sample would be anecdote dressed as evidence. Model-domain outcomes resolved: 0 — the editorial forecasts on the public register (6 resolved of 12 issued) are tracked separately and are not evidence for the models. Unresolved (open) forecasts accrue privately per user; aggregate counts will be published when consented aggregation ships. Names: the model shown as Major purchase on every page is opportunity@1.0.0 in the receipts, the ledger schema and this page’s tables — one model, one internal id.
Methodology — runway model v1.0.0
The business cash-runway model answers one resolvable question: will cash stay above zero for the stated horizon?
- Simulation: 5,000 monthly cash paths, run in antithetic pairs. Revenue follows a log-normal multiplicative step about the user’s stated growth rate — exp(ln(1+g) − σ²/2 + σ·ε), ε ~ N(0,1) — whose expected value is exactly 1+g at any volatility and which is strictly positive, so no month can drive revenue to zero; costs grow at the stated rate; committed one-off cash events land in their stated months.
- Determinism: the random seed is derived from the inputs (or supplied explicitly), so identical inputs always produce identical results. This is enforced by automated tests.
- Uncertainty display: probabilities are shown as bands rounded to 5-point steps, widened by Monte-Carlo standard error and by input quality. Raw values are stored only for calibration scoring.
- Evidence gates: the model refuses to run when core inputs are marked unknown, or when every input is an assumption. It does not fabricate values.
- Confidence: derived structurally from how many core inputs are facts versus assumptions — never from model enthusiasm.
- Adversarial review: a separate challenge module (different method: straight-line burn heuristic plus a rule library) attacks every result; material disagreement between the two models is recorded and shown.
- Resolution: the outcome criterion and due date are frozen at prediction time and cannot be edited afterwards. Outcomes are scored with the Brier rule, which is defined with the rest of the vocabulary in the decision glossary.
Methodology — career model v1.1.0
The career-change model answers only the resolvable question: if you make the change, does your financial buffer survive the income transition? Income ramps linearly to your expected level over your stated months, with monthly volatility shocks; essential costs drain the buffer; 5,000 seeded paths. Same disciplines as runway: deterministic seeds, evidence gates, banded display, frozen resolution criteria. Version 1.1 delivers the independent-heuristic cross-check promised at 1.0: a deterministic straight-line drain model (separate code path, no Monte-Carlo) attacks every result, and material disagreement between the two is recorded and shown.
Stated honestly: no defensible base rate exists for career-change success (definitions vary too much to pool), so this model does not offer one. Context figure only: ~2.9m UK workers changed jobs in 2025 (Indeed Hiring Lab, citing ONS). Known weaknesses: partner income, benefits during transition, emergency costs and the return-to-employment option are not modelled.
Methodology — major-purchase model (opportunity@1.0.0)
You will see this model called major purchase in the product and opportunity@1.0.0 in receipts, version tables and the changelog. They are the same model: opportunity is the internal identifier it was released under, and renaming it would break every reproducibility id already issued against it.
The major-purchase model answers only the resolvable question: if you take on this upfront cost and recurring commitment now, does your financial buffer survive the horizon? The upfront cost leaves the buffer at month zero; thereafter income varies with monthly volatility shocks against fixed essential costs plus the commitment; 5,000 seeded paths. It also reports the largest monthly commitment at which your median path still survives, and a waiting analysis that honestly notes it does not model price rises while you wait. Same guarantees as the other models, including a paired independent surplus-arithmetic cross-check.
Stated honestly: no defensible base rate exists for whether a major purchase "works out" (outcomes vary too much to pool), so this model does not offer one. Context only: 13.1 million UK adults (24%) had low financial resilience in May 2024, and 42% could not cover 3+ months of living costs if their main income stopped (FCA Financial Lives 2024). Known weaknesses: price and interest-rate changes while deciding, one-off repair shocks, resale/exit value, and essential-cost inflation are not modelled.
How accuracy will be reported
- Calibration by domain (runway, career, opportunity) and by horizon bucket (0–3m, 3–12m, 12m+) — never one universal percentage.
- Reliability tables (mean forecast vs actual frequency per probability bin), sample sizes, and confidence intervals.
- Per model version, so improvements and regressions are visible across versions.
- Retractions and corrections, listed on this page permanently.
The three clocks — how the register avoids hindsight
An honest forecasting record has to keep three different times apart: when something happened, when the figure about it was published, and when this register learned it. Mixing them is how systems quietly cheat — a “prediction” scored against a revised figure the forecaster could never have seen is not a prediction.
Every ledger entry carries the three clocks as separate frozen fields, and the resolution tool enforces them: issued_at and data_cutoff record what the forecaster could know and when; source_publishes_at records when the deciding figure becomes public — resolving before it is refused outright as “a guess wearing a timestamp”; and the outcome records source_published_at and resolved_at with the figure as first published, so a later revision can never rewrite a score. A worked example from the register: the July 2026 CPI forecast was issued 14 August with a cutoff of the same day, its source published 19 August at 07:00, and it resolved that morning against the first-published 2.9% — the criterion says in terms that later revisions do not change the resolution.
The reference-class figures shown beside results follow the same rule from the other side: each is a bundled snapshot with its publication and retrieval dates displayed, marked as context that does not enter the calculation — and never described as live.
Outcome verification grades
Not all resolutions are equally strong, so every resolved outcome now carries a verification grade, computed from the evidence the ledger actually holds — never assigned by hand. A means automatically resolved from an authoritative source by a connector: impossible in this architecture, because no connectors exist, so no outcome can carry it yet and none does. B means manually resolved against an authoritative published figure — the exact first-published sentence, its URL on the frozen source host and its publication time all sealed in the ledger; all six resolved outcomes on this register are grade B. C and D cover corroborated and unverified user reports — private workspace outcomes, which never enter this public record. E is unresolved or ambiguous. The grades, with each outcome’s grade and the full model coverage and release registry — every model LEXUN ships, every model it merely proposes, and the wall between them — ship machine-readable at /model-registry.json, derived from the model code itself and drift-blocked by its own release gate.
The three histories — reality, knowledge, expectation
Most forecasting systems keep one history: what happened. An accountable one needs three, kept apart. Reality history is what later evidence says happened — here, every resolved outcome as first published, sealed with its source sentence and URL, never rewritten by revision. Knowledge history is what could genuinely be known at the time — every entry’s data_cutoff, the world-state snapshot frozen into each decision, and the receipt that proves nothing later leaked in. Expectation history is what LEXUN believed would happen next — and it is the one histories quietly lose, because keeping it means leaving your worst calls on display.
This register keeps it. Every probability band ever issued stays published; a change of mind is a new entry that names the old one and the reason, never an edit. The table below is that expectation record, regenerated from the ledger at every release — the corrections with the reasons in the record’s own words:
| Earlier belief | What changed, in the record’s own words | Standing belief |
|---|---|---|
| 001 issued 2026-08-14 · 62–78% · withdrawn | Issued and PUBLISHED on lexun.co.uk on 14 August 2026, then withdrawn the same day. The reasoning treated the residual probability as cut risk. The July 2026 MPC minutes record a 6-3 vote with three members — Greene, Mann and Pill — voting to INCREASE Bank Rate to 4%, and the Committee judging… | 007 70–84% · open |
| 005 issued 2026-08-14 · 80–90% · withdrawn | Issued and PUBLISHED on lexun.co.uk on 14 August 2026, then withdrawn the same day. The due date was set to 2026-08-18T08:00:00Z, which is 09:00 BST. The Insolvency Service publishes its monthly statistics at 09:30 BST — verified from the GOV.UK content API, which gives first_published_at… | 009 80–90% · resolved |
| 006 issued 2026-08-14 · 88–96% · withdrawn | Issued and PUBLISHED on lexun.co.uk on 14 August 2026, then withdrawn the same day. It said 88-96% that Bank Rate would not rise by 5 November, on the reasoning that every indicator pointed away from tightening. The July 2026 MPC minutes show three of nine members already voting to hike and risks… | 008 38–56% · withdrawn |
| 008 issued 2026-08-14 · 38–56% · withdrawn | Issued 14 August 2026 and withdrawn the same day, before deployment. The due date was 2026-11-05T12:00:00Z. The Bank of England publishes the Monetary Policy Summary and minutes at 12 noon UK, and British Summer Time ends on 25 October 2026, so on 5 November 12:00 UK IS 12:00 UTC. The deadline was… | 010 38–56% · open |
12 entries ever issued · 4 since withdrawn · 0 deleted · 0 edited. A register that could quietly rewrite its expectation history would have nothing to be scored against.
What this page does not hold: an archive of other forecasters’ expectations — IMF and central-bank projection vintages, professional survey rounds. That is real evidence and a real gap, recorded as such in the gap registry rather than papered over; until it closes, the expectation history here is LEXUN’s own, complete and uncut.
Anatomy of one decision record — the nine fields
What the engine writes for every run, in the order it writes it. Every field below is live: it is filled by the engine or by you, and nothing on this list is planned or partial. See the exported JSON →
- The question, as you asked itLive
- Every input, classed fact or assumptionLive
- Probability band and confidenceLive
- Resolution criterion and due date, frozen at prediction timeLive
- The outcome, written once when reality answersLive
- Brier score and the domain calibration it updatesLive
- The action you actually took, in your wordsLive
- Why reality differed, in your wordsLive
- Lessons carried into the next predictionLive
Reproducibility — a worked receipt
The example on the homepage is a real engine run. Anyone entering the same inputs reproduces it exactly:
model: runway@1.0.0 · seed: 564 (input-derived) · paths: 5,000
reproducibility_id: rw_52f9a7c3_234
output: survival 60–70% (band) · runway p10/p50 9/12+ months (12+ = still solvent when the 12-month window closed) · verdict "Probably safe — watch it"
Reference data & source registry
The registry is also shipped machine-readable at /source-registry.json — generated from the same code the engines load (never hand-typed twice), with each source’s authority tier (1 official · 2 recognised research · 3 reputable industry · 4 secondary reporting), licence, publication and retrieval dates, and the register’s resolved-outcome sources. It states plainly that this static site has zero live data connectors: every figure is a bundled, dated snapshot. A release gate blocks any drift between the shipped file and the code. The file also carries a coverage matrix (every figure LEXUN actually holds, with tier, dates and where it is used) and a blind-spot register (what it lacks, classified honestly — from “discoverable and open” to “unavailable in the current architecture” — each with why and the acquisition path).
| Source | Figure used | Published | Retrieved | Status |
|---|---|---|---|---|
| ONS Business Demography, UK: 2024 OGL v3.0 · official statistics |
5-year survival of UK businesses born 2019: 38.4% | 2025-11-20 | 2026-07-30 | Bundled snapshot — not live |
| Indeed Hiring Lab — Job Switching in the UK industry source, citing ONS |
UK workers who changed jobs in 2025: ~2.9 million (context only) | 2026-03-04 | 2026-07-27 | Bundled snapshot — not live |
| FCA — Financial Lives survey 2024: key findings OGL v3.0 · official survey |
UK adults with low financial resilience, May 2024: 13.1m (24%); limited savings buffer: 42% (context only) | 2025-05-16 | 2026-07-30 | Bundled snapshot — not live |
| Bank of England — Official Bank Rate history BoE Database terms · official series |
Every rate change 1975–2025: 258 observations, shipped at /data/boe-bank-rate.json (reference context only — not a model input) | maintained series | 2026-08-27 | Bundled snapshot — not live |
| ONS — CPI annual rate (D7G7, MM23) OGL v3.0 · official statistics |
Yearly averages 1989–2025: 37 observations, shipped at /data/ons-cpi-annual.json (current vintage only, stated in the file; not a model input) | 2026-08-19 | 2026-08-27 | Bundled snapshot — not live |
No source is ever described as "live" unless an actual connection exists and has recently refreshed. Today, none are live; all reference data is bundled snapshots with retrieval dates.
Rate regimes — historical context, not prediction
The first Historical Intelligence dataset: the Bank of England’s full official Bank Rate change history — every change since 1975, bundled with its source, licence and retrieval date, and six anchor values cross-checked against independently known history before bundling. It exists to answer one honest question: is today’s rate environment ordinary or unusual by the record’s own standard? No released model takes the rate as an input, so this series is context beside results, never inside them.
LEXUN groups the record into named bands — near-zero (below 1%), low (1–3%), moderate (3–6%), restrictive (6–10%), extreme (10% and above). The thresholds are a naming convention, stated so you can disagree with them; they are not analysis.
Computing from the bundled series…
Two variables against history — an honest analogue
With two bundled series — the Bank Rate record and the ONS CPI yearly averages (37 years, 1989–2025, shipped at /data/ons-cpi-annual.json, current vintage only and the file says so) — LEXUN can ask a narrow, honest question: which past years most resembled today on these two variables alone? Today’s pair is real on both sides: the 12-month CPI rate as resolved by the public register against its first-published source (July 2026: 2.9%), and the current Bank Rate from the bundled record. One stated mismatch: today’s 12-month rate is compared with calendar-year averages — said here rather than hidden.
Computing from the bundled series…
Two variables are not a regime. A real regime comparison needs many more series — unemployment, credit, wages, housing, energy — which this static bundle does not hold; the blind-spot register says exactly that. And even a full state vector would not make history a prophecy: the years listed above went on to different futures for reasons these two numbers never carried.
Similarity is not repetition. This series records where the rate has been; it says nothing about where it will go. Two periods that look alike in one variable can end differently for reasons no single series carries — which is exactly why the register freezes forecasts and scores them, rather than reading history as prophecy.
The baseline every model must beat — scored on 36 years
Forecast science starts with a humiliating rule: before any model earns trust, it must beat the do-nothing line. Below, the simplest possible CPI forecast — next year’s average will equal this year’s — is walked forward through the bundled ONS series with no hindsight: each year’s prediction uses only the previous year’s figure, exactly as a forecaster standing in that year could have. Computed in your browser from the same file you can download.
Computing from the bundled series…
What this exhibit is for. In quiet years the naive line is very hard to beat — and in regime breaks it fails catastrophically, which is the honest shape of the forecasting problem. Any model LEXUN ever promotes for a variable like this must beat this baseline out of sample, and the years in the misses table are why “history as prophecy” is refused everywhere on this site: the past is the best guide available right up until the moment it is the worst one.
Known weaknesses (runway v1.0.0)
- Revenue shocks are modelled as smooth monthly volatility; single lumpy events (losing an anchor client, a tax bill) are only captured if entered as committed cash events.
- Receivables timing (booked vs banked revenue) is not modelled.
- The bundled reference class pools all UK industries and sizes; sector-specific base rates are not yet included, and the 1-year survival rate was not captured in our snapshot, so it is not shown.
- Outcome data is self-reported until verified-evidence attachment ships; self-reporting can bias calibration optimistically.
- No calibration history exists yet — the model's real-world track record is unproven, which is exactly what the outcome loop exists to fix.
Model versions
| Version | Released | Notes |
|---|---|---|
| runway@1.0.0 · core@1.0.0 | 2026-07 | Initial release. |
| career@1.0.0 · learning@1.0.0 · intelligence@1.0.0 · governor@1.0.0 | 2026-07 | Career financial-resilience model; outcome-gated learning layer; live governor checks. |
| career@1.1.0 · opportunity@1.0.0 · challenge@1.1.0 | 2026-07 | Career independent-heuristic cross-check delivered as promised at 1.0; major-purchase affordability model with its own paired cross-check. Test suite: 120 automated checks, all passing. |
| reality@1.0.0 · core registry 7 sources | 2026-07 | Decision Reality layer: monitored assumptions, materiality checks, forecast versioning, trust receipts, .lexun export. Registry extended with four sourced Find-My-Money figures. Test suite: 233 automated checks, all passing. |
| core@1.1.0 · vault@1.0.0 | 2026-07 | The LEXUN Key: the saved record is encrypted at rest with AES-256-GCM under a PBKDF2-HMAC-SHA256 key at 600,000 iterations. No model maths changed — core gained a persistence seam only, so every engine version above is unaffected. Test suite: 390 automated checks, all passing. |
| reality@1.1.0 | 2026-08 | The decision record gained what was missing from it: a free-text account of the action actually taken and a separate free-text account of why reality differed, plus lessons carried forward into the framing screen for the domain they were learned in. Additive record fields only — no model maths, seed derivation or hash function changed, so every version above is unaffected and all 96 frozen results resolve unchanged. Test suite: 390 automated checks, all passing. |
| router@2.0.0 | 2026-08 | Decision router rebuilt as a capability firewall: three outcomes (supported, clarify, unsupported) replace “best guess” selection; restricted categories are refused with a supported reframing where one exists. Routing decides which door opens — it never touches model output, so every engine version above is unaffected and all 96 frozen results resolve unchanged. Test suite: 397 automated checks, all passing. |
| router@2.1.0 | 2026-08 | Firewall keyword widening and the advertised-question intent contract; the first public forecast resolved (editorial board — unemployment, TRUE, Brier 0.0100) with every counter reading from the ledger. No engine maths changed; 96 frozen results unchanged. Test suite: 711 automated checks, all passing. |
Boundaries
LEXUN models business and career decisions. It refuses medical, legal and immigration questions, routes personal-crisis language to human help, and provides decision intelligence — not financial, legal or medical advice.
Watch the record grow
A zero above means no outcome in that domain has resolved — not that nothing is measured. Leave an email and we will write to you when outcomes resolve and the first calibration table publishes: the misses included, because a record without misses is marketing.
One address, one purpose. It is used for record updates and nothing else — see Privacy.