All work Quant / Trading — technical case study

Most of this is unpredictable. Here is the proof.

SignalDeck is a market-research platform I built to find a directional edge. It found that it did not have one — and said so, in public, against its own claim. This page is how that happened and what caught it.

43.1%live accuracy of the flagship directional model
2,911graded observations, one per symbol-day
56.8%the majority-class baseline it had to beat

Source: STRATEGY_DECK.md §2 — the live record, generated from data/accuracy_registry.json at the grade of 2026-08-04, not typed by hand. Not a backtest: this is the live forward record.

Scroll
A scatter of daily returns forming a structureless cloud around zero, with no visible trend. daily return zero
Daily returns, plotted. The eye keeps looking for a pattern; the measurement says there isn't one at this horizon.
Act 2 — the arithmetic ceiling

70% daily accuracy isn't hard. It's arithmetic.

Sheppard's arcsin relation ties accuracy to information coefficient directly: P(correct) = ½ + arcsin(IC) / π. Feed it a world-class IC of 0.10–0.17 and you get 53–55%. Feed it the IC a realistic system actually carries and you get 51.6%.

To reach 70% you would need an IC of 0.5878 — roughly four times what the best-documented equity signals sustain. The gap is not a matter of effort or compute. It is the relation.

51.6%what an IC of 0.05 buys you
55%the practical ceiling on 1-day direction
0.5878the IC that 70% would demand

Source: SCORE_LOOP_SUPERPROMPT.md — the 1-day cap is proven two independent ways, by this bound and by an exhaustive walk-forward test over roughly a million observations. The percentages above are the relation evaluated directly.

A curve of accuracy against information coefficient, rising steeply then flattening. The 70 percent line is only reached at an IC of 0.5878, far right of the realistic range. 70% target IC 0.05 → 51.6% IC 0.5878 → 70% information coefficient →
P(correct) = ½ + arcsin(IC)/π. The realistic operating range sits at the far left of this curve; 70% lives somewhere no equity signal reliably goes.
Act 3 — the kill

The flagship model graded itself, failed, and switched itself off.

The directional ensemble — the thing the whole platform was built to produce — accrued a live forward record of 43.1% across 2,911 independent (symbol, horizon, UTC-day) observations. It was never judged against a coin flip. The honest competing model is the constant majority-class guess, which scored 56.8% on the same days.

The entire confidence interval sits below that null. That is not "no edge found" — it is significant negative skill. modelhealth set verdict: retired, emitting: false. The model is off.

43.1%live accuracy
−13.7ppskill against the majority-class null
RETIREDemitting: false, enforced in code
Inverting it is not a rescue, and the arithmetic is on record. Flipping 43.1% gives 56.9% against a 56.8% null — a tenth of a point, on a record whose interval runs from 31.1% to 55.9%. An inverted signal only clears that bar if the original sits under 43.2%, and 43.1% clears it by a rounding error rather than a margin. Trading the flip means staking a tenth of a percentage point against an interval twenty-five points wide. The API refuses emitting: true for any -inverted, -relabeled or -flipped variant unless it passes the full canary re-admission gate as a new model.

Source: STRATEGY_DECK.md §2 for the record and its interval; PREDICTION_PROCESS.md — Layer 6 and "Why inversion is not a rescue"; enforcement in daemon/internal/api/modelhealth.go.

A cumulative accuracy line drifting down from fifty percent and settling at 43.1 percent, well below both the fifty percent line and the 56.8 percent majority-class baseline. 56.8% majority-class null 50% coin flip 43.1% — RETIRED 2,911 independent symbol-days →
The live prequential record. It never crosses back above the null, and the retirement is automatic — no grace period, no re-window, no threshold revision after the evidence arrived.
Act 4 — what survived

Direction died. Structure didn't.

The same search that rejected 1-day direction across seven independent tests found things that hold up: not where price goes, but whether a structural state persists. Will this name still be on the same side of its 200-day average in 21 sessions? Will its volatility regime hold?

Drag the conviction slider. Accuracy climbs — but every band reports its own number, never the population average. A low-conviction call says 73.1%, and saying 83.3% instead is exactly how a coin flip ends up wearing someone else's accuracy.

conviction < 0.573.1%interval published only for the ≥0.9 band
mean 21-day forward return+0.41%measured, never inferred from accuracy
These are backtested claims, not a live record. Every structural predictor currently grades PENDING with 0 of the 30 required observations resolved — 3,649 trend21 forecasts are on record and none of them can be graded before 2026-08-07, with first grading under the pre-registered protocol on 2026-08-14. A 21-day horizon cannot resolve sooner, and the waiting is the point. And the null is never 50%: trend persistence alone scores 54–57%. For liquidity21, a naive persistence rule scores the same 87.6% — the accuracy holds, the skill claim does not, and it ships saying so.

Source: daemon/internal/structregime/structregime.go (band tables), data/accuracy_registry.json (PENDING status). All three kinds were replicated 2026-07-17 by an independent re-implementation within two jackknife standard errors.

Bar chart of trend21 accuracy by conviction band: 73.1 percent below 0.5, 90.0 percent from 0.5 to 0.8, 94.6 percent from 0.8 to 0.9, and 97.2 percent at 0.9 and above. 73.1% 90.0% 94.6% 97.2% <0.5 0.5–0.8 0.8–0.9 ≥0.9 trend21 per-band accuracy · top band 95% CI [0.965, 0.978]
trend21 accuracy by conviction band. Also measured: vol21 tops out at 72.0% (CI [0.674, 0.759]), trend63 at 83.7% (CI [0.789, 0.883]).
Act 5 — the trap

The most accurate band is the one that loses money.

Put forward return on the same conviction axis and the two lines cross. Accuracy climbs 73.1% → 90.0% → 94.6% → 97.2%. Mean 21-day forward return climbs with it — +0.41% → +0.58% → +0.79% — and then, at the most accurate band of all, inverts to −0.76%. At the quarterly horizon trend63's top band is worse: −1.30%.

The mechanism is mechanical, not a fluke. Conviction here is distance from the 200-day average, so the most confident calls are by construction the most extended ones — and extended names mean-revert over the following month.

97.2%accuracy at conviction ≥ 0.9
−0.76%mean 21-day forward return of that same band
−1.30%trend63's top band, 63-day horizon

Accuracy is a persistence statistic. It is never an expected return. "97.2% chance it stays above its 200-day average" and "this basket makes money" are different claims, and only the first one is true.

This number has already been corrected once, upward in severity. The platform originally disclosed −0.39%. An adversarial review claimed the real figure was −2.05%. An in-repo replication — 976 symbols, conviction ≥0.9, non-overlapping 21-session windows, one call per symbol-day, n = 18,850 — measured −0.76%, 95% CI [−1.19%, −0.32%]. The reviewer's direction replicated; their magnitude did not, and −2.05% falls outside the interval. So the served number is the measurement: not the kinder original, and not the worse unreproduced figure. The interval is what to quote — the median call is only −0.05%, so the loss lives in a tail. The band was then re-measured a third time, on a universe rebuilt to include companies that have since died: survivorship inflation came to +1.0pp overall, and this band read 97.6% on the clean data. The bias was real and it was small, which is only knowable because someone went and looked.

Source: daemon/internal/structregime/structregime.goforwardReturnFor, corrected 2026-07-26.

Two lines over the same four conviction bands. Accuracy rises steadily from 73.1 to 97.2 percent. Mean forward return rises from plus 0.41 to plus 0.79 percent and then drops sharply to minus 0.76 percent at the highest conviction band. 0% accuracy → −0.76% forward return <0.5 0.5–0.8 0.8–0.9 ≥0.9
The divergence is the finding. Anyone who reads only this figure should leave understanding why "high accuracy" and "makes money" are unrelated claims.
Act 6 — the machinery

None of that is impressive unless something forced it into the open.

Every prediction writes a row to a hash-chained ledger with its probability frozen at prediction time — 326,957 entries, tamper-evident. Six structural predictors were pre-registered on 2026-07-26, with their band tables, resolution rules, nulls and known weaknesses frozen into that chain twelve days before the first observation could resolve. The grader is pinned by commit and its file hash is frozen alongside; editing it appends a visible amendment rather than moving the goalposts.

The refusal rule was decided before the sample arrived: under 30 independent observations is INSUFFICIENT, and under 10 distinct UTC days no interval is published at all — because no interval means no verdict.

326,957rows in the hash-chained ledger
12days claims were frozen before first grading
7.9×n-inflation the independence rule strips out

The rejections, weighted equally

A search that records only its winners cannot be audited. These are in the same ledger as the findings.

RejectedThe 52-week-high magnet

Stocks within 2% of their 52-week high made new highs 76.0–83.3% of the time — against a 21.5% unconditional rate, it looked like a real anchoring effect. Then a volatility-tercile-matched null: stocks far from their high rose the same required distance 85.0–87.8% of the time. Skill of −3 to −9pp over 16,093 observations. It was proximity mechanics, not alpha — a textbook unmatched-null fake.

Pulled from productionThe gap-fill signal

Opening gaps filled at 72–87%, comfortably over the 70% product bar, and shipped. Then the conditioning bug: those rates counted day-0 fills, which are only knowable at the gap-day open. Conditional on surviving day 0 unfilled, the remaining-window fill rate is 45.1–60.6% — every bucket below the bar. Live forecasts were pulled from the platform; the unconditional statement stands, correctly scoped.

Do not shipCointegration pairs trading

Matched nulls killed it: random same-sector pairs earned +0.321%/trade and least-cointegrated pairs +0.360% against the cointegrated arm's +0.339%. Selection edge: +0.017%/trade, i.e. nothing. The mechanism finding is the useful part — correlation rank persists hard (ρ +0.725) while cointegration rank does not persist at all (ρ −0.004), so the durable thing is shared market beta, which every pair already has and no spread position can monetise.

HeldThe nightly regression that counts its own invariants

Twenty named no-lookahead, purge and embargo tests are re-run and counted nightly, so deleting an invariant fails as loudly as breaking one. The same pass re-checks the live database read-only: 151,924 raw outcomes collapse to 19,128 independent ones — a 7.9× inflation factor that, left in, would have turned noise into significance.

Source: daemon/internal/pipeline/researchledger.go (H019, gap-fill conditioning correction, pairs), PREDICTION_PROCESS.md Parts 2–3.

A vertical chain of linked blocks representing the hash-chained prediction ledger. Most links are green; three are highlighted in magenta to mark recorded rejections. rejection, recorded rejection, recorded 326,957 links tamper-evident
The prediction ledger. Rejections occupy the same chain as confirmations — that is the property that makes the record auditable rather than promotional.
Act 7

A search that can prove itself wrong.

A Go daemon ingesting and grading on a local SQLite store, a Next.js surface that refuses to render an accuracy without its band and its caveat, and a Python grader that re-runs the verdicts daily and is allowed to retire the platform's own flagship. Research corpus: 1,854,228 rows of survivorship-aware universe membership, rebuilt point-in-time.

1,854,228rows in the survivorship-aware universe
19,128independent outcomes after de-duplication
151,924raw outcomes before the independence rule
GoSQLiteNext.jsPythonWalk-forward validationWilson & Fisher-z intervals
SignalDeck runs local-only and is deliberately not deployed. It offers no signals, no predictions to act on, and nothing to sign up for. It is a research instrument and a write-up, and that is the whole of it.