Learn from every trade — safely

traditional algos hold their parameters fixed until a human changes them. ttTrader's shared learning stack closes the loop: each closed trade becomes a labelled observation, that evidence tunes the algo's own gates and exits, and no learned layer is allowed to change a live order until a promotion gate says it has earned the right to.

Shadow-first
New layers observe before they serve
Promotion-gated
PROMOTE / HOLD / DEMOTE over N realized labels
Causal
Forward-looking features are unrepresentable
decimal_t
Learning math stays exact, end to end
The safety claim, stated plainly
The platform can learn without letting a new model change a live order. Every learned layer is shadow-first and promotion-gated on purpose — it measures before it serves. That is a selling point, not a caveat: evidence, not an untested hypothesis, decides when a model is allowed to act.

The learning & feedback loop

The diagram below follows one virtual sample — algoHELIX quoting a fictional HLX-PERP market (not a production algorithm or instrument) — through a single trade and back around the loop. Follow the numbered phases from context to persistence and the dotted edges that carry state into the next session.

Virtual sample — the loop closes once per trade, and state carries into the next session.

1

Context — strictly causal

algoRegimeFeatures_c builds a 10-feature context vector where a forward-looking feature is structurally unrepresentable, and algoRegimeLabels_c reduces it to a documented, instrument-agnostic label. The look-ahead is bounded and exposed, never hidden.

2

Decide — three learned actions

algoParticipationGate_c is an L2 contextual bandit that picks stand-down, reduce-only, or serve from the shared regime. When it serves, algoValueFeedback_c exposes trustedKEdge() and maturityVerdict() so the caller knows a warm learner from a cold one. Sizing is explored by algoProbeSizing_c and algoProbeLedger_c, then served by a LinUCB policy.

3

Execute — protection is mandatory

algoProfitCapture_c ratchets profit once favorable excursion clears a threshold and is provably monotone — it never widens risk. algoExitShape_c learns one bounded stop/hold distance pair per session.

4

Outcome — trusted before it teaches

The recorded path and realized P&L pass algoCostModel_c for a net edge, then algoRealizedPnlTrust_c. A mis-scaled venue P&L is flagged WITHHOLDING — distinct from a genuine COLD — so a bad number can never poison learning.

5

Feedback — one trade, many teachers

A validated label feeds every learner at once: algoExitCounterfactual_c replays the path to find the exit that would have done best per market-condition bucket; algoExitShape_c shrinks on losing evidence and grows on winning evidence; algoValueFeedback_c compares realized against quoted edge; the online logistic classifier (and its disagreement-based ensemble) trains; the regime dataset-trainer-posterior chain updates; and the probe ledger folds the reward into the bandit arm.

6

Govern — evidence, not enthusiasm

algoPromotionPolicy_c is the single gate every learned layer passes. Over a rolling window of realized labels it returns PROMOTE, HOLD, DEMOTE, COLD, or INSUFFICIENT_EVIDENCE. A window with no closed round trip is not a label. Promotion is advisory: a person enables serving.

7

Persist — validate, then commit

algoLearningValidity_s records may this model serve, and why. Versioned, append-only, per-contract state via saveState/loadState with algoContractFingerprint and layoutHash() means a state file can never silently re-open a closed path or apply another contract's learning.

Adaptive learning & self-tuning

Components that turn realized outcomes into tuned behaviour, shared by every strategy instead of copied per algo.

algoValueFeedback_c

Self-tunes entry gates — kEdge, capture fraction and pyramid threshold — from realized edge. trustedKEdge() and maturityVerdict() let a caller tell a warm learner from a cold one before acting on it.

algoExitShape_c

Learns one bounded stop/hold distance pair per session. Losing evidence shrinks exposure; winning evidence grows it — the exit adapts without becoming unbounded.

algoExitCounterfactual_c

Replays every closed trade over its recorded path to find the exit that would have done best, per market-condition bucket — learning from roads not taken.

algoProfitCapture_c

Mandatory-protection profit ratchet: locks in profit once favorable excursion clears a threshold. Provably monotone — it can tighten but never widen risk.

algoParticipationGate_c

L2 learned participation: stand-down, reduce-only or serve, chosen by a contextual bandit from the shared regime rather than fixed thresholds.

algoProbeSizing_c + algoProbeLedger_c

Cold-start exploration plus a LinUCB sizing policy — shadow-first, so a new sizing idea is measured before it is served.

algoOnlineLogistic_c

A decimal online classifier with drift-freeze: when fast and slow loss EMAs diverge, the model stops updating instead of chasing a regime it no longer fits.

algoOnlineLogisticEnsemble_c

A bagged ensemble over one label stream. Member disagreement is a live uncertainty readout — no held-out set required — and it is designed to be promoted only after it beats the primary model on live labels.

Causal regime intelligence

A regime model is only useful if it cannot see the future. The B4 posterior is built so it can't.

algoRegimeLabels_c

A documented, instrument-agnostic regime label — RANGE, TREND_UP, TREND_DOWN, CHOP — with a bounded, explicitly exposed look-ahead.

algoRegimeFeatures_c

A strictly causal 10-feature context vector. A forward-looking feature is structurally unrepresentable — not filtered out, impossible to express.

algoRegimePosterior_c

The end of the chain: tape bars → labelled rows → four calibrated one-vs-rest bundles → a shadow gate that can only stand an algo down. No configuration can let it force a trade.

One-way safety
algoRegimeDataset_c → algoRegimeTrainer_c → algoRegimePosterior_c. The posterior is a veto, not a trigger: it may reduce or stand an algo down, never open risk on its own.

Offline training & safe deployment

Reproducible training and one promotion gate that governs every learned layer.

algoOfflineTrainer_c + algoWeightBundle_c

Deterministic time-series cross-validation, Platt scaling and Mondrian-conformal calibration, and byte-identical versioned weight bundles: the same corpus and config always produce the same bytes.

k-fold + gap conformal alpha contentHash()

algoFeatureRegistry_c

Named, unit-typed, content-hashed feature sets. A weight vector cannot be applied to a reordered layout — the failure mode that silently corrupts a model becomes a rejected apply.

algoPromotionPolicy_c

One evidence gate for every learned layer: PROMOTE / HOLD / DEMOTE over N realized labels. Quiet windows are not labels, so a no-trade stretch cannot masquerade as break-even evidence.

Trust, safety & persistence

The proof points behind the headline: learning that cannot be poisoned and state that cannot lie.

algoRealizedPnlTrust_c

A boundary that stops a mis-scaled venue P&L from poisoning learning. It raises a WITHHOLDING verdict that is deliberately distinct from COLD — "not enough data" and "this number is suspect" are different facts.

algoLearningValidity_s

Persisted may this model serve, and why metadata: label balance, warm status, drift/freeze and reset history travel with the model.

saveState / loadState

Versioned, validate-then-commit, per-contract, append-only state with algoContractFingerprint and layoutHash(). A state file can never silently re-open a closed path or apply another contract's learning.

Shadow-first

New layers observe and measure before they are allowed to serve.

Promotion-gated

No learned layer changes a live order until evidence says it may.

Fail-closed state

A rejected or mismatched state file is refused, not partly applied.

Venue & instrument adaptation

Learning reaches the execution layer too — but always with an honest, causality-respecting rule attached.

quotePolicy.h

Regime-conditional quote side (adaptiveQuoteModeFromRegime) and tick-threshold repricing (adaptiveQuoteNeedsReprice) — the quote adapts to the regime the algo is actually in.

algoFallbackInstrument_c

Session-aware fallback when the primary market is shut: closure is declared, never inferred, so the algo never trades a market it only assumes is open.

algoCotPositioning_c

Public weekly COT positioning reduced to z-score and percentile, with a publication-instant no-look-ahead rule — the positioning is only visible from the moment it is actually published.

algoCostModel_c

Fee-, spread- and latency-aware round-trip cost plus an admission gate — an edge that does not clear real cost is not admitted.

Honesty note. Most of these layers are shadow-first and promotion-gated on purpose — they measure before they serve. That is a selling point, not a caveat: the platform can learn without letting a new model change a live order. The trust, promotion and state pieces are the “why it is safe” proof behind the headline that the algo platform now learns.