Skip to content

State management

Dyson has no single state store. Different kinds of state have different owners, different durability, and different behavior across a restart — and the most common source of operational surprise is assuming one kind behaves like another. This page names each kind and says what survives.

Source references are repo-relative paths in dyson.

Where state lives

State Owner Durable Survives restart
Strategy version and configuration MongoDB Yes, immutable per version Yes
Deployment desired state MongoDB Yes, mutable Yes — reconciled at startup
Runtime lifecycle state Supervisor, in memory No No — rebuilt from desired state
Command idempotency set Supervisor, in memory No No
Strategy position and PnL Strategy engine, in memory No No — operator-seeded
Learner / bias state (kuru-quoter) StateStore on InfluxDB Yes, revisioned Yes — by state_scope
Live order book of record OrderManager, in memory No No — venue must be clean
Adapter order and fill tracking Venue adapter, in memory No No
Fill deduplication history Venue adapter, in memory, bounded No No
Execution availability Health decorator, in memory No No — starts unproven
Post-trade projections InfluxDB Yes, append-only Yes, but not authoritative
Live bus Core NATS No No
flowchart TB
    subgraph D["Durable"]
        MDB["MongoDB<br/>versions, config, desired state"]
        INF["InfluxDB<br/>projections + kuru-quoter state"]
    end
    subgraph P["Process memory — lost on restart"]
        LC["Lifecycle state<br/>+ command idempotency"]
        POS["Positions and PnL"]
        OM["OrderManager<br/>live order records"]
        ADP["Adapter tracking<br/>+ fill dedupe"]
        HL["Execution availability"]
    end
    subgraph T["Ephemeral transport"]
        NATS["Core NATS<br/>no JetStream"]
    end

    MDB -->|"load + reconcile at startup"| LC
    INF -->|"load_latest by state_scope"| POS
    POS --> INF
    OM --> INF
    NATS --> LC

    classDef durable fill:#2e7d32,stroke:#1b5e20,color:#fff
    class MDB,INF durable

Configuration and desired state

MongoDB holds two different things with different mutability. The referenced strategy version is immutable — configuration, source commit, and config hash are fixed once published, which is what makes a run reproducible. The deployment desired state is mutable and exists for restart recovery and startup reconciliation (crates/database/src/strategy.rs):

DeploymentDesiredState Meaning
Running Reconcile up to a running runtime at startup
Paused Reconcile to paused
Stopped Default — the runtime does not begin trading

At startup the supervisor compares desired state against actual lifecycle state and issues synthetic commands (desired-state:<kind>) to close the gap. This is the only state in the system that intentionally drives the process after a restart.

Runtime lifecycle state

One LifecycleState value lives in the supervisor and has exactly one owner. All inputs — market data, lifecycle commands, venue polling — are serialized through that supervisor, so there is no locking discipline to get wrong and no concurrent mutation of lifecycle state.

The full transition table and the two-phase Pause / Drain behavior are on Runtime and strategies. The part that matters for state management: a pending transition is itself state. If orders are not clear when Pause or Drain arrives, the runtime records the target state and completes the transition on a later venue poll. A process that dies mid-pend loses the pending transition entirely.

Command idempotency

The supervisor keeps a set of processed command IDs and returns AlreadyApplied for a repeat. This makes redelivery on the NATS command subject safe — but the set is in-memory and unbounded within a run. Across a restart, the same command ID is accepted again as new.

Strategy position and PnL

Positions and aggregate PnL live in the strategy engine's memory. They are seeded from configuration, not from the venue.

Positions do not recover themselves

Venue balances and collateral are readiness checks, not the signed strategy position. After a restart an operator must supply authoritative starting positions — see Multi-venue strategy.

In the multi-venue actor, per-leg positions sit behind an RwLock in the shared actor state, keyed by venue leg, and are refreshed one leg at a time on each binding callback (strategies/multi-venue/src/actor.rs).

Persisted strategy state

kuru-quoter is the one strategy that persists its own learned state across runs (strategies/kuru-quoter/src/state.rs). It is worth understanding as the template for any future strategy that needs to survive a restart:

  • StateIdentity keys the state by strategy, version, deployment, run, venue, instrument, and actor — plus a state_scope, the stable recovery key deliberately shared across process runs. Recovery matches on scope, not on run_id.
  • revision is monotonic. A non-increasing revision is rejected with StateError::NonMonotonic rather than silently overwriting.
  • InfluxStateStore is a full-snapshot authority, not a delta log. It writes the whole state as JSON, configured with partial writes disabled and WAL sync enabled — the store is treated as an authority, so a half-written state is not acceptable.
  • The same trait optionally records the decision stream alongside durable state for Grafana and post-incident reconstruction.

Order state

OrderManager (external/algo-trading/crates/execution/src/manager.rs) is the in-process book of record. It owns the client-order-ID nonce, a map of live orders, and a resting-order index keyed by (actor_id, instrument) — which is what lets one manager serve multiple venue legs without them colliding.

Orders move through nine states:

flowchart LR
    PS["PendingSubmit"] --> LV["Live"]
    PS --> RJ["Rejected"]
    LV --> PF["PartiallyFilled"]
    LV --> FL["Filled"]
    PF --> FL
    LV --> PC["PendingCancel"]
    PF --> PC
    LV --> PM["PendingModify"]
    PM --> LV
    PC --> CX["Canceled"]
    LV --> EX["Expired"]

    classDef term fill:#2e7d32,stroke:#1b5e20,color:#fff
    class FL,CX,RJ,EX term

Cancellation identity is per-instrument: a venue is addressed either by client order ID or by venue order ID, and only client-ID cancellation can cancel an order still in PendingSubmit.

Because this state is in memory, a strategy process assumes it starts with a clean venue. The multi-venue binary enforces this literally — it rejects startup if unmanaged open orders exist on Lighter.

Adapter-side state

Each adapter keeps its own tracking, and the two current adapters keep different things because their fill sources differ:

Kuru Lighter
Tracks Applied-trade history Per-order TrackedOrder plus filled quantity
Dedupe key Chain transaction / log identity (client_order_index, fill_id)
Structure Bounded VecDeque, oldest evicted Keyed by client order index
On miss Skips rather than risk double-counting Cannot attribute the fill

The Kuru dedupe window is intentionally bounded — it only needs to outlive the window in which the same trade could be seen twice, so it evicts rather than growing forever (crates/venue-adapters/src/kuru/trades.rs). Deduplication is silent by design: a trade that is not ours and a trade already applied both produce nothing.

Execution availability

ExecutionAvailability is per-route state held by the health decorator (crates/venue-adapters/src/health.rs). Its recovery rule is deliberately asymmetric:

  • Any dispatch failure marks the route Unavailable.
  • A route becomes Available again only after a successful health check. An ordinary successful request does not clear a prior failure.

This is why a leg can stay unavailable while apparently working — the health probe, not traffic, is what restores it.

Projections and transport

InfluxDB projections are append-only and non-authoritative. Post-trade records are queued and written asynchronously; the queue exposes counters for enqueued, evicted, retried, and failed batches. Eviction means projection loss is possible under sustained backpressure, so projections are for analysis and reconstruction, never for deriving live position.

NATS is core NATS, not JetStream. There is no durable subscription and no replay. Market data, commands, status, and events are live-only: a process that is down misses what was published while it was down, which is precisely why desired state lives in MongoDB rather than being inferred from the bus.