State management¶
Dyson has no single state store. Different kinds of state have different owners, different durability, and different behavior across a restart — and the most common source of operational surprise is assuming one kind behaves like another. This page names each kind and says what survives.
Source references are repo-relative paths in dyson.
Where state lives¶
| State | Owner | Durable | Survives restart |
|---|---|---|---|
| Strategy version and configuration | MongoDB | Yes, immutable per version | Yes |
| Deployment desired state | MongoDB | Yes, mutable | Yes — reconciled at startup |
| Runtime lifecycle state | Supervisor, in memory | No | No — rebuilt from desired state |
| Command idempotency set | Supervisor, in memory | No | No |
| Strategy position and PnL | Strategy engine, in memory | No | No — operator-seeded |
Learner / bias state (kuru-quoter) |
StateStore on InfluxDB |
Yes, revisioned | Yes — by state_scope |
| Live order book of record | OrderManager, in memory |
No | No — venue must be clean |
| Adapter order and fill tracking | Venue adapter, in memory | No | No |
| Fill deduplication history | Venue adapter, in memory, bounded | No | No |
| Execution availability | Health decorator, in memory | No | No — starts unproven |
| Post-trade projections | InfluxDB | Yes, append-only | Yes, but not authoritative |
| Live bus | Core NATS | No | No |
flowchart TB
subgraph D["Durable"]
MDB["MongoDB<br/>versions, config, desired state"]
INF["InfluxDB<br/>projections + kuru-quoter state"]
end
subgraph P["Process memory — lost on restart"]
LC["Lifecycle state<br/>+ command idempotency"]
POS["Positions and PnL"]
OM["OrderManager<br/>live order records"]
ADP["Adapter tracking<br/>+ fill dedupe"]
HL["Execution availability"]
end
subgraph T["Ephemeral transport"]
NATS["Core NATS<br/>no JetStream"]
end
MDB -->|"load + reconcile at startup"| LC
INF -->|"load_latest by state_scope"| POS
POS --> INF
OM --> INF
NATS --> LC
classDef durable fill:#2e7d32,stroke:#1b5e20,color:#fff
class MDB,INF durable
Configuration and desired state¶
MongoDB holds two different things with different mutability. The referenced
strategy version is immutable — configuration, source commit, and config
hash are fixed once published, which is what makes a run reproducible. The
deployment desired state is mutable and exists for restart recovery and
startup reconciliation (crates/database/src/strategy.rs):
DeploymentDesiredState |
Meaning |
|---|---|
Running |
Reconcile up to a running runtime at startup |
Paused |
Reconcile to paused |
Stopped |
Default — the runtime does not begin trading |
At startup the supervisor compares desired state against actual lifecycle state
and issues synthetic commands (desired-state:<kind>) to close the gap. This is
the only state in the system that intentionally drives the process after a
restart.
Runtime lifecycle state¶
One LifecycleState value lives in the supervisor and has exactly one owner.
All inputs — market data, lifecycle commands, venue polling — are serialized
through that supervisor, so there is no locking discipline to get wrong and no
concurrent mutation of lifecycle state.
The full transition table and the two-phase Pause / Drain behavior are on
Runtime and strategies. The part
that matters for state management: a pending transition is itself state. If
orders are not clear when Pause or Drain arrives, the runtime records the
target state and completes the transition on a later venue poll. A process that
dies mid-pend loses the pending transition entirely.
Command idempotency¶
The supervisor keeps a set of processed command IDs and returns
AlreadyApplied for a repeat. This makes redelivery on the NATS command subject
safe — but the set is in-memory and unbounded within a run. Across a restart,
the same command ID is accepted again as new.
Strategy position and PnL¶
Positions and aggregate PnL live in the strategy engine's memory. They are seeded from configuration, not from the venue.
Positions do not recover themselves
Venue balances and collateral are readiness checks, not the signed strategy position. After a restart an operator must supply authoritative starting positions — see Multi-venue strategy.
In the multi-venue actor, per-leg positions sit behind an RwLock in the shared
actor state, keyed by venue leg, and are refreshed one leg at a time on each
binding callback (strategies/multi-venue/src/actor.rs).
Persisted strategy state¶
kuru-quoter is the one strategy that persists its own learned state across
runs (strategies/kuru-quoter/src/state.rs). It is worth understanding as the
template for any future strategy that needs to survive a restart:
StateIdentitykeys the state by strategy, version, deployment, run, venue, instrument, and actor — plus astate_scope, the stable recovery key deliberately shared across process runs. Recovery matches on scope, not onrun_id.revisionis monotonic. A non-increasing revision is rejected withStateError::NonMonotonicrather than silently overwriting.InfluxStateStoreis a full-snapshot authority, not a delta log. It writes the whole state as JSON, configured with partial writes disabled and WAL sync enabled — the store is treated as an authority, so a half-written state is not acceptable.- The same trait optionally records the decision stream alongside durable state for Grafana and post-incident reconstruction.
Order state¶
OrderManager (external/algo-trading/crates/execution/src/manager.rs) is the
in-process book of record. It owns the client-order-ID nonce, a map of live
orders, and a resting-order index keyed by (actor_id, instrument) — which is
what lets one manager serve multiple venue legs without them colliding.
Orders move through nine states:
flowchart LR
PS["PendingSubmit"] --> LV["Live"]
PS --> RJ["Rejected"]
LV --> PF["PartiallyFilled"]
LV --> FL["Filled"]
PF --> FL
LV --> PC["PendingCancel"]
PF --> PC
LV --> PM["PendingModify"]
PM --> LV
PC --> CX["Canceled"]
LV --> EX["Expired"]
classDef term fill:#2e7d32,stroke:#1b5e20,color:#fff
class FL,CX,RJ,EX term
Cancellation identity is per-instrument: a venue is addressed either by client
order ID or by venue order ID, and only client-ID cancellation can cancel an
order still in PendingSubmit.
Because this state is in memory, a strategy process assumes it starts with a clean venue. The multi-venue binary enforces this literally — it rejects startup if unmanaged open orders exist on Lighter.
Adapter-side state¶
Each adapter keeps its own tracking, and the two current adapters keep different things because their fill sources differ:
| Kuru | Lighter | |
|---|---|---|
| Tracks | Applied-trade history | Per-order TrackedOrder plus filled quantity |
| Dedupe key | Chain transaction / log identity | (client_order_index, fill_id) |
| Structure | Bounded VecDeque, oldest evicted |
Keyed by client order index |
| On miss | Skips rather than risk double-counting | Cannot attribute the fill |
The Kuru dedupe window is intentionally bounded — it only needs to outlive the
window in which the same trade could be seen twice, so it evicts rather than
growing forever (crates/venue-adapters/src/kuru/trades.rs). Deduplication is
silent by design: a trade that is not ours and a trade already applied both
produce nothing.
Execution availability¶
ExecutionAvailability is per-route state held by the health decorator
(crates/venue-adapters/src/health.rs). Its recovery rule is deliberately
asymmetric:
- Any dispatch failure marks the route
Unavailable. - A route becomes
Availableagain only after a successful health check. An ordinary successful request does not clear a prior failure.
This is why a leg can stay unavailable while apparently working — the health probe, not traffic, is what restores it.
Projections and transport¶
InfluxDB projections are append-only and non-authoritative. Post-trade records are queued and written asynchronously; the queue exposes counters for enqueued, evicted, retried, and failed batches. Eviction means projection loss is possible under sustained backpressure, so projections are for analysis and reconstruction, never for deriving live position.
NATS is core NATS, not JetStream. There is no durable subscription and no replay. Market data, commands, status, and events are live-only: a process that is down misses what was published while it was down, which is precisely why desired state lives in MongoDB rather than being inferred from the bus.