- Status: Draft
- Tier: Experimental (formally tiers the previously-untiered cluster-membership surface; see Tier Classification)
- Author: AI assistant review
- Created: 2026-08-03
- Resolved: (pending)
- Language-version at effect: none (no Frozen or Stable surface touched; NUL0 wire protocol v1 unchanged)
- Supersedes: none
- Superseded by: none
Add a pluggable split-brain resolver to the cluster-membership runtime so a
partitioned cluster converges to one surviving side instead of running two
disconnected partitions indefinitely, and fix the related non-self-healing
bug: once two sides of a clean partition each mark the other Failed,
neither side ever redials the other, so the split persists until an external
rejoin. The first shipped strategy is static-quorum (configured expected
cluster size, no live member count). The resolver is opt-in and defaults to
disabled, so existing embedders see zero behavior change.
ClusterState (src/runtime/cluster.rs) has no quorum, leader-election, or
majority logic anywhere — there are zero quorum/leader/elect hits in
src/runtime/ (the only substring matches are inside "selective"). A node's
tick() sends heartbeats and gossip only to Healthy/Joining members, so
when a partition splits a cluster and both sides time each other out to
Failed (the failure detector already works: 2s heartbeat timeout →
Suspicious, +5s → Failed, 60s retention), neither side redials the other.
The split does not self-heal; it requires an external rejoin. The existing
multi-node test coverage explicitly disclaims this: tests.rs states that
split-brain (two healthy sub-clusters that can't see each other) and
asymmetric partition are "NOT covered" and are follow-up.
This matters because Nulang's distributed-actor story is a core differentiator and the current runtime cannot survive an adversarial operational review: a partitioned cluster keeps both halves serving the same durable actors, and a resolved partition silently re-merges two divergent membership views with no record of who should have won.
A resolver observes the local membership view and decides whether the local node stays up. It is a pure function of the view — no I/O, no timers — so it is trivially unit-testable and DST-drivable.
/// A snapshot of what the local node can currently see.
pub struct MembershipView {
pub local: NodeId,
pub members: Vec<NodeInfo>, // all known members with current status
}
pub enum ResolverDecision {
StayUp,
DownSelf,
}
pub trait SplitBrainResolver: Send + Sync {
fn decide(&self, view: &MembershipView) -> ResolverDecision;
}ClusterState::tick() consults the resolver once per tick (after the
failure-detection passes). StayUp changes nothing. DownSelf transitions
the local node to a new Down state and emits a new action.
pub struct StaticQuorumResolver { pub expected_nodes: usize }Rule: the local node counts the members it currently sees as reachable —
itself plus every member whose status is Healthy or Joining
(Suspicious and Failed do not count). If the reachable count is at least
floor(expected_nodes / 2) + 1, the node stays up; otherwise it downs
itself.
- Requires only the operator-configured expected cluster size — no live count, no consensus, no leader.
- 2-node caveat (documented, not fixed): with
expected_nodes = 2, both sides see only themselves during a partition (1 < 2) and both down themselves — the cluster fails closed with no survivor. This is the standard static-quorum property; the strategy is only useful forexpected_nodes >= 3. The RFC keeps this behavior because fail-closed is the safe failure mode for a 2-node cluster (a silent two-sided split is worse than a total outage). expected_nodes == 0is rejected at configuration time (treated as a configuration error, not as "disabled" — see §4).
keep-majority/keep-oldest are deliberately NOT specified here: they
depend on an accurate live member count, which partial-view membership can
undermine. They may be added later as additional implementations of the same
trait, gated on the partial-view work proving count accuracy (Phase 5
deliverable 6 of PLAN.md). The trait contract above is the extension seam.
Two new variants join ClusterAction:
pub enum ClusterAction {
// ... existing variants ...
/// The resolver decided the local node should leave the cluster.
Down { node: NodeId }, // node == local node
/// Minimal periodic liveness probe to a Failed member.
Probe { to: NodeId, addr: SocketAddr },
}Down { node }is emitted exactly once when the resolver first decidesDownSelf. The runtime marks the local nodeDown(alocal_downflag onClusterState, exposed asClusterState::is_down()); a downed node stops emitting heartbeats and gossip (tick()returns no actions), stops processing cluster packets, and keeps running local actors. Peers learn of the down via their own failure detector and mark the nodeFailednormally — no new wire message is needed. Noon_member_*callback fires for the downed node's own peers (they discover it through existing timeouts), and the downed node's local callbacks (on_member_left,on_member_failed) fire for its own entry so operators can react.Probe { to, addr }is emitted for eachFailedmember at a configurableprobe_interval(default 5s; heartbeat interval is 500ms — probes are a minimal liveness check, not full heartbeating). The runtime sends an ordinaryHeartbeatpacket to the address; the wire packet type already exists, so NUL0 v1 is untouched — no new packet type, no version bump. If the probe reaches a live node, the node's existinghandle_heartbeatpromotion logic (which already promotesSuspicious/Failedback toHealthyon any heartbeat) re-joins it: this is the self-healing fix — when the network recovers from a partition, both sides' probes succeed and the cluster re-merges with no external rejoin. A node that is truly dead never answers and staysFaileduntil the 60s retention purge, unchanged from today.
pub enum SplitBrainConfig {
Disabled, // default
StaticQuorum { expected_nodes: usize },
}
pub struct ClusterConfig {
pub split_brain: SplitBrainConfig,
pub probe_interval: Duration, // default 5s
}New additive API on Runtime:
impl Runtime {
/// Must be called before `enable_distribution`; defaults to Disabled.
pub fn set_cluster_config(&mut self, config: ClusterConfig);
}The signature of enable_distribution(addr, tls_config) is unchanged
(deliverable 4 of PLAN.md Phase 5 reshapes it under its own RFC). The
default configuration is SplitBrainConfig::Disabled with a 5s probe
interval — existing embedders get no behavior change except the self-healing
probe to Failed members (a fix, not a break: it cannot resurrect a node
that stays down, and it cannot prevent a failed node from being purged).
Landing this RFC's implementation requires, in the same change set:
- Unit tests for
StaticQuorumResolver::decideacross every reachable-count boundary (floor(N/2)+1 - 1, the threshold, the threshold + 1) and theexpected_nodes = 0configuration error. - A probe test: two
Runtimes over loopback TCP, one killed transport (hard failure), the survivor transitions the peer toFailedand then promotes it back after the transport is restored — no external rejoin. - The DST/chaos scenarios
PLAN.mdPhase 5 deliverable 2 names — mutually invisible healthy sub-clusters and asymmetric (one-way) partition on 3-node and 5-node topologies — asserting the cluster converges to exactly one surviving side per the configured strategy and never a stuck two-sided split. These land as the verification vehicle for this RFC (deliverable 2 is sequenced immediately after deliverable 1 inPLAN.md).
Cluster membership is currently untiered: GOVERNANCE.md's Frozen list
covers .nbc, NUL0 v1, value layout, Core syntax, and the IO/Spawn/
Send/Receive effects; its Stable list covers the capability lattice and
the actor surface; neither names ClusterState, enable_distribution,
join_cluster, or membership behavior generally. This RFC formally assigns
that surface (including the new resolver) to the Experimental tier:
- It is a Rust-embedder-only API with no
.nulasyntax and no CLI surface. - NUL0 wire protocol v1 is untouched (the probe reuses the existing
Heartbeatpacket type;Downis runtime-local state). - No language-version bump is triggered.
Graduation to Stable is a follow-up RFC, gated on the DST/chaos suite above proving the resolver semantics under partition (the evidence-first bar GOVERNANCE.md §3 sets). The resolver's operational blast radius — a mechanism that can autonomously shut down a node — is exactly why this RFC exists even though the surface is formally untiered.
- No existing program or embedder breaks: the resolver defaults to
Disabled, andenable_distribution's signature is unchanged. - The only behavior change for existing embedders is the self-healing probe:
Failedmembers are now contacted with a minimal heartbeat everyprobe_intervalinstead of never. A live-but-partitioned node rejoins automatically when the network recovers (previously: manual rejoin); a dead node behaves exactly as before (staysFailed, purged after 60s). - No deprecation cycle is required; nothing is removed.
- Raft/consensus-backed membership. Rejected for this RFC by reference:
PERFORMANCE_ANALYSIS.mdrow 3.4 defers native Raft ("CRDTs cover 80% of distributed state needs"), and a split-brain resolver provides partition safety without consensus. This RFC does not relitigate that deferral. keep-majority/keep-oldestnow. Rejected: both need an accurate live member count, and partial-view membership (Phase 5 deliverable 6) can make live counts unreliable. They are deferred behind the same trait until deliverable 6 proves count accuracy — shipping a resolver that silently miscounts is worse than shippingstatic-quorumalone.- External rejoin tooling (operator script). Rejected: the phase goal (PLAN.md Phase 5) is a cluster that provably converges to one surviving side, not one whose operator must notice and intervene.
- Link-quality-based partition detection (TCP keepalive tuning, ICMP). Rejected: the membership view is the system of record, and the heartbeat state machine already encodes reachability; a second, independent signal would only add disagreement between the resolver and the failure detector.
- Down semantics for local actors. This RFC keeps local actors running on a downed node (only cluster participation stops). An alternative is a full local halt. The former is less surprising for operators debugging a partition; confirm during discussion.
- Probe interval default. 5s is proposed (10× the heartbeat interval). A shorter probe speeds rejoin after recovery; a longer one reduces noise against a dead node. Not load-bearing — an operator-configurable knob in any case.
- Operator alerting on resolver decisions. Proposed: reuse the existing
on_member_failed-style callback pattern so aDownSelfdecision is observable from Rust. Whether a log line suffices for v1 is open. expected_nodesdrift. If the operator's configured expected size is wrong (too high), a healthy cluster downs itself; too low, and a partition may keep both sides up. Should the resolver warn when the reachable count exceedsexpected_nodes(a likely misconfiguration)? Non-blocking: the failure mode is already conservative (config error is loud at startup, and the strategy is opt-in).
This section extends the same Experimental-tier surface with partial-view membership: the heartbeat data plane becomes O(active view) instead of O(every member), while the membership table and gossip stay full-view. No wire change — heartbeats, gossip, and probes are unchanged packet types; the views are purely local state.
- Active view (default 4): the members we heartbeat directly. A member is admitted by heartbeating us (with room in the view), so admission is evidence of reciprocity; the failure detector watches exactly this set.
- Passive view (default 20): known members we do not heartbeat; the repair pool. Their liveness comes from gossip.
- Probation: when an active member fails, a random Healthy passive member is promoted to probation — we heartbeat it, awaiting its first reply — and confirmed into the active view on the reply. Probationary members are not watched, so a member that never reciprocates is demoted back to passive (churn), never falsely failed. Repair retries every 5 s while the view is underfull.
- Bounded reply rule: each node replies per heartbeat round to up
to
REPLY_SLOTS(4) passive members that recently heartbeated it (rotated round-robin). Without replies, a member whose active view filled up would stop heartbeating us and the detector would false-fail it; with 4 slots × 500 ms, every pinger is answered within the 2 s detection window for clusters up to ~80 nodes.
Consequences for the failure detector:
- Only active-view members are watched, so no false failure can come from a member we do not heartbeat (including members whose views filled up). A dead member is always watched by its own active partners (heartbeat links are symmetric by construction: admission requires the partner's heartbeats, which continue until death).
- The detector now bumps the entry incarnation on a
Failedtransition, so the Failed status propagates via gossip to non-watchers (previously the Failed transition never bumped, which was invisible under full-mesh detection but fatal under partial view).
merge_membership refreshes a member's last_heartbeat on an
equal-incarnation re-broadcast when the member is passive (not
watched) and live (Healthy/Joining). Watched members are refreshed
only by direct heartbeats, so the failure detector still works. Failed
entries are never refreshed — the dead-peer protection from the
original RFC 0011 work is preserved (regression-tested).
The membership view handed to the resolver marks stale-status passive
members as Suspicious (view only, table untouched): a member whose
gossip-derived status is a frozen Healthy snapshot is not reachable
evidence. This keeps static-quorum correct under partial view — an
isolated node's stale gossip does not keep it above quorum. A
reachable_count mirroring the resolver view is exposed for the DST
harness so invariant tests assert the resolver's exact semantics.
The SimCluster DST harness (Phase 5 deliverable 2) gained four
10/30-node scenarios: bounded fanout (heartbeats ≤ active + probation +
replies, never full-mesh at 30 nodes), mesh convergence, node-death
detection with no false failures (survivors keep mutual-Healthy
statuses; the dead node is Failed everywhere), and heal/rejoin with
active-view repair. Six unit tests cover the view mechanics
(admission, detector scoping, probation confirm/demote, bounded
heartbeats, join bootstrap).
(To be filled on accept/reject.)