Concurrency: what scales, what doesn’t¶
SecantusDB is a single-process embeddable MongoDB server. This page is about what that means for concurrent writers — many client connections issuing inserts/updates/deletes at the same time.
The short version: the Rust server’s fully-durable write path now
scales monotonically — ~1.9× at two writers, ~3.0× at four, ~3.3×
at eight — with no cliff (the earlier peak-then-collapse shape was
oplog-prune churn on the write path plus over-sharded oplog btrees;
both fixed). The remaining gap to mongod’s ~5.2× is a WiredTiger
ceiling — cache eviction and checkpoint pressure inside one embedded WT
connection — not a SecantusDB lock. The opt-in async + non-logged
oplog stack (mitigations 5–6 below) starts from a ~1.6× higher
single-writer base and reaches the highest absolute throughput at
every writer count, at the cost of a checkpoint-durable-only oplog
tail. The Python server barely scales at all (the GIL). If your
workload depends on write throughput that keeps climbing past eight
writers, run a real mongod instead.
Note
Not comparable to figures published before 2026-08-22. The sweep used to share one store across every writer count, so each row measured a different database — with 8 KiB documents the eight-writer row wrote into a tree holding everything the earlier rows left behind. That biased scaling downwards, and on a smaller disk it eventually exhausted space outright and took WiredTiger’s ENOSPC panic mid-sweep.
Each row now gets a fresh store. Every engine improved as a result, most at
high writer counts where the accumulation was worst — including mongod, which
is unchanged code and gained 14% at eight writers.
Measured on a dedicated cloud instance (8 vCPU) against mongod 8.0.29, as
part of cutting secantusdb-v0.5.3-beta.163 — so these are the figures for the
build you can actually download. Absolute throughput is lower than the earlier
workstation-measured numbers because the cores are slower; the scaling
ratios, which is what this page reports, are if anything better. mongod is
the control: it scales 5.15× here against 4.65× on that workstation, which
confirms eight writers on eight vCPUs is not core-starved.
What scales fine¶
Concurrent reads. Multiple
find/count/aggregatecalls against the same or different collections run in parallel under WiredTiger’s MVCC. Reads don’t block writes and don’t block other reads.Per-connection isolation. Each TCP connection gets its own server thread and its own WiredTiger session. Sessions don’t contend on each other for reads.
Single-writer throughput. A single connection driving batched
insert_manyof 8 KiB documents sustains ~12,687 docs/s on the Rust server and ~4,563 docs/s on the Python server (fully durable, WAL-logged, on an 8-vCPU cloud instance; 2026-08-26 baseline).
What doesn’t scale¶
Aggregate write throughput without bound. Both SecantusDB servers hit
a WiredTiger ceiling as writers pile up — the Rust server’s scaling
flattens between four and eight writers (~3.0× → ~3.3×, still
monotonic); a shared-table / large-row workload (the pure-C wt_poc
case below) tops out much earlier, around N≈2, and actively regresses
past its peak.
We measured this carefully because the question kept coming up. The
benchmark and the data are at bench/wt_poc/; you can re-run it on
your hardware to confirm.
The headline number¶
bench/wt_poc/run.py runs the same workload (50,000 row inserts,
each row ~1 KiB, partitioned across N writers writing to their own
table) through three paths:
N writers |
Pure-C + pthread (no Python) |
Python + WT SWIG bindings |
|---|---|---|
1 |
276,449 rows/s (1.00×) |
116,578 rows/s (1.00×) |
2 |
340,106 rows/s (1.23×) |
87,010 rows/s (0.75×) |
4 |
352,731 rows/s (1.28×) |
67,660 rows/s (0.58×) |
8 |
285,146 rows/s (1.03×) |
58,751 rows/s (0.50×) |
The pure-C column is the theoretical best case: pthreads, no GIL, no
Python on the hot path, calling libwiredtiger directly. Even
that caps at ~1.3× of single-thread aggregate throughput at N=2 and
flatlines (or regresses) past that.
The bottleneck is at the WT C library level — B-tree page locks, log
write serialisation, cache eviction, internal scheduler. It’s the
same library mongod uses, but mongod gets multi-writer scaling by
running a careful C++ scheduler above WT that takes advantage of
lower-level WT primitives (per-cursor concurrency hints, parallel
cursor batches, careful checkpoint coordination). SecantusDB doesn’t
have that scheduler — and writing one isn’t a SecantusDB project; it
would essentially be re-implementing mongod.
End-to-end: both servers vs mongod¶
Measured 2026-07-31 with the three-server harness
(uv run python -m bench.concurrency --server all --writers 1,2,4,8 --runs 3), medians of three interleaved quiesced runs: N writer
processes, each streaming insert_many batches through pymongo
against its own collection, all servers on on-disk WiredTiger. The
chart and table below are regenerated from those measurements on every
release (invoke concurrency-refresh); the committed numbers live in
bench/results/concurrency.json.
N writers |
Python server (docs/s) |
Rust server (docs/s) |
Rust — async stack (docs/s) |
mongod (docs/s) |
|---|---|---|---|---|
1 |
4,600 |
12,700 |
16,600 |
26,400 |
2 |
5,000 |
24,000 |
30,900 |
56,900 |
4 |
3,600 |
37,500 |
44,900 |
103,500 |
8 |
3,000 |
42,400 |
51,000 |
136,200 |
(The async-stack column is the opt-in async-oplog + non-logged-oplog
configuration, mitigations 5–6 below — first-class options since Phase C:
RustServer(oplog_async=True, oplog_nonlogged=True), or secantusd-rs --oplog-async --oplog-nonlogged, or the SECANTUS_OPLOG_ASYNC=1 +
SECANTUS_OPLOG_NONLOGGED=1 env vars. Scaling in the chart is relative
to each series’ own single-writer rate.)
Three different shapes:
mongod scales — 5.2× its own single-writer aggregate at N=8. That’s the C++ scheduler above WT doing its job.
The Rust server now scales monotonically — 1.9× at two writers, 3.0× at four, 3.5× at eight, with no cliff. The earlier peak-then-collapse shape (2.6× at four easing back to 1.6× at eight) was diagnosed by profiling and eliminated in two steps: the opportunistic oplog prune was consuming ~36% of the sustained write path (every sweep re-read the full value of every doomed row; it now scans keys only), and the oplog’s 16-way shard split had outlived the append hotspot it was built for (writes now route across two append-tuned btrees). Together with the per-collection write-lock split and the RecordId keying work (write amplification cut from four WT writes per document to three), the fully-durable default now holds ~half of mongod’s scaling ratio. The remaining flattening between four and eight writers is WiredTiger itself — cache eviction and checkpoint pressure inside a single embedded WT connection — not a SecantusDB lock.
The async stack delivers the highest absolute throughput (dashed line) — the opt-in async-oplog + non-logged-oplog configuration (mitigations 5–6 below) moves the oplog write off the writers’ critical path and out of the WAL: 2.4× scaling from a ~1.4× higher single-writer base (50.4k vs 35.4k docs/s), reaching ~119k docs/s at eight writers — ~1.2× the default’s 96k. The trade is that a hard crash loses the oplog tail since the last checkpoint (data stays fully durable; change streams remain exactly-once).
The Python server degrades under contention — the GIL plus the WT-binding ceiling measured above hold it to ~0.6× of its single-writer rate. It degrades gracefully, though: the write conflicts WiredTiger reports under saturation are retried with backoff, without a deadline, exactly like mongod’s
writeConflictRetry— a client never sees an error, generic or otherwise (both the swallowed-InternalErrorclassification bug and the deadline that could surfaceWriteConflicton plain writes were found by this harness and fixed).
Note too that single-writer throughput itself keeps climbing — the Rust server from ~3.5k docs/s at the 2026-07-17 baseline to ~25.7k post-RecordId to ~35.4k now (the prune fix, the append-tuned defaults, and a freshly-trained PGO profile), the Python server from ~2.9k to ~12.1k. mongod’s single-writer rate is unchanged (~106k), which is what validates the harness is measuring the same thing.
mongod pays for an oplog too¶
One asymmetry in the chart above: the mongod it benchmarks is
standalone — no replica set, so no oplog at all — while SecantusDB
always maintains one (change streams need it). Charging mongod for the
same feature changes the picture. bench/mongod_replset_ab.py runs the
identical workload against a single-node replica set (medians of 3
interleaved reps):
mongod configuration |
1 writer |
8 writers |
|---|---|---|
standalone (no oplog) |
113.2k docs/s |
503k docs/s |
replica set, explicit |
84.0k docs/s |
305k docs/s |
replica set, default write concern |
11.8k docs/s |
68.6k docs/s |
The oplog double-write costs mongod −26% (1 writer) / −39% (8 writers) —
the same structural tax SecantusDB pays, at about half the rate (its
timestamp-slot oplog admits concurrent appends without the shared-append
serialisation we shard around). The bigger surprise is the default-config
row: since MongoDB 5.0 the implicit write concern is w:majority, and on
a one-node set a majority ack waits for a journal fsync — a ÷7 cliff that
dwarfs the oplog itself. So a single-node replica-set mongod as people
actually run it for change streams writes at 11.8k / 68.6k docs/s on
this hardware — slower than the Rust server’s async + non-logged stack
(49.5k / ~125k), though with stronger per-write durability (its
acknowledged writes survive a hard crash; ours trade that per-ack fsync
away everywhere). At equal write-concern semantics mongod’s raw ingest
path is still ~3× faster per writer; that residual is its C++ ingest
machinery, not the oplog.
Why disabling logging doesn’t fix it¶
A natural follow-up: maybe the journal is the serialiser. We tested
that — same C benchmark, log=(enabled=false):
N writers |
Pure-C + pthread, no log |
|---|---|
1 |
1,007,557 rows/s (1.00×) |
2 |
1,156,150 rows/s (1.15×) |
4 |
700,035 rows/s (0.69×) |
8 |
347,176 rows/s (0.34×) |
Single-thread is much faster (~4×) but multi-thread is worse — collapses at N=4 and N=8. Disabling logging is a single-writer optimisation that loses crash durability AND fails to deliver concurrency.
What this means for your workload¶
One connection doing batched writes is a simple, fast configuration for tests / dev / single-process applications.
pymongo’sinsert_manywith batch=100 hits ~25,000 docs/s on the Rust server and ~11,000 on the Python server on commodity hardware with full durability.Many connections doing concurrent writes scale monotonically on the Rust server — ~2.7× the single-writer rate at eight writers (~96k docs/s aggregate) with no cliff, fully durable; the Python server barely scales (the GIL). The opt-in async + non-logged oplog stack (mitigations 5–6) buys another ~1.2× (~119k docs/s at eight writers) if a checkpoint-durable oplog tail is acceptable. Run a real
mongodif you need write throughput that keeps climbing steeply past a handful of writers.Many connections doing concurrent reads scales fine. Reads use MVCC snapshots and don’t contend.
Mixed read/write at moderate N works as expected: writes serialise, reads run in parallel against an MVCC snapshot.
Mitigations within SecantusDB¶
If you genuinely need higher single-process write throughput from SecantusDB, the levers are:
Batch larger.
insert_manywith batch=100 is ~2× the throughput ofinsert_one. Going larger has diminishing returns.Reduce server-side work. Drop indexes you don’t need. Each index adds per-doc encode + WT cursor write.
Disable the oplog if you don’t need change streams. Pass
replica_set_name=NonetoSecantusDBServer(or run without--authand without a replica-set advertisement). Halves per-write WT cursor traffic.writeConcern: w:0for fire-and-forget writes — pymongo doesn’t wait for the server’s ack. Throughput climbs on the client side; server-side cost is unchanged.Async oplog (Rust server, opt-in).
RustServer(oplog_async=True),secantusd-rs --oplog-async(or[storage] oplog_async = trueinsecantusd.toml), orSECANTUS_OPLOG_ASYNC=1— the option wins over the env var for that store. Moves oplog writes off the writer’s critical path onto a background drainer — ~1.6× multi-writer write throughput while keeping change streams (validated exactly-once under concurrency; a fresh change stream opens only after the drainer has caught the acknowledged tail, so it never surfaces pre-open events; multi-document transactions buffer their entries and hand them to the drainer only at commit, so a rolled-back transaction never surfaces a change event). Trade: the oplog is no longer atomic with the data, so a hard crash loses entries the drainer hadn’t yet written (the data itself stays fully durable; a clean shutdown flushes the drainer). Bounded bySECANTUS_OPLOG_ASYNC_CAP_BYTES(default 128 MB). The drainer coalesces queued batches into one WiredTiger transaction (SECANTUS_OPLOG_ASYNC_COALESCE=0disables). Default off. (The async figures in the table above were measured before the async path gained the same opportunistic oplog prune the sync path has always had; the prune’s retention enforcement costs ~5% at eight writers, so read the async column as slightly optimistic until the next release re-baseline.)Non-logged oplog tables (Rust server, opt-in).
RustServer(oplog_nonlogged=True),secantusd-rs --oplog-nonlogged(or the TOML key), orSECANTUS_OPLOG_NONLOGGED=1— applied at first open of a fresh store (the table config is create-time-sticky) — to create the oplog + pre-image tables with WAL logging disabled: the oplog becomes checkpoint-durable only — a hard crash loses the oplog tail since the last checkpoint (change-stream resume / PITR granularity), while the data tables stay fully logged and durable; a clean shutdown checkpoints a complete oplog. Alone it buys little (~4% in sync mode), but stacked with the async oplog it removes the oplog’s WAL volume from the writers’ path entirely: ~2.2× the default 8-writer throughput (measured 56k → 125k docs/s of 8 KiB documents, against a ~191k no-oplog ceiling), and ~1.9× a single writer. Change streams remain exactly-once under the stack.WiredTiger config tuning (Rust server / daemon).
SECANTUS_WT_CONFIG_EXTRAappends raw WT connection config (last-key-wins). A largercache_sizeis the strongest single knob under sustained writes (+26% at eight writers in the Finding-13 sweep) — the daemon, the PythonRustServerhandle, and the embedded RustStoragelibrary all default to a 4G cache cap (WiredTiger fills it lazily, so idle test servers stay small;--cache-size/cache_size=overrides). Two cautions from the same sweep, post prune-fix: log pre-allocation (prealloc=true) now hurts at eight writers (−8%; the earlier +8% predates the prune fix), and never turn oplog block compression off — throughput craters to ~19% of the ceiling, because bigger uncompressed pages mean more eviction IO and IO volume, not CPU, is the constraint.
The defaults already carry the measured winners: oplog writes route
across two shard tables (sixteen existed to spread an append hotspot
that the RecordId + prune work eliminated; the read side still scans
all sixteen, so any store stays readable), and the oplog/preimage
btrees are created append-tuned (split_pct=100,leaf_page_max=128KB).
With those defaults a fully-durable eight-writer load sustains ~half
of the no-oplog ceiling (~96k docs/s of 8 KiB documents on the
reference box, vs ~43%/75k before the prune fix and these defaults).
The honest ceiling: for a fully-durable, WAL-logged oplog the limit is
WiredTiger’s aggregate write rate on a single embedded process. The
async + non-logged stack (levers 5+6) trades crash-durability of the
oplog tail for the last stretch toward the no-oplog ceiling; past
that, sustained multi-writer scaling means running a real mongod
(or dropping the oplog entirely).
What we tried, what didn’t work¶
The path we took to nail this down (preserved here so future contributors don’t re-walk it):
Lock-decomposition (replace global
Storage._lockwith per-collection locks + tiny_oplog_seq_lock). Did clean up several internal correctness issues — see thetasks/wt-concurrency-plan.mdwriteup — but didn’t move multi-writer scaling. Bottleneck wasn’t the Python lock layer.Profiling the insert hot path (
bench/profile_insert.py). Showed 50%+ of wall time was in WiredTiger’s SWIG-generated Python bindings (wiredtiger/packing.py), not in our code. Suggested a Cython rebind would help.The pure-C pthread benchmark (
bench/wt_poc/). Killed the Cython rebind hypothesis: even with no Python anywhere, WT itself doesn’t scale past N≈2. The bindings are a constant overhead; removing them wouldn’t change the multi-writer story.
The artefacts of all three exploration tracks are kept in the repo as reproducible evidence. Re-run them when somebody asks “but what if we just X?” and confirm the numbers haven’t moved.
Tracking¶
tests/test_concurrency.py (which drives the Python server) is
marked xfail (expected-fail) — it encodes the goal “2 concurrent
writers >= 0.7× of one”, which the Python server’s GIL-bound path
cannot deliver (it measures ~1.17× at two writers on the current
baseline, but collapses to ~0.82× by eight). The Rust server clears
that bar comfortably (~1.9× at two writers, rising to ~3.3× at eight —
see the chart above); the xfail tracks the Python server specifically,
and stays a useful regression detector for it.