Performance
Rift delivers 20–6,000x the throughput of Mountebank on identical imposter configs, with sub-millisecond tail latency that stays flat as stub count grows.
Benchmark Summary
The suite was run on two deliberately different hosts. Publishing both is the point: the multiplier depends heavily on the machine, and a single number would overstate the result.
- Apple M4 laptop (10 cores, macOS) — the conservative read
- AMD EPYC 9V74 (16 vCPU, 62 GiB, Linux) — the server read
Rift built from master (924cf73) · Mountebank 2.9.1 · oha, 50 keep-alive connections, 20s/scenario after warmup, native processes (no Docker), each engine run alone. Every figure is the median of 3 repetitions. Measured 2026-07-20. Full method and reproduction: tests/benchmark.
| Scenario | MB (M4) | Rift (M4) | M4 speedup | MB (EPYC) | Rift (EPYC) | EPYC speedup |
|---|---|---|---|---|---|---|
| Regex (100th pattern) | 112 | 207,024 | 1,857x | 52 | 317,851 | 6,160x |
| API stub — no match (404) | 1,351 | 209,763 | 155x | 549 | 332,574 | 606x |
| API stub — last match | 1,344 | 209,523 | 156x | 542 | 322,530 | 595x |
| API stub — middle match | 3,437 | 210,151 | 61x | 1,081 | 324,067 | 300x |
| API stub — first match | 8,546 | 211,378 | 25x | 5,728 | 323,408 | 57x |
| Query-param routing | 2,751 | 164,133 | 60x | 1,112 | 211,748 | 190x |
| Header routing | 3,016 | 158,596 | 53x | 1,202 | 201,940 | 168x |
| Complex AND/OR predicates | 4,703 | 191,987 | 41x | 1,814 | 259,548 | 143x |
| JSONPath predicates | 4,312 | 199,404 | 46x | 1,921 | 304,796 | 159x |
| XPath predicates | 5,542 | 187,869 | 34x | 1,966 | 247,897 | 126x |
| JSON body matching | 7,611 | 199,670 | 26x | 2,730 | 294,294 | 108x |
| Template responses | 9,022 | 194,236 | 22x | 3,152 | 283,815 | 90x |
| Simple static stub | 8,898 | 214,818 | 24x | 5,982 | 324,952 | 54x |
How to read the two columns
Going from the laptop to the 16-vCPU server, Rift gets faster (215k → 325k) and Mountebank gets slower (8,898 → 5,982). Mountebank is single-threaded, so it can only use one core, and this server’s individual cores are slower than the M4’s — it gains nothing from the other 15. Rift uses them all.
That means the EPYC multipliers are inflated at both ends, and the honest headline is the M4 column. It is still 22x–1,857x.
Tail latency is the more stable comparison: Rift’s p99 is 0.43–0.97 ms on both hosts, while Mountebank’s ranges from 2.9 ms to 1.7 seconds depending on scenario.
Measurement caveat: a laptop thermally throttles under a 30-minute run — both engines lost ~7% between the first and last repetition, and per-scenario spread reached 12% on the M4 versus 5% on EPYC. Treat M4 figures as ±10%.
Why throughput stays flat
Rift holds ~210k RPS (M4) / ~325k RPS (EPYC) whether the matching stub is first, middle, or last — and on a no-match 404 — while Mountebank degrades linearly with stub count:
| API stub position | Mountebank (RPS) | Rift (RPS) | Speedup |
|---|---|---|---|
| First | 8,546 | 211,378 | 25x |
| Middle | 3,437 | 210,151 | 61x |
| Last | 1,344 | 209,523 | 156x |
| No match (404) | 1,351 | 209,763 | 155x |
Regex used to be the exception on Rift’s side too — it can’t be hash-dispatched, and at the 100th pattern Rift managed ~54k RPS against Mountebank’s 106. The candidate-bitset matching framework removed that cliff: regex now runs at 207k RPS, in line with every other predicate type. Mountebank’s per-stub JS RegExp scan still collapses to 112 RPS at the 100th pattern, so the gap widened from 515x to 1,857x — not because Mountebank got slower, but because Rift stopped having a slow path.
On the admin control plane, creating 1,000 fully-overlapping stubs (the O(n²) case issue #423 fixed) takes Rift 6.6ms vs Mountebank’s 114.7ms, and grows memory +9MB vs +51MB — while Rift additionally computes stub-overlap warnings Mountebank does not.
Why Is Rift Faster?
Architecture Comparison
| Aspect | Mountebank | Rift |
|---|---|---|
| Language | Node.js (JavaScript) | Rust |
| Concurrency | Single-threaded event loop | Multi-threaded (Tokio) |
| Memory Model | Garbage collected | Zero-copy, no GC |
| Regex Engine | JavaScript RegExp | Rust regex crate |
| JSON Parsing | JavaScript JSON | serde_json (SIMD) |
| Stub Matching | Linear scan | Optimized matching |
Key Optimizations
-
Native Code: Rust compiles to native machine code, avoiding interpreter overhead.
-
Async I/O: Tokio runtime provides efficient async networking with work-stealing scheduler.
-
Zero-Copy Parsing: serde_json parses JSON without unnecessary allocations.
-
Efficient Regex: Rust’s regex crate uses finite automata for O(n) matching.
-
Connection Pooling: Reuses connections to upstream services.
-
Thread Pool: Dedicated workers for script execution.
Performance Characteristics
Latency (p99)
| Scenario | Mountebank | Rift |
|---|---|---|
| Exact stub match (last of 500) | 40ms | 0.6ms |
| Complex AND/OR predicate | 17ms | 0.8ms |
| JSONPath match | 17ms | 1.0ms |
| Regex (100th pattern) | 641ms | 1.8ms |
Throughput Scaling
Rift maintains consistent throughput regardless of:
- Stub count (500+ stubs with minimal degradation)
- Stub position (first vs last stub match)
- Predicate complexity
Mountebank shows linear degradation as stub count increases.
Running Benchmarks
The suite runs both engines as native processes, one at a time on disjoint ports, and posts byte-identical imposter JSON to each. See tests/benchmark/README.md for full details.
Prerequisites
cargo build --release -p rift-http-proxy # build Rift from source
cargo install oha # load generator
npm install --prefix ~/bench-mb mountebank@2.9.1 # reference engine
Run the suite
cd tests/benchmark
# Serving throughput + tail latency
python3 scripts/bench_direct.py --run-all \
--duration 20s --warmup 3s --connections 50 \
--rift-bin ../../target/release/rift-http-proxy \
--mb-bin ~/bench-mb/node_modules/mountebank/bin/mb
cat results/DIRECT_BENCHMARK_REPORT.md
# Admin create/read (imposter creation + overlap analysis)
python3 scripts/bench_admin.py --run-all \
--rift-bin ../../target/release/rift-http-proxy \
--mb-bin ~/bench-mb/node_modules/mountebank/bin/mb
cat results/ADMIN_BENCHMARK_REPORT.md
ohareads the macOS keychain to initialise TLS even for plain-HTTP targets — run outside a restricted sandbox.
Optimization Tips
For Maximum Throughput
- Use specific predicates -
equalsis faster thanmatches - Order stubs by frequency - Most-matched stubs first
- Avoid unnecessary behaviors - Each behavior adds overhead
- Use native formats - JSON body predicates are faster than string matching
For Script Fault Injection
Script fault decisions are memoized in a decision cache, keyed on the request. By default the key includes every request header. That is always correct, but if your traffic carries a per-request-unique header — x-request-id, traceparent, x-amzn-trace-id, date — then every key is unique, nothing ever hits, and the cache becomes pure overhead: it pays hashing, allocation and lock traffic on the hot path and returns nothing.
Rift cannot narrow the key for you: the cached value is your script’s decision, and your script is handed every header, so it may branch on any of them. Dropping a header from the key that your script actually reads would serve one request’s decision to a different request. So the allowlist is opt-in — it is your assertion about what your scripts read:
# Proxy config (the same file that carries `script_rules`) — NOT the imposter `_rift` block.
listen:
port: 8080
script_rules:
- # ...
decision_cache:
enabled: true
max_size: 10000
ttl_seconds: 300
key_headers: ["X-Tenant", "X-Feature-Flag"]
Only the listed headers enter the cache key; names are matched case-insensitively, and an empty list ([]) declares that no header affects your decisions. Your scripts still receive all headers either way — this only changes what makes two requests “the same” for caching.
If the cache degenerates to a ~0% hit rate, Rift logs a warning once per process telling you so, rather than silently burning CPU.
What makes two requests “the same”
The key is the method, the path, the query string, the key_headers above, the rule id, and the body.
The query is keyed on its raw spelling, so ?a=1&b=2 and ?b=2&a=1 are two entries even though they mean the same thing. That is deliberate: it can only cost you a cache miss, whereas keying on the parsed form could hand one request another’s decision. Clients serialize query strings deterministically, so in practice it costs nothing.
How the body counts depends on whether it is JSON:
- JSON — keyed structurally, so whitespace and key order do not split the key. Two requests whose bodies parse to the same value share one entry. The corollary: a script that branches on the raw formatting of a valid-JSON body is outside the cache-key contract, the same way one that reads a header you left out of
key_headersis. - Anything else — binary, plain text, malformed JSON, or an empty body — is keyed on its raw bytes, which is what your script reads via
ctx.request.raw_body. Two different uploads are two different keys.
The two are kept in separate hash domains, so a JSON null body, an empty body, and a binary body are always three distinct keys.
The cache is only consulted on the fault-injection proxy path with
script_rulesconfigured and flow state not configured — stateful scripts are never cached.
For Lowest Latency
- Minimize stub count - Fewer stubs = faster matching
- Use simple responses - Static
isresponses are fastest - Avoid injection - JavaScript execution adds latency
- Enable connection pooling - Reuse upstream connections
Resource Allocation
# Recommended for high throughput
resources:
requests:
cpu: 1000m
memory: 256Mi
limits:
cpu: 2000m
memory: 512Mi
Comparison with Alternatives
| Tool | Language | Typical RPS | Best For |
|---|---|---|---|
| Rift | Rust | 200,000+ | High-performance mocking |
| Mountebank | Node.js | 500-2,000 | Feature-rich service virtualization |
| WireMock | Java | 1,000-5,000 | Java ecosystem integration |
| MockServer | Java | 1,000-3,000 | Contract testing |
Rift provides 20-6,000x better performance while maintaining Mountebank compatibility. (Rift’s figure is native/unconstrained; it scales with hardware.)
Runtime Socket Tuning
Rift tunes accepted sockets for low latency out of the box and exposes a couple of knobs via environment variables:
| Variable | Default | Effect |
|---|---|---|
RIFT_TCP_NODELAY | on | TCP_NODELAY is set on every accepted socket (disables Nagle’s algorithm) for lower request latency. Set false/0/off to disable. |
RIFT_TCP_BACKLOG | 1024 | Listen backlog (queue depth) for the accept loop. A larger backlog absorbs bigger connection bursts. Non-positive or unparsable values fall back to the default. |
These apply to both the imposter and proxy accept loops.
Memory Allocator (mimalloc)
The rift-http-proxy binary uses the mimalloc global allocator by default — it improves throughput under the allocation-heavy request path. It is a Cargo feature named mimalloc, enabled in the binary’s default feature set:
# Default build — mimalloc is on
cargo build --release
# Drop it (e.g. for a cross-compile or FFI build) by opting out of default features
cargo build --release --no-default-features --features redis-backend,javascript
# Or swap in jemalloc (bake-off candidate, issue #717)
cargo build --release --no-default-features --features redis-backend,javascript,jemalloc
An opt-in jemalloc feature builds the binary with tikv-jemallocator instead, for A/B allocator comparison; if both allocator features are enabled (e.g. --all-features), mimalloc takes precedence. The startup log reports which allocator is active (Global allocator: …), and the benchmark harness automates the three-way comparison — see the allocator bake-off section in tests/benchmark/README.md.
Only the rift-http-proxy binary is affected; rift-mock-core and the FFI crate use the system allocator.
Runtime Topology (per-core, experimental)
By default the rift-http-proxy binary runs one multi-threaded, work-stealing Tokio runtime that serves everything — imposter accept loops, per-connection work, the admin API, and metrics. That is the right default and is unchanged. For Linux hosts under high connection counts, an opt-in alternative topology (RFC-712) trades a little complexity for materially lower tail latency:
# Default — one work-stealing runtime (unchanged behaviour)
rift-http-proxy --runtime work-stealing
# Per-core: N single-threaded runtimes, N = physical cores
rift-http-proxy --runtime per-core
# …or pin the worker count explicitly
rift-http-proxy --runtime per-core=8
# Env-var equivalent (the CLI flag wins if both are set)
RIFT_RUNTIME=per-core rift-http-proxy
In per-core mode each imposter port binds one SO_REUSEPORT listener per worker runtime, and each accept loop runs on its own single-threaded runtime. The kernel spreads incoming connections across the listeners by 4-tuple hash, so a connection lives and dies on one core — no cross-core wake-ups and no work-stealing overhead. The control plane (admin API, metrics, imposter create/delete) stays on a small shared runtime; only the request-serving accept loops fan out.
At startup the binary reports the topology it actually resolved to, next to the allocator line:
INFO rift: Runtime topology: per-core x8
What it actually buys you
Measured on Linux x86-64 (AMD EPYC, engine pinned to 2/4/8 vCPU with the load generator on disjoint physical cores, 3 repetitions, 14 scenarios — issue #746):
| per-core vs work-stealing | |
|---|---|
| p99 latency | 18–35% lower at every core count tested, at both 256 and 512 connections |
| p999 latency | lower in every scenario measured (84/84 points) |
| Oversubscription | at 2 vCPU / 512 connections work-stealing hit a ~20 ms p99 cliff; per-core stayed at ~5.6 ms |
| Throughput | +1–4% — at or below run-to-run noise; treat it as unchanged |
| Scaling with cores | no measured difference: both topologies scaled ~4.2× for 4× the cores |
The headline is tail latency, not throughput. If your mock server’s p99 shows up in someone’s CI timing budget, per-core is worth benchmarking; if you are chasing raw RPS, it will not move.
When to use it
- Use per-core on a Linux host that serves high connection counts and where tail latency matters — and measure it for your workload before committing (see Running Benchmarks; the harness’s
--runtimeflag benches both). - Keep the default on small hosts, low-concurrency workloads, or any non-Linux platform.
- Do not switch expecting more throughput, or better scaling as you add cores. Neither was observed.
Experimental. Per-core mode is opt-in and off by default. Its functional behaviour is validated on Linux and the latency benefit above is measured, but on a single machine class at up to 4 physical cores. Behaviour on much larger hosts is not yet characterised (#774) — benchmark it for your workload rather than enabling it blanket.
Platform matrix
| Platform | Per-core mode | Behaviour |
|---|---|---|
| Linux (x86-64 / aarch64) | First-class | SO_REUSEPORT balances accepts across the listener group by 4-tuple hash — the design’s premise. |
| macOS | Falls back, with a warning | BSD/XNU SO_REUSEPORT does not hash-balance TCP accepts across the group (they skew to one socket), so per-core would funnel most connections to one worker — worse than work-stealing. The binary logs the fallback and runs work-stealing; dev boxes lose nothing. |
| Windows | Not offered | No SO_REUSEPORT semantics; the flag is rejected at startup. |
Because macOS silently falls back, always confirm the effective topology from the startup Runtime topology: line rather than assuming the requested mode took effect.
CPU affinity
--runtime-affinity (or RIFT_RUNTIME_AFFINITY=1) pins each per-core worker thread to a CPU core. It is off by default and only meaningful with --runtime per-core; the effect is real on Linux and advisory elsewhere. Leave it off when other processes contend for the same cores — pinning under contention hurts tail latency more than the cache-locality gain is worth.
Blocking pool
Each per-core runtime owns its own spawn_blocking pool (used by JavaScript inject scripts and blocking flow-store backends). To keep the total thread count near a single runtime’s, each worker’s pool is clamped rather than defaulting to 512 threads apiece — so N workers do not multiply into N×512 blocking threads. Note that a few synchronous script paths — notably a JavaScript wait function that computes a delay — run inline on the calling worker rather than on the blocking pool, so keep such scripts cheap under per-core.
Observing load spread
SO_REUSEPORT balances by connection 4-tuple, so a load generator using few source addresses (or few connections) can leave workers unevenly loaded. Benchmark with many connections (≥256) and watch the per-worker accept counter to see the real spread:
curl -s localhost:9090/metrics | grep rift_accepted_connections_total
# rift_accepted_connections_total{worker="0"} 63
# rift_accepted_connections_total{worker="1"} 54
# rift_accepted_connections_total{worker="2"} 75
# rift_accepted_connections_total{worker="3"} 64
The worker label is the accept-loop slot — the worker index under per-core, or a single 0 in the default topology. See Metrics for the full metric set, and the CLI Reference for --runtime / --runtime-affinity and their env-var aliases.
Build Tuning
The shipped release profile is already aggressive:
[profile.release]
opt-level = 3
lto = "thin"
codegen-units = 1
strip = true
For the last few percent on self-hosted deployments you can tune the build further. These are opt-in because they trade portability or compile time for throughput.
target-cpu=native (recommended for self-hosted)
Build for the exact CPU you run on so the compiler can use the newest SIMD/AVX instructions:
RUSTFLAGS="-C target-cpu=native" cargo build --release
or persist it in .cargo/config.toml:
[build]
rustflags = ["-C", "target-cpu=native"]
Caveat: the resulting binary is not portable — it may crash with SIGILL on an older or different CPU. Use it only when you build on (or for) the same microarchitecture you deploy to; the published release artifacts deliberately omit it so they run everywhere.
lto = "fat"
Fat LTO optimizes across the whole dependency graph rather than per-crate (thin). Expect small, single-digit-percent gains at the cost of a substantially longer release build. It is not enabled by default: the compile-time cost is not worth it for CI/release, and the win should be confirmed against the performance regression gate (see the CI perf gate) before adopting. To try it locally, set lto = "fat" under [profile.release].
panic = "abort" — not adopted
panic = "abort" removes unwinding machinery (smaller binary, marginally faster). It is deliberately not used: Rift runs each script (Boa) on a spawn_blocking worker so a buggy or non-yielding script is isolated, and a panic there is contained by the async runtime as a JoinError rather than crashing the server — which relies on unwinding. Under panic = "abort" a single bad script would abort the whole process. Adopting it would require re-validating the scripting and fault paths first, so it stays off pending that work.