ArcherDB Benchmark Guide
Guide for running, interpreting, and tracking ArcherDB performance benchmarks.
Overview
ArcherDB benchmarks measure:
- Throughput: Events processed per second
- Read latency: Query response time percentiles
- Write latency: Insert response time percentiles
- Mixed workload: Combined read/write performance
- Scaling: Performance across topologies (1/3/5/6 nodes)
Performance Targets
The repository uses the following comparison gates on comparable hardware profiles:
| Metric | Baseline Target | Stretch Target |
|---|---|---|
| 3-node throughput | >=770K events/sec | >=1M events/sec |
| Read latency P95 | <1ms | <0.5ms |
| Read latency P99 | <10ms | <5ms |
| Write latency P95 | <10ms | <5ms |
| Write latency P99 | <50ms | <25ms |
Running Benchmarks Locally
Prerequisites
# Install benchmark dependencies
pip install -r test_infrastructure/requirements.txt
# Build ArcherDB (lite config for testing)
./zig/zig build -j4 -Dconfig=liteQuick Run (Single Topology)
# Run all benchmark types on 3-node cluster
python3 test_infrastructure/benchmarks/cli.py run --topology 3
# With time limit
python3 test_infrastructure/benchmarks/cli.py run --topology 3 --time-limit 60
# With operation count limit
python3 test_infrastructure/benchmarks/cli.py run --topology 3 --op-count 10000Full Suite (All Topologies)
# Run complete benchmark suite (1/3/5/6 node topologies)
python3 test_infrastructure/benchmarks/cli.py run --full-suite
# Exclude mixed workload tests (faster)
python3 test_infrastructure/benchmarks/cli.py run --full-suite --no-mixedMixed Workload Benchmarks
Control the read/write ratio:
# 80% reads, 20% writes (default)
python3 test_infrastructure/benchmarks/cli.py run --topology 3 --read-write-ratio 0.8
# 50% reads, 50% writes
python3 test_infrastructure/benchmarks/cli.py run --topology 3 --read-write-ratio 0.5
# Write-heavy (20% reads, 80% writes)
python3 test_infrastructure/benchmarks/cli.py run --topology 3 --read-write-ratio 0.2
# Read-only
python3 test_infrastructure/benchmarks/cli.py run --topology 3 --read-write-ratio 1.0Compare to Baseline
# Compare current run to stored baseline
python3 test_infrastructure/benchmarks/cli.py compare baseline.json current.jsonInterpreting Results
Throughput
Events processed per second. Higher is better.
Throughput: <measured events/sec>
Compare against: the latest checked-in baseline for the same hardware/profile
Status: PASS when the run meets or exceeds the local comparison gate
Key factors:
- Batch size (larger batches = higher throughput)
- Network latency (lower = higher throughput)
- Node count (more nodes = higher total throughput, but overhead)
Latency Percentiles
Query response times. Lower is better.
Read Latency:
P50: 0.3ms (median)
P95: 0.8ms (95% of requests)
P99: 4.2ms (99% of requests)
Target: P95 <1ms, P99 <10ms
Status: PASS
Percentile meanings:
- P50 (median): Typical user experience
- P95: 95% of requests are this fast or faster
- P99: Captures tail latency, important for SLAs
Confidence Intervals
All means are reported with 95% confidence intervals:
P95: 0.8ms +/- 0.1ms (95% CI)
Narrower intervals = more stable measurements. Wide intervals suggest:
- Insufficient samples
- High variance in measurements
- System noise
Coefficient of Variation (CV)
Measures result stability:
CV: 8.2% (target: <10%)
- <10%: Results are stable, trustworthy
- 10-20%: Somewhat noisy, consider more samples
- >20%: High variance, investigate system state
Regression Detection
Threshold
A regression is detected when:
- Performance degrades by >10% from baseline
- Statistical test confirms significance (p < 0.05)
Statistical Method
We use Welch’s t-test (unequal variance):
from scipy.stats import ttest_ind
t_stat, p_value = ttest_ind(baseline, current, equal_var=False)Benefits:
- Does not assume equal variance between runs
- Robust to different sample sizes
- Standard statistical rigor
Comparison Report
Regression Analysis
==================
Baseline: previous checked-in artifact
Current: current run artifact
Change: <measured delta>
p-value: <measured significance>
Status: REGRESSION DETECTED when the comparison crosses the configured threshold
Recommendation: Investigate recent changes when the current run underperforms the baseline
Historical Tracking
Manual Publication Runs
Maintainers can run the publication workflow manually and promote benchmark results into checked-in history:
benchmarks/history/
2026-01-05.json
2026-01-12.json
2026-01-19.json
2026-01-26.json
2026-02-02.json
...
Baseline Files
Local baselines for regression detection live under:
reports/baselines/
baseline-1node-*.json
baseline-3node-*.json
baseline-5node-*.json
baseline-6node-*.json
Visualization
Results are visualized using github-action-benchmark:
- Throughput graph: Events/sec over time
- Latency graph: P95/P99 over time
- Scaling graph: Performance vs node count
View at:
https://github.com/[org]/archerdb/benchmarks
CI Integration
Publication Workflow
The benchmark-weekly.yml workflow:
- Spins up clusters (1/3/5/6 nodes)
- Runs full benchmark suite
- Compares to baseline
- Alerts on >10% regression
- Can promote approved results into
benchmarks/history/ - Updates benchmark graphs
Alerts
On regression detection:
- Workflow fails (visible in GitHub)
- Comment posted on triggering commit
- GitHub issue created with details
- Slack notification (if configured)
Manual Trigger
# Trigger weekly benchmark manually
gh workflow run benchmark-weekly.ymlProgrammatic Usage
from test_infrastructure.benchmarks import BenchmarkOrchestrator, BenchmarkConfig
# Create orchestrator
orchestrator = BenchmarkOrchestrator()
# Configure benchmark
config = BenchmarkConfig(
topology=3,
time_limit_sec=60,
op_count_limit=10_000,
read_write_ratio=0.8, # 80% reads, 20% writes
)
# Run individual benchmarks
throughput = orchestrator.run_throughput_benchmark(3, config)
read_latency = orchestrator.run_latency_read_benchmark(3, config)
write_latency = orchestrator.run_latency_write_benchmark(3, config)
mixed = orchestrator.run_mixed_workload_benchmark(3, config)
# Access results
print(f"Throughput: {throughput['throughput_events_per_sec']}")
print(f"Read P95: {read_latency['p95_ms']}ms")
print(f"Write P95: {write_latency['p95_ms']}ms")
# Run full suite
results = orchestrator.run_full_suite(
topologies=[1, 3, 5, 6],
include_mixed=True,
)Output Formats
| Format | Location | Purpose |
|---|---|---|
| JSON | reports/benchmarks/*.json |
CI automation, data processing |
| CSV | reports/benchmarks/*.csv |
Spreadsheet analysis |
| Terminal | stdout | Interactive feedback |
| Markdown | Updated docs | Human review |
JSON Format
{
"timestamp": "2026-02-01T02:00:00Z",
"topology": 3,
"throughput": {
"events_per_sec": 823456,
"target": 770000,
"passed": true
},
"read_latency": {
"p50_ms": 0.3,
"p95_ms": 0.8,
"p99_ms": 4.2,
"samples": 10000
},
"write_latency": {
"p50_ms": 2.1,
"p95_ms": 8.5,
"p99_ms": 42.3,
"samples": 2000
},
"metadata": {
"version": "1.0.0",
"git_sha": "abc1234",
"runner": "ubuntu-latest-8-cores"
}
}Best Practices
Consistent Environment
- Use dedicated hardware or CI runners
- Close other applications during local runs
- Use the same build configuration tier (for example, lite vs standard)
Warm-up
SDKs with JIT compilation (Java, Node.js) need warm-up:
| SDK | Recommended Warm-up |
|---|---|
| Java | 500 iterations |
| Node.js | 200 iterations |
| Python | 100 iterations |
| Go | 100 iterations |
| C | 50 iterations |
Sample Size
- Minimum 1000 samples for percentile accuracy
- Continue until CV < 10% (stability check)
- Maximum 10 stability check rounds
Fresh Cluster
Each benchmark run should use a fresh cluster to ensure isolated measurements without accumulated state affecting results.
Troubleshooting
Results Vary Widely
- Increase sample count
- Check for background processes
- Verify network stability
- Use constrained build (
-Dconfig=lite)
Benchmarks Hang
- Check server health:
curl http://localhost:3001/ping - Verify cluster formed:
curl http://localhost:3001/topology - Check logs for errors
Results Don’t Match CI
- Use same hardware profile as CI
- Use same build configuration
- Account for warm-up differences
See Also
- Single-Node Evidence, August 2026 - Measured ArcherDB vs Valkey vs PostGIS comparison with reproduction steps
- Detailed Benchmark Framework - Statistical methodology
- Testing Guide - Running tests locally
- CI Tiers - Weekly benchmark workflow
- Performance Tuning - Optimization guidance
Last updated: 2026-02-01
Edit this page