Skip to main content

Limitations

Disclaimer: What This Benchmark Does Not Deliver​

A Production Workload​

Transactions are synthetic proxies: deliberately balanced mixes of PostgreSQL subsystems. See Workloads for details.

This benchmark does not predict the specific performance of any given application. It instead gives a general sense of the relative RDBMS (Relational Database Management System) performance expected of a server type.

Disk I/O Speed​

The scores do not measure storage I/O (Input/Output) performance. Database throughput usually hinges first on disk IOPS (Input/Output Operations Per Second), then on bandwidth. In the cloud, that disk is almost always network-attached block storage, provisioned independently of the server type and entirely up to the user, so it says little about the server itself.

Deploying volumes with high-enough IOPS to never bottleneck across the more than 5,000 server types on Navigator would also be prohibitively expensive. Because of this, we exclude disk speed from the measurement and score the database engine's CPU and memory performance.

Network Performance​

The scores say nothing about a server's network throughput or latency. With a remote client, both bandwidth and especially RTT (Round-Trip Time) between client and server can dominate chatty workloads. Placing the client close to the server is not enough on its own: occasional latency glitches still distort lightweight-workload results. The workload is therefore designed so the remaining RTT is a rounding error; see CPU-Heavy Transactions.

Uniform Data​

Data distribution is uniform rather than Zipfian. The product catalog has 20,000 products, of which order items reference only the first 5,000, but customer, order, and order-item values use g % k modular arithmetic rather than a realistic power-law distribution with a few high-activity customers. See Workloads for details.

Scope Decisions​

Most published database benchmarks focus on storage, network throughput, a single database operation, or one production workload. Our approach focuses on the CPU and memory speed of the server instead: we keep the workload and the PostgreSQL major version constant and vary the hardware across thousands of server types.

Minor Engine Versions​

Only the PostgreSQL major version is fixed. DBaaS (Database as a Service) providers may apply minor PostgreSQL upgrades automatically. This benchmark does not check or enforce the version of a remote target; version selection and pinning are deployment responsibilities. Standalone runs use whichever PostgreSQL 18 minor release the postgres:18 base image had when the image was built.

Instance Size Range​

From small instances (e.g. 1 vCPU and 2 GiB of RAM) to large nodes with hundreds of vCPUs, the same workload must produce meaningful, comparable numbers.

HammerDB TPROC-C (HammerDB's transaction processing workload) and similar suites size their datasets by warehouse count, and pgbench by scale factor. Neither sizing scheme can fulfill this purpose. Because of this, we chose a fixed-size workload that can run on every server type on Navigator with at least 2 GiB of RAM, with larger servers being taxed by concurrency rather than workload size.

Engine Configuration​

The vendor tunes the managed engine, and the harness cannot assume superuser access or GUC (Grand Unified Configuration) control on DBaaS. For this benchmark, we deliberately do not tune the DBaaS engine: the tuning of the managed service is part of what is being measured, so the score reflects the vendor's configuration.

JIT and Parallel Query​

For pgbench_ro, the harness disables PostgreSQL's JIT (Just-in-Time Compilation) and parallel query, and sets work_mem to 64MB; see Workload Settings. These settings are not applied to pgbench_tpcb.

This benchmark measures raw engine and CPU behavior, not JIT compilation variance or Gather scalability. Those are treated as a separate testing axis.

A Single Monolithic Statement​

One SELECT whose eight query blocks are CTEs (Common Table Expressions) combined with UNION ALL is deliberate: it keeps one pgbench transaction equal to one network round trip. This makes the pgbench_ro workload resilient to RTT simulated with netem (Network Emulator); see Latency and pipelining for the experiments. The tradeoff is that per-block planner GUCs (e.g. forcing Merge Join specifically) are not possible without affecting every block.

Pre-Calibrated Weights​

The time shares of the eight query blocks in the pgbench_ro transaction were manually calibrated and remain fixed for all runs. See the block timings and the recalibration procedure for details.

Design Constraints​

With the help of the benchANT team, over many iterations with different tools and configurations, we identified the following core principles.

Memory-Fit, Small Dataset​

This benchmark is designed to use a small dataset of about 303 MiB. On the smallest nodes, the dataset does not fit in memory, which can add some disk overhead, so production runs this benchmark only on servers with at least 2 GiB of RAM, as reported by the vendor. There, the dataset is meant to stay in memory (shared_buffers plus the OS page cache), so that the disk is not read again after warmup. Because the dataset is small by design, large instances are exercised through concurrency rather than data volume.

Read-Only Workload​

The measured pgbench_ro transaction writes no data, so in the steady state of this benchmark it generates no WAL (Write-Ahead Logging) records.

CPU-Heavy Transactions​

Each transaction takes about 70 to 100 ms of server work per connection. With 1 to 5 ms of RTT against roughly 100 ms of server work, network time stays a small fraction of each transaction and cannot dominate. Experiments showed that lightweight read-only transactions are sensitive to network delay, so the default workload uses a heavier cached transaction; see Latency and pipelining for the results.