# QuestDB Wire Protocol (QWP)

This guide covers QWP ingress and egress from Node.js and browser applications.
It describes the supported public entry points, delivery semantics, authentication,
failure handling, and migration from the existing Node.js sender and the low-level
QWP API.

The documented exports are the public compatibility baseline. Changes to this
surface follow the package's semantic-versioning policy. Imports from internal
source paths are never supported.

## Choose an entry point

| Package                   | Runtime | Use it for                                                                                          |
| ------------------------- | ------- | --------------------------------------------------------------------------------------------------- |
| `@questdb/nodejs-client`  | Node.js | Existing `Sender`, QWP codecs, WebSocket/UDP ingress, egress, TLS, and persistent store-and-forward |
| `@questdb/browser-client` | Browser | Browser-safe QWP ingress, egress, authentication bootstrap, sessions, and codecs                    |

Each distribution exposes its complete API directly from its package root.

Browser applications should install and import `@questdb/browser-client`. Its
published module graph has no Node.js imports, Node.js typings, Node engine
requirement, `undici`, or `ws`. Node-only transports and persistence remain in
`@questdb/nodejs-client`.

QWP uses `/write/v4` for ingress and `/read/v1` for egress. A server must expose
these WebSocket routes; optional features are enabled only when negotiation confirms
that the server supports them.

## Ingress

### Node.js through the existing `Sender`

Changing `http::` or `tcp::` to `ws::` selects QWP while preserving the familiar
fluent row API:

```typescript
import { Sender } from "@questdb/nodejs-client";

const sender = await Sender.fromConfig(
  "wss::addr=questdb.example:9000;token=REST_OR_OIDC_TOKEN;auto_flush=off",
);
await sender.connect();

try {
  await sender
    .table("trades")
    .symbol("symbol", "ETH-USD")
    .floatColumn("price", 2_615.54)
    .timestampColumn("received_at", Date.now(), "ms")
    .at(Date.now(), "ms");
  await sender.flush();
} finally {
  await sender.close();
}
```

`username` plus `password` selects HTTP Basic authentication for the WebSocket
upgrade. `token` selects Bearer authentication. Use `wss::` in production.

`Sender.fromConfig()` uses the same Java-compatible `ws::`/`wss::` vocabulary
as `connectQwpNodeClient()`. Comma-separated or repeated `addr` values configure
ordered failover endpoints, and ingress, egress, pool, and reserved policy keys
are validated from one schema. The standalone sender applies ingress-owned keys;
keys owned only by egress or the pooled facade are accepted as intentional no-ops.

## Configuration-string keys

Every `ws::`/`wss::` connect string is parsed by one schema, shared with the
other QuestDB clients, whichever entry point builds the client —
`Sender.fromConfig()`, `SenderOptions.fromConfig()`, `connectQwpNodeClient()`,
or `createQwpNodeClient()`. An unrecognised key is rejected with
`unknown configuration key: <key>`; a legacy ILP key adds a hint pointing at
where it applies instead.

Keys are grouped by the component that applies them. A client applies the keys
its own side owns and accepts the rest as intentional no-ops, so one connect
string can configure a sender, a query client, or the pooled facade. Every key
also has a programmatic equivalent on the corresponding options object; the
connect string is the portable spelling.

The complete connect string is parsed and validated before typed overrides are
applied. When both forms set the same option, the typed value wins. In
particular, `Sender.fromConfig()` applies `qwp.webSocket.failoverUrls`,
`target`, `zone`, and `senderId` after URL parsing. The primary ingress URL
continues to come from `addr`, because the typed object intentionally omits
`url`.

Typed `ExtraOptions.qwp` sections are transport-specific and a mismatched section
is rejected rather than ignored: `webSocket` and `session` apply only to
`ws::`/`wss::`, `udp` applies only to `udp::`, and `sender` applies to all three
QWP ingress schemes. HTTP(S) and TCP(S) reject every present `qwp` section; an
empty `qwp: {}` remains valid for callers that build the object conditionally.

### Connection

| Key                  | Value              | Default   | Meaning                                                                                  |
| -------------------- | ------------------ | --------- | ---------------------------------------------------------------------------------------- |
| `addr`               | `host[:port]`      | port 9000 | Endpoint. Repeat the key, or comma-separate, for ordered failover.                       |
| `username`, `user`   | string             | —         | HTTP Basic user for the WebSocket upgrade.                                               |
| `password`, `pass`   | string             | —         | HTTP Basic password.                                                                     |
| `token`              | string             | —         | Bearer token; alternative to Basic.                                                      |
| `tls_verify`         | `on`, `unsafe_off` | on        | Certificate verification. `unsafe_off` disables it.                                      |
| `tls_roots`          | path               | —         | PEM file containing trusted private-CA certificates. PKCS#12 is not supported.           |
| `tls_roots_password` | string             | —         | Unsupported by Node; convert PKCS#12 roots to PEM and omit this key.                     |
| `auth_timeout_ms`    | integer ms         | `15000`   | Deadline for the upgrade and authentication exchange.                                    |
| `connect_timeout`    | integer ms         | `15000`   | Deadline for the TCP/TLS transport, and for the upgrade unless `auth_timeout_ms` is set. |

### Ingress

| Key                                     | Value            | Default   | Meaning                                                                |
| --------------------------------------- | ---------------- | --------- | ---------------------------------------------------------------------- |
| `auto_flush`                            | `on`, `off`      | on        | Master switch for all auto-flush triggers.                             |
| `auto_flush_rows`                       | integer          | `1000`    | Flush after this many staged rows.                                     |
| `auto_flush_bytes`                      | integer or `off` | off       | Flush once staged rows reach this estimated size.                      |
| `auto_flush_interval`                   | integer ms       | `100`     | Flush when this long has passed. Checked as rows are added.            |
| `close_flush_timeout_millis`            | integer ms       | `5000`    | Bound on `close()`'s ACK drain. `0` or negative is a fast close.       |
| `transaction`                           | `on`, `off`      | off       | Group each flush into a per-table transaction.                         |
| `request_durable_ack`                   | `on`, `off`      | off       | Require durable ACKs; fails if the server cannot confirm them.         |
| `durable_ack_keepalive_interval_millis` | integer ms       | —         | Poll interval for durable-ACK progress.                                |
| `max_name_len`                          | integer          | `127`     | Maximum table and column name length, in UTF-8 bytes.                  |
| `sender_id`                             | string           | `default` | Names this producer's journal slot on disk. Not sent to the server.    |
| `max_frame_rejections`                  | integer          | `4`       | Consecutive suspect outcomes for one frame before terminal escalation. |
| `poison_min_escalation_window_millis`   | integer ms       | `300000`  | Minimum connected dwell before a poison frame may escalate.            |
| `connection_listener_inbox_capacity`    | integer          | `64`      | Bound on the connection-event inbox before events are dropped.         |
| `error_inbox_capacity`                  | integer          | `256`     | Bound on the `onSenderError` inbox before events are dropped.          |

### Reconnect and failover

| Key                                | Value                       | Default            | Meaning                                                                                                   |
| ---------------------------------- | --------------------------- | ------------------ | --------------------------------------------------------------------------------------------------------- |
| `reconnect_initial_backoff_millis` | integer ms                  | `100` / `50`       | First reconnect delay; grows exponentially with jitter.                                                   |
| `reconnect_max_backoff_millis`     | integer ms                  | `5000` / `1000`    | Ceiling for one reconnect delay.                                                                          |
| `reconnect_max_duration_millis`    | integer ms ≥ 0              | `300000` / `30000` | Budget for a reconnect episode; `0` disables the deadline. The QWP replacement for ILP's `retry_timeout`. |
| `failover`                         | `on`, `off`                 | on                 | Enables endpoint failover for egress.                                                                     |
| `failover_max_attempts`            | integer ≥ 1                 | `8`                | Failover attempts before giving up.                                                                       |
| `failover_backoff_initial_ms`      | integer ms                  | `50`               | First failover delay.                                                                                     |
| `failover_backoff_max_ms`          | integer ms                  | `1000`             | Ceiling for one failover delay.                                                                           |
| `failover_max_duration_ms`         | integer ms                  | `30000`            | Budget for a failover episode.                                                                            |
| `target`                           | `any`, `primary`, `replica` | —                  | Server role this client will accept, on both ingress and egress.                                          |
| `zone`                             | string                      | —                  | Preferred topology zone when ranking endpoints, on both ingress and egress.                               |

The three `reconnect_*` defaults differ by side, shown here as ingress / egress.
Ingress additionally defaults to unlimited attempts, because a running producer
must outlast any outage; egress stops after 8. The `failover_*` keys configure
egress only and share the egress reconnect defaults. Zero disables a duration
budget on both sides, matching `QwpReconnectOptions.maxDurationMs`; the backoff
keys require a positive value, because a zero delay is a hot retry loop rather
than a documented mode.

#### Timer bounds

Every millisecond option that arms a host timer is capped at `2147483647` (about
24.8 days), the largest delay Node and browsers schedule without clamping. A
larger value is rejected with a `RangeError` rather than accepted, because hosts
clamp an over-large delay to about 1 ms — the longest budget you can ask for
would otherwise become the shortest one you get. For a capped option the cap
applies identically to the connection-string key and to the typed option that
overrides it, and to an explicit `timeoutMs` argument such as
`waitForAcknowledged(sequence, timeoutMs)`.

Exemption is a property of the typed spelling, not of the key. `idle_timeout_ms`,
`max_lifetime_ms` and `auto_flush_interval` keep the parser's integer range —
`2147483647` — even though the typed `idleTimeoutMs`, `maxLifetimeMs` and
`autoFlushIntervalMs` below accept any safe integer, because every
connection-string value must parse as an integer within a stated range.

Budgets that are only measured against an elapsed clock, or that are re-clamped
inside a rescheduling loop, accept any safe integer and are deliberately exempt:
`reconnect_max_duration_millis`, `failover_max_duration_ms`,
`poison_min_escalation_window_millis` and
`catch_up_cap_gap_min_escalation_window_millis`, together with the typed
spellings `maxDurationMs`, `poisonMinEscalationWindowMs`,
`catchUpCapGapMinEscalationWindowMs`, `idleTimeoutMs`, `maxLifetimeMs`,
`autoFlushIntervalMs` and `durableAckPollIntervalMs`.

### Store-and-forward (Node only)

Setting `sf_dir` turns on the persistent journal; the rest tune it. A default
shown as a dash is applied downstream of the connect string, by the sender or
session that consumes it.

| Key                                             | Value                          | Default       | Meaning                                                                                                                                                                                                                                                    |
| ----------------------------------------------- | ------------------------------ | ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `sf_dir`                                        | path                           | —             | Slot root. Enables store-and-forward; the journal is `<sf_dir>/<sender_id>`, or `<sf_dir>/<sender_id>-<slot>` pooled.                                                                                                                                      |
| `sf_durability`                                 | `memory`, `periodic`, `append` | `memory`      | Local durability barrier after each vectored append.                                                                                                                                                                                                       |
| `sf_max_total_bytes`                            | integer bytes                  | `10737418240` | Journal ceiling. Reaching it is the one error a producer sees. With `sf_dir`, must reserve at least one whole segment: `sf_max_segment_bytes + 32`. Without `sf_dir` it retunes the 128 MiB memory queue.                                                  |
| `sf_max_segment_bytes`                          | integer bytes                  | `4194304`     | Size of one segment file, and with it the ingress frame cap, since a frame must fit a segment. With `sf_dir`, `sf_max_total_bytes` must leave room for one whole segment of this size plus 32 bytes of headers. Set, it caps a frame without `sf_dir` too. |
| `sf_sync_interval_millis`                       | integer ms                     | —             | Checkpoint interval when `sf_durability=periodic`.                                                                                                                                                                                                         |
| `sf_append_deadline_millis`                     | integer ms                     | `30000`       | How long an append waits for space or a retryable journal fault.                                                                                                                                                                                           |
| `initial_connect_retry`                         | `off`, `sync`, `async`         | `off`         | Startup policy when the server is unreachable. Applies to the memory replay queue as well as to `sf_dir`.                                                                                                                                                  |
| `drain_orphans`                                 | `on`, `off`                    | off           | Adopt and drain journals left by crashed producers.                                                                                                                                                                                                        |
| `max_background_drainers`                       | integer                        | —             | Concurrent orphan drainers.                                                                                                                                                                                                                                |
| `catch_up_cap_gap_min_escalation_window_millis` | integer ms                     | `300000`      | Minimum dwell before an orphan symbol-dictionary cap gap is quarantined. Requires `sf_dir`.                                                                                                                                                                |

### Egress

| Key                 | Value                 | Default | Meaning                                            |
| ------------------- | --------------------- | ------- | -------------------------------------------------- |
| `max_batch_rows`    | integer, 1..1048576   | —       | Rows the server puts in one result batch.          |
| `initial_credit`    | integer ≥ 0           | `0`     | Starting flow-control credit for a query.          |
| `buffer_pool_size`  | integer ≥ 1           | `4`     | Reusable result buffers held per session.          |
| `compression`       | `raw`, `zstd`, `auto` | `raw`   | Result compression to negotiate.                   |
| `compression_level` | integer, 1..22        | —       | zstd level; requires `compression=zstd` or `auto`. |
| `client_id`         | string                | —       | Identifies this client in server-side diagnostics. |

### Pool

Applied by the pooled facade. A standalone sender or query client ignores
them, with one exception noted in the table.

| Key                       | Value       | Default   | Meaning                                                                                                                         |
| ------------------------- | ----------- | --------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `sender_pool_min`         | integer     | `1`       | Senders kept warm.                                                                                                              |
| `sender_pool_max`         | integer     | `4`       | Sender ceiling.                                                                                                                 |
| `query_pool_min`          | integer     | `1`       | Query sessions kept warm.                                                                                                       |
| `query_pool_max`          | integer     | `4`       | Query-session ceiling.                                                                                                          |
| `acquire_timeout_ms`      | integer ms  | `5000`    | How long `borrowSender()` and `borrowQuery()` wait for a free entry.                                                            |
| `query_close_timeout_ms`  | integer ms  | `5000`    | Bound on the CANCEL drain when a query session closes. Also honoured by a standalone egress session built from `egressSession`. |
| `idle_timeout_ms`         | integer ms  | `60000`   | Idle time before a pooled entry is reaped.                                                                                      |
| `max_lifetime_ms`         | integer ms  | `1800000` | Absolute lifetime of a pooled entry.                                                                                            |
| `housekeeper_interval_ms` | integer ms  | `5000`    | How often the pool reaps aged entries.                                                                                          |
| `lazy_connect`            | `on`, `off` | off       | Start without blocking on a first connection.                                                                                   |

### Reserved

`on_write_error`, `on_server_error`, `on_internal_error`, `on_parse_error`,
`on_schema_error` and `on_security_error` are part of the shared vocabulary and
are accepted, but this client does not yet apply them: server-error policy comes
from `qwpDefaultSenderErrorPolicy` and the `onSenderError` stream. They are
listed so a connect string written for another QuestDB client is not rejected.

### Node.js fire-and-forget UDP

`udp::` selects Node-only QWP v1 over IPv4 UDP while retaining the fluent row API:

```typescript
import { Sender } from "@questdb/nodejs-client";

const sender = await Sender.fromConfig(
  "udp::addr=239.1.2.3:9007;max_datagram_size=1400;multicast_ttl=1",
);
await sender.connect();
await sender
  .table("trades")
  .symbol("symbol", "ETH-USD")
  .floatColumn("price", 2615.54)
  .atNow();
await sender.close();
```

The default port is 9007, the maximum datagram size (`max_datagram_size`) is 1400
bytes, and the multicast TTL (`multicast_ttl`) is zero. Both have typed
equivalents in `ExtraOptions.qwp.udp` -- `maxDatagramSize` and `multicastTtl` --
and, like every other typed QWP section, they win when the connect string sets
the same option. `max_datagram_size` accepts
1 through 65507, the IPv4 payload maximum; a larger value is rejected when the sender
is created. Many hosts refuse datagrams well below that ceiling — macOS defaults
`net.inet.udp.maxdgram` to 9216 — so keep the value at or under the path MTU unless
the receiver is known to accept more. A datagram the operating system refuses is
discarded before transmission: it is reported through `onError` and does not advance
`publishedSequence` or `acknowledgedSequence`. Each datagram is
self-contained, contains exactly one table, and uses an inline schema plus
table-local symbol dictionaries. Batches are split at row boundaries;
`QwpUdpDatagramTooLargeError` is raised before transmission when one row cannot
fit. `connectQwpNodeUdpSender()` and `connectQwpNodeUdp()` expose the same
transport from `@questdb/nodejs-client`; `createQwpNodeUdpSender()` builds the
same sender without binding the socket yet, for a caller that wants to construct
it eagerly and connect later.

UDP provides no authentication, TLS, server or durable ACK, transactions,
reconnection, compression, or store-and-forward. Local socket errors are delivered
to `QwpNodeUdpOptions.onError`; like the Java sender, they are observational and do
not retry rows that may already have been handed to the network. UDP is unavailable
from the browser entry point.

Advanced QWP options are accepted in the second argument:

```typescript
const sender = await Sender.fromConfig(
  "wss::addr=questdb.example:9000;token=REST_OR_OIDC_TOKEN;initial_connect_retry=async",
  {
    qwp: {
      webSocket: {
        requestDurableAck: true,
        connectTimeoutMs: 5_000,
        authTimeoutMs: 15_000,
        failoverUrls: ["wss://questdb-dr.example:9000/write/v4"],
        target: "any",
        zone: "eu-west-1a",
        senderId: "producer-a",
        storeAndForward: {
          // The slot root. `senderId` names the journal inside it, so this must
          // not repeat that name or the journal lands in `.../producer-a/producer-a`.
          directory: "/var/lib/my-service/qwp-replay",
          maxBytes: 512 * 1024 * 1024,
          durability: "periodic",
          checkpointIntervalMs: 5_000,
          backpressurePolicy: "wait",
          appendDeadlineMs: 30_000,
          catchUpCapGapMinEscalationWindowMs: 300_000,
          drainOrphans: true,
          maxBackgroundDrainers: 4,
        },
      },
      sender: {
        awaitDurableAck: true,
        autoFlushRows: 10_000,
        autoFlushBytes: 4 * 1024 * 1024,
      },
      session: {
        reconnect: {
          maxAttempts: 0,
          maxDurationMs: 0,
        },
      },
    },
  },
);
```

Node bounds connection establishment in two phases. `connectTimeoutMs` covers
DNS plus the TCP/TLS connection; after that succeeds, `authTimeoutMs` covers the
authenticated HTTP request and WebSocket upgrade. Left unset, `authTimeoutMs`
inherits `connectTimeoutMs`, so narrowing only the connect deadline bounds the
whole opening instead of being exceeded by a default nobody chose; set it to give
the upgrade its own budget, and one endpoint attempt can then take up to the sum
of the two. Both default to 15 seconds. A timeout is reported as `QwpUpgradeError`
with `timeoutPhase` set to `"connect"` or `"authentication"`.
Browsers cannot observe the transport boundary, so their `connectTimeoutMs` continues
to cover the complete WebSocket opening lifecycle and they do not expose
`authTimeoutMs`.

Give each active sender its own store-and-forward directory. The Node.js journal
persists frames and their symbol dictionary before sending. Set
`initialConnectMode: "async"` when a persistent sender must start while every
endpoint is offline. Unless
`awaitServerAck: true` or `awaitDurableAck: true` is selected, `flush()` resolves once
the complete logical flush reaches the configured local journal boundary; a background
drainer then sends it in order. The default `"append"` boundary is locally durable,
while `"periodic"` and `"memory"` trade that immediate guarantee for throughput.
Applications can therefore keep publishing during an outage until the configured
`maxBytes` applies backpressure. A failed journal publication leaves the high-level
rows staged so the caller can retry. When one logical flush is split across multiple
deferred frames, the client checks that every frame can fit before journalling the
prefix; an undersized journal therefore rejects without leaving an unacknowledgeable
partial transaction behind.

`initialConnectMode` selects persistent startup behavior: `"off"` (the default)
makes one
fail-fast attempt, `"sync"` retries on the caller within the configured reconnect
budget, and `"async"` returns immediately while
the background replay loop connects. `Sender.fromConfig()` also accepts
`initial_connect_retry=off|sync|async` when `qwp.webSocket.storeAndForward` is
supplied. Initial authentication and capability failures remain terminal, as does any
upgrade rejection the whole cluster would repeat. A rejection scoped to the endpoint
that returned it — the statuses that keep the sweep walking, `404` among them — is
retried instead of ending the sender, because a producer under `"async"` has already
been handed rows by the time the first attempt runs.
When no mode is explicit, configuring any reconnect duration/backoff key promotes
the initial connection to `"sync"`, so that budget also governs startup.
After a foreground persistent sender has connected successfully at least once, every
one of these failures is retried indefinitely so credential rotation and rolling
capability changes cannot strand its journal. The configured reconnect attempt/duration budget
therefore bounds `"sync"` startup and non-persistent reconnects, not steady-state
foreground store-and-forward recovery.

The connect-string key
`catch_up_cap_gap_min_escalation_window_millis` is the equivalent of
`catchUpCapGapMinEscalationWindowMs`.

`durability` controls the local persistence barrier:

- `"append"` (the default for this object form) issues a data-only durability
  barrier after every vectored positional frame write; manifest and directory
  metadata retain full barriers; hot-spare creation and activation are durable
  before publication resolves.
- `"periodic"` checkpoints segment files, symbol metadata, and directory changes in the
  background. The default interval is 5 seconds, and `close()` performs a final
  checkpoint. A power failure can lose the most recent checkpoint window.
- `"memory"` relies on operating-system writeback. It survives an orderly close and
  normally a process failure, but it makes no power-loss durability promise.

Recovery reports what it had to abandon through `onRecoveryDataLoss`, and logs it when
no handler is supplied, so journal loss is never silent. Trailing records that never
reached disk leave no trace in the segment itself -- a lost page reads back exactly like
the reservation a segment was created with -- so they are detected by comparing what
recovery could read against the append high-water mark in the journal's own metadata.
That mark lives beside the acknowledgement watermark and advances whenever one is
persisted, so detection is bounded by what that pair currently records. Two shapes fall
outside it. A slot whose very first frames are lost before any ACK has nothing to
compare against. So does a slot that has fully drained: an empty journal keeps no
watermark, both values reset, and the records appended after the last acknowledgement
are again undetectable until the next one is persisted. A healthy producer drains
continuously, so this is the steady state rather than an edge case, and it is why
`onRecoveryDataLoss` reporting nothing is not by itself proof that nothing was lost.
`sf_durability=append` avoids the shape altogether by making each append durable before
it is reported as accepted.

The two surfaces do not share a default. `storeAndForward.durability` above
defaults to `"append"`, while the `sf_durability` connect-string key defaults to
`"memory"` — so a journal configured with `sf_dir=` alone never issues a
per-append barrier, and a host crash can lose whatever writeback had not yet
reached disk. Set `sf_durability=append` explicitly when a connect-string journal
has to survive power loss.

`backpressurePolicy: "error"` preserves the existing immediate
`QwpReplayStoreFullError` behavior. Set it to `"wait"` to pause publication until an
ACK advances the checksummed cursor, then a bounded background trimmer deletes fully
drained segments.
`appendDeadlineMs` bounds each such pause and retries of transient journal faults
such as a briefly read-only, full, or descriptor-starved filesystem (30 seconds
by default). Expiry raises `QwpReplayStoreAppendTimeoutError`. A split logical
batch that cannot fit even in an empty journal generation instead fails
immediately with `QwpReplayStoreBatchTooLargeError`; that fit includes the
segment-rounded live-frame allowance preserved by a retained symbol dictionary,
because no ACK or trim is needed to consume it. Waiting appenders do not hold the
journal mutation queue, so ACK cleanup and checkpoint recovery can continue. Corruption and loss of the journal lock remain
immediate failures.
Direct users of `QwpNodeFileReplayStore` can inspect `metrics` for pending records
and segments, checkpoint work, checkpoint failures, active waiters, stalls, and
timeouts.

The persisted symbol dictionary is monotonic for one open journal generation and
cannot be reclaimed by an ACK alone. It counts toward the `maxBytes` target together
with each complete fixed-segment reservation, including the hot spare. While a
dictionary generation is retained, frame segments have an independent rounded
allowance of
`S * max(floor(maxBytes / S), ceil(min(maxBytes, 32 MiB) / S))`, where `S` is one
complete fixed-segment reservation. Dictionary persistence itself is never rejected
by the target, so actual disk usage can exceed it by the retained dictionary and by
the frames that close an open transaction: QuestDB withholds a deferred frame's ACK
until its commit arrives, so the commit is journalled even when the deferred prefix
already fills the journal. Reservations are cumulatively capped at
`S * (floor(maxBytes / S) + max(floor(maxBytes / S), ceil(min(maxBytes, 32 MiB) / S)))`,
saturated at `Number.MAX_SAFE_INTEGER`; a closing batch must still fit the standalone
segment allowance on its own. The retained dictionary is additive to this segment
ceiling. When `S` divides `maxBytes`, accounted physical use may reach the previous
exact result of `dictionaryFileSize + 2 * maxBytes`. Beyond the segment cap, frame
growth remains backpressured until background ACK trimming frees complete segments.
A partly acknowledged segment remains charged to the disk budget until its last live
record is acknowledged.
Once every frame is acknowledged,
`close()` removes the dictionary under the journal lock; the next clean start uses a
fresh symbol-ID space. A partially drained close retains the dictionary required by
the surviving frames.

The journal takes an exclusive lock when it is loaded and holds it until the sender
or session closes. A second live Node.js process using the same directory fails with
`QwpReplayStoreLockedError` before recovery or cleanup can mutate journal contents.
Ownership is held by a `.lock.owner` directory created next to the slot: `mkdir` is
the only exclusive-by-construction filesystem operation available on every supported
platform without a native addon, so exactly one process can create it. The holder PID
is recorded in `.lock.pid` for diagnostics, and the stable `.lock` file is created and
left in place so a slot keeps the on-disk shape a Java client expects. Short-lived
locks under the shared parent directory's `.slot-locks` child serialize orphan
adoption with close/rename/recreate quarantine transitions.

**A Node.js client and a Java client must not use one persistence directory at the
same time.** The Java client locks `.lock` with `flock` on Unix and `LockFileEx` on
Windows. The Node.js client does not participate in those kernel locks, so the two
runtimes will not see each other's lock and can both open the same slot, corrupting
the journal. The persistence format itself remains cross-client: a directory written
by one runtime can be handed to the other once the first has closed it. Only
concurrent access is unsupported, and only between runtimes — two Node.js processes
still exclude each other correctly.

A kernel lock disappears the instant its holder dies; a directory does not. The holder
therefore refreshes the owner directory's mtime every 5 seconds, but a lapsed timestamp
alone never authorizes takeover: the process may be suspended inside a filesystem
write and later resume through an open descriptor. A contender reclaims immediately
only when the owner record names a same-host process that no longer exists, which is
the common case after a crash. A record naming a live PID is never reclaimed automatically, even after its heartbeat
lapses. PID reuse, a container where the app is always PID 1, and `worker_threads`
registries are indistinguishable from a stalled live holder while that PID remains
alive. As with remote-host owners, an operator must confirm that the prior holder is
gone and remove `.lock.owner` before a same-PID restart can adopt the slot. A defunct owner directory is renamed
aside before removal, so two contenders racing to reclaim one slot cannot both win it.
Each acquisition also writes a token into the owner record and checks it before
removing anything, so a release can never take away a directory that has since been
handed to somebody else.

If a holder is paused long enough for its heartbeat to lapse — `SIGSTOP`, a suspended
VM, a stalled filesystem, or any synchronous section that blocks the event loop for
more than 15 seconds — contenders continue to refuse the slot while its recorded
process is alive. The holder fences its first mutation after resuming, revalidates its
acquisition token, and refreshes the heartbeat before continuing. This conservative
rule avoids a proof-to-write race: a successor cannot accept a frame while an older
process still has a writable descriptor for the same journal inode.

New journals use the cross-client SFA persistence layout. Fixed-size
`sf-<generation>.sfa` files have the Java/Rust 24-byte `SF01` header and
`[crc32c, payloadLength, payload]` frame envelope. `sf-manifest.bin` and
`.ack-watermark` use the shared dual-slot checksummed metadata layout, while
`.symbol-dict` uses the shared chunked `SYD1` representation. TypeScript tests load
Java-produced segment and dictionary fixtures and compare TypeScript output with the
same normalized bytes.

Each segment reserves `maxSegmentBytes` of target payload data (4 MiB by default)
plus one frame header so a maximum-sized frame fits. A journal must therefore be
able to reserve at least one whole segment: `sf_max_total_bytes` must be at least
`sf_max_segment_bytes + 32`, counting the 24-byte segment header and the 8-byte
frame header. A smaller total is rejected at configuration time, because no
append could ever reserve its first segment and no acknowledgement could ever
free room for one.

The check reads configuration only and runs before the journal is opened, so it
also applies to a directory that already holds unsent frames: raising
`sf_max_segment_bytes` without raising `sf_max_total_bytes` makes an existing
journal fail to open, leaving its backlog on disk until the total is raised.
Raise `sf_max_total_bytes` first, or drain the directory before changing the
segment size. Two or more segments are recommended, so a rotation can
overlap with a pre-provisioned hot spare instead of provisioning one
synchronously. The active segment and one
pre-sized temporary hot spare keep open file handles; rotation activates the spare.
A process-wide, unreferenced worker provisions replacements, checkpoints dirty paths,
and performs ACK-driven unlink and directory barriers. ACK trimming advances the
durable manifest head before handing removal to that worker and runs in bounded
background batches. Frame append uses a vectored header-plus-payload write, avoiding
an additional payload-sized journal buffer.

A background provisioning, checkpoint or trim failure is parked on the store and
raised from the next journal call, then cleared by the next successful batch. Because
such a fault is transient — a briefly full, read-only or descriptor-starved volume —
reaching one while applying a server acknowledgement reconnects and replays rather
than ending the sender: a filesystem hiccup must not cost a running producer. The same
applies to the advisory lock. Reading its owner record is the only heartbeat step that
needs a file descriptor, so descriptor pressure anywhere in the host process, an `EIO`,
or an NFS `ESTALE` can leave a holder temporarily unable to prove ownership without
anything having taken it; that is reported as `QwpReplayStoreLockUnprovableError`,
which is retryable and parks the same way. Failures
that are verdicts on the journal itself carry `retryable: false` and stay terminal;
today those are `QwpReplayStoreCorruptionError` and `QwpReplayStoreLockLostError`, the
latter only once a takeover has actually been established. The
store persists its acknowledgement cursor before it mutates anything, so a fault at
that moment leaves exactly the state a crash at that moment would leave, and replay
resumes from the persisted watermark.

Recovery validates segment CRCs with a reusable 64 KiB scanner and indexes only frame
sequence, file offset, and payload length. The reconnect loop reads one payload from
its retained segment handle when it is ready to send it; it does not materialize the
complete persisted backlog. Fresh background store-and-forward frames likewise drop
their resident payload after journal publication and are read back on demand. Memory
therefore scales with the active encoding/send window rather than total disk backlog.

Recovery also handles the canonical creation crash window in which a valid SFA
segment becomes durable before its manifest.

A partial final append is repaired only when no valid record follows it. If the
scanner finds an intact CRC-verified record after a damaged record, the slot is
treated as structurally corrupt and preserved through the quarantine path below;
recovery never erases that intact suffix in place.

On startup, a dictionary sidecar truncated at a complete-block boundary is rebuilt
from the ordered symbol deltas embedded in surviving committed frames and healed
before replay. A corrupt or stale dictionary sidecar is replaced when those committed
frames independently reconstruct a complete dense dictionary from ID zero. If the
frame journal is structurally corrupt, or the surviving deltas contain a dictionary
gap or conflict that cannot be reconstructed, the foreground slot is renamed to
`<slot>.unreplayable-N`, marked with `.failed`, and preserved for inspection. The
sender then starts once with a clean slot at the configured path.
`onRecoveryQuarantine` receives the original and quarantine paths plus the terminal
cause and a typed `senderError`. The shared `onSenderError` callback receives the same
`data-loss` / `abandoned` verdict and its `quarantinedPath`. This build-time recovery
notification is synchronous because no connected sender dispatcher exists yet;
callback failures cannot interrupt recovery. Quarantined paths are never adopted by
the orphan scanner. Operational filesystem errors are not quarantined and still fail
startup, so a temporary permissions or disk problem cannot be mistaken for data
corruption.

For a standalone sender, `drainOrphans: true` scans sibling directories beneath the
configured journal directory's parent, excludes the sender's own directory, and
adopts record-bearing slots left by failed producers. Adoption is lock-protected and
uses an independent QWP connection per slot, bounded by `maxBackgroundDrainers` (4 by
default). The scanner runs immediately and then every 30 seconds; set
`orphanScanIntervalMs: 0` for a startup-only scan. Terminal recovery failures create
`.failed` in the slot so a corrupt or permanently rejected head cannot cause a hot
retry loop. If that marker cannot be written, the drainer retains the slot in memory
and retries only the marker on later scans; it does not report abandonment or replay
the terminal head again. After inspection or repair, call
`retryQwpNodeOrphanSlot(slotDirectory)` to make it eligible again.
`onOrphanDrainEvent` reports discovery, drain, lock
contention, quarantine, scanner failures, durable-ACK capability gaps, and transient
all-replica windows through a bounded asynchronous inbox. An abandoned slot also
reports a typed `data-loss` sender error. Callback exceptions cannot interrupt
recovery.

Blocking (`off` or `sync`) foreground startup fails immediately if every usable
endpoint lacks durable-ACK support. Asynchronous foreground startup and steady-state
store-and-forward reconnects retain their records and retry through rolling upgrades.
An orphan slot retries a consecutive durable-ACK capability-gap episode until either
16 connection sweeps or the configured reconnect `maxDurationMs` is reached, then it
is quarantined behind `.failed` (`maxDurationMs: 0` disables only the time half of
the budget). A transport outage or an all-replica window resets both halves of this
orphan budget; neither transient condition can itself quarantine persisted data. The
`durable-ack-unavailable`,
`durable-ack-persistent-failure`, and `primary-unavailable` orphan events expose the
distinction to operators.

A foreground sender retries a symbol-dictionary catch-up entry that is too large for
the current target forever because a larger-cap node may return. An orphan drainer
quarantines that slot only after 16 consecutive incompatible-cap observations and a
minimum five-minute dwell. Tune the dwell with
`catchUpCapGapMinEscalationWindowMs`; an unrelated transport or upgrade failure resets
the episode so outage time cannot accidentally satisfy it.

Keep sibling adoption off unless the parent is a dedicated store-and-forward group:
every record-bearing child directory that is not the foreground slot is considered
eligible. Browser senders never scan or persist local slots.

An offline sender cannot inspect the server-advertised batch cap before its first
publication. Set `qwp.session.maxBatchSizeBytes` to a value no greater than the
smallest target node's cap when offline startup is required.

Set `awaitServerAck: true` when a particular flush must observe QuestDB's protocol ACK
before returning. `awaitDurableAck: true` implies server-ACK waiting and additionally
waits for replicated/durable progress. Browser senders use the in-memory replay
publication boundary by default and do not offer persistent disk publication.

A crash after the server accepts a frame but before local acknowledgement cleanup can
replay that frame, so delivery is at least once. Applications that require exactly-once
effects should use their own stable event key or another idempotency strategy. Closing
a persistent sender stops its drainer but preserves published, unacknowledged frames for
the next sender using that directory.

### Direct high-level API

Use `QwpSender` directly when QWP-only column types or detailed session controls are
needed:

```typescript
import * as qwp from "@questdb/nodejs-client";
import { connectQwpNodeSender } from "@questdb/nodejs-client";

const sender = await connectQwpNodeSender(
  {
    url: "wss://questdb.example:9000/write/v4",
    authorization: `Bearer ${token}`,
  },
  {
    autoFlushRows: 5_000,
    autoFlushBytes: 4 * 1024 * 1024,
    autoFlushIntervalMs: 1_000,
    encode: { symbolDictionary: "delta", gorilla: true },
  },
);

try {
  await sender
    .table("telemetry")
    .symbol("device", "sensor-7")
    .longColumn("sequence", 42n)
    .uuidColumn("event_id", "9f1c96b2-54b8-4d85-bb24-e82c6f1ac120")
    .at(1_775_000_000_000, "ms");
  await sender.flush();
} finally {
  await sender.close();
}
```

A row in progress is the columns staged so far plus the table selected by
`table()`. When a setter or `at()` rejects a value, the sender discards both, so a
half-built row can never reach QuestDB and the next row starts from `table()` again:

```typescript
for (const reading of readings) {
  try {
    await sender
      .table("telemetry")
      .symbol("device", reading.device)
      .floatColumn("value", reading.value)
      .at(reading.timestamp, "ms");
  } catch (error) {
    // Only this row is gone. Rows staged earlier stay pending.
    log.warn(error);
  }
}
await sender.flush();
```

Setters called after a failure raise `table name must be set before adding columns`
rather than quietly joining a fresh row. `cancelRow()` discards a row in progress the
same way without an error, and `reset()` remains the heavier option that also drops
every row staged since the last flush.

### Compiled object-row writers

For repeated rows with one table schema, compile a table-bound writer instead of
sharing the fluent row-builder state:

```typescript
const trades = sender.writer("trades", {
  symbol: qwp.symbol(),
  side: qwp.symbol(),
  price: qwp.double(),
  quantity: qwp.long(),
  timestamp: qwp.designatedTimestamp("ns"),
});

await trades.row({
  symbol: "ETH-USD",
  side: "sell",
  price: 2615.54,
  quantity: 42n,
  timestamp: 1_723_000_000_000_000_000n,
});

await trades.rows([
  {
    symbol: "BTC-USD",
    side: "buy",
    price: 39_269.98,
    quantity: 7n,
    timestamp: 1_723_000_001_000_000_000n,
  },
]);
```

`rows()` accepts `Iterable` and `AsyncIterable` sources and applies the sender's
normal auto-flush, batch-cap, backpressure, transaction, symbol-dictionary, and ACK
settings. The schema is validated once.

`writer()` returns a `QwpTableWriter`, exported from both package roots; it is the
type to annotate with when a compiled writer is stored in a field or passed
around. It is generic over the schema, so `row()` and `rows()` stay typed per
column rather than accepting a bare object.

The schema vocabulary covers every column type the fluent row API can write:

| Field                       | QuestDB type         | Accepted row values                                                                    |
| --------------------------- | -------------------- | -------------------------------------------------------------------------------------- |
| `symbol()`                  | SYMBOL               | `string`                                                                               |
| `varchar()`                 | VARCHAR              | `string`                                                                               |
| `char()`                    | CHAR                 | `string` of one UTF-16 code unit                                                       |
| `bool()`                    | BOOLEAN              | `boolean`                                                                              |
| `byte()`                    | BYTE                 | `number`                                                                               |
| `short()`                   | SHORT                | `number`                                                                               |
| `int32()`                   | INT                  | `number`; `-2_147_483_648` is the NULL sentinel                                        |
| `int64()`, `long()`         | LONG                 | `bigint`; `-9_223_372_036_854_775_808n` is the NULL sentinel                           |
| `float32()`                 | FLOAT                | `number`                                                                               |
| `float64()`, `double()`     | DOUBLE               | `number`                                                                               |
| `timestamp(unit)`           | TIMESTAMP            | `number` or `bigint`; `"ns"` requires `bigint`                                         |
| `designatedTimestamp(unit)` | designated TIMESTAMP | as above, required in every row                                                        |
| `date()`                    | DATE                 | epoch milliseconds; `-9_223_372_036_854_775_808n` is the NULL sentinel                 |
| `binary()`                  | BINARY               | `Uint8Array`, copied on append                                                         |
| `uuid()`                    | UUID                 | canonical UUID text, 16 canonical big-endian bytes, or `{ low, high }`                 |
| `long256()`                 | LONG256              | unsigned 256-bit `bigint`, `0x` hex text, four little-endian words, or `{ words }`     |
| `ipv4()`                    | IPV4                 | dotted-quad text or signed/unsigned packed address; `0.0.0.0` is the NULL sentinel     |
| `geohash(precisionBits)`    | GEOHASH              | raw bits, base-32 text of `precisionBits / 5` characters, or `{ bits, precisionBits }` |
| `decimal64(scale)`          | DECIMAL64            | unscaled `bigint`, decimal text, `number`, or `{ unscaled, scale }`                    |
| `decimal128(scale)`         | DECIMAL128           | as above, scale up to 38                                                               |
| `decimal256(scale)`         | DECIMAL256           | as above, scale up to 76                                                               |
| `doubleArray()`             | DOUBLE[]             | uniform nested arrays or `{ dimensions, values }`; 1 to 32 dimensions                  |
| `longArray()`               | LONG[]               | encodes the protocol type, but current QuestDB servers reject ingestion                |

LONG, LONG256, and nanosecond timestamp inputs are `bigint` so they cannot silently
lose precision. The record forms are exactly what the egress result views hand back,
so a query result value can be written straight into a row without conversion.

Current QuestDB servers accept only DOUBLE arrays for ingestion. `longArrayColumn()`
and `longArray()` remain available for Java-client and protocol parity and encode the
QWP LONG_ARRAY type, but flushing one is rejected by the server with `long arrays are
not supported, only double arrays`. Decoding LONG_ARRAY values in query results remains
supported. QWP arrays may have between 1 and 32 dimensions, and each axis is limited to
268,435,455 elements, the length QuestDB can materialise; the client rejects a larger
rank or axis before encoding a frame, on both ingress and egress.

QuestDB reserves the minimum signed value as NULL for INT, LONG, and DATE. Both the
fluent setters (`int32Column()`, `longColumn()`, and `dateColumn()`) and the compiled
writer fields above follow the Java QWP client and server convention: passing
`-2_147_483_648` to INT, or `-9_223_372_036_854_775_808n` to LONG or DATE, stores
NULL. These sentinel values cannot be stored as ordinary numeric values. Passing
`null` or `undefined`, or omitting a compiled-writer field, also writes NULL.

Widths are spelled out deliberately. The fluent row API predates these names and its
`floatColumn()` and `intColumn()` are 64-bit despite reading as 32-bit, with
`float32Column()` and `int32Column()` as the narrow forms. Compiled writers avoid the
ambiguity: `float32()`/`float64()` and `int32()`/`int64()` mean exactly what they say.

Geohash precision and decimal scale belong to the column, not the value, so they are
fixed when the schema is compiled and validated against the sender's staged schema on
every append. On the fluent row API there is no schema to fix them, so a decimal
column instead locks its scale on the first value staged for it and holds that lock
for the rest of the frame; the lock is released once those rows are published, so the
next frame's first value sets it afresh. Decimal text and `{ unscaled, scale }` values are rescaled to the
column's scale when that is exact, and rejected when it would round: at
`decimal64(2)`, `"1.50"` stages as `150n` and `"1.005"` raises `QwpWriterRowError`.
The `scale` of an `{ unscaled, scale }` value must itself be between 0 and 76,
DECIMAL256's maximum, whatever the column's width; the value is rescaled onto the
column's scale from there.
Decimal text follows the same grammar as ILP. Scientific notation is accepted, so
`"1e3"`, `"-2.5e2"` and `"1.5E-3"` are all valid spellings, and either side of the
decimal point may be omitted, so `".5"` stages as `5` at scale 1 and `"5."` as `5`
at scale 0. Text carrying no digit at all -- `"."`, `"+"`, `".e1"` -- is rejected,
as are `"NaN"` and `"Infinity"`, which ILP accepts but which no
`{ unscaled, scale }` pair can name. An exponent beyond ±1024 is rejected, because
no DECIMAL256 value can name one and expanding it would be an unbounded string
allocation.
Base-32 geohash text carries five bits per character, so `geohash(20)` accepts
`"u33d"` and rejects `"u33"`.

Regular fields may be absent, `null`, or `undefined`, which writes a NULL. A schema
may contain at most one designated timestamp and, when present, that field is required
in every row. Unknown object keys and type mismatches raise `QwpWriterRowError`; bulk
errors include the zero-based row index. A failing row is never partly staged. Rows
successfully completed before a later iterable row fails remain available to flush.

Compiled writers are also available through the regular Node `Sender` when it uses a
QWP transport. Calling `writer()` for an HTTP or TCP ILP sender raises an error. A
writer obtained from a pooled sender lease cannot be used after the lease is closed.

Like the Java QWP sender, `flush()` and `commit()` resolve after the complete
logical flush reaches the local ingress/replay publication boundary. They do
not wait for a server ACK by default. Set `awaitServerAck: true` for an
implicit ACK barrier, or use the explicit sequence API below.

For producer-controlled acknowledgement barriers, publish first and wait for the
cumulative ACK watermark separately:

```typescript
await sender
  .table("telemetry")
  .symbol("device", "sensor-7")
  .longColumn("sequence", 43n)
  .atNow();

const sequence = await sender.flushAndGetSequence();
await sender.waitForAcknowledged(sequence, 5_000);
```

`flushAndGetSequence()` always resolves at the publication boundary, independently
of `awaitServerAck`, and returns the highest stable frame sequence published by that
call. It returns `-1n` when there was nothing to publish. `publishedSequence` and
`acknowledgedSequence` expose the current immutable watermarks. ACK waits are
cumulative, so one later acknowledgement resolves all covered waits and callers may
wait for different sequences concurrently. When durable ACK is being tracked, the
acknowledged watermark advances only after QuestDB reports durable progress;
otherwise it follows ordinary protocol OK responses. Tracking needs both halves:
the caller has to ask for durable progress (`requestDurableAck`, or
`durableAckKeepaliveMs` directly) and the server has to confirm it. Negotiation
alone is not enough, because nothing polls for durable progress that was never
requested — a server that offers the capability unasked leaves the watermark on
ordinary OK ACKs rather than stalling it. A deadline failure raises
`QwpIngressAckTimeoutError` without closing an otherwise healthy session.
If crash recovery retires an incomplete deferred transaction that QuestDB never
received, its frame range does not advance this watermark. Waiting on one of those
frames rejects with `QwpIngressAckAbandonedError`; later frames can still be sent and
acknowledged normally.

Rows are staged until an auto-flush boundary or an explicit `flush()`. A `null` or
`undefined` column value omits that column from the row. `atNow()` asks QuestDB to
assign the designated timestamp; `at(value, unit)` sends an explicit `ns`, `us`, or
`ms` timestamp. `close()` publishes completed rows and waits for the committed-frame
ACK watermark for up to `closeFlushTimeoutMs` (5 seconds by default, matching the
Java client). Set it to `0` or a negative value for a fast close, which publishes
without the ACK drain; publication itself stays bounded, so `close()` always
returns. An unfinished row is still discarded with a warning, including a row
opened by `table()` whose every attempted value was nullish and therefore left it
with zero columns. The configuration-string equivalent is
`close_flush_timeout_millis`.

These warnings, and the other sender-level diagnostics, go to `QwpSenderOptions.log`
when it is supplied and to the same console-backed default logger the rest of the
client uses when it is not, on every entry point: `Sender`, the direct
`createQwpNodeSender()`/`createQwpBrowserSender()` factories, and pooled leases. A
top-level `log` wins over `qwp.sender.log`; pass an explicit no-op to silence them.
`debug` messages stay below the default level, so the per-row staging diagnostics
are not printed unless a supplied logger records them.

QWP is columnar, so a row whose values were all nullish is still a row, and
QuestDB can store it with NULL in every column. What reaches the wire depends on
the closer: `at(value, unit)` carries the designated timestamp as a column, so a
row closed that way always has at least one column, while `atNow()` leaves the
timestamp to the server and can encode the row with a **column count of zero**.
That shape is deliberate and is accepted by WebSocket QWP.

Node UDP matches Java's `QwpUdpSender`: `atNow()` rejects with
`no columns were provided` when this sender currently knows no non-null column
for the selected table. An explicit `at()` row, or a later all-nullish row after the
table schema is known, remains valid. Compiled writers enforce the same UDP rule.
ILP senders reject every field-less row with
`The row must have a symbol or column set before it is closed` -- see the nullish
section of `README.md` for the full per-protocol comparison. Call `cancelRow()`
before the closer when a permitted all-nullish row should be dropped instead of
stored.

`autoFlushBytes` is a soft threshold over estimated raw column-buffer storage and is
disabled by default (`0`). It combines with `autoFlushRows` and
`autoFlushIntervalMs`: reaching any enabled threshold flushes after the completed row,
so one row of overshoot is possible. Once connected, an enabled byte threshold is
clamped to 90% of the server-advertised batch cap. Schema and symbol-dictionary
overhead make this an estimate; exact encoded-size enforcement and automatic frame
splitting remain the ingress session's responsibility. `sender.metrics.pendingBytes`
and `sender.metrics.effectiveAutoFlushBytes` expose the live estimate and applied
threshold. Configuration strings use `auto_flush_bytes=N`; `off` is equivalent to
zero.

The sender automatically maintains connection-scoped symbol IDs, emits dictionary
deltas, tracks acknowledgements, and splits multi-row batches at the smaller of the
client cap and the server-advertised cap. One row that cannot fit is rejected with
`QwpBatchTooLargeError` before it is sent.

Low-level Node sessions expose `publishFrame()`, `publishTables()`, and
`publishTablesDelta()` for local-publication semantics. Their `send*()` counterparts
continue to return the server ACK. Use the publication methods only with persistent
store-and-forward when local durability is the intended completion boundary.
`sendFrameWithPublication()`, `sendTablesWithPublication()`, and
`sendTablesDeltaWithPublication()` expose both boundaries from one operation: await
`publication` before releasing retryable source rows, then await `acknowledgement`
when server acceptance is also required. If a split logical batch cannot be fully
journaled, its unattempted suffix is suppressed and the operation's publication
promise rejects.

A custom `QwpSenderSession` with no optional publication method uses its required
`sendTables()` method's ACK promise as the publication boundary, so an asynchronous
failure retains the sender's rows for retry. Such a session cannot support deferred
transactional auto-flush: the server withholds that ACK until commit, so transactional
mode also requires one of the explicit publication-boundary methods. Automatic
symbol-delta planning is serialized by `sendTablesDelta()` and
`publishTablesDelta()`. The synchronous `sendTablesDeltaWithPublication()` form
cannot wait behind an in-flight delta publication and rejects an overlapping call;
await its `publication` promise before starting another.

Leaving `acknowledgement` unawaited is safe. The session observes it, so an ACK
deadline or a `close()` that rejects a frame still in flight cannot surface as an
unhandled rejection, and both still reach `onError` and the metrics snapshot.

### Browser ingress

Browser applications must use the browser entry point and a same-origin WebSocket
route (directly or through a reverse proxy):

```typescript
import { connectQwpBrowserSender } from "@questdb/browser-client";

const url = new URL("/write/v4", location.href);
url.protocol = location.protocol === "https:" ? "wss:" : "ws:";

const sender = await connectQwpBrowserSender({ url }, { autoFlush: false });

try {
  await sender.table("page_events").symbol("kind", "view").atNow();
  await sender.flush();
} finally {
  await sender.close();
}
```

The browser WebSocket API cannot set `Authorization` or arbitrary `X-QWP-*`
upgrade headers. When authentication is enabled, create QuestDB's HttpOnly session
cookies over REST before opening the WebSocket:

```typescript
import {
  bootstrapQwpBrowserSession,
  connectQwpBrowserSender,
} from "@questdb/browser-client";

await bootstrapQwpBrowserSession({
  url: new URL("/exec", location.href),
  authentication: { type: "bearer", token: oidcOrRestAccessToken },
  // QuestDB Enterprise only; omit this to use the logged-in principal.
  serviceAccount: "market_data_writer",
});

const sender = await connectQwpBrowserSender({ url });
```

Basic authentication is also accepted as `{ type: "basic", username, password }`.
The application obtains OIDC tokens from its identity provider; this package does
not run an interactive OIDC flow. The bootstrap request uses
`credentials: "include"`. REST and WebSocket endpoints therefore need the same
browser origin, or correctly configured credentialed CORS and cookie attributes.
JavaScript never reads `qdb_session` or the Enterprise `qdbServiceAccount` cookie.

Set `sessionBootstrap` on the WebSocket options to repeat bootstrap before every
initial, reconnect, and failover attempt:

```typescript
const sender = await connectQwpBrowserSender({
  url,
  sessionBootstrap: {
    authentication: { type: "bearer", token: oidcOrRestAccessToken },
    serviceAccount: "market_data_writer",
  },
});
```

### Transactions and durable acknowledgement

Transactional auto-flush keeps automatically emitted frames in an open server-side
transaction. `commit()` (an alias for `flush()`) publishes the group-closing frame.
The example also waits for its cumulative durable acknowledgement because it enables
`awaitDurableAck`:

```typescript
const sender = await connectQwpBrowserSender(
  { url, requestDurableAck: true },
  {
    transactional: true,
    autoFlushRows: 10_000,
    awaitDurableAck: true,
    durableAckTimeoutMs: 30_000,
  },
);

for (const event of events) {
  await sender
    .table("events")
    .symbol("source", event.source)
    .longColumn("value", event.value)
    .at(event.timestamp, "ms");
}
await sender.commit();
```

Transactions are atomic per table, not across all tables in one flush. Closing a
sender publishes locally staged transactional rows but does not implicitly commit;
QuestDB rolls the open server transaction back. The sender logs a warning in this case.

In browsers, durable ACK capability is negotiated with a WebSocket subprotocol;
Node.js uses upgrade headers. Setting `awaitDurableAck` automatically requests the
capability unless `requestDurableAck` was set explicitly. The connection fails with
`QwpDurableAckUnavailableError` when the server does not confirm it. Browser durable
tracking is in memory only. Persistent store-and-forward is intentionally Node-only.

Browser ingress adds `qwp_browser_handshake=v1` to the WebSocket URL. Compatible
servers send a small `SERVER_INFO` message immediately after the upgrade, and the
sender uses its exact ingress payload cap for automatic splitting. Older servers
ignore the query parameter; after a bounded 250 ms negotiation window the client
continues in unknown-cap mode. Set `ingressNegotiationTimeoutMs` to tune that window,
or keep using `maxBatchSizeBytes` as a local compatibility limit.

### Reconnect, failover, and roles

The preferred URL and `failoverUrls` form one endpoint set. Endpoints are ranked by
observed health (`healthy`, unknown, transient rejection, transport error, topology
rejection) and then by zone affinity; configuration order breaks ties. Health outranks
zone, so a known healthy cross-zone node is preferred to an untried local node. Every
connection sweep can still try every endpoint, allowing role and health changes to
recover. A non-orderly close demotes the selected endpoint before the next sweep.
Each standalone Node sender/drainer family, and each pooled orphan scanner, shares one
live health ledger among its walkers while keeping independent sweep cursors, so
concurrent drainers cannot consume one another's endpoint attempts. After a foreground
round is exhausted, stale classifications are reset while learned zone tiers persist;
the most recent successful same-zone endpoint remains sticky. Background orphan
drainers publish health observations but never reset foreground classifications.
Ingress reconnect is enabled by default for factory-created browser and Node sessions.
Unacknowledged frames are retained in memory and replayed at least once after a
transport failure. The built-in memory replay queue is capped at 128 MiB. When the
cap is full, publication waits for ACK-driven trimming for at most 30 seconds, then
rejects with `QwpMemoryReplayAppendTimeoutError`; a single frame that can never fit
is rejected immediately with `QwpMemoryReplayFrameTooLargeError`, and a split logical
batch that can never fit is rejected before its first frame with
`QwpMemoryReplayBatchTooLargeError`. Set
`memoryReplayMaxBytes` and `memoryReplayAppendDeadlineMs` on ingress session options
to tune these bounds. The accounting includes a fixed per-frame allowance so many
small frames cannot bypass the byte cap.

The default memory policy uses full-jitter backoff from 100 ms to 5 seconds and a
five-minute per-outage deadline; the initial connection remains fail-fast. Set
`reconnect: false` for one fixed connection. Supplying a `reconnect` object tunes the
bounds, emits lifecycle events through `onEvent`, and retains the earlier opt-in
behavior of retrying initial connection establishment.

QuestDB stops processing a connection's later frames after any ingress NACK so a
cumulative ACK cannot advance across the rejected sequence. Reconnecting sessions
recycle that connection and replay from their last ACK. A fixed `reconnect: false`
session instead becomes terminal and closes immediately after reporting the NACK;
create a new session before sending more rows.

Each retry delay is selected between zero and the current exponential ceiling,
preventing clients disconnected together from retrying in lockstep. Configured attempt
and duration bounds apply to browser/memory reconnect and Node `"sync"` startup. A
Node foreground store-and-forward replay loop remains unbounded after startup. Without
`storeAndForward`, both Node and browser ingress replay only for the lifetime of the
process or page; configuring a Node directory makes the same replay crash-safe.

Ingress also detects a replay head that is repeatedly NACKed or followed by a
non-orderly WebSocket close. `maxFrameRejections` defaults to 4 consecutive strikes,
and `poisonMinEscalationWindowMs` defaults to 5 minutes. Both conditions must be met
before escalation. The window measures _connected_ dwell only: time spent unable to
reach a server is banked and withheld, so an outage never supplies the dwell, and the
strikes a frame has already earned survive the reconnect it caused. Normal (1000),
going-away (1001), service-restart (1012), and try-again-later (1013) closes,
`NOT_WRITABLE`, `DICTIONARY_GAP`, and retriable symbol-dictionary catch-up rejections
reset the strike episode. Abnormal closes (1006), internal-error closes (1011), and transport errors
without close information may count when an unacknowledged replay head exists.
Escalation is terminal for that producer; store-and-forward retains and quarantines
the affected rows for explicit `retryQwpNodeOrphanSlot()` recovery rather than
silently discarding them.

Node.js sees the rejected upgrade status and `X-QuestDB-Role`, so a read-only replica
or catching-up primary can be classified and skipped. Browsers deliberately expose
an opaque upgrade error because their WebSocket API hides the HTTP response. Avoid
placing ingress replica endpoints in a browser endpoint list unless the proxy routes
writers to a primary.

On Node.js a `401` or `403` on the upgrade is classified as an authentication
failure, and it is the one endpoint verdict that is terminal for the entire endpoint
set rather than for the endpoint that returned it. It short-circuits the sweep: the
endpoints ranked after it are never tried, and the reconnect loop rethrows before the
attempt and duration budgets are consulted, so no reconnect setting extends it. A
credential is cluster-wide, so a node rejecting it reports a configuration error that
walking on to a peer would only mask; the Java client applies the same rule. Every
other rejected status keeps the sweep walking, including `404`, which one node can
return mid-deploy while its peers are healthy, and a sweep mixing such attempts stays
retryable if any one of them was.

Note how this composes with health ranking. A non-orderly close demotes the endpoint
it happened on, so a peer that answers `401` can rank ahead of the endpoint that just
dropped; that sweep then ends without the dropped endpoint being retried at all, and
the sender stays terminal even after it recovers. Two cases are exempt. A Node
foreground store-and-forward sender retries these failures indefinitely once it has
connected at least once, so a credential can rotate under a running producer without
losing journaled rows; before that first connection it still retries the statuses
that keep the sweep walking, and only a cluster-wide verdict such as `401` is
terminal. A browser cannot distinguish them at all, because its upgrade
error carries no status; browser authentication failures surface from the REST
session bootstrap instead.

### Observability

Use immutable metrics snapshots for polling and callbacks for event-driven telemetry:

```typescript
import {
  QWP_INGRESS_PROGRESS_KIND,
  createQwpNodeSender,
} from "@questdb/nodejs-client";

const sender = createQwpNodeSender(
  { url: "ws://localhost:9000/write/v4" },
  {},
  {
    reconnect: {
      onEvent: (event) => console.info("QWP connection", event),
    },
    onProgress: (event) => {
      if (event.kind === QWP_INGRESS_PROGRESS_KIND.ACKNOWLEDGED) {
        console.info("accepted through", event.sequence);
      }
    },
    onError: (event) => console.error("QWP ingress", event.error),
    onSenderError: (error) => {
      console.error(
        "QWP rejection",
        error.category,
        error.appliedPolicy,
        error.fromFsn,
        error.toFsn,
      );
    },
  },
);

await sender.connect();
console.info(sender.metrics);
```

Callbacks are placed on bounded asynchronous inboxes and never invoked inside ACK,
reconnect, or orphan-recovery protocol stacks, on ingress and egress alike; the
drop counters below are reported in the ingress metrics snapshot, and
`QwpEgressSession.metrics` reports the egress connection counters
(`deliveredConnectionNotifications` and `droppedConnectionNotifications`) the
same way. Connection events default to 64 retained
entries and errors to 256; `connectionListenerInboxCapacity` and
`errorInboxCapacity` (or their snake-case unified-string keys) tune those bounds.
Overflow drops the oldest pending entry and retains the newest state. Inspect
`droppedProgressNotifications`, `droppedConnectionNotifications`, and
`droppedErrorNotifications` in the immutable ingress metrics; non-zero values mean an
observer is not keeping up. Callback failures are contained. Callbacks still execute on
the JavaScript event loop, so CPU-bound synchronous work should be moved to an
application worker.

`onSenderError` is the Java-parity rejection stream. Its immutable payload includes
`category`, applied policy, raw server status/message, wire message sequence, inclusive
stable `[fromFsn, toFsn]` correlation range, optional single-table attribution, and
`quarantinedPath` for abandoned persistent data. The legacy `onError` callback remains
available for timeouts and general session failures; classified NACK events also expose
the same payload as `event.senderError`. When `onSenderError` is omitted, QWP logs
retriable rejections at `warn` and terminal rejections or abandoned data at `error`.
General asynchronous session failures are likewise logged when `onError` is omitted,
so a background store-and-forward failure is never silent by default. Reconnect and
orphan-drain fallbacks use the same bounded asynchronous error inbox; direct session
fallback logging adds no callback or close-time dependency. Both paths work in browsers
and Node.js.

## Egress

QWP egress streams typed result batches. One connection executes one active query at
a time.

```typescript
import { connectQwpNodeEgress } from "@questdb/nodejs-client";

const session = await connectQwpNodeEgress(
  {
    url: "wss://questdb.example:9000/read/v1",
    failoverUrls: [
      "wss://questdb-replica-2.example:9000/read/v1",
      "wss://questdb-primary.example:9000/read/v1",
    ],
    target: "replica",
    zone: "eu-west-1a",
    authorization: `Bearer ${token}`,
    compression: "zstd",
    compressionLevel: 3,
    maxBatchRows: 4096,
  },
  { queryTimeoutMs: 30_000, bufferPoolSize: 4 },
);

try {
  const query = await session.query(
    "select timestamp, symbol, price from trades where symbol = $1",
    {
      binds: (binds) => binds.setVarchar(0, "ETH-USD"),
      initialCredit: 1024 * 1024,
    },
  );

  for await (const batch of query) {
    console.info(batch.columns);
    for (const row of batch.rows()) console.info(row);
  }

  const completion = await query.completion;
  console.info(completion);
} finally {
  await session.close();
}
```

### Bounded reusable result views

`query()` keeps its convenient materialized batches. For hot paths, `queryViews()`
avoids allocating a JavaScript value array for every column and delivers one
reusable batch view through an awaited callback. The value handed to that callback
is a `QwpResultBatchView`, exported from both package roots:

```typescript
const query = await session.queryViews(
  "select timestamp, symbol, price from trades",
  async (batch) => {
    const timestamp = batch.column(0);
    const symbol = batch.column(1);
    const price = batch.column(2);

    // Fixed-width values are read directly from the QWP little-endian bytes.
    for (let row = 0; row < batch.rowCount; row++) {
      if (!price.isNull(row)) {
        consume(
          timestamp.getLong(row),
          symbol.getSymbol(row),
          price.getDouble(row),
        );
      }
    }

    // Raw views are available for vectorized consumers.
    consumePackedDoubles(price.valuesBytes()!);
  },
  { initialCredit: 256 * 1024 },
);
await query.completion;
```

The callback receives bounded control operations such as `cancel()`,
`grantCredit()`, and `awaitCompletion(timeoutMs)`. The unbounded `completion`
promise is available only on the returned query handle and must be awaited
outside the callback, after its current batch has been released.

For conventional row-major processing, the same batch also owns one reusable
`QwpResultRowView`:

```typescript
batch.forEachRow((row) => {
  if (!row.isNull(2)) {
    consume(row.getLong(0), row.getSymbol(1), row.getDouble(2));
  }
});

// Direct indexed access uses the same flyweight.
const first = batch.row(0);
consume(first.rowIndex, first.getString(1));
```

`forEachRow()` is synchronous, visits rows in index order, propagates callback
exceptions, and re-points the same row object on every iteration. Do not retain
the row object or any zero-copy value returned from it; copy the value inside the
current invocation when it must survive. Calling `batch.row(index)` also returns
that shared object, re-pointed to the requested row.

The batch, its column objects, and every `Uint8Array`/`Int32Array` returned by a
column or row are valid only until the callback settles. The decoder reuses those
objects and its NULL-index, symbol-ID, array-offset, and Gorilla-timestamp scratch
storage for later batches. Copy an individual byte view with `.slice()`, or call
`batch.materialize()` inside the callback, when data must be retained.

Raw fixed-width, NULL, VARCHAR/BINARY, and array data views point into the current
decoded frame; Zstd results point into that batch's decompressed buffer. Accessors
such as `getString()` and `get()` decode or construct only the requested cell. The
callback is awaited before automatic credit is replenished, so the configured
credit window bounds server read-ahead while application work is in progress.

`target` accepts `any` (the default), `primary`, or `replica`. Primary routing also
accepts standalone servers and a primary completing catch-up, matching the Java
client. Both keys apply to ingress and egress on Node.js. In browsers they apply
to egress only: ingress cannot learn a server's role or zone there, because the
WebSocket API hides the upgrade response that carries them. `zone` is an opaque,
case-insensitive preference for `any` and `replica`;
cross-zone endpoints remain eligible. It is ignored for `primary`, which must be
followed across zones. The client validates the authoritative role and zone from the
first QWP `SERVER_INFO` frame before accepting an endpoint, so the same guarantees
work in browsers even though browser WebSocket APIs hide upgrade response headers.

Bind indexes are zero-based in the client: index `0` is SQL placeholder `$1`.
`QwpBindValues` supports booleans, integer and floating-point values, dates,
microsecond and nanosecond timestamps, strings, UUIDs, LONG256, geohashes,
decimals, and typed nulls. Set values in ascending index order. `bindPayload` and
`bindCount` remain advanced escape hatches for pre-encoded data.

Set per-query `resetDictionary: true` to ask the server to reset its
connection-scoped egress symbol dictionary before execution. The client sends the
flag only when `SERVER_INFO` advertises `QUERY_FLAGS`; older servers receive the
same flag-free request as the default path, so this option remains safe during a
rolling upgrade.

Matching Java, the high-level client defaults `initialCredit` to zero, allowing
unbounded server send-ahead. Set a positive session-level or per-query value to bound
wire buffering, particularly in browsers. With positive credit, the exact wire size
of each consumed batch is replenished automatically. Set `autoCredit: false` and call
`query.grantCredit()` for manual control.

Both materialized `query()` results and zero-copy `queryViews()` use a client-side
decoded-batch pool with four slots by default. Set the session-level
`bufferPoolSize` to tune this bound. Materialized decoding pauses when all slots are
queued until iteration requests another batch. For `queryViews()`, callbacks remain
serial and callback-scoped, while the receive loop continues decoding into the other
reusable slots; a slow callback stalls decoding only after the pool fills. This bound
is independent of QWP credit, so `initialCredit: 0` no longer permits an unbounded
queue of decoded batches. Protocol credit remains the stronger end-to-end bound,
particularly in browsers where the WebSocket implementation may buffer raw frames
before JavaScript reads them.

A session `queryTimeoutMs` supplies the default deadline; per-query `timeoutMs`
overrides it, and zero disables it. Expiry rejects iteration and `completion` with
`QwpEgressQueryTimeoutError`, sends QWP `CANCEL`, and drains the terminal response
before the connection accepts another query. Breaking out of `for await` early also
discards buffered batches, restores their flow-control credit, sends `CANCEL`, and
rejects `completion` with `QwpEgressQueryAbandonedError`. Call `query.cancel()` for
explicit cancellation.

`await query.awaitCompletion(timeoutMs)` bounds only the caller's wait and returns
`false` without cancelling when the timeout expires, matching Java
`Completion.await(timeout, unit)`. `query.isDone()` reports terminal state. Use the
query deadline options only when timeout should actively cancel the server query.
The initial and reconnect `SERVER_INFO` timeout defaults to five seconds, matching
Java, and remains configurable through `serverInfoTimeoutMs`.

Cancellation draining is bounded by `cancelDrainTimeoutMs` (5 seconds by default).
Late batches are decoded and credited while the terminal response is pending. If the
server does not terminate the query within the bound, the client fails with
`QwpEgressQueryCancelTimeoutError` and closes the unusable connection instead of
leaving the session permanently occupied.

Node.js and browsers can request Zstd with `compression: "zstd"` or `"auto"` and a
level from 1 through 22. Raw remains the compatibility default. Node uses
`X-QWP-Accept-Encoding`; browsers send the same preference in the URL's
`qwp_accept_encoding` parameter. Compatible servers report the effective codec and
operator-forced level in the existing egress `SERVER_INFO` message. Check
`session.negotiatedCompression` after the handshake. Older servers ignore the query
parameter and safely remain raw. The decoder handles raw and Zstd batches in both
runtimes.

Set transport-level `maxBatchRows` from 1 through 1,048,576 to ask QuestDB for
smaller `RESULT_BATCH` messages. The server clamps the request to its hard cap. Node
sends `X-QWP-Max-Batch-Rows`; browsers use the `qwp_max_batch_rows` URL parameter,
which requires a server that supports browser QWP negotiation. Older servers ignore
the browser parameter and keep their configured batch size.

The connect helpers also enforce that request on what comes back: a `RESULT_BATCH`
declaring more rows than were asked for is rejected as a `QwpProtocolError` before
any column is read. Decoder scratch is sized from the declared row count and
retained per buffer-pool slot for reuse, so an answer above the request would set
the session's memory floor for its lifetime. Set `maxBatchRows` on the session
options to bound a session built directly from a connection; left unset, the cell
cap below is the only bound.

A single `RESULT_BATCH` may declare at most `QWP_MAX_CELLS_PER_BATCH` cells --
32Mi, its rows multiplied by its columns. The row and column caps bound each
dimension on its own, and a compressed body detaches the grid they describe from
the bytes on the wire: an all-NULL column is one bit per cell before Zstd, so
without this bound a few kilobytes of RLE-compressed bitmap declares a result no
heap can hold. The bound is checked before any column is read, and 32Mi cells sits
far above any plausible result -- the widest supported table at 16k rows, or a full
1,048,576-row batch at 32 columns. Lower `maxBatchRows` for genuinely wide tables.

Opening a connection runs under two deadlines. `connect_timeout` covers the TCP/TLS
transport, and `auth_timeout_ms` takes over for the WebSocket upgrade and the
authentication exchange as soon as the transport connects; both default to 15
seconds. Setting only `connect_timeout` bounds both phases with that value, so an
endpoint that accepts TCP and never answers the upgrade -- a stalled proxy or load
balancer -- fails inside the budget you asked for rather than 15 seconds later. Set
`auth_timeout_ms` as well when the upgrade legitimately needs longer than the
transport.

Egress failover is enabled by default in Node.js and browsers. A transport failure or
invalid protocol response closes and deprioritizes that endpoint, reconnects, resets
connection-scoped decoding state, and re-executes the active query. The default policy
uses eight connection sweeps, full-jitter backoff starting at 50 ms and capped at one
second, and a 30-second outage deadline. `QUERY_ERROR` remains a query result and does
not trigger failover.

`session.ready` resolves once with the initial `SERVER_INFO`. Read
`session.serverInfo` for the immutable snapshot from the currently bound endpoint:
role, zone, cluster and node IDs, epoch, capabilities, server clock, and negotiated
compression. Reading the property is non-perturbing and never initiates a failover
walk. If an endpoint dies, it continues to report the previous snapshot until the
transport successfully rebinds, then refreshes to the new endpoint.

Re-execution is at least once: a statement may have completed before its response was
lost, and a consumer may already have observed a prefix of SELECT rows. Queued but
unconsumed batches are discarded automatically. Configure `onReplayReset` when the
application must clear an accumulated prefix before batches restart at sequence zero;
the callback is an optional notification, not an opt-in. Set `reconnect: false` to use
one fixed connection and surface failures without replay. Supplying a `reconnect`
object tunes the failover bounds and also retains the earlier opt-in behavior of
retrying initial connection establishment.

Browser egress uses the same session API:

```typescript
import { connectQwpBrowserEgress } from "@questdb/browser-client";

const readUrl = new URL("/read/v1", location.href);
readUrl.protocol = location.protocol === "https:" ? "wss:" : "ws:";

const session = await connectQwpBrowserEgress({
  url: readUrl,
  failoverUrls: ["wss://replica-2.example/read/v1"],
  target: "replica",
  zone: "eu-west-1a",
  compression: "zstd",
  compressionLevel: 3,
  sessionBootstrap: {
    authentication: { type: "bearer", token: oidcOrRestAccessToken },
  },
});
```

## Combined pooled client

Use `QwpClient` when one long-lived application component needs both ingestion
and concurrent queries. The Node and browser entry points provide configured
factories; each borrowed handle exclusively owns one pooled WebSocket until its
`close()` returns it:

For Node, the recommended common-case API accepts one Java-style
`ws::`/`wss::` cluster string. Every `addr` entry is shared by ingress and
egress; the facade derives `/write/v4` and `/read/v1`, applies the same
authentication and TLS configuration to both sides, and validates ingress,
egress, and pool settings before opening a socket:

```typescript
import { connectQwpNodeClient } from "@questdb/nodejs-client";

const db = await connectQwpNodeClient(
  "wss::" +
    "addr=node-a.example:9000,node-b.example:9000;" +
    `token=${token};` +
    "target=replica;zone=eu-west-1a;" +
    "sender_pool_max=2;query_pool_max=8;",
);
```

Repeated `addr=` keys also accumulate endpoints. Programmatic overrides for
callbacks, custom agents, store-and-forward, sender/session settings, and pool
sizes may be passed as the second argument. The whole string is still validated
before overrides are applied, matching the Java builder's fail-fast behavior.

A custom `wss://` agent is the WebSocket upgrade's sole TLS channel, so it
carries its own certificate verification and cannot be combined with
`tls_verify`, `tls_roots`, or `tls_roots_password` — that combination is
rejected rather than silently dropping either. Configure verification on the
agent instead. Agents are validated per endpoint. A `ws` endpoint takes a plain
`http.Agent`; an `https.Agent` there is rejected. A `wss` endpoint takes any
agent that can serve the scheme, so a tunnelling agent such as
`https-proxy-agent`, `socks-proxy-agent` or `proxy-agent` works even though it
extends `http.Agent` rather than `https.Agent`. Node performs that check itself
and reports a mismatched agent as `ERR_INVALID_PROTOCOL`, which the client
treats as a configuration fault and never retries. For mixed `ws`/`wss`
endpoint sets, omit the shared agent or use homogeneous schemes so failover does
not skip endpoints whose scheme is incompatible with it.

The Node client accepts `tls_roots` only as valid PEM-encoded CA certificates.
Password-protected PKCS#12 trust stores and `tls_roots_password` are rejected:
Node's `pfx` option represents client private-key/certificate identity, not
additional trusted roots. Export the CA certificates to PEM and omit the
password key.

Set `lazy_connect=on` to tolerate an unavailable cluster during startup. In the
JavaScript client, ingress uses memory replay by default, or persistent replay when
`sf_dir` is present, with `initial_connect_retry=async`; egress uses
`query_pool_min=0` and connects on the first query. Explicit
`initial_connect_retry=off|sync` or a positive `query_pool_min` conflicts with
`lazy_connect` and is rejected before the client is created:

```typescript
const db = await connectQwpNodeClient(
  "wss::addr=node-a.example,node-b.example;" + "lazy_connect=on;",
);
```

For unified strings with `sf_dir`, Java-compatible defaults apply: memory
durability, a 10 GiB total journal cap, 4 MiB journal segments, a 30-second
capacity wait, a 5-second close drain, and fail-fast initial connection. Set
`sender_id` to name the disk slot base; pooled senders use `<sender_id>-<slot>`.
A frame must fit a segment, so with `sf_dir` the 4 MiB segment default is also
the ingress frame cap from the first publication onward, before the server has
advertised its own: a row batch above it fails with `QwpBatchTooLargeError`.
Raise `sf_max_segment_bytes`, or lower the cap with
`qwp.session.maxBatchSizeBytes`, to choose a different bound. `sf_max_total_bytes`
must leave room for at least one `sf_max_segment_bytes` segment plus 32 bytes of
headers, so with the 4 MiB segment default `sf_max_total_bytes=4m` is rejected and
`sf_max_total_bytes=1m` requires lowering `sf_max_segment_bytes` too. Without `sf_dir`
there is no default frame cap, and `sf_max_total_bytes`,
`sf_append_deadline_millis` and `sf_max_segment_bytes` retune the built-in
memory replay queue instead -- note that queue's own ceiling is 128 MiB, not the
10 GiB the journal defaults to, while the 30-second append deadline is the same
on both.
The parser also supports `max_name_len` and the Java listener/error inbox
capacity keys. Those capacities actively bound asynchronous connection and
typed-error delivery and are reflected in ingress drop counters.

The object form remains available for cases where constructing the two sides
separately is useful:

```typescript
import { connectQwpNodeClient } from "@questdb/nodejs-client";

const db = await connectQwpNodeClient({
  ingress: {
    url: "wss://questdb.example:9000/write/v4",
    authorization: `Bearer ${token}`,
  },
  egress: {
    url: "wss://questdb.example:9000/read/v1",
    authorization: `Bearer ${token}`,
    target: "replica",
    zone: "eu-west-1a",
  },
  pool: {
    senderPoolMin: 1,
    senderPoolMax: 2,
    queryPoolMin: 1,
    queryPoolMax: 8,
    acquireTimeoutMs: 5_000,
    idleTimeoutMs: 60_000,
    maxLifetimeMs: 30 * 60_000,
    housekeepingIntervalMs: 5_000,
  },
});

try {
  const sender = await db.borrowSender();
  try {
    await sender.table("trades").symbol("symbol", "ETH-USD").atNow();
  } finally {
    // Flushes completed rows and returns the sender; the socket stays pooled.
    await sender.close();
  }

  // Each borrowQuery() resolves to a QwpQueryLease, exported from both roots.
  const [prices, volumes] = await Promise.all([
    db.borrowQuery(),
    db.borrowQuery(),
  ]);
  try {
    // These use independent egress WebSockets and may execute concurrently.
    const drain = async (lease, sql) => {
      const query = await lease.query(sql);
      for await (const batch of query) consume(batch);
      await query.completion;
    };
    await Promise.all([
      drain(prices, "select * from latest_prices"),
      drain(volumes, "select * from hourly_volumes"),
    ]);
  } finally {
    await Promise.all([prices.close(), volumes.close()]);
  }
} finally {
  await db.close();
}
```

Browser applications can likewise describe the cluster, REST/OIDC
authentication bootstrap, and failover order once. A cluster URL may be an
origin, a reverse-proxy base path, or an existing `/write/v4` or `/read/v1`
endpoint; the facade derives both protocol routes while preserving query
parameters. Omit `sessionBootstrap.url` to derive the matching `/exec` route
for every failover endpoint:

```typescript
import { connectQwpBrowserClient } from "@questdb/browser-client";

const db = await connectQwpBrowserClient({
  cluster: {
    url: "wss://node-a.example/qdb",
    failoverUrls: ["wss://node-b.example/qdb"],
    sessionBootstrap: {
      authentication: { type: "bearer", token: oidcOrRestAccessToken },
      serviceAccount: "analytics",
    },
  },
  ingress: { requestDurableAck: true },
  egress: {
    target: "replica",
    zone: "eu-west-1a",
    compression: "zstd",
  },
  pool: { senderPoolMax: 2, queryPoolMax: 8 },
});
```

`url`, `failoverUrls`, and `sessionBootstrap` belong to `cluster` in this
unified form and are rejected if repeated under `ingress` or `egress`.
Side-specific timeouts, WebSocket factories, durable-ACK settings, routing, and
compression remain available as explicit overrides. The original split object
form with complete `ingress` and `egress` trees remains supported for advanced
cases that intentionally connect the two sides differently.

`createQwpNodeClient()` and `createQwpBrowserClient()` build the pooled facade
without contacting the server, leaving the first connect to `connect()` or to the
first borrow. The connecting forms `connectQwpNodeClient()` and
`connectQwpBrowserClient()` prewarm each configured
pool minimum. A prewarm that
fails rejects but does not close the client: connections it did establish stay
pooled, and calling `connect()` again makes a fresh attempt, so a transient
outage at start-up can be retried rather than requiring a new client. Pools grow to
their maximum under concurrent borrows and apply one FIFO acquisition deadline;
exhaustion raises `QwpPoolAcquireTimeoutError`. Query handles are single-flight,
but separate borrowed handles run concurrently. Returning a handle with an active
query sends `CANCEL` and waits for the session's bounded cancellation drain; a
connection that cannot drain is closed instead of being handed to another borrower.
Each query lease exposes the same refreshed snapshot as `lease.serverInfo`; accessing
it after returning the lease raises `QwpClientClosedError` rather than exposing a
pooled connection now owned by another borrower.
The shared housekeeper closes excess connections after `idleTimeoutMs` and recycles
connections older than `maxLifetimeMs` once they are idle, while always retaining
each configured pool minimum. Set either timeout to zero to disable that policy;
`housekeepingIntervalMs` controls how quickly an expired idle connection is noticed.
Prefer returning application-owned leases before calling `QwpClient.close()`.
If shutdown races a borrower, it rejects queued borrowers, closes idle connections,
and cancels active queries before closing every borrowed query connection. A query
lease that is never returned therefore cannot retain a WebSocket after client
shutdown; subsequent operations on it fail as closed. Borrowed senders remain under
their producer's ownership: shutdown waits up to `acquireTimeoutMs` (capped at five
seconds) for them to return and never closes a sender underneath its borrower. A
sender returned during or after shutdown is closed instead of re-entering the pool,
while a sender that outlives the bounded wait owns its eventual teardown.

Pooled sender `close()` flushes completed rows, discards an unfinished row with a
warning, and resets staging before reuse. With Node store-and-forward enabled, the
configured directory is treated as a pool root and each stable sender slot owns a
`sender-N` child directory, avoiding journal lock conflicts. The configured
`senderPoolMin` remains authoritative. A client-level recovery scanner reserves and
drains inactive canonical slots independently of foreground pool connections, both
inside the current range and outside it after `senderPoolMax` is reduced. Foreground
creation and recovery share an atomic slot coordinator, so neither can acquire a
managed journal while the other owns it. This managed-slot recovery is automatic;
`drainOrphans: true` additionally adopts noncanonical sibling slots beneath the pool
root.

## Error handling and cleanup

The public error classes preserve enough context for policy decisions:

| Error                                   | Meaning                                                                                                     |
| --------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `QwpUpgradeError`                       | Classified authentication, role, version, capability, timeout, transport, or browser-opaque upgrade failure |
| `QwpRoleMismatchError`                  | A connected endpoint's advertised role does not satisfy the requested egress target                         |
| `QwpVersionMismatchError`               | The server advertised a QWP version this client does not support                                            |
| `QwpFailoverError`                      | Every eligible endpoint in one connection sweep failed                                                      |
| `QwpBrowserSessionBootstrapError`       | The browser REST session bootstrap was rejected; carries the HTTP status and body                           |
| `QwpPoolAcquireTimeoutError`            | Every pooled connection is leased beyond the configured acquisition deadline                                |
| `QwpPoolResourceError`                  | Creating a new pooled sender or query connection failed                                                     |
| `QwpClientClosedError`                  | The pooled client or an individual returned lease is already closed                                         |
| `QwpDurableAckUnavailableError`         | Durable acknowledgement was required but not negotiated                                                     |
| `QwpProtocolError`                      | A QWP payload is malformed, truncated, or unsupported                                                       |
| `QwpSendError`                          | Base class for a frame that could not be handed to the transport                                            |
| `QwpSendClosedError`                    | The WebSocket was closed, or not open, when a frame was sent                                                |
| `QwpSendTimeoutError`                   | A send did not drain before its deadline; delivery is unknown                                               |
| `QwpSenderCloseTimeoutError`            | Sender shutdown could not publish and ACK-drain all committed ingress frames within its deadline            |
| `QwpIngressNackError`                   | QuestDB rejected an ingress frame                                                                           |
| `QwpIngressAckTimeoutError`             | The cumulative ingress ACK watermark did not reach the requested sequence before its deadline               |
| `QwpIngressAckAbandonedError`           | A recovered frame was deliberately retired without a server ACK                                             |
| `QwpIngressSessionClosedError`          | The ingress session is closed; frames still in flight are rejected with it                                  |
| `QwpBatchTooLargeError`                 | One encoded row cannot fit the effective ingress cap                                                        |
| `QwpWriterRowError`                     | A compiled object-row writer rejected a value, naming its table, column and row                             |
| `QwpUdpDatagramTooLargeError`           | One encoded row exceeds `max_datagram_size` and is rejected before transmission                             |
| `QwpMemoryReplayFrameTooLargeError`     | One frame cannot fit the in-memory replay budget                                                            |
| `QwpMemoryReplayBatchTooLargeError`     | One split logical batch cannot fit the in-memory replay budget                                              |
| `QwpMemoryReplayAppendTimeoutError`     | The in-memory replay queue did not regain capacity before its append deadline                               |
| `QwpReconnectExhaustedError`            | The configured reconnect boundary was reached                                                               |
| `QwpReplayRejectedError`                | A replayed frame was rejected and retained for inspection                                                   |
| `QwpReplayDictionaryError`              | A replay store cannot preserve the dictionary its delta frames require                                      |
| `QwpReplayDictionaryPersistenceError`   | A dictionary sidecar append failed before its delta frame was published; retrying the batch is safe         |
| `QwpUnrecoverableReplayDictionaryError` | The persisted dictionary cannot be restored, so recovered delta frames cannot be replayed                   |
| `QwpReplayStoreFullError`               | The Node.js replay journal reached its configured size                                                      |
| `QwpReplayStoreBatchTooLargeError`      | One split logical batch cannot fit in an empty Node.js replay journal                                       |
| `QwpReplayStoreAppendTimeoutError`      | The Node.js replay journal did not regain capacity before the configured append deadline                    |
| `QwpReplayStoreCheckpointError`         | A periodic Node.js replay-journal checkpoint failed; operations fail closed until a retry succeeds          |
| `QwpReplayStoreLockedError`             | Another process owns the configured Node.js replay directory                                                |
| `QwpReplayStoreError`                   | Base class for every Node.js replay-journal failure                                                         |
| `QwpReplayStoreSegmentTooLargeError`    | One frame exceeds `sf_max_segment_bytes` and can never be journaled                                         |
| `QwpReplayStoreCorruptionError`         | Durable journal bytes are structurally corrupt and cannot be replayed                                       |
| `QwpReplayStoreQuarantinedError`        | Recovery preserved an unreplayable slot and continued on a fresh one                                        |
| `QwpReplayStoreLockLostError`           | The directory lock token changed, so appending could overwrite the new owner's frames                       |
| `QwpReplayStoreLockUnprovableError`     | The lock owner record is unreadable, so ownership could not be re-proved                                    |
| `QwpEgressQueryError`                   | QuestDB returned a terminal query error                                                                     |
| `QwpEgressSessionClosedError`           | The egress session or its connection is closed                                                              |
| `QwpEgressQueryAbandonedError`          | Result iteration ended before the server completed the query                                                |
| `QwpEgressQueryTimeoutError`            | The client deadline expired and cancellation began                                                          |
| `QwpEgressQueryCancelTimeoutError`      | A cancelled query did not produce a terminal server response before the drain deadline                      |
| `QwpEgressReplayRequiredError`          | Deprecated compatibility type from the former explicit replay opt-in                                        |

Always close senders and sessions in `finally`. For a standalone sender, publication
plus ACK draining is bounded by `closeFlushTimeoutMs`; the subsequent WebSocket closing
handshake is bounded by `closeTimeoutMs`. A sender from `borrowSender()` is a _lease_:
its `close()` flushes and returns the slot rather than closing the socket, so it is
bounded by the session's `ackTimeoutMs` (15 seconds by default, and programmatic-only)
rather than by `closeFlushTimeoutMs`. To bound a pooled hand-back yourself, use
`flushAndGetSequence()` followed by `waitForAcknowledged(sequence, timeoutMs)` before
returning the lease. In Node, `connectTimeoutMs` bounds the transport
connection and `authTimeoutMs` the authenticated upgrade, the latter inheriting
the former unless it is set. `closeTimeoutMs` inherits it too: a peer that accepted the
upgrade and then stopped answering makes the closing handshake run to its full budget,
and only `connect_timeout` has a configuration-string key, so an explicit connect budget
bounds shutdown as well unless `closeTimeoutMs` is set. `sendTimeoutMs`, acknowledgement
timeouts, and query deadlines cover later lifecycle phases; configure each according
to the deployment rather than using one very large catch-all value.

## Migration guide

### Existing Node.js `Sender`

For the common fluent API, migration is primarily a transport change:

```diff
- const sender = await Sender.fromConfig("http::addr=localhost:9000");
+ const sender = await Sender.fromConfig("ws::addr=localhost:9000");
```

Review these behavioral differences before rollout:

- QWP `flush()` uses the Java-compatible local-publication boundary by default in
  browsers and Node.js. Set `awaitServerAck` for a protocol ACK barrier, or
  `awaitDurableAck` to wait through durable upload. With Node persistent
  store-and-forward, local publication means durable journal append.
- QWP symbol dictionaries are connection-scoped and automatic.
- Table and column identifiers are rejected locally using the Java client's rules;
  column identity is case-insensitive and preserves the spelling first declared.
- Large batches are split to the negotiated WebSocket payload cap.
- QWP transactional auto-flush is per table and must be explicitly committed.
- Browser and Node QWP ingress reconnect by default with in-memory, at-least-once
  replay. That queue has a 128 MiB cap and a bounded 30-second capacity wait by
  default. Configure Node store-and-forward when replay must survive process failure.
- HTTP/TCP-only keys do not carry over to `ws::`; use the unified QWP connect-string
  vocabulary. Programmatic callbacks, custom agents, and other non-string hooks remain
  available under `extraOptions.qwp`.
- Auto-flush defaults differ from the ILP transports, matching the Java client's
  separate WebSocket defaults: `auto_flush_rows` is `1000` where `http::` uses
  `75000` and `tcp::` uses `600`, and `auto_flush_interval` is `100` ms where both
  use `1000` ms. A workload migrated on the one-line change above therefore sends
  smaller batches far more often; set both keys explicitly to keep its previous
  batching.

Roll out `ws::` per sender instance so the existing protocols can remain in service
during migration.

### Low-level QWP ingress

Code that manually creates `QwpTableBuffer` and calls
`QwpIngressSession.sendTables()` can normally move to `connectQwpNodeSender()` or
`connectQwpBrowserSender()`. Keep low-level sessions only when an application needs
to produce encoded table buffers itself. The high-level sender owns batching, symbol
deltas, ACK tracking, auto-flush, transactions, and durable waits.

When a bare session really is what you want, `connectQwpNodeIngress()` and
`connectQwpBrowserIngress()` open one and hand back a connected
`QwpIngressSession`. Both take the same runtime-specific connection options as
their `*Sender()` counterparts -- so `requestDurableAck` belongs with the
connection, not with the session options -- plus optional session options and an
`AbortSignal` that cancels a first connect still negotiating:

```typescript
import {
  QWP_COLUMN_TYPE,
  QwpTableBuffer,
  connectQwpNodeIngress,
} from "@questdb/nodejs-client";

const session = await connectQwpNodeIngress({
  url: "wss://questdb.example:9000/write/v4",
  authorization: `Bearer ${token}`,
  requestDurableAck: true,
});

try {
  const trades = new QwpTableBuffer("trades");
  // getOrCreateColumn() reserves this row's cell and returns the column to
  // append the value to. It returns null when the row already set that
  // column, because the first value of a row wins.
  trades
    .getOrCreateColumn("symbol", QWP_COLUMN_TYPE.SYMBOL)!
    .values.push("ETH-USD");
  trades.getOrCreateColumn("price", QWP_COLUMN_TYPE.DOUBLE)!.values.push(2615.54);
  // The designated timestamp is the column with an empty name.
  trades
    .getOrCreateColumn("", QWP_COLUMN_TYPE.TIMESTAMP_NANOS)!
    .values.push(1_723_000_000_000_000_000n);
  trades.nextRow();

  // Resolves on the server's acknowledgement of this batch.
  const response = await session.sendTables([trades]);
  if (response.sequence !== null) {
    // Redundant straight after sendTables(); use it to wait for a watermark
    // reached by sends this code did not await. The target is a bigint.
    await session.waitForAcknowledged(response.sequence);
  }
} finally {
  await session.close();
}
```

These two entry points are the supported way to obtain a connected
`QwpIngressSession`, and the one to prefer. The class and the connection
factories are exported too, so `new QwpIngressSession(connection)` and
`QwpIngressSession.connect(factory)` over `connectQwpNodeWebSocket()` /
`createQwpNodeConnectionFactory()` (and their browser counterparts) are
supported as well; they exist for applications that supply their own transport.
Importing from an internal path instead of a package root is what is not
supported. `parseQwpNodeClientConfig()` is the matching low-level helper for
turning a `ws::`/`wss::` connect string into the typed options object those
constructors take, and `scanQwpNodeOrphanSlots()` lists the store-and-forward
slots under a parent directory without starting a drainer.

Low-level `LONG`, `DATE`, and timestamp cells accept either a `bigint` within the
signed 64-bit range or a safe integer `number`. `LONG_ARRAY` applies the same rule to
every element. Coercible values such as booleans and numeric strings, unsafe or
fractional numbers, and out-of-range bigints are rejected before encoding. The
root-exported `qwpGorillaSize()` and `encodeQwpGorilla()` helpers enforce the same
signed 64-bit timestamp bounds for direct codec integrations.

### Java client concepts

The TypeScript high-level sender follows the Java client's core model—fluent rows,
automatic batching, connection-scoped symbol dictionaries, negotiated caps, durable
acknowledgement, and persistent replay—but uses runtime-specific connection factories:

| Java client concept          | TypeScript API                                                |
| ---------------------------- | ------------------------------------------------------------- |
| Sender/builder configuration | `Sender.fromConfig()` in Node.js, or `connectQwp*Sender()`    |
| Fluent table row             | `table()`, typed column methods, `at()` / `atNow()`           |
| Local publish/commit         | `flush()` / `commit()`                                        |
| Explicit ACK barrier         | `flushAndGetSequence()` plus `waitForAcknowledged()`          |
| Durable delivery             | `requestDurableAck` plus `awaitDurableAck`                    |
| Store-and-forward            | Node `storeAndForward`; intentionally unavailable in browsers |
| Fire-and-forget UDP ingress  | Node `udp::` or `connectQwpNodeUdpSender()`                   |
| Query parameters             | `session.query(sql, { binds })`                               |
| Bare ingress session         | `connectQwpNodeIngress()` / `connectQwpBrowserIngress()`      |
| Materialized result batches  | `for await (const batch of query)`                            |
| Reusable result views        | `queryViews()` with column views or `forEachRow()` row views  |
| Egress row/buffer bounds     | `maxBatchRows` and session `bufferPoolSize`                   |

Unlike Java's dedicated dispatcher threads, TypeScript callback inboxes schedule work on
later JavaScript event-loop turns. This keeps user callbacks out of protocol call stacks,
but CPU-bound callback code still blocks the runtime and belongs in a Worker or
`worker_threads` task.

## Development benchmarks

The repository includes diagnostic QWP benchmarks for ingress encoding, fluent sender
construction, symbol dictionaries, egress materialization and reusable views, Zstd,
store-and-forward persistence/recovery, and live completion-boundary latency. See
[`benchmarks/README.md`](benchmarks/README.md) for commands and result interpretation.
They are intentionally not CI performance gates.

## Public API policy

Only the two package roots listed at the top are public: `@questdb/nodejs-client`
and `@questdb/browser-client`. Each declares exactly one `exports` subpath, so
those two specifiers are the whole supported surface. In particular, paths
containing `internal`, `qwp-node`, `client-core`, or `src` are implementation
details even if a bundler can resolve them, and `@questdb/client-core` is a private
workspace package that is never published. The compatibility contract checks the
documented high-level constructors, session classes, errors, constants, and option
signatures exported from both package roots. Additional low-level codec exports
share those roots and are intended for advanced integrations; prefer the high-level
APIs when no custom encoder or transport is required.
