Skip to content

Upgrades, bugfixes and improvements#76

Open
rnbrady wants to merge 39 commits into
bitauth:masterfrom
ParyonUSD:master
Open

Upgrades, bugfixes and improvements#76
rnbrady wants to merge 39 commits into
bitauth:masterfrom
ParyonUSD:master

Conversation

@rnbrady

@rnbrady rnbrady commented May 10, 2026

Copy link
Copy Markdown

Addresses:
#71
#72
#73
#74
#75
#55
#77

Leaving these here for others who may find them helpful, courtesy of ParyonUSD.

Code by Codex GPT 5.5 xhigh and Claude Fable 5.

rnbrady added 27 commits May 7, 2026 13:22
rnbrady and others added 9 commits June 12, 2026 15:11
The default RollingUpdate strategy runs old and new agent pods
concurrently until the new pod passes readiness (minutes of node
init). The agent is a singleton writer — inserts are idempotent but
reorg handling assumes exclusive access to node_block, so an
overlapping agent can transiently resurrect stale-chain acceptance
rows. Recreate guarantees at most one agent at the cost of a brief
indexing gap per rollout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
3.5 years of CVE and performance fixes over v2.16.1, staying on the v2
line. First startup against an existing database performs a one-way
metadata catalog upgrade (47 -> 48); v2.16.1 will not start against
the upgraded catalog, so snapshot hdb_catalog before deploying.
Rehearsed against a schema-only copy of the production database
(including its catalog-47 hdb_catalog): migration completes, metadata
is consistent, and existing query shapes execute unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Restore hashes with encode(hash, 'hex') in Postgres instead of
materializing ~1.5M Buffers and hex-encoding each on the agent event
loop, which dominated startup on large databases. Log the restore
count and duration on completion. Replace the misleading "database
configuration or connectivity problem" warning — which fired after
just 6 seconds of a legitimately slow restore — with an accurate
"still restoring" message gated to 15s. e2e: verify the SQL-side
encoding matches the previous client-side conversion exactly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The "misconfigured or unresponsive" warning fired for nodes whose P2P
connection was ready but whose chain state was still being restored
from the database — legitimate work proportional to chain length
(~9s for a 960k-block chain sharing the pool with the block hash
restore). Split the warning: nodes without a ready P2P connection keep
the misconfigured/unresponsive text; connected nodes get an accurate
"still restoring chain state" message. Log each node's registration +
chain restore duration on completion, and let the block hash restore
warning use the same 5-second base patience as node warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Node 18 has been EOL since April 2025. The e2e suite already runs
under Node 24 locally (89 tests passing), and the agent boots cleanly
in the node:24-alpine image. engines raised to >=22.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Same-major driver bump carrying 3.5 years of fixes; e2e suite (89
tests, real sync against Postgres 18) passes unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
/health-check answers ~1s after start, but both probes waited 30s
before the first check, so pods showed NotReady for 30+ seconds.
Readiness now 5s delay / 10s period / 5s timeout; liveness stays
coarse so event-loop stalls during catch-up don't get pods killed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant