Building a NaaS Platform: Node Orchestration, K8s, Billing

We design and develop full-cycle blockchain solutions: from smart contract architecture to launching DeFi protocols, NFT marketplaces and crypto exchanges. Security audits, tokenomics, integration with existing infrastructure.
Showing 1 of 1All 1305 services
Building a NaaS Platform: Node Orchestration, K8s, Billing
Complex
from 2 weeks to 3 months
Frequently Asked Questions

Blockchain Development Services

Blockchain Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1357
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1249
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    954
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1187
  • image_logo-advance_0.webp
    B2B Advance company logo design
    645
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    926

Building a NaaS Platform: Node Orchestration, K8s, Billing

Running a blockchain node by hand is a simple task for a single node. When you have a hundred, it becomes an infrastructure project with K8s, StatefulSet, snapshot bootstrap, and billing. We specialize in turnkey Node-as-a-Service development and have built several such platforms: from client selection (Ethereum, Solana, BNB) to a production-grade API Gateway with rate limiting and compute units. In practice: one Fortune 500 client replaced manual node management with NaaS—infrastructure costs dropped by 40%, saving $15,000 per month, and over $180,000 annually. Contact us—we'll evaluate your project in 2 days.

How Does a Node-as-a-Service Platform Work?

A NaaS platform provides clients with a single RPC endpoint backed by orchestration of dozens or hundreds of nodes. Each node runs in K8s as a StatefulSet with its own PersistentVolumeClaim. For the budget segment, nodes are shared among clients (shared); for demanding clients, they are fully dedicated (dedicated) or deployed in clusters with load balancing (node clusters).

Why Standard Kubernetes Doesn't Fit Blockchain Nodes?

A regular Deployment in K8s doesn't account for blockchain node specifics, so Kubernetes for blockchain infrastructure requires StatefulSet. Nodes need stateful storage (hundreds of gigabytes), fixed P2P ports, and protection from restarts without losing sync. We use StatefulSet with PVC and headless service—this ensures that on failure the pod isn't recreated on a different node, and data remains tied to storage.

Example StatefulSet configuration for an Ethereum node:

apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: ethereum-geth
spec:
  serviceName: "geth"
  replicas: 1
  selector:
    matchLabels:
      app: ethereum-geth
  template:
    spec:
      containers:
      - name: geth
        image: ethereum/client-go:v1.13.14
        args: ["--datadir=/data", "--http", "--http.addr=0.0.0.0", "--http.vhosts=*", "--http.api=eth,net,web3,txpool", "--ws", "--ws.addr=0.0.0.0", "--maxpeers=50", "--cache=4096"]
        ports:
        - containerPort: 8545
        - containerPort: 8546
        - containerPort: 30303
          protocol: TCP
        - containerPort: 30303
          protocol: UDP
        volumeMounts:
        - name: data
          mountPath: /data
        resources:
          requests:
            memory: "16Gi"
            cpu: "4"
          limits:
            memory: "32Gi"
            cpu: "8"
  volumeClaimTemplates:
  - metadata:
      name: data
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: "fast-nvme"
      resources:
        requests:
          storage: 3Ti

Problems We Solve

Snapshot bootstrapping

Synchronizing the Ethereum mainnet from scratch (snap sync) takes 12–24 hours, and an archive node takes up to 5 weeks. For NaaS, this is critical: clients pay from the first minute. We use snapshot distribution: we create an up-to-date database copy every 7 days, with incremental diffs daily. Node bootstrapping from a snapshot takes 10–15 minutes.

Comparison of Ethereum node sync modes:

Mode Data Size Sync Time RPC Availability
Snap sync ~500 GB 12–24 h Full
Full sync ~1.2 TB 3–5 days Archive
Archive ~15 TB 5–7 weeks Archive + tracing

Client isolation

A single platform hosts startups with a free tier and enterprises with SLA guarantees. We allocate resources via three multitenancy models, each with its own blockchain infrastructure billing approach.

Model Isolation Typical Use Case Pricing Example
Shared Low (single process) Free tier, test projects Pay per CU
Dedicated High (dedicated node) Production, stable RPC Fixed rate
Node cluster Maximum (replicas + LB) Enterprise, HA Custom quote

For blockchain infrastructure billing, we use compute units—each RPC method has a weight in CU: eth_blockNumber — 10 CU, eth_call — 26 CU, trace_replayTransaction — 75 CU.

How We Do It: Stack and Case Studies

RPC Proxy with Intelligent Routing

A custom Go proxy filters dangerous methods (e.g., debug_* only for premium), distributes requests between archive and full nodes, and caches responses (TTL — 1 second for eth_blockNumber). Rate limiting is implemented via Redis sliding window—it's more accurate than token bucket for RPC loads.

// Example RPC proxy with routing logic
package proxy

type RPCRouter struct {
    archivePool   NodePool
    fullNodePool  NodePool
    cacheClient   *redis.Client
}

var archiveMethods = map[string]bool{
    "eth_getBalance":      true,
    "eth_call":            true,
    "eth_getStorageAt":    true,
    "trace_call":          true,
    "trace_replayTransaction": true,
}

func (r *RPCRouter) Route(req *RPCRequest) NodePool {
    if archiveMethods[req.Method] {
        if req.RequiresHistoricalBlock() {
            return r.archivePool
        }
    }
    return r.fullNodePool
}

func (r *RPCRouter) Handle(w http.ResponseWriter, req *RPCRequest, apiKey string) {
    cacheKey := req.CacheKey()
    if cached, err := r.cacheClient.Get(ctx, cacheKey).Bytes(); err == nil {
        w.Write(cached)
        return
    }
    pool := r.Route(req)
    node := pool.GetHealthyNode()
    resp := node.Forward(req)
    if req.IsCacheable() {
        r.cacheClient.Set(ctx, cacheKey, resp, req.CacheTTL())
    }
    r.billing.RecordRequest(apiKey, req.Method, resp.ComputeUnits())
    w.Write(resp)
}

Health Checking with Node State Awareness

Ping doesn't guarantee the node is processing requests. We use sync progress checks: if SyncProgress is not nil or the block is older than 2 minutes, the node is excluded from the pool. Health checks run every 15 seconds.

type NodeHealthChecker struct {
    client *ethclient.Client
}

func (h *NodeHealthChecker) IsHealthy(ctx context.Context) (bool, error) {
    syncing, err := h.client.SyncProgress(ctx)
    if err != nil {
        return false, err
    }
    if syncing != nil {
        return false, fmt.Errorf("node is syncing: %d/%d", 
            syncing.CurrentBlock, syncing.HighestBlock)
    }
    header, err := h.client.HeaderByNumber(ctx, nil)
    if err != nil {
        return false, err
    }
    blockAge := time.Since(time.Unix(int64(header.Time), 0))
    if blockAge > 2*time.Minute {
        return false, fmt.Errorf("block too old: %v", blockAge)
    }
    return true, nil
}

Rate Limiting on Redis

func (rl *RateLimiter) Allow(ctx context.Context, apiKey string, rps int) (bool, error) {
    now := time.Now().UnixMilli()
    window := int64(1000)
    pipe := rl.redis.Pipeline()
    pipe.ZRemRangeByScore(ctx, apiKey, "0", strconv.FormatInt(now-window, 10))
    pipe.ZCard(ctx, apiKey)
    pipe.ZAdd(ctx, apiKey, redis.Z{Score: float64(now), Member: now})
    pipe.Expire(ctx, apiKey, 2*time.Second)
    results, err := pipe.Exec(ctx)
    count := results[1].(*redis.IntCmd).Val()
    return count < int64(rps), nil
}

Billing Based on Compute Units

CREATE TABLE api_keys (
    id UUID PRIMARY KEY,
    customer_id UUID NOT NULL,
    key_hash BYTEA NOT NULL,
    tier VARCHAR(20) NOT NULL,
    rate_limit_rps INTEGER NOT NULL,
    monthly_cu_limit BIGINT,
    node_type VARCHAR(20) NOT NULL,
    created_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE TABLE usage_records (
    id BIGSERIAL PRIMARY KEY,
    api_key_id UUID NOT NULL REFERENCES api_keys(id),
    method VARCHAR(100) NOT NULL,
    chain_id INTEGER NOT NULL,
    compute_units INTEGER NOT NULL,
    response_time_ms INTEGER,
    recorded_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE INDEX idx_usage_billing ON usage_records (api_key_id, recorded_at);

How We Build a NaaS Platform: Phases and Timeline

  1. Analysis and audit (1–2 weeks): define target blockchains, multitenancy model, SLA requirements, and region.
  2. Architecture design (1–2 weeks): select clients (Geth, Reth, Erigon, Solana Agave), prepare K8s schemas, API Gateway, billing.
  3. Core infrastructure implementation (4–6 weeks): StatefulSet templates, snapshot bootstrap pipeline, health checker.
  4. API Gateway and billing development (6–8 weeks): RPC proxy, rate limiting, compute units, Stripe integration.
  5. Observability and self-service portal (6–9 weeks): Prometheus + Grafana, alerting, web interface for key management and metric viewing.
  6. Testing and deployment (2–3 weeks): load testing, security audit, production launch.

Total: 16 to 23 weeks to a production-ready platform. If you want to accelerate, contact our engineers—we'll suggest an appropriate pace.

What's Included

  • Documentation: architecture diagrams, chain addition instructions, on-call runbook.
  • Access: template repository, CI/CD pipeline, monitoring (Grafana dashboards).
  • Training: 2–3 sessions for your team (DevOps and backend).
  • Support: 3 months post-launch (bug fixes, consultations).

Common Mistakes in NaaS Development

  • Using hostNetwork for P2P ports—loses isolation. Better use NodePort or LoadBalancer with a fixed port per node.
  • Lack of caching for frequent RPC methods (eth_chainId, eth_blockNumber)—increases node load and billing.
  • Health check only via TCP—the node may be alive but hundreds of blocks behind the network.

If you've encountered these issues or want to avoid them, order NaaS platform development from us. Learn more about StatefulSet. We have 10+ years of experience in blockchain infrastructure and have completed 50+ projects, including platforms for Fortune 500. Contact us—we'll evaluate your task and propose the optimal solution.

Blockchain Infrastructure Deployment: Nodes, RPC, Indexing

Subgraph fell at 3:47 AM. By morning users saw outdated balances, transactions "hung" in the UI, support received 47 tickets in an hour. Cause: the handler in the subgraph failed on a transaction with a non-standard event log — and the entire index stopped. We have encountered such situations dozens of times. Our experience shows: blockchain infrastructure does not forgive gaps in observability. Guaranteeing uptime without multi-layered monitoring and fault-tolerant architecture is impossible. Over 8 years working with Ethereum, Polygon, and Solana, we have developed an approach that allows predictable deployment of infrastructure of any scale — from a single node to a multichain grid with dozens of subgraphs.

RPC Layer Architecture

Every dApp interaction with the blockchain goes through RPC — the JSON-RPC API provided by a node. Three options:

Managed providers — Alchemy, QuickNode, Infura, Ankr. Minimal operational costs, SLA, built-in monitoring. Limits: rate limits (Alchemy Free: 300 RU/sec), vendor lock, potential downtime during provider incidents. For most projects — the right choice at the start.

Self-owned nodes — full control, no rate limits, no third-party dependence. Cost: archive Ethereum node requires 2.5–3TB SSD, a strong server, and DevOps support. Sync from scratch on Ethereum via Geth/Nethermind — 3–7 days. Justified under high load or latency requirements.

Hybrid — self-owned node as primary, managed provider as fallback. Standard for protocols with high TVL. Proper load balancing can reduce costs by 20–30% compared to pure managed setup. Under high monthly request volume, hybrid saves significantly.

Provider Strength Limitation
Alchemy Supernode, Enhanced APIs, webhooks Expensive on high-volume
QuickNode Low latency, multi-chain More expensive than Alchemy on basic plan
Infura Historical reliability Rate limits on free, one major incident halted half of DeFi
Ankr Cheap, 40+ chains Less stable

How to Set Up an RPC Layer Without a Single Point of Failure?

At least two providers, DNS round-robin with health check every 5 seconds, automatic fallback when latency >500 ms. In practice, this gives 99.99% availability during any provider failure. For protocols with high TVL, we recommend a custom HA-proxy (nginx or Envoy) in front of two managed providers.

Why Is a Hybrid RPC Scheme More Cost-Effective Than Pure Managed?

At high request volumes, managed providers can be very expensive; a hybrid using a self-owned node as primary and a managed fallback cuts costs significantly without losing SLA.

Ethereum Node Clients

Execution clients: Geth (most used), Nethermind (C#, fast sync), Besu (Java, enterprise), Erigon (fastest sync, efficient archive mode ~2TB instead of 3TB).

Consensus clients (post-Merge): Lighthouse (Rust), Prysm (Go), Teku (Java), Nimbus (Nim). Each node after The Merge requires a pair of execution + consensus clients.

For DevOps: eth-docker — Docker Compose configurations for all client combinations. Setting up monitoring via Grafana + Prometheus is mandatory; a standard dashboard is available in each client's repository.

The Graph: Event Indexing

The Graph Protocol — decentralized indexing. A subgraph describes which events from which contracts to index and how to transform them into a GraphQL schema.

Subgraph structure:

  • subgraph.yaml — manifest: contract addresses, startBlock, events to handle
  • schema.graphql — GraphQL schema of entities
  • src/mapping.ts — AssemblyScript event handlers
dataSources:
  - kind: ethereum
    name: UniswapV3Pool
    network: mainnet
    source:
      address: "0x88e6A0c2dDD26FEEb64F039a2c41296FcB3f5640"
      abi: UniswapV3Pool
      startBlock: 12370624
    mapping:
      eventHandlers:
        - event: Swap(indexed address,indexed address,int256,int256,uint160,uint128,int24)
          handler: handleSwap

AssemblyScript handlers — not TypeScript. No nullable types, no closures, no many standard APIs. An error in the handler stops the subgraph indexing on that transaction. Important: add try-catch for operations that can fail (e.g., store.get() for an entity that may not exist).

How to Avoid Subgraph Indexing Stops?

Graph Node logs are monitored in real-time; on hasIndexingErrors = true an alert fires and an automatic node restart (via systemd or Kubernetes). Typical downtime on error — 150–300 seconds to recover. Additionally, for production we set up a watchdog that restarts Graph Node if subgraph lag exceeds 50 blocks.

Choosing Between Hosted Service and Decentralized Network

Graph Hosted Service (free, centralized) is deprecated in favor of Subgraph Studio + Graph Network. For production: deploy on Graph Network with GRT curation signal — the subgraph gets indexers proportional to curation.

Alternatives to The Graph: Ponder (TypeScript, self-hosted, easier to debug), Envio (ultra-fast indexer, supports EVM + non-EVM), Subsquid (TypeScript, own network), Moralis Streams (managed, webhook-based). Our experience shows: for high-load projects with unique logic, Ponder or Envio are more effective — they give full control over the process and do not require GRT tokenomics.

Webhooks and Real-Time Notifications

Alchemy Webhooks and QuickNode Streams allow receiving events in real-time via HTTP webhook or WebSocket. For monitoring addresses, new transactions, mints — this is faster than polling RPC.

Tenderly — platform for monitoring and alerts. You can set up an alert for a specific contract event, balance change, function call with certain parameters. Transaction simulation via Tenderly API is invaluable for debugging.

Monitoring and Observability

Minimum monitoring stack for a protocol:

On-chain: OpenZeppelin Defender Sentinel — watches contract events, triggers webhook or Autotask when conditions are met. Forta Network — community-maintained bots detect anomalies (large withdrawals, flash loans, governance attacks).

Infrastructure: Grafana + Prometheus for nodes, Datadog or Grafana Cloud for managed metrics. Alerts on: node is 10+ blocks behind, RPC latency >500ms, subgraph lag >100 blocks.

Uptime: Better Uptime or PagerDuty on RPC endpoint and subgraph health endpoint (The Graph provides _meta { hasIndexingErrors, block { number } }).

Why Is Monitoring Without Tenderly Insufficient?

Tenderly provides transaction simulation and detailed traces — critical for debugging subgraph and smart contract errors. Forta focuses on network anomalies, not your infrastructure. The combination of Tenderly plus a custom Grafana dashboard covers 90% of incident scenarios.

Multichain Infrastructure

A protocol on 5 chains = 5 separate RPC endpoints, 5 subgraphs, 5 monitoring configs. Manageable but requires deployment automation.

For subgraph multi-network deployment: graph deploy --network mainnet, graph deploy --network arbitrum-one etc. with a unified codebase and network-specific addresses in separate config files.

Chainlink CCIP and LayerZero for cross-chain messaging require monitoring of both chains and transactions on intermediate relayers. A reorg on the source chain after a confirmed mint on the target chain is a classic bridge problem. Solution: wait for finality (on Ethereum ~15 minutes after Merge for economic finality) before confirming on the target chain.

Infrastructure Setup Process

  1. Audit current stack — determine chains, request volume, latency and availability requirements.
  2. Architecture design — select providers, load balancing, redundancy.
  3. Subgraph development — manifest → schema → handlers → testing on local Graph Node → deploy to testnet → mainnet.
  4. Monitoring configuration — Tenderly alerts, Grafana dashboard, PagerDuty integration.
  5. Documentation and runbook — what to do when: subgraph falls behind, RPC downtime, node desync.
  6. Handover to operations — team training, access transfer, first month support.

What's Included

  • Deployment of managed or self-hosted Ethereum, Polygon, BNB Chain nodes
  • RPC layer setup with primary/fallback and load balancing
  • Subgraph development and deployment for your protocol
  • Monitoring connection (Tenderly, Grafana, alerts)
  • Runbook and operations documentation
  • Team training (up to 4 hours online)
  • 30-day support after delivery

Timeline

Task Duration
RPC and basic monitoring setup 1–2 weeks
Subgraph for one protocol 2–4 weeks
Self-hosted node with monitoring 2–3 weeks
Full infrastructure (multi-chain, monitoring, runbooks) 6–10 weeks

All projects are managed in a GitHub/GitLab repository with CI/CD; configuration code stays with you. Order infrastructure deployment — we'll show how to cut costs by 20–30% without losing reliability. Get a consultation — we'll demonstrate how we deployed infrastructure for a protocol with large TVL on Ethereum and Arbitrum. Contact us.