High Availability#

To gain a higher level of availability for your Instance, you can

  • create more Kubernetes Cluster Nodes
  • create more replicas of the nscale and nplus components
  • distribute those replicas across multiple nodes using anti-affinities

This is how:

helm install \
  --values samples/ha/values.yaml
  --values samples/environment/demo.yaml \
  sample-ha nplus/nplus-instance

The license and the storage layer server ids#

A nscale Server Storage Layer runs with a server id, and in a high availability setup each node needs its own. The sample therefore sets them explicitly, 4711 for nstla and 4712 for nstlb:

nstla:
  serverID: 4711
nstlb:
  serverID: 4712

For that to work, the license must not carry a server id of its own. If it does, the storage layer refuses to start and says so clearly:

Server-ID 4711 in config is different from Server-ID 1514574720 from license
Start aborted due to configuration errors

A license with a ServerId property pins the storage layer to exactly that one id and makes a multi node setup impossible — in that case, ask Ceyoniq for a license without it, rather than bending the server ids to match the license.

The essents of the values file is this:

  • We use three (3) nscale Server Application Layer, two dedicated to user access, one dedicated to jobs
  • if the jobs node fails, the user nodes take the jobs (handled by priority)
  • if one of the user nodes fail, the other one handles the load
  • Kubernetes takes care of restarting nodes should that happen
  • All components run with two replicas
  • Pod anti-affinities handle the distribution
  • any administration component only connects to the jobs nappl, leaving the user nodes to the users
  • PodDisruptionBudgets are defined for the crutial components. These are set via minReplicaCount for the components that can support multiple replicas, and minReplicaCountType for the first replicaSet of the components that do not support replicas, in this case nstla.
web:
  replicaCount: 2
  minReplicaCount: 1
rs:
  replicaCount: 2
  minReplicaCount: 1
ilm:
  replicaCount: 2
  minReplicaCount: 1
cmis:
  replicaCount: 2
  minReplicaCount: 1
webdav:
  replicaCount: 2
  minReplicaCount: 1
nstla:
  minReplicaCountType: 1
administrator:
  nappl:
    host: "{{ .component.prefix }}nappljobs.{{ .Release.Namespace }}"
  waitFor:
    - "-service {{ .component.prefix }}nappljobs.{{ .Release.Namespace }}.svc.cluster.local:{{ .this.nappl.port }} -timeout 600"
pam:
  nappl:
    host: "{{ .component.prefix }}nappljobs.{{ .Release.Namespace }}"
  waitFor:
    - "-service {{ .component.prefix }}nappljobs.{{ .Release.Namespace }}.svc.cluster.local:{{ .this.nappl.port }} -timeout 600"
nappl:
  replicaCount: 2
  minReplicaCount: 1
  jobs: false
  waitFor:
    - "-service {{ .component.prefix }}nappljobs.{{ .Release.Namespace }}.svc.cluster.local:{{ .this.nappl.port }} -timeout 600"
nappljobs:
  replicaCount: 1
  jobs: true
  disableSessionReplication: true
  ingress:
    enabled: false
  snc:
    enabled: true
  waitFor:
    - "-service {{ .component.prefix }}database.{{ .Release.Namespace }}.svc.cluster.local:5432 -timeout 600"
application:
  nstl: 
    host: "{{ .component.prefix }}nstl-cluster.{{ .Release.Namespace }}"
  nappl:
    host: "{{ .component.prefix }}nappljobs.{{ .Release.Namespace }}"

Files#

Download all files of this sample

build.sh#

#!/bin/bash
#
# This sample script builds the example as described. It is also used to build the test environment in our lab,
# so it should be well tested.
#

# Make sure it fails immediately, if anything goes wrong
set -e

# -- ENVironment variables:
# CHARTS: The path to the source code
# DEST: The path to the build destination
# SAMPLE: The directory of the sample
# NAME: The name of the sample, used as the .Release.Name
# KUBE_CONTEXT: The name of the kube context, used to build this sample depending on where you run it against. You might have different Environments such as lab, dev, qa, prod, demo, local, ...

# Check, if we have the source code available
if [ ! -d "$CHARTS" ]; then
    echo "ERROR Building $SAMPLE example: The Charts Sources folder is not set. Please make sure to run this script with the full Source Code available"
    exit 1
fi
if [ ! -d "$DEST" ]; then
    echo "ERROR Building $SAMPLE example: DEST folder not found."
    exit 1
fi
if [ ! -d "$CHARTS/instance" ]; then
    echo "ERROR Building $SAMPLE example: Chart Sources in $CHARTS/instance not found. Are you running this script as a subscriber?"
    exit 1
fi

# Set the Variables
SAMPLE="ha"
NAME="sample-$SAMPLE"

# Output what is happening
echo "Building $NAME"

# Create the manifest
mkdir -p $DEST/instance
helm template --debug \
     --values $SAMPLES/ha/values.yaml \
     --values $SAMPLES/hid/values.yaml \
     --values $SAMPLES/application/empty.yaml \
     --values $SAMPLES/environment/$KUBE_CONTEXT.yaml \
     --values $SAMPLES/resources/$KUBE_CONTEXT.yaml \
     $NAME $CHARTS/instance > $DEST/instance/$SAMPLE.yaml

# creating the Argo manifest
mkdir -p $DEST/instance-argo
helm template --debug \
     --values $SAMPLES/ha/values.yaml \
     --values $SAMPLES/hid/values.yaml \
     --values $SAMPLES/application/empty.yaml \
     --values $SAMPLES/environment/$KUBE_CONTEXT.yaml \
     --values $SAMPLES/resources/$KUBE_CONTEXT.yaml \
     $NAME-argo $CHARTS/instance-argo > $DEST/instance-argo/$SAMPLE-argo.yaml

values.yaml#

components:
  nappl: true
  nappljobs: true
  web: true
  mon: true
  rs: true
  ilm: true
  erpproxy: true
  erpcmis: true
  cmis: true
  database: true
  nstl: false
  nstla: true
  nstlb: true
  pipeliner: true
  application: true
  administrator: true
  webdav: true
  rms: false
  pam: true
web:
  replicaCount: 2
  minReplicaCount: 1

rs:
  replicaCount: 2
  minReplicaCount: 1

ilm:
  replicaCount: 2
  minReplicaCount: 1

erpproxy:
  replicaCount: 2
  minReplicaCount: 1

erpcmis:
  replicaCount: 2
  minReplicaCount: 1

cmis:
  replicaCount: 2
  minReplicaCount: 1

webdav:
  replicaCount: 2
  minReplicaCount: 1

administrator:
  nappl:
    host: "{{ .component.prefix }}nappljobs.{{ .Release.Namespace }}"
  waitFor:
    - "-service {{ .component.prefix }}nappljobs.{{ .Release.Namespace }}.svc.cluster.local:{{ .this.nappl.port }} -timeout 600"

pam:
  nappl:
    host: "{{ .component.prefix }}nappljobs.{{ .Release.Namespace }}"
  waitFor:
    - "-service {{ .component.prefix }}nappljobs.{{ .Release.Namespace }}.svc.cluster.local:{{ .this.nappl.port }} -timeout 600"

nappl:
  replicaCount: 2
  minReplicaCount: 1
  jobs: false
  waitFor:
    - "-service {{ .component.prefix }}nappljobs.{{ .Release.Namespace }}.svc.cluster.local:{{ .this.nappl.port }} -timeout 600"
nappljobs:
  replicaCount: 1
  jobs: true
  disableSessionReplication: true
  ingress:
    enabled: false
  snc:
    enabled: true
  waitFor:
    - "-service {{ .component.prefix }}database.{{ .Release.Namespace }}.svc.cluster.local:5432 -timeout 600"

application:
  nstl: 
    host: "{{ .component.prefix }}nstl-cluster.{{ .Release.Namespace }}"
  nappl:
    host: "{{ .component.prefix }}nappljobs.{{ .Release.Namespace }}"
  waitFor:
    - "-service {{ .component.prefix }}nappljobs.{{ .Release.Namespace }}.svc.cluster.local:{{ .this.nappl.port }} -timeout 1800"

nstla:
  minReplicaCountType: 1
  accounting: true
  logForwarder:
    - name: Accounting
      path: "/opt/ceyoniq/nscale-server/storage-layer/accounting/*.csv"
  serverID: 4711
  env:
    NSTL_REMOTESERVER_MAINTAINCONNECTION: 1
    NSTL_REMOTESERVER_SERVERID: 4712
    NSTL_REMOTESERVER_ADDRESS: "nstlb"
    NSTL_REMOTESERVER_NAME: "nstla"
    NSTL_REMOTESERVER_USERNAME: "admin"
    NSTL_REMOTESERVER_PASSWORD: "admin"
    NSTL_REMOTESERVER_MAXCONNECTIONS: 10
    NSTL_REMOTESERVER_MAXARCCONNECTIONS: 1
    NSTL_REMOTESERVER_FORWARDDELETEJOBS: 0
    NSTL_REMOTESERVER_ACCEPTRETRIEVAL: 1
    NSTL_REMOTESERVER_ACCEPTDOCS: 1
    NSTL_REMOTESERVER_ACCEPTDOCSWITHTHISSERVERID: 1
    NSTL_REMOTESERVER_PERMANENTMIGRATION: 1
nstlb:
  accounting: true
  logForwarder:
    - name: Accounting
      path: "/opt/ceyoniq/nscale-server/storage-layer/accounting/*.csv"
  serverID: 4712
  env:
    NSTL_REMOTESERVER_MAINTAINCONNECTION: 1
    NSTL_REMOTESERVER_SERVERID: 4711
    NSTL_REMOTESERVER_ADDRESS: "nstla"
    NSTL_REMOTESERVER_NAME: "nstla"
    NSTL_REMOTESERVER_USERNAME: "admin"
    NSTL_REMOTESERVER_PASSWORD: "admin"
    NSTL_REMOTESERVER_MAXCONNECTIONS: 10
    NSTL_REMOTESERVER_MAXARCCONNECTIONS: 1
    NSTL_REMOTESERVER_FORWARDDELETEJOBS: 0
    NSTL_REMOTESERVER_ACCEPTRETRIEVAL: 1
    NSTL_REMOTESERVER_ACCEPTDOCS: 1
    NSTL_REMOTESERVER_ACCEPTDOCSWITHTHISSERVERID: 1
    NSTL_REMOTESERVER_PERMANENTMIGRATION: 1
global:
  revisionHistoryLimit: 2