Chapter 3. CHASE Software Installation

The CHASE software installation is achieved by installing CHASE-BASE CD and then the other CD. During installation, some RPMs have to be installed before the others since there are some dependencies among them. These dependencies are described below.

The sysadm_base-lib RPM has to be installed first. It provides libraries which are used by all other RPMs. The the basic CHASE RPMs cluster_admin, cluster_services, and failsafe are to be installed. Following that, the GUI RPMs (sysadm_*) can be installed. Now, depending on which of the other two CDs are to be used, the packages and then the agents are installed. Example: For the CHASE-WEBSERVER CD, we install the apache-1.3.rpm for the particular Linux distribution and then the agent.

INSTALL Script

The installation described above is carried out by an INSTALL executable, that can be invoked by the user. The CHASE-BASE CD INSTALL script finishes with the installation from the CHASE-BASE CD and then waits for the user to input the second CD, where upon it calls the INSTALL from the other CD. INSTALL from the second CD can be run independently too; however the RPMs from the first CD need to be installed beforehand. Refer to the above paragraph on dependencies among the RPMs. After the second CD INSTALL script finishes installation, it tries to create a new cluster database and delete any existing ones.

Example 3-1. Sample Output from the INSTALL Script

CASE: Some RPMs are already installed.

#>./INSTALL

        ---- Begin CHASE Installation ----

  Checking for file /sbin/portmap ...
  Checking for package sysadm_base-lib ....
  Checking for package cluster_admin ....
  Checking for package cluster_services ....
  Checking for package failsafe ....
  Checking for package sysadm_base-tcpmux ....
  Checking for package sysadm_base-server ....
  Checking for package sysadm_base-client ....
  Checking for package sysadm_failsafe-client ....
  Checking for package sysadm_failsafe-server ....
  Checking for package IBMJava118-JRE ....

         Packages  sysadm_base-lib cluster_admin cluster_services 
failsafe sysadm_base-tcpmux sysadm_base-server sysadm_base-client 
sysadm_failsafe-client sysadm_failsafe-server IBMJava118-JRE already installed


        ---- Installing RPMs...

package sysadm_base-lib-1.3.6-1 is already installed
package cluster_admin-1.0.1-1 is already installed
package cluster_services-1.0.1-1 is already installed
package failsafe-1.0.1-1 is already installed
package sysadm_base-tcpmux-1.3.6-1 is already installed
package sysadm_base-server-1.3.6-1 is already installed
package sysadm_base-client-1.3.6-1 is already installed
package sysadm_failsafe-client-0.9-3 is already installed
package sysadm_failsafe-server-0.9-3 is already installed
package IBMJava118-JRE-1.1.8-3.0 is already installed

        ---- Checking packages installation...

  Checking for package sysadm_base-lib ....
  Checking for package cluster_admin ....
  Checking for package cluster_services ....
  Checking for package failsafe ....
  Checking for package sysadm_base-tcpmux ....
  Checking for package sysadm_base-server ....
  Checking for package sysadm_base-client ....
  Checking for package sysadm_failsafe-client ....
  Checking for package sysadm_failsafe-server ....
  Checking for package IBMJava118-JRE ....

        Do You want to install documentation on Failsafe (Y)/N ?


n

        ---- Updating /etc/services file...


        /etc/services update okay


        ---- Updating startup flags...


        Startup flags update okay


        ---- Creating links in init level directories...

  Creating link /etc/rc.d/rc3.d/S36fs_cluster
  Creating link /etc/rc.d/rc3.d/K39fs_cluster
  Creating link /etc/rc.d/rc3.d/S37failsafe
  Creating link /etc/rc.d/rc3.d/K37failsafe
  Creating link /etc/rc.d/rc5.d/S36fs_cluster
  Creating link /etc/rc.d/rc5.d/K39fs_cluster
  Creating link /etc/rc.d/rc5.d/S37failsafe
  Creating link /etc/rc.d/rc5.d/K37failsafe

        Links creation okay


        Insert 2nd CD now and enter when done ....

        ---- Begin CD 2 Installation ----

  Checking for package perl ....
  Checking for package sysadm_base-lib ....
  Checking for package cluster_admin ....
  Checking for package cluster_services ....
  Checking for package failsafe ....
  Checking for package apache ....

         Packages  apache already installed

WARNING: Some of the packages are of a different version.

        The Packages are - apache

         Overwrite them (Y)/N ?

n

WARNING: Proper working of software cannot be guaranteed with
                     older version of packages.
  Checking for package Apache-plugin ....

        ---- Installing RPMs...


        ---- Installing RPMs...

Apache-plugin               ############################################
Adding Apache resource type to CDB under cluster space

        ---- Checking installation...

  Checking for package Apache-plugin ....


        ---- Creating an empty database...

Assuming /var/lib/failsafe/cdb/cdb.db as the database
Preparing to delete database at /var/lib/failsafe/cdb/cdb.db! Continue [y]/n  

n

Response was not yes, exiting...

        ---- CHASE Installation DONE ----


A Brief Description of INSTALL Script


Note: The Linux FailSafe base CD requires about 25 MB.

To install the software, follow these steps:

  1. Make sure all servers in the cluster are running a supported release of Linux.

  2. Depending on the servers and storage in the configuration and the Linux revision level, install the latest install patches that are required for the platform and applications.

  3. On each system in the pool, install the version of the multiplexer driver that is appropriate to the operating system. Use the CD that accompanies the multiplexer. Reboot the system after installation.

  4. On each node that is part of the pool, install the following software, in order:

    1. sysadm_base-tcpmux

    2. sysadm_base-lib

    3. sysadm_base-server

    4. cluster_admin

    5. cluster_services

    6. failsafe

    7. sysadm_failsafe-server


    Note: You must install the sysadm_base-tcpmux, sysadm_base-server, and sysadm_failsafe packages on those nodes from which you want to run the FailSafe GUI. If you do not want to run the GUI on a specific node, you do not need to install these software packages on that node.


  5. If the pool nodes are to be administered by a Web-based version of the Linux FailSafe Cluster Manager GUI, install the following subsystems, in order:

    1. IBMJava118-JRE

    2. sysadm_base-client

    3. sysadm_failsafe-web

      If the workstation launches the GUI client from a Web browser that supports Java, install: java_plugin from the Linux FailSafe CD.

      If the Java plug-in is not installed when the Linux FailSafe Manager GUI is run from a browser, the browser is redirected to http://java.sun.com/products/plugin/1.1/plugin-install.html.

      After installing the Java plug-in, you must close all browser windows and restart the browser.

      For a non-Linux workstation, download the Java plug-in from http://java.sun.com/products/plugin/1.1/plugin-install.html

      If the Java plug-in is not installed when the Linux FailSafe Manager GUI is run from a browser, the browser is redirected to this site.

    4. sysadm_failsafe-client

  6. Install software on the administrative workstation (GUI client).

    If the workstation runs the GUI client from a Linux desktop, install these subsystems:

    1. IBMJava118-JRE

    2. sysadm_base-client

  7. On the appropriate servers, install other optional software, such as storage management or network board software.

  8. Install patches that are required for the platform and applications.

Brief Description of Second CD INSTALL Script

Chase supports two variants: CHASE-Web and CHASE-FS. In either case, the package/application and agents for each have to be installed.

For each distribution, currently only SuSE and Red Hat, and relevant release, find the package in the second CD. Install it using the RPM directive rpm -i.

For example, the CHASE-Web CD has the package apache.*.rpm in the directory /dev/cdrom/sgi/packages. Install it with the command rpm -i apache.*.rpm.

For the CHASE-FS CD, the packages are samba and nfs.

These packages may already exist on the server and may be of a different version. In that case, the older packages may have to be removed and the newer packages from the CHASE CD installed.

Installing Agents

For each CD, we have a list of agents in the directory /dev/cdrom/sgi/RPMS. Install them using the RPM command rpm -i. These agents are meant for the packages mentioned above.

Overview of Configuring Nodes for Linux FailSafe

Performing the system administration procedures required to prepare nodes for Linux FailSafe involves these steps:

  1. Install required software, as described above using the CHASE CD's INSTALL script.

  2. Configure the system files on each node, as described in “Configuring System Files”.

  3. Configure the network interfaces on the nodes using the procedure in “Configuring Network Interfaces”.

  4. When you are ready configure the nodes so that Linux FailSafe software starts up when they are rebooted.

To complete the configuration of nodes for Linux FailSafe, you must configure the components of the Linux FailSafe system, as described in Chapter 5, Linux FailSafe Cluster Configuration of the Linux FailSafe Administrator's Guide.

Configuring System Files

When you install the Linux FailSafe Software, there are some system file considerations you must take into account. This section describes the required and optional changes you make to the following files for every node in the pool:

  • /etc/services

  • /etc/failsafe/config/cad.options

  • /etc/failsafe/config/cdbd.options

  • /etc/failsafe/config/cmond.options

Configuring /etc/services for Linux FailSafe

The /etc/services file must contain entries for sgi-cmsd, sgi-crsd, sgi-gcd, and sgi-cad on each node before starting HA services in the node. The port numbers assigned for these processes must be the same in all nodes in the cluster. Note that sgi-cad requires a TCP port.

The following shows an example of /etc/services entries for sgi-cmsd, sgi-crsd, sgi-gcd and sgi-cad:

sgi-cmsd   7000/udp           # SGI Cluster Membership Daemon
sgi-crsd   17001/udp           # Cluster reset services daemon
sgi-gcd    17002/udp           # SGI Group Communication Daemon
sgi-cad    17003/tcp           # Cluster Admin daemon

Configuring /etc/failsafe/config/cad.options for Linux FailSafe

The /etc/failsafe/config/cad.options file contains the list of parameters that the cluster administration daemon (CAD) reads when the process is started. The CAD provides cluster information to the Linux FailSafe Cluster Manager GUI.

The following options can be set in the cad.options file:

--append_log
 

Append CAD logging information to the CAD log file instead of overwriting it.

--log_file filename
 

CAD log file name. Alternately, this can be specified as -lf filename.

-vvvv
 

Verbosity level. The number of “v”s indicates the level of logging. Setting -v logs the fewest messages. Setting -vvvv logs the highest number of messages.

The following example shows an /etc/failsafe/config/cad.options file:

-vv -lf /var/log/failsafe/cad_nodename --append_log

When you change the cad.options file, you must restart the CAD processes with the /etc/rc.d/init.d/fs_cluster restart command for those changes to take affect.

Configuring /etc/failsafe/config/cdbd.options for Linux FailSafe

The /etc/failsafe/config/cdbd.options file contains the list of parameters that the cdbd daemon reads when the process is started. The cdbd daemon is the configuration database daemon that manages the distribution of cluster configuration database (CDB) across the nodes in the pool.

The following options can be set in the cdbd.options file:

-logevents eventname
 

Log selected events. These event names may be used: all, internal, args, attach, chandle, node, tree, lock, datacon, trap, notify, access, storage.

The default value for this option is all.

-logdest log_destination
 

Set log destination. These log destinations may be used: all, stdout, stderr, syslog, logfile. If multiple destinations are specified, the log messages are written to all of them. If logfile is specified, it has no effect unless the -logfile option is also specified. If the log destination is stderr or stdout, logging is then disabled if cdbd runs as a daemon, because stdout and stderr are closed when cdbd is running as a daemon.

The default value for this option is logfile.

-logfile filename
 

Set log file name.

The default value is /var/log/failsafe/cdbd_log

-logfilemax maximum_size
 

Set log file maximum size (in bytes). If the file exceeds the maximum size, any preexisting filename.old will be deleted, the current file will be renamed to filename.old, and a new file will be created. A single message will not be split across files.

If -logfile is set, the default value for this option is 10000000.

-loglevel log level
 

Set log level. These log levels may be used: always, critical, error, warning, info, moreinfo, freq, morefreq, trace, busy.

The default value for this option is info.

-trace trace class
 

Trace selected events. These trace classes may be used: all, rpcs, updates, transactions, monitor. No tracing is done, even if it is requested for one or more classes of events, unless either or both of -tracefile or -tracelog is specified.

The default value for this option is transactions.

-tracefile filename
 

Set trace filename.

-tracefilemax maximum size
 

Set trace file maximum size (in bytes). If the file exceeds the maximum size, any preexisting filename.old will be deleted, the current file will be renamed to filename.old.

-[no]tracelog
 

[Do not] trace to log destination. When this option is set, tracing messages are directed to the log destination or destinations. If there is also a trace file, the tracing messages are written there as well.

-[no]parent_timer
 

[Do not] exit when parent exits.

The default value for this option is -noparent_timer.

-[no]daemonize
 

[Do not] run as a daemon.

The default value for this option is -daemonize.

-l
 

Do not run as a daemon.

-h
 

Print usage message.

-o help
 

Print usage message.

Note that if you use the default values for these options, the system will be configured so that all log messages of level info or less, and all trace messages for transaction events to file /var/log/failsafe/cdbd_log. When the file size reaches 10MB, this file will be moved to its namesake with the .old extension, and logging will roll over to a new file of the same name. A single message will not be split across files.

The following example shows an /etc/failsafe/config/cdbd.options file that directs all cdbd logging information to /var/log/messages, and all cdbd tracing information to /var/log/failsafe/cdbd_ops1. All log events are being logged, and the following trace events are being logged: RPCs, updates, and transactions. When the size of the tracefile /var/log/failsafe/cdbd_ops1 exceeds 100000000, this file is renamed to /var/log/failsafe/cdbd_ops1.old and a new file /var/log/failsafe/cdbd_ops1 is created. A single message is not split across files.

-logevents all -loglevel trace -logdest syslog -trace rpcs -trace 
updates -trace transactions -tracefile /var/log/failsafe/cdbd_ops1 
-tracefilemax 100000000

The following example shows an /etc/failsafe/config/cdbd.options file that directs all log and trace messages into one file, /var/log/failsafe/cdbd_chaos6, for which a maximum size of 100000000 is specified. -tracelog directs the tracing to the log file.

-logevents all -loglevel trace -trace rpcs -trace updates -trace 
transactions -tracelog -logfile /var/log/failsafe/cdbd_chaos6 
-logfilemax 100000000 -logdest logfile.

When you change the cdbd.options file, you must restart the cdbd processes with the /etc/rc.d/init.d/fs_cluster restart command for those changes to take affect.

Configuring /etc/failsafe/config/cmond.options for Linux FailSafe

The/etc/failsafe/config/cmond.options file contains the list of parameters that the cluster monitor daemon (cmond) reads when the process is started. It also specifies the name of the file that logs cmond events. The cluster monitor daemon provides a framework for starting, stopping, and monitoring process groups. See the cmond man page for information on the cluster monitor daemon.

The following options can be set in the cmond.options file:

-L loglevel
 

Set log level to loglevel.

-d
 

Run in debug mode.

-l
 

Lazy mode, where cmond does not validate its connection to the cluster database.

-t napinterval
 

The time interval in milliseconds after which cmond checks for liveliness of process groups it is monitoring.

-s [eventname]
 

Log messages to stderr.

A default cmond.options file is shipped with the following options. This default options file logs cmond events to the /var/log/failsafe/cmond_log file.

-L info -f /var/log/failsafe/cmond_log

Configuring Network Interfaces

The procedure in this section describes how to configure the network interfaces on the nodes in a Linux FailSafe cluster. The example shown in Figure 3-1 is used in the procedure.

Figure 3-1. Example Interface Configuration

Example Interface Configuration

  1. If possible, add every IP address, IP name, and IP alias for the nodes to /etc/hosts on one node.

    190.0.2.1 fs-ha1.company.com fs-ha1
    190.0.2.3 stocks
    190.0.3.1 priv-fs-ha1
    190.0.2.2 fs-ha2.company.com fs-ha2
    190.0.2.4 bonds
    190.0.3.2 priv-fs-ha2


    Note: IP aliases that are used exclusively by highly available services should not be added to system configuration files. These aliases will be added and removed by Linux FailSafe.


  2. Add all of the IP addresses from Step 1 to /etc/hosts on the other nodes in the cluster.

  3. If there are IP addresses, IP names, or IP aliases that you did not add to /etc/hosts in Steps 1 and 2, verify that NIS is configured on all nodes in the cluster.

    If the ypbind is off, you must start NIS. See your distribution's documentation for details.

  4. For IP addresses, IP names, and IP aliases that you did not add to /etc/hosts on the nodes in Steps 1 and 2, verify that they are in the NIS database by entering this command for each address:

    # ypmatch address mapname
    190.0.2.1 fs-ha1.company.com fs-ha1

    address is an IP address, IP name, or IP alias. mapname is hosts.byaddr if address is an IP address; otherwise, it is hosts. If ypmatch reports that address doesn't match, it must be added to the NIS database. See your distribution's documentation for details.

  5. On one node, statically configure that node's interface and IP address with the provided distribution tools.

    For the example in Figure 3-1, on a SuSE system, the public interface name and IP address lines are configured into /etc/rc.config in the following variables. Please note that YaST is the preferred method for modifying these variables. In any event, you should refer to the documentation of your distribution for help here:

    NETDEV_0=eth0
    IPADDR_0=$HOSTNAME

    $HOSTNAME is an alias for an IP address that appears in /etc/hosts.

    If there are additional public interfaces, their interface names and IP addresses appear on lines like these:

    NETDEV_1=
    IPADDR_1=

    In the example, the control network name and IP address are

    NETDEV_2=eth3
    IPADDR_3=priv-$HOSTNAME

    The control network IP address in this example, priv-$HOSTNAME, is an alias for an IP address that appears in /etc/hosts.

  6. Repeat Steps 5 and 6 on the other nodes.

  7. Verify that Linux FailSafe is off on each node:

    # /usr/lib/failsafe/bin/fsconfig failsafe
    # if [ $? -eq 1 ]; then echo off; else echo on; fi

    If failsafe is on on any node, enter this command on that node:

    # /usr/lib/failsafe/bin/fsconfig failsafe off

  8. Configure an e-mail alias on each node that sends the Linux FailSafe e-mail notifications of cluster transitions to a user outside the Linux FailSafe cluster and to a user on the other nodes in the cluster. For example, if there are two nodes called fs-ha1 and fs-ha2, in /etc/aliases on fs-ha1, add

    fsafe_admin:[email protected],[email protected] 

    On fs-ha2, add this line to /etc/aliases:

    fsafe_admin:[email protected],[email protected] 

    The alias you choose, fsafe_admin in this case, is the value you will use for the mail destination address when you configure your system. In this example, operations is the user outside the cluster and admin_user is a user on each node.

  9. If the nodes use NIS, ypbind is enabled to start at boot time, or the BIND domain name server (DNS), switching to local name resolution is recommended. Additionally, you should modify the /etc/nsswitch.conf file so that it reads as follows:

    hosts:                  files nis dns 


    Note: Exclusive use of NIS or DNS for IP address lookup for the cluster nodes has been shown to reduce availability in situations where the NIS service becomes unreliable.


  10. Reboot both nodes to put the new network configuration into effect.