Chapter 5. Troubleshooting

The 2 Gb TP9100 storage system includes a processor and associated monitoring and control logic that allows it to diagnose problems within the storage system's power, cooling, and drive systems.

SES (SCSI enclosure services) communications are used between the storage system and the RAID controllers. Status information on power, cooling, and thermal conditions is communicated to the controllers and is displayed in the management software interface.

The enclosure services processor is housed in the ESI/ops panel module. The sensors for power, cooling, and thermal conditions are housed within the power supply/cooling modules. Each module in the storage system is monitored independently.


Note: For instructions on opening the rear door of the rack, see “Opening and Closing the Rear Rack Door” in Chapter 1.

This chapter contains the following sections:

RAID Guidelines

RAID stands for “redundant array of independent disks”. In a RAID system multiple disk drives are grouped into arrays. Each array is configured as system drives consisting of one or more disk drives. A small, but important set of guidelines should be followed when connecting devices and configuring them to work with a controller.

Follow these guidelines when configuring a RAID system:

  • Distribute the disk drives equally among all the drive channels on the controller. This results in better performance. The TP9100 has two drive channels.

  • A drive pack can contain a maximum of 16 drives.

  • A drive pack can contain drives that are on any drive channel.

  • If configuring an online spare disk drive, ensure that the spare disk drive capacity is greater than or equal to the capacity of the largest disk drive in all redundant drive packs.

  • When replacing a failed disk drive, ensure that the replacement disk drive capacity is greater than or equal to the capacity of the failed disk drive in the affected drive pack.

Solving Initial Startup Problems

If cords are missing or damaged, plugs are incorrect, or cables are too short, contact your supplier for a replacement.

If the alarm sounds when you power on the storage system, one of the following conditions exists:

If the SGI server does not recognize the storage system, check the following:

  • Ensure that the device driver for the host bus adapter board has been installed. If the HBA was installed at the factory, this software is in place; if not, check the HBA and the server documentation for information on the device driver.

  • Ensure the FC-AL interface cables from the LRC I/O module to the Fibre Channel board in the host computer are installed correctly.

  • Check the selector switches on the ops panels of the storage system as follows:

    • On a tower or a RAID enclosure the ops panel should be set to address 1.

    • On the first expansion enclosure attached to a RAID system, the ops panel should be set to address 2.

    • On the first enclosure in a JBOD system, the ops panel should be set to address 1. Other enclosures that are daisy chained to the first enclosure should be addressed sequentially (2-7).

  • Ensure that the LEDs on all installed drive carrier modules are green. Note that the drive LEDs flash during drive spinup.

  • Check that all drive carrier modules are correctly installed.

  • If an amber disk drive module LED drive fault is on, there is a drive fault. See Table 5-7.

If the SGI server connected to the storage system is reporting multiple Hard Error (SCS_DATA_UNDERRUN) errors in the /var/adm/SYSLOG, the cabling connected to the controller reporting the errors requires cleaning or replacement. For more information on cleaning cables, refer to “Care and Cleaning of Optical Cables”.

Using Storage System LEDs for Troubleshooting

This section summarizes LED functions and gives instructions for solving storage system problems in these subsections:

ESI/Ops Panel LEDs and Switches

Figure 5-1 shows details of the ESI/ops panel .

Figure 5-1. ESI/Ops Panel Indicators and Switches

ESI/Ops Panel Indicators and Switches

Table 5-1 summarizes functions of the LEDs on the ESI/ops panel

Table 5-1. ESI/Ops Panel LEDs

LED

Description

Corrective Action

Power on

This LED illuminates green when power is applied to the enclosure.

N/A

Invalid address

This LED flashes amber when the enclosure is set to an invalid address mode.

Change the enclosure address thumb wheel to the proper setting. If the problem persists, contact your service provider.

System/ESI fault

This LED illuminates amber and the audible alarm sounds when the ESI processor detects an internal problem. This LED flashes when an over- or under-temperature condition exists.

Contact your service provider.

PSU/cooling/
temperature fault

This LED illuminates amber if an over- or under-temperature condition exists. This LED flashes if there is an ESI communications failure.

Check for proper airflow clearances and remove any obstructions. If the problem persists, lower the ambient temperature. In case of ESI communications failure, contact your service provider.

Hub mode

This LED illuminates green when the host side switch is enabled (RAID only).

N/A

2-Gb link speed

This LED illuminates green when 2-Gb link speed is detected.

N/A

The Ops panel switch settings for a JBOD enclosure are listed in Table 5-2. The ops panel switch settings for a RAID enclosure are listed in Table 5-3. The switches are read only during the power-on cycle.

Table 5-2. Ops Panel Configuration Switch Settings for JBOD

Switch
Number

Function

Function When Off

 

Function When On

1 (On[a])

Loop select single (1x16) or dual (2x8)

LRC operates as 2 loops of 8 drives (2x8). Refer to drive addressing mode 2.

 

LRC operates as 1 loop of 16 drives
(1x16 loop mode)

2 (On)

Loop terminate mode

If no signal is present on external FC port, the loop is left open

 

If no signal is present on external FC port, then the loop is closed

3 (Off)[b]

N/A

N/A

 

N/A

4 (Off)

N/A

N/A

 

N/A

5 and 6

RAID host hub speed select

Used only for RAID configurations.

Sw 5

Sw 6

Function

 

 

Off

Off

Force 1 Gb/s

 

 

On

Off

Force 2 Gb/s

 

 

Off

On

Reserved

 

 

On

On

Auto loop speed detect based on LRC port signals
Note: This feature is not supported

7 and 8

Drive loop speed select

Sw 7

Sw 8

Function

 

 

Off

Off

Force 1 Gb/s

 

 

On

Off

Force 2 Gb/s

 

 

Off

On

Speed selected by EEPROM bit

 

 

On

On

Auto loop speed detect based on LRC port signals
Note: This feature is not supported

9 and 10

Drive addressing mode

Sw 9

Sw 10

Function

 

 

On

On

Mode 0 - Single loop, base 16, offset of 4,
7 address ranges[c]

 

 

Off

On

Mode 1 - Single loop, base 20, 6 address ranges

 

 

On[d]

Off

Mode 2 - JBOD, dual loop, base 8,
15 address ranges

 

 

Off

Off

Mode 3 (Not used)

11 (On)

Soft select

Selects switch values stored in EEPROM

 

Selects switch values from hardware switches
Note: Soft select switch must be set to On

12 (Off)

Not used

Not used

 

Not used

[a] Bolded entries indicate default switch settings for a 1X16 JBOD. Set switches 1 and 10 to Off for 2X8 JBOD.

[b] Switches 3, 5, and 6 are used in RAID configurations

[c] Mode 0 (switches 9 and 10 set to On) is the SGI default factory setting

[d] Selecting mode 2 forces 2x8 dual loop selection


Table 5-3. Ops Panel Configuration Switch Settings for RAID

Switch Number

Function

Function when Off

 

Function when On

1 (On[a])

Loop select single (1x16) or dual (2x8)

LRC operates as two loops of 8 drives

 

LRC operates as 1 loop of 16 drives
(1x16 loop mode)

2 (On)

Loop terminate mode

If no signal is present on the external FC port, then the loop is left open

 

If no signal is present on the external FC port then the loop is closed internally

3 (Off)

Hub mode select
(RAID only)

Hub ports connect independently

 

RAID host FC ports are linked together internally

4 (Off)

Not used.

 

 

 

5 and 6

RAID host hub speed select

Note: Set switches 5 and 6 to Off to force 1 Gb/s if connecting RAID controllers to 1-Gb/s HBAs or switches

Sw 5

Sw 6

Function

 

 

Off

Off

Force 1 Gb/s

 

 

On

Off

Force 2 Gb/s

 

 

Off

On

Reserved

 

 

On

On

Auto loop speed detect based on LRC port signals
Note: This feature is not supported

7 and 8

Drive loop speed select

Sw 7

Sw 8

Function

 

 

Off

Off

Force 1 Gb/s

 

 

On

Off

Force 2 Gb/s

 

 

Off

On

Speed selected by EEPROM bit

 

 

On

On

Auto loop speed detect based on LRC port signals
Note: This feature is not supported

9 and 10

Drive addressing mode

Sw 9

Sw 10

Function

 

 

On

On

Mode 0 - Single loop, base 16, offset of 4,
7 address ranges[b]

 

 

Off

On

Mode 1 - Single loop, base 20, 6 address ranges

 

 

On

Off

Mode 2 - JBOD, dual loop, base 8,
15 address ranges[c]

 

 

Off

Off

Mode 3 (Not used)

11 (On)

Soft select

Selects switch values stored in EEPROM

 

Selects switch values from hardware switches
Note: Soft select switch must be set to On

12 (Off)

Not used

Not used

 

Not used

[a] Bolded entries indicate SGI's default switch settings for RAID

[b] Mode 0 (switches 9 and 10 set to On) is the SGI default factory setting

[c] Mode 2 (2x8) is not supported in a RAID configuration

Note the following:

Power Supply/Cooling Module LEDs

Figure 5-2 shows the meanings of the LEDs on the power supply/cooling module.

Figure 5-2. Power Supply/Cooling Module LED

Power Supply/Cooling Module LED

If the green “PSU good” LED is not lit during operation, or if the power/cooling LED on the ESI/ops panel is amber and the alarm is sounding, contact your service provider.

RAID LRC I/O Module LEDs

Figure 5-3 shows the LEDs on the dual-port RAID LRC I/O module.

Figure 5-3. Dual-port RAID LRC I/O Module LEDs

Dual-port RAID LRC I/O Module LEDs

Table 5-4 explains what the LEDs in Figure 5-3 indicate.

Table 5-4. Dual-port RAID LRC I/O Module LEDs

LED

Description

Corrective Action

ESI fault

This LED illuminates amber and the audible alarm sounds when the ESI processor detects an internal problem.

Check for mixed single-port and dual-port modules within an enclosure. Also check the drive carrier modules and PSU/cooling modules for faults. If the problem persists, contact your service provider.

RAID fault

This LED illuminates amber when a problem with the RAID controller is detected.

Contact your service provider.

RAID activity

This LED flashes green when the RAID controller is active.

N/A

Cache active

This LED flashes green when data is read into the cache.

N/A

Host port signal good (1 and 2)

This LED illuminates green when the port is connected to a host.

Check both ends of the cable and ensure that they are properly seated. If the problem persists, contact your service provider.

Drive loop signal good

This LED illuminates green when the expansion port is connected to an expansion enclosure.

Check both ends of the cable and ensure that they are properly seated. If the problem persists, contact your service provider.

Figure 5-4 shows the LEDs on the single-port RAID LRC I/O module.

Figure 5-4. Single-port RAID LRC I/O Module LEDs

Single-port RAID LRC I/O Module LEDs

Table 5-5 explains what the LEDs in Figure 5-4 indicate.

Table 5-5. Single-port RAID LRC I/O Module LEDs

LED

Description

Corrective Action

ESI fault

This LED illuminates amber and the audible alarm sounds when the ESI processor detects an internal problem.

Check the drive carrier modules and PSU/cooling modules. If the problem persists, contact your service provider.

RAID fault

This LED illuminates amber when a problem with the RAID controller is detected.

Contact your service provider.

RAID activity

This LED flashes green when the RAID controller is active.

N/A

Cache active

This LED flashes green when data is read into the cache.

N/A

Host port signal good

This LED illuminates green when the port is connected to a host.

Check both ends of the cable and ensure that they are properly seated. If the problem persists, contact your service provider.

Drive loop signal good

This LED illuminates green when the expansion port is connected to an expansion enclosure.

Check both ends of the cable and ensure that they are properly seated. If the problem persists, contact your service provider.


RAID Loopback LRC I/O Module LEDs

The LEDs on the rear of the RAID loopback LRC I/O module function similarly to those on the RAID LRC I/O modules. See “RAID LRC I/O Module LEDs” for more information.

JBOD LRC I/O Module LEDs

Figure 5-5 shows the JBOD LRC I/O module LEDs.

Figure 5-5. JBOD LRC I/O Module LEDs

JBOD LRC I/O Module LEDs

Table 5-6 explain what the LEDs in Figure 5-5 indicate.

Table 5-6. JBOD LRC I/O Module LEDs

LED

Description

Corrective Action

ESI fault

This LED illuminates amber and the audible alarm sounds when the ESI processor detects an internal problem.

Check the drive carrier modules and PSU/cooling modules. If the problem persists, contact your service provider.

FC-AL signal present

These LEDs illuminate green when the port is connected to an FC-AL.

Check the cable connections. If the problem persists, contact your service provider.


Drive Carrier Module LEDs

Each disk drive module has two LEDs, an upper (green) and a lower (amber), as shown in Figure 5-6.

Figure 5-6. Drive Carrier Module LEDs

Drive Carrier Module LEDs

Table 5-7 explains what the LEDs in Figure 5-6 indicate.

Table 5-7. Disk Drive LED Function

Green LED

Amber LED

State

Remedy

Off

Off

Disk drive not connected; the drive is not fully seated.

Check that the drive is fully seated

On

Off

Disk drive power is on, but the drive is not active.

N/A

Blinking

Off

Disk drive is active.
(LED might be off during power-on.)

N/A

Flashing at 2-second intervals

On

Disk drive fault (SES function).

Contact your service provider for a replacement drive and follow instructions in Chapter 6, “Installing and Replacing Drive Carrier Modules”

.

N/A

Flashing at half-second intervals

Disk drive identify (SES function).

N/A

In addition, the amber drive LED on the ESI/ops panel alternates between on and off every 10 seconds when a drive fault is present. 

Using the Alarm for Troubleshooting

The ESI/ops panel includes an audible alarm that indicates when a fault state is present. The following conditions activate the audible alarm:

  • RAID controller fault

  • Fan slows down

  • Voltage out of range

  • Over-temperature

  • Storage system fault

You can mute the audible alarm by pressing the alarm mute button for about a second, until you hear a double beep. The mute button is beneath the indicators on the ESI/ops panel (see Figure 5-1).

When the alarm is muted, it continues to sound with short intermittent beeps to indicate that a problem still exists. It is silenced when all problems are cleared.


Note: If a new fault condition is detected, the alarm mute is disabled.


Solving Storage System Temperature Issues

This section explains storage system temperature conditions and problems in these subsections:

Thermal Control

The storage system uses extensive thermal monitoring and ensures that component temperatures are kept low and acoustic noise is minimized. Airflow is from front to rear of the storage system. Dummy modules for unoccupied bays in enclosures and blanking panels for unoccupied bays in the rack must be in place for proper operation.

If the ambient air is cool (below 25 °C or 77 °F) and you can hear that the fans have sped up by their noise level and tone, then some restriction on airflow might be raising the storage system's internal temperature. The first stage in the thermal control process is for the fans to automatically increase in speed when a thermal threshold is reached. This might be a normal reaction to higher ambient temperatures in the local environment. The thermal threshold changes according to the number of drives and power supplies fitted.

If fans are speeding up, follow these steps:

  1. Check that there is clear, uninterrupted airflow at the front and rear of the storage system.

  2. Check for restrictions due to dust buildup; clean as appropriate.

  3. Check for excessive recirculation of heated air from the rear of the storage system to the front.

  4. Check that all blank plates and dummy disk drives are in place.

  5. Reduce the ambient temperature.

Thermal Alarm

The four types of thermal alarms and the associated corrective actions are described in Table 5-8.

Table 5-8. Thermal Alarms

Alarm Type

Indicators

Solutions

High temp warning
begins at 54°C (129°F)

Audible alarm sounds

Ops panel system fault LED flashes

Fans run at higher speed than normal

SES temperature status is non-critical

PSU Fault led is lit

If possible, power down the enclosure; then check the following:

  • Ensure that the local ambient temperature meets the specifications outlined in

“Environmental Requirements” in Appendix A

  • Ensure that the proper clearances are provided at the front and rear of the rack.

  • Ensure that the airflow through the rack is not obstructed.

    If you are unable to determine the cause of the alarm, please contact your service provider.

High temp failure
begins at 58°C (136°F)

Audible alarm sounds

Ops panel system fault LED flashes

Fans run at higher speed than normal

SES temperature status is critical

PSU Fault led is lit

 

Low temp warning
begins at 10°C (50°F)

Audible alarm sounds

Ops panel system fault is lit

SES temperature status is non-critical

 

Low temp failure
begins at 0°C (32°F)

Audible alarm sounds

Ops panel system fault is lit

SES temperature status is critical

 


Using Test Mode

When no faults are present in the storage system, you can run test mode to check the LEDs and the audible alarm on the ESI/ops panel. In this mode, the amber and green LEDs on each of the drive carrier modules and the ESI/ops panel flash on and off in sequence; the alarm beeps twice when test mode is entered and exited.

To activate test mode, press the alarm mute button until you hear a double beep. The LEDs then flash until the storage system is reset, either when you press the alarm mute button again or if an actual fault occurs.

Care and Cleaning of Optical Cables


Warning: Never look into the end of a fiber optic cable to confirm that light is being emitted (or for any other reason). Most fiber optic laser wavelengths (1300 nm and 1550 nm) are invisible to the eye and cause permanent eye damage. Shorter wavelength lasers (for example, 780 nm) are visible and can cause significant eye damage. Use only an optical power meter to verify light output.



Warning: Never look into the end of a fiber optic cable on a powered device with any type of magnifying device, such as a microscope, eye loupe, or magnifying glass. Such activity can cause a permanent burn on the retina of the eye. Optical signal cannot be determined by looking into the fiber end.

Fiber optic cable connectors must be kept clean to ensure long life and to minimize transmission loss at the connection points. When the cables are not in use, replace the caps to prevent deposits and films from adhering to the fiber. A single dust particle caught between two connectors will cause significant signal loss. In addition to causing signal loss, dust particles can scratch the polished fiber end, resulting in permanent damage. Do not touch the connector end or the ferrules; your fingers will leave an oily deposit on the fiber. Do not allow uncapped connectors to rest on the floor.

If a fiber connector becomes visibly dirty or exhibits high signal loss, carefully clean the entire ferrule and end face with special lint-free pads and isopropyl alcohol. The end face in a bulkhead adapter on test equipment can also be cleaned with special lint-free swabs and isopropyl alcohol. In extreme cases, a test unit may need to be returned to the factory for a more thorough cleaning.

Never use cotton, paper, or solvents to clean fiber optic connectors; these materials may leave behind particles or residues. Instead, use a fiber optic cleaning kit especially made for cleaning optical connectors, and follow the directions. Some kits come with canned air to blow any dust out of the bulkhead adapters. Be cautious, as canned air can damage the fiber if not used properly. Always follow the directions that come with the cleaning kit.